System
The system addresses the challenge of discovering new content by extracting and visually mapping relationships, ensuring users find relevant and up-to-date information across various content types.
Patent Information
- Application Number
- JP2024124052
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-07-30
- Publication Date
- 2026-02-12
AI Technical Summary
Users face challenges in efficiently discovering new content due to the vast amount of information available, and existing systems struggle to visually display relationships between different types of content, limiting the discovery of relevant materials in educational and research institutions.
A system that extracts meta-information from collected text data, maps relationships between content items, and visually displays these relationships, while continuously updating content data from domestic and international databases to provide the latest information.
Enables users to efficiently discover new content by intuitively understanding relationships between different types of content, ensuring freshness and diversity through continuous updates.
Smart Images

Figure 2026022535000001_ABST
Abstract
Description
[Technical Field]
[0001] The technology of the present disclosure relates to a system. [Background technology]
[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]
[0004] In modern society, the variety of content is rapidly increasing, making it difficult for users to properly search for information and discover new content. Therefore, there is a need for a system that can efficiently find relevant content based on users' interests. Furthermore, while efficient resource selection is required in educational and research institutions, finding the necessary materials from vast amounts of information is a major challenge. The purpose of this invention is to solve these problems and help users efficiently discover new content. [Means for solving the problem]
[0005] The present invention provides a system that includes a means for extracting meta-information from collected text data, a means for mapping relationships between content items based on the extracted meta-information, and a means for visually displaying the mapped relationships. Furthermore, by including a means for acquiring and updating content data in conjunction with domestic and international databases, and a means for finding commonalities between content items and displaying them as links, the system continuously provides the latest information and enables users to efficiently discover new content items that interest them.
[0006] "Collection" is the act of gathering data for a specific purpose.
[0007] "Text data" refers to data that includes character information such as books, lyrics, data after converting audio data into text, and text extracted from images using OCR.
[0008] "Meta information" is auxiliary information that describes the characteristics and content of content, such as keywords, tags, content names, important people, places, events, and themes.
[0009] "Extraction" is the act of extracting specific information from a large amount of data.
[0010] A "relationship" is a commonality or connection that exists between different pieces of content.
[0011] "Mapping" is the act of linking elements of data with elements of other data.
[0012] "Visually displayed" means that information is displayed using visual elements such as graphs, charts, interfaces, etc.
[0013] "Domestic and international databases" are systems for storing various types of information that exist both domestically and internationally.
[0014] "Acquisition" is the act of obtaining specific information.
[0015] "Update" is the act of updating existing information.
[0016] A "link" is a means of connection that connects different data or information.
[0017] A "system" is a collection of interrelated elements or devices that accomplish a specific purpose.
[0018] "User interface" refers to the means, screens, and methods of operation that a user uses to interact with a system. [Brief explanation of the drawings]
[0019] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11]FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION
[0020] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.
[0021] First, the terms used in the following description will be explained.
[0022] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).
[0023] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.
[0024] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.
[0025] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.
[0026] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."
[0027] [First embodiment]
[0028] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0029] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0030] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0031] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.
[0032] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0033] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0034] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.
[0035] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0036] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0037] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0038] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0039] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0040] This invention is a system that uses AI and natural language processing technology to extract and visualize the relationships between a wide variety of content, helping users discover new content that they are interested in. The basic process for implementing this system is shown below.
[0041] Content Collection
[0042] The server collects a variety of content, such as books, lyrics, audio data, image data, etc. At this stage, data may be retrieved from online databases via APIs, or files may be uploaded directly to the server.
[0043] Example: A server calls the Google Books API to retrieve book data related to a specific keyword. If a user wants to collect data related to "mystery novels," the server uses the Google Books API to retrieve book information related to "mystery novels." This information includes the book's title, author, publication year, summary of the content, etc.
[0044] Text analytics
[0045] The server sends the collected text data to a natural language processing engine to extract meta-information (keywords, tags, content names, etc.). This process uses technologies such as morphological analysis, entity recognition, and relationship extraction.
[0046] Example: The server analyzes data from the novel "Sherlock Holmes" and extracts character names such as "Sherlock Holmes" and "John Watson" and keywords such as "London" and "detective."
[0047] Keyword tag generation
[0048] Based on the extracted meta information, the server generates keywords and tags related to each piece of content, including important people, places, events, and themes.
[0049] Example: The server assigns tags such as "detective," "mystery," and "deduction" to "Sherlock Holmes" novels.
[0050] Relationship Mapping
[0051] The server maps the relationships between pieces of content based on keywords and tags, finding commonalities between pieces of content and creating a data structure that visually displays them as links.
[0052] Example: The server detects that the novels "Sherlock Holmes" and "Hercule Poirot" share the common keyword "detective" and links them together.
[0053] Database integration and updates
[0054] The server works in conjunction with other databases to regularly update and supplement the content information, ensuring that the most up-to-date and detailed information is always available.
[0055] Example: The server works with the Open Library API to periodically retrieve and update data on newly released "mystery novels."
[0056] User interface provided
[0057] The server provides a user interface that includes a relationship map and keyword search functionality, allowing users to enter keywords of interest and visually display related content.
[0058] Example: When a user searches for "detective" on their device, the server displays a visual map of related content, such as "Sherlock Holmes" and "Hercule Poirot."
[0059] User Interactions
[0060] Users can click on the displayed relationship map to view more detailed information, and can also enter new search keywords to find related content.
[0061] Example: When a user clicks on "Sherlock Holmes," the server displays detailed information and reviews about the work, as well as information about related TV shows and movies.
[0062] As described above, the system of the present invention provides efficient information provision and allows users to encounter new content through each stage of collection, analysis, mapping, display, and updating.
[0063] The processing flow will be explained below.
[0064] Step 1: Gather content
[0065] The server collects a variety of content such as books, lyrics, audio data, and image data.
[0066] Specifically, it retrieves data from online databases via API and also receives files uploaded by users.
[0067] Example: The server calls the Google Books API to retrieve metadata for books related to "fantasy novels."
[0068] Step 2: Registering text data storage
[0069] The server stores the collected text data in storage.
[0070] The data to be saved includes books, lyrics, text converted from audio data, and text extracted from images using OCR.
[0071] Example: A server stores the book data for "Harry Potter and the Philosopher's Stone" in a database.
[0072] Step 3: Text analysis
[0073] The server sends the stored text data to a natural language processing engine to extract meta-information.
[0074] It uses morphological analysis, entity recognition, and relationship extraction to identify characters, places, important events, etc.
[0075] Example: The server uses a natural language processing engine to extract character names such as "Sherlock Holmes" and "John Watson" from the text "Sherlock Holmes."
[0076] Step 4: Generate keywords and tags
[0077] The server generates keywords and tags related to each piece of content based on the extracted meta information.
[0078] This includes important people, places, events, themes, etc.
[0079] Example: The server assigns tags such as "detective," "mystery," and "deduction" to "Sherlock Holmes."
[0080] Step 5: Relationship mapping
[0081] The server maps the relationships between content based on the generated keywords and tags.
[0082] Create a data structure that finds commonalities between pieces of content and visually displays them as links.
[0083] Example: The server discovers that "Sherlock Holmes" and "Hercule Poirot" share the tag "detective" and links them together.
[0084] Step 6: Database integration and updates
[0085] The server cooperates with other databases to periodically retrieve and update content information.
[0086] This ensures that the latest information is always provided.
[0087] Example: The server works with the Open Library API to periodically retrieve and update data on newly released mystery novels.
[0088] Step 7: Provide the user interface
[0089] The server provides a user interface that includes a relationship map and keyword search functionality.
[0090] Users can enter keywords that interest them and relevant content will be displayed visually.
[0091] For example, if a user searches for "detective" on their device, the server will display related content such as "Sherlock Holmes" or "Hercule Poirot."
[0092] Step 8: User Interaction
[0093] Users can click on the displayed relationship map to view more detailed information.
[0094] Additionally, users can enter new search keywords to find related content.
[0095] For example, if a user clicks on "Sherlock Holmes," the server will display detailed information and reviews about the work, as well as information about related TV shows and movies.
[0096] Example 1
[0097] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0098] Conventional content management systems have difficulty efficiently extracting and visualizing the relationships between different types of content. Furthermore, the means by which users can discover new related content are limited, which does not sufficiently improve the user experience. Furthermore, there is insufficient integration with other databases to continuously acquire and update the latest information, making it difficult to ensure the freshness and diversity of content.
[0099] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0100] In this invention, the server includes a means for collecting information, a means for preprocessing the collected text data and extracting meta information, a means for generating keywords and tags based on the extracted meta information, a means for mapping relationships between content using the generated keywords and tags, and a means for visually displaying the mapped relationships. This enables efficient extraction and visualization of relationships between content of different formats. Furthermore, by including a means for acquiring and updating content data from domestic and international databases in cooperation with the server, it becomes possible to continuously acquire and provide the latest and most diverse information to users. Furthermore, by including a means for finding commonalities between mapped content and displaying them as links, users can intuitively discover new related content.
[0101] "Means for collecting information" refers to the means for collecting various content such as books, lyrics, audio data, and image data on a server through methods such as online databases and file uploads.
[0102] "Preprocessing" refers to the process of removing unnecessary characters and duplication from the collected content data, formatting the text data, and preparing it for passing to a natural language processing engine.
[0103] "Meta information" is summary information of data such as keywords, tags, and content names extracted from collected text data, and is important information for understanding the meaning and substance of the content.
[0104] The "means for generating keywords and tags" refers to a means for automatically generating keywords and tags related to each piece of content based on meta information obtained from the natural language processing engine.
[0105] A "means for mapping relationships between content" is a means for finding commonalities and relationships between multiple pieces of content based on generated keywords and tags, and visually displaying them as links.
[0106] A "visual display means" is a means for displaying the relationships between the mapped content in a graphical format that allows users to intuitively understand the relationships.
[0107] "Means for obtaining and updating content data from domestic and international databases in cooperation with the server" refers to means for the server to periodically communicate with external databases to obtain new content data and update existing data.
[0108] "Means for finding commonalities and displaying them as links" refers to means for finding common keywords and tags between multiple pieces of content and visually displaying the relationships between them as links.
[0109] This invention is a system that efficiently extracts and visualizes the relationships between different types of content, allowing users to discover new related content. This system is run by a server, a terminal, and user operations.
[0110] Hardware and software used
[0111] This system is composed of the following hardware and software components:
[0112] 1. Hardware
[0113] Server: Collects, processes, and stores data.
[0114] Terminal: A device (e.g., PC, smartphone, tablet) that a user uses to access and operate the system.
[0115] 2. Software
[0116] API: Application Program Interface for connecting with external services for data collection.
[0117] Natural language processing engine: An engine for analyzing text data and extracting meta-information (e.g., morphological analysis engine, entity recognition engine).
[0118] Database: A database management system for storing and managing the collected and generated data.
[0119] Visualization libraries: Libraries for visually displaying relationships between content (e.g., D3.js).
[0120] Specific processing contents of the program
[0121] First, the server collects various content such as books, lyrics, audio data, image data, etc. through APIs and file upload functions. For example, if a user wants to collect data related to "mystery novels," the server calls the API of an online database and retrieves book information as a result.
[0122] The server then preprocesses the collected text data and sends it to a natural language processing engine to extract meta-information, using a morphological analysis engine and entity recognition engine to extract important keywords and tags from the text.
[0123] The server then automatically generates related keywords and tags based on the extracted meta information and maps the relationships between content items. For example, if the novels of "Sherlock Holmes" and "Hercule Poirot" share the keyword "detective," these pieces of content are linked.
[0124] The server also periodically connects to external databases (e.g., Open Library API) to retrieve new content data and update the database, ensuring that users always have access to the latest information.
[0125] When a user accesses the system using a terminal and searches for a specific keyword, the server visually displays the data mapping the relationships. For example, if a user searches for "detective," related content will be displayed as a visual map.
[0126] Additionally, when a user clicks on a particular piece of content on the relationship map, the server provides more information about it. For example, if a user clicks on "Sherlock Holmes," the server displays more information about the work, related reviews, and other related content.
[0127] Prompt Sentence Examples
[0128] "Generate a program to collect, analyze, and display data to map relationships in a detective novel."
[0129] "Please explain in detail the process of the system that extracts relationships between content based on meta information and displays them visually."
[0130] The system is designed to help users discover new and relevant content and improve the overall user experience.
[0131] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0132] Step 1:
[0133] Content Collection
[0134] The server collects various content, such as books, lyrics, audio data, and image data, through API calls or file uploads. It receives keywords and files specified by the user as input. The server saves the collected data in its internal storage. It also uses APIs to retrieve data related to specific keywords from external databases. For example, if a user enters "mystery novels," the server calls the Google Books API and retrieves book information as a result. The output is the collected raw data.
[0135] Step 2:
[0136] Text Extraction and Preprocessing
[0137] The server extracts text from the collected content data and performs preprocessing. It receives the collected raw data as input. Specifically, it performs formatting processing by removing unnecessary characters and duplication from the text file. It then prepares the data for passing to the natural language processing engine. The output is preprocessed text data.
[0138] Step 3:
[0139] Extracting Meta Information
[0140] The server sends the preprocessed text data to a natural language processing engine to extract meta-information. The server receives the preprocessed text data as input. Specifically, it analyzes the text using a morphological analysis engine (e.g., Mecab) to extract meta-information such as keywords, tags, and content names. The output is the extracted meta-information.
[0141] Step 4:
[0142] Keyword tag generation
[0143] The server generates keywords and tags appropriate for each piece of content based on the extracted meta information. It receives meta information as input. Specifically, it extracts important people, places, events, themes, etc., and generates keywords and tags based on them. For example, it generates tags such as "detective," "mystery," and "deduction" from the text "Sherlock Holmes." The output is the generated keywords and tags.
[0144] Step 5:
[0145] Relationship mapping
[0146] The server maps the relationships between different content based on the generated keywords and tags. It receives the generated keywords and tags as input. Specifically, it links content that has common keywords or tags together to create a relational data structure. For example, if the texts "Sherlock Holmes" and "Hercule Poirot" share the tag "detective," they are linked together. The output is the mapped relationship data.
[0147] Step 6:
[0148] Visually mapping your data
[0149] The server performs visualization processing to visually display the relationship data. It receives the mapped relationship data as input. Specifically, it uses a visualization library such as D3.js to generate a relationship map in HTML format and sends it to the user's device. The output is a visual relationship map.
[0150] Step 7:
[0151] Database integration and updates
[0152] The server interacts with an external database to retrieve new content data and update existing data. It receives new data retrieved from the external database's API as input. Specifically, it periodically sends API requests to add new book and content data to the database and update existing data. The output is an updated database.
[0153] Step 8:
[0154] Providing a user interface
[0155] The server provides an interface that allows users to manipulate the relationship map and search results. It receives search keywords and click operations from the user as input. Specific operations include searching for related content based on the user's input and displaying it on the device as a visual map. The output is the visual map and search results displayed on the user interface.
[0156] Step 9:
[0157] Viewing detailed information
[0158] When a user clicks on a specific piece of content on the relationship map, the server displays its detailed information. It receives the user's click as input and retrieves detailed information about the related content from the database. Specifically, it provides the user with detailed information about the content, reviews, and other related content. The output is an interface displaying the detailed information.
[0159] (Application example 1)
[0160] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0161] In recent years, a wide variety of content has become available on the Internet, but it is not easy for users to discover new content that interests them. In particular, there is a need to efficiently search for content that matches users' interests from a vast amount of content data and show its relevance. Conventional methods have difficulty visually displaying the relationships between content, and there is a lack of systems that guide users to new content that interests them. To solve this problem, it is necessary to concretely visualize the abstract relationships between content and provide an interface that users can intuitively understand.
[0162] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0163] In this invention, the server includes means for extracting meta-information from collected text data, means for mapping relationships between content items based on the extracted meta-information, means for visually displaying the mapped relationships, means for vectorizing the collected text data and calculating the similarity between the vectors, and means for generating a network graph based on the similarity. This allows users to easily find new content items related to their interests from among a vast amount of content items and intuitively understand their relationships.
[0164] "Collected text data" refers to text information in a variety of formats, such as books, audio data, and image data, obtained from the Internet or other databases.
[0165] "Meta information" refers to accompanying information such as keywords, tags, and content names extracted from collected text data.
[0166] "Relationships between contents" is information that indicates commonalities and interrelationships between different contents.
[0167] "Mapping means" refers to a technical method for visually showing the relationships between content items based on extracted meta-information.
[0168] "Visual display means" refers to a method of presenting the relationships between content to users in a visual format such as a graph or map.
[0169] "Vectorization" means converting collected text data into numerical vectors, making it mathematically processable.
[0170] The "means for calculating similarity" is a method for calculating the similarity between vectorized text data and evaluating the numerical relationship.
[0171] A "network graph" is a structure that visually represents the relationships between content using nodes and edges.
[0172] This invention is a system that extracts relationships between various contents and supports users in discovering new contents. The system consists of a server, a terminal, and a user interface.
[0173] Content Collection
[0174] The server collects various content, such as books, audio data, and image data, from the Internet and other databases. Data is periodically retrieved from online databases via APIs and stored on the server as text data. For example, data can be collected using an information search API or a book information API.
[0175] Text Analysis
[0176] The server sends the collected text data to a natural language processing engine to extract meta-information (keywords, tags, content names, etc.). This process utilizes technologies such as morphological analysis, entity recognition, and relationship extraction. Specifically, it uses natural language processing toolkits and entity recognition software.
[0177] Meta information generation
[0178] The server automatically generates keywords and tags related to each piece of content based on the extracted meta information, including important people, places, events, themes, etc. For example, tags are generated based on the characters and themes of a book.
[0179] Relationship mapping
[0180] The server maps the relationships between content based on the generated keywords and tags. At this stage, the similarity between vectorized data is calculated and the numerical relationships are evaluated. Specifically, tools such as TfidfVectorizer and Cosine Similarity are used for vectorization and similarity calculation.
[0181] Visualizing Relationships
[0182] The server generates a network graph to visually display the results of the relationship mapping. The generated graph is displayed in a visual format that allows users to intuitively understand it. Data visualization tools such as NetworkX and matplotlib are used to generate the network graph.
[0183] User interface provided
[0184] The server provides a user interface with a relationship map and keyword search functionality. Users can enter keywords of interest and visually display related content. The user interface is web-based and uses web frameworks such as Flask.
[0185] User Interactions
[0186] Users can click on the visually displayed relationship map to view more detailed information, and can also enter new search keywords to discover and enjoy new related content.
[0187] Examples of concrete examples and prompts
[0188] For example, if a user searches for "suspense movies," you can visualize related movies, TV shows, and even related podcasts and books. Visually display movies, TV shows, podcasts, and books related to "suspense movies."
[0189] This system enables users to efficiently discover new content that interests them and intuitively understand its relationships.
[0190] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0191] Step 1:
[0192] The server collects content data from the Internet and other databases. Specifically, it uses APIs (such as information search APIs and book information APIs) to obtain a variety of content data, including books, audio data, and image data, and stores it on the server. The input required is a search query containing keywords that interest the user, and the output is a list of collected text data.
[0193] Step 2:
[0194] The server sends the collected text data to a natural language processing engine to extract meta-information. This process uses techniques such as morphological analysis, entity recognition, and relationship extraction. Specifically, a natural language processing toolkit is used to analyze the text and extract keywords, tags, content names, etc. The input is the text data obtained in step 1, and the output is the extracted meta-information.
[0195] Step 3:
[0196] The server generates keywords and tags related to each piece of content based on the extracted meta information. Tags such as important people, places, events, and themes are automatically added. The input is the meta information obtained in step 2, and the output is a list of generated keywords and tags.
[0197] Step 4:
[0198] The server maps the relationships between content items based on the generated keywords and tags. It evaluates the similarity between content items by vectorizing the collected text data (using TfidfVectorizer) and calculating the similarity between each vector (using Cosine Similarity). The input is the keywords and tags obtained in Step 3 and the text data, and the output is a similarity score.
[0199] Step 5:
[0200] The server generates a network graph based on the similarity. It uses data visualization tools such as NetworkX and matplotlib to visualize the relationships between content. The input is the similarity scores obtained in step 4, and the output is the visualized network graph.
[0201] Step 6:
[0202] The server provides a user-facing interface. Users can enter keywords of interest and relevant content is displayed visually. A web framework such as Flask is used to build the web-based interface. The input is the keywords entered by the user, and the output is a visually displayed relationship map.
[0203] Step 7:
[0204] Users can click on the visually displayed relationship map to view more information, and can also enter new search keywords to discover new related content. The input is the user's actions on the relationship map, and the output is the discovery of more information and new content.
[0205] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.
[0206] This invention is a system that uses AI and natural language processing technology to extract and visualize the relationships between a wide variety of content, and by combining this with an emotion engine that recognizes the user's emotions, helps users discover new content that they are interested in. The basic process for implementing this system is shown below.
[0207] Content Collection
[0208] The server collects various content such as books, lyrics, audio data, image data, etc. At this stage, it retrieves data from online databases via APIs and may also receive files uploaded by users.
[0209] Example: A server uses the Google Books API to retrieve book data related to a specific keyword. If a user wants to collect data related to "mystery novels," the server retrieves book information related to "mystery novels" via the Google Books API. This information includes the book title, author, publication year, and summary of the contents.
[0210] Text analytics
[0211] The server uses a natural language processing engine to extract meta-information (keywords, tags, content names, etc.) from the collected text data. This process involves morphological analysis, entity recognition, and relationship extraction.
[0212] Example: The server analyzes text data from "Sherlock Holmes" and extracts character names such as "Sherlock Holmes" and "John Watson" and keywords such as "London" and "detective."
[0213] Keyword tag generation
[0214] The server generates keywords and tags related to each piece of content based on the extracted meta information.
[0215] Example: The server assigns tags such as "detective," "mystery," and "deduction" to "Sherlock Holmes" novels.
[0216] Relationship Mapping
[0217] The server maps the relationships between content based on the generated keywords and tags, finding commonalities between the content and visually displaying them as links.
[0218] Example: The server discovers that "Sherlock Holmes" and "Hercule Poirot" share the tag "detective" and links them together.
[0219] Database integration and updates
[0220] The server periodically retrieves and updates content information in cooperation with other databases, ensuring that the latest information is always provided.
[0221] Example: The server works with the Open Library API to periodically retrieve and update data on newly released "mystery novels."
[0222] Incorporating an emotion engine
[0223] The server analyzes the user's feelings about the content using an emotion engine that recognizes the user's emotions.
[0224] Example: When a user searches for "mystery novels" on their device, the emotion engine analyzes the user's facial expressions and voice to determine their emotions. For example, if the user shows expressions of surprise or interest, relevant content will be recommended based on that.
[0225] Emotion-based content recommendation
[0226] The server dynamically updates and recommends content that reflects the user's emotional state based on the analysis results from the emotion engine.
[0227] Example: If a user expresses the emotion of "surprise" while browsing "Sherlock Holmes," the server will recommend "mystery novels" or "suspense movies" that are likely to induce similar feelings of "surprise."
[0228] User interface provided
[0229] The server provides a user interface that includes a relationship map and keyword search functionality, allowing users to enter keywords of interest and visually display related content.
[0230] Example: When a user searches for "detective" on their device, the server displays a visual map of related content, such as "Sherlock Holmes" and "Hercule Poirot."
[0231] Users can click on the displayed relationship map to view detailed information about their interactions, or enter new search keywords to find related content.
[0232] Example: When a user clicks on "Sherlock Holmes," the server displays detailed information and reviews about the work, as well as information about related TV shows and movies.
[0233] As described above, the system of the present invention provides efficient information provision and encounters with new content through each stage of collection, analysis, mapping, display, update, and emotion recognition.
[0234] The processing flow will be explained below.
[0235] Step 1: Gather content
[0236] The server collects a variety of content such as books, lyrics, audio data, and image data.
[0237] Specifically, it retrieves data from an online database via API and stores files uploaded by users in storage.
[0238] Example: A server uses the Google Books API to retrieve metadata for books related to "fantasy novels."
[0239] Step 2: Registering text data storage
[0240] The server stores the collected text data in a database.
[0241] The data stored includes books, lyrics, text converted from audio data, and text extracted from images.
[0242] Example: A server stores the book data for "Harry Potter and the Philosopher's Stone" in a database.
[0243] Step 3: Text analysis
[0244] The server sends the stored text data to a natural language processing engine to extract meta-information.
[0245] It uses morphological analysis, entity recognition, and relationship extraction to identify characters, places, important events, etc.
[0246] Example: The server analyzes the text data of "Sherlock Holmes" and extracts character names such as "Sherlock Holmes" and "John Watson" as well as keywords such as "London" and "detective."
[0247] Step 4: Generate keywords and tags
[0248] The server generates keywords and tags related to each piece of content based on the extracted meta information.
[0249] This includes important people, places, events, themes, etc.
[0250] Example: The server tags "detective," "mystery," and "deduction" to "Sherlock Holmes" novels.
[0251] Step 5: Relationship mapping
[0252] The server maps the relationships between content based on the generated keywords and tags.
[0253] Create a data structure that finds commonalities between pieces of content and visually displays them as links.
[0254] Example: The server discovers that "Sherlock Holmes" and "Hercule Poirot" share the tag "detective" and links them together.
[0255] Step 6: Database integration and updates
[0256] The server cooperates with other databases to periodically retrieve and update content information.
[0257] This ensures that the latest information is always provided.
[0258] Example: The server works with the Open Library API to periodically retrieve and update data on newly released mystery novels.
[0259] Step 7: Incorporating the Emotion Engine
[0260] The server analyzes the user's feelings about the content using an emotion engine that recognizes the user's emotions.
[0261] The emotion engine recognizes the user's emotions from data such as text, voice, and facial expressions.
[0262] Example: When a user searches for "mystery novels" on their device, the emotion engine analyzes the user's emotions from their facial expressions and voice.
[0263] Step 8: Emotion-based content recommendation
[0264] The server dynamically updates and recommends content that reflects the user's emotional state based on the analysis results from the emotion engine.
[0265] For example, if a user expresses the emotion of "surprise" while browsing "Sherlock Holmes," the server will recommend "mystery novels" or "suspense movies" that are likely to induce similar feelings of "surprise."
[0266] Step 9: Provide the user interface
[0267] The server provides a user interface that includes a relationship map and keyword search functionality.
[0268] Users can enter keywords that interest them and relevant content will be displayed visually.
[0269] For example, if a user searches for "detective" on their device, the server will display a visual map of related content such as "Sherlock Holmes" and "Hercule Poirot."
[0270] Step 10: User Interaction
[0271] Users can click on the displayed relationship map to view more detailed information.
[0272] Additionally, users can enter new search keywords to find related content.
[0273] For example, if a user clicks on "Sherlock Holmes," the server will display detailed information and reviews about the work, as well as information about related TV shows and movies.
[0274] Example 2
[0275] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0276] Conventional information gathering and recommendation systems have struggled to accurately recommend content that users are truly interested in. Furthermore, there are limited ways to visually display the relationships between vast amounts of content in a way that is easy for users to understand. Furthermore, there is a lack of a dynamic content recommendation function based on user emotions, so there is a need for an improved user experience.
[0277] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0278] In this invention, the server includes means for collecting text data, means for extracting meta information from the collected text data, means for mapping relationships between content items based on the extracted meta information, means for visually displaying the mapped relationships, means for analyzing user emotions, and means for recommending content items based on the user emotions. This allows for accurate recommendation of content items that interest the user, facilitating understanding of the content, and enabling dynamic content recommendation based on emotions.
[0279] "Text data" refers to the content of sentences, books, articles, documents, etc. stored in digital format.
[0280] "Meta information" refers to additional information about content, such as keywords, tags, and names.
[0281] "Relationships between content" refers to the connections based on the characteristics and tags shared between multiple pieces of content.
[0282] "Relationship mapping" refers to a data structure or graph that visually represents commonalities and connections between content.
[0283] "Visual display means" refers to a method of displaying information in a way that is easy for users to understand, such as in the form of graphs or maps.
[0284] "Analysis of user emotions" refers to the means of detecting and recognizing a user's emotional state based on their facial expressions, voice, and behavior.
[0285] "Means for recommending content" refers to methods for presenting appropriate content to users based on analysis results and relationship mapping.
[0286] A "database" refers to a systematic storage system that facilitates the collection, management, and retrieval of data.
[0287] "Domestic and international databases" refers to databases that exist and are accessible domestically and internationally.
[0288] "Commonalities" refer to characteristics or features shared by multiple pieces of content.
[0289] "Link" refers to a connection used to indicate a relationship between pieces of content.
[0290] This invention is a system that uses AI and natural language processing technology to extract and visualize the relationships between a wide variety of content, and combines this with an emotion engine that recognizes the user's emotions to help users discover new content that interests them. The following hardware and software are used to implement this invention.
[0291] Hardware:
[0292] Server: This is the central hardware for processing data collection, analysis, mapping, display, updates, and emotion recognition.
[0293] User terminal: A device used to view content and input emotions. Specifically, this includes PCs, smartphones, tablets, etc.
[0294] software:
[0295] API: An interface for retrieving data from online databases. Examples include the Google Books API and the Open Library API.
[0296] Natural language processing engine: A tool for extracting meta-information (keywords, tags, people's names, etc.) from text data.
[0297] Emotion engine: Software for analyzing the user's emotional state.
[0298] Data processing and calculation:
[0299] 1. Content Collection:
[0300] The server uses the Google Books API and Open Library API to collect various content, such as book data. When a user enters a specific keyword, related data is automatically retrieved. For example, to retrieve book data related to "mystery novels," the server calls the Google Books API and obtains information such as the title, author, publication year, and summary of the related book.
[0301] 2. Text Analysis:
[0302] The server runs the collected text data through a natural language processing engine, performing morphological analysis and entity recognition to extract meta-information. For example, it analyzes the text of the novel "Sherlock Holmes" and extracts keywords such as "Sherlock Holmes," "John Watson," and the place name "London."
[0303] 3. Keyword tag generation:
[0304] Based on the extracted meta information, the server generates keywords and tags related to each piece of content and stores them in a database. For example, a "Sherlock Holmes" novel might be tagged with "detective," "mystery," and "deduction."
[0305] 4. Relationship Mapping:
[0306] Based on the generated keywords and tags, the server analyzes and maps the relationships between multiple pieces of content. For example, "Sherlock Holmes" and "Hercule Poirot" both share the tag "detective," so they are linked together.
[0307] 5. Database integration and updates:
[0308] The server periodically connects to an external database to obtain the latest content information and update the internal database, thereby ensuring that the latest information is always available to users.
[0309] 6. Incorporating an emotional engine:
[0310] Facial expression and voice data acquired from the user's device is input into the emotion engine, which analyzes the user's emotional state in real time. For example, if the user makes an expression showing surprise or interest, that emotional information is sent to the server, and related content is recommended.
[0311] 7. Emotion-based content recommendation:
[0312] Based on the analysis results obtained from the emotion engine, the server dynamically recommends content according to the user's emotional state. For example, if a user expresses surprise while browsing "Sherlock Holmes," the server will recommend mystery novels or suspense movies.
[0313] Specific examples
[0314] Example prompt sentence:
[0315] "If a user wants to collect data related to 'mystery novels,' the server retrieves book information related to 'mystery novels' via the Google Books API. This information includes the book's title, author, publication year, and summary of the contents."
[0316] In this way, the system of the present invention provides users with efficient information provision and enables them to encounter new content through each stage of collection, analysis, mapping, display, update, and emotion recognition.
[0317] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0318] Step 1:
[0319] The server collects the content data.
[0320] Input: User search keywords (e.g., "mystery novel")
[0321] The server uses the Google Books API or Open Library API to obtain book data related to the relevant keywords.
[0322] Output: A book dataset containing book titles, authors, publication years, and summary summaries.
[0323] Specific operation: The server generates a request to call the API and stores the obtained data in an internal database.
[0324] Step 2:
[0325] The server analyzes the collected text data using natural language processing.
[0326] Input: A collected book dataset
[0327] The server uses a natural language processing engine to extract meta-information from the collected text data (e.g., keywords, tags, names).
[0328] Output: Extracted meta information (e.g. "Sherlock Holmes", "Detective", "London")
[0329] How it works: The server uses morphological analyzers to break down text into words and entity recognition algorithms to identify important names and places.
[0330] Step 3:
[0331] The server generates keywords and tags based on the extracted meta information.
[0332] Input: Meta information (keywords, tags, name)
[0333] The server generates keywords and tags related to each piece of content and stores them in a database.
[0334] Output: Keywords and tags associated with the content (e.g., "detective," "mystery," "deduction")
[0335] Specific operation: The server executes a tag generation algorithm based on the extracted meta information and stores the generated tags in a database.
[0336] Step 4:
[0337] The server maps relationships between content based on keywords and tags.
[0338] Input: Keywords and tags associated with the content
[0339] The server analyzes the relationships between multiple pieces of content based on tags and keywords and generates links.
[0340] Output: Relationship map (e.g., a link connecting "Sherlock Holmes" and "Hercule Poirot")
[0341] Specific operation: The server uses a graph database to detect commonalities between tags and generate data to visualize them as links.
[0342] Step 5:
[0343] The server works in conjunction with an external database to periodically update the content information.
[0344] Input: Updates from an external database
[0345] The server obtains the latest content information using Open Library APIs and updates its internal database.
[0346] Output: Updated database (latest book information)
[0347] What it does: The server runs scheduled jobs to retrieve new data from external databases to keep the internal system up to date.
[0348] Step 6:
[0349] The server analyzes the user's emotions using an emotion engine.
[0350] Input: User's facial expressions and voice data
[0351] The server uses an emotion engine to analyze the user's facial expressions and voice to detect emotions.
[0352] Output: User's emotional state (e.g., "surprise," "interest")
[0353] Specific operation: Data collected through the camera and microphone of the user's device is sent to the emotion engine, and emotion analysis is performed in real time.
[0354] Step 7:
[0355] The server dynamically recommends content based on the analysis results of the emotion engine.
[0356] Input: User's emotional state from the emotion engine
[0357] The server runs an algorithm that recommends optimal content based on the user's emotional state.
[0358] Output: A list of recommended content (e.g., related "mystery novels" or "suspense movies")
[0359] Specific operation: The server searches for content that best matches the emotional state and sends a recommendation list to the user terminal.
[0360] Step 8:
[0361] The server builds a user interface that provides a relationship map and keyword search functionality.
[0362] Input: User search keywords and clicks
[0363] The server generates a relationship map to visually display related content.
[0364] Output: Visual relationship map and detailed information
[0365] Specific operation: When a user performs a search on their device, the server visually displays related content, allowing the user to access detailed information by clicking.
[0366] (Application example 2)
[0367] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0368] In today's information-overloaded environment, it is difficult for users to efficiently discover relevant content based on their interests and emotions. Conventional systems simply recommend content based on keywords and tags, but they are unable to dynamically recommend content based on users' emotions and preferences. Given this background, there is a need for a system that allows users to discover relevant content through a more personalized experience.
[0369] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[0370] In this invention, the server includes means for extracting meta-information from collected text data, means for mapping relationships between content items based on the extracted meta-information, means for visually displaying the mapped relationships, means for recognizing a user's emotion, and means for dynamically recommending content items based on the recognized emotion, thereby enabling efficient discovery and recommendation of related content items based on the user's emotion and interests.
[0371] "Collected text data" refers to text information such as books, lyrics, audio data, and image data collected from a wide variety of content.
[0372] "Meta information" refers to information such as keywords, tags, and content names extracted from text data using a natural language processing engine.
[0373] "Relationships between content" refers to relationships that extract commonalities and associations between different content based on collected meta-information and show them as links.
[0374] "Visual display means" refers to means for displaying the relationships between analyzed contents in the form of a map, graph, list, etc., so that the user can visually understand the relationships between the analyzed contents.
[0375] "Means for recognizing user emotions" refers to a means for recognizing a user's emotional state by analyzing the user's facial expressions and voice using sensors such as a camera and microphone.
[0376] "Means for dynamic recommendation based on emotions" refers to a means by which the system selects and recommends relevant content in real time based on the user's recognized emotions.
[0377] This invention is a system that recognizes the user's emotions and recommends appropriate content based on those emotions. The system mainly involves three parties: a server, a terminal, and a user.
[0378] First, the server is equipped with a means for extracting meta-information from the collected text data. This meta-information includes keywords, tags, content names, etc. Specifically, the text data is analyzed using technologies such as natural language processing engines, morphological analysis, and entity recognition to extract information. For example, the contents of books are collected as text data, and then analyzed to extract keywords such as "mystery" and "detective."
[0379] Next, the server has a means for mapping the relationships between content items based on the extracted meta information. This involves visually showing the commonalities and connections between different pieces of content using the extracted keywords and tags. Possible visual display methods include maps, graphs, and lists. For example, multiple books or video works with the keyword "detective" could be linked together to show the user their relationships.
[0380] Furthermore, the device is equipped with a means of recognizing the user's emotions. This is achieved using the smartphone's camera and microphone. Specific software uses OpenCV facial recognition technology and EmotionRecognizer to recognize emotions from the user's facial expressions and voice. For example, if a user shows a surprised expression while reading a mystery novel, the device recognizes that emotion and sends it to the server.
[0381] The server has a means for receiving the user's emotional information and dynamically recommending content based on the recognition results. This allows content related to the user's emotions to be recommended in real time. For example, if the user expresses surprise, mystery novels or suspense movies that are even more surprising and intriguing will be recommended. This allows users to efficiently discover content that matches their emotions and preferences.
[0382] As a specific example, when a user searches for "mystery novels," multiple related content results are displayed. The emotion recognition engine analyzes the user's facial expressions and voice to indicate interest, and the server then recommends further related content based on the results. For example, if a user expresses the emotion "surprise," the server can further recommend "suspense movies," thereby personalizing the user experience. The following are examples of input prompts for the generative AI model:
[0383] "When a user searches for 'mystery novels,' get the 10 most relevant books from the Google Books API, analyze the content of each, and generate keywords and tags. If you detect the emotion 'surprise' from the user's facial expression, also recommend 'suspense movies.'"
[0384] As described above, the system of the present invention provides efficient information provision and enables users to encounter new content through each stage of collection, analysis, mapping, display, emotion recognition, and dynamic recommendation.
[0385] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0386] Step 1:
[0387] The user enters a keyword
[0388] The user enters keywords into the search bar on the device, and the device sends a search request to the server based on this input.
[0389] Step 2:
[0390] The server collects the content
[0391] The server retrieves content data related to keywords from an online database via API. This data includes a wide variety of content, such as books, movies, and music, and is sent to the server. For example, if a user searches for "mystery novels," related book information is retrieved from the Google Books API. The server then stores this retrieved data.
[0392] Step 3:
[0393] The server parses the text data
[0394] The server uses a natural language processing engine to analyze the collected content data. Specifically, it performs morphological analysis and entity recognition to extract keywords and tags. The results of this analysis are saved as meta-information. For example, keywords such as "detective" and "deduction" are extracted from the text data of "Sherlock Holmes."
[0395] Step 4:
[0396] The server maps the relationships between content
[0397] The server maps the relationships between content based on the extracted meta information, finding commonalities in keywords and tags and visualizing them as links. For example, it links works with the common tag "detective" and generates data for displaying them as a visual map.
[0398] Step 5:
[0399] The server visually displays the relationships
[0400] The device displays the relationship map generated by the server to the user, allowing the user to visually check related content. For example, if a user searches for "detective," related books and movies are displayed in map format. The user can click on this map to obtain more information.
[0401] Step 6:
[0402] The device recognizes the user's emotions
[0403] While a user is viewing content displayed on the device, the device uses a camera and microphone to analyze the user's facial expressions and voice. An emotion recognition engine is used to recognize the user's emotional state (surprise, interest, etc.). The results of this analysis are sent to the server.
[0404] Step 7:
[0405] The server recommends content based on emotions.
[0406] The server dynamically recommends content appropriate to the user's emotional state based on the results of the emotion recognition it receives. For example, if the user expresses surprise, the server generates data to recommend similarly interesting suspense movies or mystery novels and sends it to the device.
[0407] Step 8:
[0408] The device displays recommended content to the user
[0409] The device receives recommended content from the server and displays it to the user. The user can browse this recommended content and discover new content that piques their interest. For example, by expressing the emotion "surprise" while browsing "Sherlock Holmes," the device will recommend "suspense movies" and display them to the user.
[0410] These are the specific processing steps and their flow. Through these processes, personalized content recommendations based on the user's emotions and interests are realized.
[0411] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0412] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0413] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.
[0414] [Second embodiment]
[0415] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.
[0416] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0417] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0418] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.
[0419] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0420] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0421] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0422] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0423] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0424] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0425] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0426] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."
[0427] This invention is a system that uses AI and natural language processing technology to extract and visualize the relationships between a wide variety of content, helping users discover new content that they are interested in. The basic process for implementing this system is shown below.
[0428] Content Collection
[0429] The server collects a variety of content, such as books, lyrics, audio data, image data, etc. At this stage, data may be retrieved from online databases via APIs, or files may be uploaded directly to the server.
[0430] Example: A server calls the Google Books API to retrieve book data related to a specific keyword. If a user wants to collect data related to "mystery novels," the server uses the Google Books API to retrieve book information related to "mystery novels." This information includes the book's title, author, publication year, summary of the content, etc.
[0431] Text analytics
[0432] The server sends the collected text data to a natural language processing engine to extract meta-information (keywords, tags, content names, etc.). This process uses technologies such as morphological analysis, entity recognition, and relationship extraction.
[0433] Example: The server analyzes data from the novel "Sherlock Holmes" and extracts character names such as "Sherlock Holmes" and "John Watson" and keywords such as "London" and "detective."
[0434] Keyword tag generation
[0435] Based on the extracted meta information, the server generates keywords and tags related to each piece of content, including important people, places, events, and themes.
[0436] Example: The server assigns tags such as "detective," "mystery," and "deduction" to "Sherlock Holmes" novels.
[0437] Relationship Mapping
[0438] The server maps the relationships between pieces of content based on keywords and tags, finding commonalities between pieces of content and creating a data structure that visually displays them as links.
[0439] Example: The server detects that the novels "Sherlock Holmes" and "Hercule Poirot" share the common keyword "detective" and links them together.
[0440] Database integration and updates
[0441] The server works in conjunction with other databases to regularly update and supplement the content information, ensuring that the most up-to-date and detailed information is always available.
[0442] Example: The server works with the Open Library API to periodically retrieve and update data on newly released "mystery novels."
[0443] User interface provided
[0444] The server provides a user interface that includes a relationship map and keyword search functionality, allowing users to enter keywords of interest and visually display related content.
[0445] Example: When a user searches for "detective" on their device, the server displays a visual map of related content, such as "Sherlock Holmes" and "Hercule Poirot."
[0446] User Interactions
[0447] Users can click on the displayed relationship map to view more detailed information, and can also enter new search keywords to find related content.
[0448] Example: When a user clicks on "Sherlock Holmes," the server displays detailed information and reviews about the work, as well as information about related TV shows and movies.
[0449] As described above, the system of the present invention provides efficient information provision and encounters with new content through each stage of collection, analysis, mapping, display, and update.
[0450] The processing flow will be explained below.
[0451] Step 1: Gather content
[0452] The server collects a variety of content such as books, lyrics, audio data, and image data.
[0453] Specifically, it retrieves data from online databases via API and also receives files uploaded by users.
[0454] Example: The server calls the Google Books API to retrieve metadata for books related to "fantasy novels."
[0455] Step 2: Registering text data storage
[0456] The server stores the collected text data in storage.
[0457] The data to be saved includes books, lyrics, text converted from audio data, and text extracted from images using OCR.
[0458] Example: A server stores the book data for "Harry Potter and the Philosopher's Stone" in a database.
[0459] Step 3: Text analysis
[0460] The server sends the stored text data to a natural language processing engine to extract meta-information.
[0461] It uses morphological analysis, entity recognition, and relationship extraction to identify characters, places, important events, etc.
[0462] Example: The server uses a natural language processing engine to extract character names such as "Sherlock Holmes" and "John Watson" from the text "Sherlock Holmes."
[0463] Step 4: Generate keywords and tags
[0464] The server generates keywords and tags related to each piece of content based on the extracted meta information.
[0465] This includes important people, places, events, themes, etc.
[0466] Example: The server assigns tags such as "detective," "mystery," and "deduction" to "Sherlock Holmes."
[0467] Step 5: Relationship mapping
[0468] The server maps the relationships between content based on the generated keywords and tags.
[0469] Create a data structure that finds commonalities between pieces of content and visually displays them as links.
[0470] Example: The server discovers that "Sherlock Holmes" and "Hercule Poirot" share the tag "detective" and links them together.
[0471] Step 6: Database integration and updates
[0472] The server cooperates with other databases to periodically retrieve and update content information.
[0473] This ensures that the latest information is always provided.
[0474] Example: The server works with the Open Library API to periodically retrieve and update data on newly released mystery novels.
[0475] Step 7: Provide the user interface
[0476] The server provides a user interface that includes a relationship map and keyword search functionality.
[0477] Users can enter keywords that interest them and relevant content will be displayed visually.
[0478] For example, if a user searches for "detective" on their device, the server will display related content such as "Sherlock Holmes" or "Hercule Poirot."
[0479] Step 8: User Interaction
[0480] Users can click on the displayed relationship map to view more detailed information.
[0481] Additionally, users can enter new search keywords to find related content.
[0482] For example, if a user clicks on "Sherlock Holmes," the server will display detailed information and reviews about the work, as well as information about related TV shows and movies.
[0483] Example 1
[0484] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0485] Conventional content management systems have difficulty efficiently extracting and visualizing the relationships between different types of content. Furthermore, the means by which users can discover new related content are limited, which does not sufficiently improve the user experience. Furthermore, there is insufficient integration with other databases to continuously acquire and update the latest information, making it difficult to ensure the freshness and diversity of content.
[0486] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0487] In this invention, the server includes a means for collecting information, a means for preprocessing the collected text data and extracting meta information, a means for generating keywords and tags based on the extracted meta information, a means for mapping relationships between content using the generated keywords and tags, and a means for visually displaying the mapped relationships. This enables efficient extraction and visualization of relationships between content of different formats. Furthermore, by including a means for acquiring and updating content data from domestic and international databases in cooperation with the server, it becomes possible to continuously acquire and provide the latest and most diverse information to users. Furthermore, by including a means for finding commonalities between mapped content and displaying them as links, users can intuitively discover new related content.
[0488] "Means for collecting information" refers to the means for collecting various content such as books, lyrics, audio data, and image data on a server through methods such as online databases and file uploads.
[0489] "Preprocessing" refers to the process of removing unnecessary characters and duplication from collected content data, formatting the text data, and preparing it for passing to a natural language processing engine.
[0490] "Meta information" is summary information of data such as keywords, tags, and content names extracted from collected text data, and is important information for understanding the meaning and substance of the content.
[0491] The "means for generating keywords and tags" refers to a means for automatically generating keywords and tags related to each piece of content based on meta information obtained from the natural language processing engine.
[0492] A "means for mapping relationships between content" is a means for finding commonalities and relationships between multiple pieces of content based on generated keywords and tags, and visually displaying them as links.
[0493] A "visual display means" is a means for displaying the relationships between the mapped content in a graphical format that allows users to intuitively understand the relationships.
[0494] "Means for obtaining and updating content data from domestic and international databases in cooperation with the server" refers to means for the server to periodically communicate with external databases to obtain new content data and update existing data.
[0495] "Means for finding commonalities and displaying them as links" refers to means for finding common keywords and tags between multiple pieces of content and visually displaying the relationships between them as links.
[0496] This invention is a system that efficiently extracts and visualizes the relationships between different types of content, allowing users to discover new related content. This system is run by a server, a terminal, and user operations.
[0497] Hardware and software used
[0498] This system is composed of the following hardware and software components:
[0499] 1. Hardware
[0500] Server: Collects, processes, and stores data.
[0501] Terminal: A device (e.g., PC, smartphone, tablet) that a user uses to access and operate the system.
[0502] 2. Software
[0503] API: Application Program Interface for connecting with external services for data collection.
[0504] Natural language processing engine: An engine for analyzing text data and extracting meta-information (e.g., morphological analysis engine, entity recognition engine).
[0505] Database: A database management system for storing and managing the collected and generated data.
[0506] Visualization libraries: Libraries for visually displaying relationships between content (e.g., D3.js).
[0507] Specific processing content of the program
[0508] First, the server collects various content such as books, lyrics, audio data, image data, etc. through APIs and file upload functions. For example, if a user wants to collect data related to "mystery novels," the server calls the API of an online database and retrieves book information as a result.
[0509] The server then preprocesses the collected text data and sends it to a natural language processing engine to extract meta-information, using a morphological analysis engine and entity recognition engine to extract important keywords and tags from the text.
[0510] The server then automatically generates related keywords and tags based on the extracted meta information and maps the relationships between content items. For example, if the novels of "Sherlock Holmes" and "Hercule Poirot" share the keyword "detective," these pieces of content are linked.
[0511] The server also periodically connects to external databases (e.g., Open Library API) to retrieve new content data and update the database, ensuring that users always have access to the latest information.
[0512] When a user accesses the system using a terminal and searches for a specific keyword, the server visually displays the data mapping the relationships. For example, if a user searches for "detective," related content will be displayed as a visual map.
[0513] Additionally, when a user clicks on a particular piece of content on the relationship map, the server provides more information about it. For example, if a user clicks on "Sherlock Holmes," the server displays more information about the work, related reviews, and other related content.
[0514] Prompt Sentence Examples
[0515] "Generate a program to collect, analyze, and display data to map relationships in a detective novel."
[0516] "Please explain in detail the process of the system that extracts relationships between content based on meta information and displays them visually."
[0517] The system is designed to help users discover new and relevant content and improve the overall user experience.
[0518] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0519] Step 1:
[0520] Content Collection
[0521] The server collects various content, such as books, lyrics, audio data, and image data, through API calls or file uploads. It receives keywords and files specified by the user as input. The server saves the collected data in its internal storage. It also uses APIs to retrieve data related to specific keywords from external databases. For example, if a user enters "mystery novels," the server calls the Google Books API and retrieves book information as a result. The output is the collected raw data.
[0522] Step 2:
[0523] Text Extraction and Preprocessing
[0524] The server extracts text from the collected content data and performs preprocessing. It receives the collected raw data as input. Specifically, it performs formatting processing by removing unnecessary characters and duplication from the text file. It then prepares the data for passing to the natural language processing engine. The output is preprocessed text data.
[0525] Step 3:
[0526] Extracting Meta Information
[0527] The server sends the preprocessed text data to a natural language processing engine to extract meta-information. The server receives the preprocessed text data as input. Specifically, it analyzes the text using a morphological analysis engine (e.g., Mecab) to extract meta-information such as keywords, tags, and content names. The output is the extracted meta-information.
[0528] Step 4:
[0529] Keyword tag generation
[0530] The server generates keywords and tags appropriate for each piece of content based on the extracted meta information. It receives meta information as input. Specifically, it extracts important people, places, events, themes, etc., and generates keywords and tags based on them. For example, it generates tags such as "detective," "mystery," and "deduction" from the text "Sherlock Holmes." The output is the generated keywords and tags.
[0531] Step 5:
[0532] Relationship mapping
[0533] The server maps the relationships between different content based on the generated keywords and tags. It receives the generated keywords and tags as input. Specifically, it links content that has common keywords or tags together to create a relational data structure. For example, if the texts "Sherlock Holmes" and "Hercule Poirot" share the tag "detective," they are linked together. The output is the mapped relationship data.
[0534] Step 6:
[0535] Visually mapping your data
[0536] The server performs visualization processing to visually display the relationship data. It receives the mapped relationship data as input. Specifically, it uses a visualization library such as D3.js to generate a relationship map in HTML format and sends it to the user's device. The output is a visual relationship map.
[0537] Step 7:
[0538] Database integration and updates
[0539] The server interacts with an external database to retrieve new content data and update existing data. It receives new data retrieved from the external database's API as input. Specifically, it periodically sends API requests to add new book and content data to the database and update existing data. The output is an updated database.
[0540] Step 8:
[0541] Providing a user interface
[0542] The server provides an interface that allows users to manipulate the relationship map and search results. It receives search keywords and click operations from the user as input. Specific operations include searching for related content based on the user's input and displaying it on the device as a visual map. The output is the visual map and search results displayed on the user interface.
[0543] Step 9:
[0544] Viewing detailed information
[0545] When a user clicks on a specific piece of content on the relationship map, the server displays its detailed information. It receives the user's click as input and retrieves detailed information about the related content from the database. Specifically, it provides the user with detailed information about the content, reviews, and other related content. The output is an interface displaying the detailed information.
[0546] (Application example 1)
[0547] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0548] In recent years, a wide variety of content has become available on the Internet, but it is not easy for users to discover new content that interests them. In particular, there is a need to efficiently search for content that matches users' interests from a vast amount of content data and show its relevance. Conventional methods have difficulty visually displaying the relationships between content, and there is a lack of systems that guide users to new content that interests them. To solve this problem, it is necessary to concretely visualize the abstract relationships between content and provide an interface that users can intuitively understand.
[0549] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0550] In this invention, the server includes means for extracting meta-information from collected text data, means for mapping relationships between content items based on the extracted meta-information, means for visually displaying the mapped relationships, means for vectorizing the collected text data and calculating the similarity between the vectors, and means for generating a network graph based on the similarity. This allows users to easily find new content items related to their interests from among a vast amount of content items and intuitively understand their relationships.
[0551] "Collected text data" refers to text information in a variety of formats, such as books, audio data, and image data, obtained from the Internet or other databases.
[0552] "Meta information" refers to accompanying information such as keywords, tags, and content names extracted from collected text data.
[0553] "Relationships between contents" is information that indicates commonalities and interrelationships between different contents.
[0554] "Mapping means" refers to a technical method for visually showing the relationships between content items based on extracted meta-information.
[0555] "Visual display means" refers to a method of presenting the relationships between content to users in a visual format such as a graph or map.
[0556] "Vectorization" means converting collected text data into numerical vectors, making it mathematically processable.
[0557] The "means for calculating similarity" is a method for calculating the similarity between vectorized text data and evaluating the numerical relationship.
[0558] A "network graph" is a structure that visually represents the relationships between content using nodes and edges.
[0559] This invention is a system that extracts relationships between various contents and supports users in discovering new contents. The system consists of a server, a terminal, and a user interface.
[0560] Content Collection
[0561] The server collects various content, such as books, audio data, and image data, from the Internet and other databases. Data is periodically retrieved from online databases via APIs and stored on the server as text data. For example, data can be collected using an information search API or a book information API.
[0562] Text Analysis
[0563] The server sends the collected text data to a natural language processing engine to extract meta-information (keywords, tags, content names, etc.). This process utilizes technologies such as morphological analysis, entity recognition, and relationship extraction. Specifically, it uses natural language processing toolkits and entity recognition software.
[0564] Meta information generation
[0565] The server automatically generates keywords and tags related to each piece of content based on the extracted meta information, including important people, places, events, themes, etc. For example, tags are generated based on the characters and themes of a book.
[0566] Relationship mapping
[0567] The server maps the relationships between content based on the generated keywords and tags. At this stage, the similarity between vectorized data is calculated and the numerical relationships are evaluated. Specifically, tools such as TfidfVectorizer and Cosine Similarity are used for vectorization and similarity calculation.
[0568] Visualizing Relationships
[0569] The server generates a network graph to visually display the results of the relationship mapping. The generated graph is displayed in a visual format that allows users to intuitively understand it. Data visualization tools such as NetworkX and matplotlib are used to generate the network graph.
[0570] User interface provided
[0571] The server provides a user interface with a relationship map and keyword search functionality. Users can enter keywords of interest and visually display related content. The user interface is web-based and uses web frameworks such as Flask.
[0572] User Interactions
[0573] Users can click on the visually displayed relationship map to view more detailed information, and can also enter new search keywords to discover and enjoy new related content.
[0574] Examples of concrete examples and prompts
[0575] For example, if a user searches for "suspense movies," you can visualize related movies, TV shows, and even related podcasts and books. Visually display movies, TV shows, podcasts, and books related to "suspense movies."
[0576] This system enables users to efficiently discover new content that interests them and intuitively understand its relationships.
[0577] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0578] Step 1:
[0579] The server collects content data from the Internet and other databases. Specifically, it uses APIs (such as information search APIs and book information APIs) to obtain a variety of content data, including books, audio data, and image data, and stores it on the server. The input required is a search query containing keywords that interest the user, and the output is a list of collected text data.
[0580] Step 2:
[0581] The server sends the collected text data to a natural language processing engine to extract meta-information. This process uses techniques such as morphological analysis, entity recognition, and relationship extraction. Specifically, a natural language processing toolkit is used to analyze the text and extract keywords, tags, content names, etc. The input is the text data obtained in step 1, and the output is the extracted meta-information.
[0582] Step 3:
[0583] The server generates keywords and tags related to each piece of content based on the extracted meta information. Tags such as important people, places, events, and themes are automatically added. The input is the meta information obtained in step 2, and the output is a list of generated keywords and tags.
[0584] Step 4:
[0585] The server maps the relationships between content items based on the generated keywords and tags. It evaluates the similarity between content items by vectorizing the collected text data (using TfidfVectorizer) and calculating the similarity between each vector (using Cosine Similarity). The input is the keywords and tags obtained in Step 3 and the text data, and the output is a similarity score.
[0586] Step 5:
[0587] The server generates a network graph based on the similarity. It uses data visualization tools such as NetworkX and matplotlib to visualize the relationships between content. The input is the similarity scores obtained in step 4, and the output is the visualized network graph.
[0588] Step 6:
[0589] The server provides a user-facing interface. Users can enter keywords of interest and relevant content is displayed visually. A web framework such as Flask is used to build the web-based interface. The input is the keywords entered by the user, and the output is a visually displayed relationship map.
[0590] Step 7:
[0591] Users can click on the visually displayed relationship map to view more information, and can also enter new search keywords to discover new related content. The input is the user's actions on the relationship map, and the output is the discovery of more information and new content.
[0592] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0593] This invention is a system that uses AI and natural language processing technology to extract and visualize the relationships between a wide variety of content, and by combining this with an emotion engine that recognizes the user's emotions, helps users discover new content that they are interested in. The basic process for implementing this system is shown below.
[0594] Content Collection
[0595] The server collects various content such as books, lyrics, audio data, image data, etc. At this stage, it retrieves data from online databases via APIs and may also receive files uploaded by users.
[0596] Example: A server uses the Google Books API to retrieve book data related to a specific keyword. If a user wants to collect data related to "mystery novels," the server retrieves book information related to "mystery novels" via the Google Books API. This information includes the book title, author, publication year, and summary of the contents.
[0597] Text analytics
[0598] The server uses a natural language processing engine to extract meta-information (keywords, tags, content names, etc.) from the collected text data. This process involves morphological analysis, entity recognition, and relationship extraction.
[0599] Example: The server analyzes text data from "Sherlock Holmes" and extracts character names such as "Sherlock Holmes" and "John Watson" and keywords such as "London" and "detective."
[0600] Keyword tag generation
[0601] The server generates keywords and tags related to each piece of content based on the extracted meta information.
[0602] Example: The server assigns tags such as "detective," "mystery," and "deduction" to "Sherlock Holmes" novels.
[0603] Relationship Mapping
[0604] The server maps the relationships between content based on the generated keywords and tags, finding commonalities between the content and visually displaying them as links.
[0605] Example: The server discovers that "Sherlock Holmes" and "Hercule Poirot" share the tag "detective" and links them together.
[0606] Database integration and updates
[0607] The server periodically retrieves and updates content information in cooperation with other databases, ensuring that the latest information is always provided.
[0608] Example: The server works with the Open Library API to periodically retrieve and update data on newly released "mystery novels."
[0609] Incorporating an emotion engine
[0610] The server analyzes the user's feelings about the content using an emotion engine that recognizes the user's emotions.
[0611] Example: When a user searches for "mystery novels" on their device, the emotion engine analyzes the user's facial expressions and voice to determine their emotions. For example, if the user shows expressions of surprise or interest, relevant content will be recommended based on that.
[0612] Emotion-based content recommendation
[0613] The server dynamically updates and recommends content that reflects the user's emotional state based on the analysis results from the emotion engine.
[0614] Example: If a user expresses the emotion of "surprise" while browsing "Sherlock Holmes," the server will recommend "mystery novels" or "suspense movies" that are likely to induce similar feelings of "surprise."
[0615] User interface provided
[0616] The server provides a user interface that includes a relationship map and keyword search functionality, allowing users to enter keywords of interest and visually display related content.
[0617] Example: When a user searches for "detective" on their device, the server displays a visual map of related content, such as "Sherlock Holmes" and "Hercule Poirot."
[0618] Users can click on the displayed relationship map to view detailed information about their interactions, or enter new search keywords to find related content.
[0619] Example: When a user clicks on "Sherlock Holmes," the server displays detailed information and reviews about the work, as well as information about related TV shows and movies.
[0620] As described above, the system of the present invention provides efficient information provision and encounters with new content through each stage of collection, analysis, mapping, display, update, and emotion recognition.
[0621] The processing flow will be explained below.
[0622] Step 1: Gather content
[0623] The server collects a variety of content such as books, lyrics, audio data, and image data.
[0624] Specifically, it retrieves data from an online database via API and stores files uploaded by users in storage.
[0625] Example: A server uses the Google Books API to retrieve metadata for books related to "fantasy novels."
[0626] Step 2: Registering text data storage
[0627] The server stores the collected text data in a database.
[0628] The data stored includes books, lyrics, text converted from audio data, and text extracted from images.
[0629] Example: A server stores the book data for "Harry Potter and the Philosopher's Stone" in a database.
[0630] Step 3: Text analysis
[0631] The server sends the stored text data to a natural language processing engine to extract meta-information.
[0632] It uses morphological analysis, entity recognition, and relationship extraction to identify characters, places, important events, etc.
[0633] Example: The server analyzes the text data of "Sherlock Holmes" and extracts character names such as "Sherlock Holmes" and "John Watson" as well as keywords such as "London" and "detective."
[0634] Step 4: Generate keywords and tags
[0635] The server generates keywords and tags related to each piece of content based on the extracted meta information.
[0636] This includes important people, places, events, themes, etc.
[0637] Example: The server tags "detective," "mystery," and "deduction" to "Sherlock Holmes" novels.
[0638] Step 5: Relationship mapping
[0639] The server maps the relationships between content based on the generated keywords and tags.
[0640] Create a data structure that finds commonalities between pieces of content and visually displays them as links.
[0641] Example: The server discovers that "Sherlock Holmes" and "Hercule Poirot" share the tag "detective" and links them together.
[0642] Step 6: Database integration and updates
[0643] The server cooperates with other databases to periodically retrieve and update content information.
[0644] This ensures that the latest information is always provided.
[0645] Example: The server works with the Open Library API to periodically retrieve and update data on newly released mystery novels.
[0646] Step 7: Incorporating the Emotion Engine
[0647] The server analyzes the user's feelings about the content using an emotion engine that recognizes the user's emotions.
[0648] The emotion engine recognizes the user's emotions from data such as text, voice, and facial expressions.
[0649] Example: When a user searches for "mystery novels" on their device, the emotion engine analyzes the user's emotions from their facial expressions and voice.
[0650] Step 8: Emotion-based content recommendation
[0651] The server dynamically updates and recommends content that reflects the user's emotional state based on the analysis results from the emotion engine.
[0652] For example, if a user expresses the emotion of "surprise" while browsing "Sherlock Holmes," the server will recommend "mystery novels" or "suspense movies" that are likely to induce similar feelings of "surprise."
[0653] Step 9: Provide the user interface
[0654] The server provides a user interface that includes a relationship map and keyword search functionality.
[0655] Users can enter keywords that interest them and relevant content will be displayed visually.
[0656] For example, if a user searches for "detective" on their device, the server will display a visual map of related content such as "Sherlock Holmes" and "Hercule Poirot."
[0657] Step 10: User Interaction
[0658] Users can click on the displayed relationship map to view more detailed information.
[0659] Additionally, users can enter new search keywords to find related content.
[0660] For example, if a user clicks on "Sherlock Holmes," the server will display detailed information and reviews about the work, as well as information about related TV shows and movies.
[0661] Example 2
[0662] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0663] Conventional information gathering and recommendation systems have struggled to accurately recommend content that users are truly interested in. Furthermore, there are limited ways to visually display the relationships between vast amounts of content in a way that is easy for users to understand. Furthermore, there is a lack of a dynamic content recommendation function based on user emotions, so there is a need for an improved user experience.
[0664] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0665] In this invention, the server includes means for collecting text data, means for extracting meta information from the collected text data, means for mapping relationships between content items based on the extracted meta information, means for visually displaying the mapped relationships, means for analyzing user emotions, and means for recommending content items based on the user emotions. This allows for accurate recommendation of content items that interest the user, facilitating understanding of the content, and enabling dynamic content recommendation based on emotions.
[0666] "Text data" refers to the content of sentences, books, articles, documents, etc. stored in digital format.
[0667] "Meta information" refers to additional information about content, such as keywords, tags, and names.
[0668] "Relationships between content" refers to the connections based on the characteristics and tags shared between multiple pieces of content.
[0669] "Relationship mapping" refers to a data structure or graph that visually represents commonalities and connections between content.
[0670] "Visual display means" refers to a method of displaying information in a way that is easy for users to understand, such as in the form of graphs or maps.
[0671] "Analysis of user emotions" refers to the means of detecting and recognizing a user's emotional state based on their facial expressions, voice, and behavior.
[0672] "Means for recommending content" refers to methods for presenting appropriate content to users based on analysis results and relationship mapping.
[0673] A "database" refers to a systematic storage system that facilitates the collection, management, and retrieval of data.
[0674] "Domestic and international databases" refers to databases that exist and are accessible domestically and internationally.
[0675] "Commonalities" refer to characteristics or features shared by multiple pieces of content.
[0676] "Link" refers to a connection used to indicate a relationship between pieces of content.
[0677] This invention is a system that uses AI and natural language processing technology to extract and visualize the relationships between a wide variety of content, and combines this with an emotion engine that recognizes the user's emotions to help users discover new content that interests them. The following hardware and software are used to implement this invention.
[0678] Hardware:
[0679] Server: This is the central hardware for processing data collection, analysis, mapping, display, updates, and emotion recognition.
[0680] User terminal: A device used to view content and input emotions. Specifically, this includes PCs, smartphones, tablets, etc.
[0681] software:
[0682] API: An interface for retrieving data from online databases. Examples include the Google Books API and the Open Library API.
[0683] Natural language processing engine: A tool for extracting meta-information (keywords, tags, people's names, etc.) from text data.
[0684] Emotion engine: Software for analyzing the user's emotional state.
[0685] Data processing and calculation:
[0686] 1. Content Collection:
[0687] The server uses the Google Books API and Open Library API to collect various content, such as book data. When a user enters a specific keyword, related data is automatically retrieved. For example, to retrieve book data related to "mystery novels," the server calls the Google Books API and obtains information such as the title, author, publication year, and summary of the related book.
[0688] 2. Text Analysis:
[0689] The server runs the collected text data through a natural language processing engine, performing morphological analysis and entity recognition to extract meta-information. For example, it analyzes the text of the novel "Sherlock Holmes" and extracts keywords such as "Sherlock Holmes," "John Watson," and the place name "London."
[0690] 3. Keyword tag generation:
[0691] Based on the extracted meta information, the server generates keywords and tags related to each piece of content and stores them in a database. For example, a "Sherlock Holmes" novel might be tagged with "detective," "mystery," and "deduction."
[0692] 4. Relationship Mapping:
[0693] Based on the generated keywords and tags, the server analyzes and maps the relationships between multiple pieces of content. For example, "Sherlock Holmes" and "Hercule Poirot" both share the tag "detective," so they are linked together.
[0694] 5. Database integration and updates:
[0695] The server periodically connects to an external database to obtain the latest content information and update its internal database, thereby ensuring that users are always provided with the latest information.
[0696] 6. Incorporating an Emotion Engine:
[0697] Facial expression and voice data acquired from the user's device is input into the emotion engine, which analyzes the user's emotional state in real time. For example, if the user makes an expression showing surprise or interest, that emotional information is sent to the server, and related content is recommended.
[0698] 7. Emotion-based content recommendation:
[0699] Based on the analysis results obtained from the emotion engine, the server dynamically recommends content according to the user's emotional state. For example, if a user expresses surprise while browsing "Sherlock Holmes," the server will recommend mystery novels or suspense movies.
[0700] Specific examples
[0701] Example prompt sentence:
[0702] "If a user wants to collect data related to 'mystery novels,' the server retrieves book information related to 'mystery novels' via the Google Books API. This information includes the book's title, author, publication year, and summary of the contents."
[0703] In this way, the system of the present invention provides users with efficient information provision and enables them to encounter new content through each stage of collection, analysis, mapping, display, update, and emotion recognition.
[0704] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0705] Step 1:
[0706] The server collects the content data.
[0707] Input: User search keywords (e.g., "mystery novel")
[0708] The server uses the Google Books API or Open Library API to obtain book data related to the relevant keywords.
[0709] Output: A book dataset containing book titles, authors, publication years, and summary summaries.
[0710] Specific operation: The server generates a request to call the API and stores the obtained data in an internal database.
[0711] Step 2:
[0712] The server analyzes the collected text data using natural language processing.
[0713] Input: A collected book dataset
[0714] The server uses a natural language processing engine to extract meta-information from the collected text data (e.g., keywords, tags, names).
[0715] Output: Extracted meta information (e.g. "Sherlock Holmes", "Detective", "London")
[0716] How it works: The server uses morphological analyzers to break down text into words and entity recognition algorithms to identify important names and places.
[0717] Step 3:
[0718] The server generates keywords and tags based on the extracted meta information.
[0719] Input: Meta information (keywords, tags, name)
[0720] The server generates keywords and tags related to each piece of content and stores them in a database.
[0721] Output: Keywords and tags associated with the content (e.g., "detective," "mystery," "deduction")
[0722] Specific operation: The server executes a tag generation algorithm based on the extracted meta information and stores the generated tags in a database.
[0723] Step 4:
[0724] The server maps relationships between content based on keywords and tags.
[0725] Input: Keywords and tags associated with the content
[0726] The server analyzes the relationships between multiple pieces of content based on tags and keywords and generates links.
[0727] Output: Relationship map (e.g., a link connecting "Sherlock Holmes" and "Hercule Poirot")
[0728] Specific operation: The server uses a graph database to detect commonalities between tags and generate data to visualize them as links.
[0729] Step 5:
[0730] The server works in conjunction with an external database to periodically update the content information.
[0731] Input: Updates from an external database
[0732] The server obtains the latest content information using Open Library APIs and updates its internal database.
[0733] Output: Updated database (latest book information)
[0734] What it does: The server runs scheduled jobs to retrieve new data from external databases to keep the internal system up to date.
[0735] Step 6:
[0736] The server analyzes the user's emotions using an emotion engine.
[0737] Input: User's facial expressions and voice data
[0738] The server uses an emotion engine to analyze the user's facial expressions and voice to detect emotions.
[0739] Output: User's emotional state (e.g., "surprise," "interest")
[0740] Specific operation: Data collected through the camera and microphone of the user's device is sent to the emotion engine, and emotion analysis is performed in real time.
[0741] Step 7:
[0742] The server dynamically recommends content based on the analysis results of the emotion engine.
[0743] Input: User's emotional state from the emotion engine
[0744] The server runs an algorithm that recommends optimal content based on the user's emotional state.
[0745] Output: A list of recommended content (e.g., related "mystery novels" or "suspense movies")
[0746] Specific operation: The server searches for content that best matches the emotional state and sends a recommendation list to the user terminal.
[0747] Step 8:
[0748] The server builds a user interface that provides a relationship map and keyword search functionality.
[0749] Input: User search keywords and clicks
[0750] The server generates a relationship map to visually display related content.
[0751] Output: Visual relationship map and detailed information
[0752] Specific operation: When a user performs a search on their device, the server visually displays related content, allowing the user to access detailed information by clicking.
[0753] (Application example 2)
[0754] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0755] In today's information-overloaded environment, it is difficult for users to efficiently discover relevant content based on their interests and emotions. Conventional systems simply recommend content based on keywords and tags, but they are unable to dynamically recommend content based on users' emotions and preferences. Given this background, there is a need for a system that allows users to discover relevant content through a more personalized experience.
[0756] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[0757] In this invention, the server includes means for extracting meta-information from collected text data, means for mapping relationships between content items based on the extracted meta-information, means for visually displaying the mapped relationships, means for recognizing a user's emotion, and means for dynamically recommending content items based on the recognized emotion, thereby enabling efficient discovery and recommendation of related content items based on the user's emotion and interests.
[0758] "Collected text data" refers to text information such as books, lyrics, audio data, and image data collected from a wide variety of content.
[0759] "Meta information" refers to information such as keywords, tags, and content names extracted from text data using a natural language processing engine.
[0760] "Relationships between content" refers to relationships that extract commonalities and associations between different content based on collected meta-information and show them as links.
[0761] "Visual display means" refers to means for displaying the relationships between analyzed contents in the form of a map, graph, list, etc., so that the user can visually understand the relationships between the analyzed contents.
[0762] "Means for recognizing user emotions" refers to a means for recognizing a user's emotional state by analyzing the user's facial expressions and voice using sensors such as a camera and microphone.
[0763] "Means for dynamic recommendation based on emotions" refers to a means by which the system selects and recommends relevant content in real time based on the user's recognized emotions.
[0764] This invention is a system that recognizes the user's emotions and recommends appropriate content based on those emotions. The system mainly involves three parties: a server, a terminal, and a user.
[0765] First, the server is equipped with a means for extracting meta-information from the collected text data. This meta-information includes keywords, tags, content names, etc. Specifically, the text data is analyzed using technologies such as natural language processing engines, morphological analysis, and entity recognition to extract information. For example, the contents of books are collected as text data, and then analyzed to extract keywords such as "mystery" and "detective."
[0766] Next, the server has a means for mapping the relationships between content items based on the extracted meta information. This involves visually showing the commonalities and connections between different pieces of content using the extracted keywords and tags. Possible visual display methods include maps, graphs, and lists. For example, multiple books or video works with the keyword "detective" could be linked together to show the user their relationships.
[0767] Furthermore, the device is equipped with a means of recognizing the user's emotions. This is achieved using the smartphone's camera and microphone. Specific software uses OpenCV facial recognition technology and EmotionRecognizer to recognize emotions from the user's facial expressions and voice. For example, if a user shows a surprised expression while reading a mystery novel, the device recognizes that emotion and sends it to the server.
[0768] The server has a means for receiving the user's emotional information and dynamically recommending content based on the recognition results. This allows content related to the user's emotions to be recommended in real time. For example, if the user expresses surprise, mystery novels or suspense movies that are even more surprising and intriguing will be recommended. This allows users to efficiently discover content that matches their emotions and preferences.
[0769] As a specific example, when a user searches for "mystery novels," multiple related content results are displayed. The emotion recognition engine analyzes the user's facial expressions and voice to indicate interest, and the server then recommends further related content based on the results. For example, if a user expresses the emotion "surprise," the server can further recommend "suspense movies," thereby personalizing the user experience. The following are examples of input prompts for the generative AI model:
[0770] "When a user searches for 'mystery novels,' get the 10 most relevant books from the Google Books API, analyze the content of each, and generate keywords and tags. If you detect the emotion 'surprise' from the user's facial expression, also recommend 'suspense movies.'"
[0771] As described above, the system of the present invention provides efficient information provision and enables users to encounter new content through each stage of collection, analysis, mapping, display, emotion recognition, and dynamic recommendation.
[0772] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0773] Step 1:
[0774] The user enters a keyword
[0775] The user enters keywords into the search bar on the device, and the device sends a search request to the server based on this input.
[0776] Step 2:
[0777] The server collects the content
[0778] The server retrieves content data related to keywords from an online database via API. This data includes a wide variety of content, such as books, movies, and music, and is sent to the server. For example, if a user searches for "mystery novels," related book information is retrieved from the Google Books API. The server then stores this retrieved data.
[0779] Step 3:
[0780] The server parses the text data
[0781] The server uses a natural language processing engine to analyze the collected content data. Specifically, it performs morphological analysis and entity recognition to extract keywords and tags. The results of this analysis are saved as meta-information. For example, keywords such as "detective" and "deduction" are extracted from the text data of "Sherlock Holmes."
[0782] Step 4:
[0783] The server maps the relationships between content
[0784] The server maps the relationships between content based on the extracted meta information, finding commonalities in keywords and tags and visualizing them as links. For example, it links works with the common tag "detective" and generates data for displaying them as a visual map.
[0785] Step 5:
[0786] The server visually displays the relationships
[0787] The device displays the relationship map generated by the server to the user, allowing the user to visually check related content. For example, if a user searches for "detective," related books and movies are displayed in map format. The user can click on this map to obtain more information.
[0788] Step 6:
[0789] The device recognizes the user's emotions
[0790] While a user is viewing content displayed on the device, the device uses a camera and microphone to analyze the user's facial expressions and voice. An emotion recognition engine is used to recognize the user's emotional state (surprise, interest, etc.). The results of this analysis are sent to the server.
[0791] Step 7:
[0792] The server recommends content based on emotions.
[0793] The server dynamically recommends content appropriate to the user's emotional state based on the results of the emotion recognition it receives. For example, if the user expresses surprise, the server generates data to recommend similarly interesting suspense movies or mystery novels and sends it to the device.
[0794] Step 8:
[0795] The device displays recommended content to the user
[0796] The device receives recommended content from the server and displays it to the user. The user can browse this recommended content and discover new content that piques their interest. For example, by expressing the emotion "surprise" while viewing "Sherlock Holmes," the device will recommend "suspense movies" and display them to the user.
[0797] These are the specific processing steps and their flow. Through these processes, personalized content recommendations based on the user's emotions and interests are realized.
[0798] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0799] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0800] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.
[0801] [Third embodiment]
[0802] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.
[0803] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.
[0804] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0805] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.
[0806] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0807] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0808] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0809] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0810] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0811] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0812] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0813] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."
[0814] This invention is a system that uses AI and natural language processing technology to extract and visualize the relationships between a wide variety of content, helping users discover new content that they are interested in. The basic process for implementing this system is shown below.
[0815] Content Collection
[0816] The server collects a variety of content, such as books, lyrics, audio data, image data, etc. At this stage, data may be retrieved from online databases via APIs, or files may be uploaded directly to the server.
[0817] Example: A server calls the Google Books API to retrieve book data related to a specific keyword. If a user wants to collect data related to "mystery novels," the server uses the Google Books API to retrieve book information related to "mystery novels." This information includes the book's title, author, publication year, summary of the content, etc.
[0818] Text analytics
[0819] The server sends the collected text data to a natural language processing engine to extract meta-information (keywords, tags, content names, etc.). This process uses technologies such as morphological analysis, entity recognition, and relationship extraction.
[0820] Example: The server analyzes data from the novel "Sherlock Holmes" and extracts character names such as "Sherlock Holmes" and "John Watson" and keywords such as "London" and "detective."
[0821] Keyword tag generation
[0822] Based on the extracted meta information, the server generates keywords and tags related to each piece of content, including important people, places, events, and themes.
[0823] Example: The server assigns tags such as "detective," "mystery," and "deduction" to "Sherlock Holmes" novels.
[0824] Relationship Mapping
[0825] The server maps the relationships between pieces of content based on keywords and tags, finding commonalities between pieces of content and creating a data structure that visually displays them as links.
[0826] Example: The server detects that the novels "Sherlock Holmes" and "Hercule Poirot" share the common keyword "detective" and links them together.
[0827] Database integration and updates
[0828] The server works in conjunction with other databases to regularly update and supplement the content information, ensuring that the most up-to-date and detailed information is always available.
[0829] Example: The server works with the Open Library API to periodically retrieve and update data on newly released "mystery novels."
[0830] User interface provided
[0831] The server provides a user interface that includes a relationship map and keyword search functionality, allowing users to enter keywords of interest and visually display related content.
[0832] Example: When a user searches for "detective" on their device, the server displays a visual map of related content, such as "Sherlock Holmes" and "Hercule Poirot."
[0833] User Interactions
[0834] Users can click on the displayed relationship map to view more detailed information, and can also enter new search keywords to find related content.
[0835] Example: When a user clicks on "Sherlock Holmes," the server displays detailed information and reviews about the work, as well as information about related TV shows and movies.
[0836] As described above, the system of the present invention provides efficient information provision and encounters with new content through each stage of collection, analysis, mapping, display, and update.
[0837] The processing flow will be explained below.
[0838] Step 1: Gather content
[0839] The server collects a variety of content such as books, lyrics, audio data, and image data.
[0840] Specifically, it retrieves data from online databases via API and also receives files uploaded by users.
[0841] Example: The server calls the Google Books API to retrieve metadata for books related to "fantasy novels."
[0842] Step 2: Registering text data storage
[0843] The server stores the collected text data in storage.
[0844] The data to be saved includes books, lyrics, text converted from audio data, and text extracted from images using OCR.
[0845] Example: A server stores the book data for "Harry Potter and the Philosopher's Stone" in a database.
[0846] Step 3: Text analysis
[0847] The server sends the stored text data to a natural language processing engine to extract meta-information.
[0848] It uses morphological analysis, entity recognition, and relationship extraction to identify characters, places, important events, etc.
[0849] Example: The server uses a natural language processing engine to extract character names such as "Sherlock Holmes" and "John Watson" from the text "Sherlock Holmes."
[0850] Step 4: Generate keywords and tags
[0851] The server generates keywords and tags related to each piece of content based on the extracted meta information.
[0852] This includes important people, places, events, themes, etc.
[0853] Example: The server assigns tags such as "detective," "mystery," and "deduction" to "Sherlock Holmes."
[0854] Step 5: Relationship mapping
[0855] The server maps the relationships between content based on the generated keywords and tags.
[0856] Create a data structure that finds commonalities between pieces of content and visually displays them as links.
[0857] Example: The server discovers that "Sherlock Holmes" and "Hercule Poirot" share the tag "detective" and links them together.
[0858] Step 6: Database integration and updates
[0859] The server cooperates with other databases to periodically retrieve and update content information.
[0860] This ensures that the latest information is always provided.
[0861] Example: The server works with the Open Library API to periodically retrieve and update data on newly released mystery novels.
[0862] Step 7: Provide the user interface
[0863] The server provides a user interface that includes a relationship map and keyword search functionality.
[0864] Users can enter keywords that interest them and relevant content will be displayed visually.
[0865] For example, if a user searches for "detective" on their device, the server will display related content such as "Sherlock Holmes" or "Hercule Poirot."
[0866] Step 8: User Interaction
[0867] Users can click on the displayed relationship map to view more detailed information.
[0868] Additionally, users can enter new search keywords to find related content.
[0869] For example, if a user clicks on "Sherlock Holmes," the server will display detailed information and reviews about the work, as well as information about related TV shows and movies.
[0870] Example 1
[0871] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0872] Conventional content management systems have difficulty efficiently extracting and visualizing the relationships between different types of content. Furthermore, the means by which users can discover new related content are limited, which does not sufficiently improve the user experience. Furthermore, there is insufficient integration with other databases to continuously acquire and update the latest information, making it difficult to ensure the freshness and diversity of content.
[0873] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0874] In this invention, the server includes a means for collecting information, a means for preprocessing the collected text data and extracting meta information, a means for generating keywords and tags based on the extracted meta information, a means for mapping relationships between content using the generated keywords and tags, and a means for visually displaying the mapped relationships. This enables efficient extraction and visualization of relationships between content of different formats. Furthermore, by including a means for acquiring and updating content data from domestic and international databases in cooperation with the server, it becomes possible to continuously acquire and provide the latest and most diverse information to users. Furthermore, by including a means for finding commonalities between mapped content and displaying them as links, users can intuitively discover new related content.
[0875] "Means for collecting information" refers to the means for collecting various content such as books, lyrics, audio data, and image data on a server through methods such as online databases and file uploads.
[0876] "Preprocessing" refers to the process of removing unnecessary characters and duplication from collected content data, formatting the text data, and preparing it for passing to a natural language processing engine.
[0877] "Meta information" is summary information of data such as keywords, tags, and content names extracted from collected text data, and is important information for understanding the meaning and substance of the content.
[0878] The "means for generating keywords and tags" refers to a means for automatically generating keywords and tags related to each piece of content based on meta information obtained from the natural language processing engine.
[0879] A "means for mapping relationships between content" is a means for finding commonalities and relationships between multiple pieces of content based on generated keywords and tags, and visually displaying them as links.
[0880] A "visual display means" is a means for displaying the relationships between the mapped content in a graphical format that allows users to intuitively understand the relationships.
[0881] "Means for obtaining and updating content data from domestic and international databases in cooperation with the server" refers to means for the server to periodically communicate with external databases to obtain new content data and update existing data.
[0882] "Means for finding commonalities and displaying them as links" refers to means for finding common keywords and tags between multiple pieces of content and visually displaying the relationships between them as links.
[0883] This invention is a system that efficiently extracts and visualizes the relationships between different types of content, allowing users to discover new related content. This system is run by a server, a terminal, and user operations.
[0884] Hardware and software used
[0885] This system is composed of the following hardware and software components:
[0886] 1. Hardware
[0887] Server: Collects, processes, and stores data.
[0888] Terminal: A device (e.g., PC, smartphone, tablet) that a user uses to access and operate the system.
[0889] 2. Software
[0890] API: Application Program Interface for connecting with external services for data collection.
[0891] Natural language processing engine: An engine for analyzing text data and extracting meta-information (e.g., morphological analysis engine, entity recognition engine).
[0892] Database: A database management system for storing and managing the collected and generated data.
[0893] Visualization libraries: Libraries for visually displaying relationships between content (e.g., D3.js).
[0894] Specific processing content of the program
[0895] First, the server collects various content such as books, lyrics, audio data, image data, etc. through APIs and file upload functions. For example, if a user wants to collect data related to "mystery novels," the server calls the API of an online database and retrieves book information as a result.
[0896] The server then preprocesses the collected text data and sends it to a natural language processing engine to extract meta-information, using a morphological analysis engine and entity recognition engine to extract important keywords and tags from the text.
[0897] The server then automatically generates related keywords and tags based on the extracted meta information and maps the relationships between content items. For example, if the novels of "Sherlock Holmes" and "Hercule Poirot" share the keyword "detective," these pieces of content are linked.
[0898] The server also periodically connects to external databases (e.g., Open Library API) to retrieve new content data and update the database, ensuring that users always have access to the latest information.
[0899] When a user accesses the system using a terminal and searches for a specific keyword, the server visually displays the data mapping the relationships. For example, if a user searches for "detective," related content will be displayed as a visual map.
[0900] Additionally, when a user clicks on a particular piece of content on the relationship map, the server provides more information about it. For example, if a user clicks on "Sherlock Holmes," the server displays more information about the work, related reviews, and other related content.
[0901] Prompt Sentence Examples
[0902] "Generate a program to collect, analyze, and display data to map relationships in a detective novel."
[0903] "Please explain in detail the process of the system that extracts relationships between content based on meta information and displays them visually."
[0904] The system is designed to help users discover new and relevant content and improve the overall user experience.
[0905] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0906] Step 1:
[0907] Content Collection
[0908] The server collects various content, such as books, lyrics, audio data, and image data, through API calls or file uploads. It receives keywords and files specified by the user as input. The server saves the collected data in its internal storage. It also uses APIs to retrieve data related to specific keywords from external databases. For example, if a user enters "mystery novels," the server calls the Google Books API and retrieves book information as a result. The output is the collected raw data.
[0909] Step 2:
[0910] Text Extraction and Preprocessing
[0911] The server extracts text from the collected content data and performs preprocessing. It receives the collected raw data as input. Specifically, it performs formatting processing by removing unnecessary characters and duplication from the text file. It then prepares the data for passing to the natural language processing engine. The output is preprocessed text data.
[0912] Step 3:
[0913] Extracting Meta Information
[0914] The server sends the preprocessed text data to a natural language processing engine to extract meta-information. The server receives the preprocessed text data as input. Specifically, it analyzes the text using a morphological analysis engine (e.g., Mecab) to extract meta-information such as keywords, tags, and content names. The output is the extracted meta-information.
[0915] Step 4:
[0916] Keyword tag generation
[0917] The server generates keywords and tags appropriate for each piece of content based on the extracted meta information. It receives meta information as input. Specifically, it extracts important people, places, events, themes, etc., and generates keywords and tags based on them. For example, it generates tags such as "detective," "mystery," and "deduction" from the text "Sherlock Holmes." The output is the generated keywords and tags.
[0918] Step 5:
[0919] Relationship mapping
[0920] The server maps the relationships between different content based on the generated keywords and tags. It receives the generated keywords and tags as input. Specifically, it links content that has common keywords or tags together to create a relational data structure. For example, if the texts "Sherlock Holmes" and "Hercule Poirot" share the tag "detective," they are linked together. The output is the mapped relationship data.
[0921] Step 6:
[0922] Visually mapping your data
[0923] The server performs visualization processing to visually display the relationship data. It receives the mapped relationship data as input. Specifically, it uses a visualization library such as D3.js to generate a relationship map in HTML format and sends it to the user's device. The output is a visual relationship map.
[0924] Step 7:
[0925] Database integration and updates
[0926] The server interacts with an external database to retrieve new content data and update existing data. It receives new data retrieved from the external database's API as input. Specifically, it periodically sends API requests to add new book and content data to the database and update existing data. The output is an updated database.
[0927] Step 8:
[0928] Providing a user interface
[0929] The server provides an interface that allows users to manipulate the relationship map and search results. It receives search keywords and click operations from the user as input. Specific operations include searching for related content based on the user's input and displaying it on the device as a visual map. The output is the visual map and search results displayed on the user interface.
[0930] Step 9:
[0931] Viewing detailed information
[0932] When a user clicks on a specific piece of content on the relationship map, the server displays its detailed information. It receives the user's click as input and retrieves detailed information about the related content from the database. Specifically, it provides the user with detailed information about the content, reviews, and other related content. The output is an interface displaying the detailed information.
[0933] (Application example 1)
[0934] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0935] In recent years, a wide variety of content has become available on the Internet, but it is not easy for users to discover new content that interests them. In particular, there is a need to efficiently search for content that matches users' interests from a vast amount of content data and show its relevance. Conventional methods have difficulty visually displaying the relationships between content, and there is a lack of systems that guide users to new content that interests them. To solve this problem, it is necessary to concretely visualize the abstract relationships between content and provide an interface that users can intuitively understand.
[0936] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0937] In this invention, the server includes means for extracting meta-information from collected text data, means for mapping relationships between content items based on the extracted meta-information, means for visually displaying the mapped relationships, means for vectorizing the collected text data and calculating the similarity between the vectors, and means for generating a network graph based on the similarity. This allows users to easily find new content items related to their interests from among a vast amount of content items and intuitively understand their relationships.
[0938] "Collected text data" refers to text information in a variety of formats, such as books, audio data, and image data, obtained from the Internet or other databases.
[0939] "Meta information" refers to accompanying information such as keywords, tags, and content names extracted from collected text data.
[0940] "Relationships between contents" is information that indicates commonalities and interrelationships between different contents.
[0941] "Mapping means" refers to a technical method for visually showing the relationships between content items based on extracted meta-information.
[0942] "Visual display means" refers to a method of presenting the relationships between content to users in a visual format such as a graph or map.
[0943] "Vectorization" means converting collected text data into numerical vectors, making it mathematically processable.
[0944] The "means for calculating similarity" is a method for calculating the similarity between vectorized text data and evaluating the numerical relationship.
[0945] A "network graph" is a structure that visually represents the relationships between content using nodes and edges.
[0946] This invention is a system that extracts relationships between various contents and supports users in discovering new contents. The system consists of a server, a terminal, and a user interface.
[0947] Content Collection
[0948] The server collects various content, such as books, audio data, and image data, from the Internet and other databases. Data is periodically retrieved from online databases via APIs and stored on the server as text data. For example, data can be collected using an information search API or a book information API.
[0949] Text Analysis
[0950] The server sends the collected text data to a natural language processing engine to extract meta-information (keywords, tags, content names, etc.). This process utilizes technologies such as morphological analysis, entity recognition, and relationship extraction. Specifically, it uses natural language processing toolkits and entity recognition software.
[0951] Meta information generation
[0952] The server automatically generates keywords and tags related to each piece of content based on the extracted meta information, including important people, places, events, themes, etc. For example, tags are generated based on the characters and themes of a book.
[0953] Relationship mapping
[0954] The server maps the relationships between content based on the generated keywords and tags. At this stage, the similarity between vectorized data is calculated and the numerical relationships are evaluated. Specifically, tools such as TfidfVectorizer and Cosine Similarity are used for vectorization and similarity calculation.
[0955] Visualizing Relationships
[0956] The server generates a network graph to visually display the results of the relationship mapping. The generated graph is displayed in a visual format that allows users to intuitively understand it. Data visualization tools such as NetworkX and matplotlib are used to generate the network graph.
[0957] User interface provided
[0958] The server provides a user interface with a relationship map and keyword search functionality. Users can enter keywords of interest and visually display related content. The user interface is web-based and uses web frameworks such as Flask.
[0959] User Interactions
[0960] Users can click on the visually displayed relationship map to view more detailed information, and can also enter new search keywords to discover and enjoy new related content.
[0961] Examples of concrete examples and prompts
[0962] For example, if a user searches for "suspense movies," you can visualize related movies, TV shows, and even related podcasts and books. Visually display movies, TV shows, podcasts, and books related to "suspense movies."
[0963] This system enables users to efficiently discover new content that interests them and intuitively understand its relationships.
[0964] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0965] Step 1:
[0966] The server collects content data from the Internet and other databases. Specifically, it uses APIs (such as information search APIs and book information APIs) to obtain a variety of content data, including books, audio data, and image data, and stores it on the server. The input required is a search query containing keywords that interest the user, and the output is a list of collected text data.
[0967] Step 2:
[0968] The server sends the collected text data to a natural language processing engine to extract meta-information. This process uses techniques such as morphological analysis, entity recognition, and relationship extraction. Specifically, a natural language processing toolkit is used to analyze the text and extract keywords, tags, content names, etc. The input is the text data obtained in step 1, and the output is the extracted meta-information.
[0969] Step 3:
[0970] The server generates keywords and tags related to each piece of content based on the extracted meta information. Tags such as important people, places, events, and themes are automatically added. The input is the meta information obtained in step 2, and the output is a list of generated keywords and tags.
[0971] Step 4:
[0972] The server maps the relationships between content items based on the generated keywords and tags. It evaluates the similarity between content items by vectorizing the collected text data (using TfidfVectorizer) and calculating the similarity between each vector (using Cosine Similarity). The input is the keywords and tags obtained in Step 3 and the text data, and the output is a similarity score.
[0973] Step 5:
[0974] The server generates a network graph based on the similarity. It uses data visualization tools such as NetworkX and matplotlib to visualize the relationships between content. The input is the similarity scores obtained in step 4, and the output is the visualized network graph.
[0975] Step 6:
[0976] The server provides a user-facing interface. Users can enter keywords of interest and relevant content is displayed visually. A web framework such as Flask is used to build the web-based interface. The input is the keywords entered by the user, and the output is a visually displayed relationship map.
[0977] Step 7:
[0978] Users can click on the visually displayed relationship map to view more information, and can also enter new search keywords to discover new related content. The input is the user's actions on the relationship map, and the output is the discovery of more information and new content.
[0979] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0980] This invention is a system that uses AI and natural language processing technology to extract and visualize the relationships between a wide variety of content, and by combining this with an emotion engine that recognizes the user's emotions, helps users discover new content that they are interested in. The basic process for implementing this system is shown below.
[0981] Content Collection
[0982] The server collects various content such as books, lyrics, audio data, image data, etc. At this stage, it retrieves data from online databases via APIs and may also receive files uploaded by users.
[0983] Example: A server uses the Google Books API to retrieve book data related to a specific keyword. If a user wants to collect data related to "mystery novels," the server retrieves book information related to "mystery novels" via the Google Books API. This information includes the book title, author, publication year, and summary of the contents.
[0984] Text analytics
[0985] The server uses a natural language processing engine to extract meta-information (keywords, tags, content names, etc.) from the collected text data. This process involves morphological analysis, entity recognition, and relationship extraction.
[0986] Example: The server analyzes text data from "Sherlock Holmes" and extracts character names such as "Sherlock Holmes" and "John Watson" and keywords such as "London" and "detective."
[0987] Keyword tag generation
[0988] The server generates keywords and tags related to each piece of content based on the extracted meta information.
[0989] Example: The server assigns tags such as "detective," "mystery," and "deduction" to "Sherlock Holmes" novels.
[0990] Relationship Mapping
[0991] The server maps the relationships between content based on the generated keywords and tags, finding commonalities between the content and visually displaying them as links.
[0992] Example: The server discovers that "Sherlock Holmes" and "Hercule Poirot" share the tag "detective" and links them together.
[0993] Database integration and updates
[0994] The server periodically retrieves and updates content information in cooperation with other databases, ensuring that the latest information is always provided.
[0995] Example: The server works with the Open Library API to periodically retrieve and update data on newly released "mystery novels."
[0996] Incorporating an emotion engine
[0997] The server analyzes the user's feelings about the content using an emotion engine that recognizes the user's emotions.
[0998] Example: When a user searches for "mystery novels" on their device, the emotion engine analyzes the user's facial expressions and voice to determine their emotions. For example, if the user shows expressions of surprise or interest, relevant content will be recommended based on that.
[0999] Emotion-based content recommendation
[1000] The server dynamically updates and recommends content that reflects the user's emotional state based on the analysis results from the emotion engine.
[1001] Example: If a user expresses the emotion of "surprise" while browsing "Sherlock Holmes," the server will recommend "mystery novels" or "suspense movies" that are likely to induce similar feelings of "surprise."
[1002] User interface provided
[1003] The server provides a user interface that includes a relationship map and keyword search functionality, allowing users to enter keywords of interest and visually display related content.
[1004] Example: When a user searches for "detective" on their device, the server displays a visual map of related content, such as "Sherlock Holmes" and "Hercule Poirot."
[1005] Users can click on the displayed relationship map to view detailed information about their interactions, or enter new search keywords to find related content.
[1006] Example: When a user clicks on "Sherlock Holmes," the server displays detailed information and reviews about the work, as well as information about related TV shows and movies.
[1007] As described above, the system of the present invention provides efficient information provision and encounters with new content through each stage of collection, analysis, mapping, display, update, and emotion recognition.
[1008] The processing flow will be explained below.
[1009] Step 1: Gather content
[1010] The server collects a variety of content such as books, lyrics, audio data, and image data.
[1011] Specifically, it retrieves data from an online database via API and stores files uploaded by users in storage.
[1012] Example: A server uses the Google Books API to retrieve metadata for books related to "fantasy novels."
[1013] Step 2: Registering text data storage
[1014] The server stores the collected text data in a database.
[1015] The data stored includes books, lyrics, text converted from audio data, and text extracted from images.
[1016] Example: A server stores the book data for "Harry Potter and the Philosopher's Stone" in a database.
[1017] Step 3: Text analysis
[1018] The server sends the stored text data to a natural language processing engine to extract meta-information.
[1019] It uses morphological analysis, entity recognition, and relationship extraction to identify characters, places, important events, etc.
[1020] Example: The server analyzes the text data of "Sherlock Holmes" and extracts character names such as "Sherlock Holmes" and "John Watson" as well as keywords such as "London" and "detective."
[1021] Step 4: Generate keywords and tags
[1022] The server generates keywords and tags related to each piece of content based on the extracted meta information.
[1023] This includes important people, places, events, themes, etc.
[1024] Example: The server tags "detective," "mystery," and "deduction" to "Sherlock Holmes" novels.
[1025] Step 5: Relationship mapping
[1026] The server maps the relationships between content based on the generated keywords and tags.
[1027] Create a data structure that finds commonalities between pieces of content and visually displays them as links.
[1028] Example: The server discovers that "Sherlock Holmes" and "Hercule Poirot" share the tag "detective" and links them together.
[1029] Step 6: Database integration and updates
[1030] The server cooperates with other databases to periodically retrieve and update content information.
[1031] This ensures that the latest information is always provided.
[1032] Example: The server works with the Open Library API to periodically retrieve and update data on newly released mystery novels.
[1033] Step 7: Incorporating the Emotion Engine
[1034] The server analyzes the user's feelings about the content using an emotion engine that recognizes the user's emotions.
[1035] The emotion engine recognizes the user's emotions from data such as text, voice, and facial expressions.
[1036] Example: When a user searches for "mystery novels" on their device, the emotion engine analyzes the user's emotions from their facial expressions and voice.
[1037] Step 8: Emotion-based content recommendation
[1038] The server dynamically updates and recommends content that reflects the user's emotional state based on the analysis results from the emotion engine.
[1039] For example, if a user expresses the emotion of "surprise" while browsing "Sherlock Holmes," the server will recommend "mystery novels" or "suspense movies" that are likely to induce similar feelings of "surprise."
[1040] Step 9: Provide the user interface
[1041] The server provides a user interface that includes a relationship map and keyword search functionality.
[1042] Users can enter keywords that interest them and relevant content will be displayed visually.
[1043] For example, if a user searches for "detective" on their device, the server will display a visual map of related content such as "Sherlock Holmes" and "Hercule Poirot."
[1044] Step 10: User Interaction
[1045] Users can click on the displayed relationship map to view more detailed information.
[1046] Additionally, users can enter new search keywords to find related content.
[1047] For example, if a user clicks on "Sherlock Holmes," the server will display detailed information and reviews about the work, as well as information about related TV shows and movies.
[1048] Example 2
[1049] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1050] Conventional information gathering and recommendation systems have struggled to accurately recommend content that users are truly interested in. Furthermore, there are limited ways to visually display the relationships between vast amounts of content in a way that is easy for users to understand. Furthermore, there is a lack of a dynamic content recommendation function based on user emotions, so there is a need for an improved user experience.
[1051] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[1052] In this invention, the server includes means for collecting text data, means for extracting meta information from the collected text data, means for mapping relationships between content items based on the extracted meta information, means for visually displaying the mapped relationships, means for analyzing user emotions, and means for recommending content items based on the user emotions. This allows for accurate recommendation of content items that interest the user, facilitating understanding of the content, and enabling dynamic content recommendation based on emotions.
[1053] "Text data" refers to the content of sentences, books, articles, documents, etc. stored in digital format.
[1054] "Meta information" refers to additional information about content, such as keywords, tags, and names.
[1055] "Relationships between content" refers to the connections based on the characteristics and tags shared between multiple pieces of content.
[1056] "Relationship mapping" refers to a data structure or graph that visually represents commonalities and connections between content.
[1057] "Visual display means" refers to a method of displaying information in a way that is easy for users to understand, such as in the form of graphs or maps.
[1058] "Analysis of user emotions" refers to the means of detecting and recognizing a user's emotional state based on their facial expressions, voice, and behavior.
[1059] "Means for recommending content" refers to methods for presenting appropriate content to users based on analysis results and relationship mapping.
[1060] A "database" refers to a systematic storage system that facilitates the collection, management, and retrieval of data.
[1061] "Domestic and international databases" refers to databases that exist and are accessible domestically and internationally.
[1062] "Commonalities" refer to characteristics or features shared by multiple pieces of content.
[1063] "Link" refers to a connection used to indicate a relationship between pieces of content.
[1064] This invention is a system that uses AI and natural language processing technology to extract and visualize the relationships between a wide variety of content, and combines this with an emotion engine that recognizes the user's emotions to help users discover new content that interests them. The following hardware and software are used to implement this invention.
[1065] Hardware:
[1066] Server: This is the central hardware for processing data collection, analysis, mapping, display, updates, and emotion recognition.
[1067] User terminal: A device used to view content and input emotions. Specifically, this includes PCs, smartphones, tablets, etc.
[1068] software:
[1069] API: An interface for retrieving data from online databases. Examples include the Google Books API and the Open Library API.
[1070] Natural language processing engine: A tool for extracting meta-information (keywords, tags, people's names, etc.) from text data.
[1071] Emotion engine: Software for analyzing the user's emotional state.
[1072] Data processing and calculation:
[1073] 1. Content Collection:
[1074] The server uses the Google Books API and Open Library API to collect various content, such as book data. When a user enters a specific keyword, related data is automatically retrieved. For example, to retrieve book data related to "mystery novels," the server calls the Google Books API and obtains information such as the title, author, publication year, and summary of the related book.
[1075] 2. Text Analysis:
[1076] The server runs the collected text data through a natural language processing engine, performing morphological analysis and entity recognition to extract meta-information. For example, it analyzes the text of the novel "Sherlock Holmes" and extracts keywords such as "Sherlock Holmes," "John Watson," and the place name "London."
[1077] 3. Keyword tag generation:
[1078] Based on the extracted meta information, the server generates keywords and tags related to each piece of content and stores them in a database. For example, a "Sherlock Holmes" novel might be tagged with "detective," "mystery," and "deduction."
[1079] 4. Relationship Mapping:
[1080] Based on the generated keywords and tags, the server analyzes and maps the relationships between multiple pieces of content. For example, "Sherlock Holmes" and "Hercule Poirot" both share the tag "detective," so they are linked together.
[1081] 5. Database integration and updates:
[1082] The server periodically connects to an external database to obtain the latest content information and update its internal database, thereby ensuring that users are always provided with the latest information.
[1083] 6. Incorporating an Emotion Engine:
[1084] Facial expression and voice data acquired from the user's device is input into the emotion engine, which analyzes the user's emotional state in real time. For example, if the user makes an expression showing surprise or interest, that emotional information is sent to the server, and related content is recommended.
[1085] 7. Emotion-based content recommendation:
[1086] Based on the analysis results obtained from the emotion engine, the server dynamically recommends content according to the user's emotional state. For example, if a user expresses surprise while browsing "Sherlock Holmes," the server will recommend mystery novels or suspense movies.
[1087] Specific examples
[1088] Example prompt sentence:
[1089] "If a user wants to collect data related to 'mystery novels,' the server retrieves book information related to 'mystery novels' via the Google Books API. This information includes the book's title, author, publication year, and summary of the contents."
[1090] In this way, the system of the present invention provides users with efficient information provision and enables them to encounter new content through each stage of collection, analysis, mapping, display, update, and emotion recognition.
[1091] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1092] Step 1:
[1093] The server collects the content data.
[1094] Input: User search keywords (e.g., "mystery novel")
[1095] The server uses the Google Books API or Open Library API to obtain book data related to the relevant keywords.
[1096] Output: A book dataset containing book titles, authors, publication years, and summary summaries.
[1097] Specific operation: The server generates a request to call the API and stores the obtained data in an internal database.
[1098] Step 2:
[1099] The server analyzes the collected text data using natural language processing.
[1100] Input: A collected book dataset
[1101] The server uses a natural language processing engine to extract meta-information from the collected text data (e.g., keywords, tags, names).
[1102] Output: Extracted meta information (e.g. "Sherlock Holmes", "Detective", "London")
[1103] How it works: The server uses morphological analyzers to break down text into words and entity recognition algorithms to identify important names and places.
[1104] Step 3:
[1105] The server generates keywords and tags based on the extracted meta information.
[1106] Input: Meta information (keywords, tags, name)
[1107] The server generates keywords and tags related to each piece of content and stores them in a database.
[1108] Output: Keywords and tags associated with the content (e.g., "detective," "mystery," "deduction")
[1109] Specific operation: The server executes a tag generation algorithm based on the extracted meta information and stores the generated tags in a database.
[1110] Step 4:
[1111] The server maps relationships between content based on keywords and tags.
[1112] Input: Keywords and tags associated with the content
[1113] The server analyzes the relationships between multiple pieces of content based on tags and keywords and generates links.
[1114] Output: Relationship map (e.g., a link connecting "Sherlock Holmes" and "Hercule Poirot")
[1115] Specific operation: The server uses a graph database to detect commonalities between tags and generate data to visualize them as links.
[1116] Step 5:
[1117] The server works in conjunction with an external database to periodically update the content information.
[1118] Input: Updates from an external database
[1119] The server obtains the latest content information using Open Library APIs and updates its internal database.
[1120] Output: Updated database (latest book information)
[1121] What it does: The server runs scheduled jobs to retrieve new data from external databases to keep the internal system up to date.
[1122] Step 6:
[1123] The server analyzes the user's emotions using an emotion engine.
[1124] Input: User's facial expressions and voice data
[1125] The server uses an emotion engine to analyze the user's facial expressions and voice to detect emotions.
[1126] Output: User's emotional state (e.g., "surprise," "interest")
[1127] Specific operation: Data collected through the camera and microphone of the user's device is sent to the emotion engine, and emotion analysis is performed in real time.
[1128] Step 7:
[1129] The server dynamically recommends content based on the analysis results of the emotion engine.
[1130] Input: User's emotional state from the emotion engine
[1131] The server runs an algorithm that recommends optimal content based on the user's emotional state.
[1132] Output: A list of recommended content (e.g., related "mystery novels" or "suspense movies")
[1133] Specific operation: The server searches for content that best matches the emotional state and sends a recommendation list to the user terminal.
[1134] Step 8:
[1135] The server builds a user interface that provides a relationship map and keyword search functionality.
[1136] Input: User search keywords and clicks
[1137] The server generates a relationship map to visually display related content.
[1138] Output: Visual relationship map and detailed information
[1139] Specific operation: When a user performs a search on their device, the server visually displays related content, allowing the user to access detailed information by clicking.
[1140] (Application example 2)
[1141] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1142] In today's information-overloaded environment, it is difficult for users to efficiently discover relevant content based on their interests and emotions. Conventional systems simply recommend content based on keywords and tags, but they are unable to dynamically recommend content based on users' emotions and preferences. Given this background, there is a need for a system that allows users to discover relevant content through a more personalized experience.
[1143] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[1144] In this invention, the server includes means for extracting meta-information from collected text data, means for mapping relationships between content items based on the extracted meta-information, means for visually displaying the mapped relationships, means for recognizing a user's emotion, and means for dynamically recommending content items based on the recognized emotion, thereby enabling efficient discovery and recommendation of related content items based on the user's emotion and interests.
[1145] "Collected text data" refers to text information such as books, lyrics, audio data, and image data collected from a wide variety of content.
[1146] "Meta information" refers to information such as keywords, tags, and content names extracted from text data using a natural language processing engine.
[1147] "Relationships between content" refers to relationships that extract commonalities and associations between different content based on collected meta-information and show them as links.
[1148] "Visual display means" refers to means for displaying the relationships between analyzed contents in the form of a map, graph, list, etc., so that the user can visually understand the relationships between the analyzed contents.
[1149] "Means for recognizing user emotions" refers to a means for recognizing a user's emotional state by analyzing the user's facial expressions and voice using sensors such as a camera and microphone.
[1150] "Means for dynamic recommendation based on emotions" refers to a means by which the system selects and recommends relevant content in real time based on the user's perceived emotions.
[1151] This invention is a system that recognizes the user's emotions and recommends appropriate content based on those emotions. The system mainly involves three parties: a server, a terminal, and a user.
[1152] First, the server is equipped with a means for extracting meta-information from the collected text data. This meta-information includes keywords, tags, content names, etc. Specifically, the text data is analyzed using technologies such as natural language processing engines, morphological analysis, and entity recognition to extract information. For example, the contents of books are collected as text data, and then analyzed to extract keywords such as "mystery" and "detective."
[1153] Next, the server has a means for mapping the relationships between content items based on the extracted meta information. This involves visually showing the commonalities and connections between different pieces of content using the extracted keywords and tags. Possible visual display methods include maps, graphs, and lists. For example, multiple books or video works with the keyword "detective" could be linked together to show the user their relationships.
[1154] Furthermore, the device is equipped with a means of recognizing the user's emotions. This is achieved using the smartphone's camera and microphone. Specific software uses OpenCV facial recognition technology and EmotionRecognizer to recognize emotions from the user's facial expressions and voice. For example, if a user shows a surprised expression while reading a mystery novel, the device recognizes that emotion and sends it to the server.
[1155] The server has a means for receiving the user's emotional information and dynamically recommending content based on the recognition results. This allows content related to the user's emotions to be recommended in real time. For example, if the user expresses surprise, mystery novels or suspense movies that are even more surprising and intriguing will be recommended. This allows users to efficiently discover content that matches their emotions and preferences.
[1156] As a specific example, when a user searches for "mystery novels," multiple related content results are displayed. The emotion recognition engine analyzes the user's facial expressions and voice to indicate interest, and the server then recommends further related content based on the results. For example, if a user expresses the emotion "surprise," the server can further recommend "suspense movies," thereby personalizing the user experience. The following are examples of input prompts for the generative AI model:
[1157] "When a user searches for 'mystery novels,' get the 10 most relevant books from the Google Books API, analyze the content of each, and generate keywords and tags. If you detect the emotion 'surprise' from the user's facial expression, also recommend 'suspense movies.'"
[1158] As described above, the system of the present invention provides efficient information provision and enables users to encounter new content through each stage of collection, analysis, mapping, display, emotion recognition, and dynamic recommendation.
[1159] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1160] Step 1:
[1161] The user enters a keyword
[1162] The user enters keywords into the search bar on the device, and the device sends a search request to the server based on this input.
[1163] Step 2:
[1164] The server collects the content
[1165] The server retrieves content data related to keywords from an online database via API. This data includes a wide variety of content, such as books, movies, and music, and is sent to the server. For example, if a user searches for "mystery novels," related book information is retrieved from the Google Books API. The server then stores this retrieved data.
[1166] Step 3:
[1167] The server parses the text data
[1168] The server uses a natural language processing engine to analyze the collected content data. Specifically, it performs morphological analysis and entity recognition to extract keywords and tags. The results of this analysis are saved as meta-information. For example, keywords such as "detective" and "deduction" are extracted from the text data of "Sherlock Holmes."
[1169] Step 4:
[1170] The server maps the relationships between content
[1171] The server maps the relationships between content based on the extracted meta information, finding commonalities in keywords and tags and visualizing them as links. For example, it links works with the common tag "detective" and generates data for displaying them as a visual map.
[1172] Step 5:
[1173] The server visually displays the relationships
[1174] The device displays the relationship map generated by the server to the user, allowing the user to visually check related content. For example, if a user searches for "detective," related books and movies are displayed in map format. The user can click on this map to obtain more information.
[1175] Step 6:
[1176] The device recognizes the user's emotions
[1177] While a user is viewing content displayed on the device, the device uses a camera and microphone to analyze the user's facial expressions and voice. An emotion recognition engine is used to recognize the user's emotional state (surprise, interest, etc.). The results of this analysis are sent to the server.
[1178] Step 7:
[1179] The server recommends content based on emotions.
[1180] The server dynamically recommends content appropriate to the user's emotional state based on the results of the emotion recognition it receives. For example, if the user expresses surprise, the server generates data to recommend similarly interesting suspense movies or mystery novels and sends it to the device.
[1181] Step 8:
[1182] The device displays recommended content to the user
[1183] The device receives recommended content from the server and displays it to the user. The user can browse this recommended content and discover new content that piques their interest. For example, by expressing the emotion "surprise" while browsing "Sherlock Holmes," the device will recommend "suspense movies" and display them to the user.
[1184] These are the specific processing steps and their flow. Through these processes, personalized content recommendations based on the user's emotions and interests are realized.
[1185] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1186] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1187] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.
[1188] [Fourth embodiment]
[1189] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.
[1190] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[1191] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1192] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.
[1193] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1194] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1195] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1196] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.
[1197] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1198] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1199] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1200] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1201] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1202] This invention is a system that uses AI and natural language processing technology to extract and visualize the relationships between a wide variety of content, helping users discover new content that they are interested in. The basic process for implementing this system is shown below.
[1203] Content Collection
[1204] The server collects a variety of content, such as books, lyrics, audio data, image data, etc. At this stage, data may be retrieved from online databases via APIs, or files may be uploaded directly to the server.
[1205] Example: A server calls the Google Books API to retrieve book data related to a specific keyword. If a user wants to collect data related to "mystery novels," the server uses the Google Books API to retrieve book information related to "mystery novels." This information includes the book's title, author, publication year, summary of the content, etc.
[1206] Text analytics
[1207] The server sends the collected text data to a natural language processing engine to extract meta-information (keywords, tags, content names, etc.). This process uses technologies such as morphological analysis, entity recognition, and relationship extraction.
[1208] Example: The server analyzes data from the novel "Sherlock Holmes" and extracts character names such as "Sherlock Holmes" and "John Watson" and keywords such as "London" and "detective."
[1209] Keyword tag generation
[1210] Based on the extracted meta information, the server generates keywords and tags related to each piece of content, including important people, places, events, and themes.
[1211] Example: The server assigns tags such as "detective," "mystery," and "deduction" to "Sherlock Holmes" novels.
[1212] Relationship Mapping
[1213] The server maps the relationships between pieces of content based on keywords and tags, finding commonalities between pieces of content and creating a data structure that visually displays them as links.
[1214] Example: The server detects that the novels "Sherlock Holmes" and "Hercule Poirot" share the common keyword "detective" and links them together.
[1215] Database integration and updates
[1216] The server works in conjunction with other databases to regularly update and supplement the content information, ensuring that the most up-to-date and detailed information is always available.
[1217] Example: The server works with the Open Library API to periodically retrieve and update data on newly released "mystery novels."
[1218] User interface provided
[1219] The server provides a user interface that includes a relationship map and keyword search functionality, allowing users to enter keywords of interest and visually display related content.
[1220] Example: When a user searches for "detective" on their device, the server displays a visual map of related content, such as "Sherlock Holmes" and "Hercule Poirot."
[1221] User Interactions
[1222] Users can click on the displayed relationship map to view more detailed information, and can also enter new search keywords to find related content.
[1223] Example: When a user clicks on "Sherlock Holmes," the server displays detailed information and reviews about the work, as well as information about related TV shows and movies.
[1224] As described above, the system of the present invention provides efficient information provision and encounters with new content through each stage of collection, analysis, mapping, display, and update.
[1225] The processing flow will be explained below.
[1226] Step 1: Gather content
[1227] The server collects a variety of content such as books, lyrics, audio data, and image data.
[1228] Specifically, it retrieves data from online databases via API and also receives files uploaded by users.
[1229] Example: The server calls the Google Books API to retrieve metadata for books related to "fantasy novels."
[1230] Step 2: Registering text data storage
[1231] The server stores the collected text data in storage.
[1232] The data to be saved includes books, lyrics, text converted from audio data, and text extracted from images using OCR.
[1233] Example: A server stores the book data for "Harry Potter and the Philosopher's Stone" in a database.
[1234] Step 3: Text analysis
[1235] The server sends the stored text data to a natural language processing engine to extract meta-information.
[1236] It uses morphological analysis, entity recognition, and relationship extraction to identify characters, places, important events, etc.
[1237] Example: The server uses a natural language processing engine to extract character names such as "Sherlock Holmes" and "John Watson" from the text "Sherlock Holmes."
[1238] Step 4: Generate keywords and tags
[1239] The server generates keywords and tags related to each piece of content based on the extracted meta information.
[1240] This includes important people, places, events, themes, etc.
[1241] Example: The server assigns tags such as "detective," "mystery," and "deduction" to "Sherlock Holmes."
[1242] Step 5: Relationship mapping
[1243] The server maps the relationships between content based on the generated keywords and tags.
[1244] Create a data structure that finds commonalities between pieces of content and visually displays them as links.
[1245] Example: The server discovers that "Sherlock Holmes" and "Hercule Poirot" share the tag "detective" and links them together.
[1246] Step 6: Database integration and updates
[1247] The server cooperates with other databases to periodically retrieve and update content information.
[1248] This ensures that the latest information is always provided.
[1249] Example: The server works with the Open Library API to periodically retrieve and update data on newly released mystery novels.
[1250] Step 7: Provide the user interface
[1251] The server provides a user interface that includes a relationship map and keyword search functionality.
[1252] Users can enter keywords that interest them and relevant content will be displayed visually.
[1253] For example, if a user searches for "detective" on their device, the server will display related content such as "Sherlock Holmes" or "Hercule Poirot."
[1254] Step 8: User Interaction
[1255] Users can click on the displayed relationship map to view more detailed information.
[1256] Additionally, users can enter new search keywords to find related content.
[1257] For example, if a user clicks on "Sherlock Holmes," the server will display detailed information and reviews about the work, as well as information about related TV shows and movies.
[1258] Example 1
[1259] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1260] Conventional content management systems have difficulty efficiently extracting and visualizing the relationships between different types of content. Furthermore, the means by which users can discover new related content are limited, which does not sufficiently improve the user experience. Furthermore, there is insufficient integration with other databases to continuously acquire and update the latest information, making it difficult to ensure the freshness and diversity of content.
[1261] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1262] In this invention, the server includes a means for collecting information, a means for preprocessing the collected text data and extracting meta information, a means for generating keywords and tags based on the extracted meta information, a means for mapping relationships between content using the generated keywords and tags, and a means for visually displaying the mapped relationships. This enables efficient extraction and visualization of relationships between content of different formats. Furthermore, by including a means for acquiring and updating content data from domestic and international databases in cooperation with the server, it becomes possible to continuously acquire and provide the latest and most diverse information to users. Furthermore, by including a means for finding commonalities between mapped content and displaying them as links, users can intuitively discover new related content.
[1263] "Means for collecting information" refers to the means for collecting various content such as books, lyrics, audio data, and image data on a server through methods such as online databases and file uploads.
[1264] "Preprocessing" refers to the process of removing unnecessary characters and duplication from the collected content data, formatting the text data, and preparing it for passing to a natural language processing engine.
[1265] "Meta information" is summary information of data such as keywords, tags, and content names extracted from collected text data, and is important information for understanding the meaning and substance of the content.
[1266] The "means for generating keywords and tags" refers to a means for automatically generating keywords and tags related to each piece of content based on meta information obtained from the natural language processing engine.
[1267] A "means for mapping relationships between content" is a means for finding commonalities and relationships between multiple pieces of content based on generated keywords and tags, and visually displaying them as links.
[1268] A "visual display means" is a means for displaying the relationships between the mapped content in a graphical format that allows users to intuitively understand the relationships.
[1269] "Means for obtaining and updating content data from domestic and international databases in cooperation with the server" refers to means for the server to periodically communicate with external databases to obtain new content data and update existing data.
[1270] "Means for finding commonalities and displaying them as links" refers to means for finding common keywords and tags between multiple pieces of content and visually displaying the relationships between them as links.
[1271] This invention is a system that efficiently extracts and visualizes the relationships between different types of content, allowing users to discover new related content. This system is run by a server, a terminal, and user operations.
[1272] Hardware and software used
[1273] This system is composed of the following hardware and software components:
[1274] 1. Hardware
[1275] Server: Collects, processes, and stores data.
[1276] Terminal: A device (e.g., PC, smartphone, tablet) that a user uses to access and operate the system.
[1277] 2. Software
[1278] API: Application Program Interface for connecting with external services for data collection.
[1279] Natural language processing engine: An engine for analyzing text data and extracting meta-information (e.g., morphological analysis engine, entity recognition engine).
[1280] Database: A database management system for storing and managing the collected and generated data.
[1281] Visualization libraries: Libraries for visually displaying relationships between content (e.g., D3.js).
[1282] Specific processing content of the program
[1283] First, the server collects various content such as books, lyrics, audio data, image data, etc. through APIs and file upload functions. For example, if a user wants to collect data related to "mystery novels," the server calls the API of an online database and retrieves book information as a result.
[1284] The server then preprocesses the collected text data and sends it to a natural language processing engine to extract meta-information, using a morphological analysis engine and entity recognition engine to extract important keywords and tags from the text.
[1285] The server then automatically generates related keywords and tags based on the extracted meta information and maps the relationships between content items. For example, if the novels of "Sherlock Holmes" and "Hercule Poirot" share the keyword "detective," these pieces of content are linked.
[1286] The server also periodically connects to external databases (e.g., Open Library API) to retrieve new content data and update the database, ensuring that users always have access to the latest information.
[1287] When a user accesses the system using a terminal and searches for a specific keyword, the server visually displays the data mapping the relationships. For example, if a user searches for "detective," related content will be displayed as a visual map.
[1288] Additionally, when a user clicks on a particular piece of content on the relationship map, the server provides more information about it. For example, if a user clicks on "Sherlock Holmes," the server displays more information about the work, related reviews, and other related content.
[1289] Prompt Sentence Examples
[1290] "Generate a program to collect, analyze, and display data to map relationships in a detective novel."
[1291] "Please explain in detail the process of the system that extracts relationships between content based on meta information and displays them visually."
[1292] The system is designed to help users discover new and relevant content and improve the overall user experience.
[1293] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1294] Step 1:
[1295] Content Collection
[1296] The server collects various content, such as books, lyrics, audio data, and image data, through API calls or file uploads. It receives keywords and files specified by the user as input. The server saves the collected data in its internal storage. It also uses APIs to retrieve data related to specific keywords from external databases. For example, if a user enters "mystery novels," the server calls the Google Books API and retrieves book information as a result. The output is the collected raw data.
[1297] Step 2:
[1298] Text Extraction and Preprocessing
[1299] The server extracts text from the collected content data and performs preprocessing. It receives the collected raw data as input. Specifically, it performs formatting processing by removing unnecessary characters and duplication from the text file. It then prepares the data for passing to the natural language processing engine. The output is preprocessed text data.
[1300] Step 3:
[1301] Extracting Meta Information
[1302] The server sends the preprocessed text data to a natural language processing engine to extract meta-information. The server receives the preprocessed text data as input. Specifically, it analyzes the text using a morphological analysis engine (e.g., Mecab) to extract meta-information such as keywords, tags, and content names. The output is the extracted meta-information.
[1303] Step 4:
[1304] Keyword tag generation
[1305] The server generates keywords and tags appropriate for each piece of content based on the extracted meta information. It receives meta information as input. Specifically, it extracts important people, places, events, themes, etc., and generates keywords and tags based on them. For example, it generates tags such as "detective," "mystery," and "deduction" from the text "Sherlock Holmes." The output is the generated keywords and tags.
[1306] Step 5:
[1307] Relationship mapping
[1308] The server maps the relationships between different content based on the generated keywords and tags. It receives the generated keywords and tags as input. Specifically, it links content that has common keywords or tags together to create a relational data structure. For example, if the texts "Sherlock Holmes" and "Hercule Poirot" share the tag "detective," they are linked together. The output is the mapped relationship data.
[1309] Step 6:
[1310] Visually mapping your data
[1311] The server performs visualization processing to visually display the relationship data. It receives the mapped relationship data as input. Specifically, it uses a visualization library such as D3.js to generate a relationship map in HTML format and sends it to the user's device. The output is a visual relationship map.
[1312] Step 7:
[1313] Database integration and updates
[1314] The server interacts with an external database to retrieve new content data and update existing data. It receives new data retrieved from the external database's API as input. Specifically, it periodically sends API requests to add new book and content data to the database and update existing data. The output is an updated database.
[1315] Step 8:
[1316] Providing a user interface
[1317] The server provides an interface that allows users to manipulate the relationship map and search results. It receives search keywords and click operations from the user as input. Specific operations include searching for related content based on the user's input and displaying it on the device as a visual map. The output is the visual map and search results displayed on the user interface.
[1318] Step 9:
[1319] Viewing detailed information
[1320] When a user clicks on a specific piece of content on the relationship map, the server displays its detailed information. It receives the user's click as input and retrieves detailed information about the related content from the database. Specifically, it provides the user with detailed information about the content, reviews, and other related content. The output is an interface displaying the detailed information.
[1321] (Application example 1)
[1322] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1323] In recent years, a wide variety of content has become available on the Internet, but it is not easy for users to discover new content that interests them. In particular, there is a need to efficiently search for content that matches users' interests from a vast amount of content data and show its relevance. Conventional methods have difficulty visually displaying the relationships between content, and there is a lack of systems that guide users to new content that interests them. To solve this problem, it is necessary to concretely visualize the abstract relationships between content and provide an interface that users can intuitively understand.
[1324] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1325] In this invention, the server includes means for extracting meta-information from collected text data, means for mapping relationships between content items based on the extracted meta-information, means for visually displaying the mapped relationships, means for vectorizing the collected text data and calculating the similarity between the vectors, and means for generating a network graph based on the similarity. This allows users to easily find new content items related to their interests from among a vast amount of content items and intuitively understand their relationships.
[1326] "Collected text data" refers to text information in a variety of formats, such as books, audio data, and image data, obtained from the Internet or other databases.
[1327] "Meta information" refers to accompanying information such as keywords, tags, and content names extracted from collected text data.
[1328] "Relationships between contents" is information that indicates commonalities and interrelationships between different contents.
[1329] "Mapping means" refers to a technical method for visually showing the relationships between content items based on extracted meta-information.
[1330] "Visual display means" refers to a method of presenting the relationships between content to users in a visual format such as a graph or map.
[1331] "Vectorization" means converting collected text data into numerical vectors, making it mathematically processable.
[1332] The "means for calculating similarity" is a method for calculating the similarity between vectorized text data and evaluating the numerical relationship.
[1333] A "network graph" is a structure that visually represents the relationships between content using nodes and edges.
[1334] This invention is a system that extracts relationships between various contents and supports users in discovering new contents. The system consists of a server, a terminal, and a user interface.
[1335] Content Collection
[1336] The server collects various content, such as books, audio data, and image data, from the Internet and other databases. Data is periodically retrieved from online databases via APIs and stored on the server as text data. For example, data can be collected using an information search API or a book information API.
[1337] Text Analysis
[1338] The server sends the collected text data to a natural language processing engine to extract meta-information (keywords, tags, content names, etc.). This process utilizes technologies such as morphological analysis, entity recognition, and relationship extraction. Specifically, it uses natural language processing toolkits and entity recognition software.
[1339] Meta information generation
[1340] The server automatically generates keywords and tags related to each piece of content based on the extracted meta information, including important people, places, events, themes, etc. For example, tags are generated based on the characters and themes of a book.
[1341] Relationship mapping
[1342] The server maps the relationships between content based on the generated keywords and tags. At this stage, the similarity between vectorized data is calculated and the numerical relationships are evaluated. Specifically, tools such as TfidfVectorizer and Cosine Similarity are used for vectorization and similarity calculation.
[1343] Visualizing Relationships
[1344] The server generates a network graph to visually display the results of the relationship mapping. The generated graph is displayed in a visual format that allows users to intuitively understand it. Data visualization tools such as NetworkX and matplotlib are used to generate the network graph.
[1345] User interface provided
[1346] The server provides a user interface with a relationship map and keyword search functionality. Users can enter keywords of interest and visually display related content. The user interface is web-based and uses web frameworks such as Flask.
[1347] User Interactions
[1348] Users can click on the visually displayed relationship map to view more detailed information, and can also enter new search keywords to discover and enjoy new related content.
[1349] Examples of concrete examples and prompts
[1350] For example, if a user searches for "suspense movies," you can visualize related movies, TV shows, and even related podcasts and books. Visually display movies, TV shows, podcasts, and books related to "suspense movies."
[1351] This system enables users to efficiently discover new content that interests them and intuitively understand its relationships.
[1352] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1353] Step 1:
[1354] The server collects content data from the Internet and other databases. Specifically, it uses APIs (such as information search APIs and book information APIs) to obtain a variety of content data, including books, audio data, and image data, and stores it on the server. The input required is a search query containing keywords that interest the user, and the output is a list of collected text data.
[1355] Step 2:
[1356] The server sends the collected text data to a natural language processing engine to extract meta-information. This process uses techniques such as morphological analysis, entity recognition, and relationship extraction. Specifically, a natural language processing toolkit is used to analyze the text and extract keywords, tags, content names, etc. The input is the text data obtained in step 1, and the output is the extracted meta-information.
[1357] Step 3:
[1358] The server generates keywords and tags related to each piece of content based on the extracted meta information. Tags such as important people, places, events, and themes are automatically added. The input is the meta information obtained in step 2, and the output is a list of generated keywords and tags.
[1359] Step 4:
[1360] The server maps the relationships between content items based on the generated keywords and tags. It evaluates the similarity between content items by vectorizing the collected text data (using TfidfVectorizer) and calculating the similarity between each vector (using Cosine Similarity). The input is the keywords and tags obtained in Step 3 and the text data, and the output is a similarity score.
[1361] Step 5:
[1362] The server generates a network graph based on the similarity. It uses data visualization tools such as NetworkX and matplotlib to visualize the relationships between content. The input is the similarity scores obtained in step 4, and the output is the visualized network graph.
[1363] Step 6:
[1364] The server provides a user-facing interface. Users can enter keywords of interest and relevant content is displayed visually. A web framework such as Flask is used to build the web-based interface. The input is the keywords entered by the user, and the output is a visually displayed relationship map.
[1365] Step 7:
[1366] Users can click on the visually displayed relationship map to view more information, and can also enter new search keywords to discover new related content. The input is the user's actions on the relationship map, and the output is the discovery of more information and new content.
[1367] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1368] This invention is a system that uses AI and natural language processing technology to extract and visualize the relationships between a wide variety of content, and by combining this with an emotion engine that recognizes the user's emotions, helps users discover new content that they are interested in. The basic process for implementing this system is shown below.
[1369] Content Collection
[1370] The server collects various content such as books, lyrics, audio data, image data, etc. At this stage, it retrieves data from online databases via APIs and may also receive files uploaded by users.
[1371] Example: A server uses the Google Books API to retrieve book data related to a specific keyword. If a user wants to collect data related to "mystery novels," the server retrieves book information related to "mystery novels" via the Google Books API. This information includes the book title, author, publication year, and summary of the contents.
[1372] Text analytics
[1373] The server uses a natural language processing engine to extract meta-information (keywords, tags, content names, etc.) from the collected text data. This process involves morphological analysis, entity recognition, and relationship extraction.
[1374] Example: The server analyzes text data from "Sherlock Holmes" and extracts character names such as "Sherlock Holmes" and "John Watson" and keywords such as "London" and "detective."
[1375] Keyword tag generation
[1376] The server generates keywords and tags related to each piece of content based on the extracted meta information.
[1377] Example: The server assigns tags such as "detective," "mystery," and "deduction" to "Sherlock Holmes" novels.
[1378] Relationship Mapping
[1379] The server maps the relationships between content based on the generated keywords and tags, finding commonalities between the content and visually displaying them as links.
[1380] Example: The server discovers that "Sherlock Holmes" and "Hercule Poirot" share the tag "detective" and links them together.
[1381] Database integration and updates
[1382] The server periodically retrieves and updates content information in cooperation with other databases, ensuring that the latest information is always provided.
[1383] Example: The server works with the Open Library API to periodically retrieve and update data on newly released "mystery novels."
[1384] Incorporating an emotion engine
[1385] The server analyzes the user's feelings about the content using an emotion engine that recognizes the user's emotions.
[1386] Example: When a user searches for "mystery novels" on their device, the emotion engine analyzes the user's facial expressions and voice to determine their emotions. For example, if the user shows expressions of surprise or interest, relevant content will be recommended based on that.
[1387] Emotion-based content recommendation
[1388] The server dynamically updates and recommends content that reflects the user's emotional state based on the analysis results from the emotion engine.
[1389] Example: If a user expresses the emotion of "surprise" while browsing "Sherlock Holmes," the server will recommend "mystery novels" or "suspense movies" that are likely to induce similar feelings of "surprise."
[1390] User interface provided
[1391] The server provides a user interface that includes a relationship map and keyword search functionality, allowing users to enter keywords of interest and visually display related content.
[1392] Example: When a user searches for "detective" on their device, the server displays a visual map of related content, such as "Sherlock Holmes" and "Hercule Poirot."
[1393] Users can click on the displayed relationship map to view detailed information about their interactions, or enter new search keywords to find related content.
[1394] Example: When a user clicks on "Sherlock Holmes," the server displays detailed information and reviews about the work, as well as information about related TV shows and movies.
[1395] As described above, the system of the present invention provides efficient information provision and encounters with new content through each stage of collection, analysis, mapping, display, update, and emotion recognition.
[1396] The processing flow will be explained below.
[1397] Step 1: Gather content
[1398] The server collects a variety of content such as books, lyrics, audio data, and image data.
[1399] Specifically, it retrieves data from an online database via API and stores files uploaded by users in storage.
[1400] Example: A server uses the Google Books API to retrieve metadata for books related to "fantasy novels."
[1401] Step 2: Registering text data storage
[1402] The server stores the collected text data in a database.
[1403] The data stored includes books, lyrics, text converted from audio data, and text extracted from images.
[1404] Example: A server stores the book data for "Harry Potter and the Philosopher's Stone" in a database.
[1405] Step 3: Text analysis
[1406] The server sends the stored text data to a natural language processing engine to extract meta-information.
[1407] It uses morphological analysis, entity recognition, and relationship extraction to identify characters, places, important events, etc.
[1408] Example: The server analyzes the text data of "Sherlock Holmes" and extracts character names such as "Sherlock Holmes" and "John Watson" as well as keywords such as "London" and "detective."
[1409] Step 4: Generate keywords and tags
[1410] The server generates keywords and tags related to each piece of content based on the extracted meta information.
[1411] This includes important people, places, events, themes, etc.
[1412] Example: The server tags "detective," "mystery," and "deduction" to "Sherlock Holmes" novels.
[1413] Step 5: Relationship mapping
[1414] The server maps the relationships between content based on the generated keywords and tags.
[1415] Create a data structure that finds commonalities between pieces of content and visually displays them as links.
[1416] Example: The server discovers that "Sherlock Holmes" and "Hercule Poirot" share the tag "detective" and links them together.
[1417] Step 6: Database integration and updates
[1418] The server cooperates with other databases to periodically retrieve and update content information.
[1419] This ensures that the latest information is always provided.
[1420] Example: The server works with the Open Library API to periodically retrieve and update data on newly released mystery novels.
[1421] Step 7: Incorporating the Emotion Engine
[1422] The server analyzes the user's feelings about the content using an emotion engine that recognizes the user's emotions.
[1423] The emotion engine recognizes the user's emotions from data such as text, voice, and facial expressions.
[1424] Example: When a user searches for "mystery novels" on their device, the emotion engine analyzes the user's emotions from their facial expressions and voice.
[1425] Step 8: Emotion-based content recommendation
[1426] The server dynamically updates and recommends content that reflects the user's emotional state based on the analysis results from the emotion engine.
[1427] For example, if a user expresses the emotion of "surprise" while browsing "Sherlock Holmes," the server will recommend "mystery novels" or "suspense movies" that are likely to induce similar feelings of "surprise."
[1428] Step 9: Provide the user interface
[1429] The server provides a user interface that includes a relationship map and keyword search functionality.
[1430] Users can enter keywords that interest them and relevant content will be displayed visually.
[1431] For example, if a user searches for "detective" on their device, the server will display a visual map of related content such as "Sherlock Holmes" and "Hercule Poirot."
[1432] Step 10: User Interaction
[1433] Users can click on the displayed relationship map to view more detailed information.
[1434] Additionally, users can enter new search keywords to find related content.
[1435] For example, if a user clicks on "Sherlock Holmes," the server will display detailed information and reviews about the work, as well as information about related TV shows and movies.
[1436] Example 2
[1437] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1438] Conventional information gathering and recommendation systems have struggled to accurately recommend content that users are truly interested in. Furthermore, there are limited ways to visually display the relationships between vast amounts of content in a way that is easy for users to understand. Furthermore, there is a lack of a dynamic content recommendation function based on user emotions, so there is a need for an improved user experience.
[1439] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[1440] In this invention, the server includes means for collecting text data, means for extracting meta information from the collected text data, means for mapping relationships between content items based on the extracted meta information, means for visually displaying the mapped relationships, means for analyzing user emotions, and means for recommending content items based on the user emotions. This allows for accurate recommendation of content items that interest the user, facilitating understanding of the content, and enabling dynamic content recommendation based on emotions.
[1441] "Text data" refers to the content of sentences, books, articles, documents, etc. stored in digital format.
[1442] "Meta information" refers to additional information about content, such as keywords, tags, and names.
[1443] "Relationships between content" refers to the connections based on the characteristics and tags shared between multiple pieces of content.
[1444] "Relationship mapping" refers to a data structure or graph that visually represents commonalities and connections between content.
[1445] "Visual display means" refers to a method of displaying information in a way that is easy for users to understand, such as in the form of graphs or maps.
[1446] "Analysis of user emotions" refers to the means of detecting and recognizing a user's emotional state based on their facial expressions, voice, and behavior.
[1447] "Means for recommending content" refers to methods for presenting appropriate content to users based on analysis results and relationship mapping.
[1448] A "database" refers to a systematic storage system that facilitates the collection, management, and retrieval of data.
[1449] "Domestic and international databases" refers to databases that exist and are accessible domestically and internationally.
[1450] "Commonalities" refer to characteristics or features shared by multiple pieces of content.
[1451] "Link" refers to a connection used to indicate a relationship between pieces of content.
[1452] This invention is a system that uses AI and natural language processing technology to extract and visualize the relationships between a wide variety of content, and combines this with an emotion engine that recognizes the user's emotions to help users discover new content that interests them. The following hardware and software are used to implement this invention.
[1453] Hardware:
[1454] Server: This is the central hardware for processing data collection, analysis, mapping, display, updates, and emotion recognition.
[1455] User terminal: A device used to view content and input emotions. Specifically, this includes PCs, smartphones, tablets, etc.
[1456] software:
[1457] API: An interface for retrieving data from online databases. Examples include the Google Books API and the Open Library API.
[1458] Natural language processing engine: A tool for extracting meta-information (keywords, tags, people's names, etc.) from text data.
[1459] Emotion engine: Software for analyzing the user's emotional state.
[1460] Data processing and calculation:
[1461] 1. Content Collection:
[1462] The server uses the Google Books API and Open Library API to collect various content, such as book data. When a user enters a specific keyword, related data is automatically retrieved. For example, to retrieve book data related to "mystery novels," the server calls the Google Books API and obtains information such as the title, author, publication year, and summary of the related book.
[1463] 2. Text Analysis:
[1464] The server runs the collected text data through a natural language processing engine, performing morphological analysis and entity recognition to extract meta-information. For example, it analyzes the text of the novel "Sherlock Holmes" and extracts keywords such as "Sherlock Holmes," "John Watson," and the place name "London."
[1465] 3. Keyword tag generation:
[1466] Based on the extracted meta information, the server generates keywords and tags related to each piece of content and stores them in a database. For example, a "Sherlock Holmes" novel might be tagged with "detective," "mystery," and "deduction."
[1467] 4. Relationship Mapping:
[1468] Based on the generated keywords and tags, the server analyzes and maps the relationships between multiple pieces of content. For example, "Sherlock Holmes" and "Hercule Poirot" both share the tag "detective," so they are linked together.
[1469] 5. Database integration and updates:
[1470] The server periodically connects to an external database to obtain the latest content information and update its internal database, thereby ensuring that users are always provided with the latest information.
[1471] 6. Incorporating an Emotion Engine:
[1472] Facial expression and voice data acquired from the user's device is input into the emotion engine, which analyzes the user's emotional state in real time. For example, if the user makes an expression showing surprise or interest, that emotional information is sent to the server, and related content is recommended.
[1473] 7. Emotion-based content recommendation:
[1474] Based on the analysis results obtained from the emotion engine, the server dynamically recommends content according to the user's emotional state. For example, if a user expresses surprise while browsing "Sherlock Holmes," the server will recommend mystery novels or suspense movies.
[1475] Specific examples
[1476] Example prompt sentence:
[1477] "If a user wants to collect data related to 'mystery novels,' the server retrieves book information related to 'mystery novels' via the Google Books API. This information includes the book's title, author, publication year, and summary of the contents."
[1478] In this way, the system of the present invention provides users with efficient information provision and enables them to encounter new content through each stage of collection, analysis, mapping, display, update, and emotion recognition.
[1479] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1480] Step 1:
[1481] The server collects the content data.
[1482] Input: User search keywords (e.g., "mystery novel")
[1483] The server uses the Google Books API or Open Library API to obtain book data related to the relevant keywords.
[1484] Output: A book dataset containing book titles, authors, publication years, and summary summaries.
[1485] Specific operation: The server generates a request to call the API and stores the obtained data in an internal database.
[1486] Step 2:
[1487] The server analyzes the collected text data using natural language processing.
[1488] Input: A collected book dataset
[1489] The server uses a natural language processing engine to extract meta-information from the collected text data (e.g., keywords, tags, names).
[1490] Output: Extracted meta information (e.g. "Sherlock Holmes", "Detective", "London")
[1491] How it works: The server uses morphological analyzers to break down text into words and entity recognition algorithms to identify important names and places.
[1492] Step 3:
[1493] The server generates keywords and tags based on the extracted meta information.
[1494] Input: Meta information (keywords, tags, name)
[1495] The server generates keywords and tags related to each piece of content and stores them in a database.
[1496] Output: Keywords and tags associated with the content (e.g., "detective," "mystery," "deduction")
[1497] Specific operation: The server executes a tag generation algorithm based on the extracted meta information and stores the generated tags in a database.
[1498] Step 4:
[1499] The server maps relationships between content based on keywords and tags.
[1500] Input: Keywords and tags associated with the content
[1501] The server analyzes the relationships between multiple pieces of content based on tags and keywords and generates links.
[1502] Output: Relationship map (e.g., a link connecting "Sherlock Holmes" and "Hercule Poirot")
[1503] Specific operation: The server uses a graph database to detect commonalities between tags and generate data to visualize them as links.
[1504] Step 5:
[1505] The server works in conjunction with an external database to periodically update the content information.
[1506] Input: Updates from an external database
[1507] The server obtains the latest content information using Open Library APIs and updates its internal database.
[1508] Output: Updated database (latest book information)
[1509] What it does: The server runs scheduled jobs to retrieve new data from external databases to keep the internal system up to date.
[1510] Step 6:
[1511] The server analyzes the user's emotions using an emotion engine.
[1512] Input: User's facial expressions and voice data
[1513] The server uses an emotion engine to analyze the user's facial expressions and voice to detect emotions.
[1514] Output: User's emotional state (e.g., "surprise," "interest")
[1515] Specific operation: Data collected through the camera and microphone of the user's device is sent to the emotion engine, and emotion analysis is performed in real time.
[1516] Step 7:
[1517] The server dynamically recommends content based on the analysis results of the emotion engine.
[1518] Input: User's emotional state from the emotion engine
[1519] The server runs an algorithm that recommends optimal content based on the user's emotional state.
[1520] Output: A list of recommended content (e.g., related "mystery novels" or "suspense movies")
[1521] Specific operation: The server searches for content that best matches the emotional state and sends a recommendation list to the user terminal.
[1522] Step 8:
[1523] The server builds a user interface that provides a relationship map and keyword search functionality.
[1524] Input: User search keywords and clicks
[1525] The server generates a relationship map to visually display related content.
[1526] Output: Visual relationship map and detailed information
[1527] Specific operation: When a user performs a search on their device, the server visually displays related content, allowing the user to access detailed information by clicking.
[1528] (Application example 2)
[1529] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1530] In today's information-overloaded environment, it is difficult for users to efficiently discover relevant content based on their interests and emotions. Conventional systems simply recommend content based on keywords and tags, but they are unable to dynamically recommend content based on users' emotions and preferences. Given this background, there is a need for a system that allows users to discover relevant content through a more personalized experience.
[1531] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[1532] In this invention, the server includes means for extracting meta-information from collected text data, means for mapping relationships between content items based on the extracted meta-information, means for visually displaying the mapped relationships, means for recognizing a user's emotion, and means for dynamically recommending content items based on the recognized emotion, thereby enabling efficient discovery and recommendation of related content items based on the user's emotion and interests.
[1533] "Collected text data" refers to text information such as books, lyrics, audio data, and image data collected from a wide variety of content.
[1534] "Meta information" refers to information such as keywords, tags, and content names extracted from text data using a natural language processing engine.
[1535] "Relationships between content" refers to relationships that extract commonalities and associations between different content based on collected meta-information and show them as links.
[1536] "Visual display means" refers to means for displaying the relationships between analyzed contents in the form of a map, graph, list, etc., so that the user can visually understand the relationships between the analyzed contents.
[1537] "Means for recognizing user emotions" refers to a means for recognizing a user's emotional state by analyzing the user's facial expressions and voice using sensors such as a camera and microphone.
[1538] "Means for dynamic recommendation based on emotions" refers to a means by which the system selects and recommends relevant content in real time based on the user's perceived emotions.
[1539] This invention is a system that recognizes the user's emotions and recommends appropriate content based on those emotions. The system mainly involves three parties: a server, a terminal, and a user.
[1540] First, the server is equipped with a means for extracting meta-information from the collected text data. This meta-information includes keywords, tags, content names, etc. Specifically, the text data is analyzed using technologies such as natural language processing engines, morphological analysis, and entity recognition to extract information. For example, the contents of books are collected as text data, and then analyzed to extract keywords such as "mystery" and "detective."
[1541] Next, the server has a means for mapping the relationships between content items based on the extracted meta information. This involves visually showing the commonalities and connections between different pieces of content using the extracted keywords and tags. Possible visual display methods include maps, graphs, and lists. For example, multiple books or video works with the keyword "detective" could be linked together to show the user their relationships.
[1542] Furthermore, the device is equipped with a means of recognizing the user's emotions. This is achieved using the smartphone's camera and microphone. Specific software uses OpenCV facial recognition technology and EmotionRecognizer to recognize emotions from the user's facial expressions and voice. For example, if a user shows a surprised expression while reading a mystery novel, the device recognizes that emotion and sends it to the server.
[1543] The server has a means for receiving the user's emotional information and dynamically recommending content based on the recognition results. This allows content related to the user's emotions to be recommended in real time. For example, if the user expresses surprise, mystery novels or suspense movies that are even more surprising and intriguing will be recommended. This allows users to efficiently discover content that matches their emotions and preferences.
[1544] As a specific example, when a user searches for "mystery novels," multiple related content results are displayed. The emotion recognition engine analyzes the user's facial expressions and voice to indicate interest, and the server then recommends further related content based on the results. For example, if a user expresses the emotion "surprise," the server can further recommend "suspense movies," thereby personalizing the user experience. The following are examples of input prompts for the generative AI model:
[1545] "When a user searches for 'mystery novels,' get the 10 most relevant books from the Google Books API, analyze the content of each, and generate keywords and tags. If you detect the emotion 'surprise' from the user's facial expression, also recommend 'suspense movies.'"
[1546] As described above, the system of the present invention provides efficient information provision and enables users to encounter new content through each stage of collection, analysis, mapping, display, emotion recognition, and dynamic recommendation.
[1547] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1548] Step 1:
[1549] The user enters a keyword
[1550] The user enters keywords into the search bar on the device, and the device sends a search request to the server based on this input.
[1551] Step 2:
[1552] The server collects the content
[1553] The server retrieves content data related to keywords from an online database via API. This data includes a wide variety of content, such as books, movies, and music, and is sent to the server. For example, if a user searches for "mystery novels," related book information is retrieved from the Google Books API. The server then stores this retrieved data.
[1554] Step 3:
[1555] The server parses the text data
[1556] The server uses a natural language processing engine to analyze the collected content data. Specifically, it performs morphological analysis and entity recognition to extract keywords and tags. The results of this analysis are saved as meta-information. For example, keywords such as "detective" and "deduction" are extracted from the text data of "Sherlock Holmes."
[1557] Step 4:
[1558] The server maps the relationships between content
[1559] The server maps the relationships between content based on the extracted meta information, finding commonalities in keywords and tags and visualizing them as links. For example, it links works with the common tag "detective" and generates data for displaying them as a visual map.
[1560] Step 5:
[1561] The server visually displays the relationships
[1562] The device displays the relationship map generated by the server to the user, allowing the user to visually check related content. For example, if a user searches for "detective," related books and movies are displayed in map format. The user can click on this map to obtain more information.
[1563] Step 6:
[1564] The device recognizes the user's emotions
[1565] While a user is viewing content displayed on the device, the device uses a camera and microphone to analyze the user's facial expressions and voice. An emotion recognition engine is used to recognize the user's emotional state (surprise, interest, etc.). The results of this analysis are sent to the server.
[1566] Step 7:
[1567] The server recommends content based on emotions.
[1568] The server dynamically recommends content appropriate to the user's emotional state based on the results of the emotion recognition it receives. For example, if the user expresses surprise, the server generates data to recommend similarly interesting suspense movies or mystery novels and sends it to the device.
[1569] Step 8:
[1570] The device displays recommended content to the user
[1571] The device receives recommended content from the server and displays it to the user. The user can browse this recommended content and discover new content that piques their interest. For example, by expressing the emotion "surprise" while browsing "Sherlock Holmes," the device will recommend "suspense movies" and display them to the user.
[1572] These are the specific processing steps and their flow. Through these processes, personalized content recommendations based on the user's emotions and interests are realized.
[1573] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.
[1574] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1575] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.
[1576] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[1577] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.
[1578] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.
[1579] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).
[1580] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.
[1581] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."
[1582] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.
[1583] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).
[1584] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.
[1585] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.
[1586] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[1587] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.
[1588] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.
[1589] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.
[1590] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.
[1591] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.
[1592] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.
[1593] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.
[1594] The following is further disclosed regarding the above embodiment.
[1595] (Claim 1)
[1596] A means for extracting meta-information from the collected text data;
[1597] A means for mapping the relationships between content based on the extracted meta information;
[1598] means for visually displaying the mapped relationships;
[1599] A system including:
[1600] (Claim 2)
[1601] A means for acquiring and updating content data in cooperation with domestic and international databases;
[1602] 10. The system of claim 1.
[1603] (Claim 3)
[1604] A way to find commonalities between pieces of content and display them as links;
[1605] 10. The system of claim 1.
[1606] "Example 1"
[1607] (Claim 1)
[1608] means of collecting information;
[1609] A means for preprocessing the collected text data and extracting meta-information;
[1610] a means for generating keywords and tags based on the extracted meta information;
[1611] A means of mapping relationships between content using generated keywords and tags;
[1612] a means for visually displaying the mapped relationships; and
[1613] A system including:
[1614] (Claim 2)
[1615] A means for acquiring and updating content data from domestic and international databases in cooperation with a server;
[1616] 10. The system of claim 1.
[1617] (Claim 3)
[1618] A means to find commonalities between the mapped content and display them as links;
[1619] 10. The system of claim 1.
[1620] "Application Example 1"
[1621] (Claim 1)
[1622] A means for extracting meta-information from the collected text data;
[1623] A means for mapping the relationships between content based on the extracted meta information;
[1624] means for visually displaying the mapped relationships;
[1625] A means for vectorizing the collected text data and calculating the similarity thereof;
[1626] means for generating a network graph based on the similarity;
[1627] A system including:
[1628] (Claim 2)
[1629] A means for acquiring and updating content data in cooperation with domestic and international databases;
[1630] 10. The system of claim 1.
[1631] (Claim 3)
[1632] A way to find commonalities between pieces of content and display them as links;
[1633] 10. The system of claim 1.
[1634] "Example 2: Combining Emotion Engines"
[1635] (Claim 1)
[1636] a means for collecting text data;
[1637] A means for extracting meta-information from the collected text data;
[1638] A means for mapping the relationships between content based on the extracted meta information;
[1639] a means for visually displaying the mapped relationships; and
[1640] A means of analyzing user emotions,
[1641] A means of recommending content based on user sentiment;
[1642] …
[1643] A system including:
[1644] (Claim 2)
[1645] A means for acquiring and updating content data in cooperation with domestic and international databases;
[1646] 10. The system of claim 1.
[1647] (Claim 3)
[1648] A way to find commonalities between pieces of content and display them as links;
[1649] 10. The system of claim 1.
[1650] "Application example 2 when combining emotion engines"
[1651] (Claim 1)
[1652] A means for extracting meta-information from the collected text data;
[1653] A means for mapping the relationships between content based on the extracted meta information;
[1654] means for visually displaying the mapped relationships;
[1655] a means of recognizing a user's emotions;
[1656] A means for dynamically recommending content based on the recognized sentiment;
[1657] A system including:
[1658] (Claim 2)
[1659] A means for acquiring and updating content data in cooperation with domestic and international databases;
[1660] 10. The system of claim 1.
[1661] (Claim 3)
[1662] A way to find commonalities between pieces of content and display them as links;
[1663] 10. The system of claim 1. [Explanation of symbols]
[1664] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>
Claims
1. A means for extracting meta-information from the collected text data; A means for mapping the relationships between content based on the extracted meta information; means for visually displaying the mapped relationships; A system including:
2. A means for acquiring and updating content data in cooperation with domestic and international databases; The system of claim 1 .
3. A way to find commonalities between pieces of content and display them as links; The system of claim 1 .
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A