System
The system addresses the challenge of misinformation by using fact-checking data from public institutions to evaluate the reliability of search results, allowing users to make informed decisions.
Patent Information
- Application Number
- JP2024116417
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-07-19
- Publication Date
- 2026-01-29
AI Technical Summary
Conventional search engines lack a means for users to easily determine the reliability of information, making it difficult to distinguish between reliable and unreliable information, especially in the context of increased spread of conspiracy theories and misinformation.
A system that includes means for acquiring fact-checking data from public institutions, extracting keywords and assertions from search results, comparing them with fact-checking data, and calculating a reliability score for each search result item.
Enables users to easily check the reliability of search results and make informed decisions based on accurate information.
Smart Images

Figure 2026014943000001_ABST
Abstract
Description
[Technical Field]
[0001] The technology of the present disclosure relates to a system. [Background technology]
[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]
[0004] Due to the COVID-19 pandemic and other factors, there has been an increase in problems such as family members and friends becoming obsessed with conspiracy theories and becoming aggressive. This has made it difficult to distinguish between reliable and unreliable information. Conventional search engines lack a means for users to easily determine the reliability of the information they obtain, making it easy for them to act on inaccurate information. The present invention aims to solve these problems and enable users to easily obtain reliable information. [Means for solving the problem]
[0005] The present invention provides a system including: means for acquiring fact-checking data from a database provided by a public institution; means for receiving a search query from a user; means for acquiring related search results based on the search query; means for extracting important keywords and assertions by text-analyzing the content of the search results; means for comparing the extracted keywords and assertions with the fact-checking data; and means for calculating and displaying a reliability score for each search result item based on the comparison result. This allows users to easily check the reliability of search results and make decisions based on reliable information.
[0006] "Public organizations" are governments, international organizations, and similar organizations that provide reliable information.
[0007] "Fact-checking data" refers to documents and statements provided by public institutions that serve as reliable sources of information.
[0008] A "search query" is a keyword or phrase that a user types into a search engine.
[0009] "Search results" are lists of relevant web pages and videos provided by a search engine based on a search query.
[0010] "Text analytics" is the process of extracting important information from text data using natural language processing techniques.
[0011] "Keywords and assertions" are words and phrases that are central to the content and are extracted through text analysis.
[0012] A "trust score" is a numerical indicator of the reliability of information calculated based on the results of a fact check. [Brief explanation of the drawings]
[0013] [Figure 1]1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION
[0014] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.
[0015] First, the terms used in the following description will be explained.
[0016] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).
[0017] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.
[0018] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.
[0019] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.
[0020] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."
[0021] [First embodiment]
[0022] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0023] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0024] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0025] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.
[0026] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0027] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0028] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.
[0029] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0030] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0031] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0032] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0033] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0034] The present invention is a search system that performs fact-checking based on highly reliable data provided by public institutions. The programs required to implement this system and their processing are shown below.
[0035] System configuration
[0036] The system mainly consists of the following components:
[0037] 1. User Device
[0038] 2. Search Server
[0039] 3. Fact-checking databases
[0040] Program processing
[0041] 1. Receiving user input
[0042] The user enters a search query, such as "coronavirus conspiracy theory," into the device and presses the send button.
[0043] 2. Processing search queries
[0044] The search server analyzes the search query received from the device and extracts related keywords. Through this analysis, the system obtains keywords such as "coronavirus" and "conspiracy theory."
[0045] 3. Obtaining search results
[0046] The search server searches the index database based on the extracted keywords to retrieve a list of related websites and videos. For example, it retrieves a website titled "Is Corona a Man-Made Virus?" and a video titled "The Truth About Conspiracy Theories."
[0047] 4. Obtaining data for fact-checking
[0048] The search server accesses official databases to retrieve the latest data for fact-checking, such as documents containing official statements from the WHO and international health organizations, and stores them in a local cache.
[0049] 5. Content Analysis
[0050] The search server scrapes the content of each website and video and collects this content as text data. For example, it extracts the content of an article on a website titled "Is Corona a Man-Made Virus?"
[0051] 6. Text analysis and keyword extraction
[0052] The search server uses text analysis algorithms to process the collected text data and extract important keywords and assertions, such as "population virus," "evidence," and "research paper" from the content.
[0053] 7. Conduct fact-checks
[0054] The search server compares the extracted keywords and claims against fact-checking data—for example, comparing the claim "man-made virus" with the WHO statement "coronavirus is naturally occurring"—to identify inconsistencies, and calculates a confidence score for each search result based on this.
[0055] 8. Score tagging and search result generation
[0056] The search server tags each search result item with a confidence score and generates a search result list for display to the user. For example, a website titled "Is Corona a Man-Made Virus?" would be tagged with a low confidence score.
[0057] 9. Submitting Search Results
[0058] The search server sends the search results, along with the reliability scores, to the terminal, which then displays the received search results to the user.
[0059] 10. User Browsing and Judgment
[0060] Users review the search results and determine the reliability of the information based on the reliability score. For example, they may prioritize viewing items marked "High Reliability."
[0061] Specific examples
[0062] For example, if a user searches for "COVID-19 conspiracy theory," the search server retrieves related websites and videos, analyzes their content, compares them with official sources, calculates a reliability score, and displays it, allowing users to easily select reliable information.
[0063] In this way, the system of the present invention provides an environment in which users can easily obtain highly reliable information.
[0064] The processing flow will be explained below.
[0065] Step 1:
[0066] The user enters a search query into the device, for example, the keyword "coronavirus conspiracy theory."
[0067] Step 2:
[0068] The terminal transmits the input search query to the server.
[0069] Step 3:
[0070] The server analyzes the received search query and extracts related keywords, for example, "coronavirus" and "conspiracy theory."
[0071] Step 4:
[0072] The server searches an index database based on the extracted keywords to retrieve a list of related websites and videos, such as a website titled "Is COVID-19 a man-made virus?" and a video titled "The Truth About Conspiracy Theories."
[0073] Step 5:
[0074] The server accesses official databases to retrieve the latest data for fact-checking, such as official statements and documents from the WHO.
[0075] Step 6:
[0076] The server scrapes the content of each website and video and collects it as text data. For example, it extracts the content of an article on a website titled "Is Corona a Man-Made Virus?"
[0077] Step 7:
[0078] The server processes the collected text data with a text analysis algorithm to extract key keywords and claims, such as "population virus," "evidence," and "research paper."
[0079] Step 8:
[0080] The server compares the extracted keywords and claims against public fact-checking databases, for example, comparing the claim "man-made virus" with the WHO statement "coronavirus is naturally occurring."
[0081] Step 9:
[0082] The server calculates a confidence score for each search result item based on the comparison, for example, assigning a low confidence score to the "man-made virus" claim because it contradicts official data.
[0083] Step 10:
[0084] The server generates a search result list by tagging each search result with a confidence score. For example, a website asking "Is COVID-19 a man-made virus?" might be tagged with "Low confidence."
[0085] Step 11:
[0086] The server sends the search result list with the confidence scores attached to it to the terminal.
[0087] Step 12:
[0088] The terminal displays the received search result list to the user. The user checks the displayed search results and judges the reliability of the information based on the reliability score. For example, the user may preferentially view items marked "High Reliability."
[0089] This series of processes allows users to easily distinguish highly reliable information and makes decisions based on accurate information.
[0090] Example 1
[0091] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0092] In the current Internet environment, unreliable and false information is often spread, making it difficult for users to find accurate information, especially when it comes to information about serious issues. For this reason, there is a need for a system that allows users to quickly and easily obtain reliable information. In addition, there is a need for a method that effectively utilizes data provided by public institutions and objectively evaluates the reliability of information.
[0093] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0094] In this invention, the server includes means for receiving a search query from a user terminal, means for analyzing the received search query to extract related keywords, means for obtaining a list of related websites and videos based on the extracted keywords, means for obtaining fact-checking data from a database provided by a public institution, means for scraping the obtained content and collecting it as text data, means for analyzing the collected text data using a text analysis algorithm to extract important keywords and claims, means for comparing the extracted keywords and claims with the fact-checking data, and means for calculating and displaying a reliability score for each search result item based on the comparison result, thereby enabling users to quickly and accurately obtain reliable information.
[0095] "User terminal" means a physical device through which a user enters a search query, including a personal computer, smartphone, tablet, etc.
[0096] A "search query" is text data entered by a user to search for specific information.
[0097] A "search server" is a centralized computer system that receives, analyzes, and processes search queries from user terminals.
[0098] "Keywords" are related words or phrases extracted from a search query.
[0099] An "index database" is a database in which information about related websites, videos, etc. is organized and stored in a searchable format.
[0100] "Fact-checking data" is data obtained from databases provided by public institutions to assess the reliability of information, including official statements and research papers.
[0101] "Scraping" is the technique of extracting information from web pages using automated programs.
[0102] A "text analysis algorithm" is a computational method for analyzing text data to extract important information and patterns.
[0103] The "trust score" is a numerical evaluation of the reliability of the information in the search results, and is a value that users can use as a reference when judging the information.
[0104] "Comparison" is the process of comparing extracted keywords and claims with fact-checking data and analyzing for matches and contradictions.
[0105] The present invention is a search system that performs fact-checking based on reliable data provided by public institutions. To implement this system, a user terminal, a search server, and a fact-checking database are used as main components.
[0106] System Configuration
[0107] User Device
[0108] A device on which a user enters a search query and sends it to a search server. User terminals can be personal computers, smartphones, tablets, etc. A web browser is installed on these terminals, and users enter and send search queries through a web interface.
[0109] Search Server
[0110] The search server is the central component that analyzes search queries received from user devices and retrieves and processes relevant information. It is recommended to use a high-performance cloud server (e.g., one of the common cloud server services) for the search server. The following software and algorithms are required:
[0111] Natural Language Processing (NLP) libraries (e.g., NLP libraries commonly used in technical literature)
[0112] Scraping tools (e.g., common open source libraries)
[0113] Database management systems (e.g., commonly used open-source database management systems)
[0114] Fact-checking database
[0115] Fact-checking databases store reliable data provided by public institutions. Specifically, frequently updated data is retrieved from accessible data sources such as international health organizations and government statistical databases, and cached locally within the search server.
[0116] Example of operation
[0117] 1. A user enters a search query, for example, "coronavirus conspiracy theory," into the search bar on their device and presses the send button.
[0118] 2. The search server analyzes the received search query and extracts the related keywords "coronavirus" and "conspiracy theory."
[0119] 3. The search server searches the index database based on the extracted keywords and retrieves relevant search results (websites and videos).
[0120] 4. The search server also retrieves the latest fact-checking data from official databases and stores it in a local cache.
[0121] 5. The search server scrapes the content of the search results and collects it as text data, such as the main text of website articles and video descriptions.
[0122] 6. The search server uses text analysis algorithms to extract important keywords and assertions from the collected text data.
[0123] 7. The search server compares the extracted keywords and claims with fact-checking data, for example, comparing the claim "man-made virus" with the WHO statement "coronavirus is naturally occurring" to identify inconsistencies.
[0124] 8. The search server calculates a confidence score for each search result and displays it as a tag on each search result item.
[0125] Examples of prompt statements
[0126] "What are the steps to verify the credibility of conspiracy theories related to COVID-19?"
[0127] "Please explain specifically how you will conduct fact-checking."
[0128] This system allows users to obtain reliable information quickly and accurately, minimizing the impact of uncertain or false information.
[0129] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0130] Step 1:
[0131] A user enters a search query, such as "COVID-19 conspiracy theory," into the device's search bar and clicks the submit button. The device then sends the search query to a search server. The input data is free-form text, and the output is an HTTP request to the server.
[0132] Step 2:
[0133] The search server analyzes the received search query and extracts relevant keywords. Specifically, it uses the analysis module of a natural language processing (NLP) library (e.g., spaCy) to extract important keywords such as "coronavirus" and "conspiracy theory." The input is free-form text data, and the output is a list of extracted keywords.
[0134] Step 3:
[0135] The search server searches an index database based on the extracted keywords to obtain a list of related websites and videos. The database used for this process is one that has been pre-indexed by a web crawling tool. The input is a list of keywords, and the output is a list of related URLs and titles. Specifically, it includes a website titled "Is Corona a Man-Made Virus?" and a video titled "The Truth About Conspiracy Theories."
[0136] Step 4:
[0137] The search server makes HTTP requests to official databases (e.g., international health organizations or government agencies) to retrieve the latest data for fact-checking. The retrieved data is stored in a local cache for use in the next step. The input is the official API endpoint, and the output is a dataset of trusted documents and statements.
[0138] Step 5:
[0139] The search server uses a scraping tool (e.g., BeautifulSoup) to collect the content of the retrieved websites and videos. Specifically, it analyzes the HTML code of the webpage and extracts the article text and metadata as text data. The input is a list of URLs, and the output is content data in text format. For example, it collects the content of articles on a website titled "Is Corona a Man-Made Virus?"
[0140] Step 6:
[0141] The search server analyzes the collected text data using natural language processing (NLP) algorithms to extract important keywords and claims. Specifically, it uses an NLP library to extract frequently occurring words and phrases in the text and identify important keywords such as "population virus," "evidence," and "research paper." The input is text content data, and the output is a list of important keywords.
[0142] Step 7:
[0143] The search server uses the extracted keywords and claims to compare them with fact-checking data obtained from official databases. For example, it compares the claim "man-made virus" with the WHO statement "coronavirus is naturally occurring" and analyzes contradictions and similarities. The input is a list of keywords and reliable document data, and the output is the comparison results and a reliability score.
[0144] Step 8:
[0145] The search server tags each search result item with the calculated confidence score and generates a list of search results to display to the user. Specifically, the process involves adding the confidence score to the title and URL of each search result to create a list. The input is the comparison result and confidence score, and the output is a tagged list of search results.
[0146] Step 9:
[0147] The search server sends the tagged search result list to the terminal, which receives it and generates an HTML page to display to the user. The input is the tagged search result list, and the output is the search result page displayed to the user.
[0148] Step 10:
[0149] Users check the displayed search results and judge the reliability of the information based on the reliability score. Specifically, they prioritize browsing items with high reliability and acquire or share the information. The input is the displayed search result page, and the output is the user's browsing actions and judgment results.
[0150] This allows users to obtain reliable information quickly and accurately.
[0151] (Application example 1)
[0152] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0153] In recent years, the amount of information on the Internet has increased dramatically, including information of unknown authenticity and potentially misleading information. This has made it difficult for users to easily obtain reliable information. Furthermore, in virtual stores, it is difficult for consumers to judge the reliability of product information themselves, which risks leading to inaccurate purchasing decisions. Therefore, there is a need for a system that can automatically determine the reliability of information and present a reliability score when consumers search for product information in virtual stores.
[0154] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0155] In this invention, the server includes means for acquiring fact-checking data from a database provided by a public institution, means for receiving a search query from a user, means for acquiring related search results based on the search query, means for extracting important keywords and assertions by text-analyzing the content of the search results, means for comparing the extracted keywords and assertions with the fact-checking data, means for calculating and displaying a reliability score for each search result item based on the comparison result, and means for tagging product information with the reliability score, allowing users to judge reliability when making purchasing decisions. This allows consumers to easily judge the reliability of product information in virtual stores and make purchasing decisions with confidence.
[0156] "Public institution" refers to a public institution, such as a government or local government, that provides public functions or services.
[0157] A "database" is an information repository that organizes and stores a collection of data so that it can be efficiently searched and used.
[0158] "Fact-checking data" refers to reliable data used to verify the veracity of information.
[0159] "User" means any individual person or entity that uses the System or Services.
[0160] A "search query" refers to a question or keyword that a user enters into a search engine or system.
[0161] "Search Results" means the information returned by a search engine or system based on a search query.
[0162] "Content" refers to any piece of information or data present on a website or digital media.
[0163] "Text analytics" refers to the techniques and processes used to analyze text data and understand its content.
[0164] "Keywords" refer to key words that play an important role in a text or search query.
[0165] An "argument" is a statement of opinion or perspective on a certain matter.
[0166] "Comparison" refers to the act of comparing two or more things to see their differences and similarities.
[0167] A "trust score" is a numerical indicator that evaluates the reliability of information or data.
[0168] "Display" refers to the means by which information is visually conveyed to the user.
[0169] "Tagging" refers to the act of adding labels and scores to information.
[0170] "Product information" refers to data including descriptions, features, specifications, etc. of the target product.
[0171] "Consumption behavior" refers to the series of actions that consumers take to purchase goods and services.
[0172] This invention describes a system that automatically determines the reliability of product information when a user searches for it in a virtual store and presents a reliability score. This system is composed of the following main components:
[0173] System configuration
[0174] The system is implemented primarily using the following hardware and software:
[0175] 1. User terminal: a smartphone, smart glasses, head-mounted display, or other display device.
[0176] 2. Search server: A server responsible for data analysis and fact-checking.
[0177] 3. Fact-checking database: A database that stores reliable data provided by public institutions.
[0178] Program processing
[0179] User Roles
[0180] A user inputs a search query, such as "antioxidant supplements," into a user terminal and presses a search button. This operation sends the query to a search server.
[0181] The role of the search server
[0182] The search server analyzes the search query received from the user terminal and extracts related keywords, which allows the system to obtain keywords such as "antioxidant" and "supplement."
[0183] Furthermore, the search server searches the virtual store database based on the extracted keywords to obtain a list of related products. For example, product information such as "Antioxidant Supplement A" and "Supplement B" is obtained.
[0184] The search server then accesses databases provided by public organizations (e.g., government agencies and research institutes) to retrieve reliable information and retrieve the latest data for fact-checking, including official reports and the latest research papers.
[0185] Content analysis and fact-checking
[0186] The search server scrapes the descriptions and reviews of each product and collects them as text data. For example, it analyzes the description of "Antioxidant Supplement A" and extracts important keywords and claims.
[0187] Furthermore, the search server compares the extracted keywords and claims with information from databases of public institutions, and calculates a credibility score to evaluate the reliability of the product. For example, a credibility score for a product is determined by comparing the claim "has antioxidant properties."
[0188] Visibility and User Roles
[0189] Finally, the search server tags each product with the calculated reliability score and displays it as information to help users judge reliability when deciding on their purchasing behavior, allowing users to select products with confidence based on reliable information.
[0190] Examples and prompts
[0191] For example, if a user searches for "antioxidant supplements," the search server retrieves related product information, analyzes the descriptions of each product, and calculates a reliability score. As a result, products with a "High Reliability" rating can be presented to the user preferentially.
[0192] Here are some examples of prompts:
[0193] "Antioxidant supplement A has been proven effective in clinical trials. Please compare this information with data from official institutions and calculate a reliability score."
[0194] In this way, this system provides an environment in which consumers can easily judge the reliability of product information in a virtual store.
[0195] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0196] Step 1: Receiving User Input
[0197] A user enters a search query such as "antioxidant supplements" into a device such as a smartphone or head-mounted display and presses the search button. The device then sends this query to a search server. The input data is the user's search query, and the output data is the query information sent to the search server.
[0198] Step 2: Processing the search query
[0199] The search server analyzes the search query received from the device and extracts the related keywords "antioxidant" and "supplement." Specifically, it uses a natural language processing algorithm to tokenize the search query and extract important keywords. The input data is the received query information, and the output data is the extracted keywords.
[0200] Step 3: Getting search results
[0201] The search server searches the virtual store database based on the extracted keywords and retrieves a list of related products. For example, product information such as "antioxidant supplement A" and "supplement B" is retrieved. The input data are the extracted keywords, and the output data is a list of products.
[0202] Step 4: Obtaining fact-checking data
[0203] The search server accesses public databases to retrieve the latest data for fact-checking, including official reports from government agencies and research institutes and the latest research papers. The input data is the access request to the public database, and the output data is the retrieved data for fact-checking.
[0204] Step 5: Analyze your content
[0205] The search server scrapes the descriptions and reviews of each product and collects this content as text data. Specifically, it extracts the text of articles and reviews from each page. The input data is a list of products, and the output data is the collected text data.
[0206] Step 6: Text analysis and keyword extraction
[0207] The search server processes the collected text data using a text analysis algorithm to extract key keywords and claims. Specifically, it uses TF-IDF (Term Frequency-Inverse Document Frequency) to identify important keywords. The input data is the collected text data, and the output data is the extracted keywords and claims.
[0208] Step 7: Conduct a fact check
[0209] The search server compares the extracted keywords and claims with the fact-checking data, checking for keyword matches and inconsistencies, and calculating a reliability score for each product. The input data are the extracted keywords and claims and the fact-checking data, and the output data is a reliability score for each product.
[0210] Step 8: Score tagging and search result generation
[0211] The search server tags each product information with the calculated confidence score and generates a search result list to display to the user. Specifically, it assigns a confidence score to each product information and converts it into a format that can be displayed in the user interface. The input data are the confidence scores and the product list, and the output data is the tagged search result list.
[0212] Step 9: Submit search results
[0213] The search server sends the tagged search results to the terminal, which then displays the received search results to the user. The input data is the tagged search result list, and the output data is the displayed search results.
[0214] Step 10: User View and Decision
[0215] The user checks the displayed search results and judges the reliability of the product information based on the reliability score. Specifically, the user may take action such as prioritizing the purchase of products displayed as "High reliability." The input data are the displayed search results, and the output data are the user's purchasing behavior.
[0216] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.
[0217] This invention combines a search system that performs fact-checking based on highly reliable data provided by public institutions with an emotion engine that recognizes user emotions. The programs required to implement this system and their processing are shown below.
[0218] System configuration
[0219] The system mainly consists of the following components:
[0220] 1. User Device
[0221] 2. Search Server
[0222] 3. Fact-checking databases
[0223] 4. Emotion Engine
[0224] Program processing
[0225] 1. Receiving user input
[0226] The user enters a search query, such as "coronavirus conspiracy theory," into the device and presses the send button.
[0227] 2. Processing search queries
[0228] The terminal transmits the input search query to the server.
[0229] 3. Search Query Analysis
[0230] The server analyzes the received search query and extracts related keywords, for example, "coronavirus" and "conspiracy theory."
[0231] 4. Obtaining search results
[0232] The server searches an index database based on the extracted keywords to retrieve a list of related websites and videos, such as a website titled "Is COVID-19 a man-made virus?" and a video titled "The Truth About Conspiracy Theories."
[0233] 5. Acquisition of data from public institutions
[0234] The server accesses official databases to retrieve the latest data for fact-checking, such as official statements and documents from the WHO.
[0235] 6. Content Analysis
[0236] The server scrapes the content of each website and video and collects it as text data. For example, it extracts the content of an article on a website titled "Is Corona a Man-Made Virus?"
[0237] 7. Text Analysis and Keyword Extraction
[0238] The server processes the collected text data with a text analysis algorithm to extract key keywords and claims, such as "population virus," "evidence," and "research paper."
[0239] 8. Fact Check
[0240] The server compares the extracted keywords and claims against public fact-checking databases—for example, comparing the claim "man-made virus" with the WHO statement "coronavirus is naturally occurring"—and calculates a confidence score for each search result.
[0241] 9. Emotion Recognition with Emotion Engine
[0242] The device uses an emotion engine to analyze the user's emotions based on the user's input and operations (e.g., keystroke speed, input content, facial recognition, etc.). For example, it determines whether the user is feeling stressed.
[0243] 10. Emotional Feedback
[0244] The server then adjusts the reliability score and display order of search results based on the analysis results of the emotion engine. For example, if a user is feeling stressed, more reliable information will be displayed first.
[0245] 11. Score tagging and search result generation
[0246] The server tags each search result item with a confidence score and generates a search result list to display to the user. For example, a site asking "Is COVID-19 a man-made virus?" might be given a "Low confidence" rating.
[0247] 12. Submitting Search Results
[0248] The server sends the search result list with the confidence scores attached to it to the terminal.
[0249] 13. Display of search results
[0250] The device displays the received search results to the user. The user checks the displayed search results and determines the reliability of the information based on the reliability score. For example, the user may prioritize viewing items marked "High Reliability."
[0251] Specific examples
[0252] For example, if a user searches for "COVID-19 conspiracy theory," the search server retrieves related websites and videos and analyzes their content. It then compares them with documents from official institutions and calculates a reliability score. The device also uses an emotion engine to analyze the user's emotions (e.g., anxiety, stress) and adjusts the display order of search results based on the results. As a result, users can easily select reliable information and browse with peace of mind.
[0253] In this way, the system of the present invention provides an environment in which the user can easily obtain highly reliable information, and further provides information according to the user's emotional state.
[0254] The processing flow will be explained below.
[0255] Step 1:
[0256] The user enters a search query into the device, for example, the keyword "coronavirus conspiracy theory."
[0257] Step 2:
[0258] The terminal transmits the input search query to the server.
[0259] Step 3:
[0260] The server analyzes the received search query and extracts related keywords, for example, "coronavirus" and "conspiracy theory."
[0261] Step 4:
[0262] The server searches an index database based on the extracted keywords to retrieve a list of related websites and videos, such as a website titled "Is COVID-19 a man-made virus?" and a video titled "The Truth About Conspiracy Theories."
[0263] Step 5:
[0264] The server accesses official databases to retrieve the latest data for fact-checking, such as official statements and documents from the WHO.
[0265] Step 6:
[0266] The server scrapes the content of each website and video and collects it as text data. For example, it extracts the content of an article on a website titled "Is Corona a Man-Made Virus?"
[0267] Step 7:
[0268] The server processes the collected text data with a text analysis algorithm to extract key keywords and claims, such as "population virus," "evidence," and "research paper."
[0269] Step 8:
[0270] The server compares the extracted keywords and claims against public fact-checking databases, for example, comparing the claim "man-made virus" with the WHO statement "coronavirus is naturally occurring."
[0271] Step 9:
[0272] The server calculates a confidence score for each search result item based on the comparison, for example, assigning a low confidence score to the "man-made virus" claim because it contradicts official data.
[0273] Step 10:
[0274] The device uses an emotion engine to analyze the user's emotions based on the user's input and operations (e.g., keystroke speed, input content, facial recognition, etc.). For example, it can determine whether the user is feeling stressed.
[0275] Step 11:
[0276] The server then adjusts the reliability score and display order of search results based on the analysis results of the emotion engine. For example, if a user is feeling stressed, more reliable information will be displayed first.
[0277] Step 12:
[0278] The server tags each search result item with a confidence score and generates a search result list to display to the user. For example, a site asking "Is COVID-19 a man-made virus?" might be given a "Low confidence" rating.
[0279] Step 13:
[0280] The server sends the search result list with the confidence scores attached to it to the terminal.
[0281] Step 14:
[0282] The terminal displays the received search result list to the user. The user checks the displayed search results and judges the reliability of the information based on the reliability score. For example, the user may preferentially view items marked "High Reliability."
[0283] This series of processes allows the user to easily distinguish highly reliable information, and also makes it possible to provide information that is appropriate for the user's emotional state.
[0284] Example 2
[0285] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0286] In today's world, there is a huge amount of information available on the Internet, but it is not easy to find accurate and reliable information. Furthermore, providing information without considering the user's emotional state can increase stress and anxiety. Therefore, there is a need for a search system that takes into account the user's emotional state while ensuring the reliability of the information.
[0287] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0288] In this invention, the server includes means for acquiring fact-checking data from a database provided by a public institution, means for receiving a search query from a user, means for acquiring related search results based on the search query, means for analyzing the content of the search results to extract important keywords and claims, means for comparing the extracted keywords and claims with the fact-checking data, means for calculating and displaying a reliability score for each search result item based on the comparison result, means for analyzing user emotions, and means for adjusting the display order of the search results based on the emotion analysis results.
[0289] This allows the user to easily obtain highly reliable information and to be provided with information that corresponds to the user's own emotional state.
[0290] A "database" is a storage device that stores highly reliable data provided by public institutions.
[0291] "Fact-checking data" refers to documents and statements released by public institutions, and is basic data used to verify the accuracy of information.
[0292] A "search query" is a question or keyword that a user enters into the system, and is the search condition that the system investigates.
[0293] "Search results" are lists of information retrieved from databases or the Internet based on a search query.
[0294] "Text analysis" is a technique that uses natural language processing technology to understand the content of text data and extract important keywords and assertions.
[0295] "Keywords" are important words or phrases extracted from a search query or analyzed text.
[0296] "Arguments" are important opinions or points of view in a text that are revealed through text analysis.
[0297] "Comparison" is the process of comparing the fact-checking data with the extracted keywords and claims to confirm whether they match or disagree.
[0298] The "trust score" is a numerical representation of the reliability of information based on the results of fact-checking, and is an indicator of the reliability of search results.
[0299] "Emotion analysis" is the process of analyzing a user's emotional state based on their input and actions, and is a technology that determines what emotions the user is feeling.
[0300] "Display ranking" refers to the order in which each item is displayed in the search result list, and is determined based on confidence scores and sentiment analysis results.
[0301] This invention combines a search system that performs fact-checking based on reliable data provided by public institutions with an emotion engine that recognizes user emotions. This system mainly consists of the following components: a user terminal, a search server, a fact-checking database, and an emotion engine.
[0302] System configuration
[0303] To implement this invention, the following hardware and software are used:
[0304] 1. User Device
[0305] Hardware: Personal computers, smartphones
[0306] Software: Web browser, input interface
[0307] 2. Search Server
[0308] Hardware: High-performance server
[0309] Software: Natural language processing libraries (e.g., NLTK, SpaCy, BERT), search engines (e.g., Elasticsearch, Solr), scraping tools (e.g., BeautifulSoup, Selenium), sentiment analysis APIs (e.g., Microsoft Azure Cognitive Services, Emotion API)
[0310] 3. Fact-checking databases
[0311] Hardware: Database Server
[0312] Software: Public Sector API Access
[0313] 4. Emotion Engine
[0314] Hardware: Embedded systems, cameras and keyboards attached to user devices
[0315] Software: Emotion analysis algorithms, face tracking software (e.g. OpenCV)
[0316] Implementation Procedure
[0317] 1. Receiving user input
[0318] The user enters "COVID-19 conspiracy theory" into the search field and clicks the search button. The device receives this input, temporarily stores it in its internal memory, and prepares to send it to the server.
[0319] 2. Processing search queries
[0320] The device sends the search query entered by the user to the server as an HTTP request, which includes the query text data and metadata such as the user ID.
[0321] 3. Search Query Analysis
[0322] The server parses the incoming HTTP request and processes the query text with a natural language processing algorithm (e.g., Python's NLTK library), which identifies the main keywords "coronavirus" and "conspiracy theory."
[0323] 4. Obtaining search results
[0324] The server queries an index database using a search engine (e.g., Elasticsearch or Solr) based on the extracted keywords, which retrieves a list of related websites and videos.
[0325] 5. Acquisition of data from public institutions
[0326] The server sends API requests to public databases to retrieve the latest, most reliable information, including official statements and research papers from organizations like the WHO and CDC.
[0327] 6. Content Analysis
[0328] The server uses scraping technology (e.g., BeautifulSoup or Selenium) to collect the content of each site and video as text data, extracting the content of the article on the site "Is Corona a Man-Made Virus?" and the subtitles of the video.
[0329] 7. Text Analysis and Keyword Extraction
[0330] The server then processes the collected text data again using a natural language processing algorithm (e.g., SpaCy or BERT) to extract important keywords and assertions, such as "population virus," "evidence," and "research paper."
[0331] 8. Fact Check
[0332] The server compares the extracted keywords and claims with official data it has obtained—for example, comparing the claim "man-made virus" with the WHO statement "coronavirus is naturally occurring"—and then uses an algorithm to calculate a confidence score for each search result, which is calculated on a scale from 0 to 100.
[0333] 9. Emotion Recognition with Emotion Engine
[0334] The device uses an emotion engine to analyze the user's emotions based on the user's input and operations, such as keystroke speed and facial recognition (using OpenCV), to determine whether the user is feeling stressed or anxious.
[0335] 10. Emotional Feedback
[0336] The server then adjusts the reliability score and display order of search results based on the analysis results of the emotion engine. For example, if a user is feeling stressed, an algorithm is applied to prioritize displaying information with high reliability.
[0337] 11. Score tagging and search result generation
[0338] The server tags each search result with a confidence score and generates a list of search results to display to the user. For example, a website asking "Is COVID-19 a man-made virus?" might be tagged with "Confidence: 20 (Low)."
[0339] 12. Submission and Display of Search Results
[0340] The server sends the search result list, with assigned reliability scores, to the terminal as an HTTP response. The terminal displays the received search result list on its user interface. The user checks this display and decides which information to view based on the reliability score. For example, they may prioritize clicking and viewing items with a "Reliability: 90" rating.
[0341] Specific examples
[0342] For example, if a user searches for "coronavirus conspiracy theory," the following specific actions occur: The user enters "coronavirus conspiracy theory" into the search bar and clicks the search button. The device receives the input, generates an HTTP request, and sends it to the server. The server analyzes the query and extracts related keywords. The server uses a search engine to retrieve relevant sites and videos. The server retrieves reliable data from public institution databases via an API request. The server scrapes the content of the retrieved sites and videos and collects it as text data. The server uses a natural language processing algorithm to extract important keywords and claims from the collected text data. The server compares the extracted keywords and claims with public institution data and calculates a confidence score. The device analyzes the user's emotions using an emotion engine and sends the results to the server. The server adjusts the confidence score and display ranking of the search results based on the user's emotions. The server tags the confidence score and generates a search result list. The server sends the final search result list to the device. The device displays the search result list on a user interface, and the user can review and select results.
[0343] Example prompt sentence:
[0344] "New coronavirus conspiracy theory"
[0345] Result: A list of highly reliable information is displayed, including an official statement from a public institution that "coronavirus is a natural occurrence." Because the emotion engine determined that the user was highly stressed, highly reliable information is displayed first.
[0346] In this way, the system of the present invention provides an environment in which the user can easily obtain highly reliable information, and further provides information according to the user's emotional state.
[0347] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0348] Step 1:
[0349] The user enters "COVID-19 conspiracy theory" in the search field and clicks the search button. This causes the device to receive a search query. The input is the search query entered by the user, and the output is the query temporarily stored in the internal memory. The device prepares to send this data to the server.
[0350] Step 2:
[0351] The device sends the search query entered by the user to the server as an HTTP request. The input is the search query stored in the internal memory, and the output is the HTTP request sent to the server. The request includes the text data of the query and metadata such as the user ID.
[0352] Step 3:
[0353] The server parses the received HTTP request and processes the query text with an NLP algorithm (e.g., Python's NLTK library). The input is the HTTP request sent to the server, and the output is the parsed search query and keywords extracted from it, e.g., "coronavirus" and "conspiracy theory."
[0354] Step 4:
[0355] The server queries the index database using Elasticsearch or Solr based on the extracted keywords. The input is the extracted keywords, and the output is a list of related websites and videos, such as a website called "Is COVID-19 a man-made virus?" or a video called "The Truth About Conspiracy Theories."
[0356] Step 5:
[0357] The server sends API requests to public databases to retrieve the latest, most reliable information. The input is the API request, and the output is the retrieved public data, such as an official WHO statement or a research paper.
[0358] Step 6:
[0359] The server collects the content of each website and video as text data using scraping technology (e.g., BeautifulSoup or Selenium). The input is the URL of the relevant website or video, and the output is the scraped text data, such as the content of an article on the website "Is Coronavirus a Man-Made Virus?"
[0360] Step 7:
[0361] The server processes the collected text data again with an NLP algorithm (e.g., SpaCy or BERT) to extract important keywords and claims. The input is the collected text data, and the output is the extracted important keywords and claims, such as "population virus," "evidence," and "research paper."
[0362] Step 8:
[0363] The server compares the extracted keywords and claims with the retrieved official data. The input is the extracted keywords and claims and the official data, and the output is a confidence score for each search result, e.g., "Confidence: 20 (Low)."
[0364] Step 9:
[0365] The device uses an emotion engine to analyze the user's emotions based on their input and operations. Specifically, it uses keystroke speed and facial recognition technology (e.g., OpenCV). The input is the user's input data and operation data, and the output is the analyzed user's emotional state, such as "high stress."
[0366] Step 10:
[0367] The server adjusts the reliability score and display order of search results based on the analysis results of the emotion engine. The input is the emotion analysis result and reliability score, and the output is an adjusted search result list, for example, a list in which highly reliable information is displayed at the top.
[0368] Step 11:
[0369] The server tags each search result with a confidence score and generates a final search result list to display to the user. The input is the adjusted search result list, and the output is a search result list with a confidence score, e.g., "Confidence: 90 (High)".
[0370] Step 12:
[0371] The server sends the final search result list to the terminal. The input is the search result list with the assigned confidence scores, and the output is the HTTP response sent to the terminal.
[0372] Step 13:
[0373] The terminal displays the received search result list on the user interface. The input is the HTTP response sent to the terminal, and the output is the search result list displayed on the screen. The user checks the results and decides which information to view based on the confidence score.
[0374] (Application example 2)
[0375] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0376] In today's information society, it is extremely important for users to have access to reliable information. However, while a vast amount of information is easily available, it is often difficult to verify its reliability. Furthermore, there are no established methods for providing appropriate information to users who are feeling stressed or anxious. Therefore, an objective of the present invention is to provide a system that provides users with reliable information and also provides information that is tailored to the user's emotional state.
[0377] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for acquiring fact-checking data from a database provided by a public institution, means for receiving a search query from a user, means for acquiring related search results based on the search query, means for text-analyzing the content of the search results to extract important keywords and assertions, means for comparing the extracted keywords and assertions with the fact-checking data, means for calculating and displaying a reliability score for each search result item based on the comparison result, means for recognizing passenger emotions using an emotion engine, and means for adjusting the reliability score and information display based on the recognized emotions. This allows the user to appropriately acquire reliable information and provides information according to the user's emotional state.
[0378] A "public institution" is an organization that operates for the public interest, such as a government or local government, and issues official data and statements.
[0379] A "database" is a collection of information designed to efficiently store, retrieve, and update data.
[0380] "Fact-checking data" is reliable information and materials collected and provided for the purpose of fact-checking.
[0381] "User" means an individual or organization that uses a system or application.
[0382] A "search query" is a string of characters or keywords that a user enters when searching for information.
[0383] "Search results" are a collection of information or data retrieved based on a search query.
[0384] "Text analytics" refers to techniques and methods for extracting important information and patterns from text data.
[0385] "Keywords" are important words or phrases used to represent specific information.
[0386] An "assertion" is an expression that expresses a particular opinion or point of view.
[0387] "Comparison" is the act of comparing two or more pieces of information or data and evaluating their differences and similarities.
[0388] A "trust score" is a numerical indicator used to evaluate the reliability of information.
[0389] "Display" is the act of visually presenting acquired information or results to the user.
[0390] An "emotion engine" is software or hardware that analyzes and recognizes a user's emotions.
[0391] "Recognition" is the act of analyzing, understanding, and grasping information and data.
[0392] "Adjustment" is the act of changing information or settings to accommodate specific conditions or circumstances.
[0393] "Means" are the methods or techniques used to achieve a particular goal.
[0394] The present invention is a system for providing highly reliable information to a user and providing information according to the user's emotional state. Specific embodiments of the present invention will be described below.
[0395] System Configuration
[0396] Hardware
[0397] 1. User device: A device such as a smartphone, tablet, or PC that accepts user input and operations. It also has a camera and microphone to collect data for emotion recognition.
[0398] 2. Search Server: A central server that searches, analyzes, and fact-checks various data.
[0399] 3. Fact-checking database: A database that stores reliable data provided by public institutions.
[0400] 4. Emotion engine: Software for analyzing the user's emotional state.
[0401] software
[0402] 1. EmotionEngine: Software that analyzes data input from cameras and microphones in real time to recognize the user's emotions. It uses image processing libraries such as OpenCV.
[0403] 2. FactCheckEngine: An engine that cross-checks the information obtained with public data and evaluates its reliability.
[0404] 3. PublicDataFetcher: Software for retrieving up-to-date and reliable data from public databases.
[0405] System Operation
[0406] The server first receives a search query from a user device, then retrieves relevant search results based on the search query, performs text analysis on the content, extracts important keywords and claims, and compares the extracted keywords and claims with fact-checking data to calculate their reliability score.
[0407] At the same time, EmotionEngine recognizes the user's emotions in real time using the camera and microphone installed on the user's device. Based on the recognized emotions, the server adjusts the reliability score and information display to provide the user with appropriate information.
[0408] Specific examples
[0409] For example, if a user enters the search query "traffic congestion current situation," the server retrieves related websites and news articles, performs fact-checking, and calculates and displays a reliability score. If the emotion engine recognizes that the user is feeling stressed, it prioritizes information with high reliability and suggests relaxing music and videos.
[0410] Prompt Sentence Examples
[0411] Provide information to promote safe driving when users are under stress, such as prioritizing public traffic and weather information and playing relaxing music and scenic videos.
[0412] In this way, a system is realized that allows users to easily obtain highly reliable information and receive information that is optimal for their emotional state.
[0413] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0414] Step 1:
[0415] User enters search query and submits
[0416] A user inputs a query, for example, "traffic congestion current situation," into the terminal and presses the send button. After the input is made, the terminal sends the search query to the server. The input is in the form of text, and the output is the search query sent to the server.
[0417] Step 2:
[0418] The server parses the search query
[0419] The server analyzes the received search query and extracts relevant keywords. For example, it extracts the keywords "traffic congestion" and "current situation." The input is the search query text data, and the output is the analyzed keyword list. The server performs this process using a text analysis algorithm.
[0420] Step 3:
[0421] The server retrieves the search results
[0422] The server searches the index database based on the extracted keywords to obtain a list of relevant sites and news articles. The input is the parsed keyword list, and the output is a list of relevant search results. The server uses the index database to perform this process.
[0423] Step 4:
[0424] Server acquires data from public institutions
[0425] The server accesses a public database to retrieve the latest data for fact-checking, for example, data from a traffic information center. The input is the URL or API key of the public database to access, and the output is the trusted data. The server performs this process using an HTTP request.
[0426] Step 5:
[0427] The server analyzes the content
[0428] The server scrapes the content of each website or news article and collects it as text data. The input is a list of relevant search results, and the output is the retrieved text data. The server performs this process using web scraping technology.
[0429] Step 6:
[0430] The server performs text analysis and keyword extraction.
[0431] The server processes the collected text data with text analysis algorithms to extract important keywords and assertions. The input is the collected text data, and the output is the extracted keywords and assertions. The server performs this processing using natural language processing (NLP) techniques.
[0432] Step 7:
[0433] Server fact-checking
[0434] The server compares the extracted keywords and claims against public fact-checking databases and calculates a confidence score. The inputs are the extracted keywords and data from the public databases, and the output is a confidence score. The server performs this process using database queries and comparison algorithms.
[0435] Step 8:
[0436] The device uses an emotion engine to recognize emotions.
[0437] The device analyzes the user's emotions using a camera and microphone. The input is video and audio data, and the output is the recognized emotion data. The device performs this processing using the EmotionEngine.
[0438] Step 9:
[0439] Server provides feedback based on emotions
[0440] The server adjusts the confidence score and information display based on the analysis results of the emotion engine. The input is the recognized emotion data and confidence score, and the output is the adjusted information display ranking and feedback. The server performs this process using an algorithm.
[0441] Step 10:
[0442] The server generates a search result list tagged with a confidence score
[0443] The server tags each search result item with a confidence score and generates a search result list for display to the user. The input is the confidence score and the list of search results, and the output is the tagged search result list. The server performs this process using a list generation algorithm.
[0444] Step 11:
[0445] The server sends the search result list
[0446] The server sends the search result list with attached confidence scores to the user device. The input is the tagged search result list, and the output is the data sent to the user device. The server performs this process using a network protocol.
[0447] Step 12:
[0448] Your device will display the search results
[0449] The terminal displays the received search results to the user. The user reviews the displayed search results and determines the reliability of the information based on the reliability score. The input is the search result list received from the server, and the output is the displayed search results. The terminal performs this process using a graphical user interface (GUI).
[0450] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0451] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0452] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.
[0453] [Second embodiment]
[0454] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.
[0455] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0456] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0457] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.
[0458] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0459] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0460] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0461] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0462] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0463] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0464] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0465] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."
[0466] The present invention is a search system that performs fact-checking based on highly reliable data provided by public institutions. The programs required to implement this system and their processing are shown below.
[0467] System configuration
[0468] The system mainly consists of the following components:
[0469] 1. User Device
[0470] 2. Search Server
[0471] 3. Fact-checking databases
[0472] Program processing
[0473] 1. Receiving user input
[0474] The user enters a search query, such as "coronavirus conspiracy theory," into the device and presses the send button.
[0475] 2. Processing search queries
[0476] The search server analyzes the search query received from the device and extracts related keywords. Through this analysis, the system obtains keywords such as "coronavirus" and "conspiracy theory."
[0477] 3. Obtaining search results
[0478] The search server searches the index database based on the extracted keywords to retrieve a list of related websites and videos. For example, it retrieves a website titled "Is Corona a Man-Made Virus?" and a video titled "The Truth About Conspiracy Theories."
[0479] 4. Obtaining data for fact-checking
[0480] The search server accesses official databases to retrieve the latest data for fact-checking, such as documents containing official statements from the WHO and international health organizations, and stores them in a local cache.
[0481] 5. Content Analysis
[0482] The search server scrapes the content of each website and video and collects this content as text data. For example, it extracts the content of an article on a website titled "Is Corona a Man-Made Virus?"
[0483] 6. Text analysis and keyword extraction
[0484] The search server uses text analysis algorithms to process the collected text data and extract important keywords and assertions, such as "population virus," "evidence," and "research paper" from the content.
[0485] 7. Conduct fact-checks
[0486] The search server compares the extracted keywords and claims against fact-checking data—for example, comparing the claim "man-made virus" with the WHO statement "coronavirus is naturally occurring"—to identify inconsistencies, and calculates a confidence score for each search result based on this.
[0487] 8. Score tagging and search result generation
[0488] The search server tags each search result item with a confidence score and generates a search result list for display to the user. For example, a website titled "Is Corona a Man-Made Virus?" would be tagged with a low confidence score.
[0489] 9. Submitting Search Results
[0490] The search server sends the search results, along with the reliability scores, to the terminal, which then displays the received search results to the user.
[0491] 10. User Browsing and Judgment
[0492] Users review the search results and determine the reliability of the information based on the reliability score. For example, they may prioritize viewing items marked "High Reliability."
[0493] Specific examples
[0494] For example, if a user searches for "COVID-19 conspiracy theory," the search server retrieves related websites and videos, analyzes their content, compares them with official sources, calculates a reliability score, and displays it, allowing users to easily select reliable information.
[0495] In this way, the system of the present invention provides an environment in which users can easily obtain highly reliable information.
[0496] The processing flow will be explained below.
[0497] Step 1:
[0498] The user enters a search query into the device, for example, the keyword "coronavirus conspiracy theory."
[0499] Step 2:
[0500] The terminal transmits the input search query to the server.
[0501] Step 3:
[0502] The server analyzes the received search query and extracts related keywords, for example, "coronavirus" and "conspiracy theory."
[0503] Step 4:
[0504] The server searches an index database based on the extracted keywords to retrieve a list of related websites and videos, such as a website titled "Is COVID-19 a man-made virus?" and a video titled "The Truth About Conspiracy Theories."
[0505] Step 5:
[0506] The server accesses official databases to retrieve the latest data for fact-checking, such as official statements and documents from the WHO.
[0507] Step 6:
[0508] The server scrapes the content of each website and video and collects it as text data. For example, it extracts the content of an article on a website titled "Is Corona a Man-Made Virus?"
[0509] Step 7:
[0510] The server processes the collected text data with a text analysis algorithm to extract key keywords and claims, such as "population virus," "evidence," and "research paper."
[0511] Step 8:
[0512] The server compares the extracted keywords and claims against public fact-checking databases, for example, comparing the claim "man-made virus" with the WHO statement "coronavirus is naturally occurring."
[0513] Step 9:
[0514] The server calculates a confidence score for each search result item based on the comparison, for example, assigning a low confidence score to the "man-made virus" claim because it contradicts official data.
[0515] Step 10:
[0516] The server generates a search result list by tagging each search result with a confidence score. For example, a website asking "Is COVID-19 a man-made virus?" might be tagged with "Low confidence."
[0517] Step 11:
[0518] The server sends the search result list with the confidence scores attached to it to the terminal.
[0519] Step 12:
[0520] The terminal displays the received search result list to the user. The user checks the displayed search results and judges the reliability of the information based on the reliability score. For example, the user may preferentially view items marked "High Reliability."
[0521] This series of processes allows users to easily distinguish highly reliable information and makes decisions based on accurate information.
[0522] Example 1
[0523] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0524] In the current Internet environment, unreliable and false information is often spread, making it difficult for users to find accurate information, especially when it comes to information about serious issues. For this reason, there is a need for a system that allows users to quickly and easily obtain reliable information. In addition, there is a need for a method that effectively utilizes data provided by public institutions and objectively evaluates the reliability of information.
[0525] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0526] In this invention, the server includes means for receiving a search query from a user terminal, means for analyzing the received search query to extract related keywords, means for obtaining a list of related websites and videos based on the extracted keywords, means for obtaining fact-checking data from a database provided by a public institution, means for scraping the obtained content and collecting it as text data, means for analyzing the collected text data using a text analysis algorithm to extract important keywords and claims, means for comparing the extracted keywords and claims with the fact-checking data, and means for calculating and displaying a reliability score for each search result item based on the comparison result, thereby enabling users to quickly and accurately obtain reliable information.
[0527] "User terminal" means a physical device through which a user enters a search query, including a personal computer, smartphone, tablet, etc.
[0528] A "search query" is text data entered by a user to search for specific information.
[0529] A "search server" is a centralized computer system that receives, analyzes, and processes search queries from user terminals.
[0530] "Keywords" are related words or phrases extracted from a search query.
[0531] An "index database" is a database in which information about related websites, videos, etc. is organized and stored in a searchable format.
[0532] "Fact-checking data" is data obtained from databases provided by public institutions to assess the reliability of information, including official statements and research papers.
[0533] "Scraping" is the technique of extracting information from web pages using automated programs.
[0534] A "text analysis algorithm" is a computational method for analyzing text data to extract important information and patterns.
[0535] The "trust score" is a numerical evaluation of the reliability of the information in the search results, and is a value that users can use as a reference when judging the information.
[0536] "Comparison" is the process of comparing extracted keywords and claims with fact-checking data and analyzing for matches and contradictions.
[0537] The present invention is a search system that performs fact-checking based on reliable data provided by public institutions. To implement this system, a user terminal, a search server, and a fact-checking database are used as main components.
[0538] System Configuration
[0539] User Device
[0540] A device on which a user enters a search query and sends it to a search server. User terminals can be personal computers, smartphones, tablets, etc. A web browser is installed on these terminals, and users enter and send search queries through a web interface.
[0541] Search Server
[0542] The search server is the central component that analyzes search queries received from user devices and retrieves and processes relevant information. It is recommended to use a high-performance cloud server (e.g., one of the common cloud server services) for the search server. The following software and algorithms are required:
[0543] Natural Language Processing (NLP) libraries (e.g., NLP libraries commonly used in technical literature)
[0544] Scraping tools (e.g., common open source libraries)
[0545] Database management systems (e.g., commonly used open-source database management systems)
[0546] Fact-checking database
[0547] Fact-checking databases store reliable data provided by public institutions. Specifically, frequently updated data is retrieved from accessible data sources such as international health organizations and government statistical databases, and cached locally within the search server.
[0548] Example of operation
[0549] 1. A user enters a search query, for example, "coronavirus conspiracy theory," into the search bar on their device and presses the send button.
[0550] 2. The search server analyzes the received search query and extracts the related keywords "coronavirus" and "conspiracy theory."
[0551] 3. The search server searches the index database based on the extracted keywords and retrieves relevant search results (websites and videos).
[0552] 4. The search server also retrieves the latest fact-checking data from official databases and stores it in a local cache.
[0553] 5. The search server scrapes the content of the search results and collects it as text data, such as the main text of website articles and video descriptions.
[0554] 6. The search server uses text analysis algorithms to extract important keywords and assertions from the collected text data.
[0555] 7. The search server compares the extracted keywords and claims with fact-checking data, for example, comparing the claim "man-made virus" with the WHO statement "coronavirus is naturally occurring" to identify inconsistencies.
[0556] 8. The search server calculates a confidence score for each search result and displays it as a tag on each search result item.
[0557] Examples of prompt statements
[0558] "What are the steps to verify the credibility of conspiracy theories related to COVID-19?"
[0559] "Please explain specifically how you will conduct fact-checking."
[0560] This system allows users to obtain reliable information quickly and accurately, minimizing the impact of uncertain or false information.
[0561] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0562] Step 1:
[0563] A user enters a search query, such as "COVID-19 conspiracy theory," into the device's search bar and clicks the submit button. The device then sends the search query to a search server. The input data is free-form text, and the output is an HTTP request to the server.
[0564] Step 2:
[0565] The search server analyzes the received search query and extracts relevant keywords. Specifically, it uses the analysis module of a natural language processing (NLP) library (e.g., spaCy) to extract important keywords such as "coronavirus" and "conspiracy theory." The input is free-form text data, and the output is a list of extracted keywords.
[0566] Step 3:
[0567] The search server searches an index database based on the extracted keywords to obtain a list of related websites and videos. The database used for this process is one that has been pre-indexed by a web crawling tool. The input is a list of keywords, and the output is a list of related URLs and titles. Specifically, it includes a website titled "Is Corona a Man-Made Virus?" and a video titled "The Truth About Conspiracy Theories."
[0568] Step 4:
[0569] The search server makes HTTP requests to official databases (e.g., international health organizations or government agencies) to retrieve the latest data for fact-checking. The retrieved data is stored in a local cache for use in the next step. The input is the official API endpoint, and the output is a dataset of trusted documents and statements.
[0570] Step 5:
[0571] The search server uses a scraping tool (e.g., BeautifulSoup) to collect the content of the retrieved websites and videos. Specifically, it analyzes the HTML code of the webpage and extracts the article text and metadata as text data. The input is a list of URLs, and the output is content data in text format. For example, it collects the content of articles on a website titled "Is Corona a Man-Made Virus?"
[0572] Step 6:
[0573] The search server analyzes the collected text data using natural language processing (NLP) algorithms to extract important keywords and claims. Specifically, it uses an NLP library to extract frequently occurring words and phrases in the text and identify important keywords such as "population virus," "evidence," and "research paper." The input is text content data, and the output is a list of important keywords.
[0574] Step 7:
[0575] The search server uses the extracted keywords and claims to compare them with fact-checking data obtained from official databases. For example, it compares the claim "man-made virus" with the WHO statement "coronavirus is naturally occurring" and analyzes contradictions and similarities. The input is a list of keywords and reliable document data, and the output is the comparison results and a reliability score.
[0576] Step 8:
[0577] The search server tags each search result item with the calculated confidence score and generates a list of search results to display to the user. Specifically, the process involves adding the confidence score to the title and URL of each search result to create a list. The input is the comparison result and confidence score, and the output is a tagged list of search results.
[0578] Step 9:
[0579] The search server sends the tagged search result list to the terminal, which receives it and generates an HTML page to display to the user. The input is the tagged search result list, and the output is the search result page displayed to the user.
[0580] Step 10:
[0581] Users check the displayed search results and judge the reliability of the information based on the reliability score. Specifically, they prioritize browsing items with high reliability and acquire or share the information. The input is the displayed search result page, and the output is the user's browsing actions and judgment results.
[0582] This allows users to obtain reliable information quickly and accurately.
[0583] (Application example 1)
[0584] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0585] In recent years, the amount of information on the Internet has increased dramatically, including information of unknown authenticity and potentially misleading information. This has made it difficult for users to easily obtain reliable information. Furthermore, in virtual stores, it is difficult for consumers to judge the reliability of product information themselves, which risks leading to inaccurate purchasing decisions. Therefore, there is a need for a system that can automatically determine the reliability of information and present a reliability score when consumers search for product information in virtual stores.
[0586] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0587] In this invention, the server includes means for acquiring fact-checking data from a database provided by a public institution, means for receiving a search query from a user, means for acquiring related search results based on the search query, means for extracting important keywords and assertions by text-analyzing the content of the search results, means for comparing the extracted keywords and assertions with the fact-checking data, means for calculating and displaying a reliability score for each search result item based on the comparison result, and means for tagging product information with the reliability score, allowing users to judge reliability when making purchasing decisions. This allows consumers to easily judge the reliability of product information in virtual stores and make purchasing decisions with confidence.
[0588] "Public institution" refers to a public institution, such as a government or local government, that provides public functions or services.
[0589] A "database" is an information repository that organizes and stores a collection of data so that it can be efficiently searched and used.
[0590] "Fact-checking data" refers to reliable data used to verify the veracity of information.
[0591] "User" means any individual person or entity that uses the System or Services.
[0592] A "search query" refers to a question or keyword that a user enters into a search engine or system.
[0593] "Search Results" means the information returned by a search engine or system based on a search query.
[0594] "Content" refers to any piece of information or data present on a website or digital media.
[0595] "Text analytics" refers to the techniques and processes used to analyze text data and understand its content.
[0596] "Keywords" refer to key words that play an important role in a text or search query.
[0597] An "argument" is a statement of opinion or perspective on a certain matter.
[0598] "Comparison" refers to the act of comparing two or more things to see their differences and similarities.
[0599] A "trust score" is a numerical indicator that evaluates the reliability of information or data.
[0600] "Display" refers to the means by which information is visually conveyed to the user.
[0601] "Tagging" refers to the act of adding labels and scores to information.
[0602] "Product information" refers to data including descriptions, features, specifications, etc. of the target product.
[0603] "Consumption behavior" refers to the series of actions that consumers take to purchase goods and services.
[0604] This invention describes a system that automatically determines the reliability of product information when a user searches for it in a virtual store and presents a reliability score. This system is composed of the following main components:
[0605] System configuration
[0606] The system is implemented primarily using the following hardware and software:
[0607] 1. User terminal: a smartphone, smart glasses, head-mounted display, or other display device.
[0608] 2. Search server: A server responsible for data analysis and fact-checking.
[0609] 3. Fact-checking database: A database that stores reliable data provided by public institutions.
[0610] Program processing
[0611] User Roles
[0612] A user inputs a search query, such as "antioxidant supplements," into a user terminal and presses a search button. This operation sends the query to a search server.
[0613] The role of the search server
[0614] The search server analyzes the search query received from the user terminal and extracts related keywords, which allows the system to obtain keywords such as "antioxidant" and "supplement."
[0615] Furthermore, the search server searches the virtual store database based on the extracted keywords to obtain a list of related products. For example, product information such as "Antioxidant Supplement A" and "Supplement B" is obtained.
[0616] The search server then accesses databases provided by public organizations (e.g., government agencies and research institutes) to retrieve reliable information and retrieve the latest data for fact-checking, including official reports and the latest research papers.
[0617] Content analysis and fact-checking
[0618] The search server scrapes the descriptions and reviews of each product and collects them as text data. For example, it analyzes the description of "Antioxidant Supplement A" and extracts important keywords and claims.
[0619] Furthermore, the search server compares the extracted keywords and claims with information from databases of public institutions, and calculates a credibility score to evaluate the reliability of the product. For example, a credibility score for a product is determined by comparing the claim "has antioxidant properties."
[0620] Visibility and User Roles
[0621] Finally, the search server tags each product with the calculated reliability score and displays it as information to help users judge reliability when deciding on their purchasing behavior, allowing users to select products with confidence based on reliable information.
[0622] Examples and prompts
[0623] For example, if a user searches for "antioxidant supplements," the search server retrieves related product information, analyzes the descriptions of each product, and calculates a reliability score. As a result, products with a "High Reliability" rating can be presented to the user preferentially.
[0624] Here are some examples of prompts:
[0625] "Antioxidant supplement A has been proven effective in clinical trials. Please compare this information with data from official institutions and calculate a reliability score."
[0626] In this way, this system provides an environment in which consumers can easily judge the reliability of product information in a virtual store.
[0627] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0628] Step 1: Receiving User Input
[0629] A user enters a search query such as "antioxidant supplements" into a device such as a smartphone or head-mounted display and presses the search button. The device then sends this query to a search server. The input data is the user's search query, and the output data is the query information sent to the search server.
[0630] Step 2: Processing the search query
[0631] The search server analyzes the search query received from the device and extracts the related keywords "antioxidant" and "supplement." Specifically, it uses a natural language processing algorithm to tokenize the search query and extract important keywords. The input data is the received query information, and the output data is the extracted keywords.
[0632] Step 3: Getting search results
[0633] The search server searches the virtual store database based on the extracted keywords and retrieves a list of related products. For example, product information such as "antioxidant supplement A" and "supplement B" is retrieved. The input data are the extracted keywords, and the output data is a list of products.
[0634] Step 4: Obtaining fact-checking data
[0635] The search server accesses public databases to retrieve the latest data for fact-checking, including official reports from government agencies and research institutes and the latest research papers. The input data is the access request to the public database, and the output data is the retrieved data for fact-checking.
[0636] Step 5: Analyze your content
[0637] The search server scrapes the descriptions and reviews of each product and collects this content as text data. Specifically, it extracts the text of articles and reviews from each page. The input data is a list of products, and the output data is the collected text data.
[0638] Step 6: Text analysis and keyword extraction
[0639] The search server processes the collected text data using a text analysis algorithm to extract key keywords and claims. Specifically, it uses TF-IDF (Term Frequency-Inverse Document Frequency) to identify important keywords. The input data is the collected text data, and the output data is the extracted keywords and claims.
[0640] Step 7: Conduct a fact check
[0641] The search server compares the extracted keywords and claims with the fact-checking data, checking for keyword matches and inconsistencies, and calculating a reliability score for each product. The input data are the extracted keywords and claims and the fact-checking data, and the output data is a reliability score for each product.
[0642] Step 8: Score tagging and search result generation
[0643] The search server tags each product information with the calculated confidence score and generates a search result list to display to the user. Specifically, it assigns a confidence score to each product information and converts it into a format that can be displayed in the user interface. The input data are the confidence scores and the product list, and the output data is the tagged search result list.
[0644] Step 9: Submit search results
[0645] The search server sends the tagged search results to the terminal, which then displays the received search results to the user. The input data is the tagged search result list, and the output data is the displayed search results.
[0646] Step 10: User View and Decision
[0647] The user checks the displayed search results and judges the reliability of the product information based on the reliability score. Specifically, the user may take action such as prioritizing the purchase of products displayed as "High reliability." The input data are the displayed search results, and the output data are the user's purchasing behavior.
[0648] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0649] This invention combines a search system that performs fact-checking based on highly reliable data provided by public institutions with an emotion engine that recognizes user emotions. The programs required to implement this system and their processing are shown below.
[0650] System configuration
[0651] The system mainly consists of the following components:
[0652] 1. User Device
[0653] 2. Search Server
[0654] 3. Fact-checking databases
[0655] 4. Emotion Engine
[0656] Program processing
[0657] 1. Receiving user input
[0658] The user enters a search query, such as "coronavirus conspiracy theory," into the device and presses the send button.
[0659] 2. Processing search queries
[0660] The terminal transmits the input search query to the server.
[0661] 3. Search Query Analysis
[0662] The server analyzes the received search query and extracts related keywords, for example, "coronavirus" and "conspiracy theory."
[0663] 4. Obtaining search results
[0664] The server searches an index database based on the extracted keywords to retrieve a list of related websites and videos, such as a website titled "Is COVID-19 a man-made virus?" and a video titled "The Truth About Conspiracy Theories."
[0665] 5. Acquisition of data from public institutions
[0666] The server accesses official databases to retrieve the latest data for fact-checking, such as official statements and documents from the WHO.
[0667] 6. Content Analysis
[0668] The server scrapes the content of each website and video and collects it as text data. For example, it extracts the content of an article on a website titled "Is Corona a Man-Made Virus?"
[0669] 7. Text Analysis and Keyword Extraction
[0670] The server processes the collected text data with a text analysis algorithm to extract key keywords and claims, such as "population virus," "evidence," and "research paper."
[0671] 8. Fact Check
[0672] The server compares the extracted keywords and claims against public fact-checking databases—for example, comparing the claim "man-made virus" with the WHO statement "coronavirus is naturally occurring"—and calculates a confidence score for each search result.
[0673] 9. Emotion Recognition with Emotion Engine
[0674] The device uses an emotion engine to analyze the user's emotions based on the user's input and operations (e.g., keystroke speed, input content, facial recognition, etc.). For example, it determines whether the user is feeling stressed.
[0675] 10. Emotional Feedback
[0676] The server then adjusts the reliability score and display order of search results based on the analysis results of the emotion engine. For example, if a user is feeling stressed, more reliable information will be displayed first.
[0677] 11. Score tagging and search result generation
[0678] The server tags each search result item with a confidence score and generates a search result list to display to the user. For example, a site asking "Is COVID-19 a man-made virus?" might be given a "Low confidence" rating.
[0679] 12. Submitting Search Results
[0680] The server sends the search result list with the confidence scores attached to it to the terminal.
[0681] 13. Display of search results
[0682] The device displays the received search results to the user. The user checks the displayed search results and determines the reliability of the information based on the reliability score. For example, the user may prioritize viewing items marked "High Reliability."
[0683] Specific examples
[0684] For example, if a user searches for "COVID-19 conspiracy theory," the search server retrieves related websites and videos and analyzes their content. It then compares them with documents from official institutions and calculates a reliability score. The device also uses an emotion engine to analyze the user's emotions (e.g., anxiety, stress) and adjusts the display order of search results based on the results. As a result, users can easily select reliable information and browse with peace of mind.
[0685] In this way, the system of the present invention provides an environment in which the user can easily obtain highly reliable information, and further provides information according to the user's emotional state.
[0686] The processing flow will be explained below.
[0687] Step 1:
[0688] The user enters a search query into the device, for example, the keyword "coronavirus conspiracy theory."
[0689] Step 2:
[0690] The terminal transmits the input search query to the server.
[0691] Step 3:
[0692] The server analyzes the received search query and extracts related keywords, for example, "coronavirus" and "conspiracy theory."
[0693] Step 4:
[0694] The server searches an index database based on the extracted keywords to retrieve a list of related websites and videos, such as a website titled "Is COVID-19 a man-made virus?" and a video titled "The Truth About Conspiracy Theories."
[0695] Step 5:
[0696] The server accesses official databases to retrieve the latest data for fact-checking, such as official statements and documents from the WHO.
[0697] Step 6:
[0698] The server scrapes the content of each website and video and collects it as text data. For example, it extracts the content of an article on a website titled "Is Corona a Man-Made Virus?"
[0699] Step 7:
[0700] The server processes the collected text data with a text analysis algorithm to extract key keywords and claims, such as "population virus," "evidence," and "research paper."
[0701] Step 8:
[0702] The server compares the extracted keywords and claims against public fact-checking databases, for example, comparing the claim "man-made virus" with the WHO statement "coronavirus is naturally occurring."
[0703] Step 9:
[0704] The server calculates a confidence score for each search result item based on the comparison, for example, assigning a low confidence score to the "man-made virus" claim because it contradicts official data.
[0705] Step 10:
[0706] The device uses an emotion engine to analyze the user's emotions based on the user's input and operations (e.g., keystroke speed, input content, facial recognition, etc.). For example, it can determine whether the user is feeling stressed.
[0707] Step 11:
[0708] The server then adjusts the reliability score and display order of search results based on the analysis results of the emotion engine. For example, if a user is feeling stressed, more reliable information will be displayed first.
[0709] Step 12:
[0710] The server tags each search result item with a confidence score and generates a search result list to display to the user. For example, a site asking "Is COVID-19 a man-made virus?" might be given a "Low confidence" rating.
[0711] Step 13:
[0712] The server sends the search result list with the confidence scores attached to it to the terminal.
[0713] Step 14:
[0714] The terminal displays the received search result list to the user. The user checks the displayed search results and judges the reliability of the information based on the reliability score. For example, the user may preferentially view items marked "High Reliability."
[0715] This series of processes allows the user to easily distinguish highly reliable information, and also makes it possible to provide information that is appropriate for the user's emotional state.
[0716] Example 2
[0717] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0718] In today's world, there is a huge amount of information available on the Internet, but it is not easy to find accurate and reliable information. Furthermore, providing information without considering the user's emotional state can increase stress and anxiety. Therefore, there is a need for a search system that takes into account the user's emotional state while ensuring the reliability of the information.
[0719] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0720] In this invention, the server includes means for acquiring fact-checking data from a database provided by a public institution, means for receiving a search query from a user, means for acquiring related search results based on the search query, means for analyzing the content of the search results to extract important keywords and claims, means for comparing the extracted keywords and claims with the fact-checking data, means for calculating and displaying a reliability score for each search result item based on the comparison result, means for analyzing user emotions, and means for adjusting the display order of the search results based on the emotion analysis results.
[0721] This allows the user to easily obtain highly reliable information and to be provided with information that corresponds to the user's own emotional state.
[0722] A "database" is a storage device that stores highly reliable data provided by public institutions.
[0723] "Fact-checking data" refers to documents and statements released by public institutions, and is basic data used to verify the accuracy of information.
[0724] A "search query" is a question or keyword that a user enters into the system, and is the search condition that the system investigates.
[0725] "Search results" are lists of information retrieved from databases or the Internet based on a search query.
[0726] "Text analysis" is a technique that uses natural language processing technology to understand the content of text data and extract important keywords and assertions.
[0727] "Keywords" are important words or phrases extracted from a search query or analyzed text.
[0728] "Arguments" are important opinions or points of view in a text that are revealed through text analysis.
[0729] "Comparison" is the process of comparing the fact-checking data with the extracted keywords and claims to confirm whether they match or disagree.
[0730] The "trust score" is a numerical representation of the reliability of information based on the results of fact-checking, and is an indicator of the reliability of search results.
[0731] "Emotion analysis" is the process of analyzing a user's emotional state based on their input and actions, and is a technology that determines what emotions the user is feeling.
[0732] "Display ranking" refers to the order in which each item is displayed in the search result list, and is determined based on confidence scores and sentiment analysis results.
[0733] This invention combines a search system that performs fact-checking based on reliable data provided by public institutions with an emotion engine that recognizes user emotions. This system mainly consists of the following components: a user terminal, a search server, a fact-checking database, and an emotion engine.
[0734] System configuration
[0735] To implement this invention, the following hardware and software are used:
[0736] 1. User Device
[0737] Hardware: Personal computers, smartphones
[0738] Software: Web browser, input interface
[0739] 2. Search Server
[0740] Hardware: High-performance server
[0741] Software: Natural language processing libraries (e.g., NLTK, SpaCy, BERT), search engines (e.g., Elasticsearch, Solr), scraping tools (e.g., BeautifulSoup, Selenium), sentiment analysis APIs (e.g., Microsoft Azure Cognitive Services, Emotion API)
[0742] 3. Fact-checking databases
[0743] Hardware: Database Server
[0744] Software: Public Sector API Access
[0745] 4. Emotion Engine
[0746] Hardware: Embedded systems, cameras and keyboards attached to user devices
[0747] Software: Emotion analysis algorithms, face tracking software (e.g. OpenCV)
[0748] Implementation Procedure
[0749] 1. Receiving user input
[0750] The user enters "COVID-19 conspiracy theory" into the search field and clicks the search button. The device receives this input, temporarily stores it in its internal memory, and prepares to send it to the server.
[0751] 2. Processing search queries
[0752] The device sends the search query entered by the user to the server as an HTTP request, which includes the query text data and metadata such as the user ID.
[0753] 3. Search Query Analysis
[0754] The server parses the incoming HTTP request and processes the query text with a natural language processing algorithm (e.g., Python's NLTK library), which identifies the main keywords "coronavirus" and "conspiracy theory."
[0755] 4. Obtaining search results
[0756] The server queries an index database using a search engine (e.g., Elasticsearch or Solr) based on the extracted keywords, which retrieves a list of related websites and videos.
[0757] 5. Acquisition of data from public institutions
[0758] The server sends API requests to public databases to retrieve the latest, most reliable information, including official statements and research papers from organizations like the WHO and CDC.
[0759] 6. Content Analysis
[0760] The server uses scraping technology (e.g., BeautifulSoup or Selenium) to collect the content of each site and video as text data, extracting the content of the article on the site "Is Corona a Man-Made Virus?" and the subtitles of the video.
[0761] 7. Text Analysis and Keyword Extraction
[0762] The server then processes the collected text data again using a natural language processing algorithm (e.g., SpaCy or BERT) to extract important keywords and assertions, such as "population virus," "evidence," and "research paper."
[0763] 8. Fact Check
[0764] The server compares the extracted keywords and claims with official data it has obtained—for example, comparing the claim "man-made virus" with the WHO statement "coronavirus is naturally occurring"—and then uses an algorithm to calculate a confidence score for each search result, which is calculated on a scale from 0 to 100.
[0765] 9. Emotion Recognition with Emotion Engine
[0766] The device uses an emotion engine to analyze the user's emotions based on the user's input and operations, such as keystroke speed and facial recognition (using OpenCV), to determine whether the user is feeling stressed or anxious.
[0767] 10. Emotional Feedback
[0768] The server then adjusts the reliability score and display order of search results based on the analysis results of the emotion engine. For example, if a user is feeling stressed, an algorithm is applied to prioritize displaying information with high reliability.
[0769] 11. Score tagging and search result generation
[0770] The server tags each search result with a confidence score and generates a list of search results to display to the user. For example, a website asking "Is COVID-19 a man-made virus?" might be tagged with "Confidence: 20 (Low)."
[0771] 12. Submission and Display of Search Results
[0772] The server sends the search result list, with assigned reliability scores, to the terminal as an HTTP response. The terminal displays the received search result list on its user interface. The user checks this display and decides which information to view based on the reliability score. For example, they may prioritize clicking and viewing items with a "Reliability: 90" rating.
[0773] Specific examples
[0774] For example, if a user searches for "coronavirus conspiracy theory," the following specific actions occur: The user enters "coronavirus conspiracy theory" into the search bar and clicks the search button. The device receives the input, generates an HTTP request, and sends it to the server. The server analyzes the query and extracts related keywords. The server uses a search engine to retrieve relevant sites and videos. The server retrieves reliable data from public institution databases via an API request. The server scrapes the content of the retrieved sites and videos and collects it as text data. The server uses a natural language processing algorithm to extract important keywords and claims from the collected text data. The server compares the extracted keywords and claims with public institution data and calculates a confidence score. The device analyzes the user's emotions using an emotion engine and sends the results to the server. The server adjusts the confidence score and display ranking of the search results based on the user's emotions. The server tags the confidence score and generates a search result list. The server sends the final search result list to the device. The device displays the search result list on a user interface, and the user can review and select results.
[0775] Example prompt sentence:
[0776] "New coronavirus conspiracy theory"
[0777] Result: A list of highly reliable information is displayed, including an official statement from a public institution that "coronavirus is a natural occurrence." Because the emotion engine determined that the user was highly stressed, highly reliable information is displayed first.
[0778] In this way, the system of the present invention provides an environment in which the user can easily obtain highly reliable information, and further provides information according to the user's emotional state.
[0779] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0780] Step 1:
[0781] The user enters "COVID-19 conspiracy theory" in the search field and clicks the search button. This causes the device to receive a search query. The input is the search query entered by the user, and the output is the query temporarily stored in the internal memory. The device prepares to send this data to the server.
[0782] Step 2:
[0783] The device sends the search query entered by the user to the server as an HTTP request. The input is the search query stored in the internal memory, and the output is the HTTP request sent to the server. The request includes the text data of the query and metadata such as the user ID.
[0784] Step 3:
[0785] The server parses the received HTTP request and processes the query text with an NLP algorithm (e.g., Python's NLTK library). The input is the HTTP request sent to the server, and the output is the parsed search query and keywords extracted from it, e.g., "coronavirus" and "conspiracy theory."
[0786] Step 4:
[0787] The server queries the index database using Elasticsearch or Solr based on the extracted keywords. The input is the extracted keywords, and the output is a list of related websites and videos, such as a website called "Is COVID-19 a man-made virus?" or a video called "The Truth About Conspiracy Theories."
[0788] Step 5:
[0789] The server sends API requests to public databases to retrieve the latest, most reliable information. The input is the API request, and the output is the retrieved public data, such as an official WHO statement or a research paper.
[0790] Step 6:
[0791] The server collects the content of each website and video as text data using scraping technology (e.g., BeautifulSoup or Selenium). The input is the URL of the relevant website or video, and the output is the scraped text data, such as the content of an article on the website "Is Coronavirus a Man-Made Virus?"
[0792] Step 7:
[0793] The server processes the collected text data again with an NLP algorithm (e.g., SpaCy or BERT) to extract important keywords and claims. The input is the collected text data, and the output is the extracted important keywords and claims, such as "population virus," "evidence," and "research paper."
[0794] Step 8:
[0795] The server compares the extracted keywords and claims with the retrieved official data. The input is the extracted keywords and claims and the official data, and the output is a confidence score for each search result, e.g., "Confidence: 20 (Low)."
[0796] Step 9:
[0797] The device uses an emotion engine to analyze the user's emotions based on their input and operations. Specifically, it uses keystroke speed and facial recognition technology (e.g., OpenCV). The input is the user's input data and operation data, and the output is the analyzed user's emotional state, such as "high stress."
[0798] Step 10:
[0799] The server adjusts the reliability score and display order of search results based on the analysis results of the emotion engine. The input is the emotion analysis result and reliability score, and the output is an adjusted search result list, for example, a list in which highly reliable information is displayed at the top.
[0800] Step 11:
[0801] The server tags each search result with a confidence score and generates a final search result list to display to the user. The input is the adjusted search result list, and the output is a search result list with a confidence score, e.g., "Confidence: 90 (High)".
[0802] Step 12:
[0803] The server sends the final search result list to the terminal. The input is the search result list with the assigned confidence scores, and the output is the HTTP response sent to the terminal.
[0804] Step 13:
[0805] The terminal displays the received search result list on the user interface. The input is the HTTP response sent to the terminal, and the output is the search result list displayed on the screen. The user checks the results and decides which information to view based on the confidence score.
[0806] (Application example 2)
[0807] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0808] In today's information society, it is extremely important for users to have access to reliable information. However, while a vast amount of information is easily available, it is often difficult to verify its reliability. Furthermore, there are no established methods for providing appropriate information to users who are feeling stressed or anxious. Therefore, an objective of the present invention is to provide a system that provides users with reliable information and also provides information that is tailored to the user's emotional state.
[0809] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for acquiring fact-checking data from a database provided by a public institution, means for receiving a search query from a user, means for acquiring related search results based on the search query, means for text-analyzing the content of the search results to extract important keywords and assertions, means for comparing the extracted keywords and assertions with the fact-checking data, means for calculating and displaying a reliability score for each search result item based on the comparison result, means for recognizing passenger emotions using an emotion engine, and means for adjusting the reliability score and information display based on the recognized emotions. This allows the user to appropriately acquire reliable information and provides information according to the user's emotional state.
[0810] A "public institution" is an organization that operates for the public interest, such as a government or local government, and issues official data and statements.
[0811] A "database" is a collection of information designed to efficiently store, retrieve, and update data.
[0812] "Fact-checking data" is reliable information and materials collected and provided for the purpose of fact-checking.
[0813] "User" means an individual or organization that uses a system or application.
[0814] A "search query" is a string of characters or keywords that a user enters when searching for information.
[0815] "Search results" are a collection of information or data retrieved based on a search query.
[0816] "Text analytics" refers to techniques and methods for extracting important information and patterns from text data.
[0817] "Keywords" are important words or phrases used to represent specific information.
[0818] An "assertion" is an expression that expresses a particular opinion or point of view.
[0819] "Comparison" is the act of comparing two or more pieces of information or data and evaluating their differences and similarities.
[0820] A "trust score" is a numerical indicator used to evaluate the reliability of information.
[0821] "Display" is the act of visually presenting acquired information or results to the user.
[0822] An "emotion engine" is software or hardware that analyzes and recognizes a user's emotions.
[0823] "Recognition" is the act of analyzing, understanding, and grasping information and data.
[0824] "Adjustment" is the act of changing information or settings to accommodate specific conditions or circumstances.
[0825] "Means" are the methods or techniques used to achieve a particular goal.
[0826] The present invention is a system for providing highly reliable information to a user and providing information according to the user's emotional state. Specific embodiments of the present invention will be described below.
[0827] System Configuration
[0828] Hardware
[0829] 1. User device: A device such as a smartphone, tablet, or PC that accepts user input and operations. It also has a camera and microphone to collect data for emotion recognition.
[0830] 2. Search Server: A central server that searches, analyzes, and fact-checks various data.
[0831] 3. Fact-checking database: A database that stores reliable data provided by public institutions.
[0832] 4. Emotion engine: Software for analyzing the user's emotional state.
[0833] software
[0834] 1. EmotionEngine: Software that analyzes data input from cameras and microphones in real time to recognize the user's emotions. It uses image processing libraries such as OpenCV.
[0835] 2. FactCheckEngine: An engine that cross-checks the information obtained with public data and evaluates its reliability.
[0836] 3. PublicDataFetcher: Software for retrieving up-to-date and reliable data from public databases.
[0837] System Operation
[0838] The server first receives a search query from a user device, then retrieves relevant search results based on the search query, performs text analysis on the content, extracts important keywords and claims, and compares the extracted keywords and claims with fact-checking data to calculate their reliability score.
[0839] At the same time, EmotionEngine recognizes the user's emotions in real time using the camera and microphone installed on the user's device. Based on the recognized emotions, the server adjusts the reliability score and information display to provide the user with appropriate information.
[0840] Specific examples
[0841] For example, if a user enters the search query "traffic congestion current situation," the server retrieves related websites and news articles, performs fact-checking, and calculates and displays a reliability score. If the emotion engine recognizes that the user is feeling stressed, it prioritizes information with high reliability and suggests relaxing music and videos.
[0842] Prompt Sentence Examples
[0843] Provide information to promote safe driving when users are under stress, such as prioritizing public traffic and weather information and playing relaxing music and scenic videos.
[0844] In this way, a system is realized that allows users to easily obtain highly reliable information and receive information that is optimal for their emotional state.
[0845] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0846] Step 1:
[0847] User enters search query and submits
[0848] A user inputs a query, for example, "traffic congestion current situation," into the terminal and presses the send button. After the input is made, the terminal sends the search query to the server. The input is in the form of text, and the output is the search query sent to the server.
[0849] Step 2:
[0850] The server parses the search query
[0851] The server analyzes the received search query and extracts relevant keywords. For example, it extracts the keywords "traffic congestion" and "current situation." The input is the search query text data, and the output is the analyzed keyword list. The server performs this process using a text analysis algorithm.
[0852] Step 3:
[0853] The server retrieves the search results
[0854] The server searches the index database based on the extracted keywords to obtain a list of relevant sites and news articles. The input is the parsed keyword list, and the output is a list of relevant search results. The server uses the index database to perform this process.
[0855] Step 4:
[0856] Server acquires data from public institutions
[0857] The server accesses a public database to retrieve the latest data for fact-checking, for example, data from a traffic information center. The input is the URL or API key of the public database to access, and the output is the trusted data. The server performs this process using an HTTP request.
[0858] Step 5:
[0859] The server analyzes the content
[0860] The server scrapes the content of each website or news article and collects it as text data. The input is a list of relevant search results, and the output is the retrieved text data. The server performs this process using web scraping technology.
[0861] Step 6:
[0862] The server performs text analysis and keyword extraction.
[0863] The server processes the collected text data with text analysis algorithms to extract important keywords and assertions. The input is the collected text data, and the output is the extracted keywords and assertions. The server performs this processing using natural language processing (NLP) techniques.
[0864] Step 7:
[0865] Server fact-checking
[0866] The server compares the extracted keywords and claims against public fact-checking databases and calculates a confidence score. The inputs are the extracted keywords and data from the public databases, and the output is a confidence score. The server performs this process using database queries and comparison algorithms.
[0867] Step 8:
[0868] The device uses an emotion engine to recognize emotions.
[0869] The device analyzes the user's emotions using a camera and microphone. The input is video and audio data, and the output is the recognized emotion data. The device performs this processing using the EmotionEngine.
[0870] Step 9:
[0871] Server provides feedback based on emotions
[0872] The server adjusts the confidence score and information display based on the analysis results of the emotion engine. The input is the recognized emotion data and confidence score, and the output is the adjusted information display ranking and feedback. The server performs this process using an algorithm.
[0873] Step 10:
[0874] The server generates a search result list tagged with a confidence score
[0875] The server tags each search result item with a confidence score and generates a search result list for display to the user. The input is the confidence score and the list of search results, and the output is the tagged search result list. The server performs this process using a list generation algorithm.
[0876] Step 11:
[0877] The server sends the search result list
[0878] The server sends the search result list with attached confidence scores to the user device. The input is the tagged search result list, and the output is the data sent to the user device. The server performs this process using a network protocol.
[0879] Step 12:
[0880] Your device will display the search results
[0881] The terminal displays the received search results to the user. The user reviews the displayed search results and determines the reliability of the information based on the reliability score. The input is the search result list received from the server, and the output is the displayed search results. The terminal performs this process using a graphical user interface (GUI).
[0882] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0883] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0884] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.
[0885] [Third embodiment]
[0886] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.
[0887] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.
[0888] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0889] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.
[0890] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0891] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0892] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0893] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0894] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0895] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0896] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0897] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."
[0898] The present invention is a search system that performs fact-checking based on highly reliable data provided by public institutions. The programs required to implement this system and their processing are shown below.
[0899] System configuration
[0900] The system mainly consists of the following components:
[0901] 1. User Device
[0902] 2. Search Server
[0903] 3. Fact-checking databases
[0904] Program processing
[0905] 1. Receiving user input
[0906] The user enters a search query, such as "coronavirus conspiracy theory," into the device and presses the send button.
[0907] 2. Processing search queries
[0908] The search server analyzes the search query received from the device and extracts related keywords. Through this analysis, the system obtains keywords such as "coronavirus" and "conspiracy theory."
[0909] 3. Obtaining search results
[0910] The search server searches the index database based on the extracted keywords to retrieve a list of related websites and videos. For example, it retrieves a website titled "Is Corona a Man-Made Virus?" and a video titled "The Truth About Conspiracy Theories."
[0911] 4. Obtaining data for fact-checking
[0912] The search server accesses official databases to retrieve the latest data for fact-checking, such as documents containing official statements from the WHO and international health organizations, and stores them in a local cache.
[0913] 5. Content Analysis
[0914] The search server scrapes the content of each website and video and collects this content as text data. For example, it extracts the content of an article on a website titled "Is Corona a Man-Made Virus?"
[0915] 6. Text analysis and keyword extraction
[0916] The search server uses text analysis algorithms to process the collected text data and extract important keywords and assertions, such as "population virus," "evidence," and "research paper" from the content.
[0917] 7. Conduct fact-checks
[0918] The search server compares the extracted keywords and claims against fact-checking data—for example, comparing the claim "man-made virus" with the WHO statement "coronavirus is naturally occurring"—to identify inconsistencies, and calculates a confidence score for each search result based on this.
[0919] 8. Score tagging and search result generation
[0920] The search server tags each search result item with a confidence score and generates a search result list for display to the user. For example, a website titled "Is Corona a Man-Made Virus?" would be tagged with a low confidence score.
[0921] 9. Submitting Search Results
[0922] The search server sends the search results, along with the reliability scores, to the terminal, which then displays the received search results to the user.
[0923] 10. User Browsing and Judgment
[0924] Users review the search results and determine the reliability of the information based on the reliability score. For example, they may prioritize viewing items marked "High Reliability."
[0925] Specific examples
[0926] For example, if a user searches for "COVID-19 conspiracy theory," the search server retrieves related websites and videos, analyzes their content, compares them with official sources, calculates a reliability score, and displays it, allowing users to easily select reliable information.
[0927] In this way, the system of the present invention provides an environment in which users can easily obtain highly reliable information.
[0928] The processing flow will be explained below.
[0929] Step 1:
[0930] The user enters a search query into the device, for example, the keyword "coronavirus conspiracy theory."
[0931] Step 2:
[0932] The terminal transmits the input search query to the server.
[0933] Step 3:
[0934] The server analyzes the received search query and extracts related keywords, for example, "coronavirus" and "conspiracy theory."
[0935] Step 4:
[0936] The server searches an index database based on the extracted keywords to retrieve a list of related websites and videos, such as a website titled "Is COVID-19 a man-made virus?" and a video titled "The Truth About Conspiracy Theories."
[0937] Step 5:
[0938] The server accesses official databases to retrieve the latest data for fact-checking, such as official statements and documents from the WHO.
[0939] Step 6:
[0940] The server scrapes the content of each website and video and collects it as text data. For example, it extracts the content of an article on a website titled "Is Corona a Man-Made Virus?"
[0941] Step 7:
[0942] The server processes the collected text data with a text analysis algorithm to extract key keywords and claims, such as "population virus," "evidence," and "research paper."
[0943] Step 8:
[0944] The server compares the extracted keywords and claims against public fact-checking databases, for example, comparing the claim "man-made virus" with the WHO statement "coronavirus is naturally occurring."
[0945] Step 9:
[0946] The server calculates a confidence score for each search result item based on the comparison, for example, assigning a low confidence score to the "man-made virus" claim because it contradicts official data.
[0947] Step 10:
[0948] The server generates a search result list by tagging each search result with a confidence score. For example, a website asking "Is COVID-19 a man-made virus?" might be tagged with "Low confidence."
[0949] Step 11:
[0950] The server sends the search result list with the confidence scores attached to it to the terminal.
[0951] Step 12:
[0952] The terminal displays the received search result list to the user. The user checks the displayed search results and judges the reliability of the information based on the reliability score. For example, the user may preferentially view items marked "High Reliability."
[0953] This series of processes allows users to easily distinguish highly reliable information and makes decisions based on accurate information.
[0954] Example 1
[0955] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0956] In the current Internet environment, unreliable and false information is often spread, making it difficult for users to find accurate information, especially when it comes to information about serious issues. For this reason, there is a need for a system that allows users to quickly and easily obtain reliable information. In addition, there is a need for a method that effectively utilizes data provided by public institutions and objectively evaluates the reliability of information.
[0957] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0958] In this invention, the server includes means for receiving a search query from a user terminal, means for analyzing the received search query to extract related keywords, means for obtaining a list of related websites and videos based on the extracted keywords, means for obtaining fact-checking data from a database provided by a public institution, means for scraping the obtained content and collecting it as text data, means for analyzing the collected text data using a text analysis algorithm to extract important keywords and claims, means for comparing the extracted keywords and claims with the fact-checking data, and means for calculating and displaying a reliability score for each search result item based on the comparison result, thereby enabling users to quickly and accurately obtain reliable information.
[0959] "User terminal" means a physical device through which a user enters a search query, including a personal computer, smartphone, tablet, etc.
[0960] A "search query" is text data entered by a user to search for specific information.
[0961] A "search server" is a centralized computer system that receives, analyzes, and processes search queries from user terminals.
[0962] "Keywords" are related words or phrases extracted from a search query.
[0963] An "index database" is a database in which information about related websites, videos, etc. is organized and stored in a searchable format.
[0964] "Fact-checking data" is data obtained from databases provided by public institutions to assess the reliability of information, including official statements and research papers.
[0965] "Scraping" is the technique of extracting information from web pages using automated programs.
[0966] A "text analysis algorithm" is a computational method for analyzing text data to extract important information and patterns.
[0967] The "trust score" is a numerical evaluation of the reliability of the information in the search results, and is a value that users can use as a reference when judging the information.
[0968] "Comparison" is the process of comparing extracted keywords and claims with fact-checking data and analyzing for matches and contradictions.
[0969] The present invention is a search system that performs fact-checking based on reliable data provided by public institutions. To implement this system, a user terminal, a search server, and a fact-checking database are used as main components.
[0970] System Configuration
[0971] User Device
[0972] A device on which a user enters a search query and sends it to a search server. User terminals can be personal computers, smartphones, tablets, etc. A web browser is installed on these terminals, and users enter and send search queries through a web interface.
[0973] Search Server
[0974] The search server is the central component that analyzes search queries received from user devices and retrieves and processes relevant information. It is recommended to use a high-performance cloud server (e.g., one of the common cloud server services) for the search server. The following software and algorithms are required:
[0975] Natural Language Processing (NLP) libraries (e.g., NLP libraries commonly used in technical literature)
[0976] Scraping tools (e.g., common open source libraries)
[0977] Database management systems (e.g., commonly used open-source database management systems)
[0978] Fact-checking database
[0979] Fact-checking databases store reliable data provided by public institutions. Specifically, frequently updated data is retrieved from accessible data sources such as international health organizations and government statistical databases, and cached locally within the search server.
[0980] Example of operation
[0981] 1. A user enters a search query, for example, "coronavirus conspiracy theory," into the search bar on their device and presses the send button.
[0982] 2. The search server analyzes the received search query and extracts the related keywords "coronavirus" and "conspiracy theory."
[0983] 3. The search server searches the index database based on the extracted keywords and retrieves relevant search results (websites and videos).
[0984] 4. The search server also retrieves the latest fact-checking data from official databases and stores it in a local cache.
[0985] 5. The search server scrapes the content of the search results and collects it as text data, such as the main text of website articles and video descriptions.
[0986] 6. The search server uses text analysis algorithms to extract important keywords and assertions from the collected text data.
[0987] 7. The search server compares the extracted keywords and claims with fact-checking data, for example, comparing the claim "man-made virus" with the WHO statement "coronavirus is naturally occurring" to identify inconsistencies.
[0988] 8. The search server calculates a confidence score for each search result and displays it as a tag on each search result item.
[0989] Examples of prompt statements
[0990] "What are the steps to verify the credibility of conspiracy theories related to COVID-19?"
[0991] "Please explain specifically how you will conduct fact-checking."
[0992] This system allows users to obtain reliable information quickly and accurately, minimizing the impact of uncertain or false information.
[0993] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0994] Step 1:
[0995] A user enters a search query, such as "COVID-19 conspiracy theory," into the device's search bar and clicks the submit button. The device then sends the search query to a search server. The input data is free-form text, and the output is an HTTP request to the server.
[0996] Step 2:
[0997] The search server analyzes the received search query and extracts relevant keywords. Specifically, it uses the analysis module of a natural language processing (NLP) library (e.g., spaCy) to extract important keywords such as "coronavirus" and "conspiracy theory." The input is free-form text data, and the output is a list of extracted keywords.
[0998] Step 3:
[0999] The search server searches an index database based on the extracted keywords to obtain a list of related websites and videos. The database used for this process is one that has been pre-indexed by a web crawling tool. The input is a list of keywords, and the output is a list of related URLs and titles. Specifically, it includes a website titled "Is Corona a Man-Made Virus?" and a video titled "The Truth About Conspiracy Theories."
[1000] Step 4:
[1001] The search server makes HTTP requests to official databases (e.g., international health organizations or government agencies) to retrieve the latest data for fact-checking. The retrieved data is stored in a local cache for use in the next step. The input is the official API endpoint, and the output is a dataset of trusted documents and statements.
[1002] Step 5:
[1003] The search server uses a scraping tool (e.g., BeautifulSoup) to collect the content of the retrieved websites and videos. Specifically, it analyzes the HTML code of the webpage and extracts the article text and metadata as text data. The input is a list of URLs, and the output is content data in text format. For example, it collects the content of articles on a website titled "Is Corona a Man-Made Virus?"
[1004] Step 6:
[1005] The search server analyzes the collected text data using natural language processing (NLP) algorithms to extract important keywords and claims. Specifically, it uses an NLP library to extract frequently occurring words and phrases in the text and identify important keywords such as "population virus," "evidence," and "research paper." The input is text content data, and the output is a list of important keywords.
[1006] Step 7:
[1007] The search server uses the extracted keywords and claims to compare them with fact-checking data obtained from official databases. For example, it compares the claim "man-made virus" with the WHO statement "coronavirus is naturally occurring" and analyzes contradictions and similarities. The input is a list of keywords and reliable document data, and the output is the comparison results and a reliability score.
[1008] Step 8:
[1009] The search server tags each search result item with the calculated confidence score and generates a list of search results to display to the user. Specifically, the process involves adding the confidence score to the title and URL of each search result to create a list. The input is the comparison result and confidence score, and the output is a tagged list of search results.
[1010] Step 9:
[1011] The search server sends the tagged search result list to the terminal, which receives it and generates an HTML page to display to the user. The input is the tagged search result list, and the output is the search result page displayed to the user.
[1012] Step 10:
[1013] Users check the displayed search results and judge the reliability of the information based on the reliability score. Specifically, they prioritize browsing items with high reliability and acquire or share the information. The input is the displayed search result page, and the output is the user's browsing actions and judgment results.
[1014] This allows users to obtain reliable information quickly and accurately.
[1015] (Application example 1)
[1016] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1017] In recent years, the amount of information on the Internet has increased dramatically, including information of unknown authenticity and potentially misleading information. This has made it difficult for users to easily obtain reliable information. Furthermore, in virtual stores, it is difficult for consumers to judge the reliability of product information themselves, which risks leading to inaccurate purchasing decisions. Therefore, there is a need for a system that can automatically determine the reliability of information and present a reliability score when consumers search for product information in virtual stores.
[1018] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1019] In this invention, the server includes means for acquiring fact-checking data from a database provided by a public institution, means for receiving a search query from a user, means for acquiring related search results based on the search query, means for extracting important keywords and assertions by text-analyzing the content of the search results, means for comparing the extracted keywords and assertions with the fact-checking data, means for calculating and displaying a reliability score for each search result item based on the comparison result, and means for tagging product information with the reliability score, allowing users to judge reliability when making purchasing decisions. This allows consumers to easily judge the reliability of product information in virtual stores and make purchasing decisions with confidence.
[1020] "Public institution" refers to a public institution, such as a government or local government, that provides public functions or services.
[1021] A "database" is an information repository that organizes and stores a collection of data so that it can be efficiently searched and used.
[1022] "Fact-checking data" refers to reliable data used to verify the veracity of information.
[1023] "User" means any individual person or entity that uses the System or Services.
[1024] A "search query" refers to a question or keyword that a user enters into a search engine or system.
[1025] "Search Results" means the information returned by a search engine or system based on a search query.
[1026] "Content" refers to any piece of information or data present on a website or digital media.
[1027] "Text analytics" refers to the techniques and processes used to analyze text data and understand its content.
[1028] "Keywords" refer to key words that play an important role in a text or search query.
[1029] An "argument" is a statement of opinion or perspective on a certain matter.
[1030] "Comparison" refers to the act of comparing two or more things to see their differences and similarities.
[1031] A "trust score" is a numerical indicator that evaluates the reliability of information or data.
[1032] "Display" refers to the means by which information is visually conveyed to the user.
[1033] "Tagging" refers to the act of adding labels and scores to information.
[1034] "Product information" refers to data including descriptions, features, specifications, etc. of the target product.
[1035] "Consumption behavior" refers to the series of actions that consumers take to purchase goods and services.
[1036] This invention describes a system that automatically determines the reliability of product information when a user searches for it in a virtual store and presents a reliability score. This system is composed of the following main components:
[1037] System configuration
[1038] The system is implemented primarily using the following hardware and software:
[1039] 1. User terminal: a smartphone, smart glasses, head-mounted display, or other display device.
[1040] 2. Search server: A server responsible for data analysis and fact-checking.
[1041] 3. Fact-checking database: A database that stores reliable data provided by public institutions.
[1042] Program processing
[1043] User Roles
[1044] A user inputs a search query, such as "antioxidant supplements," into a user terminal and presses a search button. This operation sends the query to a search server.
[1045] The role of the search server
[1046] The search server analyzes the search query received from the user terminal and extracts related keywords, which allows the system to obtain keywords such as "antioxidant" and "supplement."
[1047] Furthermore, the search server searches the virtual store database based on the extracted keywords to obtain a list of related products. For example, product information such as "Antioxidant Supplement A" and "Supplement B" is obtained.
[1048] The search server then accesses databases provided by public organizations (e.g., government agencies and research institutes) to retrieve reliable information and retrieve the latest data for fact-checking, including official reports and the latest research papers.
[1049] Content analysis and fact-checking
[1050] The search server scrapes the descriptions and reviews of each product and collects them as text data. For example, it analyzes the description of "Antioxidant Supplement A" and extracts important keywords and claims.
[1051] Furthermore, the search server compares the extracted keywords and claims with information from databases of public institutions, and calculates a credibility score to evaluate the reliability of the product. For example, a credibility score for a product is determined by comparing the claim "has antioxidant properties."
[1052] Visibility and User Roles
[1053] Finally, the search server tags each product with the calculated reliability score and displays it as information to help users judge reliability when deciding on their purchasing behavior, allowing users to select products with confidence based on reliable information.
[1054] Examples and prompts
[1055] For example, if a user searches for "antioxidant supplements," the search server retrieves related product information, analyzes the descriptions of each product, and calculates a reliability score. As a result, products with a "High Reliability" rating can be presented to the user preferentially.
[1056] Here are some examples of prompts:
[1057] "Antioxidant supplement A has been proven effective in clinical trials. Please compare this information with data from official institutions and calculate a reliability score."
[1058] In this way, this system provides an environment in which consumers can easily judge the reliability of product information in a virtual store.
[1059] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1060] Step 1: Receiving User Input
[1061] A user enters a search query such as "antioxidant supplements" into a device such as a smartphone or head-mounted display and presses the search button. The device then sends this query to a search server. The input data is the user's search query, and the output data is the query information sent to the search server.
[1062] Step 2: Processing the search query
[1063] The search server analyzes the search query received from the device and extracts the related keywords "antioxidant" and "supplement." Specifically, it uses a natural language processing algorithm to tokenize the search query and extract important keywords. The input data is the received query information, and the output data is the extracted keywords.
[1064] Step 3: Getting search results
[1065] The search server searches the virtual store database based on the extracted keywords and retrieves a list of related products. For example, product information such as "antioxidant supplement A" and "supplement B" is retrieved. The input data are the extracted keywords, and the output data is a list of products.
[1066] Step 4: Obtaining fact-checking data
[1067] The search server accesses public databases to retrieve the latest data for fact-checking, including official reports from government agencies and research institutes and the latest research papers. The input data is the access request to the public database, and the output data is the retrieved data for fact-checking.
[1068] Step 5: Analyze your content
[1069] The search server scrapes the descriptions and reviews of each product and collects this content as text data. Specifically, it extracts the text of articles and reviews from each page. The input data is a list of products, and the output data is the collected text data.
[1070] Step 6: Text analysis and keyword extraction
[1071] The search server processes the collected text data using a text analysis algorithm to extract key keywords and claims. Specifically, it uses TF-IDF (Term Frequency-Inverse Document Frequency) to identify important keywords. The input data is the collected text data, and the output data is the extracted keywords and claims.
[1072] Step 7: Conduct a fact check
[1073] The search server compares the extracted keywords and claims with the fact-checking data, checking for keyword matches and inconsistencies, and calculating a reliability score for each product. The input data are the extracted keywords and claims and the fact-checking data, and the output data is a reliability score for each product.
[1074] Step 8: Score tagging and search result generation
[1075] The search server tags each product information with the calculated confidence score and generates a search result list to display to the user. Specifically, it assigns a confidence score to each product information and converts it into a format that can be displayed in the user interface. The input data are the confidence scores and the product list, and the output data is the tagged search result list.
[1076] Step 9: Submit search results
[1077] The search server sends the tagged search results to the terminal, which then displays the received search results to the user. The input data is the tagged search result list, and the output data is the displayed search results.
[1078] Step 10: User View and Decision
[1079] The user checks the displayed search results and judges the reliability of the product information based on the reliability score. Specifically, the user may take action such as prioritizing the purchase of products displayed as "High reliability." The input data are the displayed search results, and the output data are the user's purchasing behavior.
[1080] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1081] This invention combines a search system that performs fact-checking based on highly reliable data provided by public institutions with an emotion engine that recognizes user emotions. The programs required to implement this system and their processing are shown below.
[1082] System configuration
[1083] The system mainly consists of the following components:
[1084] 1. User Device
[1085] 2. Search Server
[1086] 3. Fact-checking databases
[1087] 4. Emotion Engine
[1088] Program processing
[1089] 1. Receiving user input
[1090] The user enters a search query, such as "coronavirus conspiracy theory," into the device and presses the send button.
[1091] 2. Processing search queries
[1092] The terminal transmits the input search query to the server.
[1093] 3. Search Query Analysis
[1094] The server analyzes the received search query and extracts related keywords, for example, "coronavirus" and "conspiracy theory."
[1095] 4. Obtaining search results
[1096] The server searches an index database based on the extracted keywords to retrieve a list of related websites and videos, such as a website titled "Is COVID-19 a man-made virus?" and a video titled "The Truth About Conspiracy Theories."
[1097] 5. Acquisition of data from public institutions
[1098] The server accesses official databases to retrieve the latest data for fact-checking, such as official statements and documents from the WHO.
[1099] 6. Content Analysis
[1100] The server scrapes the content of each website and video and collects it as text data. For example, it extracts the content of an article on a website titled "Is Corona a Man-Made Virus?"
[1101] 7. Text Analysis and Keyword Extraction
[1102] The server processes the collected text data with a text analysis algorithm to extract key keywords and claims, such as "population virus," "evidence," and "research paper."
[1103] 8. Fact Check
[1104] The server compares the extracted keywords and claims against public fact-checking databases—for example, comparing the claim "man-made virus" with the WHO statement "coronavirus is naturally occurring"—and calculates a confidence score for each search result.
[1105] 9. Emotion Recognition with Emotion Engine
[1106] The device uses an emotion engine to analyze the user's emotions based on the user's input and operations (e.g., keystroke speed, input content, facial recognition, etc.). For example, it determines whether the user is feeling stressed.
[1107] 10. Emotional Feedback
[1108] The server then adjusts the reliability score and display order of search results based on the analysis results of the emotion engine. For example, if a user is feeling stressed, more reliable information will be displayed first.
[1109] 11. Score tagging and search result generation
[1110] The server tags each search result item with a confidence score and generates a search result list to display to the user. For example, a site asking "Is COVID-19 a man-made virus?" might be given a "Low confidence" rating.
[1111] 12. Submitting Search Results
[1112] The server sends the search result list with the confidence scores attached to it to the terminal.
[1113] 13. Display of search results
[1114] The device displays the received search results to the user. The user checks the displayed search results and determines the reliability of the information based on the reliability score. For example, the user may prioritize viewing items marked "High Reliability."
[1115] Specific examples
[1116] For example, if a user searches for "COVID-19 conspiracy theory," the search server retrieves related websites and videos and analyzes their content. It then compares them with documents from official institutions and calculates a reliability score. The device also uses an emotion engine to analyze the user's emotions (e.g., anxiety, stress) and adjusts the display order of search results based on the results. As a result, users can easily select reliable information and browse with peace of mind.
[1117] In this way, the system of the present invention provides an environment in which the user can easily obtain highly reliable information, and further provides information according to the user's emotional state.
[1118] The processing flow will be explained below.
[1119] Step 1:
[1120] The user enters a search query into the device, for example, the keyword "coronavirus conspiracy theory."
[1121] Step 2:
[1122] The terminal transmits the input search query to the server.
[1123] Step 3:
[1124] The server analyzes the received search query and extracts related keywords, for example, "coronavirus" and "conspiracy theory."
[1125] Step 4:
[1126] The server searches an index database based on the extracted keywords to retrieve a list of related websites and videos, such as a website titled "Is COVID-19 a man-made virus?" and a video titled "The Truth About Conspiracy Theories."
[1127] Step 5:
[1128] The server accesses official databases to retrieve the latest data for fact-checking, such as official statements and documents from the WHO.
[1129] Step 6:
[1130] The server scrapes the content of each website and video and collects it as text data. For example, it extracts the content of an article on a website titled "Is Corona a Man-Made Virus?"
[1131] Step 7:
[1132] The server processes the collected text data with a text analysis algorithm to extract key keywords and claims, such as "population virus," "evidence," and "research paper."
[1133] Step 8:
[1134] The server compares the extracted keywords and claims against public fact-checking databases, for example, comparing the claim "man-made virus" with the WHO statement "coronavirus is naturally occurring."
[1135] Step 9:
[1136] The server calculates a confidence score for each search result item based on the comparison, for example, assigning a low confidence score to the "man-made virus" claim because it contradicts official data.
[1137] Step 10:
[1138] The device uses an emotion engine to analyze the user's emotions based on the user's input and operations (e.g., keystroke speed, input content, facial recognition, etc.). For example, it can determine whether the user is feeling stressed.
[1139] Step 11:
[1140] The server then adjusts the reliability score and display order of search results based on the analysis results of the emotion engine. For example, if a user is feeling stressed, more reliable information will be displayed first.
[1141] Step 12:
[1142] The server tags each search result item with a confidence score and generates a search result list to display to the user. For example, a site asking "Is COVID-19 a man-made virus?" might be given a "Low confidence" rating.
[1143] Step 13:
[1144] The server sends the search result list with the confidence scores attached to it to the terminal.
[1145] Step 14:
[1146] The terminal displays the received search result list to the user. The user checks the displayed search results and judges the reliability of the information based on the reliability score. For example, the user may preferentially view items marked "High Reliability."
[1147] This series of processes allows the user to easily distinguish highly reliable information, and also makes it possible to provide information that is appropriate for the user's emotional state.
[1148] Example 2
[1149] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1150] In today's world, there is a huge amount of information available on the Internet, but it is not easy to find accurate and reliable information. Furthermore, providing information without considering the user's emotional state can increase stress and anxiety. Therefore, there is a need for a search system that takes into account the user's emotional state while ensuring the reliability of the information.
[1151] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[1152] In this invention, the server includes means for acquiring fact-checking data from a database provided by a public institution, means for receiving a search query from a user, means for acquiring related search results based on the search query, means for analyzing the content of the search results to extract important keywords and claims, means for comparing the extracted keywords and claims with the fact-checking data, means for calculating and displaying a reliability score for each search result item based on the comparison result, means for analyzing user emotions, and means for adjusting the display order of the search results based on the emotion analysis results.
[1153] This allows the user to easily obtain highly reliable information and to be provided with information that corresponds to the user's own emotional state.
[1154] A "database" is a storage device that stores highly reliable data provided by public institutions.
[1155] "Fact-checking data" refers to documents and statements released by public institutions, and is basic data used to verify the accuracy of information.
[1156] A "search query" is a question or keyword that a user enters into the system, and is the search condition that the system investigates.
[1157] "Search results" are lists of information retrieved from databases or the Internet based on a search query.
[1158] "Text analysis" is a technique that uses natural language processing technology to understand the content of text data and extract important keywords and assertions.
[1159] "Keywords" are important words or phrases extracted from a search query or analyzed text.
[1160] "Arguments" are important opinions or points of view in a text that are revealed through text analysis.
[1161] "Comparison" is the process of comparing the fact-checking data with the extracted keywords and claims to confirm whether they match or disagree.
[1162] The "trust score" is a numerical representation of the reliability of information based on the results of fact-checking, and is an indicator of the reliability of search results.
[1163] "Emotion analysis" is the process of analyzing a user's emotional state based on their input and actions, and is a technology that determines what emotions the user is feeling.
[1164] "Display ranking" refers to the order in which each item is displayed in the search result list, and is determined based on confidence scores and sentiment analysis results.
[1165] This invention combines a search system that performs fact-checking based on reliable data provided by public institutions with an emotion engine that recognizes user emotions. This system mainly consists of the following components: a user terminal, a search server, a fact-checking database, and an emotion engine.
[1166] System configuration
[1167] To implement this invention, the following hardware and software are used:
[1168] 1. User Device
[1169] Hardware: Personal computers, smartphones
[1170] Software: Web browser, input interface
[1171] 2. Search Server
[1172] Hardware: High-performance server
[1173] Software: Natural language processing libraries (e.g., NLTK, SpaCy, BERT), search engines (e.g., Elasticsearch, Solr), scraping tools (e.g., BeautifulSoup, Selenium), sentiment analysis APIs (e.g., Microsoft Azure Cognitive Services, Emotion API)
[1174] 3. Fact-checking databases
[1175] Hardware: Database Server
[1176] Software: Public Sector API Access
[1177] 4. Emotion Engine
[1178] Hardware: Embedded systems, cameras and keyboards attached to user devices
[1179] Software: Emotion analysis algorithms, face tracking software (e.g. OpenCV)
[1180] Implementation Procedure
[1181] 1. Receiving user input
[1182] The user enters "COVID-19 conspiracy theory" into the search field and clicks the search button. The device receives this input, temporarily stores it in its internal memory, and prepares to send it to the server.
[1183] 2. Processing search queries
[1184] The device sends the search query entered by the user to the server as an HTTP request, which includes the query text data and metadata such as the user ID.
[1185] 3. Search Query Analysis
[1186] The server parses the incoming HTTP request and processes the query text with a natural language processing algorithm (e.g., Python's NLTK library), which identifies the main keywords "coronavirus" and "conspiracy theory."
[1187] 4. Obtaining search results
[1188] The server queries an index database using a search engine (e.g., Elasticsearch or Solr) based on the extracted keywords, which retrieves a list of related websites and videos.
[1189] 5. Acquisition of data from public institutions
[1190] The server sends API requests to public databases to retrieve the latest, most reliable information, including official statements and research papers from organizations like the WHO and CDC.
[1191] 6. Content Analysis
[1192] The server uses scraping technology (e.g., BeautifulSoup or Selenium) to collect the content of each site and video as text data, extracting the content of the article on the site "Is Corona a Man-Made Virus?" and the subtitles of the video.
[1193] 7. Text Analysis and Keyword Extraction
[1194] The server then processes the collected text data again using a natural language processing algorithm (e.g., SpaCy or BERT) to extract important keywords and assertions, such as "population virus," "evidence," and "research paper."
[1195] 8. Fact Check
[1196] The server compares the extracted keywords and claims with official data it has obtained—for example, comparing the claim "man-made virus" with the WHO statement "coronavirus is naturally occurring"—and then uses an algorithm to calculate a confidence score for each search result, which is calculated on a scale from 0 to 100.
[1197] 9. Emotion Recognition with Emotion Engine
[1198] The device uses an emotion engine to analyze the user's emotions based on the user's input and operations, such as keystroke speed and facial recognition (using OpenCV), to determine whether the user is feeling stressed or anxious.
[1199] 10. Emotional Feedback
[1200] The server then adjusts the reliability score and display order of search results based on the analysis results of the emotion engine. For example, if a user is feeling stressed, an algorithm is applied to prioritize displaying information with high reliability.
[1201] 11. Score tagging and search result generation
[1202] The server tags each search result with a confidence score and generates a list of search results to display to the user. For example, a website asking "Is COVID-19 a man-made virus?" might be tagged with "Confidence: 20 (Low)."
[1203] 12. Submission and Display of Search Results
[1204] The server sends the search result list, with assigned reliability scores, to the terminal as an HTTP response. The terminal displays the received search result list on its user interface. The user checks this display and decides which information to view based on the reliability score. For example, they may prioritize clicking and viewing items with a "Reliability: 90" rating.
[1205] Specific examples
[1206] For example, if a user searches for "coronavirus conspiracy theory," the following specific actions occur: The user enters "coronavirus conspiracy theory" into the search bar and clicks the search button. The device receives the input, generates an HTTP request, and sends it to the server. The server analyzes the query and extracts related keywords. The server uses a search engine to retrieve relevant sites and videos. The server retrieves reliable data from public institution databases via an API request. The server scrapes the content of the retrieved sites and videos and collects it as text data. The server uses a natural language processing algorithm to extract important keywords and claims from the collected text data. The server compares the extracted keywords and claims with public institution data and calculates a confidence score. The device analyzes the user's emotions using an emotion engine and sends the results to the server. The server adjusts the confidence score and display ranking of the search results based on the user's emotions. The server tags the confidence score and generates a search result list. The server sends the final search result list to the device. The device displays the search result list on a user interface, and the user can review and select results.
[1207] Example prompt sentence:
[1208] "New coronavirus conspiracy theory"
[1209] Result: A list of highly reliable information is displayed, including an official statement from a public institution that "coronavirus is a natural occurrence." Because the emotion engine determined that the user was highly stressed, highly reliable information is displayed first.
[1210] In this way, the system of the present invention provides an environment in which the user can easily obtain highly reliable information, and further provides information according to the user's emotional state.
[1211] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1212] Step 1:
[1213] The user enters "COVID-19 conspiracy theory" in the search field and clicks the search button. This causes the device to receive a search query. The input is the search query entered by the user, and the output is the query temporarily stored in the internal memory. The device prepares to send this data to the server.
[1214] Step 2:
[1215] The device sends the search query entered by the user to the server as an HTTP request. The input is the search query stored in the internal memory, and the output is the HTTP request sent to the server. The request includes the text data of the query and metadata such as the user ID.
[1216] Step 3:
[1217] The server parses the received HTTP request and processes the query text with an NLP algorithm (e.g., Python's NLTK library). The input is the HTTP request sent to the server, and the output is the parsed search query and keywords extracted from it, e.g., "coronavirus" and "conspiracy theory."
[1218] Step 4:
[1219] The server queries the index database using Elasticsearch or Solr based on the extracted keywords. The input is the extracted keywords, and the output is a list of related websites and videos, such as a website called "Is COVID-19 a man-made virus?" or a video called "The Truth About Conspiracy Theories."
[1220] Step 5:
[1221] The server sends API requests to public databases to retrieve the latest, most reliable information. The input is the API request, and the output is the retrieved public data, such as an official WHO statement or a research paper.
[1222] Step 6:
[1223] The server collects the content of each website and video as text data using scraping technology (e.g., BeautifulSoup or Selenium). The input is the URL of the relevant website or video, and the output is the scraped text data, such as the content of an article on the website "Is Coronavirus a Man-Made Virus?"
[1224] Step 7:
[1225] The server processes the collected text data again with an NLP algorithm (e.g., SpaCy or BERT) to extract important keywords and claims. The input is the collected text data, and the output is the extracted important keywords and claims, such as "population virus," "evidence," and "research paper."
[1226] Step 8:
[1227] The server compares the extracted keywords and claims with the retrieved official data. The input is the extracted keywords and claims and the official data, and the output is a confidence score for each search result, e.g., "Confidence: 20 (Low)."
[1228] Step 9:
[1229] The device uses an emotion engine to analyze the user's emotions based on their input and operations. Specifically, it uses keystroke speed and facial recognition technology (e.g., OpenCV). The input is the user's input data and operation data, and the output is the analyzed user's emotional state, such as "high stress."
[1230] Step 10:
[1231] The server adjusts the reliability score and display order of search results based on the analysis results of the emotion engine. The input is the emotion analysis result and reliability score, and the output is an adjusted search result list, for example, a list in which highly reliable information is displayed at the top.
[1232] Step 11:
[1233] The server tags each search result with a confidence score and generates a final search result list to display to the user. The input is the adjusted search result list, and the output is a search result list with a confidence score, e.g., "Confidence: 90 (High)".
[1234] Step 12:
[1235] The server sends the final search result list to the terminal. The input is the search result list with the assigned confidence scores, and the output is the HTTP response sent to the terminal.
[1236] Step 13:
[1237] The terminal displays the received search result list on the user interface. The input is the HTTP response sent to the terminal, and the output is the search result list displayed on the screen. The user checks the results and decides which information to view based on the confidence score.
[1238] (Application example 2)
[1239] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1240] In today's information society, it is extremely important for users to have access to reliable information. However, while a vast amount of information is easily available, it is often difficult to verify its reliability. Furthermore, there are no established methods for providing appropriate information to users who are feeling stressed or anxious. Therefore, an objective of the present invention is to provide a system that provides users with reliable information and also provides information that is tailored to the user's emotional state.
[1241] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for acquiring fact-checking data from a database provided by a public institution, means for receiving a search query from a user, means for acquiring related search results based on the search query, means for text-analyzing the content of the search results to extract important keywords and assertions, means for comparing the extracted keywords and assertions with the fact-checking data, means for calculating and displaying a reliability score for each search result item based on the comparison result, means for recognizing passenger emotions using an emotion engine, and means for adjusting the reliability score and information display based on the recognized emotions. This allows the user to appropriately acquire reliable information and provides information according to the user's emotional state.
[1242] A "public institution" is an organization that operates for the public interest, such as a government or local government, and issues official data and statements.
[1243] A "database" is a collection of information designed to efficiently store, retrieve, and update data.
[1244] "Fact-checking data" is reliable information and materials collected and provided for the purpose of fact-checking.
[1245] "User" means an individual or organization that uses a system or application.
[1246] A "search query" is a string of characters or keywords that a user enters when searching for information.
[1247] "Search results" are a collection of information or data retrieved based on a search query.
[1248] "Text analytics" refers to techniques and methods for extracting important information and patterns from text data.
[1249] "Keywords" are important words or phrases used to represent specific information.
[1250] An "assertion" is an expression that expresses a particular opinion or point of view.
[1251] "Comparison" is the act of comparing two or more pieces of information or data and evaluating their differences and similarities.
[1252] A "trust score" is a numerical indicator used to evaluate the reliability of information.
[1253] "Display" is the act of visually presenting acquired information or results to the user.
[1254] An "emotion engine" is software or hardware that analyzes and recognizes a user's emotions.
[1255] "Recognition" is the act of analyzing, understanding, and grasping information and data.
[1256] "Adjustment" is the act of changing information or settings to accommodate specific conditions or circumstances.
[1257] "Means" are the methods or techniques used to achieve a particular goal.
[1258] The present invention is a system for providing highly reliable information to a user and providing information according to the user's emotional state. Specific embodiments of the present invention will be described below.
[1259] System Configuration
[1260] Hardware
[1261] 1. User device: A device such as a smartphone, tablet, or PC that accepts user input and operations. It also has a camera and microphone to collect data for emotion recognition.
[1262] 2. Search Server: A central server that searches, analyzes, and fact-checks various data.
[1263] 3. Fact-checking database: A database that stores reliable data provided by public institutions.
[1264] 4. Emotion engine: Software for analyzing the user's emotional state.
[1265] software
[1266] 1. EmotionEngine: Software that analyzes data input from cameras and microphones in real time to recognize the user's emotions. It uses image processing libraries such as OpenCV.
[1267] 2. FactCheckEngine: An engine that cross-checks the information obtained with public data and evaluates its reliability.
[1268] 3. PublicDataFetcher: Software for retrieving up-to-date and reliable data from public databases.
[1269] System Operation
[1270] The server first receives a search query from a user device, then retrieves relevant search results based on the search query, performs text analysis on the content, extracts important keywords and claims, and compares the extracted keywords and claims with fact-checking data to calculate their reliability score.
[1271] At the same time, EmotionEngine recognizes the user's emotions in real time using the camera and microphone installed on the user's device. Based on the recognized emotions, the server adjusts the reliability score and information display to provide the user with appropriate information.
[1272] Specific examples
[1273] For example, if a user enters the search query "traffic congestion current situation," the server retrieves related websites and news articles, performs fact-checking, and calculates and displays a reliability score. If the emotion engine recognizes that the user is feeling stressed, it prioritizes information with high reliability and suggests relaxing music and videos.
[1274] Prompt Sentence Examples
[1275] Provide information to promote safe driving when users are under stress, such as prioritizing public traffic and weather information and playing relaxing music and scenic videos.
[1276] In this way, a system is realized that allows users to easily obtain highly reliable information and receive information that is optimal for their emotional state.
[1277] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1278] Step 1:
[1279] User enters search query and submits
[1280] A user inputs a query, for example, "traffic congestion current situation," into the terminal and presses the send button. After the input is made, the terminal sends the search query to the server. The input is in the form of text, and the output is the search query sent to the server.
[1281] Step 2:
[1282] The server parses the search query
[1283] The server analyzes the received search query and extracts relevant keywords. For example, it extracts the keywords "traffic congestion" and "current situation." The input is the search query text data, and the output is the analyzed keyword list. The server performs this process using a text analysis algorithm.
[1284] Step 3:
[1285] The server retrieves the search results
[1286] The server searches the index database based on the extracted keywords to obtain a list of relevant sites and news articles. The input is the parsed keyword list, and the output is a list of relevant search results. The server uses the index database to perform this process.
[1287] Step 4:
[1288] Server acquires data from public institutions
[1289] The server accesses a public database to retrieve the latest data for fact-checking, for example, data from a traffic information center. The input is the URL or API key of the public database to access, and the output is the trusted data. The server performs this process using an HTTP request.
[1290] Step 5:
[1291] The server analyzes the content
[1292] The server scrapes the content of each website or news article and collects it as text data. The input is a list of relevant search results, and the output is the retrieved text data. The server performs this process using web scraping technology.
[1293] Step 6:
[1294] The server performs text analysis and keyword extraction.
[1295] The server processes the collected text data with text analysis algorithms to extract important keywords and assertions. The input is the collected text data, and the output is the extracted keywords and assertions. The server performs this processing using natural language processing (NLP) techniques.
[1296] Step 7:
[1297] Server fact-checking
[1298] The server compares the extracted keywords and claims against public fact-checking databases and calculates a confidence score. The inputs are the extracted keywords and data from the public databases, and the output is a confidence score. The server performs this process using database queries and comparison algorithms.
[1299] Step 8:
[1300] The device uses an emotion engine to recognize emotions.
[1301] The device analyzes the user's emotions using a camera and microphone. The input is video and audio data, and the output is the recognized emotion data. The device performs this processing using the EmotionEngine.
[1302] Step 9:
[1303] Server provides feedback based on emotions
[1304] The server adjusts the confidence score and information display based on the analysis results of the emotion engine. The input is the recognized emotion data and confidence score, and the output is the adjusted information display ranking and feedback. The server performs this process using an algorithm.
[1305] Step 10:
[1306] The server generates a search result list tagged with a confidence score
[1307] The server tags each search result item with a confidence score and generates a search result list for display to the user. The input is the confidence score and the list of search results, and the output is the tagged search result list. The server performs this process using a list generation algorithm.
[1308] Step 11:
[1309] The server sends the search result list
[1310] The server sends the search result list with attached confidence scores to the user device. The input is the tagged search result list, and the output is the data sent to the user device. The server performs this process using a network protocol.
[1311] Step 12:
[1312] Your device will display the search results
[1313] The terminal displays the received search results to the user. The user reviews the displayed search results and determines the reliability of the information based on the reliability score. The input is the search result list received from the server, and the output is the displayed search results. The terminal performs this process using a graphical user interface (GUI).
[1314] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1315] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1316] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.
[1317] [Fourth embodiment]
[1318] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.
[1319] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[1320] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1321] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.
[1322] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1323] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1324] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1325] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.
[1326] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1327] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1328] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1329] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1330] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1331] The present invention is a search system that performs fact-checking based on highly reliable data provided by public institutions. The programs required to implement this system and their processing are shown below.
[1332] System configuration
[1333] The system mainly consists of the following components:
[1334] 1. User Device
[1335] 2. Search Server
[1336] 3. Fact-checking databases
[1337] Program processing
[1338] 1. Receiving user input
[1339] The user enters a search query, such as "coronavirus conspiracy theory," into the device and presses the send button.
[1340] 2. Processing search queries
[1341] The search server analyzes the search query received from the device and extracts related keywords. Through this analysis, the system obtains keywords such as "coronavirus" and "conspiracy theory."
[1342] 3. Obtaining search results
[1343] The search server searches the index database based on the extracted keywords to retrieve a list of related websites and videos. For example, it retrieves a website titled "Is Corona a Man-Made Virus?" and a video titled "The Truth About Conspiracy Theories."
[1344] 4. Obtaining data for fact-checking
[1345] The search server accesses official databases to retrieve the latest data for fact-checking, such as documents containing official statements from the WHO and international health organizations, and stores them in a local cache.
[1346] 5. Content Analysis
[1347] The search server scrapes the content of each website and video and collects this content as text data. For example, it extracts the content of an article on a website titled "Is Corona a Man-Made Virus?"
[1348] 6. Text analysis and keyword extraction
[1349] The search server uses text analysis algorithms to process the collected text data and extract important keywords and assertions, such as "population virus," "evidence," and "research paper" from the content.
[1350] 7. Conduct fact-checks
[1351] The search server compares the extracted keywords and claims against fact-checking data—for example, comparing the claim "man-made virus" with the WHO statement "coronavirus is naturally occurring"—to identify inconsistencies, and calculates a confidence score for each search result based on this.
[1352] 8. Score tagging and search result generation
[1353] The search server tags each search result item with a confidence score and generates a search result list for display to the user. For example, a website titled "Is Corona a Man-Made Virus?" would be tagged with a low confidence score.
[1354] 9. Submitting Search Results
[1355] The search server sends the search results, along with the reliability scores, to the terminal, which then displays the received search results to the user.
[1356] 10. User Browsing and Judgment
[1357] Users review the search results and determine the reliability of the information based on the reliability score. For example, they may prioritize viewing items marked "High Reliability."
[1358] Specific examples
[1359] For example, if a user searches for "COVID-19 conspiracy theory," the search server retrieves related websites and videos, analyzes their content, compares them with official sources, calculates a reliability score, and displays it, allowing users to easily select reliable information.
[1360] In this way, the system of the present invention provides an environment in which users can easily obtain highly reliable information.
[1361] The processing flow will be explained below.
[1362] Step 1:
[1363] The user enters a search query into the device, for example, the keyword "coronavirus conspiracy theory."
[1364] Step 2:
[1365] The terminal transmits the input search query to the server.
[1366] Step 3:
[1367] The server analyzes the received search query and extracts related keywords, for example, "coronavirus" and "conspiracy theory."
[1368] Step 4:
[1369] The server searches an index database based on the extracted keywords to retrieve a list of related websites and videos, such as a website titled "Is COVID-19 a man-made virus?" and a video titled "The Truth About Conspiracy Theories."
[1370] Step 5:
[1371] The server accesses official databases to retrieve the latest data for fact-checking, such as official statements and documents from the WHO.
[1372] Step 6:
[1373] The server scrapes the content of each website and video and collects it as text data. For example, it extracts the content of an article on a website titled "Is Corona a Man-Made Virus?"
[1374] Step 7:
[1375] The server processes the collected text data with a text analysis algorithm to extract key keywords and claims, such as "population virus," "evidence," and "research paper."
[1376] Step 8:
[1377] The server compares the extracted keywords and claims against public fact-checking databases, for example, comparing the claim "man-made virus" with the WHO statement "coronavirus is naturally occurring."
[1378] Step 9:
[1379] The server calculates a confidence score for each search result item based on the comparison, for example, assigning a low confidence score to the "man-made virus" claim because it contradicts official data.
[1380] Step 10:
[1381] The server generates a search result list by tagging each search result with a confidence score. For example, a website asking "Is COVID-19 a man-made virus?" might be tagged with "Low confidence."
[1382] Step 11:
[1383] The server sends the search result list with the confidence scores attached to it to the terminal.
[1384] Step 12:
[1385] The terminal displays the received search result list to the user. The user checks the displayed search results and judges the reliability of the information based on the reliability score. For example, the user may preferentially view items marked "High Reliability."
[1386] This series of processes allows users to easily distinguish highly reliable information and makes decisions based on accurate information.
[1387] Example 1
[1388] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1389] In the current Internet environment, unreliable and false information is often spread, making it difficult for users to find accurate information, especially when it comes to information about serious issues. For this reason, there is a need for a system that allows users to quickly and easily obtain reliable information. In addition, there is a need for a method that effectively utilizes data provided by public institutions and objectively evaluates the reliability of information.
[1390] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1391] In this invention, the server includes means for receiving a search query from a user terminal, means for analyzing the received search query to extract related keywords, means for obtaining a list of related websites and videos based on the extracted keywords, means for obtaining fact-checking data from a database provided by a public institution, means for scraping the obtained content and collecting it as text data, means for analyzing the collected text data using a text analysis algorithm to extract important keywords and claims, means for comparing the extracted keywords and claims with the fact-checking data, and means for calculating and displaying a reliability score for each search result item based on the comparison result, thereby enabling users to quickly and accurately obtain reliable information.
[1392] "User terminal" means a physical device through which a user enters a search query, including a personal computer, smartphone, tablet, etc.
[1393] A "search query" is text data entered by a user to search for specific information.
[1394] A "search server" is a centralized computer system that receives, analyzes, and processes search queries from user terminals.
[1395] "Keywords" are related words or phrases extracted from a search query.
[1396] An "index database" is a database in which information about related websites, videos, etc. is organized and stored in a searchable format.
[1397] "Fact-checking data" is data obtained from databases provided by public institutions to assess the reliability of information, including official statements and research papers.
[1398] "Scraping" is the technique of extracting information from web pages using automated programs.
[1399] A "text analysis algorithm" is a computational method for analyzing text data to extract important information and patterns.
[1400] The "trust score" is a numerical evaluation of the reliability of the information in the search results, and is a value that users can use as a reference when judging the information.
[1401] "Comparison" is the process of comparing extracted keywords and claims with fact-checking data and analyzing for matches and contradictions.
[1402] The present invention is a search system that performs fact-checking based on reliable data provided by public institutions. To implement this system, a user terminal, a search server, and a fact-checking database are used as main components.
[1403] System Configuration
[1404] User Device
[1405] A device on which a user enters a search query and sends it to a search server. User terminals can be personal computers, smartphones, tablets, etc. A web browser is installed on these terminals, and users enter and send search queries through a web interface.
[1406] Search Server
[1407] The search server is the central component that analyzes search queries received from user devices and retrieves and processes relevant information. It is recommended to use a high-performance cloud server (e.g., one of the common cloud server services) for the search server. The following software and algorithms are required:
[1408] Natural Language Processing (NLP) libraries (e.g., NLP libraries commonly used in technical literature)
[1409] Scraping tools (e.g., common open source libraries)
[1410] Database management systems (e.g., commonly used open-source database management systems)
[1411] Fact-checking database
[1412] Fact-checking databases store reliable data provided by public institutions. Specifically, frequently updated data is retrieved from accessible data sources such as international health organizations and government statistical databases, and cached locally within the search server.
[1413] Example of operation
[1414] 1. A user enters a search query, for example, "coronavirus conspiracy theory," into the search bar on their device and presses the send button.
[1415] 2. The search server analyzes the received search query and extracts the related keywords "coronavirus" and "conspiracy theory."
[1416] 3. The search server searches the index database based on the extracted keywords and retrieves relevant search results (websites and videos).
[1417] 4. The search server also retrieves the latest fact-checking data from official databases and stores it in a local cache.
[1418] 5. The search server scrapes the content of the search results and collects it as text data, such as the main text of website articles and video descriptions.
[1419] 6. The search server uses text analysis algorithms to extract important keywords and assertions from the collected text data.
[1420] 7. The search server compares the extracted keywords and claims with fact-checking data, for example, comparing the claim "man-made virus" with the WHO statement "coronavirus is naturally occurring" to identify inconsistencies.
[1421] 8. The search server calculates a confidence score for each search result and displays it as a tag on each search result item.
[1422] Examples of prompt statements
[1423] "What are the steps to verify the credibility of conspiracy theories related to COVID-19?"
[1424] "Please explain specifically how you will conduct fact-checking."
[1425] This system allows users to obtain reliable information quickly and accurately, minimizing the impact of uncertain or false information.
[1426] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1427] Step 1:
[1428] A user enters a search query, such as "COVID-19 conspiracy theory," into the device's search bar and clicks the submit button. The device then sends the search query to a search server. The input data is free-form text, and the output is an HTTP request to the server.
[1429] Step 2:
[1430] The search server analyzes the received search query and extracts relevant keywords. Specifically, it uses the analysis module of a natural language processing (NLP) library (e.g., spaCy) to extract important keywords such as "coronavirus" and "conspiracy theory." The input is free-form text data, and the output is a list of extracted keywords.
[1431] Step 3:
[1432] The search server searches an index database based on the extracted keywords to obtain a list of related websites and videos. The database used for this process is one that has been pre-indexed by a web crawling tool. The input is a list of keywords, and the output is a list of related URLs and titles. Specifically, it includes a website titled "Is Corona a Man-Made Virus?" and a video titled "The Truth About Conspiracy Theories."
[1433] Step 4:
[1434] The search server makes HTTP requests to official databases (e.g., international health organizations or government agencies) to retrieve the latest data for fact-checking. The retrieved data is stored in a local cache for use in the next step. The input is the official API endpoint, and the output is a dataset of trusted documents and statements.
[1435] Step 5:
[1436] The search server uses a scraping tool (e.g., BeautifulSoup) to collect the content of the retrieved websites and videos. Specifically, it analyzes the HTML code of the webpage and extracts the article text and metadata as text data. The input is a list of URLs, and the output is content data in text format. For example, it collects the content of articles on a website titled "Is Corona a Man-Made Virus?"
[1437] Step 6:
[1438] The search server analyzes the collected text data using natural language processing (NLP) algorithms to extract important keywords and claims. Specifically, it uses an NLP library to extract frequently occurring words and phrases in the text and identify important keywords such as "population virus," "evidence," and "research paper." The input is text content data, and the output is a list of important keywords.
[1439] Step 7:
[1440] The search server uses the extracted keywords and claims to compare them with fact-checking data obtained from official databases. For example, it compares the claim "man-made virus" with the WHO statement "coronavirus is naturally occurring" and analyzes contradictions and similarities. The input is a list of keywords and reliable document data, and the output is the comparison results and a reliability score.
[1441] Step 8:
[1442] The search server tags each search result item with the calculated confidence score and generates a list of search results to display to the user. Specifically, the process involves adding the confidence score to the title and URL of each search result to create a list. The input is the comparison result and confidence score, and the output is a tagged list of search results.
[1443] Step 9:
[1444] The search server sends the tagged search result list to the terminal, which receives it and generates an HTML page to display to the user. The input is the tagged search result list, and the output is the search result page displayed to the user.
[1445] Step 10:
[1446] Users check the displayed search results and judge the reliability of the information based on the reliability score. Specifically, they prioritize browsing items with high reliability and acquire or share the information. The input is the displayed search result page, and the output is the user's browsing actions and judgment results.
[1447] This allows users to obtain reliable information quickly and accurately.
[1448] (Application example 1)
[1449] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1450] In recent years, the amount of information on the Internet has increased dramatically, including information of unknown authenticity and potentially misleading information. This has made it difficult for users to easily obtain reliable information. Furthermore, in virtual stores, it is difficult for consumers to judge the reliability of product information themselves, which risks leading to inaccurate purchasing decisions. Therefore, there is a need for a system that can automatically determine the reliability of information and present a reliability score when consumers search for product information in virtual stores.
[1451] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1452] In this invention, the server includes means for acquiring fact-checking data from a database provided by a public institution, means for receiving a search query from a user, means for acquiring related search results based on the search query, means for extracting important keywords and assertions by text-analyzing the content of the search results, means for comparing the extracted keywords and assertions with the fact-checking data, means for calculating and displaying a reliability score for each search result item based on the comparison result, and means for tagging product information with the reliability score, allowing users to judge reliability when making purchasing decisions. This allows consumers to easily judge the reliability of product information in virtual stores and make purchasing decisions with confidence.
[1453] "Public institution" refers to a public institution, such as a government or local government, that provides public functions or services.
[1454] A "database" is an information repository that organizes and stores a collection of data so that it can be efficiently searched and used.
[1455] "Fact-checking data" refers to reliable data used to verify the veracity of information.
[1456] "User" means any individual person or entity that uses the System or Services.
[1457] A "search query" refers to a question or keyword that a user enters into a search engine or system.
[1458] "Search Results" means the information returned by a search engine or system based on a search query.
[1459] "Content" refers to any piece of information or data present on a website or digital media.
[1460] "Text analytics" refers to the techniques and processes used to analyze text data and understand its content.
[1461] "Keywords" refer to key words that play an important role in a text or search query.
[1462] An "argument" is a statement of opinion or perspective on a certain matter.
[1463] "Comparison" refers to the act of comparing two or more things to see their differences and similarities.
[1464] A "trust score" is a numerical indicator that evaluates the reliability of information or data.
[1465] "Display" refers to the means by which information is visually conveyed to the user.
[1466] "Tagging" refers to the act of adding labels and scores to information.
[1467] "Product information" refers to data including descriptions, features, specifications, etc. of the target product.
[1468] "Consumption behavior" refers to the series of actions that consumers take to purchase goods and services.
[1469] This invention describes a system that automatically determines the reliability of product information when a user searches for it in a virtual store and presents a reliability score. This system is composed of the following main components:
[1470] System configuration
[1471] The system is implemented primarily using the following hardware and software:
[1472] 1. User terminal: a smartphone, smart glasses, head-mounted display, or other display device.
[1473] 2. Search server: A server responsible for data analysis and fact-checking.
[1474] 3. Fact-checking database: A database that stores reliable data provided by public institutions.
[1475] Program processing
[1476] User Roles
[1477] A user inputs a search query, such as "antioxidant supplements," into a user terminal and presses a search button. This operation sends the query to a search server.
[1478] The role of the search server
[1479] The search server analyzes the search query received from the user terminal and extracts related keywords, which allows the system to obtain keywords such as "antioxidant" and "supplement."
[1480] Furthermore, the search server searches the virtual store database based on the extracted keywords to obtain a list of related products. For example, product information such as "Antioxidant Supplement A" and "Supplement B" is obtained.
[1481] The search server then accesses databases provided by public organizations (e.g., government agencies and research institutes) to retrieve reliable information and retrieve the latest data for fact-checking, including official reports and the latest research papers.
[1482] Content analysis and fact-checking
[1483] The search server scrapes the descriptions and reviews of each product and collects them as text data. For example, it analyzes the description of "Antioxidant Supplement A" and extracts important keywords and claims.
[1484] Furthermore, the search server compares the extracted keywords and claims with information from databases of public institutions, and calculates a credibility score to evaluate the reliability of the product. For example, a credibility score for a product is determined by comparing the claim "has antioxidant properties."
[1485] Visibility and User Roles
[1486] Finally, the search server tags each product with the calculated reliability score and displays it as information to help users judge reliability when deciding on their purchasing behavior, allowing users to select products with confidence based on reliable information.
[1487] Examples and prompts
[1488] For example, if a user searches for "antioxidant supplements," the search server retrieves related product information, analyzes the descriptions of each product, and calculates a reliability score. As a result, products with a "High Reliability" rating can be presented to the user preferentially.
[1489] Here are some examples of prompts:
[1490] "Antioxidant supplement A has been proven effective in clinical trials. Please compare this information with data from official institutions and calculate a reliability score."
[1491] In this way, this system provides an environment in which consumers can easily judge the reliability of product information in a virtual store.
[1492] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1493] Step 1: Receiving User Input
[1494] A user enters a search query such as "antioxidant supplements" into a device such as a smartphone or head-mounted display and presses the search button. The device then sends this query to a search server. The input data is the user's search query, and the output data is the query information sent to the search server.
[1495] Step 2: Processing the search query
[1496] The search server analyzes the search query received from the device and extracts the related keywords "antioxidant" and "supplement." Specifically, it uses a natural language processing algorithm to tokenize the search query and extract important keywords. The input data is the received query information, and the output data is the extracted keywords.
[1497] Step 3: Getting search results
[1498] The search server searches the virtual store database based on the extracted keywords and retrieves a list of related products. For example, product information such as "antioxidant supplement A" and "supplement B" is retrieved. The input data are the extracted keywords, and the output data is a list of products.
[1499] Step 4: Obtaining fact-checking data
[1500] The search server accesses public databases to retrieve the latest data for fact-checking, including official reports from government agencies and research institutes and the latest research papers. The input data is the access request to the public database, and the output data is the retrieved data for fact-checking.
[1501] Step 5: Analyze your content
[1502] The search server scrapes the descriptions and reviews of each product and collects this content as text data. Specifically, it extracts the text of articles and reviews from each page. The input data is a list of products, and the output data is the collected text data.
[1503] Step 6: Text analysis and keyword extraction
[1504] The search server processes the collected text data using a text analysis algorithm to extract key keywords and claims. Specifically, it uses TF-IDF (Term Frequency-Inverse Document Frequency) to identify important keywords. The input data is the collected text data, and the output data is the extracted keywords and claims.
[1505] Step 7: Conduct a fact check
[1506] The search server compares the extracted keywords and claims with the fact-checking data, checking for keyword matches and inconsistencies, and calculating a reliability score for each product. The input data are the extracted keywords and claims and the fact-checking data, and the output data is a reliability score for each product.
[1507] Step 8: Score tagging and search result generation
[1508] The search server tags each product information with the calculated confidence score and generates a search result list to display to the user. Specifically, it assigns a confidence score to each product information and converts it into a format that can be displayed in the user interface. The input data are the confidence scores and the product list, and the output data is the tagged search result list.
[1509] Step 9: Submit search results
[1510] The search server sends the tagged search results to the terminal, which then displays the received search results to the user. The input data is the tagged search result list, and the output data is the displayed search results.
[1511] Step 10: User View and Decision
[1512] The user checks the displayed search results and judges the reliability of the product information based on the reliability score. Specifically, the user may take action such as prioritizing the purchase of products displayed as "High reliability." The input data are the displayed search results, and the output data are the user's purchasing behavior.
[1513] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1514] This invention combines a search system that performs fact-checking based on highly reliable data provided by public institutions with an emotion engine that recognizes user emotions. The programs required to implement this system and their processing are shown below.
[1515] System configuration
[1516] The system mainly consists of the following components:
[1517] 1. User Device
[1518] 2. Search Server
[1519] 3. Fact-checking databases
[1520] 4. Emotion Engine
[1521] Program processing
[1522] 1. Receiving user input
[1523] The user enters a search query, such as "coronavirus conspiracy theory," into the device and presses the send button.
[1524] 2. Processing search queries
[1525] The terminal transmits the input search query to the server.
[1526] 3. Search Query Analysis
[1527] The server analyzes the received search query and extracts related keywords, for example, "coronavirus" and "conspiracy theory."
[1528] 4. Obtaining search results
[1529] The server searches an index database based on the extracted keywords to retrieve a list of related websites and videos, such as a website titled "Is COVID-19 a man-made virus?" and a video titled "The Truth About Conspiracy Theories."
[1530] 5. Acquisition of data from public institutions
[1531] The server accesses official databases to retrieve the latest data for fact-checking, such as official statements and documents from the WHO.
[1532] 6. Content Analysis
[1533] The server scrapes the content of each website and video and collects it as text data. For example, it extracts the content of an article on a website titled "Is Corona a Man-Made Virus?"
[1534] 7. Text Analysis and Keyword Extraction
[1535] The server processes the collected text data with a text analysis algorithm to extract key keywords and claims, such as "population virus," "evidence," and "research paper."
[1536] 8. Fact Check
[1537] The server compares the extracted keywords and claims against public fact-checking databases—for example, comparing the claim "man-made virus" with the WHO statement "coronavirus is naturally occurring"—and calculates a confidence score for each search result.
[1538] 9. Emotion Recognition with Emotion Engine
[1539] The device uses an emotion engine to analyze the user's emotions based on the user's input and operations (e.g., keystroke speed, input content, facial recognition, etc.). For example, it determines whether the user is feeling stressed.
[1540] 10. Emotional Feedback
[1541] The server then adjusts the reliability score and display order of search results based on the analysis results of the emotion engine. For example, if a user is feeling stressed, more reliable information will be displayed first.
[1542] 11. Score tagging and search result generation
[1543] The server tags each search result item with a confidence score and generates a search result list to display to the user. For example, a site asking "Is COVID-19 a man-made virus?" might be given a "Low confidence" rating.
[1544] 12. Submitting Search Results
[1545] The server sends the search result list with the confidence scores attached to it to the terminal.
[1546] 13. Display of search results
[1547] The device displays the received search results to the user. The user checks the displayed search results and determines the reliability of the information based on the reliability score. For example, the user may prioritize viewing items marked "High Reliability."
[1548] Specific examples
[1549] For example, if a user searches for "COVID-19 conspiracy theory," the search server retrieves related websites and videos and analyzes their content. It then compares them with documents from official institutions and calculates a reliability score. The device also uses an emotion engine to analyze the user's emotions (e.g., anxiety, stress) and adjusts the display order of search results based on the results. As a result, users can easily select reliable information and browse with peace of mind.
[1550] In this way, the system of the present invention provides an environment in which the user can easily obtain highly reliable information, and further provides information according to the user's emotional state.
[1551] The processing flow will be explained below.
[1552] Step 1:
[1553] The user enters a search query into the device, for example, the keyword "coronavirus conspiracy theory."
[1554] Step 2:
[1555] The terminal transmits the input search query to the server.
[1556] Step 3:
[1557] The server analyzes the received search query and extracts related keywords, for example, "coronavirus" and "conspiracy theory."
[1558] Step 4:
[1559] The server searches an index database based on the extracted keywords to retrieve a list of related websites and videos, such as a website titled "Is COVID-19 a man-made virus?" and a video titled "The Truth About Conspiracy Theories."
[1560] Step 5:
[1561] The server accesses official databases to retrieve the latest data for fact-checking, such as official statements and documents from the WHO.
[1562] Step 6:
[1563] The server scrapes the content of each website and video and collects it as text data. For example, it extracts the content of an article on a website titled "Is Corona a Man-Made Virus?"
[1564] Step 7:
[1565] The server processes the collected text data with a text analysis algorithm to extract key keywords and claims, such as "population virus," "evidence," and "research paper."
[1566] Step 8:
[1567] The server compares the extracted keywords and claims against public fact-checking databases, for example, comparing the claim "man-made virus" with the WHO statement "coronavirus is naturally occurring."
[1568] Step 9:
[1569] The server calculates a confidence score for each search result item based on the comparison, for example, assigning a low confidence score to the "man-made virus" claim because it contradicts official data.
[1570] Step 10:
[1571] The device uses an emotion engine to analyze the user's emotions based on the user's input and operations (e.g., keystroke speed, input content, facial recognition, etc.). For example, it can determine whether the user is feeling stressed.
[1572] Step 11:
[1573] The server then adjusts the reliability score and display order of search results based on the analysis results of the emotion engine. For example, if a user is feeling stressed, more reliable information will be displayed first.
[1574] Step 12:
[1575] The server tags each search result item with a confidence score and generates a search result list to display to the user. For example, a site asking "Is COVID-19 a man-made virus?" might be given a "Low confidence" rating.
[1576] Step 13:
[1577] The server sends the search result list with the confidence scores attached to it to the terminal.
[1578] Step 14:
[1579] The terminal displays the received search result list to the user. The user checks the displayed search results and judges the reliability of the information based on the reliability score. For example, the user may preferentially view items marked "High Reliability."
[1580] This series of processes allows the user to easily distinguish highly reliable information, and also makes it possible to provide information that is appropriate for the user's emotional state.
[1581] Example 2
[1582] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1583] In today's world, there is a huge amount of information available on the Internet, but it is not easy to find accurate and reliable information. Furthermore, providing information without considering the user's emotional state can increase stress and anxiety. Therefore, there is a need for a search system that takes into account the user's emotional state while ensuring the reliability of the information.
[1584] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[1585] In this invention, the server includes means for acquiring fact-checking data from a database provided by a public institution, means for receiving a search query from a user, means for acquiring related search results based on the search query, means for analyzing the content of the search results to extract important keywords and claims, means for comparing the extracted keywords and claims with the fact-checking data, means for calculating and displaying a reliability score for each search result item based on the comparison result, means for analyzing user emotions, and means for adjusting the display order of the search results based on the emotion analysis results.
[1586] This allows the user to easily obtain highly reliable information and to be provided with information that corresponds to the user's own emotional state.
[1587] A "database" is a storage device that stores highly reliable data provided by public institutions.
[1588] "Fact-checking data" refers to documents and statements released by public institutions, and is basic data used to verify the accuracy of information.
[1589] A "search query" is a question or keyword that a user enters into the system, and is the search condition that the system investigates.
[1590] "Search results" are lists of information retrieved from databases or the Internet based on a search query.
[1591] "Text analysis" is a technique that uses natural language processing technology to understand the content of text data and extract important keywords and assertions.
[1592] "Keywords" are important words or phrases extracted from a search query or analyzed text.
[1593] "Arguments" are important opinions or points of view in a text that are revealed through text analysis.
[1594] "Comparison" is the process of comparing the fact-checking data with the extracted keywords and claims to confirm whether they match or disagree.
[1595] The "trust score" is a numerical representation of the reliability of information based on the results of fact-checking, and is an indicator of the reliability of search results.
[1596] "Emotion analysis" is the process of analyzing a user's emotional state based on their input and actions, and is a technology that determines what emotions the user is feeling.
[1597] "Display ranking" refers to the order in which each item is displayed in the search result list, and is determined based on confidence scores and sentiment analysis results.
[1598] This invention combines a search system that performs fact-checking based on reliable data provided by public institutions with an emotion engine that recognizes user emotions. This system mainly consists of the following components: a user terminal, a search server, a fact-checking database, and an emotion engine.
[1599] System configuration
[1600] To implement this invention, the following hardware and software are used:
[1601] 1. User Device
[1602] Hardware: Personal computers, smartphones
[1603] Software: Web browser, input interface
[1604] 2. Search Server
[1605] Hardware: High-performance server
[1606] Software: Natural language processing libraries (e.g., NLTK, SpaCy, BERT), search engines (e.g., Elasticsearch, Solr), scraping tools (e.g., BeautifulSoup, Selenium), sentiment analysis APIs (e.g., Microsoft Azure Cognitive Services, Emotion API)
[1607] 3. Fact-checking databases
[1608] Hardware: Database Server
[1609] Software: Public Sector API Access
[1610] 4. Emotion Engine
[1611] Hardware: Embedded systems, cameras and keyboards attached to user devices
[1612] Software: Emotion analysis algorithms, face tracking software (e.g. OpenCV)
[1613] Implementation Procedure
[1614] 1. Receiving user input
[1615] The user enters "COVID-19 conspiracy theory" into the search field and clicks the search button. The device receives this input, temporarily stores it in its internal memory, and prepares to send it to the server.
[1616] 2. Processing search queries
[1617] The device sends the search query entered by the user to the server as an HTTP request, which includes the query text data and metadata such as the user ID.
[1618] 3. Search Query Analysis
[1619] The server parses the incoming HTTP request and processes the query text with a natural language processing algorithm (e.g., Python's NLTK library), which identifies the main keywords "coronavirus" and "conspiracy theory."
[1620] 4. Obtaining search results
[1621] The server queries an index database using a search engine (e.g., Elasticsearch or Solr) based on the extracted keywords, which retrieves a list of related websites and videos.
[1622] 5. Acquisition of data from public institutions
[1623] The server sends API requests to public databases to retrieve the latest, most reliable information, including official statements and research papers from organizations like the WHO and CDC.
[1624] 6. Content Analysis
[1625] The server uses scraping technology (e.g., BeautifulSoup or Selenium) to collect the content of each site and video as text data, extracting the content of the article on the site "Is Corona a Man-Made Virus?" and the subtitles of the video.
[1626] 7. Text Analysis and Keyword Extraction
[1627] The server then processes the collected text data again using a natural language processing algorithm (e.g., SpaCy or BERT) to extract important keywords and assertions, such as "population virus," "evidence," and "research paper."
[1628] 8. Fact Check
[1629] The server compares the extracted keywords and claims with official data it has obtained—for example, comparing the claim "man-made virus" with the WHO statement "coronavirus is naturally occurring"—and then uses an algorithm to calculate a confidence score for each search result, which is calculated on a scale from 0 to 100.
[1630] 9. Emotion Recognition with Emotion Engine
[1631] The device uses an emotion engine to analyze the user's emotions based on the user's input and operations, such as keystroke speed and facial recognition (using OpenCV), to determine whether the user is feeling stressed or anxious.
[1632] 10. Emotional Feedback
[1633] The server then adjusts the reliability score and display order of search results based on the analysis results of the emotion engine. For example, if a user is feeling stressed, an algorithm is applied to prioritize displaying information with high reliability.
[1634] 11. Score tagging and search result generation
[1635] The server tags each search result with a confidence score and generates a list of search results to display to the user. For example, a website asking "Is COVID-19 a man-made virus?" might be tagged with "Confidence: 20 (Low)."
[1636] 12. Submission and Display of Search Results
[1637] The server sends the search result list, with assigned reliability scores, to the terminal as an HTTP response. The terminal displays the received search result list on its user interface. The user checks this display and decides which information to view based on the reliability score. For example, they may prioritize clicking and viewing items with a "Reliability: 90" rating.
[1638] Specific examples
[1639] For example, if a user searches for "coronavirus conspiracy theory," the following specific actions occur: The user enters "coronavirus conspiracy theory" into the search bar and clicks the search button. The device receives the input, generates an HTTP request, and sends it to the server. The server analyzes the query and extracts related keywords. The server uses a search engine to retrieve relevant sites and videos. The server retrieves reliable data from public institution databases via an API request. The server scrapes the content of the retrieved sites and videos and collects it as text data. The server uses a natural language processing algorithm to extract important keywords and claims from the collected text data. The server compares the extracted keywords and claims with public institution data and calculates a confidence score. The device analyzes the user's emotions using an emotion engine and sends the results to the server. The server adjusts the confidence score and display ranking of the search results based on the user's emotions. The server tags the confidence score and generates a search result list. The server sends the final search result list to the device. The device displays the search result list on a user interface, and the user can review and select results.
[1640] Example prompt sentence:
[1641] "New coronavirus conspiracy theory"
[1642] Result: A list of highly reliable information is displayed, including an official statement from a public institution that "coronavirus is a natural occurrence." Because the emotion engine determined that the user was highly stressed, highly reliable information is displayed first.
[1643] In this way, the system of the present invention provides an environment in which the user can easily obtain highly reliable information, and further provides information according to the user's emotional state.
[1644] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1645] Step 1:
[1646] The user enters "COVID-19 conspiracy theory" in the search field and clicks the search button. This causes the device to receive a search query. The input is the search query entered by the user, and the output is the query temporarily stored in the internal memory. The device prepares to send this data to the server.
[1647] Step 2:
[1648] The device sends the search query entered by the user to the server as an HTTP request. The input is the search query stored in the internal memory, and the output is the HTTP request sent to the server. The request includes the text data of the query and metadata such as the user ID.
[1649] Step 3:
[1650] The server parses the received HTTP request and processes the query text with an NLP algorithm (e.g., Python's NLTK library). The input is the HTTP request sent to the server, and the output is the parsed search query and keywords extracted from it, e.g., "coronavirus" and "conspiracy theory."
[1651] Step 4:
[1652] The server queries the index database using Elasticsearch or Solr based on the extracted keywords. The input is the extracted keywords, and the output is a list of related websites and videos, such as a website called "Is COVID-19 a man-made virus?" or a video called "The Truth About Conspiracy Theories."
[1653] Step 5:
[1654] The server sends API requests to public databases to retrieve the latest, most reliable information. The input is the API request, and the output is the retrieved public data, such as an official WHO statement or a research paper.
[1655] Step 6:
[1656] The server collects the content of each website and video as text data using scraping technology (e.g., BeautifulSoup or Selenium). The input is the URL of the relevant website or video, and the output is the scraped text data, such as the content of an article on the website "Is Coronavirus a Man-Made Virus?"
[1657] Step 7:
[1658] The server processes the collected text data again with an NLP algorithm (e.g., SpaCy or BERT) to extract important keywords and claims. The input is the collected text data, and the output is the extracted important keywords and claims, such as "population virus," "evidence," and "research paper."
[1659] Step 8:
[1660] The server compares the extracted keywords and claims with the retrieved official data. The input is the extracted keywords and claims and the official data, and the output is a confidence score for each search result, e.g., "Confidence: 20 (Low)."
[1661] Step 9:
[1662] The device uses an emotion engine to analyze the user's emotions based on their input and operations. Specifically, it uses keystroke speed and facial recognition technology (e.g., OpenCV). The input is the user's input data and operation data, and the output is the analyzed user's emotional state, such as "high stress."
[1663] Step 10:
[1664] The server adjusts the reliability score and display order of search results based on the analysis results of the emotion engine. The input is the emotion analysis result and reliability score, and the output is an adjusted search result list, for example, a list in which highly reliable information is displayed at the top.
[1665] Step 11:
[1666] The server tags each search result with a confidence score and generates a final search result list to display to the user. The input is the adjusted search result list, and the output is a search result list with a confidence score, e.g., "Confidence: 90 (High)".
[1667] Step 12:
[1668] The server sends the final search result list to the terminal. The input is the search result list with the assigned confidence scores, and the output is the HTTP response sent to the terminal.
[1669] Step 13:
[1670] The terminal displays the received search result list on the user interface. The input is the HTTP response sent to the terminal, and the output is the search result list displayed on the screen. The user checks the results and decides which information to view based on the confidence score.
[1671] (Application example 2)
[1672] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1673] In today's information society, it is extremely important for users to have access to reliable information. However, while a vast amount of information is easily available, it is often difficult to verify its reliability. Furthermore, there are no established methods for providing appropriate information to users who are feeling stressed or anxious. Therefore, an objective of the present invention is to provide a system that provides users with reliable information and also provides information that is tailored to the user's emotional state.
[1674] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for acquiring fact-checking data from a database provided by a public institution, means for receiving a search query from a user, means for acquiring related search results based on the search query, means for text-analyzing the content of the search results to extract important keywords and assertions, means for comparing the extracted keywords and assertions with the fact-checking data, means for calculating and displaying a reliability score for each search result item based on the comparison result, means for recognizing passenger emotions using an emotion engine, and means for adjusting the reliability score and information display based on the recognized emotions. This allows the user to appropriately acquire reliable information and provides information according to the user's emotional state.
[1675] A "public institution" is an organization that operates for the public interest, such as a government or local government, and issues official data and statements.
[1676] A "database" is a collection of information designed to efficiently store, retrieve, and update data.
[1677] "Fact-checking data" is reliable information and materials collected and provided for the purpose of fact-checking.
[1678] "User" means an individual or organization that uses a system or application.
[1679] A "search query" is a string of characters or keywords that a user enters when searching for information.
[1680] "Search results" are a collection of information or data retrieved based on a search query.
[1681] "Text analytics" refers to techniques and methods for extracting important information and patterns from text data.
[1682] "Keywords" are important words or phrases used to represent specific information.
[1683] An "assertion" is an expression that expresses a particular opinion or point of view.
[1684] "Comparison" is the act of comparing two or more pieces of information or data and evaluating their differences and similarities.
[1685] A "trust score" is a numerical indicator used to evaluate the reliability of information.
[1686] "Display" is the act of visually presenting acquired information or results to the user.
[1687] An "emotion engine" is software or hardware that analyzes and recognizes a user's emotions.
[1688] "Recognition" is the act of analyzing, understanding, and grasping information and data.
[1689] "Adjustment" is the act of changing information or settings to accommodate specific conditions or circumstances.
[1690] "Means" are the methods or techniques used to achieve a particular goal.
[1691] The present invention is a system for providing highly reliable information to a user and providing information according to the user's emotional state. Specific embodiments of the present invention will be described below.
[1692] System Configuration
[1693] Hardware
[1694] 1. User device: A device such as a smartphone, tablet, or PC that accepts user input and operations. It also has a camera and microphone to collect data for emotion recognition.
[1695] 2. Search Server: A central server that searches, analyzes, and fact-checks various data.
[1696] 3. Fact-checking database: A database that stores reliable data provided by public institutions.
[1697] 4. Emotion engine: Software for analyzing the user's emotional state.
[1698] software
[1699] 1. EmotionEngine: Software that analyzes data input from cameras and microphones in real time to recognize the user's emotions. It uses image processing libraries such as OpenCV.
[1700] 2. FactCheckEngine: An engine that cross-checks the information obtained with public data and evaluates its reliability.
[1701] 3. PublicDataFetcher: Software for retrieving up-to-date and reliable data from public databases.
[1702] System Operation
[1703] The server first receives a search query from a user device, then retrieves relevant search results based on the search query, performs text analysis on the content, extracts important keywords and claims, and compares the extracted keywords and claims with fact-checking data to calculate their reliability score.
[1704] At the same time, EmotionEngine recognizes the user's emotions in real time using the camera and microphone installed on the user's device. Based on the recognized emotions, the server adjusts the reliability score and information display to provide the user with appropriate information.
[1705] Specific examples
[1706] For example, if a user enters the search query "traffic congestion current situation," the server retrieves related websites and news articles, performs fact-checking, and calculates and displays a reliability score. If the emotion engine recognizes that the user is feeling stressed, it prioritizes information with high reliability and suggests relaxing music and videos.
[1707] Prompt Sentence Examples
[1708] Provide information to promote safe driving when users are under stress, such as prioritizing public traffic and weather information and playing relaxing music and scenic videos.
[1709] In this way, a system is realized that allows users to easily obtain highly reliable information and receive information that is optimal for their emotional state.
[1710] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1711] Step 1:
[1712] User enters search query and submits
[1713] A user inputs a query, for example, "traffic congestion current situation," into the terminal and presses the send button. After the input is made, the terminal sends the search query to the server. The input is in the form of text, and the output is the search query sent to the server.
[1714] Step 2:
[1715] The server parses the search query
[1716] The server analyzes the received search query and extracts relevant keywords. For example, it extracts the keywords "traffic congestion" and "current situation." The input is the search query text data, and the output is the analyzed keyword list. The server performs this process using a text analysis algorithm.
[1717] Step 3:
[1718] The server retrieves the search results
[1719] The server searches the index database based on the extracted keywords to obtain a list of relevant sites and news articles. The input is the parsed keyword list, and the output is a list of relevant search results. The server uses the index database to perform this process.
[1720] Step 4:
[1721] Server acquires data from public institutions
[1722] The server accesses a public database to retrieve the latest data for fact-checking, for example, data from a traffic information center. The input is the URL or API key of the public database to access, and the output is the trusted data. The server performs this process using an HTTP request.
[1723] Step 5:
[1724] The server analyzes the content
[1725] The server scrapes the content of each website or news article and collects it as text data. The input is a list of relevant search results, and the output is the retrieved text data. The server performs this process using web scraping technology.
[1726] Step 6:
[1727] The server performs text analysis and keyword extraction.
[1728] The server processes the collected text data with text analysis algorithms to extract important keywords and assertions. The input is the collected text data, and the output is the extracted keywords and assertions. The server performs this processing using natural language processing (NLP) techniques.
[1729] Step 7:
[1730] Server fact-checking
[1731] The server compares the extracted keywords and claims against public fact-checking databases and calculates a confidence score. The inputs are the extracted keywords and data from the public databases, and the output is a confidence score. The server performs this process using database queries and comparison algorithms.
[1732] Step 8:
[1733] The device uses an emotion engine to recognize emotions.
[1734] The device analyzes the user's emotions using a camera and microphone. The input is video and audio data, and the output is the recognized emotion data. The device performs this processing using the EmotionEngine.
[1735] Step 9:
[1736] Server provides feedback based on emotions
[1737] The server adjusts the confidence score and information display based on the analysis results of the emotion engine. The input is the recognized emotion data and confidence score, and the output is the adjusted information display ranking and feedback. The server performs this process using an algorithm.
[1738] Step 10:
[1739] The server generates a search result list tagged with a confidence score
[1740] The server tags each search result item with a confidence score and generates a search result list for display to the user. The input is the confidence score and the list of search results, and the output is the tagged search result list. The server performs this process using a list generation algorithm.
[1741] Step 11:
[1742] The server sends the search result list
[1743] The server sends the search result list with attached confidence scores to the user device. The input is the tagged search result list, and the output is the data sent to the user device. The server performs this process using a network protocol.
[1744] Step 12:
[1745] Your device will display the search results
[1746] The terminal displays the received search results to the user. The user reviews the displayed search results and determines the reliability of the information based on the reliability score. The input is the search result list received from the server, and the output is the displayed search results. The terminal performs this process using a graphical user interface (GUI).
[1747] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.
[1748] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1749] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.
[1750] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[1751] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.
[1752] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.
[1753] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).
[1754] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.
[1755] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."
[1756] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.
[1757] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).
[1758] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.
[1759] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.
[1760] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[1761] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.
[1762] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.
[1763] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.
[1764] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.
[1765] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.
[1766] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.
[1767] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.
[1768] The following is further disclosed regarding the above embodiment.
[1769] (Claim 1)
[1770] A means of obtaining fact-checking data from databases provided by public institutions; and
[1771] a means for receiving a search query from a user;
[1772] means for obtaining relevant search results based on the search query;
[1773] means for extracting important keywords and assertions by text-analyzing the content of the search results;
[1774] means for comparing the extracted keywords and assertions against the fact-checking data;
[1775] means for calculating and displaying a confidence score for each search result item based on the comparison results;
[1776] A system including:
[1777] (Claim 2)
[1778] 10. The system of claim 1, wherein the fact-checking data includes documents and statements released by public institutions.
[1779] (Claim 3)
[1780] 10. The system of claim 1, further comprising means for displaying the search result confidence score as a tag on each search result item.
[1781] "Example 1"
[1782] (Claim 1)
[1783] means for receiving a search query from a user device;
[1784] means for analyzing the received search query to extract relevant keywords;
[1785] A way to obtain a list of related websites and videos based on the extracted keywords, and
[1786] A means of obtaining fact-checking data from databases provided by public institutions; and
[1787] A means of scraping the acquired content and collecting it as text data;
[1788] A means for analyzing the collected text data using a text analysis algorithm and extracting important keywords and assertions;
[1789] A means of comparing extracted keywords and claims against fact-checking data;
[1790] means for calculating and displaying a confidence score for each search result item based on the comparison results;
[1791] A system including:
[1792] (Claim 2)
[1793] 10. The system of claim 1, wherein the fact-checking data includes documents and statements released by public institutions.
[1794] (Claim 3)
[1795] 10. The system of claim 1, further comprising means for displaying the search result confidence score as a tag on each search result item.
[1796] "Application Example 1"
[1797] (Claim 1)
[1798] A means of obtaining fact-checking data from databases provided by public institutions; and
[1799] a means for receiving a search query from a user;
[1800] means for obtaining relevant search results based on the search query;
[1801] means for extracting important keywords and assertions by text-analyzing the content of the search results;
[1802] means for comparing the extracted keywords and assertions against the fact-checking data;
[1803] means for calculating and displaying a confidence score for each search result item based on the comparison results;
[1804] a means for tagging the reliability score to product information and determining reliability when a user decides on a consumption behavior;
[1805] A system including:
[1806] (Claim 2)
[1807] 10. The system of claim 1, wherein the fact-checking data includes documents and statements released by public institutions.
[1808] (Claim 3)
[1809] 10. The system of claim 1, further comprising means for displaying the search result confidence score as a tag on each search result item.
[1810] "Example 2: Combining Emotion Engines"
[1811] (Claim 1)
[1812] A means of obtaining fact-checking data from databases provided by public institutions; and
[1813] a means for receiving a search query from a user;
[1814] means for obtaining relevant search results based on the search query;
[1815] means for extracting important keywords and assertions by text-analyzing the content of the search results;
[1816] means for comparing the extracted keywords and assertions against the fact-checking data;
[1817] means for calculating and displaying a confidence score for each search result item based on the comparison results;
[1818] means for analyzing user emotions;
[1819] means for adjusting the display order of search results based on the emotion analysis results;
[1820] A system including:
[1821] (Claim 2)
[1822] 10. The system of claim 1, wherein the fact-checking data includes documents and statements released by public institutions.
[1823] (Claim 3)
[1824] 2. The system according to claim 1, further comprising: means for displaying the confidence score of the search result as a tag on an item of each search result; and means for adjusting a display order based on the sentiment analysis result.
[1825] "Application example 2 when combining emotion engines"
[1826] (Claim 1)
[1827] A means of obtaining fact-checking data from databases provided by public institutions; and
[1828] a means for receiving a search query from a user;
[1829] means for obtaining relevant search results based on the search query;
[1830] means for extracting important keywords and assertions by text-analyzing the content of the search results;
[1831] means for comparing the extracted keywords and assertions against the fact-checking data;
[1832] means for calculating and displaying a confidence score for each search result item based on the comparison results;
[1833] means for recognizing passenger emotions using an emotion engine;
[1834] means for adjusting the confidence score and / or information display based on the recognized emotion;
[1835] A system including:
[1836] (Claim 2)
[1837] 10. The system of claim 1, wherein the fact-checking data includes data and statements released by public authorities.
[1838] (Claim 3)
[1839] 10. The system of claim 1, further comprising means for displaying the search result confidence score as a tag on each search result item. [Explanation of symbols]
[1840] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>
Claims
1. A means of obtaining fact-checking data from databases provided by public institutions; and a means for receiving a search query from a user; means for obtaining relevant search results based on the search query; means for extracting important keywords and assertions by text-analyzing the content of the search results; means for comparing the extracted keywords and assertions against the fact-checking data; means for calculating and displaying a confidence score for each search result item based on the comparison results; A system including:
2. The system of claim 1 , wherein the fact-checking data includes documents and statements released by public institutions.
3. The system of claim 1 , further comprising means for displaying the search result confidence score as a tag on each search result item.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A