system
A system for evaluating data credibility through data collection, preprocessing, and generative model analysis provides users with credibility scores, addressing the challenge of unreliable Internet data and enabling quick, accurate decision-making.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-08-20
- Publication Date
- 2026-03-05
AI Technical Summary
In today's Internet environment, the accuracy and reliability of data obtained from various sources vary, posing a risk for making decisions based on unreliable information, necessitating a means to effectively evaluate the reliability of data and make quick and accurate decisions based on highly accurate and reliable statistical data.
A system that includes data collection, extraction, preprocessing, credibility evaluation using a generative model, credibility score calculation, and user interface for visually displaying the credibility score, allowing users to make informed decisions.
Enables users to quickly and accurately evaluate the credibility of data from diverse sources and make decisions based on highly reliable information.
Smart Images

Figure 2026036293000001_ABST
Abstract
Description
[Technical Field]
[0001] The technology of the present disclosure relates to a system. [Background technology]
[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]
[0004] In today's Internet environment, the accuracy and reliability of data obtained from various sources varies, and there is a risk of making decisions based on unreliable information. In such a situation, there is a need for a means for individuals and companies to effectively evaluate the reliability of data and make quick and accurate decisions based on highly accurate and reliable statistical data. [Means for solving the problem]
[0005] The present invention provides a system including a data collection means, a means for extracting necessary statistical data from the collected data, a means for preprocessing the extracted data, a means for analyzing the preprocessed data using a generative model for evaluating the credibility of the data, a means for calculating a credibility score based on the analysis results, and a means for providing the credibility score and data to a user. This system evaluates the credibility of data by taking into account the reliability of the information source, past revision history, data consistency, etc., allowing users to make decisions based on highly reliable information. The system also has a user interface for visually displaying the credibility score, allowing users to easily confirm the credibility of the data.
[0006] "Data collection means" refers to a function for automatically acquiring data from various sources on the Internet.
[0007] "Statistical data" refers to numerical data and the results of analysis relating to specific events or phenomena.
[0008] "Extraction means" is a function that extracts necessary information from collected data.
[0009] The "preprocessing means" is a function that removes noise from the extracted data and standardizes the format to prepare the data in an analyzable format.
[0010] A "generative model" is an analytical model that uses machine learning and artificial intelligence techniques to evaluate the credibility of data.
[0011] "Analysis means" is a function that evaluates the credibility of data using a generative model.
[0012] A "credibility score" is a numerical indicator of the reliability of data, expressed on a scale of 0 to 100.
[0013] "Means to provide to users" refers to a function for communicating the analysis results and credibility scores to users.
[0014] A "user interface" is an interaction means such as a screen or input device that allows a user to interact with a system. [Brief explanation of the drawings]
[0015] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14]FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION
[0016] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.
[0017] First, the terms used in the following description will be explained.
[0018] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).
[0019] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.
[0020] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.
[0021] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.
[0022] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."
[0023] [First embodiment]
[0024] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0025] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0026] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0027] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.
[0028] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0029] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0030] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.
[0031] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0032] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0033] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0034] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0035] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0036] The system related to this invention, "DataVerity," collects data from various sources, evaluates its credibility using a generative AI model, and provides users with the data and a credibility score. The program processing content and specific examples of this system are shown below.
[0037] Data collection
[0038] The server collects data from various sources on the Internet, including public government databases, internal corporate databases, news sites, social media, etc. Using APIs and web scraping technology, the server automatically retrieves the necessary data and stores it in a cloud database.
[0039] Data extraction and preprocessing
[0040] The server extracts the necessary statistical data from the stored database. This data extraction is performed using database queries and natural language processing techniques. The extracted data undergoes preprocessing, such as noise removal, missing value completion, and format standardization. For example, the data is converted into an analyzable format by standardizing date formats and converting numerical units.
[0041] Analysis for credibility assessment
[0042] The server uses a generative AI model to evaluate the credibility of the pre-processed data. The model analyzes the credibility of the data based on multiple evaluation criteria, such as the reliability of the source, past revision history, and data consistency, and then calculates a credibility score for each piece of data.
[0043] Providing and displaying credibility scores
[0044] When a user wants to check data from their device, they send a request. Based on the request, the server searches the database for the relevant data and its credibility score, and sends it back to the device. The device then visually displays the received data and credibility score to the user. The display format is easy to understand using graphs and tables, and the credibility score is color-coded. This allows the user to visually check the credibility of the data.
[0045] Specific examples
[0046] For example, consider the case of collecting data on the number of new coronavirus infections and evaluating its credibility. The server collects data on the number of new coronavirus infections from official government databases, major news sites, and social media. Next, the server extracts the number of new infections, deaths, and recovered cases from the collected data, standardizing the date format and removing noise.
[0047] The server then uses a generative AI model to analyze the credibility of the data from each source. Data from public databases is assigned a high score (e.g., 90 points or higher) because it has high reliability, data from news sites is assigned a medium score (e.g., 70 to 80 points), and data from anonymous posts on social media is assigned a low score (e.g., 40 to 60 points).
[0048] Finally, when a user requests data on the number of coronavirus cases from their device, the server sends the data and credibility score to the device based on the request. The device then displays this information visually in graphs and tables, and users can click on specific data to view detailed analysis results.
[0049] As described above, this system enables users to easily evaluate the credibility of data collected from a wide variety of information sources and make decisions quickly and accurately based on highly reliable data.
[0050] The processing flow will be explained below.
[0051] Program processing flow
[0052] Step 1:
[0053] The server collects data from designated sources (public government databases, internal corporate databases, news sites, social media, etc.), automatically retrieves the data using APIs and web scraping technology, and stores it in a cloud database.
[0054] Step 2:
[0055] The server extracts the statistical data the user needs from the data stored in the database, using database queries and natural language processing techniques to retrieve information based on specific keywords or categories.
[0056] Step 3:
[0057] The server preprocesses the extracted data, removing noise (e.g., removing irrelevant and duplicate data), imputing missing values, and standardizing data formats (e.g., converting date formats and numeric units).
[0058] Step 4:
[0059] The server inputs the preprocessed data into a generative AI model to evaluate its credibility. The generative AI model analyzes the data based on multiple criteria, including the reliability of the source, past revision history, and data consistency.
[0060] Step 5:
[0061] The server calculates a credibility score for each piece of data based on the analysis results of the generative AI model. The credibility score is expressed in the range of 0 to 100 and is assigned individually to each piece of data.
[0062] Step 6:
[0063] The server stores the calculated credibility score and the reason for the evaluation in a database, which allows for quick responses to subsequent user requests.
[0064] Step 7:
[0065] When a user wants to check statistical data from their device, they send a request for specific information, including search keywords and the range of data they require.
[0066] Step 8:
[0067] The server searches the database for relevant statistical data and credibility scores based on the user request, and returns them to the terminal together with detailed analysis results.
[0068] Step 9:
[0069] The device visually displays the received data and the credibility score to the user in the form of graphs and tables, and the credibility score is color-coded for easy understanding.
[0070] Step 10:
[0071] If a user wants to check the details of a particular piece of data, they can click on it to see detailed analysis results and the reasons for the evaluation, allowing users to understand the credibility of the data and make decisions based on reliable information.
[0072] Through these processing steps, the DataVerity system helps users effectively evaluate the credibility of data from a wide variety of sources and make decisions quickly and accurately.
[0073] Example 1
[0074] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0075] There is a need for an appropriate system that can quickly and accurately evaluate the credibility of data collected from various sources and provide the results visually to users. However, current systems require a lot of manual work in the collection, preprocessing, credibility evaluation, and provision of data, which is time-consuming and labor-intensive, and has low reliability. Therefore, an efficient and reliable data collection and evaluation system is needed.
[0076] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0077] In this invention, the server includes a data collection means, a means for extracting necessary information from the collected data, a means for preprocessing the extracted data, a means for analyzing the preprocessed data using a generative model for evaluating the credibility of the data, a means for calculating a credibility score based on the analysis result, and a means for providing the credibility score and data to a user. This enables a user to quickly and accurately collect data from various information sources, evaluate the credibility of the data, and then visually confirm it.
[0078] "Data collection means" refers to a function for automatically acquiring data from various sources on the Internet.
[0079] "Means for extracting necessary information from collected data" refers to a function for selecting specific information from collected data and formatting it into a format suitable for data analysis.
[0080] "Means for preprocessing extracted data" refers to functions for removing noise, filling in missing values, standardizing formats, etc. from the extracted data, and preparing it in a format suitable for analysis.
[0081] "Means for performing analysis using a generative model to assess the credibility of preprocessed data" refers to a function that performs data analysis using a generative AI model to assess the reliability of preprocessed data.
[0082] The "means for calculating the credibility score" is a function for quantifying and indicating the reliability of data evaluated using a generative model.
[0083] "Means for providing credibility scores and data to users" refers to a function for providing analyzed credibility scores and related data to users in a visually verifiable format.
[0084] The system of the present invention collects data from various sources, evaluates its credibility using a generative AI model, and provides users with the data and a credibility score. Specific embodiments and methods of use are described below.
[0085] The system includes the following main elements:
[0086] Data collection methods
[0087] A means of extracting the necessary information from the collected data
[0088] A means of preprocessing the extracted data
[0089] A means of analyzing using generative models to assess the veracity of preprocessed data
[0090] A method for calculating a credibility score based on the analysis results
[0091] Means of providing credibility scores and data to users
[0092] Data collection methods
[0093] The server automatically collects data from various sources on the Internet using API access and web scraping techniques. For example, it uses the Python libraries BeautifulSoup and Scrapy to retrieve data from public databases, news sites, and social media.
[0094] A means of extracting the necessary information from the collected data
[0095] The server selects specific information from the collected data and formats it into a format suitable for data analysis. Specifically, it uses SQL queries to extract data such as the number of new infections, deaths, and recovered cases from the database.
[0096] A means of preprocessing the extracted data
[0097] The server performs preprocessing on the extracted data, specifically removing noise, filling in missing values, standardizing formats, etc. By using the Python Pandas library, the server consistently standardizes date formats and converts numeric units.
[0098] A means of analyzing using generative models to assess the veracity of preprocessed data
[0099] The server then inputs the preprocessed data into a generative AI model to assess its credibility. This generative model takes into account the reliability of the source, past revision history, and data consistency, and uses natural language processing (NLP) techniques, such as Transformer-based models (such as BERT and GPT).
[0100] A method for calculating a credibility score based on the analysis results
[0101] The server calculates a credibility score based on the analysis results of the generative model. This score is expressed as a number and serves as a standard for evaluating the reliability of the data. For example, public databases have a high score (90 points or above), news sites have a medium score (70-80 points), and social media data has a low score (40-60 points).
[0102] Means of providing credibility scores and data to users
[0103] When a user uses a device to send a data verification request, the server searches for and returns the relevant data and credibility score. The device then visually displays this data. Specifically, the received data is displayed in graphs and tables, and the credibility score is color-coded. Users can click on the data to view detailed analysis results.
[0104] Examples of concrete examples and prompts
[0105] For example, consider the case of evaluating the credibility of data on the number of people infected with the new coronavirus. In this case, the server collects relevant data from official government databases, major news sites, and social media, and inputs it into a generative AI model for analysis. After calculating a credibility score, it displays it on the device. The following is an example of a specific prompt sentence:
[0106] "Collect data on the number of COVID-19 cases and rate its credibility. Collect data from each source, use a generative AI model to score its credibility, and display it visually."
[0107] As described above, the present invention enables users to easily evaluate the credibility of data collected from a variety of information sources and make decisions quickly and accurately.
[0108] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0109] Step 1:
[0110] The server collects data from various sources on the Internet, specifically using API access and web scraping technology. It uses a list of URLs from public government databases, news sites, social media, etc. as input, and automatically retrieves the latest data from them. It stores the retrieved data in a cloud database. The output is a collection of the collected raw data.
[0111] Specifically, data extraction is performed automatically using Python libraries such as BeautifulSoup and Scrapy.
[0112] Step 2:
[0113] The server extracts the necessary information from the collected raw data. It uses the raw data as input and uses database queries and natural language processing (NLP) techniques to select specific data such as the number of new cases, deaths, and recoveries. The output is the extracted data.
[0114] Specifically, it uses SQL queries to extract specific columns and rows from a database and then uses NLP techniques to analyze the text data.
[0115] Step 3:
[0116] The server performs preprocessing on the extracted data. It uses the extracted data as input and performs noise removal, missing value imputation, format standardization, etc. The output is preprocessed data that can be analyzed.
[0117] Specifically, it uses Python's Pandas library to clean the data and standardize date formats and numerical units.
[0118] Step 4:
[0119] The server inputs the preprocessed data into a generative AI model to analyze it and evaluate its credibility. As input, the preprocessed data is provided to a generative model (e.g., BERT or GPT), and the model evaluates the credibility of the data. The output is a credibility score for each data point.
[0120] Specifically, it uses a generative model to extract features from the data and quantify their reliability.
[0121] Step 5:
[0122] The server calculates a credibility score based on the analysis results of the generative model. It uses the evaluation results of the generative model as input and applies an algorithm to calculate the credibility score. The output is a dataset containing the calculated credibility scores.
[0123] Specifically, the evaluation algorithm is executed to quantify the reliability and store it in a database.
[0124] Step 6:
[0125] When a user uses a terminal to send a request for data verification, the server searches for the relevant data and credibility score and returns it.The server receives the user's request as input and searches the database for the corresponding data and credibility score.The output is the search result: the data and credibility score.
[0126] Specifically, it receives an HTTP request from a user and executes a corresponding SQL query on a database.
[0127] Step 7:
[0128] The terminal visually displays the received data and credibility scores. It uses the data and credibility scores received from the server as input and displays them in the form of graphs and tables. The output is a visual representation of the data provided to the user.
[0129] Specifically, the system uses JavaScript (registered trademark) and HTML to display data on a web page, color-coding the reliability to make it visually easy to understand, and users can click to see detailed analysis results.
[0130] (Application example 1)
[0131] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0132] In modern society, there is a huge amount of information available on the Internet, making it extremely difficult to obtain accurate and reliable information from all that information. News and important data, in particular, often contain unreliable information, requiring users to expend considerable effort to determine the accuracy of the information. Furthermore, many news articles are constantly being updated, creating a lack of means to assess their credibility in real time. Therefore, there is a need for a system that can quickly and accurately assess the credibility of information and provide it to users.
[0133] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0134] In this invention, the server includes a data collection means, a means for extracting necessary statistical data from the collected data, a means for pre-processing the extracted data, a means for analyzing the pre-processed data using a generative model for evaluating the credibility of the data, a means for calculating a credibility score based on the analysis result, and a means for providing the credibility score and data to a user in a news feed format, thereby enabling the user to evaluate the credibility of news articles in real time and make quick and accurate decisions based on highly reliable information.
[0135] "Data collection means" refers to technology for automatically obtaining necessary data from various information sources on the Internet.
[0136] "Means for extracting statistical data" refers to techniques for extracting necessary statistical information from collected data.
[0137] "Preprocessing means" refers to technology that converts extracted data into an analyzable format by removing noise, filling in missing values, and standardizing the format.
[0138] "Means for performing analysis using generative models" refers to techniques for evaluating the veracity of data that has been preprocessed using a generative AI model.
[0139] The "means for calculating the credibility score" is a technology that quantifies the reliability of each piece of data based on the analysis results of the generative AI model.
[0140] "Newsfeed-style provision" refers to a technology that visually provides users with credibility scores and data in real time.
[0141] "Source credibility" is a criterion for evaluating the past performance and reliability of the source of data.
[0142] "Edit history" is historical information about how data has been edited in the past.
[0143] "Data consistency" is a criterion for assessing the consistency and agreement of data obtained from multiple sources.
[0144] "User interface" refers to the screen or operating means through which a user can interact with the system and visually check the credibility score of data.
[0145] A system for realizing this invention includes a data collection means, a means for extracting statistical data, a data preprocessing means, a means for performing credibility assessment using a generative AI model, a means for calculating a credibility score, and a means for providing the data and the credibility score in a news feed format.
[0146] Program processing and technologies used
[0147] Data collection methods
[0148] The server collects data from various sources on the Internet, including public government databases, corporate databases, news sites, and social media, using APIs and web scraping techniques. Specifically, web scraping is performed using Python's requests and BeautifulSoup libraries.
[0149] A means of extracting statistical data
[0150] The server extracts the necessary statistics from the collected data, using database queries and natural language processing (NLP) techniques to extract key information such as article titles and content, for example using the Python nltk library.
[0151] Data preprocessing measures
[0152] The server removes noise from the extracted data, imputes missing values, and standardizes the format. Specifically, it standardizes date formats and converts numerical units. This converts the data into an analyzable format. For example, it uses the pandas library.
[0153] Credibility assessment tools
[0154] The server evaluates the credibility of the preprocessed data using a generative AI model, which analyzes the credibility of the data based on criteria such as the reliability of the source, past revision history, and data consistency. For example, it applies a logistic regression model using scikit-learn.
[0155] How the credibility score is calculated
[0156] The server calculates a credibility score based on the analysis results of the generative AI model. This score is a numerical representation of the reliability of the data, and is expressed, for example, in the range of 0 to 1.
[0157] A means to provide data and credibility scores in a news feed format
[0158] The system allows users to check data and credibility scores in a news feed format on their devices. The credibility scores are displayed in color, and by clicking on specific data, detailed analysis results are displayed. For example, using Django as the web application framework, the data and scores are displayed visually in graphs and tables.
[0159] Specific examples
[0160] For example, if a user wants to check the latest news on the number of new coronavirus infections, the server collects relevant data from public databases, news sites, and social media. The generative AI model assigns a credibility score to each source, assigning a high score (90 points or above) to data from public databases, a medium score (70 to 80 points) to data from news sites, and a low score (40 to 60 points) to data from social media. Users can launch the app on their smartphone and check the credibility score of each article along with the data displayed in a news feed format.
[0161] Prompt Sentence Examples
[0162] "What is the credibility score for the latest news on the number of new coronavirus cases?"
[0163] As described above, this invention collects data from various sources on the Internet, evaluates its credibility, and provides it to users, enabling them to make decisions based on highly reliable information in real time.
[0164] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0165] Step 1:
[0166] The server collects data from various sources on the Internet. This includes scraping and using APIs to obtain data from public government databases, news sites, social media, etc. Specifically, it uses Python's requests library to obtain the content of web pages and the BeautifulSoup library to extract the required information. The input is the URL or API endpoint of the source, and the output is the collected raw data.
[0167] Step 2:
[0168] The server extracts the necessary statistical data from the collected data. In this case, statistical information such as the title and body of news articles is extracted from the raw data. For example, the Python nltk library is used to parse the article text and extract important keywords and sentences. The input is the raw data, and the output is the extracted statistical data.
[0169] Step 3:
[0170] The server performs preprocessing on the extracted data. This preprocessing includes removing noise from the data, imputing missing values, standardizing date formats, and converting numerical units. Specifically, it uses the Python pandas library to create a data frame, imputing missing values, and aligning the data format. The input is the extracted statistical data, and the output is the preprocessed statistical data.
[0171] Step 4:
[0172] The server uses a generative AI model to evaluate the credibility of the preprocessed data. This evaluation includes criteria such as the reliability of the source, past revision history, and data consistency. Specifically, it uses scikit-learn to apply a logistic regression model to calculate a credibility score for each data point. The input is the preprocessed data, and the output is the credibility score.
[0173] Step 5:
[0174] The server calculates a credibility score based on the analysis results. This score is a numerical representation of the reliability of the data and is expressed in a range from 0 to 1. The calculated score is converted into a format that is easy for users to visually check. The input is the credibility evaluation result provided by the generative AI model, and the output is the credibility score.
[0175] Step 6:
[0176] The server provides the credibility score and data to the user in the form of a news feed. The user can view the news feed with the credibility score via their smartphone. The credibility score is color-coded, and by clicking on specific data, detailed analysis results are displayed. This is done using a web application framework such as Django. The input is the credibility score and data, and the output is a visually displayed news feed.
[0177] Prompt Sentence Examples
[0178] The user enters the following prompt into the system:
[0179] "What is the credibility score for the latest news on the number of new coronavirus cases?"
[0180] Users can view real-time news feeds with credibility scores.
[0181] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.
[0182] The system related to this invention, "DataVerity with Emotion Engine," collects data from various information sources, evaluates its credibility using a generative AI model, and combines it with an emotion engine that recognizes the user's emotions and adjusts the displayed content. The program processing details and specific examples of this system are shown below.
[0183] Data collection
[0184] The server collects data from various sources on the Internet, including public government databases, internal corporate databases, news sites, social media, etc. Using APIs and web scraping technology, the server automatically retrieves the necessary data and stores it in a cloud database.
[0185] Data extraction and preprocessing
[0186] The server extracts the statistical data required by the user from the data stored in the database. It uses database queries and natural language processing techniques to extract information based on specific keywords or categories. The extracted data undergoes preprocessing, such as noise removal, missing value completion, and format standardization. For example, standardizing date formats and converting numerical units converts the data into an analyzable format.
[0187] Analysis for credibility assessment
[0188] The server uses a generative AI model to evaluate the credibility of the pre-processed data. The model analyzes the data based on multiple criteria, such as the reliability of the source, past revision history, and data consistency, and calculates a credibility score for each piece of data.
[0189] Recognizing user emotions with an emotion engine
[0190] The device uses a camera and microphone to capture the user's face and voice in real time. The emotion engine uses facial recognition and voice analysis technologies to analyze the user's emotions. Specifically, it recognizes joy, sadness, surprise, anger, etc. from facial expressions and determines emotions from the tone and speed of the voice.
[0191] Providing and displaying credibility scores
[0192] When a user wants to check data from their device, they send a request. Based on the request, the server searches the database for the relevant data and its credibility score, and sends it back to the device. The device then visually displays the received data and credibility score to the user. The display format is easy to understand using graphs and tables, and the credibility score is displayed in a color-coded format.
[0193] Emotion-based content adjustment
[0194] The device adjusts the display content based on the user's emotions analyzed by the emotion engine. For example, if the user expresses surprise, the device displays additional related data or provides a detailed explanation. If the user expresses sadness or anger, the device simplifies the display content and provides an interface that reduces the user's stress.
[0195] Specific examples
[0196] For example, consider the case of collecting data on the number of new coronavirus infections and evaluating its credibility. The server collects data on the number of new coronavirus infections from official government databases, major news sites, and social media. Next, the server extracts the number of new infections, deaths, and recovered cases from the collected data, standardizing the date format and removing noise.
[0197] The server then uses a generative AI model to analyze the credibility of the data from each source. Data from public databases is assigned a high score (e.g., 90 points or higher) because it has high reliability, data from news sites is assigned a medium score (e.g., 70 to 80 points), and data from anonymous posts on social media is assigned a low score (e.g., 40 to 60 points).
[0198] Finally, when a user requests data on the number of coronavirus cases from their device, the server sends the data and credibility score based on the request. The device then visually displays this information in graphs and tables, and an emotion engine adjusts the display based on the user's emotions. For example, if the user expresses surprise, additional infection forecast data and policy response information will be displayed. If the user expresses sadness, improvements and positive information will be highlighted.
[0199] As described above, this system allows users to evaluate the credibility of data collected from a wide variety of sources while taking their emotions into account, and makes decisions quickly and accurately based on highly reliable data.
[0200] The processing flow will be explained below.
[0201] Program processing flow
[0202] Step 1:
[0203] The server collects data from designated sources (public government databases, internal corporate databases, news sites, social media, etc.), automatically retrieves the data using APIs and web scraping technology, and stores it in a cloud database.
[0204] Step 2:
[0205] The server extracts the statistical data the user needs from the data stored in the database, using database queries and natural language processing techniques to retrieve information based on specific keywords or categories.
[0206] Step 3:
[0207] The server preprocesses the extracted data, removing noise (e.g., removing irrelevant and duplicate data), imputing missing values, and standardizing data formats (e.g., converting date formats and numeric units).
[0208] Step 4:
[0209] The server inputs the preprocessed data into a generative AI model to evaluate its credibility. The generative AI model analyzes the data based on multiple criteria, including the reliability of the source, past revision history, and data consistency.
[0210] Step 5:
[0211] The server calculates a credibility score for each piece of data based on the analysis results of the generative AI model. The credibility score is expressed in the range of 0 to 100 and is assigned individually to each piece of data.
[0212] Step 6:
[0213] The server stores the calculated credibility score and the reason for the evaluation in a database, which allows for quick responses to subsequent user requests.
[0214] Step 7:
[0215] When a user wants to check statistical data from their device, they send a request for specific information, including search keywords and the range of data they require.
[0216] Step 8:
[0217] The server searches the database for relevant statistical data and credibility scores based on the user request, and returns them to the terminal together with detailed analysis results.
[0218] Step 9:
[0219] The device visually displays the received data and the credibility score to the user in the form of graphs and tables, and the credibility score is color-coded for easy understanding.
[0220] Step 10:
[0221] The device uses a camera and microphone to capture the user's face and voice in real time. The emotion engine uses facial recognition and voice analysis technologies to analyze the user's emotions. It recognizes joy, sadness, surprise, anger, etc. from facial expressions and determines emotions from the tone and speed of the voice.
[0222] Step 11:
[0223] The device adjusts the display content based on the user's emotions analyzed by the emotion engine. For example, if the user expresses surprise, the device displays additional related data or provides a detailed explanation. If the user expresses sadness or anger, the device simplifies the display content and provides an interface that reduces the user's stress.
[0224] Example 2
[0225] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0226] Conventional data collection and analysis systems lack the ability to evaluate the credibility of data from diverse sources. Furthermore, few systems adjust information based on user sentiment, and interfaces for improving the user experience are lacking. Therefore, there is a need for systems that provide highly credible data and display information that takes user sentiment into consideration.
[0227] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes a data collection means, a means for extracting necessary statistical information from the collected data, a means for pre-processing the extracted data, a means for analyzing the pre-processed data using a generative model to evaluate the credibility of the pre-processed data, a means for calculating a credibility score based on the analysis result, a means for providing the credibility score and data to the user, a means for performing face recognition and voice analysis to recognize the user's emotions, and a means for adjusting display content based on the user's emotions. This makes it possible to evaluate the credibility of data obtained from various information sources, and further, by adjusting the information display according to the user's emotions, an improved user experience is realized.
[0228] "Data collection means" refers to devices or programs that automatically acquire necessary data from various information sources and store it in a database.
[0229] "Statistical information" is numerical or textual data collected for a specific purpose and analyzed accordingly.
[0230] "Preprocessing" is the process of converting collected data into a format suitable for analysis by performing processes such as noise removal, missing value completion, and format standardization on the data.
[0231] A "generative model" is an artificial intelligence or machine learning algorithm used to input a variety of data and perform tasks such as evaluating credibility and extracting patterns.
[0232] The "credibility score" is an evaluation index that numerically represents the reliability of preprocessed data.
[0233] "Facial recognition" is a technology that uses a camera to capture a user's face and analyze their facial expressions to recognize their emotions.
[0234] "Voice analysis" is a technology that analyzes voice data collected through a microphone and recognizes emotions and intentions from the tone and speed of the voice.
[0235] "Adjusting display content" is the process of optimizing the format and content of displayed information based on the user's emotions to improve the user experience.
[0236] "Diverse sources" refers to information providers from different sources, such as public government databases, news sites, social media, and internal corporate databases.
[0237] "User interface" refers to the screens and tools that allow users to visually check and manipulate data and credibility scores.
[0238] MODE FOR CARRYING OUT THE INVENTION
[0239] The system according to the present invention, "data credibility evaluation system," is composed of the following specific hardware and software, and collects, analyzes, and displays data.
[0240] Data collection
[0241] The server collects data from various sources on the Internet. To do this, it uses APIs (Application Programming Interfaces) and web scraping technology. For example, it obtains government-issued data from the API of public government databases, or uses tools such as BeautifulSoup and Selenium to extract necessary data from websites. The collected data is stored in cloud databases such as Amazon RDS and Google® BigQuery.
[0242] Data extraction and preprocessing
[0243] The server extracts the necessary statistical information from the stored data. It uses SQL queries to filter the data based on specific conditions, and natural language processing (NLP) techniques to extract information based on specific keywords or categories. For example, libraries such as NLTK or SpaCy are used. The extracted data is preprocessed to remove noise, impute missing values, and standardize formats (for example, standardizing date formats to YYYY-MM-DD).
[0244] Analysis for credibility assessment
[0245] The server uses a generative AI model (e.g., GPT-3 (registered trademark) or BERT) to evaluate the veracity of the preprocessed data. The generative AI model is given a prompt sentence such as:
[0246] "Please rate the reliability of this data. We will calculate a reliability score for each of the data collected from official government databases, major news sites, and social media, and tell us which data you think is the most trustworthy."
[0247] The generative AI model then calculates a credibility score for each source, taking into account factors such as the source's reliability, revision history, and data consistency.
[0248] Recognizing user emotions with an emotion engine
[0249] The device uses a camera and microphone to recognize the user's emotions in real time. The camera analyzes the user's facial expressions using facial recognition technology such as OpenCV or Face++. The microphone uses voice analysis technology such as Google Speech API or IBM Watson (registered trademark) to determine the user's emotions from the tone and speed of the voice. For example, it can identify whether the user is laughing, angry, or surprised.
[0250] Providing a credibility score
[0251] When a user sends a data verification request through their device, the server searches the database for the relevant data and its credibility score and returns it to the device. This data is sent in JSON format, so the device can analyze it and display it visually in graphs and tables. For example, data with high reliability is displayed in green, and data with low reliability is displayed in red.
[0252] Emotion-based content adjustment
[0253] The device adjusts the display content based on the user's emotional data analyzed by the emotion engine. For example, if the user shows a surprised expression, the device will display additional relevant information (such as infection forecasts and policy response information). If the user is sad, the display content will be simplified and positive information will be emphasized. Adjustments will be made, such as emphasizing the increase in the number of recovered people rather than the increase in the number of infected people.
[0254] In this way, the system is designed to allow users to easily evaluate the reliability of data collected from multiple sources and to provide the most appropriate information according to the user's feelings.
[0255] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0256] Step 1:
[0257] The server collects data from various sources. Specifically, it uses APIs and web scraping technology to obtain data from public databases, news sites, social media, etc. The collected data is stored in a cloud database. For example, BeautifulSoup can be used to collect article content from news sites. The input is a list of source URLs, and the output is the collected raw data.
[0258] Step 2:
[0259] The server extracts the necessary statistical information from the raw data stored in the cloud database. Specifically, it uses SQL queries and natural language processing techniques to filter the data and extract information based on specific keywords or categories. The input is the raw data, and the output is the extracted statistical information. A specific example of how it works is to use an SQL query to extract data on the "number of new coronavirus infections."
[0260] Step 3:
[0261] The server preprocesses the extracted statistical information. Specifically, it removes noise, fills in missing values, and standardizes formats. For example, it standardizes date formats and aligns numerical units. The input is the extracted statistical information, and the output is the preprocessed data. A specific operation is to standardize all date formats to "YYYY-MM-DD".
[0262] Step 4:
[0263] The server analyzes the preprocessed data using a generative AI model to evaluate its credibility. The input is the preprocessed data, and the output is a credibility score. Specifically, the server prompts the generative AI model as follows: "Please rate the credibility of this data. Based on data collected from public government databases, major news sites, and social media, calculate a credibility score for each, and tell us which data is the most trustworthy."
[0264] Step 5:
[0265] The device uses a camera and microphone to capture the user's face and voice in real time, and analyzes them using an emotion engine. The input is the user's real-time video and audio data, and the output is the user's emotion data. Specific operations include facial recognition using OpenCV and voice analysis using the Google Speech API.
[0266] Step 6:
[0267] When a user sends a data verification request from a device, the server searches the database for the relevant data and its credibility score and returns it to the device. The input is the user's request data, and the output is the relevant data and its credibility score. Specific operations include sending data to the device in JSON format.
[0268] Step 7:
[0269] The terminal visually displays the received data and credibility score in graph and table format, and the credibility score is displayed in color. The input is the data received from the server and the credibility score, and the output is the visually displayed information. Specific operations include drawing graphs using Matplotlib.
[0270] Step 8:
[0271] The device adjusts the display content based on the user's emotional data analyzed by the emotion engine. The input is the user's emotional data, and the output is the adjusted display content. Specific operations include displaying additional information when the user expresses surprise, and emphasizing positive information when the user expresses sadness.
[0272] (Application example 2)
[0273] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0274] In modern society, there is a need to collect data from many sources and evaluate its credibility. At the same time, it is also important to provide appropriate information based on individual user emotions. However, conventional systems have had difficulty simultaneously evaluating the credibility of data and displaying information based on user emotions. Furthermore, in the advertising field, there is a need to display appropriate advertisements based on user emotions, but existing technologies have not been able to adequately meet this need. A system that can meet these complex requirements is needed.
[0275] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes a data collection means, a means for extracting necessary statistical data from the collected data, a means for pre-processing the extracted data, a means for analyzing the pre-processed data using a generative model for evaluating the credibility of the pre-processed data, a means for calculating a credibility score based on the analysis result, a means for providing the credibility score and data to the user, and a sentiment analysis means for recognizing the user's sentiment and adjusting the display content. This allows the user not only to evaluate the credibility of data collected from a wide variety of information sources, but also to display appropriate information and advertisements in real time according to their sentiments.
[0276] "Data collection means" refers to a method of automatically obtaining necessary data from various sources such as public databases, news sites, and social media via the Internet and storing it in a cloud database.
[0277] The "means for extracting statistical data" refers to a means for extracting statistical data required by a user from the collected data based on specific keywords or categories.
[0278] "Preprocessing means" refers to a means of performing preprocessing such as noise removal, missing value completion, and format standardization on the extracted data, and converting it into an analyzable format.
[0279] "Means for performing analysis using a generative model" refers to means for performing data analysis using a generative AI model to evaluate the credibility of preprocessed data.
[0280] The "means for calculating a credibility score" is a means for converting the credibility of each piece of data into a numerical value based on the analysis results and calculating a score.
[0281] The "means for providing the credibility score and data to the user" refers to a means for visually displaying the calculated credibility score and related data to the user.
[0282] The "emotion analysis means" is a means for analyzing the user's face and voice in real time using a camera and microphone, recognizing the user's emotions, and adjusting the display content accordingly.
[0283] A "generative AI model" is a machine learning model trained on a large dataset and used to assess the veracity of the data.
[0284] The system related to this invention, "AdOptimizer with Emotion Engine," consists of the following steps: data collection, data extraction, preprocessing, analysis using a generative model, credibility evaluation and score calculation, emotion analysis, and display of advertisements. This system can adjust the display content based on the user's emotions and provide optimal advertisements.
[0285] The server uses a data collection means to automatically obtain necessary data via the Internet from public databases, news sites, social media, etc., and stores it in a cloud database. Next, a means for extracting necessary statistical data from the collected data is used to extract data required by the user based on specific keywords or categories. The extracted data is then processed using a preprocessing means to remove noise, fill in missing values, standardize formats, etc.
[0286] The pre-processed data is analyzed using a generative AI model to assess the veracity of the data. Based on the analysis results, a veracity score is calculated. These scores and data are visually displayed to a user by a means for providing the veracity score and data to a user.
[0287] The device captures the user's face and voice in real time using emotion analysis means. The emotion analysis means uses a camera and microphone to analyze the user's emotions using facial recognition and voice analysis technologies. Specifically, it recognizes emotions such as joy, sadness, surprise, and anger from facial expressions and determines emotions from the tone and speed of the voice.
[0288] When a user requests advertising data from a device, the server sends the data and credibility score to the device based on the request. The device visually displays this information in the form of graphs, tables, etc., and the emotion analysis means adjusts the display content based on the user's emotions. For example, if the user expresses surprise, detailed advertising information is displayed, and if the user expresses sadness, a calm advertisement is presented.
[0289] As a concrete example, consider a robot in a store displaying advertisements. The robot's camera captures the user's facial expressions, and the microphone analyzes their voice. If the user smiles (happiness), a regular advertisement (e.g., a promotion for a new product) is displayed. On the other hand, if the user is sad, a more gentle advertisement (e.g., an introduction to a relaxing drink) is displayed.
[0290] An example of a prompt sentence is "If the user is excited, provide detailed advertising information." This prompt sentence is used as an instruction for the generative AI model to provide appropriate advertising based on the user's emotions.
[0291] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0292] Step 1:
[0293] The server uses data collection tools to automatically retrieve the necessary data from public databases, news sites, social media, etc. via the Internet and stores it in a cloud database. Specifically, data is collected using APIs and web scraping technology. The input requires the URL of the source and an API key, and the output is the retrieved raw data.
[0294] Step 2:
[0295] The server processes the collected raw data using a data extraction method to extract the statistical data required by the user. Specifically, it uses natural language processing techniques and database queries to extract information based on specific keywords or categories. The input is the raw data, and the output is the extracted statistical data.
[0296] Step 3:
[0297] The server preprocesses the extracted statistical data using preprocessing means. Specifically, it performs noise removal, missing value completion, and format standardization. For example, it standardizes date formats and converts numerical units. The input is the extracted statistical data, and the output is the preprocessed data.
[0298] Step 4:
[0299] The server uses a generative AI model to evaluate the credibility of the preprocessed data and perform analysis. Specifically, it analyzes the data based on the reliability of the source, past revision history, data consistency, etc. The input is the preprocessed data, and the output is the analysis result and a credibility score.
[0300] Step 5:
[0301] The server uses a means to calculate a credibility score based on the analysis results. Specifically, it assigns a credibility score calculated by the generative AI model to each piece of data. The input is the analysis results, and the output is data with a credibility score.
[0302] Step 6:
[0303] When a user requests advertising data from a terminal, the server uses a means to provide the user with a credibility score and the data. Specifically, it searches the database for the relevant data and sends it to the user's terminal. The input is the user's request, and the output is the data with a credibility score.
[0304] Step 7:
[0305] The device uses emotion analysis means to recognize the user's emotions. Specifically, it analyzes facial expressions captured by a camera and voice captured by a microphone, and identifies emotions using facial recognition technology and voice analysis technology. The input is the captured face and voice data, and the output is the analyzed user's emotion data.
[0306] Step 8:
[0307] The terminal adjusts the advertisement content to be displayed based on the user's emotional data. Specifically, the advertisement content is changed according to the user's emotional state. For example, when the user expresses surprise, detailed advertisement information is displayed, and when the user expresses sadness, calm advertisement content is displayed. The input is the user's emotional data, and the output is the adjusted advertisement content.
[0308] Step 9:
[0309] The terminal visually displays the adjusted advertising content to the user. Specifically, the advertising content is displayed on the screen using graphs and tables to allow the user to intuitively understand it. The input is the adjusted advertising content, and the output is the displayed advertisement.
[0310] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0311] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (registered trademark) (Internet search engine).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0312] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.
[0313] [Second embodiment]
[0314] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.
[0315] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0316] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0317] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.
[0318] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0319] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0320] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0321] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0322] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0323] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0324] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0325] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."
[0326] The system related to this invention, "DataVerity," collects data from various sources, evaluates its credibility using a generative AI model, and provides users with the data and a credibility score. The program processing content and specific examples of this system are shown below.
[0327] Data collection
[0328] The server collects data from various sources on the Internet, including public government databases, internal corporate databases, news sites, social media, etc. Using APIs and web scraping technology, the server automatically retrieves the necessary data and stores it in a cloud database.
[0329] Data extraction and preprocessing
[0330] The server extracts the necessary statistical data from the stored database. This data extraction is performed using database queries and natural language processing techniques. The extracted data undergoes preprocessing, such as noise removal, missing value completion, and format standardization. For example, the data is converted into an analyzable format by standardizing date formats and converting numerical units.
[0331] Analysis for credibility assessment
[0332] The server uses a generative AI model to evaluate the credibility of the pre-processed data. The model analyzes the credibility of the data based on multiple evaluation criteria, such as the reliability of the source, past revision history, and data consistency, and then calculates a credibility score for each piece of data.
[0333] Providing and displaying credibility scores
[0334] When a user wants to check data from their device, they send a request. Based on the request, the server searches the database for the relevant data and its credibility score, and sends it back to the device. The device then visually displays the received data and credibility score to the user. The display format is easy to understand using graphs and tables, and the credibility score is color-coded. This allows the user to visually check the credibility of the data.
[0335] Specific examples
[0336] For example, consider the case of collecting data on the number of new coronavirus infections and evaluating its credibility. The server collects data on the number of new coronavirus infections from official government databases, major news sites, and social media. Next, the server extracts the number of new infections, deaths, and recovered cases from the collected data, standardizing the date format and removing noise.
[0337] The server then uses a generative AI model to analyze the credibility of the data from each source. Data from public databases is assigned a high score (e.g., 90 points or higher) because it has high reliability, data from news sites is assigned a medium score (e.g., 70 to 80 points), and data from anonymous posts on social media is assigned a low score (e.g., 40 to 60 points).
[0338] Finally, when a user requests data on the number of coronavirus cases from their device, the server sends the data and credibility score to the device based on the request. The device then displays this information visually in graphs and tables, and users can click on specific data to view detailed analysis results.
[0339] As described above, this system enables users to easily evaluate the credibility of data collected from a wide variety of information sources and make decisions quickly and accurately based on highly reliable data.
[0340] The processing flow will be explained below.
[0341] Program processing flow
[0342] Step 1:
[0343] The server collects data from designated sources (public government databases, internal corporate databases, news sites, social media, etc.), automatically retrieves the data using APIs and web scraping technology, and stores it in a cloud database.
[0344] Step 2:
[0345] The server extracts the statistical data the user needs from the data stored in the database, using database queries and natural language processing techniques to retrieve information based on specific keywords or categories.
[0346] Step 3:
[0347] The server preprocesses the extracted data, removing noise (e.g., removing irrelevant and duplicate data), imputing missing values, and standardizing data formats (e.g., converting date formats and numeric units).
[0348] Step 4:
[0349] The server inputs the preprocessed data into a generative AI model to evaluate its credibility. The generative AI model analyzes the data based on multiple criteria, including the reliability of the source, past revision history, and data consistency.
[0350] Step 5:
[0351] The server calculates a credibility score for each piece of data based on the analysis results of the generative AI model. The credibility score is expressed in the range of 0 to 100 and is assigned individually to each piece of data.
[0352] Step 6:
[0353] The server stores the calculated credibility score and the reason for the evaluation in a database, which allows for quick responses to subsequent user requests.
[0354] Step 7:
[0355] When a user wants to check statistical data from their device, they send a request for specific information, including search keywords and the range of data they require.
[0356] Step 8:
[0357] The server searches the database for relevant statistical data and credibility scores based on the user request, and returns them to the terminal together with detailed analysis results.
[0358] Step 9:
[0359] The device visually displays the received data and the credibility score to the user in the form of graphs and tables, and the credibility score is color-coded for easy understanding.
[0360] Step 10:
[0361] If a user wants to check the details of a particular piece of data, they can click on it to see detailed analysis results and the reasons for the evaluation, allowing users to understand the credibility of the data and make decisions based on reliable information.
[0362] Through these processing steps, the DataVerity system helps users effectively evaluate the credibility of data from a wide variety of sources and make decisions quickly and accurately.
[0363] Example 1
[0364] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0365] There is a need for an appropriate system that can quickly and accurately evaluate the credibility of data collected from various sources and provide the results visually to users. However, current systems require a lot of manual work in the collection, preprocessing, credibility evaluation, and provision of data, which is time-consuming and labor-intensive, and has low reliability. Therefore, an efficient and reliable data collection and evaluation system is needed.
[0366] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0367] In this invention, the server includes a data collection means, a means for extracting necessary information from the collected data, a means for preprocessing the extracted data, a means for analyzing the preprocessed data using a generative model for evaluating the credibility of the data, a means for calculating a credibility score based on the analysis result, and a means for providing the credibility score and data to a user. This enables a user to quickly and accurately collect data from various information sources, evaluate the credibility of the data, and then visually confirm it.
[0368] "Data collection means" refers to a function for automatically acquiring data from various sources on the Internet.
[0369] "Means for extracting necessary information from collected data" refers to a function for selecting specific information from collected data and formatting it into a format suitable for data analysis.
[0370] "Means for preprocessing extracted data" refers to functions for removing noise, filling in missing values, standardizing formats, etc. from the extracted data, and preparing it in a format suitable for analysis.
[0371] "Means for performing analysis using a generative model to assess the credibility of preprocessed data" refers to a function that performs data analysis using a generative AI model to assess the reliability of preprocessed data.
[0372] The "means for calculating the credibility score" is a function for quantifying and indicating the reliability of data evaluated using a generative model.
[0373] "Means for providing credibility scores and data to users" refers to a function for providing analyzed credibility scores and related data to users in a visually verifiable format.
[0374] The system of the present invention collects data from various sources, evaluates its credibility using a generative AI model, and provides users with the data and a credibility score. Specific embodiments and methods of use are described below.
[0375] The system includes the following main elements:
[0376] Data collection methods
[0377] A means of extracting the necessary information from the collected data
[0378] A means of preprocessing the extracted data
[0379] A means of analyzing using generative models to assess the veracity of preprocessed data
[0380] A method for calculating a credibility score based on the analysis results
[0381] Means of providing credibility scores and data to users
[0382] Data collection methods
[0383] The server automatically collects data from various sources on the Internet using API access and web scraping techniques. For example, it uses the Python libraries BeautifulSoup and Scrapy to retrieve data from public databases, news sites, and social media.
[0384] A means of extracting the necessary information from the collected data
[0385] The server selects specific information from the collected data and formats it into a format suitable for data analysis. Specifically, it uses SQL queries to extract data such as the number of new infections, deaths, and recovered cases from the database.
[0386] A means of preprocessing the extracted data
[0387] The server performs preprocessing on the extracted data, specifically removing noise, filling in missing values, standardizing formats, etc. By using the Python Pandas library, the server consistently standardizes date formats and converts numeric units.
[0388] A means of analyzing using generative models to assess the veracity of preprocessed data
[0389] The server then inputs the preprocessed data into a generative AI model to assess its credibility. This generative model takes into account the reliability of the source, past revision history, and data consistency, and uses natural language processing (NLP) techniques, such as Transformer-based models (such as BERT and GPT).
[0390] A method for calculating a credibility score based on the analysis results
[0391] The server calculates a credibility score based on the analysis results of the generative model. This score is expressed as a number and serves as a standard for evaluating the reliability of the data. For example, public databases have a high score (90 points or above), news sites have a medium score (70-80 points), and social media data has a low score (40-60 points).
[0392] Means of providing credibility scores and data to users
[0393] When a user uses a device to send a data verification request, the server searches for and returns the relevant data and credibility score. The device then visually displays this data. Specifically, the received data is displayed in graphs and tables, and the credibility score is color-coded. Users can click on the data to view detailed analysis results.
[0394] Examples of concrete examples and prompts
[0395] For example, consider the case of evaluating the credibility of data on the number of people infected with the new coronavirus. In this case, the server collects relevant data from official government databases, major news sites, and social media, and inputs it into a generative AI model for analysis. After calculating a credibility score, it displays it on the device. The following is an example of a specific prompt sentence:
[0396] "Collect data on the number of COVID-19 cases and rate its credibility. Collect data from each source, use a generative AI model to score its credibility, and display it visually."
[0397] As described above, the present invention enables users to easily evaluate the credibility of data collected from a variety of information sources and make decisions quickly and accurately.
[0398] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0399] Step 1:
[0400] The server collects data from various sources on the Internet, specifically using API access and web scraping technology. It uses a list of URLs from public government databases, news sites, social media, etc. as input, and automatically retrieves the latest data from them. It stores the retrieved data in a cloud database. The output is a collection of the collected raw data.
[0401] Specifically, data extraction is performed automatically using Python libraries such as BeautifulSoup and Scrapy.
[0402] Step 2:
[0403] The server extracts the necessary information from the collected raw data. It uses the raw data as input and uses database queries and natural language processing (NLP) techniques to select specific data such as the number of new cases, deaths, and recoveries. The output is the extracted data.
[0404] Specifically, it uses SQL queries to extract specific columns and rows from a database and then uses NLP techniques to analyze the text data.
[0405] Step 3:
[0406] The server performs preprocessing on the extracted data. It uses the extracted data as input and performs noise removal, missing value imputation, format standardization, etc. The output is preprocessed data that can be analyzed.
[0407] Specifically, it uses Python's Pandas library to clean the data and standardize date formats and numerical units.
[0408] Step 4:
[0409] The server inputs the preprocessed data into a generative AI model to analyze it and evaluate its credibility. As input, the preprocessed data is provided to a generative model (e.g., BERT or GPT), and the model evaluates the credibility of the data. The output is a credibility score for each data point.
[0410] Specifically, it uses a generative model to extract features from the data and quantify their reliability.
[0411] Step 5:
[0412] The server calculates a credibility score based on the analysis results of the generative model. It uses the evaluation results of the generative model as input and applies an algorithm to calculate the credibility score. The output is a dataset containing the calculated credibility scores.
[0413] Specifically, the evaluation algorithm is executed to quantify the reliability and store it in a database.
[0414] Step 6:
[0415] When a user uses a terminal to send a request for data verification, the server searches for the relevant data and credibility score and returns it.The server receives the user's request as input and searches the database for the corresponding data and credibility score.The output is the search result: the data and credibility score.
[0416] Specifically, it receives an HTTP request from a user and executes a corresponding SQL query on a database.
[0417] Step 7:
[0418] The terminal visually displays the received data and credibility scores. It uses the data and credibility scores received from the server as input and displays them in the form of graphs and tables. The output is a visual representation of the data provided to the user.
[0419] Specifically, the data is displayed on a web page using JavaScript and HTML, and the reliability is color-coded to make it easier to understand visually. Users can click to see detailed analysis results.
[0420] (Application example 1)
[0421] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0422] In modern society, there is a huge amount of information available on the Internet, making it extremely difficult to obtain accurate and reliable information from all that information. News and important data, in particular, often contain unreliable information, requiring users to expend considerable effort to determine the accuracy of the information. Furthermore, many news articles are constantly being updated, creating a lack of means to assess their credibility in real time. Therefore, there is a need for a system that can quickly and accurately assess the credibility of information and provide it to users.
[0423] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0424] In this invention, the server includes a data collection means, a means for extracting necessary statistical data from the collected data, a means for pre-processing the extracted data, a means for analyzing the pre-processed data using a generative model for evaluating the credibility of the data, a means for calculating a credibility score based on the analysis result, and a means for providing the credibility score and data to a user in a news feed format, thereby enabling the user to evaluate the credibility of news articles in real time and make quick and accurate decisions based on highly reliable information.
[0425] "Data collection means" refers to technology for automatically obtaining necessary data from various information sources on the Internet.
[0426] "Means for extracting statistical data" refers to techniques for extracting necessary statistical information from collected data.
[0427] "Preprocessing means" refers to technology that converts extracted data into an analyzable format by removing noise, filling in missing values, and standardizing the format.
[0428] "Means for performing analysis using generative models" refers to techniques for evaluating the veracity of data that has been preprocessed using a generative AI model.
[0429] The "means for calculating the credibility score" is a technology that quantifies the reliability of each piece of data based on the analysis results of the generative AI model.
[0430] "Newsfeed-style provision" refers to a technology that visually provides users with credibility scores and data in real time.
[0431] "Source credibility" is a criterion for evaluating the past performance and reliability of the source of data.
[0432] "Edit history" is historical information about how data has been edited in the past.
[0433] "Data consistency" is a criterion for assessing the consistency and agreement of data obtained from multiple sources.
[0434] "User interface" refers to the screen or operating means through which a user can interact with the system and visually check the credibility score of data.
[0435] A system for realizing this invention includes a data collection means, a means for extracting statistical data, a data preprocessing means, a means for performing credibility assessment using a generative AI model, a means for calculating a credibility score, and a means for providing the data and the credibility score in a news feed format.
[0436] Program processing and technologies used
[0437] Data collection methods
[0438] The server collects data from various sources on the Internet, including public government databases, corporate databases, news sites, and social media, using APIs and web scraping techniques. Specifically, web scraping is performed using Python's requests and BeautifulSoup libraries.
[0439] A means of extracting statistical data
[0440] The server extracts the necessary statistics from the collected data, using database queries and natural language processing (NLP) techniques to extract key information such as article titles and content, for example using the Python nltk library.
[0441] Data preprocessing measures
[0442] The server removes noise from the extracted data, imputes missing values, and standardizes the format. Specifically, it standardizes date formats and converts numerical units. This converts the data into an analyzable format. For example, it uses the pandas library.
[0443] Credibility assessment tools
[0444] The server evaluates the credibility of the preprocessed data using a generative AI model, which analyzes the credibility of the data based on criteria such as the reliability of the source, past revision history, and data consistency. For example, it applies a logistic regression model using scikit-learn.
[0445] How the credibility score is calculated
[0446] The server calculates a credibility score based on the analysis results of the generative AI model. This score is a numerical representation of the reliability of the data, and is expressed, for example, in the range of 0 to 1.
[0447] A means to provide data and credibility scores in a news feed format
[0448] The system allows users to check data and credibility scores in a news feed format on their devices. The credibility scores are displayed in color, and by clicking on specific data, detailed analysis results are displayed. For example, using Django as the web application framework, the data and scores are displayed visually in graphs and tables.
[0449] Specific examples
[0450] For example, if a user wants to check the latest news on the number of new coronavirus infections, the server collects relevant data from public databases, news sites, and social media. The generative AI model assigns a credibility score to each source, assigning a high score (90 points or above) to data from public databases, a medium score (70 to 80 points) to data from news sites, and a low score (40 to 60 points) to data from social media. Users can launch the app on their smartphone and check the credibility score of each article along with the data displayed in a news feed format.
[0451] Prompt Sentence Examples
[0452] "What is the credibility score for the latest news on the number of new coronavirus cases?"
[0453] As described above, this invention collects data from various sources on the Internet, evaluates its credibility, and provides it to users, enabling them to make decisions based on highly reliable information in real time.
[0454] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0455] Step 1:
[0456] The server collects data from various sources on the Internet. This includes scraping and using APIs to obtain data from public government databases, news sites, social media, etc. Specifically, it uses Python's requests library to obtain the content of web pages and the BeautifulSoup library to extract the required information. The input is the URL or API endpoint of the source, and the output is the collected raw data.
[0457] Step 2:
[0458] The server extracts the necessary statistical data from the collected data. In this case, statistical information such as the title and body of news articles is extracted from the raw data. For example, the Python nltk library is used to parse the article text and extract important keywords and sentences. The input is the raw data, and the output is the extracted statistical data.
[0459] Step 3:
[0460] The server performs preprocessing on the extracted data. This preprocessing includes removing noise from the data, imputing missing values, standardizing date formats, and converting numerical units. Specifically, it uses the Python pandas library to create a data frame, imputing missing values, and aligning the data format. The input is the extracted statistical data, and the output is the preprocessed statistical data.
[0461] Step 4:
[0462] The server uses a generative AI model to evaluate the credibility of the preprocessed data. This evaluation includes criteria such as the reliability of the source, past revision history, and data consistency. Specifically, it uses scikit-learn to apply a logistic regression model to calculate a credibility score for each data point. The input is the preprocessed data, and the output is the credibility score.
[0463] Step 5:
[0464] The server calculates a credibility score based on the analysis results. This score is a numerical representation of the reliability of the data and is expressed in a range from 0 to 1. The calculated score is converted into a format that is easy for users to visually check. The input is the credibility evaluation result provided by the generative AI model, and the output is the credibility score.
[0465] Step 6:
[0466] The server provides the credibility score and data to the user in the form of a news feed. The user can view the news feed with the credibility score via their smartphone. The credibility score is color-coded, and by clicking on specific data, detailed analysis results are displayed. This is done using a web application framework such as Django. The input is the credibility score and data, and the output is a visually displayed news feed.
[0467] Prompt Sentence Examples
[0468] The user enters the following prompt into the system:
[0469] "What is the credibility score for the latest news on the number of new coronavirus cases?"
[0470] Users can view real-time news feeds with credibility scores.
[0471] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0472] The system related to this invention, "DataVerity with Emotion Engine," collects data from various information sources, evaluates its credibility using a generative AI model, and combines it with an emotion engine that recognizes the user's emotions and adjusts the displayed content. The program processing details and specific examples of this system are shown below.
[0473] Data collection
[0474] The server collects data from various sources on the Internet, including public government databases, internal corporate databases, news sites, social media, etc. Using APIs and web scraping technology, the server automatically retrieves the necessary data and stores it in a cloud database.
[0475] Data extraction and preprocessing
[0476] The server extracts the statistical data required by the user from the data stored in the database. It uses database queries and natural language processing techniques to extract information based on specific keywords or categories. The extracted data undergoes preprocessing, such as noise removal, missing value completion, and format standardization. For example, standardizing date formats and converting numerical units converts the data into an analyzable format.
[0477] Analysis for credibility assessment
[0478] The server uses a generative AI model to evaluate the credibility of the pre-processed data. The model analyzes the data based on multiple criteria, such as the reliability of the source, past revision history, and data consistency, and calculates a credibility score for each piece of data.
[0479] Recognizing user emotions with an emotion engine
[0480] The device uses a camera and microphone to capture the user's face and voice in real time. The emotion engine uses facial recognition and voice analysis technologies to analyze the user's emotions. Specifically, it recognizes joy, sadness, surprise, anger, etc. from facial expressions and determines emotions from the tone and speed of the voice.
[0481] Providing and displaying credibility scores
[0482] When a user wants to check data from their device, they send a request. Based on the request, the server searches the database for the relevant data and its credibility score, and sends it back to the device. The device then visually displays the received data and credibility score to the user. The display format is easy to understand using graphs and tables, and the credibility score is displayed in a color-coded format.
[0483] Emotion-based content adjustment
[0484] The device adjusts the display content based on the user's emotions analyzed by the emotion engine. For example, if the user expresses surprise, the device displays additional related data or provides a detailed explanation. If the user expresses sadness or anger, the device simplifies the display content and provides an interface that reduces the user's stress.
[0485] Specific examples
[0486] For example, consider the case of collecting data on the number of new coronavirus infections and evaluating its credibility. The server collects data on the number of new coronavirus infections from official government databases, major news sites, and social media. Next, the server extracts the number of new infections, deaths, and recovered cases from the collected data, standardizing the date format and removing noise.
[0487] The server then uses a generative AI model to analyze the credibility of the data from each source. Data from public databases is assigned a high score (e.g., 90 points or higher) because it has high reliability, data from news sites is assigned a medium score (e.g., 70 to 80 points), and data from anonymous posts on social media is assigned a low score (e.g., 40 to 60 points).
[0488] Finally, when a user requests data on the number of coronavirus cases from their device, the server sends the data and credibility score based on the request. The device then visually displays this information in graphs and tables, and an emotion engine adjusts the display based on the user's emotions. For example, if the user expresses surprise, additional infection forecast data and policy response information will be displayed. If the user expresses sadness, improvements and positive information will be highlighted.
[0489] As described above, this system allows users to evaluate the credibility of data collected from a wide variety of sources while taking their emotions into account, and makes decisions quickly and accurately based on highly reliable data.
[0490] The processing flow will be explained below.
[0491] Program processing flow
[0492] Step 1:
[0493] The server collects data from designated sources (public government databases, internal corporate databases, news sites, social media, etc.), automatically retrieves the data using APIs and web scraping technology, and stores it in a cloud database.
[0494] Step 2:
[0495] The server extracts the statistical data the user needs from the data stored in the database, using database queries and natural language processing techniques to retrieve information based on specific keywords or categories.
[0496] Step 3:
[0497] The server preprocesses the extracted data, removing noise (e.g., removing irrelevant and duplicate data), imputing missing values, and standardizing data formats (e.g., converting date formats and numeric units).
[0498] Step 4:
[0499] The server inputs the preprocessed data into a generative AI model to evaluate its credibility. The generative AI model analyzes the data based on multiple criteria, including the reliability of the source, past revision history, and data consistency.
[0500] Step 5:
[0501] The server calculates a credibility score for each piece of data based on the analysis results of the generative AI model. The credibility score is expressed in the range of 0 to 100 and is assigned individually to each piece of data.
[0502] Step 6:
[0503] The server stores the calculated credibility score and the reason for the evaluation in a database, which allows for quick responses to subsequent user requests.
[0504] Step 7:
[0505] When a user wants to check statistical data from their device, they send a request for specific information, including search keywords and the range of data they require.
[0506] Step 8:
[0507] The server searches the database for relevant statistical data and credibility scores based on the user request, and returns them to the terminal together with detailed analysis results.
[0508] Step 9:
[0509] The device visually displays the received data and the credibility score to the user in the form of graphs and tables, and the credibility score is color-coded for easy understanding.
[0510] Step 10:
[0511] The device uses a camera and microphone to capture the user's face and voice in real time. The emotion engine uses facial recognition and voice analysis technologies to analyze the user's emotions. It recognizes joy, sadness, surprise, anger, etc. from facial expressions and determines emotions from the tone and speed of the voice.
[0512] Step 11:
[0513] The device adjusts the display content based on the user's emotions analyzed by the emotion engine. For example, if the user expresses surprise, the device displays additional related data or provides a detailed explanation. If the user expresses sadness or anger, the device simplifies the display content and provides an interface that reduces the user's stress.
[0514] Example 2
[0515] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0516] Conventional data collection and analysis systems lack the ability to evaluate the credibility of data from diverse sources. Furthermore, few systems adjust information based on user sentiment, and interfaces for improving the user experience are lacking. Therefore, there is a need for systems that provide highly credible data and display information that takes user sentiment into consideration.
[0517] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes a data collection means, a means for extracting necessary statistical information from the collected data, a means for pre-processing the extracted data, a means for analyzing the pre-processed data using a generative model to evaluate the credibility of the pre-processed data, a means for calculating a credibility score based on the analysis result, a means for providing the credibility score and data to the user, a means for performing face recognition and voice analysis to recognize the user's emotions, and a means for adjusting display content based on the user's emotions. This makes it possible to evaluate the credibility of data obtained from various information sources, and further, by adjusting the information display according to the user's emotions, an improved user experience is realized.
[0518] "Data collection means" refers to devices or programs that automatically acquire necessary data from various information sources and store it in a database.
[0519] "Statistical information" is numerical or textual data collected for a specific purpose and analyzed accordingly.
[0520] "Preprocessing" is the process of converting collected data into a format suitable for analysis by performing processes such as noise removal, missing value completion, and format standardization on the data.
[0521] A "generative model" is an artificial intelligence or machine learning algorithm used to input a variety of data and perform tasks such as evaluating credibility and extracting patterns.
[0522] The "credibility score" is an evaluation index that numerically represents the reliability of preprocessed data.
[0523] "Facial recognition" is a technology that uses a camera to capture a user's face and analyze their facial expressions to recognize their emotions.
[0524] "Voice analysis" is a technology that analyzes voice data collected through a microphone and recognizes emotions and intentions from the tone and speed of the voice.
[0525] "Adjusting display content" is the process of optimizing the format and content of displayed information based on the user's emotions to improve the user experience.
[0526] "Diverse sources" refers to information providers from different sources, such as public government databases, news sites, social media, and internal corporate databases.
[0527] "User interface" refers to the screens and tools that allow users to visually check and manipulate data and credibility scores.
[0528] MODE FOR CARRYING OUT THE INVENTION
[0529] The system according to the present invention, "data credibility evaluation system," is composed of the following specific hardware and software, and collects, analyzes, and displays data.
[0530] Data collection
[0531] The server collects data from various sources on the Internet. To do this, it uses APIs (Application Programming Interfaces) and web scraping technology. For example, it retrieves government-issued data from the API of public government databases, or uses tools such as BeautifulSoup and Selenium to extract necessary data from websites. The collected data is then stored in cloud databases such as Amazon RDS and Google BigQuery.
[0532] Data extraction and preprocessing
[0533] The server extracts the necessary statistical information from the stored data. It uses SQL queries to filter the data based on specific conditions, and natural language processing (NLP) techniques to extract information based on specific keywords or categories. For example, libraries such as NLTK or SpaCy are used. The extracted data is preprocessed to remove noise, impute missing values, and standardize formats (for example, standardizing date formats to YYYY-MM-DD).
[0534] Analysis for credibility assessment
[0535] The server uses a generative AI model (e.g., GPT-3 or BERT) to evaluate the veracity of the preprocessed data. The generative AI model is given a prompt like this:
[0536] "Please rate the reliability of this data. We will calculate a reliability score for each of the data collected from official government databases, major news sites, and social media, and tell us which data you think is the most trustworthy."
[0537] The generative AI model then calculates a credibility score for each source, taking into account factors such as the source's reliability, revision history, and data consistency.
[0538] Recognizing user emotions with an emotion engine
[0539] The device uses a camera and microphone to recognize the user's emotions in real time. The camera analyzes the user's facial expressions using facial recognition technologies such as OpenCV and Face++. The microphone uses voice analysis technologies such as Google Speech API and IBM Watson to determine the user's emotions from the tone and speed of the voice. For example, it can identify whether the user is laughing, angry, or surprised.
[0540] Providing a credibility score
[0541] When a user sends a data verification request through their device, the server searches the database for the relevant data and its credibility score and returns it to the device. This data is sent in JSON format, so the device can analyze it and display it visually in graphs and tables. For example, data with high reliability is displayed in green, and data with low reliability is displayed in red.
[0542] Emotion-based content adjustment
[0543] The device adjusts the display content based on the user's emotional data analyzed by the emotion engine. For example, if the user shows a surprised expression, the device will display additional relevant information (such as infection forecasts and policy response information). If the user is sad, the display content will be simplified and positive information will be emphasized. Adjustments will be made, such as emphasizing the increase in the number of recovered people rather than the increase in the number of infected people.
[0544] In this way, the system is designed to allow users to easily evaluate the reliability of data collected from multiple sources and to provide the most appropriate information according to the user's feelings.
[0545] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0546] Step 1:
[0547] The server collects data from various sources. Specifically, it uses APIs and web scraping technology to obtain data from public databases, news sites, social media, etc. The collected data is stored in a cloud database. For example, BeautifulSoup can be used to collect article content from news sites. The input is a list of source URLs, and the output is the collected raw data.
[0548] Step 2:
[0549] The server extracts the necessary statistical information from the raw data stored in the cloud database. Specifically, it uses SQL queries and natural language processing techniques to filter the data and extract information based on specific keywords or categories. The input is the raw data, and the output is the extracted statistical information. A specific example of how it works is to use an SQL query to extract data on the "number of new coronavirus infections."
[0550] Step 3:
[0551] The server preprocesses the extracted statistical information. Specifically, it removes noise, fills in missing values, and standardizes formats. For example, it standardizes date formats and aligns numerical units. The input is the extracted statistical information, and the output is the preprocessed data. A specific operation is to standardize all date formats to "YYYY-MM-DD".
[0552] Step 4:
[0553] The server analyzes the preprocessed data using a generative AI model to evaluate its credibility. The input is the preprocessed data, and the output is a credibility score. Specifically, the server prompts the generative AI model as follows: "Please rate the credibility of this data. Based on data collected from public government databases, major news sites, and social media, calculate a credibility score for each, and tell us which data is the most trustworthy."
[0554] Step 5:
[0555] The device uses a camera and microphone to capture the user's face and voice in real time, and analyzes them using an emotion engine. The input is the user's real-time video and audio data, and the output is the user's emotion data. Specific operations include facial recognition using OpenCV and voice analysis using the Google Speech API.
[0556] Step 6:
[0557] When a user sends a data verification request from a device, the server searches the database for the relevant data and its credibility score and returns it to the device. The input is the user's request data, and the output is the relevant data and its credibility score. Specific operations include sending data to the device in JSON format.
[0558] Step 7:
[0559] The terminal visually displays the received data and credibility score in graph and table format, and the credibility score is displayed in color. The input is the data received from the server and the credibility score, and the output is the visually displayed information. Specific operations include drawing graphs using Matplotlib.
[0560] Step 8:
[0561] The device adjusts the display content based on the user's emotional data analyzed by the emotion engine. The input is the user's emotional data, and the output is the adjusted display content. Specific operations include displaying additional information when the user expresses surprise, and emphasizing positive information when the user expresses sadness.
[0562] (Application example 2)
[0563] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0564] In modern society, there is a need to collect data from many sources and evaluate its credibility. At the same time, it is also important to provide appropriate information based on individual user emotions. However, conventional systems have had difficulty simultaneously evaluating the credibility of data and displaying information based on user emotions. Furthermore, in the advertising field, there is a need to display appropriate advertisements based on user emotions, but existing technologies have not been able to adequately meet this need. A system that can meet these complex requirements is needed.
[0565] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes a data collection means, a means for extracting necessary statistical data from the collected data, a means for pre-processing the extracted data, a means for analyzing the pre-processed data using a generative model for evaluating the credibility of the pre-processed data, a means for calculating a credibility score based on the analysis result, a means for providing the credibility score and data to the user, and a sentiment analysis means for recognizing the user's sentiment and adjusting the display content. This allows the user not only to evaluate the credibility of data collected from a wide variety of information sources, but also to display appropriate information and advertisements in real time according to their sentiments.
[0566] "Data collection means" refers to a method of automatically obtaining necessary data from various sources such as public databases, news sites, and social media via the Internet and storing it in a cloud database.
[0567] The "means for extracting statistical data" refers to a means for extracting statistical data required by a user from the collected data based on specific keywords or categories.
[0568] "Preprocessing means" refers to a means of performing preprocessing such as noise removal, missing value completion, and format standardization on the extracted data, and converting it into an analyzable format.
[0569] "Means for performing analysis using a generative model" refers to means for performing data analysis using a generative AI model to evaluate the credibility of preprocessed data.
[0570] The "means for calculating a credibility score" is a means for converting the credibility of each piece of data into a numerical value based on the analysis results and calculating a score.
[0571] The "means for providing the credibility score and data to the user" refers to a means for visually displaying the calculated credibility score and related data to the user.
[0572] The "emotion analysis means" is a means for analyzing the user's face and voice in real time using a camera and microphone, recognizing the user's emotions, and adjusting the display content accordingly.
[0573] A "generative AI model" is a machine learning model trained on a large dataset and used to assess the veracity of the data.
[0574] The system related to this invention, "AdOptimizer with Emotion Engine," consists of the following steps: data collection, data extraction, preprocessing, analysis using a generative model, credibility evaluation and score calculation, emotion analysis, and display of advertisements. This system can adjust the display content based on the user's emotions and provide optimal advertisements.
[0575] The server uses a data collection means to automatically obtain necessary data via the Internet from public databases, news sites, social media, etc., and stores it in a cloud database. Next, a means for extracting necessary statistical data from the collected data is used to extract data required by the user based on specific keywords or categories. The extracted data is then processed using a preprocessing means to remove noise, fill in missing values, standardize formats, etc.
[0576] The pre-processed data is analyzed using a generative AI model to assess the veracity of the data. Based on the analysis results, a veracity score is calculated. These scores and data are visually displayed to a user by a means for providing the veracity score and data to a user.
[0577] The device captures the user's face and voice in real time using emotion analysis means. The emotion analysis means uses a camera and microphone to analyze the user's emotions using facial recognition and voice analysis technologies. Specifically, it recognizes emotions such as joy, sadness, surprise, and anger from facial expressions and determines emotions from the tone and speed of the voice.
[0578] When a user requests advertising data from a device, the server sends the data and credibility score to the device based on the request. The device visually displays this information in the form of graphs, tables, etc., and the emotion analysis means adjusts the display content based on the user's emotions. For example, if the user expresses surprise, detailed advertising information is displayed, and if the user expresses sadness, a calm advertisement is presented.
[0579] As a concrete example, consider a robot in a store displaying advertisements. The robot's camera captures the user's facial expressions, and the microphone analyzes their voice. If the user smiles (happiness), a regular advertisement (e.g., a promotion for a new product) is displayed. On the other hand, if the user is sad, a more gentle advertisement (e.g., an introduction to a relaxing drink) is displayed.
[0580] An example of a prompt sentence is "If the user is excited, provide detailed advertising information." This prompt sentence is used as an instruction for the generative AI model to provide appropriate advertising based on the user's emotions.
[0581] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0582] Step 1:
[0583] The server uses data collection tools to automatically retrieve the necessary data from public databases, news sites, social media, etc. via the Internet and stores it in a cloud database. Specifically, data is collected using APIs and web scraping technology. The input requires the URL of the source and an API key, and the output is the retrieved raw data.
[0584] Step 2:
[0585] The server processes the collected raw data using a data extraction method to extract the statistical data required by the user. Specifically, it uses natural language processing techniques and database queries to extract information based on specific keywords or categories. The input is the raw data, and the output is the extracted statistical data.
[0586] Step 3:
[0587] The server preprocesses the extracted statistical data using preprocessing means. Specifically, it performs noise removal, missing value completion, and format standardization. For example, it standardizes date formats and converts numerical units. The input is the extracted statistical data, and the output is the preprocessed data.
[0588] Step 4:
[0589] The server uses a generative AI model to evaluate the credibility of the preprocessed data and perform analysis. Specifically, it analyzes the data based on the reliability of the source, past revision history, data consistency, etc. The input is the preprocessed data, and the output is the analysis result and a credibility score.
[0590] Step 5:
[0591] The server uses a means to calculate a credibility score based on the analysis results. Specifically, it assigns a credibility score calculated by the generative AI model to each piece of data. The input is the analysis results, and the output is data with a credibility score.
[0592] Step 6:
[0593] When a user requests advertising data from a terminal, the server uses a means to provide the user with a credibility score and the data. Specifically, it searches the database for the relevant data and sends it to the user's terminal. The input is the user's request, and the output is the data with a credibility score.
[0594] Step 7:
[0595] The device uses emotion analysis means to recognize the user's emotions. Specifically, it analyzes facial expressions captured by a camera and voice captured by a microphone, and identifies emotions using facial recognition technology and voice analysis technology. The input is the captured face and voice data, and the output is the analyzed user's emotion data.
[0596] Step 8:
[0597] The terminal adjusts the advertisement content to be displayed based on the user's emotional data. Specifically, the advertisement content is changed according to the user's emotional state. For example, when the user expresses surprise, detailed advertisement information is displayed, and when the user expresses sadness, calm advertisement content is displayed. The input is the user's emotional data, and the output is the adjusted advertisement content.
[0598] Step 9:
[0599] The terminal visually displays the adjusted advertising content to the user. Specifically, the advertising content is displayed on the screen using graphs and tables to allow the user to intuitively understand it. The input is the adjusted advertising content, and the output is the displayed advertisement.
[0600] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0601] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0602] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.
[0603] [Third embodiment]
[0604] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.
[0605] 5, the data processing system 310 includes the data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[0606] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0607] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.
[0608] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0609] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0610] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0611] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0612] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0613] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0614] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0615] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."
[0616] The system related to this invention, "DataVerity," collects data from various sources, evaluates its credibility using a generative AI model, and provides users with the data and a credibility score. The program processing content and specific examples of this system are shown below.
[0617] Data collection
[0618] The server collects data from various sources on the Internet, including public government databases, internal corporate databases, news sites, social media, etc. Using APIs and web scraping technology, the server automatically retrieves the necessary data and stores it in a cloud database.
[0619] Data extraction and preprocessing
[0620] The server extracts the necessary statistical data from the stored database. This data extraction is performed using database queries and natural language processing techniques. The extracted data undergoes preprocessing, such as noise removal, missing value completion, and format standardization. For example, the data is converted into an analyzable format by standardizing date formats and converting numerical units.
[0621] Analysis for credibility assessment
[0622] The server uses a generative AI model to evaluate the credibility of the pre-processed data. The model analyzes the credibility of the data based on multiple evaluation criteria, such as the reliability of the source, past revision history, and data consistency, and then calculates a credibility score for each piece of data.
[0623] Providing and displaying credibility scores
[0624] When a user wants to check data from their device, they send a request. Based on the request, the server searches the database for the relevant data and its credibility score, and sends it back to the device. The device then visually displays the received data and credibility score to the user. The display format is easy to understand using graphs and tables, and the credibility score is color-coded. This allows the user to visually check the credibility of the data.
[0625] Specific examples
[0626] For example, consider the case of collecting data on the number of new coronavirus infections and evaluating its credibility. The server collects data on the number of new coronavirus infections from official government databases, major news sites, and social media. Next, the server extracts the number of new infections, deaths, and recovered cases from the collected data, standardizing the date format and removing noise.
[0627] The server then uses a generative AI model to analyze the credibility of the data from each source. Data from public databases is assigned a high score (e.g., 90 points or higher) because it has high reliability, data from news sites is assigned a medium score (e.g., 70 to 80 points), and data from anonymous posts on social media is assigned a low score (e.g., 40 to 60 points).
[0628] Finally, when a user requests data on the number of coronavirus cases from their device, the server sends the data and credibility score to the device based on the request. The device then displays this information visually in graphs and tables, and users can click on specific data to view detailed analysis results.
[0629] As described above, this system enables users to easily evaluate the credibility of data collected from a wide variety of information sources and make decisions quickly and accurately based on highly reliable data.
[0630] The processing flow will be explained below.
[0631] Program processing flow
[0632] Step 1:
[0633] The server collects data from designated sources (public government databases, internal corporate databases, news sites, social media, etc.), automatically retrieves the data using APIs and web scraping technology, and stores it in a cloud database.
[0634] Step 2:
[0635] The server extracts the statistical data the user needs from the data stored in the database, using database queries and natural language processing techniques to retrieve information based on specific keywords or categories.
[0636] Step 3:
[0637] The server preprocesses the extracted data, removing noise (e.g., removing irrelevant and duplicate data), imputing missing values, and standardizing data formats (e.g., converting date formats and numeric units).
[0638] Step 4:
[0639] The server inputs the preprocessed data into a generative AI model to evaluate its credibility. The generative AI model analyzes the data based on multiple criteria, including the reliability of the source, past revision history, and data consistency.
[0640] Step 5:
[0641] The server calculates a credibility score for each piece of data based on the analysis results of the generative AI model. The credibility score is expressed in the range of 0 to 100 and is assigned individually to each piece of data.
[0642] Step 6:
[0643] The server stores the calculated credibility score and the reason for the evaluation in a database, which allows for quick responses to subsequent user requests.
[0644] Step 7:
[0645] When a user wants to check statistical data from their device, they send a request for specific information, including search keywords and the range of data they require.
[0646] Step 8:
[0647] The server searches the database for relevant statistical data and credibility scores based on the user request, and returns them to the terminal together with detailed analysis results.
[0648] Step 9:
[0649] The device visually displays the received data and the credibility score to the user in the form of graphs and tables, and the credibility score is color-coded for easy understanding.
[0650] Step 10:
[0651] If a user wants to check the details of a particular piece of data, they can click on it to see detailed analysis results and the reasons for the evaluation, allowing users to understand the credibility of the data and make decisions based on reliable information.
[0652] Through these processing steps, the DataVerity system helps users effectively evaluate the credibility of data from a wide variety of sources and make decisions quickly and accurately.
[0653] Example 1
[0654] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0655] There is a need for an appropriate system that can quickly and accurately evaluate the credibility of data collected from various sources and provide the results visually to users. However, current systems require a lot of manual work in the collection, preprocessing, credibility evaluation, and provision of data, which is time-consuming and labor-intensive, and has low reliability. Therefore, an efficient and reliable data collection and evaluation system is needed.
[0656] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0657] In this invention, the server includes a data collection means, a means for extracting necessary information from the collected data, a means for preprocessing the extracted data, a means for analyzing the preprocessed data using a generative model for evaluating the credibility of the data, a means for calculating a credibility score based on the analysis result, and a means for providing the credibility score and data to a user. This enables a user to quickly and accurately collect data from various information sources, evaluate the credibility of the data, and then visually confirm it.
[0658] "Data collection means" refers to a function for automatically acquiring data from various sources on the Internet.
[0659] "Means for extracting necessary information from collected data" refers to a function for selecting specific information from collected data and formatting it into a format suitable for data analysis.
[0660] "Means for preprocessing extracted data" refers to functions for removing noise, filling in missing values, standardizing formats, etc. from the extracted data, and preparing it in a format suitable for analysis.
[0661] "Means for performing analysis using a generative model to assess the credibility of preprocessed data" refers to a function that performs data analysis using a generative AI model to assess the reliability of preprocessed data.
[0662] The "means for calculating the credibility score" is a function for quantifying and indicating the reliability of data evaluated using a generative model.
[0663] "Means for providing credibility scores and data to users" refers to a function for providing analyzed credibility scores and related data to users in a visually verifiable format.
[0664] The system of the present invention collects data from various sources, evaluates its credibility using a generative AI model, and provides users with the data and a credibility score. Specific embodiments and methods of use are described below.
[0665] The system includes the following main elements:
[0666] Data collection methods
[0667] A means of extracting the necessary information from the collected data
[0668] A means of preprocessing the extracted data
[0669] A means of analyzing using generative models to assess the veracity of preprocessed data
[0670] A method for calculating a credibility score based on the analysis results
[0671] Means of providing credibility scores and data to users
[0672] Data collection methods
[0673] The server automatically collects data from various sources on the Internet using API access and web scraping techniques. For example, it uses the Python libraries BeautifulSoup and Scrapy to retrieve data from public databases, news sites, and social media.
[0674] A means of extracting the necessary information from the collected data
[0675] The server selects specific information from the collected data and formats it into a format suitable for data analysis. Specifically, it uses SQL queries to extract data such as the number of new infections, deaths, and recovered cases from the database.
[0676] A means of preprocessing the extracted data
[0677] The server performs preprocessing on the extracted data, specifically removing noise, filling in missing values, standardizing formats, etc. By using the Python Pandas library, the server consistently standardizes date formats and converts numeric units.
[0678] A means of analyzing using generative models to assess the veracity of preprocessed data
[0679] The server then inputs the preprocessed data into a generative AI model to assess its credibility. This generative model takes into account the reliability of the source, past revision history, and data consistency, and uses natural language processing (NLP) techniques, such as Transformer-based models (such as BERT and GPT).
[0680] A method for calculating a credibility score based on the analysis results
[0681] The server calculates a credibility score based on the analysis results of the generative model. This score is expressed as a number and serves as a standard for evaluating the reliability of the data. For example, public databases have a high score (90 points or above), news sites have a medium score (70-80 points), and social media data has a low score (40-60 points).
[0682] Means of providing credibility scores and data to users
[0683] When a user uses a device to send a data verification request, the server searches for and returns the relevant data and credibility score. The device then visually displays this data. Specifically, the received data is displayed in graphs and tables, and the credibility score is color-coded. Users can click on the data to view detailed analysis results.
[0684] Examples of concrete examples and prompts
[0685] For example, consider the case of evaluating the credibility of data on the number of people infected with the new coronavirus. In this case, the server collects relevant data from official government databases, major news sites, and social media, and inputs it into a generative AI model for analysis. After calculating a credibility score, it displays it on the device. The following is an example of a specific prompt sentence:
[0686] "Collect data on the number of COVID-19 cases and rate its credibility. Collect data from each source, use a generative AI model to score its credibility, and display it visually."
[0687] As described above, the present invention enables users to easily evaluate the credibility of data collected from a variety of information sources and make decisions quickly and accurately.
[0688] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0689] Step 1:
[0690] The server collects data from various sources on the Internet, specifically using API access and web scraping technology. It uses a list of URLs from public government databases, news sites, social media, etc. as input, and automatically retrieves the latest data from them. It stores the retrieved data in a cloud database. The output is a collection of the collected raw data.
[0691] Specifically, data extraction is performed automatically using Python libraries such as BeautifulSoup and Scrapy.
[0692] Step 2:
[0693] The server extracts the necessary information from the collected raw data. It uses the raw data as input and uses database queries and natural language processing (NLP) techniques to select specific data such as the number of new cases, deaths, and recoveries. The output is the extracted data.
[0694] Specifically, it uses SQL queries to extract specific columns and rows from a database and then uses NLP techniques to analyze the text data.
[0695] Step 3:
[0696] The server performs preprocessing on the extracted data. It uses the extracted data as input and performs noise removal, missing value imputation, format standardization, etc. The output is preprocessed data that can be analyzed.
[0697] Specifically, it uses Python's Pandas library to clean the data and standardize date formats and numerical units.
[0698] Step 4:
[0699] The server inputs the preprocessed data into a generative AI model to analyze it and evaluate its credibility. As input, the preprocessed data is provided to a generative model (e.g., BERT or GPT), and the model evaluates the credibility of the data. The output is a credibility score for each data point.
[0700] Specifically, it uses a generative model to extract features from the data and quantify their reliability.
[0701] Step 5:
[0702] The server calculates a credibility score based on the analysis results of the generative model. It uses the evaluation results of the generative model as input and applies an algorithm to calculate the credibility score. The output is a dataset containing the calculated credibility scores.
[0703] Specifically, the evaluation algorithm is executed to quantify the reliability and store it in a database.
[0704] Step 6:
[0705] When a user uses a terminal to send a request for data verification, the server searches for the relevant data and credibility score and returns it.The server receives the user's request as input and searches the database for the corresponding data and credibility score.The output is the search result: the data and credibility score.
[0706] Specifically, it receives an HTTP request from a user and executes a corresponding SQL query on a database.
[0707] Step 7:
[0708] The terminal visually displays the received data and credibility scores. It uses the data and credibility scores received from the server as input and displays them in the form of graphs and tables. The output is a visual representation of the data provided to the user.
[0709] Specifically, the data is displayed on a web page using JavaScript and HTML, and the reliability is color-coded to make it easier to understand visually. Users can click to see detailed analysis results.
[0710] (Application example 1)
[0711] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0712] In modern society, there is a huge amount of information available on the Internet, making it extremely difficult to obtain accurate and reliable information from all that information. News and important data, in particular, often contain unreliable information, requiring users to expend considerable effort to determine the accuracy of the information. Furthermore, many news articles are constantly being updated, creating a lack of means to assess their credibility in real time. Therefore, there is a need for a system that can quickly and accurately assess the credibility of information and provide it to users.
[0713] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0714] In this invention, the server includes a data collection means, a means for extracting necessary statistical data from the collected data, a means for pre-processing the extracted data, a means for analyzing the pre-processed data using a generative model for evaluating the credibility of the data, a means for calculating a credibility score based on the analysis result, and a means for providing the credibility score and data to a user in a news feed format, thereby enabling the user to evaluate the credibility of news articles in real time and make quick and accurate decisions based on highly reliable information.
[0715] "Data collection means" refers to technology for automatically obtaining necessary data from various information sources on the Internet.
[0716] "Means for extracting statistical data" refers to techniques for extracting necessary statistical information from collected data.
[0717] "Preprocessing means" refers to technology that converts extracted data into an analyzable format by removing noise, filling in missing values, and standardizing the format.
[0718] "Means for performing analysis using generative models" refers to techniques for evaluating the veracity of data that has been preprocessed using a generative AI model.
[0719] The "means for calculating the credibility score" is a technology that quantifies the reliability of each piece of data based on the analysis results of the generative AI model.
[0720] "Newsfeed-style provision" refers to a technology that visually provides users with credibility scores and data in real time.
[0721] "Source credibility" is a criterion for evaluating the past performance and reliability of the source of data.
[0722] "Edit history" is historical information about how data has been edited in the past.
[0723] "Data consistency" is a criterion for assessing the consistency and agreement of data obtained from multiple sources.
[0724] "User interface" refers to the screen or operating means through which a user can interact with the system and visually check the credibility score of data.
[0725] A system for realizing this invention includes a data collection means, a means for extracting statistical data, a data preprocessing means, a means for performing credibility assessment using a generative AI model, a means for calculating a credibility score, and a means for providing the data and the credibility score in a news feed format.
[0726] Program processing and technologies used
[0727] Data collection methods
[0728] The server collects data from various sources on the Internet, including public government databases, corporate databases, news sites, and social media, using APIs and web scraping techniques. Specifically, web scraping is performed using Python's requests and BeautifulSoup libraries.
[0729] A means of extracting statistical data
[0730] The server extracts the necessary statistics from the collected data, using database queries and natural language processing (NLP) techniques to extract key information such as article titles and content, for example using the Python nltk library.
[0731] Data preprocessing measures
[0732] The server removes noise from the extracted data, imputes missing values, and standardizes the format. Specifically, it standardizes date formats and converts numerical units. This converts the data into an analyzable format. For example, it uses the pandas library.
[0733] Credibility assessment tools
[0734] The server evaluates the credibility of the preprocessed data using a generative AI model, which analyzes the credibility of the data based on criteria such as the reliability of the source, past revision history, and data consistency. For example, it applies a logistic regression model using scikit-learn.
[0735] How the credibility score is calculated
[0736] The server calculates a credibility score based on the analysis results of the generative AI model. This score is a numerical representation of the reliability of the data, and is expressed, for example, in the range of 0 to 1.
[0737] A means to provide data and credibility scores in a news feed format
[0738] The system allows users to check data and credibility scores in a news feed format on their devices. The credibility scores are displayed in color, and by clicking on specific data, detailed analysis results are displayed. For example, using Django as the web application framework, the data and scores are displayed visually in graphs and tables.
[0739] Specific examples
[0740] For example, if a user wants to check the latest news on the number of new coronavirus infections, the server collects relevant data from public databases, news sites, and social media. The generative AI model assigns a credibility score to each source, assigning a high score (90 points or above) to data from public databases, a medium score (70 to 80 points) to data from news sites, and a low score (40 to 60 points) to data from social media. Users can launch the app on their smartphone and check the credibility score of each article along with the data displayed in a news feed format.
[0741] Prompt Sentence Examples
[0742] "What is the credibility score for the latest news on the number of new coronavirus cases?"
[0743] As described above, this invention collects data from various sources on the Internet, evaluates its credibility, and provides it to users, enabling them to make decisions based on highly reliable information in real time.
[0744] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0745] Step 1:
[0746] The server collects data from various sources on the Internet. This includes scraping and using APIs to obtain data from public government databases, news sites, social media, etc. Specifically, it uses Python's requests library to obtain the content of web pages and the BeautifulSoup library to extract the required information. The input is the URL or API endpoint of the source, and the output is the collected raw data.
[0747] Step 2:
[0748] The server extracts the necessary statistical data from the collected data. In this case, statistical information such as the title and body of news articles is extracted from the raw data. For example, the Python nltk library is used to parse the article text and extract important keywords and sentences. The input is the raw data, and the output is the extracted statistical data.
[0749] Step 3:
[0750] The server performs preprocessing on the extracted data. This preprocessing includes removing noise from the data, imputing missing values, standardizing date formats, and converting numerical units. Specifically, it uses the Python pandas library to create a data frame, imputing missing values, and aligning the data format. The input is the extracted statistical data, and the output is the preprocessed statistical data.
[0751] Step 4:
[0752] The server uses a generative AI model to evaluate the credibility of the preprocessed data. This evaluation includes criteria such as the reliability of the source, past revision history, and data consistency. Specifically, it uses scikit-learn to apply a logistic regression model to calculate a credibility score for each data point. The input is the preprocessed data, and the output is the credibility score.
[0753] Step 5:
[0754] The server calculates a credibility score based on the analysis results. This score is a numerical representation of the reliability of the data and is expressed in a range from 0 to 1. The calculated score is converted into a format that is easy for users to visually check. The input is the credibility evaluation result provided by the generative AI model, and the output is the credibility score.
[0755] Step 6:
[0756] The server provides the credibility score and data to the user in the form of a news feed. The user can view the news feed with the credibility score via their smartphone. The credibility score is color-coded, and by clicking on specific data, detailed analysis results are displayed. This is done using a web application framework such as Django. The input is the credibility score and data, and the output is a visually displayed news feed.
[0757] Prompt Sentence Examples
[0758] The user enters the following prompt into the system:
[0759] "What is the credibility score for the latest news on the number of new coronavirus cases?"
[0760] Users can view real-time news feeds with credibility scores.
[0761] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0762] The system related to this invention, "DataVerity with Emotion Engine," collects data from various information sources, evaluates its credibility using a generative AI model, and combines it with an emotion engine that recognizes the user's emotions and adjusts the displayed content. The program processing details and specific examples of this system are shown below.
[0763] Data collection
[0764] The server collects data from various sources on the Internet, including public government databases, internal corporate databases, news sites, social media, etc. Using APIs and web scraping technology, the server automatically retrieves the necessary data and stores it in a cloud database.
[0765] Data extraction and preprocessing
[0766] The server extracts the statistical data required by the user from the data stored in the database. It uses database queries and natural language processing techniques to extract information based on specific keywords or categories. The extracted data undergoes preprocessing, such as noise removal, missing value completion, and format standardization. For example, standardizing date formats and converting numerical units converts the data into an analyzable format.
[0767] Analysis for credibility assessment
[0768] The server uses a generative AI model to evaluate the credibility of the pre-processed data. The model analyzes the data based on multiple criteria, such as the reliability of the source, past revision history, and data consistency, and calculates a credibility score for each piece of data.
[0769] Recognizing user emotions with an emotion engine
[0770] The device uses a camera and microphone to capture the user's face and voice in real time. The emotion engine uses facial recognition and voice analysis technologies to analyze the user's emotions. Specifically, it recognizes joy, sadness, surprise, anger, etc. from facial expressions and determines emotions from the tone and speed of the voice.
[0771] Providing and displaying credibility scores
[0772] When a user wants to check data from their device, they send a request. Based on the request, the server searches the database for the relevant data and its credibility score, and sends it back to the device. The device then visually displays the received data and credibility score to the user. The display format is easy to understand using graphs and tables, and the credibility score is displayed in a color-coded format.
[0773] Emotion-based content adjustment
[0774] The device adjusts the display content based on the user's emotions analyzed by the emotion engine. For example, if the user expresses surprise, the device displays additional related data or provides a detailed explanation. If the user expresses sadness or anger, the device simplifies the display content and provides an interface that reduces the user's stress.
[0775] Specific examples
[0776] For example, consider the case of collecting data on the number of new coronavirus infections and evaluating its credibility. The server collects data on the number of new coronavirus infections from official government databases, major news sites, and social media. Next, the server extracts the number of new infections, deaths, and recovered cases from the collected data, standardizing the date format and removing noise.
[0777] The server then uses a generative AI model to analyze the credibility of the data from each source. Data from public databases is assigned a high score (e.g., 90 points or higher) because it has high reliability, data from news sites is assigned a medium score (e.g., 70 to 80 points), and data from anonymous posts on social media is assigned a low score (e.g., 40 to 60 points).
[0778] Finally, when a user requests data on the number of coronavirus cases from their device, the server sends the data and credibility score based on the request. The device then visually displays this information in graphs and tables, and an emotion engine adjusts the display based on the user's emotions. For example, if the user expresses surprise, additional infection forecast data and policy response information will be displayed. If the user expresses sadness, improvements and positive information will be highlighted.
[0779] As described above, this system allows users to evaluate the credibility of data collected from a wide variety of sources while taking their emotions into account, and makes decisions quickly and accurately based on highly reliable data.
[0780] The processing flow will be explained below.
[0781] Program processing flow
[0782] Step 1:
[0783] The server collects data from designated sources (public government databases, internal corporate databases, news sites, social media, etc.), automatically retrieves the data using APIs and web scraping technology, and stores it in a cloud database.
[0784] Step 2:
[0785] The server extracts the statistical data the user needs from the data stored in the database, using database queries and natural language processing techniques to retrieve information based on specific keywords or categories.
[0786] Step 3:
[0787] The server preprocesses the extracted data, removing noise (e.g., removing irrelevant and duplicate data), imputing missing values, and standardizing data formats (e.g., converting date formats and numeric units).
[0788] Step 4:
[0789] The server inputs the preprocessed data into a generative AI model to evaluate its credibility. The generative AI model analyzes the data based on multiple criteria, including the reliability of the source, past revision history, and data consistency.
[0790] Step 5:
[0791] The server calculates a credibility score for each piece of data based on the analysis results of the generative AI model. The credibility score is expressed in the range of 0 to 100 and is assigned individually to each piece of data.
[0792] Step 6:
[0793] The server stores the calculated credibility score and the reason for the evaluation in a database, which allows for quick responses to subsequent user requests.
[0794] Step 7:
[0795] When a user wants to check statistical data from their device, they send a request for specific information, including search keywords and the range of data they require.
[0796] Step 8:
[0797] The server searches the database for relevant statistical data and credibility scores based on the user request, and returns them to the terminal together with detailed analysis results.
[0798] Step 9:
[0799] The device visually displays the received data and the credibility score to the user in the form of graphs and tables, and the credibility score is color-coded for easy understanding.
[0800] Step 10:
[0801] The device uses a camera and microphone to capture the user's face and voice in real time. The emotion engine uses facial recognition and voice analysis technologies to analyze the user's emotions. It recognizes joy, sadness, surprise, anger, etc. from facial expressions and determines emotions from the tone and speed of the voice.
[0802] Step 11:
[0803] The device adjusts the display content based on the user's emotions analyzed by the emotion engine. For example, if the user expresses surprise, the device displays additional related data or provides a detailed explanation. If the user expresses sadness or anger, the device simplifies the display content and provides an interface that reduces the user's stress.
[0804] Example 2
[0805] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0806] Conventional data collection and analysis systems lack the ability to evaluate the credibility of data from diverse sources. Furthermore, few systems adjust information based on user sentiment, and interfaces for improving the user experience are lacking. Therefore, there is a need for systems that provide highly credible data and display information that takes user sentiment into consideration.
[0807] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes a data collection means, a means for extracting necessary statistical information from the collected data, a means for pre-processing the extracted data, a means for analyzing the pre-processed data using a generative model to evaluate the credibility of the pre-processed data, a means for calculating a credibility score based on the analysis result, a means for providing the credibility score and data to the user, a means for performing face recognition and voice analysis to recognize the user's emotions, and a means for adjusting display content based on the user's emotions. This makes it possible to evaluate the credibility of data obtained from various information sources, and further, by adjusting the information display according to the user's emotions, an improved user experience is realized.
[0808] "Data collection means" refers to devices or programs that automatically acquire necessary data from various information sources and store it in a database.
[0809] "Statistical information" is numerical or textual data collected for a specific purpose and analyzed accordingly.
[0810] "Preprocessing" is the process of converting collected data into a format suitable for analysis by performing processes such as noise removal, missing value completion, and format standardization on the data.
[0811] A "generative model" is an artificial intelligence or machine learning algorithm used to input a variety of data and perform tasks such as evaluating credibility and extracting patterns.
[0812] The "credibility score" is an evaluation index that numerically represents the reliability of preprocessed data.
[0813] "Facial recognition" is a technology that uses a camera to capture a user's face and analyze their facial expressions to recognize their emotions.
[0814] "Voice analysis" is a technology that analyzes voice data collected through a microphone and recognizes emotions and intentions from the tone and speed of the voice.
[0815] "Adjusting display content" is the process of optimizing the format and content of displayed information based on the user's emotions to improve the user experience.
[0816] "Diverse sources" refers to information providers from different sources, such as public government databases, news sites, social media, and internal corporate databases.
[0817] "User interface" refers to the screens and tools that allow users to visually check and manipulate data and credibility scores.
[0818] MODE FOR CARRYING OUT THE INVENTION
[0819] The system according to the present invention, "data credibility evaluation system," is composed of the following specific hardware and software, and collects, analyzes, and displays data.
[0820] Data collection
[0821] The server collects data from various sources on the Internet. To do this, it uses APIs (Application Programming Interfaces) and web scraping technology. For example, it retrieves government-issued data from the API of public government databases, or uses tools such as BeautifulSoup and Selenium to extract necessary data from websites. The collected data is then stored in cloud databases such as Amazon RDS and Google BigQuery.
[0822] Data extraction and preprocessing
[0823] The server extracts the necessary statistical information from the stored data. It uses SQL queries to filter the data based on specific conditions, and natural language processing (NLP) techniques to extract information based on specific keywords or categories. For example, libraries such as NLTK or SpaCy are used. The extracted data is preprocessed to remove noise, impute missing values, and standardize formats (for example, standardizing date formats to YYYY-MM-DD).
[0824] Analysis for credibility assessment
[0825] The server uses a generative AI model (e.g., GPT-3 or BERT) to evaluate the veracity of the preprocessed data. The generative AI model is given a prompt like this:
[0826] "Please rate the reliability of this data. We will calculate a reliability score for each of the data collected from official government databases, major news sites, and social media, and tell us which data you think is the most trustworthy."
[0827] The generative AI model then calculates a credibility score for each source, taking into account factors such as the source's reliability, revision history, and data consistency.
[0828] Recognizing user emotions with an emotion engine
[0829] The device uses a camera and microphone to recognize the user's emotions in real time. The camera analyzes the user's facial expressions using facial recognition technologies such as OpenCV and Face++. The microphone uses voice analysis technologies such as Google Speech API and IBM Watson to determine the user's emotions from the tone and speed of the voice. For example, it can identify whether the user is laughing, angry, or surprised.
[0830] Providing a credibility score
[0831] When a user sends a data verification request through their device, the server searches the database for the relevant data and its credibility score and returns it to the device. This data is sent in JSON format, so the device can analyze it and display it visually in graphs and tables. For example, data with high reliability is displayed in green, and data with low reliability is displayed in red.
[0832] Emotion-based content adjustment
[0833] The device adjusts the display content based on the user's emotional data analyzed by the emotion engine. For example, if the user shows a surprised expression, the device will display additional relevant information (such as infection forecasts and policy response information). If the user is sad, the display content will be simplified and positive information will be emphasized. Adjustments will be made, such as emphasizing the increase in the number of recovered people rather than the increase in the number of infected people.
[0834] In this way, the system is designed to allow users to easily evaluate the reliability of data collected from multiple sources and to provide the most appropriate information according to the user's feelings.
[0835] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0836] Step 1:
[0837] The server collects data from various sources. Specifically, it uses APIs and web scraping technology to obtain data from public databases, news sites, social media, etc. The collected data is stored in a cloud database. For example, BeautifulSoup can be used to collect article content from news sites. The input is a list of source URLs, and the output is the collected raw data.
[0838] Step 2:
[0839] The server extracts the necessary statistical information from the raw data stored in the cloud database. Specifically, it uses SQL queries and natural language processing techniques to filter the data and extract information based on specific keywords or categories. The input is the raw data, and the output is the extracted statistical information. A specific example of how it works is to use an SQL query to extract data on the "number of new coronavirus infections."
[0840] Step 3:
[0841] The server preprocesses the extracted statistical information. Specifically, it removes noise, fills in missing values, and standardizes formats. For example, it standardizes date formats and aligns numerical units. The input is the extracted statistical information, and the output is the preprocessed data. A specific operation is to standardize all date formats to "YYYY-MM-DD".
[0842] Step 4:
[0843] The server analyzes the preprocessed data using a generative AI model to evaluate its credibility. The input is the preprocessed data, and the output is a credibility score. Specifically, the server prompts the generative AI model as follows: "Please rate the credibility of this data. Based on data collected from public government databases, major news sites, and social media, calculate a credibility score for each, and tell us which data is the most trustworthy."
[0844] Step 5:
[0845] The device uses a camera and microphone to capture the user's face and voice in real time, and analyzes them using an emotion engine. The input is the user's real-time video and audio data, and the output is the user's emotion data. Specific operations include facial recognition using OpenCV and voice analysis using the Google Speech API.
[0846] Step 6:
[0847] When a user sends a data verification request from a device, the server searches the database for the relevant data and its credibility score and returns it to the device. The input is the user's request data, and the output is the relevant data and its credibility score. Specific operations include sending data to the device in JSON format.
[0848] Step 7:
[0849] The terminal visually displays the received data and credibility score in graph and table format, and the credibility score is displayed in color. The input is the data received from the server and the credibility score, and the output is the visually displayed information. Specific operations include drawing graphs using Matplotlib.
[0850] Step 8:
[0851] The device adjusts the display content based on the user's emotional data analyzed by the emotion engine. The input is the user's emotional data, and the output is the adjusted display content. Specific operations include displaying additional information when the user expresses surprise, and emphasizing positive information when the user expresses sadness.
[0852] (Application example 2)
[0853] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0854] In modern society, there is a need to collect data from many sources and evaluate its credibility. At the same time, it is also important to provide appropriate information based on individual user emotions. However, conventional systems have had difficulty simultaneously evaluating the credibility of data and displaying information based on user emotions. Furthermore, in the advertising field, there is a need to display appropriate advertisements based on user emotions, but existing technologies have not been able to adequately meet this need. A system that can meet these complex requirements is needed.
[0855] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes a data collection means, a means for extracting necessary statistical data from the collected data, a means for pre-processing the extracted data, a means for analyzing the pre-processed data using a generative model for evaluating the credibility of the pre-processed data, a means for calculating a credibility score based on the analysis result, a means for providing the credibility score and data to the user, and a sentiment analysis means for recognizing the user's sentiment and adjusting the display content. This allows the user not only to evaluate the credibility of data collected from a wide variety of information sources, but also to display appropriate information and advertisements in real time according to their sentiments.
[0856] "Data collection means" refers to a method of automatically obtaining necessary data from various sources such as public databases, news sites, and social media via the Internet and storing it in a cloud database.
[0857] The "means for extracting statistical data" refers to a means for extracting statistical data required by a user from the collected data based on specific keywords or categories.
[0858] "Preprocessing means" refers to a means of performing preprocessing such as noise removal, missing value completion, and format standardization on the extracted data, and converting it into an analyzable format.
[0859] "Means for performing analysis using a generative model" refers to means for performing data analysis using a generative AI model to evaluate the credibility of preprocessed data.
[0860] The "means for calculating a credibility score" is a means for converting the credibility of each piece of data into a numerical value based on the analysis results and calculating a score.
[0861] The "means for providing the credibility score and data to the user" refers to a means for visually displaying the calculated credibility score and related data to the user.
[0862] The "emotion analysis means" is a means for analyzing the user's face and voice in real time using a camera and microphone, recognizing the user's emotions, and adjusting the display content accordingly.
[0863] A "generative AI model" is a machine learning model trained on a large dataset and used to assess the veracity of the data.
[0864] The system related to this invention, "AdOptimizer with Emotion Engine," consists of the following steps: data collection, data extraction, preprocessing, analysis using a generative model, credibility evaluation and score calculation, emotion analysis, and display of advertisements. This system can adjust the display content based on the user's emotions and provide optimal advertisements.
[0865] The server uses a data collection means to automatically obtain necessary data via the Internet from public databases, news sites, social media, etc., and stores it in a cloud database. Next, a means for extracting necessary statistical data from the collected data is used to extract data required by the user based on specific keywords or categories. The extracted data is then processed using a preprocessing means to remove noise, fill in missing values, standardize formats, etc.
[0866] The pre-processed data is analyzed using a generative AI model to assess the veracity of the data. Based on the analysis results, a veracity score is calculated. These scores and data are visually displayed to a user by a means for providing the veracity score and data to a user.
[0867] The device captures the user's face and voice in real time using emotion analysis means. The emotion analysis means uses a camera and microphone to analyze the user's emotions using facial recognition and voice analysis technologies. Specifically, it recognizes emotions such as joy, sadness, surprise, and anger from facial expressions and determines emotions from the tone and speed of the voice.
[0868] When a user requests advertising data from a device, the server sends the data and credibility score to the device based on the request. The device visually displays this information in the form of graphs, tables, etc., and the emotion analysis means adjusts the display content based on the user's emotions. For example, if the user expresses surprise, detailed advertising information is displayed, and if the user expresses sadness, a calm advertisement is presented.
[0869] As a concrete example, consider a robot in a store displaying advertisements. The robot's camera captures the user's facial expressions, and the microphone analyzes their voice. If the user smiles (happiness), a regular advertisement (e.g., a promotion for a new product) is displayed. On the other hand, if the user is sad, a more gentle advertisement (e.g., an introduction to a relaxing drink) is displayed.
[0870] An example of a prompt sentence is "If the user is excited, provide detailed advertising information." This prompt sentence is used as an instruction for the generative AI model to provide appropriate advertising based on the user's emotions.
[0871] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0872] Step 1:
[0873] The server uses data collection tools to automatically retrieve the necessary data from public databases, news sites, social media, etc. via the Internet and stores it in a cloud database. Specifically, data is collected using APIs and web scraping technology. The input requires the URL of the source and an API key, and the output is the retrieved raw data.
[0874] Step 2:
[0875] The server processes the collected raw data using a data extraction method to extract the statistical data required by the user. Specifically, it uses natural language processing techniques and database queries to extract information based on specific keywords or categories. The input is the raw data, and the output is the extracted statistical data.
[0876] Step 3:
[0877] The server preprocesses the extracted statistical data using preprocessing means. Specifically, it performs noise removal, missing value completion, and format standardization. For example, it standardizes date formats and converts numerical units. The input is the extracted statistical data, and the output is the preprocessed data.
[0878] Step 4:
[0879] The server uses a generative AI model to evaluate the credibility of the preprocessed data and perform analysis. Specifically, it analyzes the data based on the reliability of the source, past revision history, data consistency, etc. The input is the preprocessed data, and the output is the analysis result and a credibility score.
[0880] Step 5:
[0881] The server uses a means to calculate a credibility score based on the analysis results. Specifically, it assigns a credibility score calculated by the generative AI model to each piece of data. The input is the analysis results, and the output is data with a credibility score.
[0882] Step 6:
[0883] When a user requests advertising data from a terminal, the server uses a means to provide the user with a credibility score and the data. Specifically, it searches the database for the relevant data and sends it to the user's terminal. The input is the user's request, and the output is the data with a credibility score.
[0884] Step 7:
[0885] The device uses emotion analysis means to recognize the user's emotions. Specifically, it analyzes facial expressions captured by a camera and voice captured by a microphone, and identifies emotions using facial recognition technology and voice analysis technology. The input is the captured face and voice data, and the output is the analyzed user's emotion data.
[0886] Step 8:
[0887] The terminal adjusts the advertisement content to be displayed based on the user's emotional data. Specifically, the advertisement content is changed according to the user's emotional state. For example, when the user expresses surprise, detailed advertisement information is displayed, and when the user expresses sadness, calm advertisement content is displayed. The input is the user's emotional data, and the output is the adjusted advertisement content.
[0888] Step 9:
[0889] The terminal visually displays the adjusted advertising content to the user. Specifically, the advertising content is displayed on the screen using graphs and tables to allow the user to intuitively understand it. The input is the adjusted advertising content, and the output is the displayed advertisement.
[0890] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0891] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0892] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.
[0893] [Fourth embodiment]
[0894] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.
[0895] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[0896] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0897] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.
[0898] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0899] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0900] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0901] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.
[0902] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0903] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0904] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0905] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0906] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[0907] The system related to this invention, "DataVerity," collects data from various sources, evaluates its credibility using a generative AI model, and provides users with the data and a credibility score. The program processing content and specific examples of this system are shown below.
[0908] Data collection
[0909] The server collects data from various sources on the Internet, including public government databases, internal corporate databases, news sites, social media, etc. Using APIs and web scraping technology, the server automatically retrieves the necessary data and stores it in a cloud database.
[0910] Data extraction and preprocessing
[0911] The server extracts the necessary statistical data from the stored database. This data extraction is performed using database queries and natural language processing techniques. The extracted data undergoes preprocessing, such as noise removal, missing value completion, and format standardization. For example, the data is converted into an analyzable format by standardizing date formats and converting numerical units.
[0912] Analysis for credibility assessment
[0913] The server uses a generative AI model to evaluate the credibility of the pre-processed data. The model analyzes the credibility of the data based on multiple evaluation criteria, such as the reliability of the source, past revision history, and data consistency, and then calculates a credibility score for each piece of data.
[0914] Providing and displaying credibility scores
[0915] When a user wants to check data from their device, they send a request. Based on the request, the server searches the database for the relevant data and its credibility score, and sends it back to the device. The device then visually displays the received data and credibility score to the user. The display format is easy to understand using graphs and tables, and the credibility score is color-coded. This allows the user to visually check the credibility of the data.
[0916] Specific examples
[0917] For example, consider the case of collecting data on the number of new coronavirus infections and evaluating its credibility. The server collects data on the number of new coronavirus infections from official government databases, major news sites, and social media. Next, the server extracts the number of new infections, deaths, and recovered cases from the collected data, standardizing the date format and removing noise.
[0918] The server then uses a generative AI model to analyze the credibility of the data from each source. Data from public databases is assigned a high score (e.g., 90 points or higher) because it has high reliability, data from news sites is assigned a medium score (e.g., 70 to 80 points), and data from anonymous posts on social media is assigned a low score (e.g., 40 to 60 points).
[0919] Finally, when a user requests data on the number of coronavirus cases from their device, the server sends the data and credibility score to the device based on the request. The device then displays this information visually in graphs and tables, and users can click on specific data to view detailed analysis results.
[0920] As described above, this system enables users to easily evaluate the credibility of data collected from a wide variety of information sources and make decisions quickly and accurately based on highly reliable data.
[0921] The processing flow will be explained below.
[0922] Program processing flow
[0923] Step 1:
[0924] The server collects data from designated sources (public government databases, internal corporate databases, news sites, social media, etc.), automatically retrieves the data using APIs and web scraping technology, and stores it in a cloud database.
[0925] Step 2:
[0926] The server extracts the statistical data the user needs from the data stored in the database, using database queries and natural language processing techniques to retrieve information based on specific keywords or categories.
[0927] Step 3:
[0928] The server preprocesses the extracted data, removing noise (e.g., removing irrelevant and duplicate data), imputing missing values, and standardizing data formats (e.g., converting date formats and numeric units).
[0929] Step 4:
[0930] The server inputs the preprocessed data into a generative AI model to evaluate its credibility. The generative AI model analyzes the data based on multiple criteria, including the reliability of the source, past revision history, and data consistency.
[0931] Step 5:
[0932] The server calculates a credibility score for each piece of data based on the analysis results of the generative AI model. The credibility score is expressed in the range of 0 to 100 and is assigned individually to each piece of data.
[0933] Step 6:
[0934] The server stores the calculated credibility score and the reason for the evaluation in a database, which allows for quick responses to subsequent user requests.
[0935] Step 7:
[0936] When a user wants to check statistical data from their device, they send a request for specific information, including search keywords and the range of data they require.
[0937] Step 8:
[0938] The server searches the database for relevant statistical data and credibility scores based on the user request, and returns them to the terminal together with detailed analysis results.
[0939] Step 9:
[0940] The device visually displays the received data and the credibility score to the user in the form of graphs and tables, and the credibility score is color-coded for easy understanding.
[0941] Step 10:
[0942] If a user wants to check the details of a particular piece of data, they can click on it to see detailed analysis results and the reasons for the evaluation, allowing users to understand the credibility of the data and make decisions based on reliable information.
[0943] Through these processing steps, the DataVerity system helps users effectively evaluate the credibility of data from a wide variety of sources and make decisions quickly and accurately.
[0944] Example 1
[0945] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[0946] There is a need for an appropriate system that can quickly and accurately evaluate the credibility of data collected from various sources and provide the results visually to users. However, current systems require a lot of manual work in the collection, preprocessing, credibility evaluation, and provision of data, which is time-consuming and labor-intensive, and has low reliability. Therefore, an efficient and reliable data collection and evaluation system is needed.
[0947] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0948] In this invention, the server includes a data collection means, a means for extracting necessary information from the collected data, a means for preprocessing the extracted data, a means for analyzing the preprocessed data using a generative model for evaluating the credibility of the data, a means for calculating a credibility score based on the analysis result, and a means for providing the credibility score and data to a user. This enables a user to quickly and accurately collect data from various information sources, evaluate the credibility of the data, and then visually confirm it.
[0949] "Data collection means" refers to a function for automatically acquiring data from various sources on the Internet.
[0950] "Means for extracting necessary information from collected data" refers to a function for selecting specific information from collected data and formatting it into a format suitable for data analysis.
[0951] "Means for preprocessing extracted data" refers to functions for removing noise, filling in missing values, standardizing formats, etc. from the extracted data, and preparing it in a format suitable for analysis.
[0952] "Means for performing analysis using a generative model to assess the credibility of preprocessed data" refers to a function that performs data analysis using a generative AI model to assess the reliability of preprocessed data.
[0953] The "means for calculating the credibility score" is a function for quantifying and indicating the reliability of data evaluated using a generative model.
[0954] "Means for providing credibility scores and data to users" refers to a function for providing analyzed credibility scores and related data to users in a visually verifiable format.
[0955] The system of the present invention collects data from various sources, evaluates its credibility using a generative AI model, and provides users with the data and a credibility score. Specific embodiments and methods of use are described below.
[0956] The system includes the following main elements:
[0957] Data collection methods
[0958] A means of extracting the necessary information from the collected data
[0959] A means of preprocessing the extracted data
[0960] A means of analyzing using generative models to assess the veracity of preprocessed data
[0961] A method for calculating a credibility score based on the analysis results
[0962] Means of providing credibility scores and data to users
[0963] Data collection methods
[0964] The server automatically collects data from various sources on the Internet using API access and web scraping techniques. For example, it uses the Python libraries BeautifulSoup and Scrapy to retrieve data from public databases, news sites, and social media.
[0965] A means of extracting the necessary information from the collected data
[0966] The server selects specific information from the collected data and formats it into a format suitable for data analysis. Specifically, it uses SQL queries to extract data such as the number of new infections, deaths, and recovered cases from the database.
[0967] A means of preprocessing the extracted data
[0968] The server performs preprocessing on the extracted data, specifically removing noise, filling in missing values, standardizing formats, etc. By using the Python Pandas library, the server consistently standardizes date formats and converts numeric units.
[0969] A means of analyzing using generative models to assess the veracity of preprocessed data
[0970] The server then inputs the preprocessed data into a generative AI model to assess its credibility. This generative model takes into account the reliability of the source, past revision history, and data consistency, and uses natural language processing (NLP) techniques, such as Transformer-based models (such as BERT and GPT).
[0971] A method for calculating a credibility score based on the analysis results
[0972] The server calculates a credibility score based on the analysis results of the generative model. This score is expressed as a number and serves as a standard for evaluating the reliability of the data. For example, public databases have a high score (90 points or above), news sites have a medium score (70-80 points), and social media data has a low score (40-60 points).
[0973] Means of providing credibility scores and data to users
[0974] When a user uses a device to send a data verification request, the server searches for and returns the relevant data and credibility score. The device then visually displays this data. Specifically, the received data is displayed in graphs and tables, and the credibility score is color-coded. Users can click on the data to view detailed analysis results.
[0975] Examples of concrete examples and prompts
[0976] For example, consider the case of evaluating the credibility of data on the number of people infected with the new coronavirus. In this case, the server collects relevant data from official government databases, major news sites, and social media, and inputs it into a generative AI model for analysis. After calculating a credibility score, it displays it on the device. The following is an example of a specific prompt sentence:
[0977] "Collect data on the number of COVID-19 cases and rate its credibility. Collect data from each source, use a generative AI model to score its credibility, and display it visually."
[0978] As described above, the present invention enables users to easily evaluate the credibility of data collected from a variety of information sources and make decisions quickly and accurately.
[0979] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0980] Step 1:
[0981] The server collects data from various sources on the Internet, specifically using API access and web scraping technology. It uses a list of URLs from public government databases, news sites, social media, etc. as input, and automatically retrieves the latest data from them. It stores the retrieved data in a cloud database. The output is a collection of the collected raw data.
[0982] Specifically, data extraction is performed automatically using Python libraries such as BeautifulSoup and Scrapy.
[0983] Step 2:
[0984] The server extracts the necessary information from the collected raw data. It uses the raw data as input and uses database queries and natural language processing (NLP) techniques to select specific data such as the number of new cases, deaths, and recoveries. The output is the extracted data.
[0985] Specifically, it uses SQL queries to extract specific columns and rows from a database and then uses NLP techniques to analyze the text data.
[0986] Step 3:
[0987] The server performs preprocessing on the extracted data. It uses the extracted data as input and performs noise removal, missing value imputation, format standardization, etc. The output is preprocessed data that can be analyzed.
[0988] Specifically, it uses Python's Pandas library to clean the data and standardize date formats and numerical units.
[0989] Step 4:
[0990] The server inputs the preprocessed data into a generative AI model to analyze it and evaluate its credibility. As input, the preprocessed data is provided to a generative model (e.g., BERT or GPT), and the model evaluates the credibility of the data. The output is a credibility score for each data point.
[0991] Specifically, it uses a generative model to extract features from the data and quantify their reliability.
[0992] Step 5:
[0993] The server calculates a credibility score based on the analysis results of the generative model. It uses the evaluation results of the generative model as input and applies an algorithm to calculate the credibility score. The output is a dataset containing the calculated credibility scores.
[0994] Specifically, the evaluation algorithm is executed to quantify the reliability and store it in a database.
[0995] Step 6:
[0996] When a user uses a terminal to send a request for data verification, the server searches for the relevant data and credibility score and returns it.The server receives the user's request as input and searches the database for the corresponding data and credibility score.The output is the search result: the data and credibility score.
[0997] Specifically, it receives an HTTP request from a user and executes a corresponding SQL query on a database.
[0998] Step 7:
[0999] The terminal visually displays the received data and credibility scores. It uses the data and credibility scores received from the server as input and displays them in the form of graphs and tables. The output is a visual representation of the data provided to the user.
[1000] Specifically, the data is displayed on a web page using JavaScript and HTML, and the reliability is color-coded to make it easier to understand visually. Users can click to see detailed analysis results.
[1001] (Application example 1)
[1002] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1003] In modern society, there is a huge amount of information available on the Internet, making it extremely difficult to obtain accurate and reliable information from all that information. News and important data, in particular, often contain unreliable information, requiring users to expend considerable effort to determine the accuracy of the information. Furthermore, many news articles are constantly being updated, creating a lack of means to assess their credibility in real time. Therefore, there is a need for a system that can quickly and accurately assess the credibility of information and provide it to users.
[1004] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1005] In this invention, the server includes a data collection means, a means for extracting necessary statistical data from the collected data, a means for pre-processing the extracted data, a means for analyzing the pre-processed data using a generative model for evaluating the credibility of the data, a means for calculating a credibility score based on the analysis result, and a means for providing the credibility score and data to a user in a news feed format, thereby enabling the user to evaluate the credibility of news articles in real time and make quick and accurate decisions based on highly reliable information.
[1006] "Data collection means" refers to technology for automatically obtaining necessary data from various information sources on the Internet.
[1007] "Means for extracting statistical data" refers to techniques for extracting necessary statistical information from collected data.
[1008] "Preprocessing means" refers to technology that converts extracted data into an analyzable format by removing noise, filling in missing values, and standardizing the format.
[1009] "Means for performing analysis using generative models" refers to techniques for evaluating the veracity of data that has been preprocessed using a generative AI model.
[1010] The "means for calculating the credibility score" is a technology that quantifies the reliability of each piece of data based on the analysis results of the generative AI model.
[1011] "Newsfeed-style provision" refers to a technology that visually provides users with credibility scores and data in real time.
[1012] "Source credibility" is a criterion for evaluating the past performance and reliability of the source of data.
[1013] "Edit history" is historical information about how data has been edited in the past.
[1014] "Data consistency" is a criterion for assessing the consistency and agreement of data obtained from multiple sources.
[1015] "User interface" refers to the screen or operating means through which a user can interact with the system and visually check the credibility score of data.
[1016] A system for realizing this invention includes a data collection means, a means for extracting statistical data, a data preprocessing means, a means for performing credibility assessment using a generative AI model, a means for calculating a credibility score, and a means for providing the data and the credibility score in a news feed format.
[1017] Program processing and technologies used
[1018] Data collection methods
[1019] The server collects data from various sources on the Internet, including public government databases, corporate databases, news sites, and social media, using APIs and web scraping techniques. Specifically, web scraping is performed using Python's requests and BeautifulSoup libraries.
[1020] A means of extracting statistical data
[1021] The server extracts the necessary statistics from the collected data, using database queries and natural language processing (NLP) techniques to extract key information such as article titles and content, for example using the Python nltk library.
[1022] Data preprocessing measures
[1023] The server removes noise from the extracted data, imputes missing values, and standardizes the format. Specifically, it standardizes date formats and converts numerical units. This converts the data into an analyzable format. For example, it uses the pandas library.
[1024] Credibility assessment tools
[1025] The server evaluates the credibility of the preprocessed data using a generative AI model, which analyzes the credibility of the data based on criteria such as the reliability of the source, past revision history, and data consistency. For example, it applies a logistic regression model using scikit-learn.
[1026] How the credibility score is calculated
[1027] The server calculates a credibility score based on the analysis results of the generative AI model. This score is a numerical representation of the reliability of the data, and is expressed, for example, in the range of 0 to 1.
[1028] A means to provide data and credibility scores in a news feed format
[1029] The system allows users to check data and credibility scores in a news feed format on their devices. The credibility scores are displayed in color, and by clicking on specific data, detailed analysis results are displayed. For example, using Django as the web application framework, the data and scores are displayed visually in graphs and tables.
[1030] Specific examples
[1031] For example, if a user wants to check the latest news on the number of new coronavirus infections, the server collects relevant data from public databases, news sites, and social media. The generative AI model assigns a credibility score to each source, assigning a high score (90 points or above) to data from public databases, a medium score (70 to 80 points) to data from news sites, and a low score (40 to 60 points) to data from social media. Users can launch the app on their smartphone and check the credibility score of each article along with the data displayed in a news feed format.
[1032] Prompt Sentence Examples
[1033] "What is the credibility score for the latest news on the number of new coronavirus cases?"
[1034] As described above, this invention collects data from various sources on the Internet, evaluates its credibility, and provides it to users, enabling them to make decisions based on highly reliable information in real time.
[1035] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1036] Step 1:
[1037] The server collects data from various sources on the Internet. This includes scraping and using APIs to obtain data from public government databases, news sites, social media, etc. Specifically, it uses Python's requests library to obtain the content of web pages and the BeautifulSoup library to extract the required information. The input is the URL or API endpoint of the source, and the output is the collected raw data.
[1038] Step 2:
[1039] The server extracts the necessary statistical data from the collected data. In this case, statistical information such as the title and body of news articles is extracted from the raw data. For example, the Python nltk library is used to parse the article text and extract important keywords and sentences. The input is the raw data, and the output is the extracted statistical data.
[1040] Step 3:
[1041] The server performs preprocessing on the extracted data. This preprocessing includes removing noise from the data, imputing missing values, standardizing date formats, and converting numerical units. Specifically, it uses the Python pandas library to create a data frame, imputing missing values, and aligning the data format. The input is the extracted statistical data, and the output is the preprocessed statistical data.
[1042] Step 4:
[1043] The server uses a generative AI model to evaluate the credibility of the preprocessed data. This evaluation includes criteria such as the reliability of the source, past revision history, and data consistency. Specifically, it uses scikit-learn to apply a logistic regression model to calculate a credibility score for each data point. The input is the preprocessed data, and the output is the credibility score.
[1044] Step 5:
[1045] The server calculates a credibility score based on the analysis results. This score is a numerical representation of the reliability of the data and is expressed in a range from 0 to 1. The calculated score is converted into a format that is easy for users to visually check. The input is the credibility evaluation result provided by the generative AI model, and the output is the credibility score.
[1046] Step 6:
[1047] The server provides the credibility score and data to the user in the form of a news feed. The user can view the news feed with the credibility score via their smartphone. The credibility score is color-coded, and by clicking on specific data, detailed analysis results are displayed. This is done using a web application framework such as Django. The input is the credibility score and data, and the output is a visually displayed news feed.
[1048] Prompt Sentence Examples
[1049] The user enters the following prompt into the system:
[1050] "What is the credibility score for the latest news on the number of new coronavirus cases?"
[1051] Users can view real-time news feeds with credibility scores.
[1052] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1053] The system related to this invention, "DataVerity with Emotion Engine," collects data from various information sources, evaluates its credibility using a generative AI model, and combines it with an emotion engine that recognizes the user's emotions and adjusts the displayed content. The program processing details and specific examples of this system are shown below.
[1054] Data collection
[1055] The server collects data from various sources on the Internet, including public government databases, internal corporate databases, news sites, social media, etc. Using APIs and web scraping technology, the server automatically retrieves the necessary data and stores it in a cloud database.
[1056] Data extraction and preprocessing
[1057] The server extracts the statistical data required by the user from the data stored in the database. It uses database queries and natural language processing techniques to extract information based on specific keywords or categories. The extracted data undergoes preprocessing, such as noise removal, missing value completion, and format standardization. For example, standardizing date formats and converting numerical units converts the data into an analyzable format.
[1058] Analysis for credibility assessment
[1059] The server uses a generative AI model to evaluate the credibility of the pre-processed data. The model analyzes the data based on multiple criteria, such as the reliability of the source, past revision history, and data consistency, and calculates a credibility score for each piece of data.
[1060] Recognizing user emotions with an emotion engine
[1061] The device uses a camera and microphone to capture the user's face and voice in real time. The emotion engine uses facial recognition and voice analysis technologies to analyze the user's emotions. Specifically, it recognizes joy, sadness, surprise, anger, etc. from facial expressions and determines emotions from the tone and speed of the voice.
[1062] Providing and displaying credibility scores
[1063] When a user wants to check data from their device, they send a request. Based on the request, the server searches the database for the relevant data and its credibility score, and sends it back to the device. The device then visually displays the received data and credibility score to the user. The display format is easy to understand using graphs and tables, and the credibility score is displayed in a color-coded format.
[1064] Emotion-based content adjustment
[1065] The device adjusts the display content based on the user's emotions analyzed by the emotion engine. For example, if the user expresses surprise, the device displays additional related data or provides a detailed explanation. If the user expresses sadness or anger, the device simplifies the display content and provides an interface that reduces the user's stress.
[1066] Specific examples
[1067] For example, consider the case of collecting data on the number of new coronavirus infections and evaluating its credibility. The server collects data on the number of new coronavirus infections from official government databases, major news sites, and social media. Next, the server extracts the number of new infections, deaths, and recovered cases from the collected data, standardizing the date format and removing noise.
[1068] The server then uses a generative AI model to analyze the credibility of the data from each source. Data from public databases is assigned a high score (e.g., 90 points or higher) because it has high reliability, data from news sites is assigned a medium score (e.g., 70 to 80 points), and data from anonymous posts on social media is assigned a low score (e.g., 40 to 60 points).
[1069] Finally, when a user requests data on the number of coronavirus cases from their device, the server sends the data and credibility score based on the request. The device then visually displays this information in graphs and tables, and an emotion engine adjusts the display based on the user's emotions. For example, if the user expresses surprise, additional infection forecast data and policy response information will be displayed. If the user expresses sadness, improvements and positive information will be highlighted.
[1070] As described above, this system allows users to evaluate the credibility of data collected from a wide variety of sources while taking their emotions into account, and makes decisions quickly and accurately based on highly reliable data.
[1071] The processing flow will be explained below.
[1072] Program processing flow
[1073] Step 1:
[1074] The server collects data from designated sources (public government databases, internal corporate databases, news sites, social media, etc.), automatically retrieves the data using APIs and web scraping technology, and stores it in a cloud database.
[1075] Step 2:
[1076] The server extracts the statistical data the user needs from the data stored in the database, using database queries and natural language processing techniques to retrieve information based on specific keywords or categories.
[1077] Step 3:
[1078] The server preprocesses the extracted data, removing noise (e.g., removing irrelevant and duplicate data), imputing missing values, and standardizing data formats (e.g., converting date formats and numeric units).
[1079] Step 4:
[1080] The server inputs the preprocessed data into a generative AI model to evaluate its credibility. The generative AI model analyzes the data based on multiple criteria, including the reliability of the source, past revision history, and data consistency.
[1081] Step 5:
[1082] The server calculates a credibility score for each piece of data based on the analysis results of the generative AI model. The credibility score is expressed in the range of 0 to 100 and is assigned individually to each piece of data.
[1083] Step 6:
[1084] The server stores the calculated credibility score and the reason for the evaluation in a database, which allows for quick responses to subsequent user requests.
[1085] Step 7:
[1086] When a user wants to check statistical data from their device, they send a request for specific information, including search keywords and the range of data they require.
[1087] Step 8:
[1088] The server searches the database for relevant statistical data and credibility scores based on the user request, and returns them to the terminal together with detailed analysis results.
[1089] Step 9:
[1090] The device visually displays the received data and the credibility score to the user in the form of graphs and tables, and the credibility score is color-coded for easy understanding.
[1091] Step 10:
[1092] The device uses a camera and microphone to capture the user's face and voice in real time. The emotion engine uses facial recognition and voice analysis technologies to analyze the user's emotions. It recognizes joy, sadness, surprise, anger, etc. from facial expressions and determines emotions from the tone and speed of the voice.
[1093] Step 11:
[1094] The device adjusts the display content based on the user's emotions analyzed by the emotion engine. For example, if the user expresses surprise, the device displays additional related data or provides a detailed explanation. If the user expresses sadness or anger, the device simplifies the display content and provides an interface that reduces the user's stress.
[1095] Example 2
[1096] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1097] Conventional data collection and analysis systems lack the ability to evaluate the credibility of data from diverse sources. Furthermore, few systems adjust information based on user sentiment, and interfaces for improving the user experience are lacking. Therefore, there is a need for systems that provide highly credible data and display information that takes user sentiment into consideration.
[1098] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes a data collection means, a means for extracting necessary statistical information from the collected data, a means for pre-processing the extracted data, a means for analyzing the pre-processed data using a generative model to evaluate the credibility of the pre-processed data, a means for calculating a credibility score based on the analysis result, a means for providing the credibility score and data to the user, a means for performing face recognition and voice analysis to recognize the user's emotions, and a means for adjusting display content based on the user's emotions. This makes it possible to evaluate the credibility of data obtained from various information sources, and further, by adjusting the information display according to the user's emotions, an improved user experience is realized.
[1099] "Data collection means" refers to devices or programs that automatically acquire necessary data from various information sources and store it in a database.
[1100] "Statistical information" is numerical or textual data collected for a specific purpose and analyzed accordingly.
[1101] "Preprocessing" is the process of converting collected data into a format suitable for analysis by performing processes such as noise removal, missing value completion, and format standardization on the data.
[1102] A "generative model" is an artificial intelligence or machine learning algorithm used to input a variety of data and perform tasks such as evaluating credibility and extracting patterns.
[1103] The "credibility score" is an evaluation index that numerically represents the reliability of preprocessed data.
[1104] "Facial recognition" is a technology that uses a camera to capture a user's face and analyze their facial expressions to recognize their emotions.
[1105] "Voice analysis" is a technology that analyzes voice data collected through a microphone and recognizes emotions and intentions from the tone and speed of the voice.
[1106] "Adjusting display content" is the process of optimizing the format and content of displayed information based on the user's emotions to improve the user experience.
[1107] "Diverse sources" refers to information providers from different sources, such as public government databases, news sites, social media, and internal corporate databases.
[1108] "User interface" refers to the screens and tools that allow users to visually check and manipulate data and credibility scores.
[1109] MODE FOR CARRYING OUT THE INVENTION
[1110] The system according to the present invention, "data credibility evaluation system," is composed of the following specific hardware and software, and collects, analyzes, and displays data.
[1111] Data collection
[1112] The server collects data from various sources on the Internet. To do this, it uses APIs (Application Programming Interfaces) and web scraping technology. For example, it retrieves government-issued data from the API of public government databases, or uses tools such as BeautifulSoup and Selenium to extract necessary data from websites. The collected data is then stored in cloud databases such as Amazon RDS and Google BigQuery.
[1113] Data extraction and preprocessing
[1114] The server extracts the necessary statistical information from the stored data. It uses SQL queries to filter the data based on specific conditions, and natural language processing (NLP) techniques to extract information based on specific keywords or categories. For example, libraries such as NLTK or SpaCy are used. The extracted data is preprocessed to remove noise, impute missing values, and standardize formats (for example, standardizing date formats to YYYY-MM-DD).
[1115] Analysis for credibility assessment
[1116] The server uses a generative AI model (e.g., GPT-3 or BERT) to evaluate the veracity of the preprocessed data. The generative AI model is given a prompt like this:
[1117] "Please rate the reliability of this data. We will calculate a reliability score for each of the data collected from official government databases, major news sites, and social media, and tell us which data you think is the most trustworthy."
[1118] The generative AI model then calculates a credibility score for each source, taking into account factors such as the source's reliability, revision history, and data consistency.
[1119] Recognizing user emotions with an emotion engine
[1120] The device uses a camera and microphone to recognize the user's emotions in real time. The camera analyzes the user's facial expressions using facial recognition technologies such as OpenCV and Face++. The microphone uses voice analysis technologies such as Google Speech API and IBM Watson to determine the user's emotions from the tone and speed of the voice. For example, it can identify whether the user is laughing, angry, or surprised.
[1121] Providing a credibility score
[1122] When a user sends a data verification request through their device, the server searches the database for the relevant data and its credibility score and returns it to the device. This data is sent in JSON format, so the device can analyze it and display it visually in graphs and tables. For example, data with high reliability is displayed in green, and data with low reliability is displayed in red.
[1123] Emotion-based content adjustment
[1124] The device adjusts the display content based on the user's emotional data analyzed by the emotion engine. For example, if the user shows a surprised expression, the device will display additional relevant information (such as infection forecasts and policy response information). If the user is sad, the display content will be simplified and positive information will be emphasized. Adjustments will be made, such as emphasizing the increase in the number of recovered people rather than the increase in the number of infected people.
[1125] In this way, the system is designed to allow users to easily evaluate the reliability of data collected from multiple sources and to provide the most appropriate information according to the user's feelings.
[1126] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1127] Step 1:
[1128] The server collects data from various sources. Specifically, it uses APIs and web scraping technology to obtain data from public databases, news sites, social media, etc. The collected data is stored in a cloud database. For example, BeautifulSoup can be used to collect article content from news sites. The input is a list of source URLs, and the output is the collected raw data.
[1129] Step 2:
[1130] The server extracts the necessary statistical information from the raw data stored in the cloud database. Specifically, it uses SQL queries and natural language processing techniques to filter the data and extract information based on specific keywords or categories. The input is the raw data, and the output is the extracted statistical information. A specific example of how it works is to use an SQL query to extract data on the "number of new coronavirus infections."
[1131] Step 3:
[1132] The server preprocesses the extracted statistical information. Specifically, it removes noise, fills in missing values, and standardizes formats. For example, it standardizes date formats and aligns numerical units. The input is the extracted statistical information, and the output is the preprocessed data. A specific operation is to standardize all date formats to "YYYY-MM-DD".
[1133] Step 4:
[1134] The server analyzes the preprocessed data using a generative AI model to evaluate its credibility. The input is the preprocessed data, and the output is a credibility score. Specifically, the server prompts the generative AI model as follows: "Please rate the credibility of this data. Based on data collected from public government databases, major news sites, and social media, calculate a credibility score for each, and tell us which data is the most trustworthy."
[1135] Step 5:
[1136] The device uses a camera and microphone to capture the user's face and voice in real time, and analyzes them using an emotion engine. The input is the user's real-time video and audio data, and the output is the user's emotion data. Specific operations include facial recognition using OpenCV and voice analysis using the Google Speech API.
[1137] Step 6:
[1138] When a user sends a data verification request from a device, the server searches the database for the relevant data and its credibility score and returns it to the device. The input is the user's request data, and the output is the relevant data and its credibility score. Specific operations include sending data to the device in JSON format.
[1139] Step 7:
[1140] The terminal visually displays the received data and credibility score in graph and table format, and the credibility score is displayed in color. The input is the data received from the server and the credibility score, and the output is the visually displayed information. Specific operations include drawing graphs using Matplotlib.
[1141] Step 8:
[1142] The device adjusts the display content based on the user's emotional data analyzed by the emotion engine. The input is the user's emotional data, and the output is the adjusted display content. Specific operations include displaying additional information when the user expresses surprise, and emphasizing positive information when the user expresses sadness.
[1143] (Application example 2)
[1144] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1145] In modern society, there is a need to collect data from many sources and evaluate its credibility. At the same time, it is also important to provide appropriate information based on individual user emotions. However, conventional systems have had difficulty simultaneously evaluating the credibility of data and displaying information based on user emotions. Furthermore, in the advertising field, there is a need to display appropriate advertisements based on user emotions, but existing technologies have not been able to adequately meet this need. A system that can meet these complex requirements is needed.
[1146] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes a data collection means, a means for extracting necessary statistical data from the collected data, a means for pre-processing the extracted data, a means for analyzing the pre-processed data using a generative model for evaluating the credibility of the pre-processed data, a means for calculating a credibility score based on the analysis result, a means for providing the credibility score and data to the user, and a sentiment analysis means for recognizing the user's sentiment and adjusting the display content. This allows the user not only to evaluate the credibility of data collected from a wide variety of information sources, but also to display appropriate information and advertisements in real time according to their sentiments.
[1147] "Data collection means" refers to a method of automatically obtaining necessary data from various sources such as public databases, news sites, and social media via the Internet and storing it in a cloud database.
[1148] The "means for extracting statistical data" refers to a means for extracting statistical data required by a user from the collected data based on specific keywords or categories.
[1149] "Preprocessing means" refers to a means of performing preprocessing such as noise removal, missing value completion, and format standardization on the extracted data, and converting it into an analyzable format.
[1150] "Means for performing analysis using a generative model" refers to means for performing data analysis using a generative AI model to evaluate the credibility of preprocessed data.
[1151] The "means for calculating a credibility score" is a means for converting the credibility of each piece of data into a numerical value based on the analysis results and calculating a score.
[1152] The "means for providing the credibility score and data to the user" refers to a means for visually displaying the calculated credibility score and related data to the user.
[1153] The "emotion analysis means" is a means for analyzing the user's face and voice in real time using a camera and microphone, recognizing the user's emotions, and adjusting the display content accordingly.
[1154] A "generative AI model" is a machine learning model trained on a large dataset and used to assess the veracity of the data.
[1155] The system related to this invention, "AdOptimizer with Emotion Engine," consists of the following steps: data collection, data extraction, preprocessing, analysis using a generative model, credibility evaluation and score calculation, emotion analysis, and display of advertisements. This system can adjust the display content based on the user's emotions and provide optimal advertisements.
[1156] The server uses a data collection means to automatically obtain necessary data via the Internet from public databases, news sites, social media, etc., and stores it in a cloud database. Next, a means for extracting necessary statistical data from the collected data is used to extract data required by the user based on specific keywords or categories. The extracted data is then processed using a preprocessing means to remove noise, fill in missing values, standardize formats, etc.
[1157] The pre-processed data is analyzed using a generative AI model to assess the veracity of the data. Based on the analysis results, a veracity score is calculated. These scores and data are visually displayed to a user by a means for providing the veracity score and data to a user.
[1158] The device captures the user's face and voice in real time using emotion analysis means. The emotion analysis means uses a camera and microphone to analyze the user's emotions using facial recognition and voice analysis technologies. Specifically, it recognizes emotions such as joy, sadness, surprise, and anger from facial expressions and determines emotions from the tone and speed of the voice.
[1159] When a user requests advertising data from a device, the server sends the data and credibility score to the device based on the request. The device visually displays this information in the form of graphs, tables, etc., and the emotion analysis means adjusts the display content based on the user's emotions. For example, if the user expresses surprise, detailed advertising information is displayed, and if the user expresses sadness, a calm advertisement is presented.
[1160] As a concrete example, consider a robot in a store displaying advertisements. The robot's camera captures the user's facial expressions, and the microphone analyzes their voice. If the user smiles (happiness), a regular advertisement (e.g., a promotion for a new product) is displayed. On the other hand, if the user is sad, a more gentle advertisement (e.g., an introduction to a relaxing drink) is displayed.
[1161] An example of a prompt sentence is "If the user is excited, provide detailed advertising information." This prompt sentence is used as an instruction for the generative AI model to provide appropriate advertising based on the user's emotions.
[1162] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1163] Step 1:
[1164] The server uses data collection tools to automatically retrieve the necessary data from public databases, news sites, social media, etc. via the Internet and stores it in a cloud database. Specifically, data is collected using APIs and web scraping technology. The input requires the URL of the source and an API key, and the output is the retrieved raw data.
[1165] Step 2:
[1166] The server processes the collected raw data using a data extraction method to extract the statistical data required by the user. Specifically, it uses natural language processing techniques and database queries to extract information based on specific keywords or categories. The input is the raw data, and the output is the extracted statistical data.
[1167] Step 3:
[1168] The server preprocesses the extracted statistical data using preprocessing means. Specifically, it performs noise removal, missing value completion, and format standardization. For example, it standardizes date formats and converts numerical units. The input is the extracted statistical data, and the output is the preprocessed data.
[1169] Step 4:
[1170] The server uses a generative AI model to evaluate the credibility of the preprocessed data and perform analysis. Specifically, it analyzes the data based on the reliability of the source, past revision history, data consistency, etc. The input is the preprocessed data, and the output is the analysis result and a credibility score.
[1171] Step 5:
[1172] The server uses a means to calculate a credibility score based on the analysis results. Specifically, it assigns a credibility score calculated by the generative AI model to each piece of data. The input is the analysis results, and the output is data with a credibility score.
[1173] Step 6:
[1174] When a user requests advertising data from a terminal, the server uses a means to provide the user with a credibility score and the data. Specifically, it searches the database for the relevant data and sends it to the user's terminal. The input is the user's request, and the output is the data with a credibility score.
[1175] Step 7:
[1176] The device uses emotion analysis means to recognize the user's emotions. Specifically, it analyzes facial expressions captured by a camera and voice captured by a microphone, and identifies emotions using facial recognition technology and voice analysis technology. The input is the captured face and voice data, and the output is the analyzed user's emotion data.
[1177] Step 8:
[1178] The terminal adjusts the advertisement content to be displayed based on the user's emotional data. Specifically, the advertisement content is changed according to the user's emotional state. For example, when the user expresses surprise, detailed advertisement information is displayed, and when the user expresses sadness, calm advertisement content is displayed. The input is the user's emotional data, and the output is the adjusted advertisement content.
[1179] Step 9:
[1180] The terminal visually displays the adjusted advertising content to the user. Specifically, the advertising content is displayed on the screen using graphs and tables to allow the user to intuitively understand it. The input is the adjusted advertising content, and the output is the displayed advertisement.
[1181] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.
[1182] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1183] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.
[1184] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[1185] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.
[1186] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.
[1187] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).
[1188] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.
[1189] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."
[1190] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.
[1191] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).
[1192] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.
[1193] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.
[1194] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[1195] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.
[1196] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.
[1197] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.
[1198] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.
[1199] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.
[1200] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.
[1201] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.
[1202] The following is further disclosed regarding the above embodiment.
[1203] (Claim 1)
[1204] data collection means;
[1205] A means for extracting necessary statistical data from the collected data;
[1206] means for pre-processing the extracted data;
[1207] means for analyzing the preprocessed data using a generative model to assess the authenticity of the preprocessed data;
[1208] means for calculating a credibility score based on the analysis results;
[1209] means for providing the authenticity scores and data to a user;
[1210] A system including:
[1211] (Claim 2)
[1212] 10. The system of claim 1, wherein the assessment of the credibility score includes the reliability of the source, past revision history, and data consistency.
[1213] (Claim 3)
[1214] 10. The system of claim 1, further comprising a user interface for visually displaying the credibility score.
[1215] "Example 1"
[1216] (Claim 1)
[1217] data collection means;
[1218] A means for extracting necessary information from the collected data;
[1219] means for pre-processing the extracted data;
[1220] means for analyzing the preprocessed data using a generative model to assess the authenticity of the preprocessed data;
[1221] means for calculating a credibility score based on the analysis results;
[1222] a means for providing the authenticity scores and data to users;
[1223] A system including:
[1224] (Claim 2)
[1225] 10. The system of claim 1, wherein the assessment of the credibility score includes the reliability of the source, past revision history, and data consistency.
[1226] (Claim 3)
[1227] 10. The system of claim 1, further comprising a user interface for visually displaying the credibility score.
[1228] "Application Example 1"
[1229] (Claim 1)
[1230] data collection means;
[1231] A means for extracting necessary statistical data from the collected data;
[1232] means for pre-processing the extracted data;
[1233] means for analyzing the preprocessed data using a generative model to assess the authenticity of the preprocessed data;
[1234] means for calculating a credibility score based on the analysis results;
[1235] means for providing the credibility scores and data to a user in the form of a news feed;
[1236] A system including:
[1237] (Claim 2)
[1238] 10. The system of claim 1, wherein the assessment of the credibility score includes the reliability of the source, past revision history, and data consistency.
[1239] (Claim 3)
[1240] 10. The system of claim 1, further comprising a user interface for visually displaying the credibility scores and allowing a user to select particular data.
[1241] "Example 2: Combining Emotion Engines"
[1242] (Claim 1)
[1243] data collection means;
[1244] a means for extracting necessary statistical information from the collected data;
[1245] means for pre-processing the extracted data;
[1246] means for analyzing the preprocessed data using a generative model to assess the authenticity of the preprocessed data;
[1247] means for calculating a credibility score based on the analysis results;
[1248] means for providing the authenticity scores and data to a user;
[1249] means for performing facial recognition and voice analysis to recognize the user's emotions;
[1250] means for adjusting the display content based on the user's emotions;
[1251] A system including:
[1252] (Claim 2)
[1253] 10. The system of claim 1, wherein the assessment of the credibility score includes the reliability of the source, past revision history, and data consistency.
[1254] (Claim 3)
[1255] 10. The system of claim 1, further comprising a user interface for visually displaying the credibility score.
[1256] "Application example 2 when combining emotion engines"
[1257] (Claim 1)
[1258] data collection means;
[1259] A means for extracting necessary statistical data from the collected data;
[1260] means for pre-processing the extracted data;
[1261] means for analyzing the preprocessed data using a generative model to assess the authenticity of the preprocessed data;
[1262] means for calculating a credibility score based on the analysis results;
[1263] means for providing the authenticity scores and data to a user;
[1264] emotion analysis means for recognizing a user's emotion and adjusting the display content;
[1265] A system including:
[1266] (Claim 2)
[1267] 2. The system of claim 1, wherein the assessment of the credibility score includes the reliability of the source, past revision history, and data consistency.
[1268] (Claim 3)
[1269] 10. The system of claim 1, further comprising a user interface for visually displaying the credibility score, and further comprising optimizing advertising content based on user sentiment. [Explanation of symbols]
[1270] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>
Claims
1. data collection means; A means for extracting necessary statistical data from the collected data; means for pre-processing the extracted data; means for analyzing the preprocessed data using a generative model to assess the authenticity of the preprocessed data; means for calculating a credibility score based on the analysis results; means for providing the authenticity scores and data to a user; A system including:
2. The system of claim 1 , wherein the assessment of the credibility score includes the reliability of the source, past revision history, and data consistency.
3. The system of claim 1 , further comprising a user interface for visually displaying the credibility score.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A