System
The system addresses the challenge of fake news by analyzing media data for authenticity through metadata and content-based methods, ensuring rapid and accurate reliability evaluations.
Patent Information
- Application Number
- JP2024119024
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-07-24
- Publication Date
- 2026-02-05
AI Technical Summary
Current systems struggle to comprehensively evaluate the authenticity and reliability of manipulated photos and videos, leading to the rapid spread of fake news and social unrest.
A system that uploads media data from user terminals to a server for metadata, forensic, and content-based analysis, integrating results to evaluate reliability and notify users quickly.
Enables comprehensive evaluation of media authenticity, preventing the spread of fake news by providing quick and accurate reliability assessments.
Smart Images

Figure 2026017963000001_ABST
Abstract
Description
[Technical Field]
[0001] The technology of the present disclosure relates to a system. [Background technology]
[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]
[0004] Advances in digital technology have made it easy to manipulate photos and videos, resulting in the rapid spread of fake news. This raises concerns that this could lead to social unrest and increased distrust. Against this backdrop, there is a growing need for systems that can automatically determine the authenticity of manipulated photos and videos and evaluate their reliability. However, many current systems only support a limited range of analysis, making it difficult to comprehensively evaluate their reliability. [Means for solving the problem]
[0005] To solve the above problems, the present invention provides the following means: a system including means for uploading media data from a user terminal to a server, means for the server to analyze metadata of the uploaded media data, means for the server to perform forensic analysis of the uploaded media data, means for the server to perform content-based analysis of the uploaded media data, means for the server to integrate the results of each analysis and evaluate the reliability, means for the server to send the evaluation result to a user terminal, and means for the user terminal to display the evaluation result. This makes it possible to comprehensively evaluate the authenticity of photos and videos and quickly notify the user of the results, thereby effectively preventing the spread of fake news.
[0006] "Media data" refers to digital information files such as images and videos.
[0007] "User terminal" refers to a device such as a computer, smartphone, or tablet used by an individual.
[0008] A "server" refers to a computer system that provides services and data to other devices over a network.
[0009] "Upload" refers to the act of sending data from a user terminal to a server.
[0010] "Metadata" is additional information that accompanies media data such as images and videos, and includes, for example, the date and time of the photo, the location where the photo was taken, and the model of the camera.
[0011] "Forensic analysis" refers to a technical method for detecting traces of tampering or manipulation of media data.
[0012] "Content-based analysis" refers to an analysis method that evaluates media data based on its content. Specifically, it includes object recognition and scene analysis.
[0013] "Reliability assessment" refers to comprehensively judging the authenticity of media data and expressing the reliability as a score or category.
[0014] "Notification" refers to the action of transmitting the analysis results to the user terminal.
[0015] "Display" refers to visually showing the analysis results on the screen of the user terminal. [Brief explanation of the drawings]
[0016] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION
[0017] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.
[0018] First, the terms used in the following description will be explained.
[0019] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).
[0020] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.
[0021] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.
[0022] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.
[0023] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."
[0024] [First embodiment]
[0025] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0026] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0027] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0028] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.
[0029] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0030] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0031] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.
[0032] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0033] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0034] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0035] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0036] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0037] This invention relates to a system for assessing the reliability of digital media data to prevent the spread of fake news. This system is implemented by uploading images and videos from user devices to a server, which then analyzes them. Specifically, the system performs metadata analysis, forensic analysis, and content-based analysis, and integrates the results to assess reliability.
[0038] Program processing
[0039] User uploads data from device
[0040] Users can select photos and videos suspected of being fake news and upload them to the system using their own devices. A simple interface is provided on the device to support data uploading.
[0041] The server receives the data
[0042] The server receives image and video data sent by users, which is transmitted via a secure communication protocol and stored and analyzed on the server side.
[0043] Metadata Analysis
[0044] The server analyzes the metadata of the received data, which includes the date and time of the photo, the location (GPS information), the model of the camera used, and the editing software used.
[0045] Forensic Analysis
[0046] The server performs forensic analysis to detect any tampering or manipulation of the media data. The following methods are used in this step:
[0047] Pixel compatibility check: Detects discontinuities in pixel color or brightness.
[0048] Edge Detection: Detects image edges and identifies unnatural edges.
[0049] Block Inspection: Analyzes compressed formats for block artifacts.
[0050] Content-Based Analysis
[0051] The server uses AI to analyze the media data based on its content, using the following methods:
[0052] Object Recognition: Recognize objects in images and videos and match them against a database.
[0053] Context analysis: Checks the consistency of each scene in the video and detects unnatural transitions.
[0054] Integrated evaluation and notification of results
[0055] The server integrates the results of each analysis and evaluates the reliability of the media data. Based on the reliability score, the data is classified into categories such as "high reliability," "partially suspicious," and "low reliability." The server then sends the evaluation results to the user's device, where the user can confirm the results.
[0056] Specific examples
[0057] 1. A user uploads a suspicious photo they found on a news site.
[0058] The user selects a photo on their device and uploads it to the system.
[0059] The server receives the photos and performs metadata and forensic analysis.
[0060] Content-based analysis reveals that the person in the photo does not match any photos in a reliable database.
[0061] The server evaluates the reliability as "low" and sends the result to the user's device.
[0062] The user reviews the results and makes a decision about the authenticity of the photo.
[0063] 2. When a user uploads a video shared on social media
[0064] The user selects a video on their device and uploads it to the system.
[0065] The server receives the video and performs metadata and forensic analysis on each frame.
[0066] Using object recognition technology, it is discovered that scenes in the video do not match other public records.
[0067] The server evaluates it as "partially suspicious" and sends the result to the user's device.
[0068] Users review the results and make a decision about the video's authenticity.
[0069] This allows the system to help users quickly and accurately determine the authenticity of photos and videos, preventing the spread of fake news.
[0070] The processing flow will be explained below.
[0071] Step 1:
[0072] The user selects an image or video.
[0073] The user selects images or videos suspected of being fake news from their device using the device's file selection function.
[0074] Step 2:
[0075] A user uploads data into the system.
[0076] The user presses the upload button to send the selected media data to the server. The data is sent to the server using a secure communication protocol (e.g., HTTPS).
[0077] Step 3:
[0078] The server receives the data.
[0079] The server receives the media data sent by the user, stores it in a database, and checks the file format and basic attributes.
[0080] Step 4:
[0081] The server pre-processes the data.
[0082] The server performs pre-processing on the received data, specifically resizing, color space conversion (e.g., RGB to grayscale), and noise reduction.
[0083] Step 5:
[0084] The server parses the metadata.
[0085] The server extracts and analyzes the metadata of the media data, checking the shooting date and time, shooting location (GPS information), camera model, editing software usage history, etc.
[0086] Step 6:
[0087] The server performs the forensic analysis.
[0088] The server performs forensic analysis to detect any evidence of tampering or manipulation of the media data, including pixel compatibility checks, edge detection, and block inspection.
[0089] Step 7:
[0090] The server performs content-based analysis.
[0091] The server uses AI technology to perform content-based analysis, object recognition to identify objects in images and videos and match them with a trusted database, and context analysis to check the consistency of each scene in the video.
[0092] Step 8:
[0093] The server integrates the results of each analysis and evaluates their reliability.
[0094] The server evaluates all analysis results comprehensively and assigns a reliability score, which is used to classify the media data as "highly reliable," "partially suspicious," or "low reliability."
[0095] Step 9:
[0096] The server transmits the evaluation results to the user terminal.
[0097] The server then transmits the evaluation results to the user terminal, again using a secure communication protocol.
[0098] Step 10:
[0099] The terminal displays the evaluation results.
[0100] The user terminal displays the received evaluation results, including the reliability score, analysis details, and the rationale for the results.
[0101] Step 11:
[0102] The user checks the results and makes a decision.
[0103] Users can view the results displayed on their device, make a judgment about the authenticity of the photo or video, and take action based on this information to prevent the spread of fake news.
[0104] Example 1
[0105] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0106] In modern society, the extremely rapid spread of information through digital media has created the problem of fake news spreading easily and making it difficult to distinguish it from real information. Furthermore, unreliable media data can lead to incorrect decision-making and cause social unrest. To solve these problems, a system is needed that can quickly and accurately evaluate the reliability of uploaded media data and provide accurate information to users.
[0107] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0108] In this invention, the server includes means for uploading media data from a user terminal, means for analyzing metadata of the uploaded media data, means for performing forensic analysis of the uploaded media data, means for performing content-based analysis of the uploaded media data, means for integrating the results of each analysis to evaluate the reliability, means for transmitting the evaluation result to the user terminal, and means for displaying the evaluation result, which makes it possible to prevent the spread of fake news and quickly and accurately determine the authenticity of media data.
[0109] A "user terminal" is an electronic device such as a computer, smartphone, or tablet that is operated by a user.
[0110] A "server" is a computer system set up to provide a particular service over a network.
[0111] "Media data" is data stored in digital format, such as images, video, and audio.
[0112] "Metadata" refers to data recorded in addition to the information on the media data itself, and includes information such as the date and time of shooting, the location, and the model of the camera.
[0113] "Forensic analysis" is a general term for technical analytical methods used to detect tampering or manipulation of digital data.
[0114] "Content-based analysis" is an analysis method based on the content of media data, and includes object recognition and context analysis.
[0115] "Authenticity assessment" is a process for assessing the authenticity and reliability of media data, and is carried out by integrating the results of various analyses.
[0116] "Resizing" refers to the process of changing the size of an image or video.
[0117] "Color space conversion" is the process of converting a color representation to a different color space.
[0118] "Noise reduction" is a process for removing unnecessary noise from media data.
[0119] "Pixel compatibility testing" is an analysis method that detects discontinuities in color and brightness at the pixel level.
[0120] "Edge detection" is an analysis technique that detects edges within an image and identifies unnatural edges.
[0121] "Block inspection" is a technique that analyzes block artifacts in compression formats to detect signs of tampering.
[0122] "Object recognition" is a technology that identifies specific objects or people in images or videos.
[0123] "Context analysis" is an analytical method that evaluates the continuity and consistency of each scene and detects unnatural transitions and inconsistencies.
[0124] The present invention relates to a system for evaluating the reliability of digital media data and preventing the spread of fake news. This system is implemented by uploading images and videos from a user's device to a server, which then analyzes the data. The following describes how this system is specifically implemented.
[0125] Configuration and Hardware
[0126] The system of the present invention consists of the following main components:
[0127] 1. User terminal: Electronic devices such as smartphones, tablets, and PCs. They provide a GUI (graphical user interface) to facilitate data uploading.
[0128] 2. Server: A high-performance computer system with the computing power and storage capacity to analyze the received media data. The server software uses programming languages such as Python or Java, and various libraries (e.g., OpenCV, TensorFlow, etc.) for data analysis.
[0129] Software and Data Processing
[0130] The system software performs the following main tasks:
[0131] 1. Upload your data:
[0132] A user uses a smartphone application to select suspicious media data and upload it to the server by selecting the "Upload suspicious media data" button, selecting photos or videos in the file selection dialog, and pressing the "Upload" button.
[0133] 2. Metadata analysis:
[0134] The server analyzes the metadata of the received media data, including the date and time of the photo, the location (GPS information), the model of the camera used, the editing software used, etc. The server extracts the Exif data of the image or video and checks whether it matches a specific event.
[0135] 3. Forensic Analysis:
[0136] The server performs forensic analysis to detect any manipulation or alteration of media data, primarily using pixel compatibility checks, edge detection, and block checks. It analyzes the color and brightness of each pixel to detect discontinuities. It detects image edges and checks for unnatural edges. It analyzes block artifacts of compression formats within images to find signs of manipulation.
[0137] 4. Content-based analysis:
[0138] The server uses AI technology to perform object recognition and context analysis, which evaluates the consistency of objects and scenes within images and videos. It uses AI models (e.g., deep learning models) to evaluate the continuity and context of each scene within a video and detect unnatural transitions and inconsistencies.
[0139] Analysis results and reliability evaluation
[0140] The server integrates the results of each analysis and evaluates the reliability of the media data. It calculates a reliability score and classifies the data into categories of "high reliability," "partially suspicious," or "low reliability." The evaluation results are sent to the user's device, where the user can confirm the results.
[0141] Examples and prompts
[0142] 1. Example: A photo from a news site:
[0143] A user uploads a suspicious photo they found on a news site. The server performs metadata and forensic analysis, and evaluates the trustworthiness of the photo using content-based analysis. The resulting rating is "low trustworthiness," and a notification is sent to the user's device.
[0144] 2. Example: Social media sharing video:
[0145] A user uploads a suspicious video shared on a social networking site. The server analyzes each frame, performing object recognition and context analysis. The video is ultimately evaluated as "partially suspicious" and a notification is sent to the user's device.
[0146] Prompt Sentence Examples
[0147] "Please rate the reliability of this image."
[0148] "Detect unnatural elements in this video."
[0149] In this way, the system of the present invention can quickly and accurately assess the authenticity of digital media data, providing a reliable means to prevent the spread of fake news.
[0150] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0151] Step 1:
[0152] User uploads data from device
[0153] Input: User-selected image or video files
[0154] Output: Sending data to the server
[0155] Specific behavior:
[0156] The user launches an application on the smartphone.
[0157] Select the "Upload suspicious media data" button on the application interface.
[0158] Select a photo or video in the file selection dialog and press the "Upload" button.
[0159] The selected file is sent to the server using the HTTPS protocol.
[0160] Step 2:
[0161] The server receives the data
[0162] Input: Image and video files sent from the user's device
[0163] Output: Received data stored on the server
[0164] Specific behavior:
[0165] The server listens for HTTPS requests.
[0166] The data sent by the user arrives at the server.
[0167] The server stores the received data in a secure directory.
[0168] Step 3:
[0169] The server performs metadata analysis
[0170] Input: Image and video files stored on the server
[0171] Output: Parsed metadata information
[0172] Specific behavior:
[0173] The server loads the image or video file.
[0174] Parse metadata information (e.g., Exif data).
[0175] The system extracts the date and time of the photo, GPS information, the model of the camera used, and the editing software usage history, and compares them with specific events and past data.
[0176] Step 4:
[0177] The server performs forensic analysis
[0178] Input: Image and video files stored on the server
[0179] Output: Forensic analysis results
[0180] Specific behavior:
[0181] A pixel compatibility check is performed to analyze the color and brightness of each pixel to detect discontinuities.
[0182] Apply edge detection to identify edges in the image and check if any unnatural edges are present.
[0183] Performs block inspection and analyzes block artifacts in compressed formats to find signs of tampering.
[0184] Step 5:
[0185] The server performs content-based analysis
[0186] Input: Image and video files stored on the server
[0187] Output: Results of content-based analysis
[0188] Specific behavior:
[0189] Use object recognition technology to recognize objects and people in images and videos.
[0190] Recognized objects and scenes are checked against a database to see if there is a match.
[0191] For video, it analyzes each frame to assess the consistency of each scene and detects unnatural transitions and inconsistencies.
[0192] Step 6:
[0193] The server performs an integrated evaluation and notifies the results.
[0194] Input: Results of metadata analysis, forensic analysis, and content-based analysis
[0195] Output: Reliability assessment and its results
[0196] Specific behavior:
[0197] The server consolidates the analysis results.
[0198] A reliability score is calculated and the data is classified into categories of "high reliability," "partially questionable," and "low reliability."
[0199] A message for transmitting the evaluation result to the user terminal is generated and transmitted using a secure communication protocol.
[0200] The user receives the results and displays them in the application.
[0201] Through these steps, the system prevents the spread of fake news and enables users to quickly and accurately determine the authenticity of digital media data.
[0202] (Application example 1)
[0203] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0204] In recent years, the spread of fake news on social networking services and news distribution sites has become a social problem. In particular, there have been many cases where media data such as images and videos have been tampered with, resulting in the dissemination of information that differs from the truth. Existing technologies require users to manually verify the authenticity of data, which has limitations in reliability and efficiency. Therefore, the present invention aims to provide a system that allows users to easily evaluate the reliability of images and videos and prevent the spread of fake news.
[0205] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0206] In this invention, the server includes means for uploading media data from a user terminal to the server, means for the server to analyze metadata of the uploaded media data, means for the server to perform forensic analysis of the uploaded media data, means for the server to perform content-based analysis of the uploaded media data, means for the server to integrate the results of each analysis and evaluate the reliability, means for the server to transmit the evaluation result to the user terminal, means for the user terminal to display the evaluation result, and means for the user to confirm the reliability evaluation result of the media data and prevent the spread of fake news before sharing the data. This enables users to quickly evaluate the reliability of data before sharing it on social networking services or news distribution sites, and effectively prevent the spread of fake news.
[0207] "Media data" is information stored in digital form, such as images and videos.
[0208] A "user terminal" is an electronic device that can be directly operated by a user, such as a smartphone or computer.
[0209] A "server" is a computer system that processes and manages data over a network.
[0210] "Metadata" is additional information that accompanies digital media data, including the date and time the image was taken, the location where it was taken, the model of the camera used, and so on.
[0211] "Forensic analysis" refers to scientific and technical methods for detecting tampering or manipulation of media data.
[0212] "Content-based analysis" is a method of analyzing digital media data based on its content, and includes object recognition and context analysis.
[0213] "Authenticity assessment" is the process of determining the authenticity and reliability of digital media data based on analysis results.
[0214] "Sending to the user terminal" refers to the act of sending the evaluation results from the server to a terminal that the user can directly operate.
[0215] "Displaying the evaluation results" means visually displaying the results of the analysis on the user terminal.
[0216] "Preventing the spread of fake news before data is shared" means preventing the spread of misinformation by assessing the reliability of media data before it is made public on social networking services, news distribution sites, etc.
[0217] The present invention provides a system for preventing the spread of fake news on social networking services and news distribution sites, allowing users to quickly and accurately evaluate the reliability of media data such as images and videos.
[0218] System configuration
[0219] The system consists of a user device and a server. User devices include smartphones, computers, and tablets. The server has data processing capabilities and storage, and receives, analyzes, and notifies users of the results.
[0220] Hardware and Software
[0221] An application runs on the user's device to select images and videos and upload them to the server. The specific development environment uses the cross-platform React Native.
[0222] The server runs a web server using Flask, which receives the uploaded media data, analyzes it, and performs authenticity assessment using AI models (using TensorFlow or PyTorch, for example) and forensic analysis tools.
[0223] Media Data Analysis
[0224] The server performs the following analysis on the uploaded media data:
[0225] 1. Metadata analysis: Analyzes the shooting date and time, shooting location (GPS information), camera model used, editing software usage history, etc.
[0226] 2. Forensic analysis: Detecting tampering or manipulation of media data, including pixel compatibility checks, edge detection, and discontinuous block checks.
[0227] 3. Content-based analysis: Using AI models to analyze the content of images and videos, specifically through object recognition and context analysis, and matching the results with a database.
[0228] Integrated evaluation and notification of results
[0229] The server integrates the results of each analysis and evaluates the reliability of the media data. The results are expressed as a reliability score and classified into categories such as "high reliability," "partially questionable," and "low reliability." The evaluation results are sent to the user's device in real time and can be viewed through the application.
[0230] Specific use cases
[0231] Assume that users verify the authenticity of photos and videos through the application before sharing them on social media. Below is a concrete example and an example of a prompt sentence to input to the generative AI model.
[0232] Specific examples
[0233] When a user uploads a suspicious photo they find on a news site:
[0234] The user selects a photo on their device and uploads it to the system.
[0235] The server receives the photos and performs metadata and forensic analysis.
[0236] Content-based analysis matches the content of an image against a trusted database.
[0237] The server evaluates it as "partially suspicious" and sends the result to the user's device.
[0238] The user reviews the results and makes a decision about the authenticity of the photo.
[0239] Example prompt for a generative AI model:
[0240] "Rate the trustworthiness of this image by analyzing metadata such as objects shown, location and time of capture, and camera model used, and then using pixel and edge detection and unnatural blocks to provide a confidence score based on its content."
[0241] This system allows users to easily assess the reliability of data before sharing it on social networking services or news distribution sites, effectively preventing the spread of fake news.
[0242] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0243] Step 1:
[0244] The user selects media data on the terminal.
[0245] Input: The user selects a suspicious image or video and presses the upload button.
[0246] Behavior: The path of the file selected on the device is obtained and an upload operation is triggered.
[0247] Output: The file path of the media data is ready to be sent by the application to the server.
[0248] Step 2:
[0249] The terminal uploads the media data to the server.
[0250] Input: The file path of the selected media data.
[0251] Operation: When a file is selected, the terminal begins the process of sending the file to the server as form data.
[0252] Output: A binary stream of media data sent to the server.
[0253] Step 3:
[0254] The server receives and stores the media data.
[0255] Input: A binary stream of uploaded media data.
[0256] Behavior: The server saves the received data to the specified directory, using Flask's upload function to save the file name appropriately.
[0257] Output: The file path of the saved media data.
[0258] Step 4:
[0259] The server performs the metadata analysis.
[0260] Input: The file path of the saved media data.
[0261] How it works: The server uses Python libraries (e.g., PIL and ExifRead) to extract metadata from the media data (such as the date and time of the image capture, GPS information, and camera model).
[0262] Output: A set of metadata information (in dictionary format).
[0263] Step 5:
[0264] The server performs the forensic analysis.
[0265] Input: The file path of the media data.
[0266] What it does: Performs pixel compatibility and edge detection on images, and checks for discontinuous blocks. Analysis is performed using OpenCV and Scikit-Image.
[0267] Output: Forensic analysis results (whether or not the data has been tampered with or altered).
[0268] Step 6:
[0269] The server performs content-based analysis.
[0270] Input: The file path of the media data.
[0271] How it works: It uses AI models (e.g. TensorFlow or PyTorch) to perform object recognition and context analysis in images and videos, and matches them with authoritative information in a database.
[0272] Output: Content-based analysis results (object recognition results and context consistency information).
[0273] Step 7:
[0274] The server integrates the results of each analysis and evaluates their reliability.
[0275] Input: Metadata analysis results, forensic analysis results, content-based analysis results.
[0276] How it works: Each analysis result is evaluated comprehensively and a reliability score is calculated. The data is classified using a reliability evaluation algorithm.
[0277] Output: Confidence score (categorised as high, medium, low etc.).
[0278] Step 8:
[0279] The server transmits the evaluation results to the user terminal.
[0280] Input: Confidence score.
[0281] Operation: The evaluation results are sent to the user's terminal in real time, and the data is formatted so that the user can check it immediately.
[0282] Output: The trust evaluation result sent to the user terminal.
[0283] Step 9:
[0284] The user terminal displays the evaluation results.
[0285] Input: The trust evaluation result received from the server.
[0286] Operation: The application displays the received evaluation results in the user interface so that the user can check them. The evaluation results also include detailed analysis information.
[0287] Output: The reliability assessment results displayed in the user interface.
[0288] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.
[0289] This invention combines a system that evaluates the reliability of digital media data and prevents the spread of fake news with an emotion engine that recognizes user emotions. This system is realized by uploading images and videos from users' devices to a server, which then analyzes them. Furthermore, the emotion engine is used to integrate the analysis results with user emotions, allowing for more effective evaluation.
[0290] Program processing
[0291] User uploads data from device
[0292] Users can select photos and videos suspected of being fake news and upload them to the system using their own devices. A simple interface is provided on the device to support data uploading.
[0293] The server receives the data
[0294] The server receives image and video data sent by users, which is transmitted via a secure communication protocol and stored and analyzed on the server side.
[0295] Metadata Analysis
[0296] The server analyzes the metadata of the received data, which includes the date and time of the photo, the location (GPS information), the model of the camera used, and the editing software used.
[0297] Forensic Analysis
[0298] The server performs forensic analysis to detect any tampering or manipulation of the media data. The following methods are used in this step:
[0299] Pixel compatibility check: Detects discontinuities in pixel color or brightness.
[0300] Edge Detection: Detects image edges and identifies unnatural edges.
[0301] Block Inspection: Analyzes compressed formats for block artifacts.
[0302] Content-Based Analysis
[0303] The server uses AI to analyze the media data based on its content, using the following methods:
[0304] Object Recognition: Recognize objects in images and videos and match them against a database.
[0305] Context analysis: Checks the consistency of each scene in the video and detects unnatural transitions.
[0306] Emotion recognition by emotion engine
[0307] The server runs an emotion engine that recognizes the user's emotions. Based on the data sent from the user's device, the emotion engine performs the following:
[0308] Facial expression analysis: Capture the user's facial expressions with a camera and recognize the emotions they express.
[0309] Voice analysis: Analyze the user's voice and estimate their emotions.
[0310] Integrated evaluation and notification of results
[0311] The server integrates the results of each analysis and the emotion engine to evaluate the reliability of the media data. This evaluation includes the results of metadata analysis, forensic analysis, content-based analysis, and the user's emotional state. Based on the reliability score, the data is classified as "highly reliable," "partially suspicious," or "low reliability." The server then sends the evaluation results to the user's device, where the user can confirm the results.
[0312] Specific examples
[0313] 1. A user uploads a suspicious photo they found on a news site.
[0314] The user selects a photo on their device and uploads it to the system.
[0315] The server receives the photos and performs metadata and forensic analysis.
[0316] Content-based analysis reveals that the person in the photo does not match any photos in a reliable database.
[0317] The emotion engine analyzes the user's facial expressions and recognizes the user's emotions when viewing photos.
[0318] The server integrates the analysis results with the emotion recognition, evaluates it as "low reliability," and sends the result to the user's device.
[0319] The user reviews the results and makes a decision about the authenticity of the photo.
[0320] 2. When a user uploads a video shared on social media
[0321] The user selects a video on their device and uploads it to the system.
[0322] The server receives the video and performs metadata and forensic analysis on each frame.
[0323] Using object recognition technology, it is discovered that scenes in the video do not match other public records.
[0324] The emotion engine analyzes the user's voice and recognizes their emotions when watching videos.
[0325] The server integrates the analysis results with emotion recognition, evaluates the image as "partially suspicious," and sends the result to the user's device.
[0326] Users review the results and make a decision about the video's authenticity.
[0327] As described above, this system helps users quickly and accurately judge the authenticity of photos and videos, preventing the spread of fake news. Furthermore, by taking user emotions into account, it achieves more reliable evaluations.
[0328] The processing flow will be explained below.
[0329] Step 1:
[0330] The user selects an image or video.
[0331] The user selects images or videos suspected of being fake news from their device using the device's file selection function.
[0332] Step 2:
[0333] A user uploads data into the system.
[0334] The user presses the upload button to send the selected media data to the server. The data is sent to the server using a secure communication protocol (e.g., HTTPS).
[0335] Step 3:
[0336] The server receives the data.
[0337] The server receives the media data sent by the user, stores it in a database, and checks the file format and basic attributes.
[0338] Step 4:
[0339] The server pre-processes the data.
[0340] The server performs pre-processing on the received data, specifically resizing, color space conversion (e.g., RGB to grayscale), and noise reduction.
[0341] Step 5:
[0342] The server parses the metadata.
[0343] The server extracts and analyzes the metadata of the media data, checking the shooting date and time, shooting location (GPS information), camera model, editing software usage history, etc.
[0344] Step 6:
[0345] The server performs the forensic analysis.
[0346] The server performs forensic analysis to detect any evidence of tampering or manipulation of the media data, including pixel compatibility checks, edge detection, and block inspection.
[0347] Step 7:
[0348] The server performs content-based analysis.
[0349] The server uses AI technology to perform content-based analysis, object recognition to identify objects in images and videos and match them with a trusted database, and context analysis to check the consistency of each scene in the video.
[0350] Step 8:
[0351] The server runs the emotion engine.
[0352] The server recognizes the user's emotions based on the data sent from the user's device. The emotion engine estimates emotions through facial expression and voice analysis. Specifically, the server captures the user's facial expressions with a camera and records their voice with a microphone, and analyzes these to recognize emotions.
[0353] Step 9:
[0354] The server integrates the results of each analysis and evaluates their reliability.
[0355] The server comprehensively evaluates all analysis results and assigns a reliability score. Based on the reliability score, the media data is classified as "highly reliable," "partially suspicious," or "low reliability." Furthermore, the server takes into account the user's emotional state and reflects it in the evaluation.
[0356] Step 10:
[0357] The server transmits the evaluation results to the user terminal.
[0358] The server then transmits the evaluation results and sentiment analysis results to the user terminal, using a secure communication protocol.
[0359] Step 11:
[0360] The terminal displays the evaluation results.
[0361] The user terminal displays the received evaluation results, including the reliability score, analysis details, the rationale for the results, and comments based on the user's emotional state.
[0362] Step 12:
[0363] The user checks the results and makes a decision.
[0364] Users can view the results displayed on their device, make a judgment about the authenticity of the photo or video, and take action based on this information to prevent the spread of fake news.
[0365] Example 2
[0366] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0367] Fake news can have a significant impact on society, so there is a need for effective methods to detect it and prevent its spread. However, conventional systems only analyze media data and do not consider user sentiment, resulting in insufficient assessment of its reliability. In addition, it is necessary to improve the accuracy of fake news detection by integrating multiple analysis methods.
[0368] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0369] In this invention, the server includes means for analyzing metadata of media data, means for performing forensic analysis of the media data, means for performing content-based analysis of the media data, means for recognizing user emotions in the media data, and means for evaluating reliability by integrating the results of each analysis and the emotion recognition result, thereby making it possible to more accurately determine the authenticity of media data and effectively prevent the spread of fake news.
[0370] "Metadata" refers to background information that accompanies media data, and includes the date and time of shooting, the location where the footage was taken, the model of camera used, and the history of editing software used.
[0371] "Media data" refers to media such as images and videos stored in digital format.
[0372] "Forensic analysis" is a scientific analysis method for detecting tampering or manipulation of media data, and includes pixel compatibility testing, edge detection, block testing, etc.
[0373] "Content-based analysis" is a method for analyzing the authenticity of media data based on its content, and includes object recognition and context analysis.
[0374] "Emotion recognition" is a technology that detects and analyzes the emotions expressed by users in response to media data, and involves facial expression analysis and voice analysis.
[0375] "Trustworthiness assessment" is the process of integrating various analysis results and emotion recognition results to determine the authenticity and reliability of media data.
[0376] This invention is a system that evaluates the reliability of digital media data and prevents the spread of fake news. This system is realized by uploading images and videos from users' devices to a server, which then analyzes them. Furthermore, an emotion engine is used to integrate the analysis results with the user's emotions, allowing for more effective evaluation.
[0377] First, the user selects images or videos suspected of being fake news using the device. The selected media data is then uploaded to the server via a secure communication protocol (e.g., HTTPS). The device is provided with a user-friendly interface to assist in uploading the data.
[0378] The server receives the uploaded media data and analyzes its metadata, which extracts information such as the date and time of the photo shoot, the location, the model of the camera used, and the editing software used. Based on this information, the authenticity of the media data is evaluated.
[0379] The server then performs forensic analysis, using techniques such as pixel compatibility testing, edge detection, and block detection. Pixel compatibility testing detects discontinuities in pixel color or brightness, while edge detection finds image edges and identifies unnatural transitions and boundaries. Block detection analyzes block artifacts in compression formats to identify signs of image or video tampering.
[0380] Additionally, the server performs content-based analysis, which uses generative AI models (such as TensorFlow or PyTorch) to analyze media data based on its content. Specifically, it uses object recognition technology to recognize objects in images and videos and compare them with public databases. Context analysis checks the consistency of each scene in the video and detects unnatural scene transitions.
[0381] The server also operates an emotion engine to recognize the user's emotions. The emotion engine performs the following operations based on the data sent from the user's device: facial expression analysis recognizes emotions from the user's facial expression data captured by a camera, and voice analysis estimates emotions from the user's recorded voice data.
[0382] Finally, the server integrates the results of each analysis and the emotion engine to evaluate the reliability of the media data. This evaluation includes the results of metadata analysis, forensic analysis, content-based analysis, and the user's emotional state. Based on the reliability score, the media data is classified as "highly reliable," "partially suspicious," or "low reliability." The evaluation results are then sent from the server to the user's device, where the user can confirm the results.
[0383] Specific examples
[0384] 1. Uploading a suspicious photo you found on a news site
[0385] The user selects a photo on their device and uploads it to the system.
[0386] The server receives the photos and performs metadata and forensic analysis.
[0387] Content-based analysis reveals that the person in the photo does not match any photos in a reliable database.
[0388] The emotion engine analyzes the user's facial expressions and recognizes the user's emotions when viewing photos.
[0389] The server integrates the analysis results with the emotion recognition, evaluates it as "low reliability," and sends the result to the user's device.
[0390] The user reviews the results and makes a decision about the authenticity of the photo.
[0391] 2. When uploading a video shared on social media
[0392] The user selects a video on their device and uploads it to the system.
[0393] The server receives the video and performs metadata and forensic analysis on each frame.
[0394] Using object recognition technology, it is discovered that scenes in the video do not match other public records.
[0395] The emotion engine analyzes the user's voice and recognizes their emotions when watching videos.
[0396] The server integrates the analysis results with emotion recognition, evaluates the image as "partially suspicious," and sends the result to the user's device.
[0397] Users review the results and make a decision about the video's authenticity.
[0398] Prompt Sentence Examples
[0399] "Is this photo real? Rate its authenticity."
[0400] Please confirm whether the scene in this video actually happened.
[0401] "Please analyze this news image to see if it has been tampered with."
[0402] Through the above-described embodiment, the system helps users quickly and accurately determine the authenticity of photos and videos, preventing the spread of fake news. Furthermore, by taking into account the user's emotions, the system achieves more reliable evaluation.
[0403] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0404] Step 1:
[0405] User selects and uploads suspicious media
[0406] Input: Image or video data selected by the user on their device.
[0407] How it works: The user selects media suspected of being fake news using a web browser or a dedicated app on their device. The selected media file is then uploaded to a server via a secure communication protocol (e.g., HTTPS).
[0408] Output: Media data sent to the server over a secure communication channel.
[0409] Step 2:
[0410] The server receives and stores the media data.
[0411] Input: The media data uploaded in step 1.
[0412] Specific operation: The server receives image or video data sent by the user using a secure communication protocol, calculates a checksum to check the integrity and completeness of the received data, and stores it in a temporary data storage area.
[0413] Output: Media data stored on the server.
[0414] Step 3:
[0415] The server analyzes the metadata
[0416] Input: Media data stored on the server.
[0417] How it works: The server extracts and analyzes metadata from stored media files, including EXIF information such as the date and time the image was taken, the location where it was taken (GPS information), the model of the camera used, and the editing software used.
[0418] Output: Metadata analysis results.
[0419] Step 4:
[0420] Server performs forensic analysis
[0421] Input: Media data stored on the server.
[0422] What it does: The server performs forensic analysis: pixel compatibility checks to detect discontinuities in pixel color or brightness, edge detection detects image edges and identifies unnatural transitions and boundaries, and block inspection analyzes block artifacts specific to compression formats to detect signs of tampering.
[0423] Output: Forensic analysis results.
[0424] Step 5:
[0425] The server performs content-based analysis
[0426] Input: Media data stored on the server.
[0427] How it works: The server uses generative AI models (e.g., TensorFlow, PyTorch) to recognize objects in images and videos and match them against a trusted database. It also performs context analysis to check the consistency of each scene in the video and detect unnatural scene transitions.
[0428] Output: Content-based analysis results.
[0429] Step 6:
[0430] The server runs the emotion engine and performs emotion recognition.
[0431] Input: Media data stored on the server and emotion data from the user's device.
[0432] Specific operation: The server acquires data from the camera and microphone installed on the user's device. For facial expression analysis, the camera captures the user's facial expression data, and the AI model recognizes the emotion. For voice analysis, the server records the user's voice data and estimates the emotion.
[0433] Output: Emotion recognition result.
[0434] Step 7:
[0435] The server integrates the analysis results and emotion recognition to evaluate reliability
[0436] Input: Metadata analysis results, forensic analysis results, content-based analysis results, and emotion recognition results.
[0437] How it works: The server combines all the above analysis results and calculates a final reliability score, classifying the media data as "highly reliable," "partially questionable," or "low reliability."
[0438] Output: The estimated reliability score and its detailed analysis results.
[0439] Step 8:
[0440] The server sends the evaluation results to the user's device.
[0441] Input: The assessed reliability score and its detailed analysis results.
[0442] Specific operation: The server sends the trust evaluation result and its details to the user terminal using a secure communication protocol. An interface is provided that allows the user to check the result.
[0443] Output: Evaluation results are sent to the user's device, allowing the user to check the results.
[0444] The above specific processing steps enable the user to quickly and accurately judge the authenticity of the media data.
[0445] (Application example 2)
[0446] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0447] In recent years, the spread of fake news on social media and online platforms has become a social problem. Furthermore, users themselves often lack the ability to determine the veracity of information, resulting in a high risk of spreading false information. In response to this, there is a growing need for a system that can quickly and accurately evaluate the veracity of media data and provide feedback to users. Against this background, the present invention aims to provide a system for increasing the reliability of media data and suppressing the spread of fake news.
[0448] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for uploading media data from a user terminal to the server, means for the server analyzing metadata of the uploaded media data, means for the server performing forensic analysis of the uploaded media data, means for the server performing content-based analysis of the uploaded media data, means for the server acquiring user emotion data using an emotion engine that recognizes the user's emotion, means for the server integrating the analysis results and the user emotion data to evaluate reliability, means for the server transmitting the evaluation result to the user terminal, and means for the user terminal displaying the evaluation result. This makes it possible to quickly and accurately evaluate the authenticity of media data and provide the result to the user.
[0449] "Media data" is data in digital format that contains visual information such as images and videos.
[0450] A "user terminal" is a communication device, such as a smartphone, tablet, or PC, that is operated by a user to send and receive data.
[0451] A "server" is a computer system that manages and analyzes data on a network and supports communication with user terminals.
[0452] "Metadata" is supplementary information related to media data, and includes the date and time of shooting, the location of shooting, the model of the camera used, and editing history.
[0453] "Forensic analysis" is an analytical technique that technically detects whether media data has been tampered with or manipulated.
[0454] "Content-based analysis" is a technology that analyzes the content of images and videos themselves and evaluates object recognition and scene consistency.
[0455] An "emotion engine" is a technology or algorithm that recognizes a user's emotions by analyzing facial expressions and voice.
[0456] "Reliability evaluation" is the process of integrating each analysis result with the user's emotional data to determine the authenticity and credibility of the media data.
[0457] The "evaluation result" is information about the authenticity and reliability of the media data, generated based on the reliability evaluation.
[0458] "Preprocessing" refers to processing such as resizing, color space conversion, and noise removal to make the media data easier to analyze.
[0459] "Pixel compatibility testing" is a technology that detects data tampering by detecting discontinuities in the color or brightness of pixels within an image.
[0460] "Edge detection" is a technology that detects image tampering by detecting edges within an image and finding unnatural edges.
[0461] "Block inspection" is a technique that detects editing history and tampering of media data by analyzing block artifacts in compression formats.
[0462] The system embodying the present invention is implemented via a user terminal, a server, and communications between them. The overall configuration and operation procedure of the system will be described in detail below.
[0463] User terminal
[0464] The user device is a smartphone, tablet, or PC. The user uses the device to select suspicious media data (photos and videos) and upload them to the system. The device is provided with a simple interface that supports uploading media data, allowing users to upload data intuitively. The device is also equipped with a camera and microphone, which can capture the user's facial expressions and voice to collect emotional data.
[0465] server
[0466] The server is a computer system that receives and analyzes media data and emotion data sent from user terminals. The specific processing of the server for implementing the present invention is as follows.
[0467] Metadata Analysis
[0468] The server analyzes the metadata of the received media data, including the shooting date and time, shooting location (GPS information), camera model, and editing software usage history.
[0469] Forensic Analysis
[0470] The server performs forensic analysis to detect any tampering or manipulation of the media data, including pixel compatibility checks, edge detection, and block inspection.
[0471] Content-Based Analysis
[0472] The server uses AI to perform content-based analysis of the media data, including object recognition within images and videos, and contextual analysis to ensure consistency within video scenes.
[0473] Emotion Engine
[0474] The server recognizes the user's emotions using an emotion engine, which analyzes facial expression data and voice data sent by the user and estimates the user's emotions.
[0475] Integrated Evaluation
[0476] The server evaluates the reliability of the media data by integrating the results of each analysis with the user's emotional data. This evaluation includes the results of metadata analysis, forensic analysis, content-based analysis, and user emotion recognition. A reliability score is calculated and the media data is classified as "highly reliable," "partially suspicious," or "low reliability."
[0477] Result notification
[0478] The server transmits the evaluation results to the user terminal, which displays the received evaluation results so that the user can confirm them.
[0479] Specific examples
[0480] Take the example of a user uploading a suspicious video they found on a social networking site. The user selects the video on their device and uploads it to the system. The server receives the video and performs metadata and forensic analysis. At the same time, it uses object recognition technology to analyze scenes in the video and compare them with other public records. The server analyzes the user's voice and recognizes the emotions they express while watching the video. It then integrates all the analysis results, rates the video's authenticity as "partially suspicious," and sends the result to the user's device. The user then checks the results on their device and makes a decision about the video's authenticity.
[0481] Prompt Sentence Examples
[0482] "Upload your suspicious media and see the analysis results. Select a photo or video and press the upload button."
[0483] This will prevent the spread of fake news and allow users to make decisions based on accurate information.
[0484] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0485] Step 1:
[0486] A user selects suspicious media data (photos or videos) and uploads them to the system from their device. In this process, the user device provides a simple interface for acquiring media data and sending it to the server. The server receives the media data selected by the user as input and the sent media data as output.
[0487] Step 2:
[0488] The server receives media data sent from the user terminal. This process requires media data from the user terminal as input, which the server receives and stores. The stored media data is generated as output.
[0489] Step 3:
[0490] The server analyzes the metadata of the media data it receives. The stored media data is given as input, and the metadata analysis results are obtained as output. Specifically, it extracts information such as the shooting date and time, shooting location, camera model, and editing history.
[0491] Step 4:
[0492] The server performs forensic analysis of the media data. The media data obtained in the previous step is used as input, and the forensic analysis results are obtained as output. Specific analysis techniques used are pixel compatibility check, edge detection, and block inspection.
[0493] Step 5:
[0494] The server performs content-based analysis of the media data. It requires stored media data as input and produces the results of the content-based analysis as output. Specifically, it uses AI to recognize objects and evaluate scene consistency within images and videos.
[0495] Step 6:
[0496] The server uses an emotion engine to obtain the user's emotion data. Facial expression data and voice data sent from the user's device are used as input, and the emotion analysis results are obtained as output. The emotion engine recognizes the user's emotion based on facial expression and voice analysis.
[0497] Step 7:
[0498] The server integrates the results of each analysis and the user's sentiment data to assess the reliability of the media data. The results of metadata analysis, forensic analysis, content-based analysis, and sentiment analysis are used as input, and a reliability assessment score is generated as output. Based on this score, the media data is classified as "highly reliable," "partially suspicious," or "low reliability."
[0499] Step 8:
[0500] The server sends the trustworthiness evaluation results to the user terminal. The trustworthiness evaluation score is required as input, and the evaluation result is generated as output, which is received by the user terminal.
[0501] Step 9:
[0502] The user terminal displays the received evaluation results. The evaluation results sent from the server are used as input to generate the results that are displayed to the user as output. This allows the user to confirm the authenticity of the media data.
[0503] Through the above steps, the system is able to accurately and quickly evaluate the reliability of media data and notify the user of the results.
[0504] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0505] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0506] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.
[0507] [Second embodiment]
[0508] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.
[0509] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0510] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0511] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.
[0512] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0513] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0514] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0515] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0516] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0517] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0518] In the smart glasses 214, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0519] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."
[0520] This invention relates to a system for assessing the reliability of digital media data to prevent the spread of fake news. This system is implemented by uploading images and videos from user devices to a server, which then analyzes them. Specifically, the system performs metadata analysis, forensic analysis, and content-based analysis, and integrates the results to assess reliability.
[0521] Program processing
[0522] User uploads data from device
[0523] Users can select photos and videos suspected of being fake news and upload them to the system using their own devices. A simple interface is provided on the device to support data uploading.
[0524] The server receives the data
[0525] The server receives image and video data sent by users, which is transmitted via a secure communication protocol and stored and analyzed on the server side.
[0526] Metadata Analysis
[0527] The server analyzes the metadata of the received data, which includes the date and time of the photo, the location (GPS information), the model of the camera used, and the editing software used.
[0528] Forensic Analysis
[0529] The server performs forensic analysis to detect any tampering or manipulation of the media data. The following methods are used in this step:
[0530] Pixel compatibility check: Detects discontinuities in pixel color or brightness.
[0531] Edge Detection: Detects image edges and identifies unnatural edges.
[0532] Block Inspection: Analyzes compressed formats for block artifacts.
[0533] Content-Based Analysis
[0534] The server uses AI to analyze the media data based on its content, using the following methods:
[0535] Object Recognition: Recognize objects in images and videos and match them against a database.
[0536] Context analysis: Checks the consistency of each scene in the video and detects unnatural transitions.
[0537] Integrated evaluation and notification of results
[0538] The server integrates the results of each analysis and evaluates the reliability of the media data. Based on the reliability score, the data is classified into categories such as "high reliability," "partially suspicious," and "low reliability." The server then sends the evaluation results to the user's device, where the user can confirm the results.
[0539] Specific examples
[0540] 1. A user uploads a suspicious photo they found on a news site.
[0541] The user selects a photo on their device and uploads it to the system.
[0542] The server receives the photos and performs metadata and forensic analysis.
[0543] Content-based analysis reveals that the person in the photo does not match any photos in a reliable database.
[0544] The server evaluates the reliability as "low" and sends the result to the user's device.
[0545] The user reviews the results and makes a decision about the authenticity of the photo.
[0546] 2. When a user uploads a video shared on social media
[0547] The user selects a video on their device and uploads it to the system.
[0548] The server receives the video and performs metadata and forensic analysis on each frame.
[0549] Using object recognition technology, it is discovered that scenes in the video do not match other public records.
[0550] The server evaluates it as "partially suspicious" and sends the result to the user's device.
[0551] Users review the results and make a decision about the video's authenticity.
[0552] This allows the system to help users quickly and accurately determine the authenticity of photos and videos, preventing the spread of fake news.
[0553] The processing flow will be explained below.
[0554] Step 1:
[0555] The user selects an image or video.
[0556] The user selects images or videos suspected of being fake news from their device using the device's file selection function.
[0557] Step 2:
[0558] A user uploads data into the system.
[0559] The user presses the upload button to send the selected media data to the server. The data is sent to the server using a secure communication protocol (e.g., HTTPS).
[0560] Step 3:
[0561] The server receives the data.
[0562] The server receives the media data sent by the user, stores it in a database, and checks the file format and basic attributes.
[0563] Step 4:
[0564] The server pre-processes the data.
[0565] The server performs pre-processing on the received data, specifically resizing, color space conversion (e.g., RGB to grayscale), and noise reduction.
[0566] Step 5:
[0567] The server parses the metadata.
[0568] The server extracts and analyzes the metadata of the media data, checking the shooting date and time, shooting location (GPS information), camera model, editing software usage history, etc.
[0569] Step 6:
[0570] The server performs the forensic analysis.
[0571] The server performs forensic analysis to detect any evidence of tampering or manipulation of the media data, including pixel compatibility checks, edge detection, and block inspection.
[0572] Step 7:
[0573] The server performs content-based analysis.
[0574] The server uses AI technology to perform content-based analysis, object recognition to identify objects in images and videos and match them with a trusted database, and context analysis to check the consistency of each scene in the video.
[0575] Step 8:
[0576] The server integrates the results of each analysis and evaluates their reliability.
[0577] The server evaluates all analysis results comprehensively and assigns a reliability score, which is used to classify the media data as "highly reliable," "partially suspicious," or "low reliability."
[0578] Step 9:
[0579] The server transmits the evaluation results to the user terminal.
[0580] The server then transmits the evaluation results to the user terminal, again using a secure communication protocol.
[0581] Step 10:
[0582] The terminal displays the evaluation results.
[0583] The user terminal displays the received evaluation results, including the reliability score, analysis details, and the rationale for the results.
[0584] Step 11:
[0585] The user checks the results and makes a decision.
[0586] Users can view the results displayed on their device, make a judgment about the authenticity of the photo or video, and take action based on this information to prevent the spread of fake news.
[0587] Example 1
[0588] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0589] In modern society, the extremely rapid spread of information through digital media has created the problem of fake news spreading easily and making it difficult to distinguish it from real information. Furthermore, unreliable media data can lead to incorrect decision-making and cause social unrest. To solve these problems, a system is needed that can quickly and accurately evaluate the reliability of uploaded media data and provide accurate information to users.
[0590] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0591] In this invention, the server includes means for uploading media data from a user terminal, means for analyzing metadata of the uploaded media data, means for performing forensic analysis of the uploaded media data, means for performing content-based analysis of the uploaded media data, means for integrating the results of each analysis to evaluate the reliability, means for transmitting the evaluation result to the user terminal, and means for displaying the evaluation result, which makes it possible to prevent the spread of fake news and quickly and accurately determine the authenticity of media data.
[0592] A "user terminal" is an electronic device such as a computer, smartphone, or tablet that is operated by a user.
[0593] A "server" is a computer system set up to provide a particular service over a network.
[0594] "Media data" is data stored in digital format, such as images, video, and audio.
[0595] "Metadata" refers to data recorded in addition to the information on the media data itself, and includes information such as the date and time of shooting, the location, and the model of the camera.
[0596] "Forensic analysis" is a general term for technical analytical methods used to detect tampering or manipulation of digital data.
[0597] "Content-based analysis" is an analysis method based on the content of media data, and includes object recognition and context analysis.
[0598] "Authenticity assessment" is a process for assessing the authenticity and reliability of media data, and is carried out by integrating the results of various analyses.
[0599] "Resizing" refers to the process of changing the size of an image or video.
[0600] "Color space conversion" is the process of converting a color representation to a different color space.
[0601] "Noise reduction" is a process for removing unnecessary noise from media data.
[0602] "Pixel compatibility testing" is an analysis method that detects discontinuities in color and brightness at the pixel level.
[0603] "Edge detection" is an analysis technique that detects edges within an image and identifies unnatural edges.
[0604] "Block inspection" is a technique that analyzes block artifacts in compression formats to detect signs of tampering.
[0605] "Object recognition" is a technology that identifies specific objects or people in images or videos.
[0606] "Context analysis" is an analytical method that evaluates the continuity and consistency of each scene and detects unnatural transitions and inconsistencies.
[0607] The present invention relates to a system for evaluating the reliability of digital media data and preventing the spread of fake news. This system is implemented by uploading images and videos from a user's device to a server, which then analyzes the data. The following describes how this system is specifically implemented.
[0608] Configuration and Hardware
[0609] The system of the present invention consists of the following main components:
[0610] 1. User terminal: Electronic devices such as smartphones, tablets, and PCs. They provide a GUI (graphical user interface) to facilitate data uploading.
[0611] 2. Server: A high-performance computer system with the computing power and storage capacity to analyze the received media data. The server software uses programming languages such as Python or Java, and various libraries (e.g., OpenCV, TensorFlow, etc.) for data analysis.
[0612] Software and Data Processing
[0613] The system software performs the following main tasks:
[0614] 1. Upload your data:
[0615] A user uses a smartphone application to select suspicious media data and upload it to the server by selecting the "Upload suspicious media data" button, selecting photos or videos in the file selection dialog, and pressing the "Upload" button.
[0616] 2. Metadata analysis:
[0617] The server analyzes the metadata of the received media data, including the date and time of the photo, the location (GPS information), the model of the camera used, the editing software used, etc. The server extracts the Exif data of the image or video and checks whether it matches a specific event.
[0618] 3. Forensic Analysis:
[0619] The server performs forensic analysis to detect any manipulation or alteration of media data, primarily using pixel compatibility checks, edge detection, and block checks. It analyzes the color and brightness of each pixel to detect discontinuities. It detects image edges and checks for unnatural edges. It analyzes block artifacts of compression formats within images to find signs of manipulation.
[0620] 4. Content-based analysis:
[0621] The server uses AI technology to perform object recognition and context analysis, which evaluates the consistency of objects and scenes within images and videos. It uses AI models (e.g., deep learning models) to evaluate the continuity and context of each scene within a video and detect unnatural transitions and inconsistencies.
[0622] Analysis results and reliability evaluation
[0623] The server integrates the results of each analysis and evaluates the reliability of the media data. It calculates a reliability score and classifies the data into categories of "high reliability," "partially suspicious," or "low reliability." The evaluation results are sent to the user's device, where the user can confirm the results.
[0624] Examples and prompts
[0625] 1. Example: A photo from a news site:
[0626] A user uploads a suspicious photo they found on a news site. The server performs metadata and forensic analysis, and evaluates the trustworthiness of the photo using content-based analysis. The resulting rating is "low trustworthiness," and a notification is sent to the user's device.
[0627] 2. Example: Social media sharing video:
[0628] A user uploads a suspicious video shared on a social networking site. The server analyzes each frame, performing object recognition and context analysis. The video is ultimately evaluated as "partially suspicious" and a notification is sent to the user's device.
[0629] Prompt Sentence Examples
[0630] "Please rate the reliability of this image."
[0631] "Detect unnatural elements in this video."
[0632] In this way, the system of the present invention can quickly and accurately assess the authenticity of digital media data, providing a reliable means to prevent the spread of fake news.
[0633] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0634] Step 1:
[0635] User uploads data from device
[0636] Input: User-selected image or video files
[0637] Output: Sending data to the server
[0638] Specific behavior:
[0639] The user launches an application on the smartphone.
[0640] Select the "Upload suspicious media data" button on the application interface.
[0641] Select a photo or video in the file selection dialog and press the "Upload" button.
[0642] The selected file is sent to the server using the HTTPS protocol.
[0643] Step 2:
[0644] The server receives the data
[0645] Input: Image and video files sent from the user's device
[0646] Output: Received data stored on the server
[0647] Specific behavior:
[0648] The server listens for HTTPS requests.
[0649] The data sent by the user arrives at the server.
[0650] The server stores the received data in a secure directory.
[0651] Step 3:
[0652] The server performs metadata analysis
[0653] Input: Image and video files stored on the server
[0654] Output: Parsed metadata information
[0655] Specific behavior:
[0656] The server loads the image or video file.
[0657] Parse metadata information (e.g., Exif data).
[0658] The system extracts the date and time of the photo, GPS information, the model of the camera used, and the editing software usage history, and compares them with specific events and past data.
[0659] Step 4:
[0660] The server performs forensic analysis
[0661] Input: Image and video files stored on the server
[0662] Output: Forensic analysis results
[0663] Specific behavior:
[0664] A pixel compatibility check is performed to analyze the color and brightness of each pixel to detect discontinuities.
[0665] Apply edge detection to identify edges in the image and check if any unnatural edges are present.
[0666] Performs block inspection and analyzes block artifacts in compressed formats to find signs of tampering.
[0667] Step 5:
[0668] The server performs content-based analysis
[0669] Input: Image and video files stored on the server
[0670] Output: Results of content-based analysis
[0671] Specific behavior:
[0672] Use object recognition technology to recognize objects and people in images and videos.
[0673] Recognized objects and scenes are checked against a database to see if there is a match.
[0674] For video, it analyzes each frame to assess the consistency of each scene and detects unnatural transitions and inconsistencies.
[0675] Step 6:
[0676] The server performs an integrated evaluation and notifies the results.
[0677] Input: Results of metadata analysis, forensic analysis, and content-based analysis
[0678] Output: Reliability assessment and its results
[0679] Specific behavior:
[0680] The server consolidates the analysis results.
[0681] A reliability score is calculated and the data is classified into categories of "high reliability," "partially questionable," and "low reliability."
[0682] A message for transmitting the evaluation result to the user terminal is generated and transmitted using a secure communication protocol.
[0683] The user receives the results and displays them in the application.
[0684] Through these steps, the system prevents the spread of fake news and enables users to quickly and accurately determine the authenticity of digital media data.
[0685] (Application example 1)
[0686] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0687] In recent years, the spread of fake news on social networking services and news distribution sites has become a social problem. In particular, there have been many cases where media data such as images and videos have been tampered with, resulting in the dissemination of information that differs from the truth. Existing technologies require users to manually verify the authenticity of data, which has limitations in reliability and efficiency. Therefore, the present invention aims to provide a system that allows users to easily evaluate the reliability of images and videos and prevent the spread of fake news.
[0688] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0689] In this invention, the server includes means for uploading media data from a user terminal to the server, means for the server to analyze metadata of the uploaded media data, means for the server to perform forensic analysis of the uploaded media data, means for the server to perform content-based analysis of the uploaded media data, means for the server to integrate the results of each analysis and evaluate the reliability, means for the server to transmit the evaluation result to the user terminal, means for the user terminal to display the evaluation result, and means for the user to confirm the reliability evaluation result of the media data and prevent the spread of fake news before sharing the data. This enables users to quickly evaluate the reliability of data before sharing it on social networking services or news distribution sites, and effectively prevent the spread of fake news.
[0690] "Media data" is information stored in digital form, such as images and videos.
[0691] A "user terminal" is an electronic device that can be directly operated by a user, such as a smartphone or computer.
[0692] A "server" is a computer system that processes and manages data over a network.
[0693] "Metadata" is additional information that accompanies digital media data, including the date and time the image was taken, the location where it was taken, the model of the camera used, and so on.
[0694] "Forensic analysis" refers to scientific and technical methods for detecting tampering or manipulation of media data.
[0695] "Content-based analysis" is a method of analyzing digital media data based on its content, and includes object recognition and context analysis.
[0696] "Authenticity assessment" is the process of determining the authenticity and reliability of digital media data based on analysis results.
[0697] "Sending to the user terminal" refers to the act of sending the evaluation results from the server to a terminal that the user can directly operate.
[0698] "Displaying the evaluation results" means visually displaying the results of the analysis on the user terminal.
[0699] "Preventing the spread of fake news before data is shared" means preventing the spread of misinformation by assessing the reliability of media data before it is made public on social networking services, news distribution sites, etc.
[0700] The present invention provides a system for preventing the spread of fake news on social networking services and news distribution sites, allowing users to quickly and accurately evaluate the reliability of media data such as images and videos.
[0701] System configuration
[0702] The system consists of a user device and a server. User devices include smartphones, computers, and tablets. The server has data processing capabilities and storage, and receives, analyzes, and notifies users of the results.
[0703] Hardware and Software
[0704] An application runs on the user's device to select images and videos and upload them to the server. The specific development environment uses the cross-platform React Native.
[0705] The server runs a web server using Flask, which receives the uploaded media data, analyzes it, and performs authenticity assessment using AI models (using TensorFlow or PyTorch, for example) and forensic analysis tools.
[0706] Media Data Analysis
[0707] The server performs the following analysis on the uploaded media data:
[0708] 1. Metadata analysis: Analyzes the shooting date and time, shooting location (GPS information), camera model used, editing software usage history, etc.
[0709] 2. Forensic analysis: Detecting tampering or manipulation of media data, including pixel compatibility checks, edge detection, and discontinuous block checks.
[0710] 3. Content-based analysis: Using AI models to analyze the content of images and videos, specifically through object recognition and context analysis, and matching the results with a database.
[0711] Integrated evaluation and notification of results
[0712] The server integrates the results of each analysis and evaluates the reliability of the media data. The results are expressed as a reliability score and classified into categories such as "high reliability," "partially questionable," and "low reliability." The evaluation results are sent to the user's device in real time and can be viewed through the application.
[0713] Specific use cases
[0714] Assume that users verify the authenticity of photos and videos through the application before sharing them on social media. Below is a concrete example and an example of a prompt sentence to input to the generative AI model.
[0715] Specific examples
[0716] When a user uploads a suspicious photo they find on a news site:
[0717] The user selects a photo on their device and uploads it to the system.
[0718] The server receives the photos and performs metadata and forensic analysis.
[0719] Content-based analysis matches the content of an image against a trusted database.
[0720] The server evaluates it as "partially suspicious" and sends the result to the user's device.
[0721] The user reviews the results and makes a decision about the authenticity of the photo.
[0722] Example prompt for a generative AI model:
[0723] "Rate the trustworthiness of this image by analyzing metadata such as objects shown, location and time of capture, and camera model used, and then using pixel and edge detection and unnatural blocks to provide a confidence score based on its content."
[0724] This system allows users to easily assess the reliability of data before sharing it on social networking services or news distribution sites, effectively preventing the spread of fake news.
[0725] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0726] Step 1:
[0727] The user selects media data on the terminal.
[0728] Input: The user selects a suspicious image or video and presses the upload button.
[0729] Behavior: The path of the file selected on the device is obtained and an upload operation is triggered.
[0730] Output: The file path of the media data is ready to be sent by the application to the server.
[0731] Step 2:
[0732] The terminal uploads the media data to the server.
[0733] Input: The file path of the selected media data.
[0734] Operation: When a file is selected, the terminal begins the process of sending the file to the server as form data.
[0735] Output: A binary stream of media data sent to the server.
[0736] Step 3:
[0737] The server receives and stores the media data.
[0738] Input: A binary stream of uploaded media data.
[0739] Behavior: The server saves the received data to the specified directory, using Flask's upload function to save the file name appropriately.
[0740] Output: The file path of the saved media data.
[0741] Step 4:
[0742] The server performs the metadata analysis.
[0743] Input: The file path of the saved media data.
[0744] How it works: The server uses Python libraries (e.g., PIL and ExifRead) to extract metadata from the media data (such as the date and time of the image capture, GPS information, and camera model).
[0745] Output: A set of metadata information (in dictionary format).
[0746] Step 5:
[0747] The server performs the forensic analysis.
[0748] Input: The file path of the media data.
[0749] What it does: Performs pixel compatibility and edge detection on images, and checks for discontinuous blocks. Analysis is performed using OpenCV and Scikit-Image.
[0750] Output: Forensic analysis results (whether or not the data has been tampered with or altered).
[0751] Step 6:
[0752] The server performs content-based analysis.
[0753] Input: The file path of the media data.
[0754] How it works: It uses AI models (e.g. TensorFlow or PyTorch) to perform object recognition and context analysis in images and videos, and matches them with authoritative information in a database.
[0755] Output: Content-based analysis results (object recognition results and context consistency information).
[0756] Step 7:
[0757] The server integrates the results of each analysis and evaluates their reliability.
[0758] Input: Metadata analysis results, forensic analysis results, content-based analysis results.
[0759] How it works: Each analysis result is evaluated comprehensively and a reliability score is calculated. The data is classified using a reliability evaluation algorithm.
[0760] Output: Confidence score (categorised as high, medium, low etc.).
[0761] Step 8:
[0762] The server transmits the evaluation results to the user terminal.
[0763] Input: Confidence score.
[0764] Operation: The evaluation results are sent to the user's terminal in real time, and the data is formatted so that the user can check it immediately.
[0765] Output: The trust evaluation result sent to the user terminal.
[0766] Step 9:
[0767] The user terminal displays the evaluation results.
[0768] Input: The trust evaluation result received from the server.
[0769] Operation: The application displays the received evaluation results in the user interface so that the user can check them. The evaluation results also include detailed analysis information.
[0770] Output: The reliability assessment results displayed in the user interface.
[0771] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0772] This invention combines a system that evaluates the reliability of digital media data and prevents the spread of fake news with an emotion engine that recognizes user emotions. This system is realized by uploading images and videos from users' devices to a server, which then analyzes them. Furthermore, the emotion engine is used to integrate the analysis results with user emotions, allowing for more effective evaluation.
[0773] Program processing
[0774] User uploads data from device
[0775] Users can select photos and videos suspected of being fake news and upload them to the system using their own devices. A simple interface is provided on the device to support data uploading.
[0776] The server receives the data
[0777] The server receives image and video data sent by users, which is transmitted via a secure communication protocol and stored and analyzed on the server side.
[0778] Metadata Analysis
[0779] The server analyzes the metadata of the received data, which includes the date and time of the photo, the location (GPS information), the model of the camera used, and the editing software used.
[0780] Forensic Analysis
[0781] The server performs forensic analysis to detect any tampering or manipulation of the media data. The following methods are used in this step:
[0782] Pixel compatibility check: Detects discontinuities in pixel color or brightness.
[0783] Edge Detection: Detects image edges and identifies unnatural edges.
[0784] Block Inspection: Analyzes compressed formats for block artifacts.
[0785] Content-Based Analysis
[0786] The server uses AI to analyze the media data based on its content, using the following methods:
[0787] Object Recognition: Recognize objects in images and videos and match them against a database.
[0788] Context analysis: Checks the consistency of each scene in the video and detects unnatural transitions.
[0789] Emotion recognition by emotion engine
[0790] The server runs an emotion engine that recognizes the user's emotions. Based on the data sent from the user's device, the emotion engine performs the following:
[0791] Facial expression analysis: Capture the user's facial expressions with a camera and recognize the emotions they express.
[0792] Voice analysis: Analyze the user's voice and estimate their emotions.
[0793] Integrated evaluation and notification of results
[0794] The server integrates the results of each analysis and the emotion engine to evaluate the reliability of the media data. This evaluation includes the results of metadata analysis, forensic analysis, content-based analysis, and the user's emotional state. Based on the reliability score, the data is classified as "highly reliable," "partially suspicious," or "low reliability." The server then sends the evaluation results to the user's device, where the user can confirm the results.
[0795] Specific examples
[0796] 1. A user uploads a suspicious photo they found on a news site.
[0797] The user selects a photo on their device and uploads it to the system.
[0798] The server receives the photos and performs metadata and forensic analysis.
[0799] Content-based analysis reveals that the person in the photo does not match any photos in a reliable database.
[0800] The emotion engine analyzes the user's facial expressions and recognizes the user's emotions when viewing photos.
[0801] The server integrates the analysis results with the emotion recognition, evaluates it as "low reliability," and sends the result to the user's device.
[0802] The user reviews the results and makes a decision about the authenticity of the photo.
[0803] 2. When a user uploads a video shared on social media
[0804] The user selects a video on their device and uploads it to the system.
[0805] The server receives the video and performs metadata and forensic analysis on each frame.
[0806] Using object recognition technology, it is discovered that scenes in the video do not match other public records.
[0807] The emotion engine analyzes the user's voice and recognizes their emotions when watching videos.
[0808] The server integrates the analysis results with emotion recognition, evaluates the image as "partially suspicious," and sends the result to the user's device.
[0809] Users review the results and make a decision about the video's authenticity.
[0810] As described above, this system helps users quickly and accurately judge the authenticity of photos and videos, preventing the spread of fake news. Furthermore, by taking user emotions into account, it achieves more reliable evaluations.
[0811] The processing flow will be explained below.
[0812] Step 1:
[0813] The user selects an image or video.
[0814] The user selects images or videos suspected of being fake news from their device using the device's file selection function.
[0815] Step 2:
[0816] A user uploads data into the system.
[0817] The user presses the upload button to send the selected media data to the server. The data is sent to the server using a secure communication protocol (e.g., HTTPS).
[0818] Step 3:
[0819] The server receives the data.
[0820] The server receives the media data sent by the user, stores it in a database, and checks the file format and basic attributes.
[0821] Step 4:
[0822] The server pre-processes the data.
[0823] The server performs pre-processing on the received data, specifically resizing, color space conversion (e.g., RGB to grayscale), and noise reduction.
[0824] Step 5:
[0825] The server parses the metadata.
[0826] The server extracts and analyzes the metadata of the media data, checking the shooting date and time, shooting location (GPS information), camera model, editing software usage history, etc.
[0827] Step 6:
[0828] The server performs the forensic analysis.
[0829] The server performs forensic analysis to detect any evidence of tampering or manipulation of the media data, including pixel compatibility checks, edge detection, and block inspection.
[0830] Step 7:
[0831] The server performs content-based analysis.
[0832] The server uses AI technology to perform content-based analysis, object recognition to identify objects in images and videos and match them with a trusted database, and context analysis to check the consistency of each scene in the video.
[0833] Step 8:
[0834] The server runs the emotion engine.
[0835] The server recognizes the user's emotions based on the data sent from the user's device. The emotion engine estimates emotions through facial expression and voice analysis. Specifically, the server captures the user's facial expressions with a camera and records their voice with a microphone, and analyzes these to recognize emotions.
[0836] Step 9:
[0837] The server integrates the results of each analysis and evaluates their reliability.
[0838] The server comprehensively evaluates all analysis results and assigns a reliability score. Based on the reliability score, the media data is classified as "highly reliable," "partially suspicious," or "low reliability." Furthermore, the server takes into account the user's emotional state and reflects it in the evaluation.
[0839] Step 10:
[0840] The server transmits the evaluation results to the user terminal.
[0841] The server then transmits the evaluation results and sentiment analysis results to the user terminal, using a secure communication protocol.
[0842] Step 11:
[0843] The terminal displays the evaluation results.
[0844] The user terminal displays the received evaluation results, including the reliability score, analysis details, the rationale for the results, and comments based on the user's emotional state.
[0845] Step 12:
[0846] The user checks the results and makes a decision.
[0847] Users can view the results displayed on their device, make a judgment about the authenticity of the photo or video, and take action based on this information to prevent the spread of fake news.
[0848] Example 2
[0849] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0850] Fake news can have a significant impact on society, so there is a need for effective methods to detect it and prevent its spread. However, conventional systems only analyze media data and do not consider user sentiment, resulting in insufficient assessment of its reliability. In addition, it is necessary to improve the accuracy of fake news detection by integrating multiple analysis methods.
[0851] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0852] In this invention, the server includes means for analyzing metadata of media data, means for performing forensic analysis of the media data, means for performing content-based analysis of the media data, means for recognizing user emotions in the media data, and means for evaluating reliability by integrating the results of each analysis and the emotion recognition result, thereby making it possible to more accurately determine the authenticity of media data and effectively prevent the spread of fake news.
[0853] "Metadata" refers to background information that accompanies media data, and includes the date and time of shooting, the location where the footage was taken, the model of camera used, and the history of editing software used.
[0854] "Media data" refers to media such as images and videos stored in digital format.
[0855] "Forensic analysis" is a scientific analysis method for detecting tampering or manipulation of media data, and includes pixel compatibility testing, edge detection, block testing, etc.
[0856] "Content-based analysis" is a method for analyzing the authenticity of media data based on its content, and includes object recognition and context analysis.
[0857] "Emotion recognition" is a technology that detects and analyzes the emotions expressed by users in response to media data, and involves facial expression analysis and voice analysis.
[0858] "Trustworthiness assessment" is the process of integrating various analysis results and emotion recognition results to determine the authenticity and reliability of media data.
[0859] This invention is a system that evaluates the reliability of digital media data and prevents the spread of fake news. This system is realized by uploading images and videos from users' devices to a server, which then analyzes them. Furthermore, an emotion engine is used to integrate the analysis results with the user's emotions, allowing for more effective evaluation.
[0860] First, the user selects images or videos suspected of being fake news using the device. The selected media data is then uploaded to the server via a secure communication protocol (e.g., HTTPS). The device is provided with a user-friendly interface to assist in uploading the data.
[0861] The server receives the uploaded media data and analyzes its metadata, which extracts information such as the date and time of the photo shoot, the location, the model of the camera used, and the editing software used. Based on this information, the authenticity of the media data is evaluated.
[0862] The server then performs forensic analysis, using techniques such as pixel compatibility testing, edge detection, and block detection. Pixel compatibility testing detects discontinuities in pixel color or brightness, while edge detection finds image edges and identifies unnatural transitions and boundaries. Block detection analyzes block artifacts in compression formats to identify signs of image or video tampering.
[0863] Additionally, the server performs content-based analysis, which uses generative AI models (such as TensorFlow or PyTorch) to analyze media data based on its content. Specifically, it uses object recognition technology to recognize objects in images and videos and compare them with public databases. Context analysis checks the consistency of each scene in the video and detects unnatural scene transitions.
[0864] The server also operates an emotion engine to recognize the user's emotions. The emotion engine performs the following operations based on the data sent from the user's device: facial expression analysis recognizes emotions from the user's facial expression data captured by a camera, and voice analysis estimates emotions from the user's recorded voice data.
[0865] Finally, the server integrates the results of each analysis and the emotion engine to evaluate the reliability of the media data. This evaluation includes the results of metadata analysis, forensic analysis, content-based analysis, and the user's emotional state. Based on the reliability score, the media data is classified as "highly reliable," "partially suspicious," or "low reliability." The evaluation results are then sent from the server to the user's device, where the user can confirm the results.
[0866] Specific examples
[0867] 1. Uploading a suspicious photo you found on a news site
[0868] The user selects a photo on their device and uploads it to the system.
[0869] The server receives the photos and performs metadata and forensic analysis.
[0870] Content-based analysis reveals that the person in the photo does not match any photos in a reliable database.
[0871] The emotion engine analyzes the user's facial expressions and recognizes the user's emotions when viewing photos.
[0872] The server integrates the analysis results with the emotion recognition, evaluates it as "low reliability," and sends the result to the user's device.
[0873] The user reviews the results and makes a decision about the authenticity of the photo.
[0874] 2. When uploading a video shared on social media
[0875] The user selects a video on their device and uploads it to the system.
[0876] The server receives the video and performs metadata and forensic analysis on each frame.
[0877] Using object recognition technology, it is discovered that scenes in the video do not match other public records.
[0878] The emotion engine analyzes the user's voice and recognizes their emotions when watching videos.
[0879] The server integrates the analysis results with emotion recognition, evaluates the image as "partially suspicious," and sends the result to the user's device.
[0880] Users review the results and make a decision about the video's authenticity.
[0881] Prompt Sentence Examples
[0882] "Is this photo real? Rate its authenticity."
[0883] Please confirm whether the scene in this video actually happened.
[0884] "Please analyze this news image to see if it has been tampered with."
[0885] Through the above-described embodiment, the system helps users quickly and accurately determine the authenticity of photos and videos, preventing the spread of fake news. Furthermore, by taking into account the user's emotions, the system achieves more reliable evaluation.
[0886] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0887] Step 1:
[0888] User selects and uploads suspicious media
[0889] Input: Image or video data selected by the user on their device.
[0890] How it works: The user selects media suspected of being fake news using a web browser or a dedicated app on their device. The selected media file is then uploaded to a server via a secure communication protocol (e.g., HTTPS).
[0891] Output: Media data sent to the server over a secure communication channel.
[0892] Step 2:
[0893] The server receives and stores the media data.
[0894] Input: The media data uploaded in step 1.
[0895] Specific operation: The server receives image or video data sent by the user using a secure communication protocol, calculates a checksum to check the integrity and completeness of the received data, and stores it in a temporary data storage area.
[0896] Output: Media data stored on the server.
[0897] Step 3:
[0898] The server analyzes the metadata
[0899] Input: Media data stored on the server.
[0900] How it works: The server extracts and analyzes metadata from stored media files, including EXIF information such as the date and time the image was taken, the location where it was taken (GPS information), the model of the camera used, and the editing software used.
[0901] Output: Metadata analysis results.
[0902] Step 4:
[0903] Server performs forensic analysis
[0904] Input: Media data stored on the server.
[0905] What it does: The server performs forensic analysis: pixel compatibility checks to detect discontinuities in pixel color or brightness, edge detection detects image edges and identifies unnatural transitions and boundaries, and block inspection analyzes block artifacts specific to compression formats to detect signs of tampering.
[0906] Output: Forensic analysis results.
[0907] Step 5:
[0908] The server performs content-based analysis
[0909] Input: Media data stored on the server.
[0910] How it works: The server uses generative AI models (e.g., TensorFlow, PyTorch) to recognize objects in images and videos and match them against a trusted database. It also performs context analysis to check the consistency of each scene in the video and detect unnatural scene transitions.
[0911] Output: Content-based analysis results.
[0912] Step 6:
[0913] The server runs the emotion engine and performs emotion recognition.
[0914] Input: Media data stored on the server and emotion data from the user's device.
[0915] Specific operation: The server acquires data from the camera and microphone installed on the user's device. For facial expression analysis, the camera captures the user's facial expression data, and the AI model recognizes the emotion. For voice analysis, the server records the user's voice data and estimates the emotion.
[0916] Output: Emotion recognition result.
[0917] Step 7:
[0918] The server integrates the analysis results and emotion recognition to evaluate reliability
[0919] Input: Metadata analysis results, forensic analysis results, content-based analysis results, and emotion recognition results.
[0920] How it works: The server combines all the above analysis results and calculates a final reliability score, classifying the media data as "highly reliable," "partially questionable," or "low reliability."
[0921] Output: The estimated reliability score and its detailed analysis results.
[0922] Step 8:
[0923] The server sends the evaluation results to the user's device.
[0924] Input: The assessed reliability score and its detailed analysis results.
[0925] Specific operation: The server sends the trust evaluation result and its details to the user terminal using a secure communication protocol. An interface is provided that allows the user to check the result.
[0926] Output: Evaluation results are sent to the user's device, allowing the user to check the results.
[0927] The above specific processing steps enable the user to quickly and accurately judge the authenticity of the media data.
[0928] (Application example 2)
[0929] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0930] In recent years, the spread of fake news on social media and online platforms has become a social problem. Furthermore, users themselves often lack the ability to determine the veracity of information, resulting in a high risk of spreading false information. In response to this, there is a growing need for a system that can quickly and accurately evaluate the veracity of media data and provide feedback to users. Against this background, the present invention aims to provide a system for increasing the reliability of media data and suppressing the spread of fake news.
[0931] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for uploading media data from a user terminal to the server, means for the server analyzing metadata of the uploaded media data, means for the server performing forensic analysis of the uploaded media data, means for the server performing content-based analysis of the uploaded media data, means for the server acquiring user emotion data using an emotion engine that recognizes the user's emotion, means for the server integrating the analysis results and the user emotion data to evaluate reliability, means for the server transmitting the evaluation result to the user terminal, and means for the user terminal displaying the evaluation result. This makes it possible to quickly and accurately evaluate the authenticity of media data and provide the result to the user.
[0932] "Media data" is data in digital format that contains visual information such as images and videos.
[0933] A "user terminal" is a communication device, such as a smartphone, tablet, or PC, that is operated by a user to send and receive data.
[0934] A "server" is a computer system that manages and analyzes data on a network and supports communication with user terminals.
[0935] "Metadata" is supplementary information related to media data, and includes the date and time of shooting, the location of shooting, the model of the camera used, and editing history.
[0936] "Forensic analysis" is an analytical technique that technically detects whether media data has been tampered with or manipulated.
[0937] "Content-based analysis" is a technology that analyzes the content of images and videos themselves and evaluates object recognition and scene consistency.
[0938] An "emotion engine" is a technology or algorithm that recognizes a user's emotions by analyzing facial expressions and voice.
[0939] "Reliability evaluation" is the process of integrating each analysis result with the user's emotional data to determine the authenticity and credibility of the media data.
[0940] The "evaluation result" is information about the authenticity and reliability of the media data, generated based on the reliability evaluation.
[0941] "Preprocessing" refers to processing such as resizing, color space conversion, and noise removal to make the media data easier to analyze.
[0942] "Pixel compatibility testing" is a technology that detects data tampering by detecting discontinuities in the color or brightness of pixels within an image.
[0943] "Edge detection" is a technology that detects image tampering by detecting edges within an image and finding unnatural edges.
[0944] "Block inspection" is a technique that detects editing history and tampering of media data by analyzing block artifacts in compression formats.
[0945] The system embodying the present invention is implemented via a user terminal, a server, and communications between them. The overall configuration and operation procedure of the system will be described in detail below.
[0946] User terminal
[0947] The user device is a smartphone, tablet, or PC. The user uses the device to select suspicious media data (photos and videos) and upload them to the system. The device is provided with a simple interface that supports uploading media data, allowing users to upload data intuitively. The device is also equipped with a camera and microphone, which can capture the user's facial expressions and voice to collect emotional data.
[0948] server
[0949] The server is a computer system that receives and analyzes media data and emotion data sent from user terminals. The specific processing of the server for implementing the present invention is as follows.
[0950] Metadata Analysis
[0951] The server analyzes the metadata of the received media data, including the shooting date and time, shooting location (GPS information), camera model, and editing software usage history.
[0952] Forensic Analysis
[0953] The server performs forensic analysis to detect any tampering or manipulation of the media data, including pixel compatibility checks, edge detection, and block inspection.
[0954] Content-Based Analysis
[0955] The server uses AI to perform content-based analysis of the media data, including object recognition within images and videos, and contextual analysis to ensure consistency within video scenes.
[0956] Emotion Engine
[0957] The server recognizes the user's emotions using an emotion engine, which analyzes facial expression data and voice data sent by the user and estimates the user's emotions.
[0958] Integrated Evaluation
[0959] The server evaluates the reliability of the media data by integrating the results of each analysis with the user's emotional data. This evaluation includes the results of metadata analysis, forensic analysis, content-based analysis, and user emotion recognition. A reliability score is calculated and the media data is classified as "highly reliable," "partially suspicious," or "low reliability."
[0960] Result notification
[0961] The server transmits the evaluation results to the user terminal, which displays the received evaluation results so that the user can confirm them.
[0962] Specific examples
[0963] Take the example of a user uploading a suspicious video they found on a social networking site. The user selects the video on their device and uploads it to the system. The server receives the video and performs metadata and forensic analysis. At the same time, it uses object recognition technology to analyze scenes in the video and compare them with other public records. The server analyzes the user's voice and recognizes the emotions they express while watching the video. It then integrates all the analysis results, rates the video's authenticity as "partially suspicious," and sends the result to the user's device. The user then checks the results on their device and makes a decision about the video's authenticity.
[0964] Prompt Sentence Examples
[0965] "Upload your suspicious media and see the analysis results. Select a photo or video and press the upload button."
[0966] This will prevent the spread of fake news and allow users to make decisions based on accurate information.
[0967] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0968] Step 1:
[0969] A user selects suspicious media data (photos or videos) and uploads them to the system from their device. In this process, the user device provides a simple interface for acquiring media data and sending it to the server. The server receives the media data selected by the user as input and the sent media data as output.
[0970] Step 2:
[0971] The server receives media data sent from the user terminal. This process requires media data from the user terminal as input, which the server receives and stores. The stored media data is generated as output.
[0972] Step 3:
[0973] The server analyzes the metadata of the media data it receives. The stored media data is given as input, and the metadata analysis results are obtained as output. Specifically, it extracts information such as the shooting date and time, shooting location, camera model, and editing history.
[0974] Step 4:
[0975] The server performs forensic analysis of the media data. The media data obtained in the previous step is used as input, and the forensic analysis results are obtained as output. Specific analysis techniques used are pixel compatibility check, edge detection, and block inspection.
[0976] Step 5:
[0977] The server performs content-based analysis of the media data. It requires stored media data as input and produces the results of the content-based analysis as output. Specifically, it uses AI to recognize objects and evaluate scene consistency within images and videos.
[0978] Step 6:
[0979] The server uses an emotion engine to obtain the user's emotion data. Facial expression data and voice data sent from the user's device are used as input, and the emotion analysis results are obtained as output. The emotion engine recognizes the user's emotion based on facial expression and voice analysis.
[0980] Step 7:
[0981] The server integrates the results of each analysis and the user's sentiment data to assess the reliability of the media data. The results of metadata analysis, forensic analysis, content-based analysis, and sentiment analysis are used as input, and a reliability assessment score is generated as output. Based on this score, the media data is classified as "highly reliable," "partially suspicious," or "low reliability."
[0982] Step 8:
[0983] The server sends the trustworthiness evaluation results to the user terminal. The trustworthiness evaluation score is required as input, and the evaluation result is generated as output, which is received by the user terminal.
[0984] Step 9:
[0985] The user terminal displays the received evaluation results. The evaluation results sent from the server are used as input to generate the results that are displayed to the user as output. This allows the user to confirm the authenticity of the media data.
[0986] Through the above steps, the system is able to accurately and quickly evaluate the reliability of media data and notify the user of the results.
[0987] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0988] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0989] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.
[0990] [Third embodiment]
[0991] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.
[0992] 5, the data processing system 310 includes the data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[0993] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0994] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.
[0995] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0996] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0997] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0998] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0999] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1000] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1001] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1002] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."
[1003] This invention relates to a system for assessing the reliability of digital media data to prevent the spread of fake news. This system is implemented by uploading images and videos from user devices to a server, which then analyzes them. Specifically, the system performs metadata analysis, forensic analysis, and content-based analysis, and integrates the results to assess reliability.
[1004] Program processing
[1005] User uploads data from device
[1006] Users can select photos and videos suspected of being fake news and upload them to the system using their own devices. A simple interface is provided on the device to support data uploading.
[1007] The server receives the data
[1008] The server receives image and video data sent by users, which is transmitted via a secure communication protocol and stored and analyzed on the server side.
[1009] Metadata Analysis
[1010] The server analyzes the metadata of the received data, which includes the date and time of the photo, the location (GPS information), the model of the camera used, and the editing software used.
[1011] Forensic Analysis
[1012] The server performs forensic analysis to detect any tampering or manipulation of the media data. The following methods are used in this step:
[1013] Pixel compatibility check: Detects discontinuities in pixel color or brightness.
[1014] Edge Detection: Detects image edges and identifies unnatural edges.
[1015] Block Inspection: Analyzes compressed formats for block artifacts.
[1016] Content-Based Analysis
[1017] The server uses AI to analyze the media data based on its content, using the following methods:
[1018] Object Recognition: Recognize objects in images and videos and match them against a database.
[1019] Context analysis: Checks the consistency of each scene in the video and detects unnatural transitions.
[1020] Integrated evaluation and notification of results
[1021] The server integrates the results of each analysis and evaluates the reliability of the media data. Based on the reliability score, the data is classified into categories such as "high reliability," "partially suspicious," and "low reliability." The server then sends the evaluation results to the user's device, where the user can confirm the results.
[1022] Specific examples
[1023] 1. A user uploads a suspicious photo they found on a news site.
[1024] The user selects a photo on their device and uploads it to the system.
[1025] The server receives the photos and performs metadata and forensic analysis.
[1026] Content-based analysis reveals that the person in the photo does not match any photos in a reliable database.
[1027] The server evaluates the reliability as "low" and sends the result to the user's device.
[1028] The user reviews the results and makes a decision about the authenticity of the photo.
[1029] 2. When a user uploads a video shared on social media
[1030] The user selects a video on their device and uploads it to the system.
[1031] The server receives the video and performs metadata and forensic analysis on each frame.
[1032] Using object recognition technology, it is discovered that scenes in the video do not match other public records.
[1033] The server evaluates it as "partially suspicious" and sends the result to the user's device.
[1034] Users review the results and make a decision about the video's authenticity.
[1035] This allows the system to help users quickly and accurately determine the authenticity of photos and videos, preventing the spread of fake news.
[1036] The processing flow will be explained below.
[1037] Step 1:
[1038] The user selects an image or video.
[1039] The user selects images or videos suspected of being fake news from their device using the device's file selection function.
[1040] Step 2:
[1041] A user uploads data into the system.
[1042] The user presses the upload button to send the selected media data to the server. The data is sent to the server using a secure communication protocol (e.g., HTTPS).
[1043] Step 3:
[1044] The server receives the data.
[1045] The server receives the media data sent by the user, stores it in a database, and checks the file format and basic attributes.
[1046] Step 4:
[1047] The server pre-processes the data.
[1048] The server performs pre-processing on the received data, specifically resizing, color space conversion (e.g., RGB to grayscale), and noise reduction.
[1049] Step 5:
[1050] The server parses the metadata.
[1051] The server extracts and analyzes the metadata of the media data, checking the shooting date and time, shooting location (GPS information), camera model, editing software usage history, etc.
[1052] Step 6:
[1053] The server performs the forensic analysis.
[1054] The server performs forensic analysis to detect any evidence of tampering or manipulation of the media data, including pixel compatibility checks, edge detection, and block inspection.
[1055] Step 7:
[1056] The server performs content-based analysis.
[1057] The server uses AI technology to perform content-based analysis, object recognition to identify objects in images and videos and match them with a trusted database, and context analysis to check the consistency of each scene in the video.
[1058] Step 8:
[1059] The server integrates the results of each analysis and evaluates their reliability.
[1060] The server evaluates all analysis results comprehensively and assigns a reliability score, which is used to classify the media data as "highly reliable," "partially suspicious," or "low reliability."
[1061] Step 9:
[1062] The server transmits the evaluation results to the user terminal.
[1063] The server then transmits the evaluation results to the user terminal, again using a secure communication protocol.
[1064] Step 10:
[1065] The terminal displays the evaluation results.
[1066] The user terminal displays the received evaluation results, including the reliability score, analysis details, and the rationale for the results.
[1067] Step 11:
[1068] The user checks the results and makes a decision.
[1069] Users can view the results displayed on their device, make a judgment about the authenticity of the photo or video, and take action based on this information to prevent the spread of fake news.
[1070] Example 1
[1071] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1072] In modern society, the extremely rapid spread of information through digital media has created the problem of fake news spreading easily and making it difficult to distinguish it from real information. Furthermore, unreliable media data can lead to incorrect decision-making and cause social unrest. To solve these problems, a system is needed that can quickly and accurately evaluate the reliability of uploaded media data and provide accurate information to users.
[1073] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1074] In this invention, the server includes means for uploading media data from a user terminal, means for analyzing metadata of the uploaded media data, means for performing forensic analysis of the uploaded media data, means for performing content-based analysis of the uploaded media data, means for integrating the results of each analysis to evaluate the reliability, means for transmitting the evaluation result to the user terminal, and means for displaying the evaluation result, which makes it possible to prevent the spread of fake news and quickly and accurately determine the authenticity of media data.
[1075] A "user terminal" is an electronic device such as a computer, smartphone, or tablet that is operated by a user.
[1076] A "server" is a computer system set up to provide a particular service over a network.
[1077] "Media data" is data stored in digital format, such as images, video, and audio.
[1078] "Metadata" refers to data recorded in addition to the information on the media data itself, and includes information such as the date and time of shooting, the location, and the model of the camera.
[1079] "Forensic analysis" is a general term for technical analytical methods used to detect tampering or manipulation of digital data.
[1080] "Content-based analysis" is an analysis method based on the content of media data, and includes object recognition and context analysis.
[1081] "Authenticity assessment" is a process for assessing the authenticity and reliability of media data, and is carried out by integrating the results of various analyses.
[1082] "Resizing" refers to the process of changing the size of an image or video.
[1083] "Color space conversion" is the process of converting a color representation to a different color space.
[1084] "Noise reduction" is a process for removing unnecessary noise from media data.
[1085] "Pixel compatibility testing" is an analysis method that detects discontinuities in color and brightness at the pixel level.
[1086] "Edge detection" is an analysis technique that detects edges within an image and identifies unnatural edges.
[1087] "Block inspection" is a technique that analyzes block artifacts in compression formats to detect signs of tampering.
[1088] "Object recognition" is a technology that identifies specific objects or people in images or videos.
[1089] "Context analysis" is an analytical method that evaluates the continuity and consistency of each scene and detects unnatural transitions and inconsistencies.
[1090] The present invention relates to a system for evaluating the reliability of digital media data and preventing the spread of fake news. This system is implemented by uploading images and videos from a user's device to a server, which then analyzes the data. The following describes how this system is specifically implemented.
[1091] Configuration and Hardware
[1092] The system of the present invention consists of the following main components:
[1093] 1. User terminal: Electronic devices such as smartphones, tablets, and PCs. They provide a GUI (graphical user interface) to facilitate data uploading.
[1094] 2. Server: A high-performance computer system with the computing power and storage capacity to analyze the received media data. The server software uses programming languages such as Python or Java, and various libraries (e.g., OpenCV, TensorFlow, etc.) for data analysis.
[1095] Software and Data Processing
[1096] The system software performs the following main tasks:
[1097] 1. Upload your data:
[1098] A user uses a smartphone application to select suspicious media data and upload it to the server by selecting the "Upload suspicious media data" button, selecting photos or videos in the file selection dialog, and pressing the "Upload" button.
[1099] 2. Metadata analysis:
[1100] The server analyzes the metadata of the received media data, including the date and time of the photo, the location (GPS information), the model of the camera used, the editing software used, etc. The server extracts the Exif data of the image or video and checks whether it matches a specific event.
[1101] 3. Forensic Analysis:
[1102] The server performs forensic analysis to detect any manipulation or alteration of media data, primarily using pixel compatibility checks, edge detection, and block checks. It analyzes the color and brightness of each pixel to detect discontinuities. It detects image edges and checks for unnatural edges. It analyzes block artifacts of compression formats within images to find signs of manipulation.
[1103] 4. Content-based analysis:
[1104] The server uses AI technology to perform object recognition and context analysis, which evaluates the consistency of objects and scenes within images and videos. It uses AI models (e.g., deep learning models) to evaluate the continuity and context of each scene within a video and detect unnatural transitions and inconsistencies.
[1105] Analysis results and reliability evaluation
[1106] The server integrates the results of each analysis and evaluates the reliability of the media data. It calculates a reliability score and classifies the data into categories of "high reliability," "partially suspicious," or "low reliability." The evaluation results are sent to the user's device, where the user can confirm the results.
[1107] Examples and prompts
[1108] 1. Example: A photo from a news site:
[1109] A user uploads a suspicious photo they found on a news site. The server performs metadata and forensic analysis, and evaluates the trustworthiness of the photo using content-based analysis. The resulting rating is "low trustworthiness," and a notification is sent to the user's device.
[1110] 2. Example: Social media sharing video:
[1111] A user uploads a suspicious video shared on a social networking site. The server analyzes each frame, performing object recognition and context analysis. The video is ultimately evaluated as "partially suspicious" and a notification is sent to the user's device.
[1112] Prompt Sentence Examples
[1113] "Please rate the reliability of this image."
[1114] "Detect unnatural elements in this video."
[1115] In this way, the system of the present invention can quickly and accurately assess the authenticity of digital media data, providing a reliable means to prevent the spread of fake news.
[1116] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1117] Step 1:
[1118] User uploads data from device
[1119] Input: User-selected image or video files
[1120] Output: Sending data to the server
[1121] Specific behavior:
[1122] The user launches an application on the smartphone.
[1123] Select the "Upload suspicious media data" button on the application interface.
[1124] Select a photo or video in the file selection dialog and press the "Upload" button.
[1125] The selected file is sent to the server using the HTTPS protocol.
[1126] Step 2:
[1127] The server receives the data
[1128] Input: Image and video files sent from the user's device
[1129] Output: Received data stored on the server
[1130] Specific behavior:
[1131] The server listens for HTTPS requests.
[1132] The data sent by the user arrives at the server.
[1133] The server stores the received data in a secure directory.
[1134] Step 3:
[1135] The server performs metadata analysis
[1136] Input: Image and video files stored on the server
[1137] Output: Parsed metadata information
[1138] Specific behavior:
[1139] The server loads the image or video file.
[1140] Parse metadata information (e.g., Exif data).
[1141] The system extracts the date and time of the photo, GPS information, the model of the camera used, and the editing software usage history, and compares them with specific events and past data.
[1142] Step 4:
[1143] The server performs forensic analysis
[1144] Input: Image and video files stored on the server
[1145] Output: Forensic analysis results
[1146] Specific behavior:
[1147] A pixel compatibility check is performed to analyze the color and brightness of each pixel to detect discontinuities.
[1148] Apply edge detection to identify edges in the image and check if any unnatural edges are present.
[1149] Performs block inspection and analyzes block artifacts in compressed formats to find signs of tampering.
[1150] Step 5:
[1151] The server performs content-based analysis
[1152] Input: Image and video files stored on the server
[1153] Output: Results of content-based analysis
[1154] Specific behavior:
[1155] Use object recognition technology to recognize objects and people in images and videos.
[1156] Recognized objects and scenes are checked against a database to see if there is a match.
[1157] For video, it analyzes each frame to assess the consistency of each scene and detects unnatural transitions and inconsistencies.
[1158] Step 6:
[1159] The server performs an integrated evaluation and notifies the results.
[1160] Input: Results of metadata analysis, forensic analysis, and content-based analysis
[1161] Output: Reliability assessment and its results
[1162] Specific behavior:
[1163] The server consolidates the analysis results.
[1164] A reliability score is calculated and the data is classified into categories of "high reliability," "partially questionable," and "low reliability."
[1165] A message for transmitting the evaluation result to the user terminal is generated and transmitted using a secure communication protocol.
[1166] The user receives the results and displays them in the application.
[1167] Through these steps, the system prevents the spread of fake news and enables users to quickly and accurately determine the authenticity of digital media data.
[1168] (Application example 1)
[1169] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1170] In recent years, the spread of fake news on social networking services and news distribution sites has become a social problem. In particular, there have been many cases where media data such as images and videos have been tampered with, resulting in the dissemination of information that differs from the truth. Existing technologies require users to manually verify the authenticity of data, which has limitations in reliability and efficiency. Therefore, the present invention aims to provide a system that allows users to easily evaluate the reliability of images and videos and prevent the spread of fake news.
[1171] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1172] In this invention, the server includes means for uploading media data from a user terminal to the server, means for the server to analyze metadata of the uploaded media data, means for the server to perform forensic analysis of the uploaded media data, means for the server to perform content-based analysis of the uploaded media data, means for the server to integrate the results of each analysis and evaluate the reliability, means for the server to transmit the evaluation result to the user terminal, means for the user terminal to display the evaluation result, and means for the user to confirm the reliability evaluation result of the media data and prevent the spread of fake news before sharing the data. This enables users to quickly evaluate the reliability of data before sharing it on social networking services or news distribution sites, and effectively prevent the spread of fake news.
[1173] "Media data" is information stored in digital form, such as images and videos.
[1174] A "user terminal" is an electronic device that can be directly operated by a user, such as a smartphone or computer.
[1175] A "server" is a computer system that processes and manages data over a network.
[1176] "Metadata" is additional information that accompanies digital media data, including the date and time the image was taken, the location where it was taken, the model of the camera used, and so on.
[1177] "Forensic analysis" refers to scientific and technical methods for detecting tampering or manipulation of media data.
[1178] "Content-based analysis" is a method of analyzing digital media data based on its content, and includes object recognition and context analysis.
[1179] "Authenticity assessment" is the process of determining the authenticity and reliability of digital media data based on analysis results.
[1180] "Sending to the user terminal" refers to the act of sending the evaluation results from the server to a terminal that the user can directly operate.
[1181] "Displaying the evaluation results" means visually displaying the results of the analysis on the user terminal.
[1182] "Preventing the spread of fake news before data is shared" means preventing the spread of misinformation by assessing the reliability of media data before it is made public on social networking services, news distribution sites, etc.
[1183] The present invention provides a system for preventing the spread of fake news on social networking services and news distribution sites, allowing users to quickly and accurately evaluate the reliability of media data such as images and videos.
[1184] System configuration
[1185] The system consists of a user device and a server. User devices include smartphones, computers, and tablets. The server has data processing capabilities and storage, and receives, analyzes, and notifies users of the results.
[1186] Hardware and Software
[1187] An application runs on the user's device to select images and videos and upload them to the server. The specific development environment uses the cross-platform React Native.
[1188] The server runs a web server using Flask, which receives the uploaded media data, analyzes it, and performs authenticity assessment using AI models (using TensorFlow or PyTorch, for example) and forensic analysis tools.
[1189] Media Data Analysis
[1190] The server performs the following analysis on the uploaded media data:
[1191] 1. Metadata analysis: Analyzes the shooting date and time, shooting location (GPS information), camera model used, editing software usage history, etc.
[1192] 2. Forensic analysis: Detecting tampering or manipulation of media data, including pixel compatibility checks, edge detection, and discontinuous block checks.
[1193] 3. Content-based analysis: Using AI models to analyze the content of images and videos, specifically through object recognition and context analysis, and matching the results with a database.
[1194] Integrated evaluation and notification of results
[1195] The server integrates the results of each analysis and evaluates the reliability of the media data. The results are expressed as a reliability score and classified into categories such as "high reliability," "partially questionable," and "low reliability." The evaluation results are sent to the user's device in real time and can be viewed through the application.
[1196] Specific use cases
[1197] Assume that users verify the authenticity of photos and videos through the application before sharing them on social media. Below is a concrete example and an example of a prompt sentence to input to the generative AI model.
[1198] Specific examples
[1199] When a user uploads a suspicious photo they find on a news site:
[1200] The user selects a photo on their device and uploads it to the system.
[1201] The server receives the photos and performs metadata and forensic analysis.
[1202] Content-based analysis matches the content of an image against a trusted database.
[1203] The server evaluates it as "partially suspicious" and sends the result to the user's device.
[1204] The user reviews the results and makes a decision about the authenticity of the photo.
[1205] Example prompt for a generative AI model:
[1206] "Rate the trustworthiness of this image by analyzing metadata such as objects shown, location and time of capture, and camera model used, and then using pixel and edge detection and unnatural blocks to provide a confidence score based on its content."
[1207] This system allows users to easily assess the reliability of data before sharing it on social networking services or news distribution sites, effectively preventing the spread of fake news.
[1208] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1209] Step 1:
[1210] The user selects media data on the terminal.
[1211] Input: The user selects a suspicious image or video and presses the upload button.
[1212] Behavior: The path of the file selected on the device is obtained and an upload operation is triggered.
[1213] Output: The file path of the media data is ready to be sent by the application to the server.
[1214] Step 2:
[1215] The terminal uploads the media data to the server.
[1216] Input: The file path of the selected media data.
[1217] Operation: When a file is selected, the terminal begins the process of sending the file to the server as form data.
[1218] Output: A binary stream of media data sent to the server.
[1219] Step 3:
[1220] The server receives and stores the media data.
[1221] Input: A binary stream of uploaded media data.
[1222] Behavior: The server saves the received data to the specified directory, using Flask's upload function to save the file name appropriately.
[1223] Output: The file path of the saved media data.
[1224] Step 4:
[1225] The server performs the metadata analysis.
[1226] Input: The file path of the saved media data.
[1227] How it works: The server uses Python libraries (e.g., PIL and ExifRead) to extract metadata from the media data (such as the date and time of the image capture, GPS information, and camera model).
[1228] Output: A set of metadata information (in dictionary format).
[1229] Step 5:
[1230] The server performs the forensic analysis.
[1231] Input: The file path of the media data.
[1232] What it does: Performs pixel compatibility and edge detection on images, and checks for discontinuous blocks. Analysis is performed using OpenCV and Scikit-Image.
[1233] Output: Forensic analysis results (whether or not the data has been tampered with or altered).
[1234] Step 6:
[1235] The server performs content-based analysis.
[1236] Input: The file path of the media data.
[1237] How it works: It uses AI models (e.g. TensorFlow or PyTorch) to perform object recognition and context analysis in images and videos, and matches them with authoritative information in a database.
[1238] Output: Content-based analysis results (object recognition results and context consistency information).
[1239] Step 7:
[1240] The server integrates the results of each analysis and evaluates their reliability.
[1241] Input: Metadata analysis results, forensic analysis results, content-based analysis results.
[1242] How it works: Each analysis result is evaluated comprehensively and a reliability score is calculated. The data is classified using a reliability evaluation algorithm.
[1243] Output: Confidence score (categorised as high, medium, low etc.).
[1244] Step 8:
[1245] The server transmits the evaluation results to the user terminal.
[1246] Input: Confidence score.
[1247] Operation: The evaluation results are sent to the user's terminal in real time, and the data is formatted so that the user can check it immediately.
[1248] Output: The trust evaluation result sent to the user terminal.
[1249] Step 9:
[1250] The user terminal displays the evaluation results.
[1251] Input: The trust evaluation result received from the server.
[1252] Operation: The application displays the received evaluation results in the user interface so that the user can check them. The evaluation results also include detailed analysis information.
[1253] Output: The reliability assessment results displayed in the user interface.
[1254] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1255] This invention combines a system that evaluates the reliability of digital media data and prevents the spread of fake news with an emotion engine that recognizes user emotions. This system is realized by uploading images and videos from users' devices to a server, which then analyzes them. Furthermore, the emotion engine is used to integrate the analysis results with user emotions, allowing for more effective evaluation.
[1256] Program processing
[1257] User uploads data from device
[1258] Users can select photos and videos suspected of being fake news and upload them to the system using their own devices. A simple interface is provided on the device to support data uploading.
[1259] The server receives the data
[1260] The server receives image and video data sent by users, which is transmitted via a secure communication protocol and stored and analyzed on the server side.
[1261] Metadata Analysis
[1262] The server analyzes the metadata of the received data, which includes the date and time of the photo, the location (GPS information), the model of the camera used, and the editing software used.
[1263] Forensic Analysis
[1264] The server performs forensic analysis to detect any tampering or manipulation of the media data. The following methods are used in this step:
[1265] Pixel compatibility check: Detects discontinuities in pixel color or brightness.
[1266] Edge Detection: Detects image edges and identifies unnatural edges.
[1267] Block Inspection: Analyzes compressed formats for block artifacts.
[1268] Content-Based Analysis
[1269] The server uses AI to analyze the media data based on its content, using the following methods:
[1270] Object Recognition: Recognize objects in images and videos and match them against a database.
[1271] Context analysis: Checks the consistency of each scene in the video and detects unnatural transitions.
[1272] Emotion recognition by emotion engine
[1273] The server runs an emotion engine that recognizes the user's emotions. Based on the data sent from the user's device, the emotion engine performs the following:
[1274] Facial expression analysis: Capture the user's facial expressions with a camera and recognize the emotions they express.
[1275] Voice analysis: Analyze the user's voice and estimate their emotions.
[1276] Integrated evaluation and notification of results
[1277] The server integrates the results of each analysis and the emotion engine to evaluate the reliability of the media data. This evaluation includes the results of metadata analysis, forensic analysis, content-based analysis, and the user's emotional state. Based on the reliability score, the data is classified as "highly reliable," "partially suspicious," or "low reliability." The server then sends the evaluation results to the user's device, where the user can confirm the results.
[1278] Specific examples
[1279] 1. A user uploads a suspicious photo they found on a news site.
[1280] The user selects a photo on their device and uploads it to the system.
[1281] The server receives the photos and performs metadata and forensic analysis.
[1282] Content-based analysis reveals that the person in the photo does not match any photos in a reliable database.
[1283] The emotion engine analyzes the user's facial expressions and recognizes the user's emotions when viewing photos.
[1284] The server integrates the analysis results with the emotion recognition, evaluates it as "low reliability," and sends the result to the user's device.
[1285] The user reviews the results and makes a decision about the authenticity of the photo.
[1286] 2. When a user uploads a video shared on social media
[1287] The user selects a video on their device and uploads it to the system.
[1288] The server receives the video and performs metadata and forensic analysis on each frame.
[1289] Using object recognition technology, it is discovered that scenes in the video do not match other public records.
[1290] The emotion engine analyzes the user's voice and recognizes their emotions when watching videos.
[1291] The server integrates the analysis results with emotion recognition, evaluates the image as "partially suspicious," and sends the result to the user's device.
[1292] Users review the results and make a decision about the video's authenticity.
[1293] As described above, this system helps users quickly and accurately judge the authenticity of photos and videos, preventing the spread of fake news. Furthermore, by taking user emotions into account, it achieves more reliable evaluations.
[1294] The processing flow will be explained below.
[1295] Step 1:
[1296] The user selects an image or video.
[1297] The user selects images or videos suspected of being fake news from their device using the device's file selection function.
[1298] Step 2:
[1299] A user uploads data into the system.
[1300] The user presses the upload button to send the selected media data to the server. The data is sent to the server using a secure communication protocol (e.g., HTTPS).
[1301] Step 3:
[1302] The server receives the data.
[1303] The server receives the media data sent by the user, stores it in a database, and checks the file format and basic attributes.
[1304] Step 4:
[1305] The server pre-processes the data.
[1306] The server performs pre-processing on the received data, specifically resizing, color space conversion (e.g., RGB to grayscale), and noise reduction.
[1307] Step 5:
[1308] The server parses the metadata.
[1309] The server extracts and analyzes the metadata of the media data, checking the shooting date and time, shooting location (GPS information), camera model, editing software usage history, etc.
[1310] Step 6:
[1311] The server performs the forensic analysis.
[1312] The server performs forensic analysis to detect any evidence of tampering or manipulation of the media data, including pixel compatibility checks, edge detection, and block inspection.
[1313] Step 7:
[1314] The server performs content-based analysis.
[1315] The server uses AI technology to perform content-based analysis, object recognition to identify objects in images and videos and match them with a trusted database, and context analysis to check the consistency of each scene in the video.
[1316] Step 8:
[1317] The server runs the emotion engine.
[1318] The server recognizes the user's emotions based on the data sent from the user's device. The emotion engine estimates emotions through facial expression and voice analysis. Specifically, the server captures the user's facial expressions with a camera and records their voice with a microphone, and analyzes these to recognize emotions.
[1319] Step 9:
[1320] The server integrates the results of each analysis and evaluates their reliability.
[1321] The server comprehensively evaluates all analysis results and assigns a reliability score. Based on the reliability score, the media data is classified as "highly reliable," "partially suspicious," or "low reliability." Furthermore, the server takes into account the user's emotional state and reflects it in the evaluation.
[1322] Step 10:
[1323] The server transmits the evaluation results to the user terminal.
[1324] The server then transmits the evaluation results and sentiment analysis results to the user terminal, using a secure communication protocol.
[1325] Step 11:
[1326] The terminal displays the evaluation results.
[1327] The user terminal displays the received evaluation results, including the reliability score, analysis details, the rationale for the results, and comments based on the user's emotional state.
[1328] Step 12:
[1329] The user checks the results and makes a decision.
[1330] Users can view the results displayed on their device, make a judgment about the authenticity of the photo or video, and take action based on this information to prevent the spread of fake news.
[1331] Example 2
[1332] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1333] Fake news can have a significant impact on society, so there is a need for effective methods to detect it and prevent its spread. However, conventional systems only analyze media data and do not consider user sentiment, resulting in insufficient assessment of its reliability. In addition, it is necessary to improve the accuracy of fake news detection by integrating multiple analysis methods.
[1334] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[1335] In this invention, the server includes means for analyzing metadata of media data, means for performing forensic analysis of the media data, means for performing content-based analysis of the media data, means for recognizing user emotions in the media data, and means for evaluating reliability by integrating the results of each analysis and the emotion recognition result, thereby making it possible to more accurately determine the authenticity of media data and effectively prevent the spread of fake news.
[1336] "Metadata" refers to background information that accompanies media data, and includes the date and time of shooting, the location where the footage was taken, the model of camera used, and the history of editing software used.
[1337] "Media data" refers to media such as images and videos stored in digital format.
[1338] "Forensic analysis" is a scientific analysis method for detecting tampering or manipulation of media data, and includes pixel compatibility testing, edge detection, block testing, etc.
[1339] "Content-based analysis" is a method for analyzing the authenticity of media data based on its content, and includes object recognition and context analysis.
[1340] "Emotion recognition" is a technology that detects and analyzes the emotions expressed by users in response to media data, and involves facial expression analysis and voice analysis.
[1341] "Trustworthiness assessment" is the process of integrating various analysis results and emotion recognition results to determine the authenticity and reliability of media data.
[1342] This invention is a system that evaluates the reliability of digital media data and prevents the spread of fake news. This system is realized by uploading images and videos from users' devices to a server, which then analyzes them. Furthermore, an emotion engine is used to integrate the analysis results with the user's emotions, allowing for more effective evaluation.
[1343] First, the user selects images or videos suspected of being fake news using the device. The selected media data is then uploaded to the server via a secure communication protocol (e.g., HTTPS). The device is provided with a user-friendly interface to assist in uploading the data.
[1344] The server receives the uploaded media data and analyzes its metadata, which extracts information such as the date and time of the photo shoot, the location, the model of the camera used, and the editing software used. Based on this information, the authenticity of the media data is evaluated.
[1345] The server then performs forensic analysis, using techniques such as pixel compatibility testing, edge detection, and block detection. Pixel compatibility testing detects discontinuities in pixel color or brightness, while edge detection finds image edges and identifies unnatural transitions and boundaries. Block detection analyzes block artifacts in compression formats to identify signs of image or video tampering.
[1346] Additionally, the server performs content-based analysis, which uses generative AI models (such as TensorFlow or PyTorch) to analyze media data based on its content. Specifically, it uses object recognition technology to recognize objects in images and videos and compare them with public databases. Context analysis checks the consistency of each scene in the video and detects unnatural scene transitions.
[1347] The server also operates an emotion engine to recognize the user's emotions. The emotion engine performs the following operations based on the data sent from the user's device: facial expression analysis recognizes emotions from the user's facial expression data captured by a camera, and voice analysis estimates emotions from the user's recorded voice data.
[1348] Finally, the server integrates the results of each analysis and the emotion engine to evaluate the reliability of the media data. This evaluation includes the results of metadata analysis, forensic analysis, content-based analysis, and the user's emotional state. Based on the reliability score, the media data is classified as "highly reliable," "partially suspicious," or "low reliability." The evaluation results are then sent from the server to the user's device, where the user can confirm the results.
[1349] Specific examples
[1350] 1. Uploading a suspicious photo you found on a news site
[1351] The user selects a photo on their device and uploads it to the system.
[1352] The server receives the photos and performs metadata and forensic analysis.
[1353] Content-based analysis reveals that the person in the photo does not match any photos in a reliable database.
[1354] The emotion engine analyzes the user's facial expressions and recognizes the user's emotions when viewing photos.
[1355] The server integrates the analysis results with the emotion recognition, evaluates it as "low reliability," and sends the result to the user's device.
[1356] The user reviews the results and makes a decision about the authenticity of the photo.
[1357] 2. When uploading a video shared on social media
[1358] The user selects a video on their device and uploads it to the system.
[1359] The server receives the video and performs metadata and forensic analysis on each frame.
[1360] Using object recognition technology, it is discovered that scenes in the video do not match other public records.
[1361] The emotion engine analyzes the user's voice and recognizes their emotions when watching videos.
[1362] The server integrates the analysis results with emotion recognition, evaluates the image as "partially suspicious," and sends the result to the user's device.
[1363] Users review the results and make a decision about the video's authenticity.
[1364] Prompt Sentence Examples
[1365] "Is this photo real? Rate its authenticity."
[1366] Please confirm whether the scene in this video actually happened.
[1367] "Please analyze this news image to see if it has been tampered with."
[1368] Through the above-described embodiment, the system helps users quickly and accurately determine the authenticity of photos and videos, preventing the spread of fake news. Furthermore, by taking into account the user's emotions, the system achieves more reliable evaluation.
[1369] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1370] Step 1:
[1371] User selects and uploads suspicious media
[1372] Input: Image or video data selected by the user on their device.
[1373] How it works: The user selects media suspected of being fake news using a web browser or a dedicated app on their device. The selected media file is then uploaded to a server via a secure communication protocol (e.g., HTTPS).
[1374] Output: Media data sent to the server over a secure communication channel.
[1375] Step 2:
[1376] The server receives and stores the media data.
[1377] Input: The media data uploaded in step 1.
[1378] Specific operation: The server receives image or video data sent by the user using a secure communication protocol, calculates a checksum to check the integrity and completeness of the received data, and stores it in a temporary data storage area.
[1379] Output: Media data stored on the server.
[1380] Step 3:
[1381] The server analyzes the metadata
[1382] Input: Media data stored on the server.
[1383] How it works: The server extracts and analyzes metadata from stored media files, including EXIF information such as the date and time the image was taken, the location where it was taken (GPS information), the model of the camera used, and the editing software used.
[1384] Output: Metadata analysis results.
[1385] Step 4:
[1386] Server performs forensic analysis
[1387] Input: Media data stored on the server.
[1388] What it does: The server performs forensic analysis: pixel compatibility checks to detect discontinuities in pixel color or brightness, edge detection detects image edges and identifies unnatural transitions and boundaries, and block inspection analyzes block artifacts specific to compression formats to detect signs of tampering.
[1389] Output: Forensic analysis results.
[1390] Step 5:
[1391] The server performs content-based analysis
[1392] Input: Media data stored on the server.
[1393] How it works: The server uses generative AI models (e.g., TensorFlow, PyTorch) to recognize objects in images and videos and match them against a trusted database. It also performs context analysis to check the consistency of each scene in the video and detect unnatural scene transitions.
[1394] Output: Content-based analysis results.
[1395] Step 6:
[1396] The server runs the emotion engine and performs emotion recognition.
[1397] Input: Media data stored on the server and emotion data from the user's device.
[1398] Specific operation: The server acquires data from the camera and microphone installed on the user's device. For facial expression analysis, the camera captures the user's facial expression data, and the AI model recognizes the emotion. For voice analysis, the server records the user's voice data and estimates the emotion.
[1399] Output: Emotion recognition result.
[1400] Step 7:
[1401] The server integrates the analysis results and emotion recognition to evaluate reliability
[1402] Input: Metadata analysis results, forensic analysis results, content-based analysis results, and emotion recognition results.
[1403] How it works: The server combines all the above analysis results and calculates a final reliability score, classifying the media data as "highly reliable," "partially questionable," or "low reliability."
[1404] Output: The estimated reliability score and its detailed analysis results.
[1405] Step 8:
[1406] The server sends the evaluation results to the user's device.
[1407] Input: The assessed reliability score and its detailed analysis results.
[1408] Specific operation: The server sends the trust evaluation result and its details to the user terminal using a secure communication protocol. An interface is provided that allows the user to check the result.
[1409] Output: Evaluation results are sent to the user's device, allowing the user to check the results.
[1410] The above specific processing steps enable the user to quickly and accurately judge the authenticity of the media data.
[1411] (Application example 2)
[1412] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1413] In recent years, the spread of fake news on social media and online platforms has become a social problem. Furthermore, users themselves often lack the ability to determine the veracity of information, resulting in a high risk of spreading false information. In response to this, there is a growing need for a system that can quickly and accurately evaluate the veracity of media data and provide feedback to users. Against this background, the present invention aims to provide a system for increasing the reliability of media data and suppressing the spread of fake news.
[1414] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for uploading media data from a user terminal to the server, means for the server analyzing metadata of the uploaded media data, means for the server performing forensic analysis of the uploaded media data, means for the server performing content-based analysis of the uploaded media data, means for the server acquiring user emotion data using an emotion engine that recognizes the user's emotion, means for the server integrating the analysis results and the user emotion data to evaluate reliability, means for the server transmitting the evaluation result to the user terminal, and means for the user terminal displaying the evaluation result. This makes it possible to quickly and accurately evaluate the authenticity of media data and provide the result to the user.
[1415] "Media data" is data in digital format that contains visual information such as images and videos.
[1416] A "user terminal" is a communication device, such as a smartphone, tablet, or PC, that is operated by a user to send and receive data.
[1417] A "server" is a computer system that manages and analyzes data on a network and supports communication with user terminals.
[1418] "Metadata" is supplementary information related to media data, and includes the date and time of shooting, the location of shooting, the model of the camera used, and editing history.
[1419] "Forensic analysis" is an analytical technique that technically detects whether media data has been tampered with or manipulated.
[1420] "Content-based analysis" is a technology that analyzes the content of images and videos themselves and evaluates object recognition and scene consistency.
[1421] An "emotion engine" is a technology or algorithm that recognizes a user's emotions by analyzing facial expressions and voice.
[1422] "Reliability evaluation" is the process of integrating each analysis result with the user's emotional data to determine the authenticity and credibility of the media data.
[1423] The "evaluation result" is information about the authenticity and reliability of the media data, generated based on the reliability evaluation.
[1424] "Preprocessing" refers to processing such as resizing, color space conversion, and noise removal to make the media data easier to analyze.
[1425] "Pixel compatibility testing" is a technology that detects data tampering by detecting discontinuities in the color or brightness of pixels within an image.
[1426] "Edge detection" is a technology that detects image tampering by detecting edges within an image and finding unnatural edges.
[1427] "Block inspection" is a technique that detects editing history and tampering of media data by analyzing block artifacts in compression formats.
[1428] The system embodying the present invention is implemented via a user terminal, a server, and communications between them. The overall configuration and operation procedure of the system will be described in detail below.
[1429] User terminal
[1430] The user device is a smartphone, tablet, or PC. The user uses the device to select suspicious media data (photos and videos) and upload them to the system. The device is provided with a simple interface that supports uploading media data, allowing users to upload data intuitively. The device is also equipped with a camera and microphone, which can capture the user's facial expressions and voice to collect emotional data.
[1431] server
[1432] The server is a computer system that receives and analyzes media data and emotion data sent from user terminals. The specific processing of the server for implementing the present invention is as follows.
[1433] Metadata Analysis
[1434] The server analyzes the metadata of the received media data, including the shooting date and time, shooting location (GPS information), camera model, and editing software usage history.
[1435] Forensic Analysis
[1436] The server performs forensic analysis to detect any tampering or manipulation of the media data, including pixel compatibility checks, edge detection, and block inspection.
[1437] Content-Based Analysis
[1438] The server uses AI to perform content-based analysis of the media data, including object recognition within images and videos, and contextual analysis to ensure consistency within video scenes.
[1439] Emotion Engine
[1440] The server recognizes the user's emotions using an emotion engine, which analyzes facial expression data and voice data sent by the user and estimates the user's emotions.
[1441] Integrated Evaluation
[1442] The server evaluates the reliability of the media data by integrating the results of each analysis with the user's emotional data. This evaluation includes the results of metadata analysis, forensic analysis, content-based analysis, and user emotion recognition. A reliability score is calculated and the media data is classified as "highly reliable," "partially suspicious," or "low reliability."
[1443] Result notification
[1444] The server transmits the evaluation results to the user terminal, which displays the received evaluation results so that the user can confirm them.
[1445] Specific examples
[1446] Take the example of a user uploading a suspicious video they found on a social networking site. The user selects the video on their device and uploads it to the system. The server receives the video and performs metadata and forensic analysis. At the same time, it uses object recognition technology to analyze scenes in the video and compare them with other public records. The server analyzes the user's voice and recognizes the emotions they express while watching the video. It then integrates all the analysis results, rates the video's authenticity as "partially suspicious," and sends the result to the user's device. The user then checks the results on their device and makes a decision about the video's authenticity.
[1447] Prompt Sentence Examples
[1448] "Upload your suspicious media and see the analysis results. Select a photo or video and press the upload button."
[1449] This will prevent the spread of fake news and allow users to make decisions based on accurate information.
[1450] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1451] Step 1:
[1452] A user selects suspicious media data (photos or videos) and uploads them to the system from their device. In this process, the user device provides a simple interface for acquiring media data and sending it to the server. The server receives the media data selected by the user as input and the sent media data as output.
[1453] Step 2:
[1454] The server receives media data sent from the user terminal. This process requires media data from the user terminal as input, which the server receives and stores. The stored media data is generated as output.
[1455] Step 3:
[1456] The server analyzes the metadata of the media data it receives. The stored media data is given as input, and the metadata analysis results are obtained as output. Specifically, it extracts information such as the shooting date and time, shooting location, camera model, and editing history.
[1457] Step 4:
[1458] The server performs forensic analysis of the media data. The media data obtained in the previous step is used as input, and the forensic analysis results are obtained as output. Specific analysis techniques used are pixel compatibility check, edge detection, and block inspection.
[1459] Step 5:
[1460] The server performs content-based analysis of the media data. It requires stored media data as input and produces the results of the content-based analysis as output. Specifically, it uses AI to recognize objects and evaluate scene consistency within images and videos.
[1461] Step 6:
[1462] The server uses an emotion engine to obtain the user's emotion data. Facial expression data and voice data sent from the user's device are used as input, and the emotion analysis results are obtained as output. The emotion engine recognizes the user's emotion based on facial expression and voice analysis.
[1463] Step 7:
[1464] The server integrates the results of each analysis and the user's sentiment data to assess the reliability of the media data. The results of metadata analysis, forensic analysis, content-based analysis, and sentiment analysis are used as input, and a reliability assessment score is generated as output. Based on this score, the media data is classified as "highly reliable," "partially suspicious," or "low reliability."
[1465] Step 8:
[1466] The server sends the trustworthiness evaluation results to the user terminal. The trustworthiness evaluation score is required as input, and the evaluation result is generated as output, which is received by the user terminal.
[1467] Step 9:
[1468] The user terminal displays the received evaluation results. The evaluation results sent from the server are used as input to generate the results that are displayed to the user as output. This allows the user to confirm the authenticity of the media data.
[1469] Through the above steps, the system is able to accurately and quickly evaluate the reliability of media data and notify the user of the results.
[1470] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1471] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1472] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.
[1473] [Fourth embodiment]
[1474] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.
[1475] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[1476] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1477] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.
[1478] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1479] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1480] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1481] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.
[1482] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1483] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1484] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1485] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1486] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1487] This invention relates to a system for assessing the reliability of digital media data to prevent the spread of fake news. This system is implemented by uploading images and videos from user devices to a server, which then analyzes them. Specifically, the system performs metadata analysis, forensic analysis, and content-based analysis, and integrates the results to assess reliability.
[1488] Program processing
[1489] User uploads data from device
[1490] Users can select photos and videos suspected of being fake news and upload them to the system using their own devices. A simple interface is provided on the device to support data uploading.
[1491] The server receives the data
[1492] The server receives image and video data sent by users, which is transmitted via a secure communication protocol and stored and analyzed on the server side.
[1493] Metadata Analysis
[1494] The server analyzes the metadata of the received data, which includes the date and time of the photo, the location (GPS information), the model of the camera used, and the editing software used.
[1495] Forensic Analysis
[1496] The server performs forensic analysis to detect any tampering or manipulation of the media data. The following methods are used in this step:
[1497] Pixel compatibility check: Detects discontinuities in pixel color or brightness.
[1498] Edge Detection: Detects image edges and identifies unnatural edges.
[1499] Block Inspection: Analyzes compressed formats for block artifacts.
[1500] Content-Based Analysis
[1501] The server uses AI to analyze the media data based on its content, using the following methods:
[1502] Object Recognition: Recognize objects in images and videos and match them against a database.
[1503] Context analysis: Checks the consistency of each scene in the video and detects unnatural transitions.
[1504] Integrated evaluation and notification of results
[1505] The server integrates the results of each analysis and evaluates the reliability of the media data. Based on the reliability score, the data is classified into categories such as "high reliability," "partially suspicious," and "low reliability." The server then sends the evaluation results to the user's device, where the user can confirm the results.
[1506] Specific examples
[1507] 1. A user uploads a suspicious photo they found on a news site.
[1508] The user selects a photo on their device and uploads it to the system.
[1509] The server receives the photos and performs metadata and forensic analysis.
[1510] Content-based analysis reveals that the person in the photo does not match any photos in a reliable database.
[1511] The server evaluates the reliability as "low" and sends the result to the user's device.
[1512] The user reviews the results and makes a decision about the authenticity of the photo.
[1513] 2. When a user uploads a video shared on social media
[1514] The user selects a video on their device and uploads it to the system.
[1515] The server receives the video and performs metadata and forensic analysis on each frame.
[1516] Using object recognition technology, it is discovered that scenes in the video do not match other public records.
[1517] The server evaluates it as "partially suspicious" and sends the result to the user's device.
[1518] Users review the results and make a decision about the video's authenticity.
[1519] This allows the system to help users quickly and accurately determine the authenticity of photos and videos, preventing the spread of fake news.
[1520] The processing flow will be explained below.
[1521] Step 1:
[1522] The user selects an image or video.
[1523] The user selects images or videos suspected of being fake news from their device using the device's file selection function.
[1524] Step 2:
[1525] A user uploads data into the system.
[1526] The user presses the upload button to send the selected media data to the server. The data is sent to the server using a secure communication protocol (e.g., HTTPS).
[1527] Step 3:
[1528] The server receives the data.
[1529] The server receives the media data sent by the user, stores it in a database, and checks the file format and basic attributes.
[1530] Step 4:
[1531] The server pre-processes the data.
[1532] The server performs pre-processing on the received data, specifically resizing, color space conversion (e.g., RGB to grayscale), and noise reduction.
[1533] Step 5:
[1534] The server parses the metadata.
[1535] The server extracts and analyzes the metadata of the media data, checking the shooting date and time, shooting location (GPS information), camera model, editing software usage history, etc.
[1536] Step 6:
[1537] The server performs the forensic analysis.
[1538] The server performs forensic analysis to detect any evidence of tampering or manipulation of the media data, including pixel compatibility checks, edge detection, and block inspection.
[1539] Step 7:
[1540] The server performs content-based analysis.
[1541] The server uses AI technology to perform content-based analysis, object recognition to identify objects in images and videos and match them with a trusted database, and context analysis to check the consistency of each scene in the video.
[1542] Step 8:
[1543] The server integrates the results of each analysis and evaluates their reliability.
[1544] The server evaluates all analysis results comprehensively and assigns a reliability score, which is used to classify the media data as "highly reliable," "partially suspicious," or "low reliability."
[1545] Step 9:
[1546] The server transmits the evaluation results to the user terminal.
[1547] The server then transmits the evaluation results to the user terminal, again using a secure communication protocol.
[1548] Step 10:
[1549] The terminal displays the evaluation results.
[1550] The user terminal displays the received evaluation results, including the reliability score, analysis details, and the rationale for the results.
[1551] Step 11:
[1552] The user checks the results and makes a decision.
[1553] Users can view the results displayed on their device, make a judgment about the authenticity of the photo or video, and take action based on this information to prevent the spread of fake news.
[1554] Example 1
[1555] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1556] In modern society, the extremely rapid spread of information through digital media has created the problem of fake news spreading easily and making it difficult to distinguish it from real information. Furthermore, unreliable media data can lead to incorrect decision-making and cause social unrest. To solve these problems, a system is needed that can quickly and accurately evaluate the reliability of uploaded media data and provide accurate information to users.
[1557] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1558] In this invention, the server includes means for uploading media data from a user terminal, means for analyzing metadata of the uploaded media data, means for performing forensic analysis of the uploaded media data, means for performing content-based analysis of the uploaded media data, means for integrating the results of each analysis to evaluate the reliability, means for transmitting the evaluation result to the user terminal, and means for displaying the evaluation result, which makes it possible to prevent the spread of fake news and quickly and accurately determine the authenticity of media data.
[1559] A "user terminal" is an electronic device such as a computer, smartphone, or tablet that is operated by a user.
[1560] A "server" is a computer system set up to provide a particular service over a network.
[1561] "Media data" is data stored in digital format, such as images, video, and audio.
[1562] "Metadata" refers to data recorded in addition to the information on the media data itself, and includes information such as the date and time of shooting, the location, and the model of the camera.
[1563] "Forensic analysis" is a general term for technical analytical methods used to detect tampering or manipulation of digital data.
[1564] "Content-based analysis" is an analysis method based on the content of media data, and includes object recognition and context analysis.
[1565] "Authenticity assessment" is a process for assessing the authenticity and reliability of media data, and is carried out by integrating the results of various analyses.
[1566] "Resizing" refers to the process of changing the size of an image or video.
[1567] "Color space conversion" is the process of converting a color representation to a different color space.
[1568] "Noise reduction" is a process for removing unnecessary noise from media data.
[1569] "Pixel compatibility testing" is an analysis method that detects discontinuities in color and brightness at the pixel level.
[1570] "Edge detection" is an analysis technique that detects edges within an image and identifies unnatural edges.
[1571] "Block inspection" is a technique that analyzes block artifacts in compression formats to detect signs of tampering.
[1572] "Object recognition" is a technology that identifies specific objects or people in images or videos.
[1573] "Context analysis" is an analytical method that evaluates the continuity and consistency of each scene and detects unnatural transitions and inconsistencies.
[1574] The present invention relates to a system for evaluating the reliability of digital media data and preventing the spread of fake news. This system is implemented by uploading images and videos from a user's device to a server, which then analyzes the data. The following describes how this system is specifically implemented.
[1575] Configuration and Hardware
[1576] The system of the present invention consists of the following main components:
[1577] 1. User terminal: Electronic devices such as smartphones, tablets, and PCs. They provide a GUI (graphical user interface) to facilitate data uploading.
[1578] 2. Server: A high-performance computer system with the computing power and storage capacity to analyze the received media data. The server software uses programming languages such as Python or Java, and various libraries (e.g., OpenCV, TensorFlow, etc.) for data analysis.
[1579] Software and Data Processing
[1580] The system software performs the following main tasks:
[1581] 1. Upload your data:
[1582] A user uses a smartphone application to select suspicious media data and upload it to the server by selecting the "Upload suspicious media data" button, selecting photos or videos in the file selection dialog, and pressing the "Upload" button.
[1583] 2. Metadata analysis:
[1584] The server analyzes the metadata of the received media data, including the date and time of the photo, the location (GPS information), the model of the camera used, the editing software used, etc. The server extracts the Exif data of the image or video and checks whether it matches a specific event.
[1585] 3. Forensic Analysis:
[1586] The server performs forensic analysis to detect any manipulation or alteration of media data, primarily using pixel compatibility checks, edge detection, and block checks. It analyzes the color and brightness of each pixel to detect discontinuities. It detects image edges and checks for unnatural edges. It analyzes block artifacts of compression formats within images to find signs of manipulation.
[1587] 4. Content-based analysis:
[1588] The server uses AI technology to perform object recognition and context analysis, which evaluates the consistency of objects and scenes within images and videos. It uses AI models (e.g., deep learning models) to evaluate the continuity and context of each scene within a video and detect unnatural transitions and inconsistencies.
[1589] Analysis results and reliability evaluation
[1590] The server integrates the results of each analysis and evaluates the reliability of the media data. It calculates a reliability score and classifies the data into categories of "high reliability," "partially suspicious," or "low reliability." The evaluation results are sent to the user's device, where the user can confirm the results.
[1591] Examples and prompts
[1592] 1. Example: A photo from a news site:
[1593] A user uploads a suspicious photo they found on a news site. The server performs metadata and forensic analysis, and evaluates the trustworthiness of the photo using content-based analysis. The resulting rating is "low trustworthiness," and a notification is sent to the user's device.
[1594] 2. Example: Social media sharing video:
[1595] A user uploads a suspicious video shared on a social networking site. The server analyzes each frame, performing object recognition and context analysis. The video is ultimately evaluated as "partially suspicious" and a notification is sent to the user's device.
[1596] Prompt Sentence Examples
[1597] "Please rate the reliability of this image."
[1598] "Detect unnatural elements in this video."
[1599] In this way, the system of the present invention can quickly and accurately assess the authenticity of digital media data, providing a reliable means to prevent the spread of fake news.
[1600] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1601] Step 1:
[1602] User uploads data from device
[1603] Input: User-selected image or video files
[1604] Output: Sending data to the server
[1605] Specific behavior:
[1606] The user launches an application on the smartphone.
[1607] Select the "Upload suspicious media data" button on the application interface.
[1608] Select a photo or video in the file selection dialog and press the "Upload" button.
[1609] The selected file is sent to the server using the HTTPS protocol.
[1610] Step 2:
[1611] The server receives the data
[1612] Input: Image and video files sent from the user's device
[1613] Output: Received data stored on the server
[1614] Specific behavior:
[1615] The server listens for HTTPS requests.
[1616] The data sent by the user arrives at the server.
[1617] The server stores the received data in a secure directory.
[1618] Step 3:
[1619] The server performs metadata analysis
[1620] Input: Image and video files stored on the server
[1621] Output: Parsed metadata information
[1622] Specific behavior:
[1623] The server loads the image or video file.
[1624] Parse metadata information (e.g., Exif data).
[1625] The system extracts the date and time of the photo, GPS information, the model of the camera used, and the editing software usage history, and compares them with specific events and past data.
[1626] Step 4:
[1627] The server performs forensic analysis
[1628] Input: Image and video files stored on the server
[1629] Output: Forensic analysis results
[1630] Specific behavior:
[1631] A pixel compatibility check is performed to analyze the color and brightness of each pixel to detect discontinuities.
[1632] Apply edge detection to identify edges in the image and check if any unnatural edges are present.
[1633] Performs block inspection and analyzes block artifacts in compressed formats to find signs of tampering.
[1634] Step 5:
[1635] The server performs content-based analysis
[1636] Input: Image and video files stored on the server
[1637] Output: Results of content-based analysis
[1638] Specific behavior:
[1639] Use object recognition technology to recognize objects and people in images and videos.
[1640] Recognized objects and scenes are checked against a database to see if there is a match.
[1641] For video, it analyzes each frame to assess the consistency of each scene and detects unnatural transitions and inconsistencies.
[1642] Step 6:
[1643] The server performs an integrated evaluation and notifies the results.
[1644] Input: Results of metadata analysis, forensic analysis, and content-based analysis
[1645] Output: Reliability assessment and its results
[1646] Specific behavior:
[1647] The server consolidates the analysis results.
[1648] A reliability score is calculated and the data is classified into categories of "high reliability," "partially questionable," and "low reliability."
[1649] A message for transmitting the evaluation result to the user terminal is generated and transmitted using a secure communication protocol.
[1650] The user receives the results and displays them in the application.
[1651] Through these steps, the system prevents the spread of fake news and enables users to quickly and accurately determine the authenticity of digital media data.
[1652] (Application example 1)
[1653] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1654] In recent years, the spread of fake news on social networking services and news distribution sites has become a social problem. In particular, there have been many cases where media data such as images and videos have been tampered with, resulting in the dissemination of information that differs from the truth. Existing technologies require users to manually verify the authenticity of data, which has limitations in reliability and efficiency. Therefore, the present invention aims to provide a system that allows users to easily evaluate the reliability of images and videos and prevent the spread of fake news.
[1655] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1656] In this invention, the server includes means for uploading media data from a user terminal to the server, means for the server to analyze metadata of the uploaded media data, means for the server to perform forensic analysis of the uploaded media data, means for the server to perform content-based analysis of the uploaded media data, means for the server to integrate the results of each analysis and evaluate the reliability, means for the server to transmit the evaluation result to the user terminal, means for the user terminal to display the evaluation result, and means for the user to confirm the reliability evaluation result of the media data and prevent the spread of fake news before sharing the data. This enables users to quickly evaluate the reliability of data before sharing it on social networking services or news distribution sites, and effectively prevent the spread of fake news.
[1657] "Media data" is information stored in digital form, such as images and videos.
[1658] A "user terminal" is an electronic device that can be directly operated by a user, such as a smartphone or computer.
[1659] A "server" is a computer system that processes and manages data over a network.
[1660] "Metadata" is additional information that accompanies digital media data, including the date and time the image was taken, the location where it was taken, the model of the camera used, and so on.
[1661] "Forensic analysis" refers to scientific and technical methods for detecting tampering or manipulation of media data.
[1662] "Content-based analysis" is a method of analyzing digital media data based on its content, and includes object recognition and context analysis.
[1663] "Authenticity assessment" is the process of determining the authenticity and reliability of digital media data based on analysis results.
[1664] "Sending to the user terminal" refers to the act of sending the evaluation results from the server to a terminal that the user can directly operate.
[1665] "Displaying the evaluation results" means visually displaying the results of the analysis on the user terminal.
[1666] "Preventing the spread of fake news before data is shared" means preventing the spread of misinformation by assessing the reliability of media data before it is made public on social networking services, news distribution sites, etc.
[1667] The present invention provides a system for preventing the spread of fake news on social networking services and news distribution sites, allowing users to quickly and accurately evaluate the reliability of media data such as images and videos.
[1668] System configuration
[1669] The system consists of a user device and a server. User devices include smartphones, computers, and tablets. The server has data processing capabilities and storage, and receives, analyzes, and notifies users of the results.
[1670] Hardware and Software
[1671] An application runs on the user's device to select images and videos and upload them to the server. The specific development environment uses the cross-platform React Native.
[1672] The server runs a web server using Flask, which receives the uploaded media data, analyzes it, and performs authenticity assessment using AI models (using TensorFlow or PyTorch, for example) and forensic analysis tools.
[1673] Media Data Analysis
[1674] The server performs the following analysis on the uploaded media data:
[1675] 1. Metadata analysis: Analyzes the shooting date and time, shooting location (GPS information), camera model used, editing software usage history, etc.
[1676] 2. Forensic analysis: Detecting tampering or manipulation of media data, including pixel compatibility checks, edge detection, and discontinuous block checks.
[1677] 3. Content-based analysis: Using AI models to analyze the content of images and videos, specifically through object recognition and context analysis, and matching the results with a database.
[1678] Integrated evaluation and notification of results
[1679] The server integrates the results of each analysis and evaluates the reliability of the media data. The results are expressed as a reliability score and classified into categories such as "high reliability," "partially questionable," and "low reliability." The evaluation results are sent to the user's device in real time and can be viewed through the application.
[1680] Specific use cases
[1681] Assume that users verify the authenticity of photos and videos through the application before sharing them on social media. Below is a concrete example and an example of a prompt sentence to input to the generative AI model.
[1682] Specific examples
[1683] When a user uploads a suspicious photo they find on a news site:
[1684] The user selects a photo on their device and uploads it to the system.
[1685] The server receives the photos and performs metadata and forensic analysis.
[1686] Content-based analysis matches the content of an image against a trusted database.
[1687] The server evaluates it as "partially suspicious" and sends the result to the user's device.
[1688] The user reviews the results and makes a decision about the authenticity of the photo.
[1689] Example prompt for a generative AI model:
[1690] "Rate the trustworthiness of this image by analyzing metadata such as objects shown, location and time of capture, and camera model used, and then using pixel and edge detection and unnatural blocks to provide a confidence score based on its content."
[1691] This system allows users to easily assess the reliability of data before sharing it on social networking services or news distribution sites, effectively preventing the spread of fake news.
[1692] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1693] Step 1:
[1694] The user selects media data on the terminal.
[1695] Input: The user selects a suspicious image or video and presses the upload button.
[1696] Behavior: The path of the file selected on the device is obtained and an upload operation is triggered.
[1697] Output: The file path of the media data is ready to be sent by the application to the server.
[1698] Step 2:
[1699] The terminal uploads the media data to the server.
[1700] Input: The file path of the selected media data.
[1701] Operation: When a file is selected, the terminal begins the process of sending the file to the server as form data.
[1702] Output: A binary stream of media data sent to the server.
[1703] Step 3:
[1704] The server receives and stores the media data.
[1705] Input: A binary stream of uploaded media data.
[1706] Behavior: The server saves the received data to the specified directory, using Flask's upload function to save the file name appropriately.
[1707] Output: The file path of the saved media data.
[1708] Step 4:
[1709] The server performs the metadata analysis.
[1710] Input: The file path of the saved media data.
[1711] How it works: The server uses Python libraries (e.g., PIL and ExifRead) to extract metadata from the media data (such as the date and time of the image capture, GPS information, and camera model).
[1712] Output: A set of metadata information (in dictionary format).
[1713] Step 5:
[1714] The server performs the forensic analysis.
[1715] Input: The file path of the media data.
[1716] What it does: Performs pixel compatibility and edge detection on images, and checks for discontinuous blocks. Analysis is performed using OpenCV and Scikit-Image.
[1717] Output: Forensic analysis results (whether or not the data has been tampered with or altered).
[1718] Step 6:
[1719] The server performs content-based analysis.
[1720] Input: The file path of the media data.
[1721] How it works: It uses AI models (e.g. TensorFlow or PyTorch) to perform object recognition and context analysis in images and videos, and matches them with authoritative information in a database.
[1722] Output: Content-based analysis results (object recognition results and context consistency information).
[1723] Step 7:
[1724] The server integrates the results of each analysis and evaluates their reliability.
[1725] Input: Metadata analysis results, forensic analysis results, content-based analysis results.
[1726] How it works: Each analysis result is evaluated comprehensively and a reliability score is calculated. The data is classified using a reliability evaluation algorithm.
[1727] Output: Confidence score (categorised as high, medium, low etc.).
[1728] Step 8:
[1729] The server transmits the evaluation results to the user terminal.
[1730] Input: Confidence score.
[1731] Operation: The evaluation results are sent to the user's terminal in real time, and the data is formatted so that the user can check it immediately.
[1732] Output: The trust evaluation result sent to the user terminal.
[1733] Step 9:
[1734] The user terminal displays the evaluation results.
[1735] Input: The trust evaluation result received from the server.
[1736] Operation: The application displays the received evaluation results in the user interface so that the user can check them. The evaluation results also include detailed analysis information.
[1737] Output: The reliability assessment results displayed in the user interface.
[1738] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1739] This invention combines a system that evaluates the reliability of digital media data and prevents the spread of fake news with an emotion engine that recognizes user emotions. This system is realized by uploading images and videos from users' devices to a server, which then analyzes them. Furthermore, the emotion engine is used to integrate the analysis results with user emotions, allowing for more effective evaluation.
[1740] Program processing
[1741] User uploads data from device
[1742] Users can select photos and videos suspected of being fake news and upload them to the system using their own devices. A simple interface is provided on the device to support data uploading.
[1743] The server receives the data
[1744] The server receives image and video data sent by users, which is transmitted via a secure communication protocol and stored and analyzed on the server side.
[1745] Metadata Analysis
[1746] The server analyzes the metadata of the received data, which includes the date and time of the photo, the location (GPS information), the model of the camera used, and the editing software used.
[1747] Forensic Analysis
[1748] The server performs forensic analysis to detect any tampering or manipulation of the media data. The following methods are used in this step:
[1749] Pixel compatibility check: Detects discontinuities in pixel color or brightness.
[1750] Edge Detection: Detects image edges and identifies unnatural edges.
[1751] Block Inspection: Analyzes compressed formats for block artifacts.
[1752] Content-Based Analysis
[1753] The server uses AI to analyze the media data based on its content, using the following methods:
[1754] Object Recognition: Recognize objects in images and videos and match them against a database.
[1755] Context analysis: Checks the consistency of each scene in the video and detects unnatural transitions.
[1756] Emotion recognition by emotion engine
[1757] The server runs an emotion engine that recognizes the user's emotions. Based on the data sent from the user's device, the emotion engine performs the following:
[1758] Facial expression analysis: Capture the user's facial expressions with a camera and recognize the emotions they express.
[1759] Voice analysis: Analyze the user's voice and estimate their emotions.
[1760] Integrated evaluation and notification of results
[1761] The server integrates the results of each analysis and the emotion engine to evaluate the reliability of the media data. This evaluation includes the results of metadata analysis, forensic analysis, content-based analysis, and the user's emotional state. Based on the reliability score, the data is classified as "highly reliable," "partially suspicious," or "low reliability." The server then sends the evaluation results to the user's device, where the user can confirm the results.
[1762] Specific examples
[1763] 1. A user uploads a suspicious photo they found on a news site.
[1764] The user selects a photo on their device and uploads it to the system.
[1765] The server receives the photos and performs metadata and forensic analysis.
[1766] Content-based analysis reveals that the person in the photo does not match any photos in a reliable database.
[1767] The emotion engine analyzes the user's facial expressions and recognizes the user's emotions when viewing photos.
[1768] The server integrates the analysis results with the emotion recognition, evaluates it as "low reliability," and sends the result to the user's device.
[1769] The user reviews the results and makes a decision about the authenticity of the photo.
[1770] 2. When a user uploads a video shared on social media
[1771] The user selects a video on their device and uploads it to the system.
[1772] The server receives the video and performs metadata and forensic analysis on each frame.
[1773] Using object recognition technology, it is discovered that scenes in the video do not match other public records.
[1774] The emotion engine analyzes the user's voice and recognizes their emotions when watching videos.
[1775] The server integrates the analysis results with emotion recognition, evaluates the image as "partially suspicious," and sends the result to the user's device.
[1776] Users review the results and make a decision about the video's authenticity.
[1777] As described above, this system helps users quickly and accurately judge the authenticity of photos and videos, preventing the spread of fake news. Furthermore, by taking user emotions into account, it achieves more reliable evaluations.
[1778] The processing flow will be explained below.
[1779] Step 1:
[1780] The user selects an image or video.
[1781] The user selects images or videos suspected of being fake news from their device using the device's file selection function.
[1782] Step 2:
[1783] A user uploads data into the system.
[1784] The user presses the upload button to send the selected media data to the server. The data is sent to the server using a secure communication protocol (e.g., HTTPS).
[1785] Step 3:
[1786] The server receives the data.
[1787] The server receives the media data sent by the user, stores it in a database, and checks the file format and basic attributes.
[1788] Step 4:
[1789] The server pre-processes the data.
[1790] The server performs pre-processing on the received data, specifically resizing, color space conversion (e.g., RGB to grayscale), and noise reduction.
[1791] Step 5:
[1792] The server parses the metadata.
[1793] The server extracts and analyzes the metadata of the media data, checking the shooting date and time, shooting location (GPS information), camera model, editing software usage history, etc.
[1794] Step 6:
[1795] The server performs the forensic analysis.
[1796] The server performs forensic analysis to detect any evidence of tampering or manipulation of the media data, including pixel compatibility checks, edge detection, and block inspection.
[1797] Step 7:
[1798] The server performs content-based analysis.
[1799] The server uses AI technology to perform content-based analysis, object recognition to identify objects in images and videos and match them with a trusted database, and context analysis to check the consistency of each scene in the video.
[1800] Step 8:
[1801] The server runs the emotion engine.
[1802] The server recognizes the user's emotions based on the data sent from the user's device. The emotion engine estimates emotions through facial expression and voice analysis. Specifically, the server captures the user's facial expressions with a camera and records their voice with a microphone, and analyzes these to recognize emotions.
[1803] Step 9:
[1804] The server integrates the results of each analysis and evaluates their reliability.
[1805] The server comprehensively evaluates all analysis results and assigns a reliability score. Based on the reliability score, the media data is classified as "highly reliable," "partially suspicious," or "low reliability." Furthermore, the server takes into account the user's emotional state and reflects it in the evaluation.
[1806] Step 10:
[1807] The server transmits the evaluation results to the user terminal.
[1808] The server then transmits the evaluation results and sentiment analysis results to the user terminal, using a secure communication protocol.
[1809] Step 11:
[1810] The terminal displays the evaluation results.
[1811] The user terminal displays the received evaluation results, including the reliability score, analysis details, the rationale for the results, and comments based on the user's emotional state.
[1812] Step 12:
[1813] The user checks the results and makes a decision.
[1814] Users can view the results displayed on their device, make a judgment about the authenticity of the photo or video, and take action based on this information to prevent the spread of fake news.
[1815] Example 2
[1816] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1817] Fake news can have a significant impact on society, so there is a need for effective methods to detect it and prevent its spread. However, conventional systems only analyze media data and do not consider user sentiment, resulting in insufficient assessment of its reliability. In addition, it is necessary to improve the accuracy of fake news detection by integrating multiple analysis methods.
[1818] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[1819] In this invention, the server includes means for analyzing metadata of media data, means for performing forensic analysis of the media data, means for performing content-based analysis of the media data, means for recognizing user emotions in the media data, and means for evaluating reliability by integrating the results of each analysis and the emotion recognition result, thereby making it possible to more accurately determine the authenticity of media data and effectively prevent the spread of fake news.
[1820] "Metadata" refers to background information that accompanies media data, and includes the date and time of shooting, the location where the footage was taken, the model of camera used, and the history of editing software used.
[1821] "Media data" refers to media such as images and videos stored in digital format.
[1822] "Forensic analysis" is a scientific analysis method for detecting tampering or manipulation of media data, and includes pixel compatibility testing, edge detection, block testing, etc.
[1823] "Content-based analysis" is a method for analyzing the authenticity of media data based on its content, and includes object recognition and context analysis.
[1824] "Emotion recognition" is a technology that detects and analyzes the emotions expressed by users in response to media data, and involves facial expression analysis and voice analysis.
[1825] "Trustworthiness assessment" is the process of integrating various analysis results and emotion recognition results to determine the authenticity and reliability of media data.
[1826] This invention is a system that evaluates the reliability of digital media data and prevents the spread of fake news. This system is realized by uploading images and videos from users' devices to a server, which then analyzes them. Furthermore, an emotion engine is used to integrate the analysis results with the user's emotions, allowing for more effective evaluation.
[1827] First, the user selects images or videos suspected of being fake news using the device. The selected media data is then uploaded to the server via a secure communication protocol (e.g., HTTPS). The device is provided with a user-friendly interface to assist in uploading the data.
[1828] The server receives the uploaded media data and analyzes its metadata, which extracts information such as the date and time of the photo shoot, the location, the model of the camera used, and the editing software used. Based on this information, the authenticity of the media data is evaluated.
[1829] The server then performs forensic analysis, using techniques such as pixel compatibility testing, edge detection, and block detection. Pixel compatibility testing detects discontinuities in pixel color or brightness, while edge detection finds image edges and identifies unnatural transitions and boundaries. Block detection analyzes block artifacts in compression formats to identify signs of image or video tampering.
[1830] Additionally, the server performs content-based analysis, which uses generative AI models (such as TensorFlow or PyTorch) to analyze media data based on its content. Specifically, it uses object recognition technology to recognize objects in images and videos and compare them with public databases. Context analysis checks the consistency of each scene in the video and detects unnatural scene transitions.
[1831] The server also operates an emotion engine to recognize the user's emotions. The emotion engine performs the following operations based on the data sent from the user's device: facial expression analysis recognizes emotions from the user's facial expression data captured by a camera, and voice analysis estimates emotions from the user's recorded voice data.
[1832] Finally, the server integrates the results of each analysis and the emotion engine to evaluate the reliability of the media data. This evaluation includes the results of metadata analysis, forensic analysis, content-based analysis, and the user's emotional state. Based on the reliability score, the media data is classified as "highly reliable," "partially suspicious," or "low reliability." The evaluation results are then sent from the server to the user's device, where the user can confirm the results.
[1833] Specific examples
[1834] 1. Uploading a suspicious photo you found on a news site
[1835] The user selects a photo on their device and uploads it to the system.
[1836] The server receives the photos and performs metadata and forensic analysis.
[1837] Content-based analysis reveals that the person in the photo does not match any photos in a reliable database.
[1838] The emotion engine analyzes the user's facial expressions and recognizes the user's emotions when viewing photos.
[1839] The server integrates the analysis results with the emotion recognition, evaluates it as "low reliability," and sends the result to the user's device.
[1840] The user reviews the results and makes a decision about the authenticity of the photo.
[1841] 2. When uploading a video shared on social media
[1842] The user selects a video on their device and uploads it to the system.
[1843] The server receives the video and performs metadata and forensic analysis on each frame.
[1844] Using object recognition technology, it is discovered that scenes in the video do not match other public records.
[1845] The emotion engine analyzes the user's voice and recognizes their emotions when watching videos.
[1846] The server integrates the analysis results with emotion recognition, evaluates the image as "partially suspicious," and sends the result to the user's device.
[1847] Users review the results and make a decision about the video's authenticity.
[1848] Prompt Sentence Examples
[1849] "Is this photo real? Rate its authenticity."
[1850] Please confirm whether the scene in this video actually happened.
[1851] "Please analyze this news image to see if it has been tampered with."
[1852] Through the above-described embodiment, the system helps users quickly and accurately determine the authenticity of photos and videos, preventing the spread of fake news. Furthermore, by taking into account the user's emotions, the system achieves more reliable evaluation.
[1853] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1854] Step 1:
[1855] User selects and uploads suspicious media
[1856] Input: Image or video data selected by the user on their device.
[1857] How it works: The user selects media suspected of being fake news using a web browser or a dedicated app on their device. The selected media file is then uploaded to a server via a secure communication protocol (e.g., HTTPS).
[1858] Output: Media data sent to the server over a secure communication channel.
[1859] Step 2:
[1860] The server receives and stores the media data.
[1861] Input: The media data uploaded in step 1.
[1862] Specific operation: The server receives image or video data sent by the user using a secure communication protocol, calculates a checksum to check the integrity and completeness of the received data, and stores it in a temporary data storage area.
[1863] Output: Media data stored on the server.
[1864] Step 3:
[1865] The server analyzes the metadata
[1866] Input: Media data stored on the server.
[1867] How it works: The server extracts and analyzes metadata from stored media files, including EXIF information such as the date and time the image was taken, the location where it was taken (GPS information), the model of the camera used, and the editing software used.
[1868] Output: Metadata analysis results.
[1869] Step 4:
[1870] Server performs forensic analysis
[1871] Input: Media data stored on the server.
[1872] What it does: The server performs forensic analysis: pixel compatibility checks to detect discontinuities in pixel color or brightness, edge detection detects image edges and identifies unnatural transitions and boundaries, and block inspection analyzes block artifacts specific to compression formats to detect signs of tampering.
[1873] Output: Forensic analysis results.
[1874] Step 5:
[1875] The server performs content-based analysis
[1876] Input: Media data stored on the server.
[1877] How it works: The server uses generative AI models (e.g., TensorFlow, PyTorch) to recognize objects in images and videos and match them against a trusted database. It also performs context analysis to check the consistency of each scene in the video and detect unnatural scene transitions.
[1878] Output: Content-based analysis results.
[1879] Step 6:
[1880] The server runs the emotion engine and performs emotion recognition.
[1881] Input: Media data stored on the server and emotion data from the user's device.
[1882] Specific operation: The server acquires data from the camera and microphone installed on the user's device. For facial expression analysis, the camera captures the user's facial expression data, and the AI model recognizes the emotion. For voice analysis, the server records the user's voice data and estimates the emotion.
[1883] Output: Emotion recognition result.
[1884] Step 7:
[1885] The server integrates the analysis results and emotion recognition to evaluate reliability
[1886] Input: Metadata analysis results, forensic analysis results, content-based analysis results, and emotion recognition results.
[1887] How it works: The server combines all the above analysis results and calculates a final reliability score, classifying the media data as "highly reliable," "partially questionable," or "low reliability."
[1888] Output: The estimated reliability score and its detailed analysis results.
[1889] Step 8:
[1890] The server sends the evaluation results to the user's device.
[1891] Input: The assessed reliability score and its detailed analysis results.
[1892] Specific operation: The server sends the trust evaluation result and its details to the user terminal using a secure communication protocol. An interface is provided that allows the user to check the result.
[1893] Output: Evaluation results are sent to the user's device, allowing the user to check the results.
[1894] The above specific processing steps enable the user to quickly and accurately judge the authenticity of the media data.
[1895] (Application example 2)
[1896] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1897] In recent years, the spread of fake news on social media and online platforms has become a social problem. Furthermore, users themselves often lack the ability to determine the veracity of information, resulting in a high risk of spreading false information. In response to this, there is a growing need for a system that can quickly and accurately evaluate the veracity of media data and provide feedback to users. Against this background, the present invention aims to provide a system for increasing the reliability of media data and suppressing the spread of fake news.
[1898] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for uploading media data from a user terminal to the server, means for the server analyzing metadata of the uploaded media data, means for the server performing forensic analysis of the uploaded media data, means for the server performing content-based analysis of the uploaded media data, means for the server acquiring user emotion data using an emotion engine that recognizes the user's emotion, means for the server integrating the analysis results and the user emotion data to evaluate reliability, means for the server transmitting the evaluation result to the user terminal, and means for the user terminal displaying the evaluation result. This makes it possible to quickly and accurately evaluate the authenticity of media data and provide the result to the user.
[1899] "Media data" is data in digital format that contains visual information such as images and videos.
[1900] A "user terminal" is a communication device, such as a smartphone, tablet, or PC, that is operated by a user to send and receive data.
[1901] A "server" is a computer system that manages and analyzes data on a network and supports communication with user terminals.
[1902] "Metadata" is supplementary information related to media data, and includes the date and time of shooting, the location of shooting, the model of the camera used, and editing history.
[1903] "Forensic analysis" is an analytical technique that technically detects whether media data has been tampered with or manipulated.
[1904] "Content-based analysis" is a technology that analyzes the content of images and videos themselves and evaluates object recognition and scene consistency.
[1905] An "emotion engine" is a technology or algorithm that recognizes a user's emotions by analyzing facial expressions and voice.
[1906] "Reliability evaluation" is the process of integrating each analysis result with the user's emotional data to determine the authenticity and credibility of the media data.
[1907] The "evaluation result" is information about the authenticity and reliability of the media data, generated based on the reliability evaluation.
[1908] "Preprocessing" refers to processing such as resizing, color space conversion, and noise removal to make the media data easier to analyze.
[1909] "Pixel compatibility testing" is a technology that detects data tampering by detecting discontinuities in the color or brightness of pixels within an image.
[1910] "Edge detection" is a technology that detects image tampering by detecting edges within an image and finding unnatural edges.
[1911] "Block inspection" is a technique that detects editing history and tampering of media data by analyzing block artifacts in compression formats.
[1912] The system embodying the present invention is implemented via a user terminal, a server, and communications between them. The overall configuration and operation procedure of the system will be described in detail below.
[1913] User terminal
[1914] The user device is a smartphone, tablet, or PC. The user uses the device to select suspicious media data (photos and videos) and upload them to the system. The device is provided with a simple interface that supports uploading media data, allowing users to upload data intuitively. The device is also equipped with a camera and microphone, which can capture the user's facial expressions and voice to collect emotional data.
[1915] server
[1916] The server is a computer system that receives and analyzes media data and emotion data sent from user terminals. The specific processing of the server for implementing the present invention is as follows.
[1917] Metadata Analysis
[1918] The server analyzes the metadata of the received media data, including the shooting date and time, shooting location (GPS information), camera model, and editing software usage history.
[1919] Forensic Analysis
[1920] The server performs forensic analysis to detect any tampering or manipulation of the media data, including pixel compatibility checks, edge detection, and block inspection.
[1921] Content-Based Analysis
[1922] The server uses AI to perform content-based analysis of the media data, including object recognition within images and videos, and contextual analysis to ensure consistency within video scenes.
[1923] Emotion Engine
[1924] The server recognizes the user's emotions using an emotion engine, which analyzes facial expression data and voice data sent by the user and estimates the user's emotions.
[1925] Integrated Evaluation
[1926] The server evaluates the reliability of the media data by integrating the results of each analysis with the user's emotional data. This evaluation includes the results of metadata analysis, forensic analysis, content-based analysis, and user emotion recognition. A reliability score is calculated and the media data is classified as "highly reliable," "partially suspicious," or "low reliability."
[1927] Result notification
[1928] The server transmits the evaluation results to the user terminal, which displays the received evaluation results so that the user can confirm them.
[1929] Specific examples
[1930] Take the example of a user uploading a suspicious video they found on a social networking site. The user selects the video on their device and uploads it to the system. The server receives the video and performs metadata and forensic analysis. At the same time, it uses object recognition technology to analyze scenes in the video and compare them with other public records. The server analyzes the user's voice and recognizes the emotions they express while watching the video. It then integrates all the analysis results, rates the video's authenticity as "partially suspicious," and sends the result to the user's device. The user then checks the results on their device and makes a decision about the video's authenticity.
[1931] Prompt Sentence Examples
[1932] "Upload your suspicious media and see the analysis results. Select a photo or video and press the upload button."
[1933] This will prevent the spread of fake news and allow users to make decisions based on accurate information.
[1934] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1935] Step 1:
[1936] A user selects suspicious media data (photos or videos) and uploads them to the system from their device. In this process, the user device provides a simple interface for acquiring media data and sending it to the server. The server receives the media data selected by the user as input and the sent media data as output.
[1937] Step 2:
[1938] The server receives media data sent from the user terminal. This process requires media data from the user terminal as input, which the server receives and stores. The stored media data is generated as output.
[1939] Step 3:
[1940] The server analyzes the metadata of the media data it receives. The stored media data is given as input, and the metadata analysis results are obtained as output. Specifically, it extracts information such as the shooting date and time, shooting location, camera model, and editing history.
[1941] Step 4:
[1942] The server performs forensic analysis of the media data. The media data obtained in the previous step is used as input, and the forensic analysis results are obtained as output. Specific analysis techniques used are pixel compatibility check, edge detection, and block inspection.
[1943] Step 5:
[1944] The server performs content-based analysis of the media data. It requires stored media data as input and produces the results of the content-based analysis as output. Specifically, it uses AI to recognize objects and evaluate scene consistency within images and videos.
[1945] Step 6:
[1946] The server uses an emotion engine to obtain the user's emotion data. Facial expression data and voice data sent from the user's device are used as input, and the emotion analysis results are obtained as output. The emotion engine recognizes the user's emotion based on facial expression and voice analysis.
[1947] Step 7:
[1948] The server integrates the results of each analysis and the user's sentiment data to assess the reliability of the media data. The results of metadata analysis, forensic analysis, content-based analysis, and sentiment analysis are used as input, and a reliability assessment score is generated as output. Based on this score, the media data is classified as "highly reliable," "partially suspicious," or "low reliability."
[1949] Step 8:
[1950] The server sends the trustworthiness evaluation results to the user terminal. The trustworthiness evaluation score is required as input, and the evaluation result is generated as output, which is received by the user terminal.
[1951] Step 9:
[1952] The user terminal displays the received evaluation results. The evaluation results sent from the server are used as input to generate the results that are displayed to the user as output. This allows the user to confirm the authenticity of the media data.
[1953] Through the above steps, the system is able to accurately and quickly evaluate the reliability of media data and notify the user of the results.
[1954] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.
[1955] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1956] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.
[1957] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[1958] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.
[1959] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.
[1960] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).
[1961] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.
[1962] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."
[1963] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.
[1964] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).
[1965] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.
[1966] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.
[1967] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[1968] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.
[1969] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.
[1970] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.
[1971] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.
[1972] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.
[1973] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.
[1974] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.
[1975] The following is further disclosed regarding the above embodiment.
[1976] (Claim 1)
[1977] means for uploading media data from a user terminal to a server;
[1978] a means for the server to analyze metadata of the uploaded media data;
[1979] means for the server to perform forensic analysis of the uploaded media data;
[1980] means for the server to perform content-based analysis of the uploaded media data;
[1981] A means for the server to integrate each analysis result and evaluate its reliability;
[1982] A means for the server to transmit the evaluation result to a user terminal;
[1983] A means for displaying the evaluation results on the user terminal;
[1984] A system including:
[1985] (Claim 2)
[1986] 10. The system of claim 1, further comprising means for pre-processing the media data by resizing, color space conversion, and noise reduction.
[1987] (Claim 3)
[1988] 10. The system of claim 1, further comprising means for performing pixel compatibility check, edge detection, and block check as forensic analysis.
[1989] "Example 1"
[1990] (Claim 1)
[1991] means for uploading media data from a user terminal to a server;
[1992] a means for the server to analyze metadata of the uploaded media data;
[1993] means for the server to perform forensic analysis of the uploaded media data;
[1994] means for the server to perform content-based analysis of the uploaded media data;
[1995] A means for the server to integrate each analysis result and evaluate its reliability;
[1996] A means for the server to transmit the evaluation result to a user terminal;
[1997] A means for displaying the evaluation results on the user terminal;
[1998] A system including:
[1999] (Claim 2)
[2000] 10. The system of claim 1, further comprising means for pre-processing the media data by resizing, color space conversion, and noise reduction.
[2001] (Claim 3)
[2002] 10. The system of claim 1, further comprising means for performing pixel compatibility check, edge detection, and block check as forensic analysis.
[2003] (Claim 4)
[2004] 10. The system of claim 1, further comprising means for performing object recognition and context analysis as content-based analysis.
[2005] (Claim 5)
[2006] 10. The system of claim 1, further comprising means for classifying the reliability of the media data into categories of "highly reliable," "partially questionable," and "low reliability."
[2007] "Application Example 1"
[2008] (Claim 1)
[2009] means for uploading media data from a user terminal to a server;
[2010] a means for the server to analyze metadata of the uploaded media data;
[2011] means for the server to perform forensic analysis of the uploaded media data;
[2012] means for the server to perform content-based analysis of the uploaded media data;
[2013] A means for the server to integrate each analysis result and evaluate its reliability;
[2014] A means for the server to transmit the evaluation result to a user terminal;
[2015] A means for displaying the evaluation results on the user terminal;
[2016] A means for users to check the reliability evaluation results of media data and prevent the spread of fake news before sharing the data;
[2017] A system including:
[2018] (Claim 2)
[2019] 10. The system of claim 1, further comprising means for pre-processing the media data by resizing, color space conversion, and noise reduction.
[2020] (Claim 3)
[2021] 10. The system of claim 1, further comprising means for performing pixel compatibility check, edge detection, and block check as forensic analysis.
[2022] "Example 2: Combining Emotion Engines"
[2023] (Claim 1)
[2024] means for uploading media data from a user terminal to a server;
[2025] a means for the server to analyze metadata of the uploaded media data;
[2026] means for the server to perform forensic analysis of the uploaded media data;
[2027] means for the server to perform content-based analysis of the uploaded media data;
[2028] A means for the server to recognize user emotions from the uploaded media data;
[2029] A means for the server to integrate each analysis result and emotion recognition result and evaluate their reliability;
[2030] A means for the server to transmit the evaluation result to a user terminal;
[2031] A means for displaying the evaluation results on the user terminal;
[2032] A system including:
[2033] (Claim 2)
[2034] 10. The system of claim 1, further comprising means for pre-processing the media data by resizing, color space conversion, and noise reduction.
[2035] (Claim 3)
[2036] 10. The system of claim 1, further comprising means for performing pixel compatibility check, edge detection, and block check as forensic analysis.
[2037] (Claim 4)
[2038] 10. The system of claim 1, further comprising means for performing object recognition and context analysis as content-based analysis.
[2039] (Claim 5)
[2040] 2. The system according to claim 1, further comprising means for performing facial expression analysis and voice analysis as emotion recognition.
[2041] "Application example 2 when combining emotion engines"
[2042] (Claim 1)
[2043] means for uploading media data from a user terminal to a server;
[2044] a means for the server to analyze metadata of the uploaded media data;
[2045] means for the server to perform forensic analysis of the uploaded media data;
[2046] means for the server to perform content-based analysis of the uploaded media data;
[2047] A means for acquiring user emotion data by using an emotion engine that recognizes user emotions in the server;
[2048] A means for the server to integrate each analysis result and the user's emotion data and evaluate the reliability;
[2049] A means for the server to transmit the evaluation result to a user terminal;
[2050] A means for displaying the evaluation results on the user terminal;
[2051] A system including:
[2052] (Claim 2)
[2053] 10. The system of claim 1, further comprising means for pre-processing the media data by resizing, color space conversion, and noise reduction.
[2054] (Claim 3)
[2055] 10. The system of claim 1, further comprising means for performing pixel compatibility check, edge detection, and block check as forensic analysis. [Explanation of symbols]
[2056] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>
Claims
1. means for uploading media data from a user terminal to a server; a means for the server to analyze metadata of the uploaded media data; means for the server to perform forensic analysis of the uploaded media data; means for the server to perform content-based analysis of the uploaded media data; A means for the server to integrate each analysis result and evaluate its reliability; A means for the server to transmit the evaluation result to a user terminal; A means for displaying the evaluation results on the user terminal; A system including:
2. The system of claim 1 , further comprising means for pre-processing the media data by resizing, color space conversion, and noise reduction.
3. The system of claim 1 , further comprising means for performing pixel compatibility checks, edge detection, and block checks as forensic analyses.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A