System

The system addresses the challenge of deepfake misinformation by allowing users to upload content, analyze it for deep fakes, and provide a reliability score, effectively combating the spread of false information.

JP2026019716APending Publication Date: 2026-02-05SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024121464
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-07-26
Publication Date
2026-02-05

AI Technical Summary

Technical Problem

The spread of false information through deepfake technology poses a risk of social confusion and misunderstanding, necessitating a system to minimize its impact and enable users to consume accurate information.

Method used

A system that allows users to upload content, analyze it for deep fakes, calculate a reliability score, and return the results to the user, utilizing a server with deep learning frameworks to detect deep fakes and provide a confidence score.

Benefits of technology

Enables users to quickly and accurately evaluate the reliability of content, preventing the spread of false information and supporting informed decision-making.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026019716000001_ABST
    Figure 2026019716000001_ABST
Patent Text Reader

Abstract

A system is provided.SOLUTION: The system includes a means for allowing a user to upload content, a means for receiving the uploaded content and adding it to an analysis queue, a means for executing an algorithm for detecting Deepfake included in the content, a means for calculating the reliability Score of the content on the basis of a detection result, and a means for returning the reliability Score and an analysis result to the user.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The technology of the present disclosure relates to a system. [Background technology]

[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]

[0004] In today's information society, a vast amount of information is circulating through the media and social media platforms. This includes information that has been cleverly fabricated using deep fake technology, which increases the risk that viewers and readers will believe it to be fact. This not only spreads false information but can also cause social confusion and misunderstanding. There is a need to minimize the impact of such false information and enable users to make decisions based on accurate information. [Means for solving the problem]

[0005] The present invention solves the above problem with a system that includes a means for users to upload content, a means for receiving the uploaded content and adding it to an analysis queue, a means for executing an algorithm to detect deep fakes contained in the content, a means for calculating a reliability score for the content based on the detection results, and a means for returning the reliability score and analysis results to the user. This allows the system to determine the proportion of deep fakes contained in the content and provide a score indicating the reliability, enabling users to select and consume accurate information. As a result, it is possible to prevent the spread of false information and support viewers and readers in consuming accurate information.

[0006] "User" means an individual or organization that uses the system to upload content and receive the analysis results.

[0007] "Content" refers to information uploaded to a server in digital form, such as video, image, and audio files.

[0008] "Upload" refers to the act of a user sending content from their own terminal to a server.

[0009] "Receiving" refers to the action of the server receiving content sent by the user.

[0010] "Analysis queue" refers to a list or buffer where received content is queued for analysis.

[0011] "Algorithm" refers to a set of computational procedures or logical processes used to detect Deep Fakes.

[0012] "Deep fake" refers to media files containing false information created using artificial intelligence technology, such as images, videos, and audio that differ from their original content.

[0013] "Detection" refers to the act of using an algorithm to identify Deep Fakes contained within content.

[0014] "Reliability Score" is a numerical value that evaluates the accuracy of content and is used to indicate the reliability of the content as a whole.

[0015] "Return" refers to the act of the server sending back the analysis results and the reliability score to the user.

[0016] "System" refers to a collection of integrated software and hardware that provides a set of functions that allow users to upload content, analyze that content, and return the analysis results to the users. [Brief explanation of the drawings]

[0017] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION

[0018] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.

[0019] First, the terms used in the following description will be explained.

[0020] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).

[0021] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.

[0022] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.

[0023] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.

[0024] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."

[0025] [First embodiment]

[0026] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0027] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0028] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0029] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.

[0030] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0031] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0032] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.

[0033] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0034] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0035] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0036] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0037] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0038] To implement the present invention, it is necessary to build a system in which a user uploads content using a web application, a server analyzes the content, detects Deep Fake parts, calculates a reliability score, and returns the result to the user.

[0039] The server performs the following main operations:

[0040] 1. User uploads content

[0041] A user accesses the web application, selects the video, image, or audio file they want to analyze, and clicks the upload button, which causes the user's device to issue an HTTP POST request to send the selected content to the server.

[0042] 2. The server receives the uploaded content

[0043] The server receives the content sent from the user terminal, temporarily stores the content received by the server, and is ready for the next analysis step.

[0044] 3. The server adds it to the analysis queue

[0045] The server adds the received content to the analysis queue and registers it as an analysis job. This operation puts the content in a queue and allows subsequent analysis processing to begin.

[0046] 4. The server runs the Deep Fake detection algorithm

[0047] The server runs the DeepFake detection algorithm on the content queued for analysis. This algorithm performs a detailed analysis of each frame (video / image) or each second (audio) to detect DeepFake signatures.

[0048] 5. The server calculates the reliability score based on the detection result.

[0049] The server calculates the reliability score of the entire content based on the ratio of detected deep fakes. The reliability score is provided to the user as a numerical value for evaluating the accuracy of the content.

[0050] 6. The server returns the analysis results and the reliability score to the user.

[0051] The server returns the completed analysis results and the confidence score to the user, which are then displayed in the user's web application.

[0052] Specific examples

[0053] For example, suppose user "A" wants to check the reliability of a certain news video. A uploads the video through a web application, and the server receives the video and adds it to the analysis queue. The server analyzes each frame of the video to detect whether or not it contains Deep fakes. The detection results indicate that 5% of the frames are Deep fakes, and calculate a reliability score of 75%. The server notifies A of this reliability score and detailed analysis results. Based on this reliability score, A can determine that the news video is relatively reliable.

[0054] In this way, the present invention provides a system that allows users to easily evaluate the reliability of the content they consume and assists them in making accurate information selections.

[0055] The processing flow will be explained below.

[0056] Step 1:

[0057] The user selects the content and performs the upload operation. The user accesses the web application, selects the file they want to analyze, and clicks the "Upload" button. This operation sends the selected file from the user's device to the server.

[0058] Step 2:

[0059] The user device sends a file to the server via an HTTP POST request, along with metadata, structuring the file in a format that the server can receive.

[0060] Step 3:

[0061] The server receives the file sent from the user terminal, temporarily stores the file, and prepares it for the next analysis process.

[0062] Step 4:

[0063] The server adds the received file to the analysis queue. When it is registered as an analysis job, a job ID is generated and used to track the processing progress.

[0064] Step 5:

[0065] The server initializes the Deep Fake detection algorithm. In this step, the necessary models and libraries are loaded and the system is ready for analysis.

[0066] Step 6:

[0067] The server splits the file. In the case of video, it is split into frames, and in the case of images and audio, it is split into appropriate units (pixels, seconds).

[0068] Step 7:

[0069] The server analyzes each unit (frame, pixel, second) using a Deep Fake detection algorithm, which evaluates each unit individually to determine the presence of Deep Fake.

[0070] Step 8:

[0071] The server aggregates the analysis results for all units and calculates the Deep Fake ratio for the entire content based on the detection results for each unit.

[0072] Step 9:

[0073] The server calculates a trust score based on the Deep Fake ratio. The trust score is expressed between 0 and 100, with higher scores indicating more trustworthy content.

[0074] Step 10:

[0075] The server saves the analysis results and the confidence score in a database. The results are saved as results associated with the job ID, and can be returned in response to a user request.

[0076] Step 11:

[0077] The server notifies the user that the analysis results are ready, and the user can then review the results.

[0078] Step 12:

[0079] The user terminal sends a request to the server to display the analysis results. The request includes the job ID.

[0080] Step 13:

[0081] The server retrieves the analysis results corresponding to the job ID from the database and returns them to the user device along with the reliability score.

[0082] Step 14:

[0083] The user's device displays the received reliability score and analysis results, and the user evaluates the reliability of the content based on these results.

[0084] Example 1

[0085] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0086] In recent years, the development of deep fake technology has increased the risk of inaccurate information spreading. As a result, it has become more difficult for users to determine trustworthy content. Therefore, there is a need for technology that can quickly and accurately evaluate the reliability of content uploaded by users. Conventional technologies require a lot of time to analyze content, and the results are not always highly reliable. The present invention aims to solve this problem.

[0087] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0088] In this invention, the server includes a means for users to upload content, a means for receiving the uploaded content and adding it to an analysis queue, a means for executing an algorithm to detect Deep Fakes contained in the content, a means for calculating a reliability score of the content based on the detection result, a means for returning the reliability score and the analysis result to the user, and a means for displaying the analysis result, thereby enabling the user to quickly and accurately evaluate the reliability of the content.

[0089] A "user" is a person or entity that uses the system to upload content and receive the results of that analysis.

[0090] "Content" refers to the media data to be analyzed, such as videos, images, and audio files.

[0091] "Uploading means" means the functionality or interface that allows a user to submit content to the system.

[0092] "Server" means a computer system that receives, analyzes, and processes content uploaded by users.

[0093] "Analysis Queue" refers to a queue that registers and manages uploaded content for analysis in order.

[0094] "Deep fake" is fake media content generated using deep learning technology.

[0095] "Means for executing algorithms" refers to the ability to execute programs or processes to detect Deep Fakes in content.

[0096] The "Reliability Score" is a numerical value that evaluates the reliability of content, calculated based on the percentage of detected Deep Fakes.

[0097] "Analysis results" refers to the information and data obtained after the server runs the Deep Fake detection algorithm.

[0098] The "means for returning" is a function that allows the server to provide the analysis results and the reliability score to the user.

[0099] "Display means" refers to an interface or method for visually presenting the analysis results and the reliability score to the user.

[0100] A "database" is an electronic data storage system that stores analysis results and allows for searching and referencing as needed.

[0101] A "request" refers to an operation or action in which a user requests information such as analysis results from a system.

[0102] MODE FOR CARRYING OUT THE INVENTION

[0103] In the system based on this invention, a user uploads content using a web application, the server analyzes the content, detects Deep Fake parts, calculates a reliability score, and returns the result to the user. A specific embodiment of this system will be described below.

[0104] System Overview

[0105] Users can access the web application and select and upload video, image, or audio files they want to analyze. The user device issues an HTTP POST request to send the selected content to the server. The server temporarily stores the received content and then adds it to the analysis queue. The server runs the Deep Fake detection algorithm on the content in the analysis queue and calculates a confidence score based on the detection results. Finally, the server returns the analysis results and confidence score to the user and displays them in the web application.

[0106] Hardware and Software Use

[0107] server:

[0108] The server can be a cloud-based virtual machine with high-performance computing resources, such as an EC2 instance from Amazon Web Services (AWS) or a Compute Engine from Google Cloud Platform (GCP).

[0109] The server uses MySQL or PostgreSQL as a database to efficiently store and manage analysis results.

[0110] Deep Learning Frameworks:

[0111] Deep fake detection uses deep learning frameworks such as TensorFlow and PyTorch, which enable advanced image and video analysis.

[0112] User device:

[0113] A user terminal is a device that can connect to the Internet, such as a PC, smartphone, or tablet, and accesses web applications through a web browser (e.g., Google Chrome or Mozilla Firefox).

[0114] Specific examples

[0115] For example, suppose user "A" wants to check the reliability of a certain news video. A uploads the video through a web application, and the server receives the video and adds it to the analysis queue. The server analyzes each frame of the video to detect whether or not it contains Deep fakes. The detection results indicate that 5% of the frames are Deep fakes, and calculate a reliability score of 75%. The server notifies A of this reliability score and detailed analysis results. Based on this reliability score, A can determine that the news video is relatively reliable.

[0116] Example prompts to be input to the generative AI model

[0117] Enter "Please rate the reliability of the news video. Determine whether this video is a Deep Fake and calculate the reliability score." The generated results include the reliability score and the detection results for each analyzed frame.

[0118] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0119] Step 1:

[0120] A user accesses a web application and selects a media file (video, image, or audio file) to upload. When the user clicks the upload button, the device issues an HTTP POST request to send the selected content to the server. This request includes metadata (file name, file size, file type, etc.) along with the file data. The input is the media file selected by the user, and the output is the HTTP request sent to the server.

[0121] Step 2:

[0122] The server receives the HTTP POST request sent from the user device and extracts the content data. The server temporarily stores the received content in local storage, for example in the " / tmp / uploaded_files" directory. The input is the media file sent in step 1, and the output is the file stored in local storage.

[0123] Step 3:

[0124] The server adds the temporarily saved file to the analysis queue. It registers it as a job in the analysis queue management system (e.g., Redis Queue) and generates a job ID. The registered job waits for subsequent analysis processing. The input is the path of the saved file, and the output is the job ID registered in the analysis queue.

[0125] Step 4:

[0126] The server processes jobs registered in the analysis queue. It uses a deep learning framework (e.g., TensorFlow or PyTorch) to execute the Deep Fake detection algorithm. It performs a detailed analysis of each frame (video / image) or each second (audio) of the content to detect traces of Deep Fake. The input is the job ID and the corresponding file path, and the output is the Deep Fake detection result (the judgment result for each frame or second).

[0127] Step 5:

[0128] The server calculates the percentage of detected Deep fakes and calculates a confidence score. Specifically, it calculates a confidence score in the range of 0 to 100 using the percentage of Deep fakes determined to be Deep fakes for all frames and seconds. For example, if 5% of frames are Deep fakes, the confidence score is calculated as 95%. The input is the Deep fake detection result, and the output is the confidence score.

[0129] Step 6:

[0130] The server returns the reliability score and detailed analysis results to the user. The results are sent to the web application in JSON format and displayed in the user's browser. For example, a message such as "The reliability score of the news video is 75%. Click here for detailed analysis results" is displayed. The input is the reliability score and analysis results, and the output is the results displayed in the user's web browser.

[0131] (Application example 1)

[0132] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0133] In recent years, advances in deepfake technology have increased the risk of content such as video and audio being easily altered. Such alterations could have serious implications for important business decisions and communications. Therefore, verifying the authenticity of content is essential, especially for corporate security departments and journalists. However, traditional methods of manually verifying each piece of content are extremely time-consuming and inefficient. This has created a demand for an automated and reliable deepfake detection system.

[0134] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0135] In this invention, the server includes a means for users to upload content, a means for receiving the uploaded content and adding it to an analysis queue, a means for executing an algorithm to detect Deep Fakes contained in the content, a means for calculating a reliability score for the content based on the detection results, a means for analyzing recordings of corporate audio or video conferences to detect whether Deep Fake technology is used, and a means for returning the reliability score and analysis results to the user. This allows users to automatically verify whether the content they uploaded has been altered using Deep Fake technology and quickly evaluate its reliability.

[0136] "User" means a person or entity that uses the system to upload content and receive the results of that analysis.

[0137] "Content" refers to media files such as video, audio, and images uploaded by users.

[0138] "Upload" is the act of a user sending content they own to the system.

[0139] "Receiving" is the process by which the server receives content sent by the user.

[0140] The "analysis queue" is a collection of content that is waiting to be analyzed.

[0141] "Deep fake" is media content that has been artificially generated or altered using deep learning techniques.

[0142] An "algorithm" is a computational procedure or step for solving a particular problem.

[0143] The "confidence score" is a numerical assessment of the likelihood that content has been altered.

[0144] "Analysis Results" means the analytical output generated by the deepfake detection algorithm.

[0145] "Enterprise" means an entity or organization that conducts business activities.

[0146] An "audio or video conference" is a real-time audio and / or video conference between people in different locations.

[0147] "Deepfake technology" is a technology that uses deep learning technology to alter existing media content.

[0148] "Use" refers to whether a particular technology or method is used.

[0149] "Return" is the process in which the server returns the analysis results and the reliability score to the user.

[0150] To implement this invention, it is necessary to build a system and execute a series of processes for users to upload content and return the analysis results and reliability scores. This process is mainly composed of specific means including a server, a user terminal, and an analysis algorithm.

[0151] First, users use a web application to upload content (video, audio, or images). This web application is built using React and provides a user interface that allows users to intuitively upload content. When the user clicks the upload button, the user's device sends the selected content to the server as an HTTP POST request.

[0152] Next, the server is built using Node.js and Express.js and receives the content sent by the user. The received content is temporarily stored and then added to the analysis queue. This analysis queue keeps the content waiting to be analyzed in order.

[0153] The server then runs a DeepFake detection algorithm on the received content. This algorithm, built using TensorFlow and PyTorch, performs detailed analysis of each frame and each second of data to detect signs of deepfakes. Based on the results of this analysis, a confidence score is calculated. The confidence score is a numerical assessment of the accuracy of the content.

[0154] Finally, the server returns the calculated confidence score and analysis results to the user, displaying the results in the user's web application. For example, a user might upload a recording of a business meeting or presentation and check whether the recording has been altered using deepfake technology. In such a case, the server analyzes each frame of the recording, calculates a confidence score, and notifies the user.

[0155] As a concrete example, suppose a user uploads a recording of a business presentation. The server analyzes the file to detect whether deepfake technology has been used. Based on the detection result, the server calculates a confidence score and returns it to the user, such as "confidence score is 97%."

[0156] An example of a prompt might be:

[0157] Upload and analyze a recording of your business presentation, and use DeepFake technology to detect any alterations and calculate a confidence score.

[0158] This completes the description of the embodiment of the present invention. This system enables users to automatically verify whether uploaded content has been altered using deepfake technology and quickly evaluate its authenticity.

[0159] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0160] Step 1:

[0161] A user uploads content. Using the web application interface, the user selects the video, audio, or image file they want to analyze and clicks the upload button. The input is the media file selected by the user, and the output is the content sent from the user device to the server. Specific actions include selecting a file and clicking the upload button.

[0162] Step 2:

[0163] The server receives the uploaded content. The server receives the content sent as an HTTP POST request and temporarily stores it. The input is the received media file, and the output is that file saved in temporary storage on the server. Specific operations include receiving a file and saving a file.

[0164] Step 3:

[0165] The server adds the received content to the analysis queue. The server registers the received media file as an analysis job and adds it to the analysis queue. The input is a media file stored in temporary storage, and the output is that file added to the analysis queue. Specific operations include registering the file as an analysis job and updating the queue.

[0166] Step 4:

[0167] The server runs the Deep Fake detection algorithm. The server runs the Deep Fake detection algorithm using TensorFlow or PyTorch on the media files registered in the analysis queue. The media files registered in the analysis queue are used as input, and the deep fake detection results are obtained as output. Specific operations include analyzing each frame or each second and detecting traces of deep fakes.

[0168] Step 5:

[0169] The server calculates a confidence score based on the detection results. The server calculates a confidence score for the entire media file based on the percentage of deepfakes detected. The deepfake detection results are the input, and the confidence score is calculated and obtained as the output. Specific operations include calculating the percentage of deepfakes and calculating the confidence score.

[0170] Step 6:

[0171] The server returns the analysis results and confidence score to the user. The server returns the calculated confidence score and detailed analysis results to the user's web application and displays the results. The confidence score and analysis results are input, and this information is displayed to the user as output. Specific operations include returning and displaying the results.

[0172] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.

[0173] To implement this invention, a system is required in which a user uploads content using a web application, a server analyzes the content, detects Deep Fake parts, calculates a reliability score, and returns the result to the user, combined with an emotion engine that recognizes the user's emotions.

[0174] The server and user terminal perform the following main operations:

[0175] 1. User uploads content

[0176] A user accesses the web application, selects the file (video, image, audio file) they want to analyze, and clicks the upload button. This action causes the user's device to make an HTTP POST request to send the selected file to the server.

[0177] 2. The server receives the uploaded content

[0178] The server receives the file sent from the user terminal, temporarily stores the content received by the server, and prepares it for the next analysis process.

[0179] 3. Obtaining user emotion data

[0180] Using a camera and microphone installed on the user's device, emotion data is acquired from the user's facial expressions and voice. The emotion engine analyzes this data and estimates the user's emotional state.

[0181] 4. The server adds it to the analysis queue

[0182] The server adds the received file and the acquired emotion data to the analysis queue and registers it as an analysis job. This operation puts the content in a queue and allows subsequent analysis processing to begin.

[0183] 5. The server runs the Deep Fake detection algorithm

[0184] The server runs the DeepFake detection algorithm on the content queued for analysis. This algorithm performs a detailed analysis of each frame (video / image) or each second (audio) to detect DeepFake signatures.

[0185] 6. The server calculates the reliability score based on the detection result.

[0186] The server calculates the reliability score of the entire content based on the ratio of detected deep fakes, and provides the reliability score to the user as a numerical value for evaluating the accuracy of the content.

[0187] 7. The server stores the analysis results, including emotional data.

[0188] The server stores the completed analysis results, the confidence score, and the acquired emotion data in a database. These are stored in association with the job ID, and can be returned in response to subsequent requests.

[0189] 8. The server returns the analysis results and emotion data to the user.

[0190] The server returns the analysis results, a confidence score, and the user's emotion data to the user, which are then displayed in the user's web application.

[0191] Specific examples

[0192] For example, suppose user "B" wants to check the reliability of a news video. B uploads the video through a web application, and the server receives it and adds it to the analysis queue. At this time, the user's device acquires emotional data from B's facial expressions and voice, and sends this data to the server. The server analyzes each frame of the video to detect whether or not it contains Deep Fakes. The detection results indicate that 10% of the frames are Deep Fakes, with a calculated reliability score of 60%. Furthermore, suppose the emotion engine recognizes that B has expressed concerns. The server returns these analysis results to B, who then evaluates the reliability of the video based on this information. In this way, adding emotional data enables a system that enables more detailed analysis and improves the user experience.

[0193] Through this series of processes, the present invention can provide a system that not only evaluates the reliability of content consumed by a user, but also supports comprehensive judgment, including the user's emotional state.

[0194] The processing flow will be explained below.

[0195] Step 1:

[0196] The user selects the content and performs the upload operation. The user accesses the web application, selects the video, image, or audio file they want to analyze, and clicks the "Upload" button. This operation sends the selected file from the user's device to the server.

[0197] Step 2:

[0198] The user device sends a file to the server via an HTTP POST request, along with metadata, structuring the file in a format that the server can receive.

[0199] Step 3:

[0200] The server receives the file sent from the user terminal, temporarily stores the file, and prepares for the next analysis process.

[0201] Step 4:

[0202] Acquires user emotional data. Emotional data is acquired from the user's facial expressions and voice using a camera and microphone installed on the user's device. The emotion engine analyzes this data and estimates the user's emotional state.

[0203] Step 5:

[0204] The server adds the received file and the acquired emotion data to the analysis queue. At the same time as it registers it as an analysis job, a job ID is generated and used to track the progress of the processing.

[0205] Step 6:

[0206] The server initializes the Deep Fake detection algorithm. In this step, the necessary models and libraries are loaded and the system is ready for analysis.

[0207] Step 7:

[0208] The server splits the file. In the case of video, it is split into frames, and in the case of images and audio, it is split into appropriate units.

[0209] Step 8:

[0210] The server analyzes each unit (frame, pixel, second) using a Deep Fake detection algorithm, which evaluates each unit individually to determine the presence of Deep Fake.

[0211] Step 9:

[0212] The server aggregates the analysis results for all units and calculates the Deep Fake ratio for the entire content based on the detection results for each unit.

[0213] Step 10:

[0214] The server calculates a trust score based on the Deep Fake ratio. The trust score is expressed between 0 and 100, with higher scores indicating more trustworthy content.

[0215] Step 11:

[0216] The server stores the analysis results, confidence score, and acquired emotion data in a database, linked to the job ID, so that the results can be returned in response to subsequent requests.

[0217] Step 12:

[0218] The server notifies the user that the analysis results are ready, and the user can then review the results.

[0219] Step 13:

[0220] The user device sends a request to the server to display the analysis results. The request includes the job ID.

[0221] Step 14:

[0222] The server retrieves the analysis results corresponding to the job ID from the database and returns them to the user device along with the confidence score and emotion data.

[0223] Step 15:

[0224] The user's device displays the received trust score, analysis results, and emotional data. Based on these results, the user evaluates the trustworthiness of the content and their own emotional state. This information helps the user to more accurately judge the trustworthiness of the content.

[0225] Example 2

[0226] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0227] Conventional content analysis systems are specialized in detecting deep fakes and calculating their reliability, and are unable to provide comprehensive analysis results that take into account the user's emotional state. This has resulted in issues such as the inability to fully reflect the emotional impact of users when evaluating the reliability of content. Furthermore, there has been an issue with insufficient storage and management of analysis results, making it difficult for users to easily check past analysis results.

[0228] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0229] In this invention, the server includes means for a user to upload data using an electronic device, means for receiving the uploaded data and adding it to an analysis queue, means for acquiring and analyzing the user's emotional data, means for executing an algorithm for detecting deep fakes contained in the data, means for calculating a reliability score of the data based on the detection results, means for storing the analysis results including the emotional data in a database, and means for returning the reliability score and the analysis results to the user. This enables accurate detection of deep fakes in content uploaded by users and provision of comprehensive analysis results including emotional data. Furthermore, by storing the analysis results and emotional data in a database, users can easily check past analysis results.

[0230] "User" refers to an individual or entity that uses an electronic device to upload data and view analytical results.

[0231] "Electronic Device" refers to a device, such as a computer or smartphone, that a User uses to upload Content.

[0232] "Data" refers to content such as videos, images, and audio files uploaded by users.

[0233] "Means for uploading" refers to the operation performed by a user using an electronic device to send data to a server, and the software that supports that operation.

[0234] "Means for receiving" refers to the hardware and software that the server uses to receive and temporarily store data sent by the user.

[0235] "Analysis queue" refers to a waiting list used to process uploaded data in order as analysis jobs.

[0236] "Emotion data" refers to information that indicates the emotional state of a user analyzed from their facial expressions and voice.

[0237] "Means for acquiring emotional data" refers to hardware and software for collecting and analyzing the user's facial expressions and voice to estimate their emotional state.

[0238] "Deep fake" refers to the technology and products that use deep learning technology to generate realistic-looking fake video and audio.

[0239] "Means for implementing algorithms to detect Deep Fakes" refers to hardware and software that runs deep learning models and other algorithms to identify Deep Fake signatures in data.

[0240] "Reliability Score" refers to a number that indicates the reliability of the data, calculated based on the percentage of deep fakes contained in the data.

[0241] "Means for calculating the reliability score" refers to the hardware and software for calculating the reliability score based on the ratio of Deep fake detection results.

[0242] "Means for storing in a database" refers to a storage system and its management software for permanently storing analysis results and emotion data.

[0243] "Means of return" refers to software and communication means for displaying or notifying the user of the analysis results, reliability score, and emotional data.

[0244] MODE FOR CARRYING OUT THE INVENTION

[0245] This invention relates to a system in which a user uploads data using an electronic device, a server analyzes the data to detect Deep Fakes and calculate a reliability score, and also obtains the user's emotional state and returns a comprehensive analysis result.

[0246] The system includes the following major hardware and software:

[0247] User terminal

[0248] The user terminal is an electronic device such as a computer or smartphone that provides an interface for users to upload data. The user terminal is equipped with a camera and microphone, which can capture the user's facial expression and voice data. This data is analyzed by the emotion engine to identify the user's emotional state.

[0249] server

[0250] The server receives data sent by users and has a storage system for temporarily storing it. The received data is added to an analysis queue and registered as an analysis job.

[0251] Deep fake detection algorithm: The server implements a deep learning model to detect deep fakes in the data. This algorithm uses deep learning libraries such as TensorFlow and PyTorch.

[0252] Data analysis: The server uses ffmpeg to split the video into frames and then runs the Deep Fake detection process on each frame.

[0253] Calculation of confidence score: Based on the Deep Fake detection results, a score is calculated to evaluate the reliability of the entire data. This is provided as a confidence score ranging from 0% to 100%.

[0254] Database

[0255] The server is equipped with a database for permanently storing analysis results and user emotion data, making it possible to easily track past analysis results and display them again upon user request.

[0256] Data return

[0257] Once the analysis is complete, the server returns the confidence score, analysis results, and emotion data to the user, which are then displayed on the screen of the user's electronic device.

[0258] Specific examples

[0259] For example, consider a situation where user "B" wants to check the authenticity of a news video. User B accesses a web application, selects the news video file, and clicks the upload button. This action sends the video file from User B's computer to the server via an HTTP POST request.

[0260] The server receives the video file and temporarily stores it in storage. During this time, Person B's device uses the camera and microphone to capture Person B's facial expressions and voice data. The emotion engine analyzes this data and identifies Person B's emotional state (e.g., concern).

[0261] The server adds the received video file and the acquired emotion data to the analysis queue and runs the Deep Fake detection algorithm. Specifically, it divides the video into frames using "ffmpeg" and analyzes each frame using "TensorFlow." This analysis determines that 10% of the video is Deep Fake, and the confidence score is calculated as 90%.

[0262] The analysis results and emotion data are stored in a database along with the job ID. Finally, the server returns the analysis results, confidence score, and emotion data to Person B. Person B can check the analysis results, including the "concern" emotion state, along with the confidence score, in a web application.

[0263] Prompt Sentence Examples

[0264] "Please check the authenticity of this video and tell us how much Deepfake evidence there is."

[0265] Please send the Deepfake analysis results for the uploaded image along with the emotion data.

[0266] "Evaluate the reliability of the audio file and return the result. Emotional data on whether the user is concerned is also important."

[0267] This system allows users to accurately assess the reliability of the content they consume and provides comprehensive analysis based on emotional data.

[0268] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0269] Step 1:

[0270] Users upload content

[0271] Input: Data such as video, images, and audio files

[0272] How it works: A user logs into the web application, selects the file they want to analyze, and clicks the "Upload" button, which causes the user's device to send an HTTP POST request to the server with the selected file.

[0273] Output: Upload request sent to the server

[0274] Step 2:

[0275] The server receives the uploaded content

[0276] Input: HTTP POST request sent from the user terminal

[0277] Operation: The server receives the HTTP request and temporarily saves the sent file on the local disk. As an example, save the file at " / tmp / uploaded_files / video.mp4".

[0278] Output: Temporarily saved video file

[0279] Step 3:

[0280] Obtain the user's sentiment data<°

[0281] Input: The user's facial expression and voice data

[0282] Operation: Use the camera and microphone of the user terminal to obtain the user's facial expression and voice in real time. Pass these data to the sentiment engine, and the sentiment engine analyzes and estimates the user's sentiment state. For example, determine "suspicion" from a "frowning expression".

[0283] Output: Obtained sentiment data

[0284] Step 4:

[0285] The server adds to the analysis queue

[0286] Input: Temporarily saved video file, obtained sentiment data

[0287] Operation: The server adds the received file and emotion data to the analysis queue and registers it as an analysis job. Specifically, it registers the file path " / tmp / uploaded_files / video.mp4" and the emotion data "concern" in the analysis queue.

[0288] Output: Jobs waiting to be analyzed

[0289] Step 5:

[0290] The server runs the Deep Fake detection algorithm

[0291] Input: Video files queued for analysis

[0292] How it works: The server splits the video into frames using ffmpeg and runs the Deep Fake detection algorithm on each frame. It then uses a deep learning model (e.g. TensorFlow) to determine whether or not there is a Deep Fake in each frame.

[0293] Output: Deep fake detection results for each frame

[0294] Step 6:

[0295] The server calculates the reliability score based on the detection result.

[0296] Input: Deep fake detection results for each frame

[0297] Operation: The detection results are aggregated and the overall confidence score is calculated based on the percentage of Deep fake detections. Specifically, if the Deep fake portion accounts for 10% of the total, a "confidence score of 90%" is calculated.

[0298] Output: Confidence Score

[0299] Step 7:

[0300] The server stores the analysis results, including emotional data.

[0301] Input: confidence score, acquired emotion data

[0302] Operation: The server saves the analysis results and emotion data in the database. It associates them with the job ID and saves "Video.mp4", "Confidence Score: 90%", and "Emotional State: Concern".

[0303] Output: Analysis results and emotion data stored in a database

[0304] Step 8:

[0305] The server returns the analysis results and emotion data to the user.

[0306] Input: Analysis results and emotion data stored in the database

[0307] Behavior: The server generates a report of the analysis results and returns it to the user. The user's web application displays the result "Video Confidence Score: 90%, Emotional State: Concern".

[0308] Output: Analysis results and emotion data displayed on the user's electronic device

[0309] (Application example 2)

[0310] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0311] Traditional deep fake detection systems focus on assessing the trustworthiness of content uploaded by users. However, these systems provide analysis results without considering the user's emotional state, and therefore cannot provide information about how users understand and perceive the trustworthiness of content. In particular, when the results of deep fake detection are ambiguous or have low confidence, this can burden users' understanding and behavior. To solve this problem and improve user experience, it is necessary to incorporate user emotional data into the analysis.

[0312] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[0313] In this invention, the server includes a means for users to upload content, a means for receiving the uploaded content and adding it to an analysis queue, a means for executing an algorithm to detect Deep Fakes contained in the content, a means for calculating a reliability score for the content based on the detection results, a means for executing an engine that acquires and analyzes emotional data from the user terminal, and a means for providing the user with feedback based on the emotional data along with the reliability score. This allows the user to receive feedback based on their own emotional state in addition to the reliability of the content, allowing the user to make decisions based on more comprehensive information.

[0314] The "means by which users upload content" refers to an interface that users use to send data such as video, image, and audio files to be analyzed to a server via the Internet.

[0315] The "means for receiving uploaded content and adding it to the analysis queue" is a function by which the server receives content sent by a user and adds the content to a waiting list for analysis processing.

[0316] "Means for executing an algorithm to detect Deep Fakes contained in content" refers to a program that allows the server to perform detailed analysis of each frame (video / image) or each second (audio) to identify traces of Deep Fakes.

[0317] "Means for calculating the reliability score of content based on the detection results" is a function that allows the server to use the detection results of Deep Fake to calculate a numerical value (reliability score) to evaluate the reliability of the entire content.

[0318] "Means for executing an engine that acquires and analyzes emotional data from a user terminal" refers to a program that acquires data using a camera or microphone and performs emotional analysis in order to estimate the emotional state of the user from their facial expressions and voice.

[0319] "Means for providing users with feedback based on emotional data together with the reliability score" is a function in which the server combines the reliability score with the user's emotional data to provide the analysis results to the user, providing feedback in a form that is easier for the user to understand.

[0320] This invention is a system that not only evaluates the reliability of content but also takes into account the emotional state of the user to support more comprehensive judgment. This system is mainly composed of a server and a user terminal.

[0321] The server and user terminal perform the following main operations:

[0322] First, the user device provides an interface for uploading the content (video, image, audio files) that the user wants to analyze. The user selects the file through a web application or a dedicated app and clicks the upload button. The user device then makes an HTTP POST request to send the selected file to the server.

[0323] The server then receives and temporarily stores the file sent from the user terminal, and adds the received content to the analysis queue, ready for analysis processing.

[0324] Furthermore, the emotion engine uses the camera and microphone installed on the user device to acquire emotion data from the user's facial expressions and voice, and analyzes this data to estimate the user's emotional state.

[0325] The server adds the uploaded file and the acquired emotion data to the analysis queue and registers it as an analysis job, which starts the subsequent analysis process.

[0326] The server runs the DeepFake detection algorithm on the content queued for analysis. This algorithm performs detailed analysis of each frame (video / image) and each second (audio) to detect traces of DeepFake.

[0327] The server calculates the reliability score for the entire content based on the Deep Fake detection results. For example, if 10% of the frames are determined to be Deep Fake, the reliability score is calculated as 60%. Furthermore, if the emotion engine recognizes the user's emotional state (happiness, sadness, anger, anxiety, etc.), that information is also added as feedback.

[0328] The analysis results include the reliability score of the content and the user's emotional data, which are stored in a database. The server returns these results in response to a user request. The results are then displayed, for example, on the screen of a web application.

[0329] A specific example would be analyzing images captured by a user during a meeting or video call. The user uploads the image to be analyzed, and the server analyzes the image. During the analysis, the user's facial expression is captured by a camera, and the emotional data is included in the analysis results. As a result, the emotional data "the user is concerned" is provided along with the confidence score of the image.

[0330] The following software and hardware are used to configure the emotion engine and Deep Fake detection algorithm functionality:

[0331] Open Source Computer Vision Library: OpenCV

[0332] Sentiment analysis engine: EmotionEngine

[0333] Deep fake detection algorithm API

[0334] This allows users to receive detailed information about the reliability of the content and feedback that includes their own emotional state, allowing them to make decisions based on more comprehensive information.

[0335] An example of a prompt to input to a generative AI model is as follows:

[0336] "Please confirm whether this image or video frame could have been generated by AI. Please return the analysis results and a confidence score. Below is the image data for verification."

[0337] "Please estimate the user's emotional state (happiness, sadness, anger, anxiety, etc.) from this image. Please return the analysis result and the emotional state. Below is the image data for testing."

[0338] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0339] Step 1: User uploads content

[0340] The user selects the content (video, image, audio file) they want to analyze through a web application or dedicated app and clicks the upload button. The user's device sends the selected file data to the server as an HTTP POST request. The input is the file data, and the output is the status of completion of transmission to the server.

[0341] Step 2: The server receives the uploaded content

[0342] The server receives files sent from the user terminal and temporarily stores them. The input is the file data sent from the user terminal, and the output is the storage path of the temporarily stored file. The specific operation is to receive the file and store it in the specified directory.

[0343] Step 3: Obtain user emotion data

[0344] The camera and microphone installed on the user device capture the user's facial expressions and voice, and analyzes the data. The emotion engine estimates the user's emotional state (happiness, sadness, anger, anxiety, etc.). The input is the captured image and voice data, and the output is numerical data related to the emotional state. Specific operations include activating the camera and microphone, capturing data, and analyzing the obtained data.

[0345] Step 4: Server adds to analysis queue

[0346] The server adds the received file and the acquired emotion data to the analysis queue and registers it as an analysis job. The input is the file data and emotion data, and the output is the status of completion of adding it to the analysis queue. The specific operation is to add the file and emotion data to the analysis queue.

[0347] Step 5: The server runs the Deep Fake detection algorithm

[0348] The server runs the DeepFake detection algorithm on the content queued for analysis. This algorithm performs detailed analysis of each frame (video / image) and each second (audio) to detect DeepFake signatures. The input is the file data to be analyzed, and the output is the DeepFake detection results and a confidence score. Specific operations include analyzing each frame and each second individually to detect specific patterns and features.

[0349] Step 6: The server calculates the trust score

[0350] The server calculates the reliability score of the entire content based on the detection results of Deep Fake. The input is the detection result data of Deep Fake, and the output is a numerical reliability score. Specific operations include executing an algorithm that sums up the detection results and evaluates the overall reliability.

[0351] Step 7: The server stores the analysis results and emotion data

[0352] The server stores the analysis results, confidence scores, and the acquired emotion data in a database. The input is the analysis result data, confidence score data, and emotion data, and the output is the status of completion of saving to the database. Specifically, each dataset is linked to a job ID and stored in the database.

[0353] Step 8: The server returns the analysis results and emotion data to the user.

[0354] The server returns the analysis results, confidence score, and user emotion data to the user. The input is the analysis result data, confidence score data, and emotion data, and the output is data to be displayed to the user. Specific operations include data format conversion and transmission to display the results on the screen of the web application used by the user.

[0355] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0356] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0357] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.

[0358] [Second embodiment]

[0359] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.

[0360] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0361] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0362] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.

[0363] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0364] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0365] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0366] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0367] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0368] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0369] In the smart glasses 214, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0370] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."

[0371] To implement the present invention, it is necessary to build a system in which a user uploads content using a web application, a server analyzes the content, detects Deep Fake parts, calculates a reliability score, and returns the result to the user.

[0372] The server performs the following main operations:

[0373] 1. User uploads content

[0374] A user accesses the web application, selects the video, image, or audio file they want to analyze, and clicks the upload button, which causes the user's device to issue an HTTP POST request to send the selected content to the server.

[0375] 2. The server receives the uploaded content

[0376] The server receives the content sent from the user terminal, temporarily stores the content received by the server, and is ready for the next analysis step.

[0377] 3. The server adds it to the analysis queue

[0378] The server adds the received content to the analysis queue and registers it as an analysis job. This operation puts the content in a queue and allows subsequent analysis processing to begin.

[0379] 4. The server runs the Deep Fake detection algorithm

[0380] The server runs the DeepFake detection algorithm on the content queued for analysis. This algorithm performs a detailed analysis of each frame (video / image) or each second (audio) to detect DeepFake signatures.

[0381] 5. The server calculates the reliability score based on the detection result.

[0382] The server calculates the reliability score of the entire content based on the ratio of detected deep fakes. The reliability score is provided to the user as a numerical value for evaluating the accuracy of the content.

[0383] 6. The server returns the analysis results and the reliability score to the user.

[0384] The server returns the completed analysis results and the confidence score to the user, which are then displayed in the user's web application.

[0385] Specific examples

[0386] For example, suppose user "A" wants to check the reliability of a certain news video. A uploads the video through a web application, and the server receives the video and adds it to the analysis queue. The server analyzes each frame of the video to detect whether or not it contains Deep fakes. The detection results indicate that 5% of the frames are Deep fakes, and calculate a reliability score of 75%. The server notifies A of this reliability score and detailed analysis results. Based on this reliability score, A can determine that the news video is relatively reliable.

[0387] In this way, the present invention provides a system that allows users to easily evaluate the reliability of the content they consume and assists them in making accurate information selections.

[0388] The processing flow will be explained below.

[0389] Step 1:

[0390] The user selects the content and performs the upload operation. The user accesses the web application, selects the file they want to analyze, and clicks the "Upload" button. This operation sends the selected file from the user's device to the server.

[0391] Step 2:

[0392] The user device sends a file to the server via an HTTP POST request, along with metadata, structuring the file in a format that the server can receive.

[0393] Step 3:

[0394] The server receives the file sent from the user terminal, temporarily stores the file, and prepares it for the next analysis process.

[0395] Step 4:

[0396] The server adds the received file to the analysis queue. When it is registered as an analysis job, a job ID is generated and used to track the processing progress.

[0397] Step 5:

[0398] The server initializes the Deep Fake detection algorithm. In this step, the necessary models and libraries are loaded and the system is ready for analysis.

[0399] Step 6:

[0400] The server splits the file. In the case of video, it is split into frames, and in the case of images and audio, it is split into appropriate units (pixels, seconds).

[0401] Step 7:

[0402] The server analyzes each unit (frame, pixel, second) using a Deep Fake detection algorithm, which evaluates each unit individually to determine the presence of Deep Fake.

[0403] Step 8:

[0404] The server aggregates the analysis results for all units and calculates the Deep Fake ratio for the entire content based on the detection results for each unit.

[0405] Step 9:

[0406] The server calculates a trust score based on the Deep Fake ratio. The trust score is expressed between 0 and 100, with higher scores indicating more trustworthy content.

[0407] Step 10:

[0408] The server saves the analysis results and the confidence score in a database. The results are saved as results associated with the job ID, and can be returned in response to a user request.

[0409] Step 11:

[0410] The server notifies the user that the analysis results are ready, and the user can then review the results.

[0411] Step 12:

[0412] The user terminal sends a request to the server to display the analysis results. The request includes the job ID.

[0413] Step 13:

[0414] The server retrieves the analysis results corresponding to the job ID from the database and returns them to the user device along with the reliability score.

[0415] Step 14:

[0416] The user's device displays the received reliability score and analysis results, and the user evaluates the reliability of the content based on these results.

[0417] Example 1

[0418] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0419] In recent years, the development of deep fake technology has increased the risk of inaccurate information spreading. As a result, it has become more difficult for users to determine trustworthy content. Therefore, there is a need for technology that can quickly and accurately evaluate the reliability of content uploaded by users. Conventional technologies require a lot of time to analyze content, and the results are not always highly reliable. The present invention aims to solve this problem.

[0420] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0421] In this invention, the server includes a means for users to upload content, a means for receiving the uploaded content and adding it to an analysis queue, a means for executing an algorithm to detect Deep Fakes contained in the content, a means for calculating a reliability score of the content based on the detection result, a means for returning the reliability score and the analysis result to the user, and a means for displaying the analysis result, thereby enabling the user to quickly and accurately evaluate the reliability of the content.

[0422] A "user" is a person or entity that uses the system to upload content and receive the results of that analysis.

[0423] "Content" refers to the media data to be analyzed, such as videos, images, and audio files.

[0424] "Uploading means" means the functionality or interface that allows a user to submit content to the system.

[0425] "Server" means a computer system that receives, analyzes, and processes content uploaded by users.

[0426] "Analysis Queue" refers to a queue that registers and manages uploaded content for analysis in order.

[0427] "Deep fake" is fake media content generated using deep learning technology.

[0428] "Means for executing algorithms" refers to the ability to execute programs or processes to detect Deep Fakes in content.

[0429] The "Reliability Score" is a numerical value that evaluates the reliability of content, calculated based on the percentage of detected Deep Fakes.

[0430] "Analysis results" refers to the information and data obtained after the server runs the Deep Fake detection algorithm.

[0431] The "means for returning" is a function that allows the server to provide the analysis results and the reliability score to the user.

[0432] "Display means" refers to an interface or method for visually presenting the analysis results and the reliability score to the user.

[0433] A "database" is an electronic data storage system that stores analysis results and allows for searching and referencing as needed.

[0434] A "request" refers to an operation or action in which a user requests information such as analysis results from a system.

[0435] MODE FOR CARRYING OUT THE INVENTION

[0436] In the system based on this invention, a user uploads content using a web application, the server analyzes the content, detects Deep Fake parts, calculates a reliability score, and returns the result to the user. A specific embodiment of this system will be described below.

[0437] System Overview

[0438] Users can access the web application and select and upload video, image, or audio files they want to analyze. The user device issues an HTTP POST request to send the selected content to the server. The server temporarily stores the received content and then adds it to the analysis queue. The server runs the Deep Fake detection algorithm on the content in the analysis queue and calculates a confidence score based on the detection results. Finally, the server returns the analysis results and confidence score to the user and displays them in the web application.

[0439] Hardware and Software Use

[0440] server:

[0441] The server can be a cloud-based virtual machine with high-performance computing resources, such as an EC2 instance from Amazon Web Services (AWS) or a Compute Engine from Google Cloud Platform (GCP).

[0442] The server uses MySQL or PostgreSQL as a database to efficiently store and manage analysis results.

[0443] Deep Learning Frameworks:

[0444] Deep fake detection uses deep learning frameworks such as TensorFlow and PyTorch, which enable advanced image and video analysis.

[0445] User device:

[0446] A user terminal is a device that can connect to the Internet, such as a PC, smartphone, or tablet, and accesses web applications through a web browser (e.g., Google Chrome or Mozilla Firefox).

[0447] Specific examples

[0448] For example, suppose user "A" wants to check the reliability of a certain news video. A uploads the video through a web application, and the server receives the video and adds it to the analysis queue. The server analyzes each frame of the video to detect whether or not it contains Deep fakes. The detection results indicate that 5% of the frames are Deep fakes, and calculate a reliability score of 75%. The server notifies A of this reliability score and detailed analysis results. Based on this reliability score, A can determine that the news video is relatively reliable.

[0449] Example prompts to be input to the generative AI model

[0450] Enter "Please rate the reliability of the news video. Determine whether this video is a Deep Fake and calculate the reliability score." The generated results include the reliability score and the detection results for each analyzed frame.

[0451] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0452] Step 1:

[0453] A user accesses a web application and selects a media file (video, image, or audio file) to upload. When the user clicks the upload button, the device issues an HTTP POST request to send the selected content to the server. This request includes metadata (file name, file size, file type, etc.) along with the file data. The input is the media file selected by the user, and the output is the HTTP request sent to the server.

[0454] Step 2:

[0455] The server receives the HTTP POST request sent from the user device and extracts the content data. The server temporarily stores the received content in local storage, for example in the " / tmp / uploaded_files" directory. The input is the media file sent in step 1, and the output is the file stored in local storage.

[0456] Step 3:

[0457] The server adds the temporarily saved file to the analysis queue. It registers it as a job in the analysis queue management system (e.g., Redis Queue) and generates a job ID. The registered job waits for subsequent analysis processing. The input is the path of the saved file, and the output is the job ID registered in the analysis queue.

[0458] Step 4:

[0459] The server processes jobs registered in the analysis queue. It uses a deep learning framework (e.g., TensorFlow or PyTorch) to execute the Deep Fake detection algorithm. It performs a detailed analysis of each frame (video / image) or each second (audio) of the content to detect traces of Deep Fake. The input is the job ID and the corresponding file path, and the output is the Deep Fake detection result (the judgment result for each frame or second).

[0460] Step 5:

[0461] The server calculates the percentage of detected Deep fakes and calculates a confidence score. Specifically, it calculates a confidence score in the range of 0 to 100 using the percentage of Deep fakes determined to be Deep fakes for all frames and seconds. For example, if 5% of frames are Deep fakes, the confidence score is calculated as 95%. The input is the Deep fake detection result, and the output is the confidence score.

[0462] Step 6:

[0463] The server returns the reliability score and detailed analysis results to the user. The results are sent to the web application in JSON format and displayed in the user's browser. For example, a message such as "The reliability score of the news video is 75%. Click here for detailed analysis results" is displayed. The input is the reliability score and analysis results, and the output is the results displayed in the user's web browser.

[0464] (Application example 1)

[0465] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0466] In recent years, advances in deepfake technology have increased the risk of content such as video and audio being easily altered. Such alterations could have serious implications for important business decisions and communications. Therefore, verifying the authenticity of content is essential, especially for corporate security departments and journalists. However, traditional methods of manually verifying each piece of content are extremely time-consuming and inefficient. This has created a demand for an automated and reliable deepfake detection system.

[0467] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0468] In this invention, the server includes a means for users to upload content, a means for receiving the uploaded content and adding it to an analysis queue, a means for executing an algorithm to detect Deep Fakes contained in the content, a means for calculating a reliability score for the content based on the detection results, a means for analyzing recordings of corporate audio or video conferences to detect whether Deep Fake technology is used, and a means for returning the reliability score and analysis results to the user. This allows users to automatically verify whether the content they uploaded has been altered using Deep Fake technology and quickly evaluate its reliability.

[0469] "User" means a person or entity that uses the system to upload content and receive the results of that analysis.

[0470] "Content" refers to media files such as video, audio, and images uploaded by users.

[0471] "Upload" is the act of a user sending content they own to the system.

[0472] "Receiving" is the process by which the server receives content sent by the user.

[0473] The "analysis queue" is a collection of content that is waiting to be analyzed.

[0474] "Deep fake" is media content that has been artificially generated or altered using deep learning techniques.

[0475] An "algorithm" is a computational procedure or step for solving a particular problem.

[0476] The "confidence score" is a numerical assessment of the likelihood that content has been altered.

[0477] "Analysis Results" means the analytical output generated by the deepfake detection algorithm.

[0478] "Enterprise" means an entity or organization that conducts business activities.

[0479] An "audio or video conference" is a real-time audio and / or video conference between people in different locations.

[0480] "Deepfake technology" is a technology that uses deep learning technology to alter existing media content.

[0481] "Use" refers to whether a particular technology or method is used.

[0482] "Return" is the process in which the server returns the analysis results and the reliability score to the user.

[0483] To implement this invention, it is necessary to build a system and execute a series of processes for users to upload content and return the analysis results and reliability scores. This process is mainly composed of specific means including a server, a user terminal, and an analysis algorithm.

[0484] First, users use a web application to upload content (video, audio, or images). This web application is built using React and provides a user interface that allows users to intuitively upload content. When the user clicks the upload button, the user's device sends the selected content to the server as an HTTP POST request.

[0485] Next, the server is built using Node.js and Express.js and receives the content sent by the user. The received content is temporarily stored and then added to the analysis queue. This analysis queue keeps the content waiting to be analyzed in order.

[0486] The server then runs a DeepFake detection algorithm on the received content. This algorithm, built using TensorFlow and PyTorch, performs detailed analysis of each frame and each second of data to detect signs of deepfakes. Based on the results of this analysis, a confidence score is calculated. The confidence score is a numerical assessment of the accuracy of the content.

[0487] Finally, the server returns the calculated confidence score and analysis results to the user, displaying the results in the user's web application. For example, a user might upload a recording of a business meeting or presentation and check whether the recording has been altered using deepfake technology. In such a case, the server analyzes each frame of the recording, calculates a confidence score, and notifies the user.

[0488] As a concrete example, suppose a user uploads a recording of a business presentation. The server analyzes the file to detect whether deepfake technology has been used. Based on the detection result, the server calculates a confidence score and returns it to the user, such as "confidence score is 97%."

[0489] An example of a prompt might be:

[0490] Upload and analyze a recording of your business presentation, and use DeepFake technology to detect any alterations and calculate a confidence score.

[0491] This completes the description of the embodiment of the present invention. This system enables users to automatically verify whether uploaded content has been altered using deepfake technology and quickly evaluate its authenticity.

[0492] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0493] Step 1:

[0494] A user uploads content. Using the web application interface, the user selects the video, audio, or image file they want to analyze and clicks the upload button. The input is the media file selected by the user, and the output is the content sent from the user device to the server. Specific actions include selecting a file and clicking the upload button.

[0495] Step 2:

[0496] The server receives the uploaded content. The server receives the content sent as an HTTP POST request and temporarily stores it. The input is the received media file, and the output is that file saved in temporary storage on the server. Specific operations include receiving a file and saving a file.

[0497] Step 3:

[0498] The server adds the received content to the analysis queue. The server registers the received media file as an analysis job and adds it to the analysis queue. The input is a media file stored in temporary storage, and the output is that file added to the analysis queue. Specific operations include registering the file as an analysis job and updating the queue.

[0499] Step 4:

[0500] The server runs the Deep Fake detection algorithm. The server runs the Deep Fake detection algorithm using TensorFlow or PyTorch on the media files registered in the analysis queue. The media files registered in the analysis queue are used as input, and the deep fake detection results are obtained as output. Specific operations include analyzing each frame or each second and detecting traces of deep fakes.

[0501] Step 5:

[0502] The server calculates a confidence score based on the detection results. The server calculates a confidence score for the entire media file based on the percentage of deepfakes detected. The deepfake detection results are the input, and the confidence score is calculated and obtained as the output. Specific operations include calculating the percentage of deepfakes and calculating the confidence score.

[0503] Step 6:

[0504] The server returns the analysis results and confidence score to the user. The server returns the calculated confidence score and detailed analysis results to the user's web application and displays the results. The confidence score and analysis results are input, and this information is displayed to the user as output. Specific operations include returning and displaying the results.

[0505] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0506] To implement this invention, a system is required in which a user uploads content using a web application, a server analyzes the content, detects Deep Fake parts, calculates a reliability score, and returns the result to the user, combined with an emotion engine that recognizes the user's emotions.

[0507] The server and user terminal perform the following main operations:

[0508] 1. User uploads content

[0509] A user accesses the web application, selects the file (video, image, audio file) they want to analyze, and clicks the upload button. This action causes the user's device to make an HTTP POST request to send the selected file to the server.

[0510] 2. The server receives the uploaded content

[0511] The server receives the file sent from the user terminal, temporarily stores the content received by the server, and prepares it for the next analysis process.

[0512] 3. Obtaining user emotion data

[0513] Using a camera and microphone installed on the user's device, emotion data is acquired from the user's facial expressions and voice. The emotion engine analyzes this data and estimates the user's emotional state.

[0514] 4. The server adds it to the analysis queue

[0515] The server adds the received file and the acquired emotion data to the analysis queue and registers it as an analysis job. This operation puts the content in a queue and allows subsequent analysis processing to begin.

[0516] 5. The server runs the Deep Fake detection algorithm

[0517] The server runs the DeepFake detection algorithm on the content queued for analysis. This algorithm performs a detailed analysis of each frame (video / image) or each second (audio) to detect DeepFake signatures.

[0518] 6. The server calculates the reliability score based on the detection result.

[0519] The server calculates the reliability score of the entire content based on the ratio of detected deep fakes, and provides the reliability score to the user as a numerical value for evaluating the accuracy of the content.

[0520] 7. The server stores the analysis results, including emotional data.

[0521] The server stores the completed analysis results, the confidence score, and the acquired emotion data in a database. These are stored in association with the job ID, and can be returned in response to subsequent requests.

[0522] 8. The server returns the analysis results and emotion data to the user.

[0523] The server returns the analysis results, a confidence score, and the user's emotion data to the user, which are then displayed in the user's web application.

[0524] Specific examples

[0525] For example, suppose user "B" wants to check the reliability of a news video. B uploads the video through a web application, and the server receives it and adds it to the analysis queue. At this time, the user's device acquires emotional data from B's facial expressions and voice, and sends this data to the server. The server analyzes each frame of the video to detect whether or not it contains Deep Fakes. The detection results indicate that 10% of the frames are Deep Fakes, with a calculated reliability score of 60%. Furthermore, suppose the emotion engine recognizes that B has expressed concerns. The server returns these analysis results to B, who then evaluates the reliability of the video based on this information. In this way, adding emotional data enables a system that enables more detailed analysis and improves the user experience.

[0526] Through this series of processes, the present invention can provide a system that not only evaluates the reliability of content consumed by a user, but also supports comprehensive judgment, including the user's emotional state.

[0527] The processing flow will be explained below.

[0528] Step 1:

[0529] The user selects the content and performs the upload operation. The user accesses the web application, selects the video, image, or audio file they want to analyze, and clicks the "Upload" button. This operation sends the selected file from the user's device to the server.

[0530] Step 2:

[0531] The user device sends a file to the server via an HTTP POST request, along with metadata, structuring the file in a format that the server can receive.

[0532] Step 3:

[0533] The server receives the file sent from the user terminal, temporarily stores the file, and prepares for the next analysis process.

[0534] Step 4:

[0535] Acquires user emotional data. Emotional data is acquired from the user's facial expressions and voice using a camera and microphone installed on the user's device. The emotion engine analyzes this data and estimates the user's emotional state.

[0536] Step 5:

[0537] The server adds the received file and the acquired emotion data to the analysis queue. At the same time as it registers it as an analysis job, a job ID is generated and used to track the progress of the processing.

[0538] Step 6:

[0539] The server initializes the Deep Fake detection algorithm. In this step, the necessary models and libraries are loaded and the system is ready for analysis.

[0540] Step 7:

[0541] The server splits the file. In the case of video, it is split into frames, and in the case of images and audio, it is split into appropriate units.

[0542] Step 8:

[0543] The server analyzes each unit (frame, pixel, second) using a Deep Fake detection algorithm, which evaluates each unit individually to determine the presence of Deep Fake.

[0544] Step 9:

[0545] The server aggregates the analysis results for all units and calculates the Deep Fake ratio for the entire content based on the detection results for each unit.

[0546] Step 10:

[0547] The server calculates a trust score based on the Deep Fake ratio. The trust score is expressed between 0 and 100, with higher scores indicating more trustworthy content.

[0548] Step 11:

[0549] The server stores the analysis results, confidence score, and acquired emotion data in a database, linked to the job ID, so that the results can be returned in response to subsequent requests.

[0550] Step 12:

[0551] The server notifies the user that the analysis results are ready, and the user can then review the results.

[0552] Step 13:

[0553] The user device sends a request to the server to display the analysis results. The request includes the job ID.

[0554] Step 14:

[0555] The server retrieves the analysis results corresponding to the job ID from the database and returns them to the user device along with the confidence score and emotion data.

[0556] Step 15:

[0557] The user's device displays the received trust score, analysis results, and emotional data. Based on these results, the user evaluates the trustworthiness of the content and their own emotional state. This information helps the user to more accurately judge the trustworthiness of the content.

[0558] Example 2

[0559] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0560] Conventional content analysis systems are specialized in detecting deep fakes and calculating their reliability, and are unable to provide comprehensive analysis results that take into account the user's emotional state. This has resulted in issues such as the inability to fully reflect the emotional impact of users when evaluating the reliability of content. Furthermore, there has been an issue with insufficient storage and management of analysis results, making it difficult for users to easily check past analysis results.

[0561] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0562] In this invention, the server includes means for a user to upload data using an electronic device, means for receiving the uploaded data and adding it to an analysis queue, means for acquiring and analyzing the user's emotional data, means for executing an algorithm for detecting deep fakes contained in the data, means for calculating a reliability score of the data based on the detection results, means for storing the analysis results including the emotional data in a database, and means for returning the reliability score and the analysis results to the user. This enables accurate detection of deep fakes in content uploaded by users and provision of comprehensive analysis results including emotional data. Furthermore, by storing the analysis results and emotional data in a database, users can easily check past analysis results.

[0563] "User" refers to an individual or entity that uses an electronic device to upload data and view analytical results.

[0564] "Electronic Device" refers to a device, such as a computer or smartphone, that a User uses to upload Content.

[0565] "Data" refers to content such as videos, images, and audio files uploaded by users.

[0566] "Means for uploading" refers to the operation performed by a user using an electronic device to send data to a server, and the software that supports that operation.

[0567] "Means for receiving" refers to the hardware and software that the server uses to receive and temporarily store data sent by the user.

[0568] "Analysis queue" refers to a waiting list used to process uploaded data in order as analysis jobs.

[0569] "Emotion data" refers to information that indicates the emotional state of a user analyzed from their facial expressions and voice.

[0570] "Means for acquiring emotional data" refers to hardware and software for collecting and analyzing the user's facial expressions and voice to estimate their emotional state.

[0571] "Deep fake" refers to the technology and products that use deep learning technology to generate realistic-looking fake video and audio.

[0572] "Means for implementing algorithms to detect Deep Fakes" refers to hardware and software that runs deep learning models and other algorithms to identify Deep Fake signatures in data.

[0573] "Reliability Score" refers to a number that indicates the reliability of the data, calculated based on the percentage of deep fakes contained in the data.

[0574] "Means for calculating the reliability score" refers to the hardware and software for calculating the reliability score based on the ratio of Deep fake detection results.

[0575] "Means for storing in a database" refers to a storage system and its management software for permanently storing analysis results and emotion data.

[0576] "Means of return" refers to software and communication means for displaying or notifying the user of the analysis results, reliability score, and emotional data.

[0577] MODE FOR CARRYING OUT THE INVENTION

[0578] This invention relates to a system in which a user uploads data using an electronic device, a server analyzes the data to detect Deep Fakes and calculate a reliability score, and also obtains the user's emotional state and returns a comprehensive analysis result.

[0579] The system includes the following major hardware and software:

[0580] User terminal

[0581] The user terminal is an electronic device such as a computer or smartphone that provides an interface for users to upload data. The user terminal is equipped with a camera and microphone, which can capture the user's facial expression and voice data. This data is analyzed by the emotion engine to identify the user's emotional state.

[0582] server

[0583] The server receives data sent by users and has a storage system for temporarily storing it. The received data is added to an analysis queue and registered as an analysis job.

[0584] Deep fake detection algorithm: The server implements a deep learning model to detect deep fakes in the data. This algorithm uses deep learning libraries such as TensorFlow and PyTorch.

[0585] Data analysis: The server uses ffmpeg to split the video into frames and then runs the Deep Fake detection process on each frame.

[0586] Calculation of confidence score: Based on the Deep Fake detection results, a score is calculated to evaluate the reliability of the entire data. This is provided as a confidence score ranging from 0% to 100%.

[0587] Database

[0588] The server is equipped with a database for permanently storing analysis results and user emotion data, making it possible to easily track past analysis results and display them again upon user request.

[0589] Data return

[0590] Once the analysis is complete, the server returns the confidence score, analysis results, and emotion data to the user, which are then displayed on the screen of the user's electronic device.

[0591] Specific examples

[0592] For example, consider a situation where user "B" wants to check the authenticity of a news video. User B accesses a web application, selects the news video file, and clicks the upload button. This action sends the video file from User B's computer to the server via an HTTP POST request.

[0593] The server receives the video file and temporarily stores it in storage. During this time, Person B's device uses the camera and microphone to capture Person B's facial expressions and voice data. The emotion engine analyzes this data and identifies Person B's emotional state (e.g., concern).

[0594] The server adds the received video file and the acquired emotion data to the analysis queue and runs the Deep Fake detection algorithm. Specifically, it divides the video into frames using "ffmpeg" and analyzes each frame using "TensorFlow." This analysis determines that 10% of the video is Deep Fake, and the confidence score is calculated as 90%.

[0595] The analysis results and emotion data are stored in a database along with the job ID. Finally, the server returns the analysis results, confidence score, and emotion data to Person B. Person B can check the analysis results, including the "concern" emotion state, along with the confidence score, in a web application.

[0596] Prompt Sentence Examples

[0597] "Please check the authenticity of this video and tell us how much Deepfake evidence there is."

[0598] Please send the Deepfake analysis results for the uploaded image along with the emotion data.

[0599] "Evaluate the reliability of the audio file and return the result. Emotional data on whether the user is concerned is also important."

[0600] This system allows users to accurately assess the reliability of the content they consume and provides comprehensive analysis based on emotional data.

[0601] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0602] Step 1:

[0603] Users upload content

[0604] Input: Data such as video, images, and audio files

[0605] How it works: A user logs into the web application, selects the file they want to analyze, and clicks the "Upload" button, which causes the user's device to send an HTTP POST request to the server with the selected file.

[0606] Output: Upload request sent to the server

[0607] Step 2:

[0608] The server receives the uploaded content

[0609] Input: HTTP POST request sent from the user terminal

[0610] Operation: The server receives the HTTP request and temporarily saves the sent file to the local disk. As an example, save the file to " / tmp / uploaded_files / video.mp4".

[0611] Output: Temporarily saved video file

[0612] Step 3:

[0613] Obtain the user's emotional data

[0614] Input: The user's facial expression and voice data

[0615] Operation: Use the camera and microphone of the user terminal to obtain the user's facial expression and voice in real time. Pass these data to the emotion engine, and the emotion engine analyzes and estimates the user's emotional state. For example, judge "suspicion" from "frowning expression".

[0616] Output: Obtained emotional data

[0617] Step 4:

[0618] The server adds to the analysis queue

[0619] Input: Temporarily saved video file, obtained emotional data

[0620] Operation: The server adds the received file and emotion data to the analysis queue and registers it as an analysis job. Specifically, it registers the file path " / tmp / uploaded_files / video.mp4" and the emotion data "concern" in the analysis queue.

[0621] Output: Jobs waiting to be analyzed

[0622] Step 5:

[0623] The server runs the Deep Fake detection algorithm

[0624] Input: Video files queued for analysis

[0625] How it works: The server splits the video into frames using ffmpeg and runs the Deep Fake detection algorithm on each frame. It then uses a deep learning model (e.g. TensorFlow) to determine whether or not there is a Deep Fake in each frame.

[0626] Output: Deep fake detection results for each frame

[0627] Step 6:

[0628] The server calculates the reliability score based on the detection result.

[0629] Input: Deep fake detection results for each frame

[0630] Operation: The detection results are aggregated and the overall confidence score is calculated based on the percentage of Deep fake detections. Specifically, if the Deep fake portion accounts for 10% of the total, a "confidence score of 90%" is calculated.

[0631] Output: Confidence Score

[0632] Step 7:

[0633] The server stores the analysis results, including emotional data.

[0634] Input: confidence score, acquired emotion data

[0635] Operation: The server saves the analysis results and emotion data in the database. It associates them with the job ID and saves "Video.mp4", "Confidence Score: 90%", and "Emotional State: Concern".

[0636] Output: Analysis results and emotion data stored in a database

[0637] Step 8:

[0638] The server returns the analysis results and emotion data to the user.

[0639] Input: Analysis results and emotion data stored in the database

[0640] Behavior: The server generates a report of the analysis results and returns it to the user. The user's web application displays the result "Video Confidence Score: 90%, Emotional State: Concern".

[0641] Output: Analysis results and emotion data displayed on the user's electronic device

[0642] (Application example 2)

[0643] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0644] Traditional deep fake detection systems focus on assessing the trustworthiness of content uploaded by users. However, these systems provide analysis results without considering the user's emotional state, and therefore cannot provide information about how users understand and perceive the trustworthiness of content. In particular, when the results of deep fake detection are ambiguous or have low confidence, this can burden users' understanding and behavior. To solve this problem and improve user experience, it is necessary to incorporate user emotional data into the analysis.

[0645] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[0646] In this invention, the server includes a means for users to upload content, a means for receiving the uploaded content and adding it to an analysis queue, a means for executing an algorithm to detect Deep Fakes contained in the content, a means for calculating a reliability score for the content based on the detection results, a means for executing an engine that acquires and analyzes emotional data from the user terminal, and a means for providing the user with feedback based on the emotional data along with the reliability score. This allows the user to receive feedback based on their own emotional state in addition to the reliability of the content, allowing the user to make decisions based on more comprehensive information.

[0647] The "means by which users upload content" refers to an interface that users use to send data such as video, image, and audio files to be analyzed to a server via the Internet.

[0648] The "means for receiving uploaded content and adding it to the analysis queue" is a function by which the server receives content sent by a user and adds the content to a waiting list for analysis processing.

[0649] "Means for executing an algorithm to detect Deep Fakes contained in content" refers to a program that allows the server to perform detailed analysis of each frame (video / image) or each second (audio) to identify traces of Deep Fakes.

[0650] "Means for calculating the reliability score of content based on the detection results" is a function that allows the server to use the detection results of Deep Fake to calculate a numerical value (reliability score) to evaluate the reliability of the entire content.

[0651] "Means for executing an engine that acquires and analyzes emotional data from a user terminal" refers to a program that acquires data using a camera or microphone and performs emotional analysis in order to estimate the emotional state of the user from their facial expressions and voice.

[0652] "Means for providing users with feedback based on emotional data together with the reliability score" is a function in which the server combines the reliability score with the user's emotional data to provide the analysis results to the user, providing feedback in a form that is easier for the user to understand.

[0653] This invention is a system that not only evaluates the reliability of content but also takes into account the emotional state of the user to support more comprehensive judgment. This system is mainly composed of a server and a user terminal.

[0654] The server and user terminal perform the following main operations:

[0655] First, the user device provides an interface for uploading the content (video, image, audio files) that the user wants to analyze. The user selects the file through a web application or a dedicated app and clicks the upload button. The user device then makes an HTTP POST request to send the selected file to the server.

[0656] The server then receives and temporarily stores the file sent from the user terminal, and adds the received content to the analysis queue, ready for analysis processing.

[0657] Furthermore, the emotion engine uses the camera and microphone installed on the user device to acquire emotion data from the user's facial expressions and voice, and analyzes this data to estimate the user's emotional state.

[0658] The server adds the uploaded file and the acquired emotion data to the analysis queue and registers it as an analysis job, which starts the subsequent analysis process.

[0659] The server runs the DeepFake detection algorithm on the content queued for analysis. This algorithm performs detailed analysis of each frame (video / image) and each second (audio) to detect traces of DeepFake.

[0660] The server calculates the reliability score for the entire content based on the Deep Fake detection results. For example, if 10% of the frames are determined to be Deep Fake, the reliability score is calculated as 60%. Furthermore, if the emotion engine recognizes the user's emotional state (happiness, sadness, anger, anxiety, etc.), that information is also added as feedback.

[0661] The analysis results include the reliability score of the content and the user's emotional data, which are stored in a database. The server returns these results in response to a user request. The results are then displayed, for example, on the screen of a web application.

[0662] A specific example would be analyzing images captured by a user during a meeting or video call. The user uploads the image to be analyzed, and the server analyzes the image. During the analysis, the user's facial expression is captured by a camera, and the emotional data is included in the analysis results. As a result, the emotional data "the user is concerned" is provided along with the confidence score of the image.

[0663] The following software and hardware are used to configure the emotion engine and Deep Fake detection algorithm functionality:

[0664] Open Source Computer Vision Library: OpenCV

[0665] Sentiment analysis engine: EmotionEngine

[0666] Deep fake detection algorithm API

[0667] This allows users to receive detailed information about the reliability of the content and feedback that includes their own emotional state, allowing them to make decisions based on more comprehensive information.

[0668] An example of a prompt to input to a generative AI model is as follows:

[0669] "Please confirm whether this image or video frame could have been generated by AI. Please return the analysis results and a confidence score. Below is the image data for verification."

[0670] "Please estimate the user's emotional state (happiness, sadness, anger, anxiety, etc.) from this image. Please return the analysis result and the emotional state. Below is the image data for testing."

[0671] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0672] Step 1: User uploads content

[0673] The user selects the content (video, image, audio file) they want to analyze through a web application or dedicated app and clicks the upload button. The user's device sends the selected file data to the server as an HTTP POST request. The input is the file data, and the output is the status of completion of transmission to the server.

[0674] Step 2: The server receives the uploaded content

[0675] The server receives files sent from the user terminal and temporarily stores them. The input is the file data sent from the user terminal, and the output is the storage path of the temporarily stored file. The specific operation is to receive the file and store it in the specified directory.

[0676] Step 3: Obtain user emotion data

[0677] The camera and microphone installed on the user device capture the user's facial expressions and voice, and analyzes the data. The emotion engine estimates the user's emotional state (happiness, sadness, anger, anxiety, etc.). The input is the captured image and voice data, and the output is numerical data related to the emotional state. Specific operations include activating the camera and microphone, capturing data, and analyzing the obtained data.

[0678] Step 4: Server adds to analysis queue

[0679] The server adds the received file and the acquired emotion data to the analysis queue and registers it as an analysis job. The input is the file data and emotion data, and the output is the status of completion of adding it to the analysis queue. The specific operation is to add the file and emotion data to the analysis queue.

[0680] Step 5: The server runs the Deep Fake detection algorithm

[0681] The server runs the DeepFake detection algorithm on the content queued for analysis. This algorithm performs detailed analysis of each frame (video / image) and each second (audio) to detect DeepFake signatures. The input is the file data to be analyzed, and the output is the DeepFake detection results and a confidence score. Specific operations include analyzing each frame and each second individually to detect specific patterns and features.

[0682] Step 6: The server calculates the trust score

[0683] The server calculates the reliability score of the entire content based on the detection results of Deep Fake. The input is the detection result data of Deep Fake, and the output is a numerical reliability score. Specific operations include executing an algorithm that sums up the detection results and evaluates the overall reliability.

[0684] Step 7: The server stores the analysis results and emotion data

[0685] The server stores the analysis results, confidence scores, and the acquired emotion data in a database. The input is the analysis result data, confidence score data, and emotion data, and the output is the status of completion of saving to the database. Specifically, each dataset is linked to a job ID and stored in the database.

[0686] Step 8: The server returns the analysis results and emotion data to the user.

[0687] The server returns the analysis results, confidence score, and user emotion data to the user. The input is the analysis result data, confidence score data, and emotion data, and the output is data to be displayed to the user. Specific operations include data format conversion and transmission to display the results on the screen of the web application used by the user.

[0688] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0689] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0690] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.

[0691] [Third embodiment]

[0692] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.

[0693] 5, the data processing system 310 includes the data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0694] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0695] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.

[0696] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0697] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0698] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0699] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0700] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0701] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0702] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0703] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."

[0704] To implement the present invention, it is necessary to build a system in which a user uploads content using a web application, a server analyzes the content, detects Deep Fake parts, calculates a reliability score, and returns the result to the user.

[0705] The server performs the following main operations:

[0706] 1. User uploads content

[0707] A user accesses the web application, selects the video, image, or audio file they want to analyze, and clicks the upload button, which causes the user's device to issue an HTTP POST request to send the selected content to the server.

[0708] 2. The server receives the uploaded content

[0709] The server receives the content sent from the user terminal, temporarily stores the content received by the server, and is ready for the next analysis step.

[0710] 3. The server adds it to the analysis queue

[0711] The server adds the received content to the analysis queue and registers it as an analysis job. This operation puts the content in a queue and allows subsequent analysis processing to begin.

[0712] 4. The server runs the Deep Fake detection algorithm

[0713] The server runs the DeepFake detection algorithm on the content queued for analysis. This algorithm performs a detailed analysis of each frame (video / image) or each second (audio) to detect DeepFake signatures.

[0714] 5. The server calculates the reliability score based on the detection result.

[0715] The server calculates the reliability score of the entire content based on the ratio of detected deep fakes. The reliability score is provided to the user as a numerical value for evaluating the accuracy of the content.

[0716] 6. The server returns the analysis results and the reliability score to the user.

[0717] The server returns the completed analysis results and the confidence score to the user, which are then displayed in the user's web application.

[0718] Specific examples

[0719] For example, suppose user "A" wants to check the reliability of a certain news video. A uploads the video through a web application, and the server receives the video and adds it to the analysis queue. The server analyzes each frame of the video to detect whether or not it contains Deep fakes. The detection results indicate that 5% of the frames are Deep fakes, and calculate a reliability score of 75%. The server notifies A of this reliability score and detailed analysis results. Based on this reliability score, A can determine that the news video is relatively reliable.

[0720] In this way, the present invention provides a system that allows users to easily evaluate the reliability of the content they consume and assists them in making accurate information selections.

[0721] The processing flow will be explained below.

[0722] Step 1:

[0723] The user selects the content and performs the upload operation. The user accesses the web application, selects the file they want to analyze, and clicks the "Upload" button. This operation sends the selected file from the user's device to the server.

[0724] Step 2:

[0725] The user device sends a file to the server via an HTTP POST request, along with metadata, structuring the file in a format that the server can receive.

[0726] Step 3:

[0727] The server receives the file sent from the user terminal, temporarily stores the file, and prepares it for the next analysis process.

[0728] Step 4:

[0729] The server adds the received file to the analysis queue. When it is registered as an analysis job, a job ID is generated and used to track the processing progress.

[0730] Step 5:

[0731] The server initializes the Deep Fake detection algorithm. In this step, the necessary models and libraries are loaded and the system is ready for analysis.

[0732] Step 6:

[0733] The server splits the file. In the case of video, it is split into frames, and in the case of images and audio, it is split into appropriate units (pixels, seconds).

[0734] Step 7:

[0735] The server analyzes each unit (frame, pixel, second) using a Deep Fake detection algorithm, which evaluates each unit individually to determine the presence of Deep Fake.

[0736] Step 8:

[0737] The server aggregates the analysis results for all units and calculates the Deep Fake ratio for the entire content based on the detection results for each unit.

[0738] Step 9:

[0739] The server calculates a trust score based on the Deep Fake ratio. The trust score is expressed between 0 and 100, with higher scores indicating more trustworthy content.

[0740] Step 10:

[0741] The server saves the analysis results and the confidence score in a database. The results are saved as results associated with the job ID, and can be returned in response to a user request.

[0742] Step 11:

[0743] The server notifies the user that the analysis results are ready, and the user can then review the results.

[0744] Step 12:

[0745] The user terminal sends a request to the server to display the analysis results. The request includes the job ID.

[0746] Step 13:

[0747] The server retrieves the analysis results corresponding to the job ID from the database and returns them to the user device along with the reliability score.

[0748] Step 14:

[0749] The user's device displays the received reliability score and analysis results, and the user evaluates the reliability of the content based on these results.

[0750] Example 1

[0751] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0752] In recent years, the development of deep fake technology has increased the risk of inaccurate information spreading. As a result, it has become more difficult for users to determine trustworthy content. Therefore, there is a need for technology that can quickly and accurately evaluate the reliability of content uploaded by users. Conventional technologies require a lot of time to analyze content, and the results are not always highly reliable. The present invention aims to solve this problem.

[0753] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0754] In this invention, the server includes a means for users to upload content, a means for receiving the uploaded content and adding it to an analysis queue, a means for executing an algorithm to detect Deep Fakes contained in the content, a means for calculating a reliability score of the content based on the detection result, a means for returning the reliability score and the analysis result to the user, and a means for displaying the analysis result, thereby enabling the user to quickly and accurately evaluate the reliability of the content.

[0755] A "user" is a person or entity that uses the system to upload content and receive the results of that analysis.

[0756] "Content" refers to the media data to be analyzed, such as videos, images, and audio files.

[0757] "Uploading means" means the functionality or interface that allows a user to submit content to the system.

[0758] "Server" means a computer system that receives, analyzes, and processes content uploaded by users.

[0759] "Analysis Queue" refers to a queue that registers and manages uploaded content for analysis in order.

[0760] "Deep fake" is fake media content generated using deep learning technology.

[0761] "Means for executing algorithms" refers to the ability to execute programs or processes to detect Deep Fakes in content.

[0762] The "Reliability Score" is a numerical value that evaluates the reliability of content, calculated based on the percentage of detected Deep Fakes.

[0763] "Analysis results" refers to the information and data obtained after the server runs the Deep Fake detection algorithm.

[0764] The "means for returning" is a function that allows the server to provide the analysis results and the reliability score to the user.

[0765] "Display means" refers to an interface or method for visually presenting the analysis results and the reliability score to the user.

[0766] A "database" is an electronic data storage system that stores analysis results and allows for searching and referencing as needed.

[0767] A "request" refers to an operation or action in which a user requests information such as analysis results from a system.

[0768] MODE FOR CARRYING OUT THE INVENTION

[0769] In the system based on this invention, a user uploads content using a web application, the server analyzes the content, detects Deep Fake parts, calculates a reliability score, and returns the result to the user. A specific embodiment of this system will be described below.

[0770] System Overview

[0771] Users can access the web application and select and upload video, image, or audio files they want to analyze. The user device issues an HTTP POST request to send the selected content to the server. The server temporarily stores the received content and then adds it to the analysis queue. The server runs the Deep Fake detection algorithm on the content in the analysis queue and calculates a confidence score based on the detection results. Finally, the server returns the analysis results and confidence score to the user and displays them in the web application.

[0772] Hardware and Software Use

[0773] server:

[0774] The server can be a cloud-based virtual machine with high-performance computing resources, such as an EC2 instance from Amazon Web Services (AWS) or a Compute Engine from Google Cloud Platform (GCP).

[0775] The server uses MySQL or PostgreSQL as a database to efficiently store and manage analysis results.

[0776] Deep Learning Frameworks:

[0777] Deep fake detection uses deep learning frameworks such as TensorFlow and PyTorch, which enable advanced image and video analysis.

[0778] User device:

[0779] A user terminal is a device that can connect to the Internet, such as a PC, smartphone, or tablet, and accesses web applications through a web browser (e.g., Google Chrome or Mozilla Firefox).

[0780] Specific examples

[0781] For example, suppose user "A" wants to check the reliability of a certain news video. A uploads the video through a web application, and the server receives the video and adds it to the analysis queue. The server analyzes each frame of the video to detect whether or not it contains Deep fakes. The detection results indicate that 5% of the frames are Deep fakes, and calculate a reliability score of 75%. The server notifies A of this reliability score and detailed analysis results. Based on this reliability score, A can determine that the news video is relatively reliable.

[0782] Example prompts to be input to the generative AI model

[0783] Enter "Please rate the reliability of the news video. Determine whether this video is a Deep Fake and calculate the reliability score." The generated results include the reliability score and the detection results for each analyzed frame.

[0784] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0785] Step 1:

[0786] A user accesses a web application and selects a media file (video, image, or audio file) to upload. When the user clicks the upload button, the device issues an HTTP POST request to send the selected content to the server. This request includes metadata (file name, file size, file type, etc.) along with the file data. The input is the media file selected by the user, and the output is the HTTP request sent to the server.

[0787] Step 2:

[0788] The server receives the HTTP POST request sent from the user device and extracts the content data. The server temporarily stores the received content in local storage, for example in the " / tmp / uploaded_files" directory. The input is the media file sent in step 1, and the output is the file stored in local storage.

[0789] Step 3:

[0790] The server adds the temporarily saved file to the analysis queue. It registers it as a job in the analysis queue management system (e.g., Redis Queue) and generates a job ID. The registered job waits for subsequent analysis processing. The input is the path of the saved file, and the output is the job ID registered in the analysis queue.

[0791] Step 4:

[0792] The server processes jobs registered in the analysis queue. It uses a deep learning framework (e.g., TensorFlow or PyTorch) to execute the Deep Fake detection algorithm. It performs a detailed analysis of each frame (video / image) or each second (audio) of the content to detect traces of Deep Fake. The input is the job ID and the corresponding file path, and the output is the Deep Fake detection result (the judgment result for each frame or second).

[0793] Step 5:

[0794] The server calculates the percentage of detected Deep fakes and calculates a confidence score. Specifically, it calculates a confidence score in the range of 0 to 100 using the percentage of Deep fakes determined to be Deep fakes for all frames and seconds. For example, if 5% of frames are Deep fakes, the confidence score is calculated as 95%. The input is the Deep fake detection result, and the output is the confidence score.

[0795] Step 6:

[0796] The server returns the reliability score and detailed analysis results to the user. The results are sent to the web application in JSON format and displayed in the user's browser. For example, a message such as "The reliability score of the news video is 75%. Click here for detailed analysis results" is displayed. The input is the reliability score and analysis results, and the output is the results displayed in the user's web browser.

[0797] (Application example 1)

[0798] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0799] In recent years, advances in deepfake technology have increased the risk of content such as video and audio being easily altered. Such alterations could have serious implications for important business decisions and communications. Therefore, verifying the authenticity of content is essential, especially for corporate security departments and journalists. However, traditional methods of manually verifying each piece of content are extremely time-consuming and inefficient. This has created a demand for an automated and reliable deepfake detection system.

[0800] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0801] In this invention, the server includes a means for users to upload content, a means for receiving the uploaded content and adding it to an analysis queue, a means for executing an algorithm to detect Deep Fakes contained in the content, a means for calculating a reliability score for the content based on the detection results, a means for analyzing recordings of corporate audio or video conferences to detect whether Deep Fake technology is used, and a means for returning the reliability score and analysis results to the user. This allows users to automatically verify whether the content they uploaded has been altered using Deep Fake technology and quickly evaluate its reliability.

[0802] "User" means a person or entity that uses the system to upload content and receive the results of that analysis.

[0803] "Content" refers to media files such as video, audio, and images uploaded by users.

[0804] "Upload" is the act of a user sending content they own to the system.

[0805] "Receiving" is the process by which the server receives content sent by the user.

[0806] The "analysis queue" is a collection of content that is waiting to be analyzed.

[0807] "Deep fake" is media content that has been artificially generated or altered using deep learning techniques.

[0808] An "algorithm" is a computational procedure or step for solving a particular problem.

[0809] The "confidence score" is a numerical assessment of the likelihood that content has been altered.

[0810] "Analysis Results" means the analytical output generated by the deepfake detection algorithm.

[0811] "Enterprise" means an entity or organization that conducts business activities.

[0812] An "audio or video conference" is a real-time audio and / or video conference between people in different locations.

[0813] "Deepfake technology" is a technology that uses deep learning technology to alter existing media content.

[0814] "Use" refers to whether a particular technology or method is used.

[0815] "Return" is the process in which the server returns the analysis results and the reliability score to the user.

[0816] To implement this invention, it is necessary to build a system and execute a series of processes for users to upload content and return the analysis results and reliability scores. This process is mainly composed of specific means including a server, a user terminal, and an analysis algorithm.

[0817] First, users use a web application to upload content (video, audio, or images). This web application is built using React and provides a user interface that allows users to intuitively upload content. When the user clicks the upload button, the user's device sends the selected content to the server as an HTTP POST request.

[0818] Next, the server is built using Node.js and Express.js and receives the content sent by the user. The received content is temporarily stored and then added to the analysis queue. This analysis queue keeps the content waiting to be analyzed in order.

[0819] The server then runs a DeepFake detection algorithm on the received content. This algorithm, built using TensorFlow and PyTorch, performs detailed analysis of each frame and each second of data to detect signs of deepfakes. Based on the results of this analysis, a confidence score is calculated. The confidence score is a numerical assessment of the accuracy of the content.

[0820] Finally, the server returns the calculated confidence score and analysis results to the user, displaying the results in the user's web application. For example, a user might upload a recording of a business meeting or presentation and check whether the recording has been altered using deepfake technology. In such a case, the server analyzes each frame of the recording, calculates a confidence score, and notifies the user.

[0821] As a concrete example, suppose a user uploads a recording of a business presentation. The server analyzes the file to detect whether deepfake technology has been used. Based on the detection result, the server calculates a confidence score and returns it to the user, such as "confidence score is 97%."

[0822] An example of a prompt might be:

[0823] Upload and analyze a recording of your business presentation, and use DeepFake technology to detect any alterations and calculate a confidence score.

[0824] This completes the description of the embodiment of the present invention. This system enables users to automatically verify whether uploaded content has been altered using deepfake technology and quickly evaluate its authenticity.

[0825] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0826] Step 1:

[0827] A user uploads content. Using the web application interface, the user selects the video, audio, or image file they want to analyze and clicks the upload button. The input is the media file selected by the user, and the output is the content sent from the user device to the server. Specific actions include selecting a file and clicking the upload button.

[0828] Step 2:

[0829] The server receives the uploaded content. The server receives the content sent as an HTTP POST request and temporarily stores it. The input is the received media file, and the output is that file saved in temporary storage on the server. Specific operations include receiving a file and saving a file.

[0830] Step 3:

[0831] The server adds the received content to the analysis queue. The server registers the received media file as an analysis job and adds it to the analysis queue. The input is a media file stored in temporary storage, and the output is that file added to the analysis queue. Specific operations include registering the file as an analysis job and updating the queue.

[0832] Step 4:

[0833] The server runs the Deep Fake detection algorithm. The server runs the Deep Fake detection algorithm using TensorFlow or PyTorch on the media files registered in the analysis queue. The media files registered in the analysis queue are used as input, and the deep fake detection results are obtained as output. Specific operations include analyzing each frame or each second and detecting traces of deep fakes.

[0834] Step 5:

[0835] The server calculates a confidence score based on the detection results. The server calculates a confidence score for the entire media file based on the percentage of deepfakes detected. The deepfake detection results are the input, and the confidence score is calculated and obtained as the output. Specific operations include calculating the percentage of deepfakes and calculating the confidence score.

[0836] Step 6:

[0837] The server returns the analysis results and confidence score to the user. The server returns the calculated confidence score and detailed analysis results to the user's web application and displays the results. The confidence score and analysis results are input, and this information is displayed to the user as output. Specific operations include returning and displaying the results.

[0838] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0839] To implement this invention, a system is required in which a user uploads content using a web application, a server analyzes the content, detects Deep Fake parts, calculates a reliability score, and returns the result to the user, combined with an emotion engine that recognizes the user's emotions.

[0840] The server and user terminal perform the following main operations:

[0841] 1. User uploads content

[0842] A user accesses the web application, selects the file (video, image, audio file) they want to analyze, and clicks the upload button. This action causes the user's device to make an HTTP POST request to send the selected file to the server.

[0843] 2. The server receives the uploaded content

[0844] The server receives the file sent from the user terminal, temporarily stores the content received by the server, and prepares it for the next analysis process.

[0845] 3. Obtaining user emotion data

[0846] Using a camera and microphone installed on the user's device, emotion data is acquired from the user's facial expressions and voice. The emotion engine analyzes this data and estimates the user's emotional state.

[0847] 4. The server adds it to the analysis queue

[0848] The server adds the received file and the acquired emotion data to the analysis queue and registers it as an analysis job. This operation puts the content in a queue and allows subsequent analysis processing to begin.

[0849] 5. The server runs the Deep Fake detection algorithm

[0850] The server runs the DeepFake detection algorithm on the content queued for analysis. This algorithm performs a detailed analysis of each frame (video / image) or each second (audio) to detect DeepFake signatures.

[0851] 6. The server calculates the reliability score based on the detection result.

[0852] The server calculates the reliability score of the entire content based on the ratio of detected deep fakes, and provides the reliability score to the user as a numerical value for evaluating the accuracy of the content.

[0853] 7. The server stores the analysis results, including emotional data.

[0854] The server stores the completed analysis results, the confidence score, and the acquired emotion data in a database. These are stored in association with the job ID, and can be returned in response to subsequent requests.

[0855] 8. The server returns the analysis results and emotion data to the user.

[0856] The server returns the analysis results, a confidence score, and the user's emotion data to the user, which are then displayed in the user's web application.

[0857] Specific examples

[0858] For example, suppose user "B" wants to check the reliability of a news video. B uploads the video through a web application, and the server receives it and adds it to the analysis queue. At this time, the user's device acquires emotional data from B's facial expressions and voice, and sends this data to the server. The server analyzes each frame of the video to detect whether or not it contains Deep Fakes. The detection results indicate that 10% of the frames are Deep Fakes, with a calculated reliability score of 60%. Furthermore, suppose the emotion engine recognizes that B has expressed concerns. The server returns these analysis results to B, who then evaluates the reliability of the video based on this information. In this way, adding emotional data enables a system that enables more detailed analysis and improves the user experience.

[0859] Through this series of processes, the present invention can provide a system that not only evaluates the reliability of content consumed by a user, but also supports comprehensive judgment, including the user's emotional state.

[0860] The processing flow will be explained below.

[0861] Step 1:

[0862] The user selects the content and performs the upload operation. The user accesses the web application, selects the video, image, or audio file they want to analyze, and clicks the "Upload" button. This operation sends the selected file from the user's device to the server.

[0863] Step 2:

[0864] The user device sends a file to the server via an HTTP POST request, along with metadata, structuring the file in a format that the server can receive.

[0865] Step 3:

[0866] The server receives the file sent from the user terminal, temporarily stores the file, and prepares for the next analysis process.

[0867] Step 4:

[0868] Acquires user emotional data. Emotional data is acquired from the user's facial expressions and voice using a camera and microphone installed on the user's device. The emotion engine analyzes this data and estimates the user's emotional state.

[0869] Step 5:

[0870] The server adds the received file and the acquired emotion data to the analysis queue. At the same time as it registers it as an analysis job, a job ID is generated and used to track the progress of the processing.

[0871] Step 6:

[0872] The server initializes the Deep Fake detection algorithm. In this step, the necessary models and libraries are loaded and the system is ready for analysis.

[0873] Step 7:

[0874] The server splits the file. In the case of video, it is split into frames, and in the case of images and audio, it is split into appropriate units.

[0875] Step 8:

[0876] The server analyzes each unit (frame, pixel, second) using a Deep Fake detection algorithm, which evaluates each unit individually to determine the presence of Deep Fake.

[0877] Step 9:

[0878] The server aggregates the analysis results for all units and calculates the Deep Fake ratio for the entire content based on the detection results for each unit.

[0879] Step 10:

[0880] The server calculates a trust score based on the Deep Fake ratio. The trust score is expressed between 0 and 100, with higher scores indicating more trustworthy content.

[0881] Step 11:

[0882] The server stores the analysis results, confidence score, and acquired emotion data in a database, linked to the job ID, so that the results can be returned in response to subsequent requests.

[0883] Step 12:

[0884] The server notifies the user that the analysis results are ready, and the user can then review the results.

[0885] Step 13:

[0886] The user device sends a request to the server to display the analysis results. The request includes the job ID.

[0887] Step 14:

[0888] The server retrieves the analysis results corresponding to the job ID from the database and returns them to the user device along with the confidence score and emotion data.

[0889] Step 15:

[0890] The user's device displays the received trust score, analysis results, and emotional data. Based on these results, the user evaluates the trustworthiness of the content and their own emotional state. This information helps the user to more accurately judge the trustworthiness of the content.

[0891] Example 2

[0892] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0893] Conventional content analysis systems are specialized in detecting deep fakes and calculating their reliability, and are unable to provide comprehensive analysis results that take into account the user's emotional state. This has resulted in issues such as the inability to fully reflect the emotional impact of users when evaluating the reliability of content. Furthermore, there has been an issue with insufficient storage and management of analysis results, making it difficult for users to easily check past analysis results.

[0894] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0895] In this invention, the server includes means for a user to upload data using an electronic device, means for receiving the uploaded data and adding it to an analysis queue, means for acquiring and analyzing the user's emotional data, means for executing an algorithm for detecting deep fakes contained in the data, means for calculating a reliability score of the data based on the detection results, means for storing the analysis results including the emotional data in a database, and means for returning the reliability score and the analysis results to the user. This enables accurate detection of deep fakes in content uploaded by users and provision of comprehensive analysis results including emotional data. Furthermore, by storing the analysis results and emotional data in a database, users can easily check past analysis results.

[0896] "User" refers to an individual or entity that uses an electronic device to upload data and view analytical results.

[0897] "Electronic Device" refers to a device, such as a computer or smartphone, that a User uses to upload Content.

[0898] "Data" refers to content such as videos, images, and audio files uploaded by users.

[0899] "Means for uploading" refers to the operation performed by a user using an electronic device to send data to a server, and the software that supports that operation.

[0900] "Means for receiving" refers to the hardware and software that the server uses to receive and temporarily store data sent by the user.

[0901] "Analysis queue" refers to a waiting list used to process uploaded data in order as analysis jobs.

[0902] "Emotion data" refers to information that indicates the emotional state of a user analyzed from their facial expressions and voice.

[0903] "Means for acquiring emotional data" refers to hardware and software for collecting and analyzing the user's facial expressions and voice to estimate their emotional state.

[0904] "Deep fake" refers to the technology and products that use deep learning technology to generate realistic-looking fake video and audio.

[0905] "Means for implementing algorithms to detect Deep Fakes" refers to hardware and software that runs deep learning models and other algorithms to identify Deep Fake signatures in data.

[0906] "Reliability Score" refers to a number that indicates the reliability of the data, calculated based on the percentage of deep fakes contained in the data.

[0907] "Means for calculating the reliability score" refers to the hardware and software for calculating the reliability score based on the ratio of Deep fake detection results.

[0908] "Means for storing in a database" refers to a storage system and its management software for permanently storing analysis results and emotion data.

[0909] "Means of return" refers to software and communication means for displaying or notifying the user of the analysis results, reliability score, and emotional data.

[0910] MODE FOR CARRYING OUT THE INVENTION

[0911] This invention relates to a system in which a user uploads data using an electronic device, a server analyzes the data to detect Deep Fakes and calculate a reliability score, and also obtains the user's emotional state and returns a comprehensive analysis result.

[0912] The system includes the following major hardware and software:

[0913] User terminal

[0914] The user terminal is an electronic device such as a computer or smartphone that provides an interface for users to upload data. The user terminal is equipped with a camera and microphone, which can capture the user's facial expression and voice data. This data is analyzed by the emotion engine to identify the user's emotional state.

[0915] server

[0916] The server receives data sent by users and has a storage system for temporarily storing it. The received data is added to an analysis queue and registered as an analysis job.

[0917] Deep fake detection algorithm: The server implements a deep learning model to detect deep fakes in the data. This algorithm uses deep learning libraries such as TensorFlow and PyTorch.

[0918] Data analysis: The server uses ffmpeg to split the video into frames and then runs the Deep Fake detection process on each frame.

[0919] Calculation of confidence score: Based on the Deep Fake detection results, a score is calculated to evaluate the reliability of the entire data. This is provided as a confidence score ranging from 0% to 100%.

[0920] Database

[0921] The server is equipped with a database for permanently storing analysis results and user emotion data, making it possible to easily track past analysis results and display them again upon user request.

[0922] Data return

[0923] Once the analysis is complete, the server returns the confidence score, analysis results, and emotion data to the user, which are then displayed on the screen of the user's electronic device.

[0924] Specific examples

[0925] For example, consider a situation where user "B" wants to check the authenticity of a news video. User B accesses a web application, selects the news video file, and clicks the upload button. This action sends the video file from User B's computer to the server via an HTTP POST request.

[0926] The server receives the video file and temporarily stores it in storage. During this time, Person B's device uses the camera and microphone to capture Person B's facial expressions and voice data. The emotion engine analyzes this data and identifies Person B's emotional state (e.g., concern).

[0927] The server adds the received video file and the acquired emotion data to the analysis queue and runs the Deep Fake detection algorithm. Specifically, it divides the video into frames using "ffmpeg" and analyzes each frame using "TensorFlow." This analysis determines that 10% of the video is Deep Fake, and the confidence score is calculated as 90%.

[0928] The analysis results and emotion data are stored in a database along with the job ID. Finally, the server returns the analysis results, confidence score, and emotion data to Person B. Person B can check the analysis results, including the "concern" emotion state, along with the confidence score, in a web application.

[0929] Prompt Sentence Examples

[0930] "Please check the authenticity of this video and tell us how much Deepfake evidence there is."

[0931] Please send the Deepfake analysis results for the uploaded image along with the emotion data.

[0932] "Evaluate the reliability of the audio file and return the result. Emotional data on whether the user is concerned is also important."

[0933] This system allows users to accurately assess the reliability of the content they consume and provides comprehensive analysis based on emotional data.

[0934] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0935] Step 1:

[0936] Users upload content

[0937] Input: Data such as video, images, and audio files

[0938] How it works: A user logs into the web application, selects the file they want to analyze, and clicks the "Upload" button, which causes the user's device to send an HTTP POST request to the server with the selected file.

[0939] Output: Upload request sent to the server

[0940] Step 2:

[0941] The server receives the uploaded content

[0942] Input: HTTP POST request sent from the user terminal

[0943] Operation: The server receives the HTTP request and temporarily saves the sent file to the local disk. As an example, save the file to " / tmp / uploaded_files / video.mp4".

[0944] Output: Temporarily saved video file

[0945] Step 3:

[0946] Obtain the user's sentiment data

[0947] Input: The user's facial expression and voice data

[0948] Operation: Use the camera and microphone of the user terminal to obtain the user's facial expression and voice in real time. Pass these data to the sentiment engine, and the sentiment engine analyzes and estimates the user's sentiment state. For example, determine "suspicion" from a "frowning expression".

[0949] Output: Obtained sentiment data

[0950] Step 4:

[0951] The server adds to the analysis queue

[0952] Input: Temporarily saved video file, obtained sentiment data

[0953] Operation: The server adds the received file and emotion data to the analysis queue and registers it as an analysis job. Specifically, it registers the file path " / tmp / uploaded_files / video.mp4" and the emotion data "concern" in the analysis queue.

[0954] Output: Jobs waiting to be analyzed

[0955] Step 5:

[0956] The server runs the Deep Fake detection algorithm

[0957] Input: Video files queued for analysis

[0958] How it works: The server splits the video into frames using ffmpeg and runs the Deep Fake detection algorithm on each frame. It then uses a deep learning model (e.g. TensorFlow) to determine whether or not there is a Deep Fake in each frame.

[0959] Output: Deep fake detection results for each frame

[0960] Step 6:

[0961] The server calculates the reliability score based on the detection result.

[0962] Input: Deep fake detection results for each frame

[0963] Operation: The detection results are aggregated and the overall confidence score is calculated based on the percentage of Deep fake detections. Specifically, if the Deep fake portion accounts for 10% of the total, a "confidence score of 90%" is calculated.

[0964] Output: Confidence Score

[0965] Step 7:

[0966] The server stores the analysis results, including emotional data.

[0967] Input: confidence score, acquired emotion data

[0968] Operation: The server saves the analysis results and emotion data in the database. It associates them with the job ID and saves "Video.mp4", "Confidence Score: 90%", and "Emotional State: Concern".

[0969] Output: Analysis results and emotion data stored in a database

[0970] Step 8:

[0971] The server returns the analysis results and emotion data to the user.

[0972] Input: Analysis results and emotion data stored in the database

[0973] Behavior: The server generates a report of the analysis results and returns it to the user. The user's web application displays the result "Video Confidence Score: 90%, Emotional State: Concern".

[0974] Output: Analysis results and emotion data displayed on the user's electronic device

[0975] (Application example 2)

[0976] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0977] Traditional deep fake detection systems focus on assessing the trustworthiness of content uploaded by users. However, these systems provide analysis results without considering the user's emotional state, and therefore cannot provide information about how users understand and perceive the trustworthiness of content. In particular, when the results of deep fake detection are ambiguous or have low confidence, this can burden users' understanding and behavior. To solve this problem and improve user experience, it is necessary to incorporate user emotional data into the analysis.

[0978] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[0979] In this invention, the server includes a means for users to upload content, a means for receiving the uploaded content and adding it to an analysis queue, a means for executing an algorithm to detect Deep Fakes contained in the content, a means for calculating a reliability score for the content based on the detection results, a means for executing an engine that acquires and analyzes emotional data from the user terminal, and a means for providing the user with feedback based on the emotional data along with the reliability score. This allows the user to receive feedback based on their own emotional state in addition to the reliability of the content, allowing the user to make decisions based on more comprehensive information.

[0980] The "means by which users upload content" refers to an interface that users use to send data such as video, image, and audio files to be analyzed to a server via the Internet.

[0981] The "means for receiving uploaded content and adding it to the analysis queue" is a function by which the server receives content sent by a user and adds the content to a waiting list for analysis processing.

[0982] "Means for executing an algorithm to detect Deep Fakes contained in content" refers to a program that allows the server to perform detailed analysis of each frame (video / image) or each second (audio) to identify traces of Deep Fakes.

[0983] "Means for calculating the reliability score of content based on the detection results" is a function that allows the server to use the detection results of Deep Fake to calculate a numerical value (reliability score) to evaluate the reliability of the entire content.

[0984] "Means for executing an engine that acquires and analyzes emotional data from a user terminal" refers to a program that acquires data using a camera or microphone and performs emotional analysis in order to estimate the emotional state of the user from their facial expressions and voice.

[0985] "Means for providing users with feedback based on emotional data together with the reliability score" is a function in which the server combines the reliability score with the user's emotional data to provide the analysis results to the user, providing feedback in a form that is easier for the user to understand.

[0986] This invention is a system that not only evaluates the reliability of content but also takes into account the emotional state of the user to support more comprehensive judgment. This system is mainly composed of a server and a user terminal.

[0987] The server and user terminal perform the following main operations:

[0988] First, the user device provides an interface for uploading the content (video, image, audio files) that the user wants to analyze. The user selects the file through a web application or a dedicated app and clicks the upload button. The user device then makes an HTTP POST request to send the selected file to the server.

[0989] The server then receives and temporarily stores the file sent from the user terminal, and adds the received content to the analysis queue, ready for analysis processing.

[0990] Furthermore, the emotion engine uses the camera and microphone installed on the user device to acquire emotion data from the user's facial expressions and voice, and analyzes this data to estimate the user's emotional state.

[0991] The server adds the uploaded file and the acquired emotion data to the analysis queue and registers it as an analysis job, which starts the subsequent analysis process.

[0992] The server runs the DeepFake detection algorithm on the content queued for analysis. This algorithm performs detailed analysis of each frame (video / image) and each second (audio) to detect traces of DeepFake.

[0993] The server calculates the reliability score for the entire content based on the Deep Fake detection results. For example, if 10% of the frames are determined to be Deep Fake, the reliability score is calculated as 60%. Furthermore, if the emotion engine recognizes the user's emotional state (happiness, sadness, anger, anxiety, etc.), that information is also added as feedback.

[0994] The analysis results include the reliability score of the content and the user's emotional data, which are stored in a database. The server returns these results in response to a user request. The results are then displayed, for example, on the screen of a web application.

[0995] A specific example would be analyzing images captured by a user during a meeting or video call. The user uploads the image to be analyzed, and the server analyzes the image. During the analysis, the user's facial expression is captured by a camera, and the emotional data is included in the analysis results. As a result, the emotional data "the user is concerned" is provided along with the confidence score of the image.

[0996] The following software and hardware are used to configure the emotion engine and Deep Fake detection algorithm functionality:

[0997] Open Source Computer Vision Library: OpenCV

[0998] Sentiment analysis engine: EmotionEngine

[0999] Deep fake detection algorithm API

[1000] This allows users to receive detailed information about the reliability of the content and feedback that includes their own emotional state, allowing them to make decisions based on more comprehensive information.

[1001] An example of a prompt to input to a generative AI model is as follows:

[1002] "Please confirm whether this image or video frame could have been generated by AI. Please return the analysis results and a confidence score. Below is the image data for verification."

[1003] "Please estimate the user's emotional state (happiness, sadness, anger, anxiety, etc.) from this image. Please return the analysis result and the emotional state. Below is the image data for testing."

[1004] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1005] Step 1: User uploads content

[1006] The user selects the content (video, image, audio file) they want to analyze through a web application or dedicated app and clicks the upload button. The user's device sends the selected file data to the server as an HTTP POST request. The input is the file data, and the output is the status of completion of transmission to the server.

[1007] Step 2: The server receives the uploaded content

[1008] The server receives files sent from the user terminal and temporarily stores them. The input is the file data sent from the user terminal, and the output is the storage path of the temporarily stored file. The specific operation is to receive the file and store it in the specified directory.

[1009] Step 3: Obtain user emotion data

[1010] The camera and microphone installed on the user device capture the user's facial expressions and voice, and analyzes the data. The emotion engine estimates the user's emotional state (happiness, sadness, anger, anxiety, etc.). The input is the captured image and voice data, and the output is numerical data related to the emotional state. Specific operations include activating the camera and microphone, capturing data, and analyzing the obtained data.

[1011] Step 4: Server adds to analysis queue

[1012] The server adds the received file and the acquired emotion data to the analysis queue and registers it as an analysis job. The input is the file data and emotion data, and the output is the status of completion of adding it to the analysis queue. The specific operation is to add the file and emotion data to the analysis queue.

[1013] Step 5: The server runs the Deep Fake detection algorithm

[1014] The server runs the DeepFake detection algorithm on the content queued for analysis. This algorithm performs detailed analysis of each frame (video / image) and each second (audio) to detect DeepFake signatures. The input is the file data to be analyzed, and the output is the DeepFake detection results and a confidence score. Specific operations include analyzing each frame and each second individually to detect specific patterns and features.

[1015] Step 6: The server calculates the trust score

[1016] The server calculates the reliability score of the entire content based on the detection results of Deep Fake. The input is the detection result data of Deep Fake, and the output is a numerical reliability score. Specific operations include executing an algorithm that sums up the detection results and evaluates the overall reliability.

[1017] Step 7: The server stores the analysis results and emotion data

[1018] The server stores the analysis results, confidence scores, and the acquired emotion data in a database. The input is the analysis result data, confidence score data, and emotion data, and the output is the status of completion of saving to the database. Specifically, each dataset is linked to a job ID and stored in the database.

[1019] Step 8: The server returns the analysis results and emotion data to the user.

[1020] The server returns the analysis results, confidence score, and user emotion data to the user. The input is the analysis result data, confidence score data, and emotion data, and the output is data to be displayed to the user. Specific operations include data format conversion and transmission to display the results on the screen of the web application used by the user.

[1021] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[1022] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1023] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.

[1024] [Fourth embodiment]

[1025] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.

[1026] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[1027] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1028] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.

[1029] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1030] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1031] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1032] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.

[1033] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1034] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1035] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1036] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1037] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1038] To implement the present invention, it is necessary to build a system in which a user uploads content using a web application, a server analyzes the content, detects Deep Fake parts, calculates a reliability score, and returns the result to the user.

[1039] The server performs the following main operations:

[1040] 1. User uploads content

[1041] A user accesses the web application, selects the video, image, or audio file they want to analyze, and clicks the upload button, which causes the user's device to issue an HTTP POST request to send the selected content to the server.

[1042] 2. The server receives the uploaded content

[1043] The server receives the content sent from the user terminal, temporarily stores the content received by the server, and is ready for the next analysis step.

[1044] 3. The server adds it to the analysis queue

[1045] The server adds the received content to the analysis queue and registers it as an analysis job. This operation puts the content in a queue and allows subsequent analysis processing to begin.

[1046] 4. The server runs the Deep Fake detection algorithm

[1047] The server runs the DeepFake detection algorithm on the content queued for analysis. This algorithm performs a detailed analysis of each frame (video / image) or each second (audio) to detect DeepFake signatures.

[1048] 5. The server calculates the reliability score based on the detection result.

[1049] The server calculates the reliability score of the entire content based on the ratio of detected deep fakes. The reliability score is provided to the user as a numerical value for evaluating the accuracy of the content.

[1050] 6. The server returns the analysis results and the reliability score to the user.

[1051] The server returns the completed analysis results and the confidence score to the user, which are then displayed in the user's web application.

[1052] Specific examples

[1053] For example, suppose user "A" wants to check the reliability of a certain news video. A uploads the video through a web application, and the server receives the video and adds it to the analysis queue. The server analyzes each frame of the video to detect whether or not it contains Deep fakes. The detection results indicate that 5% of the frames are Deep fakes, and calculate a reliability score of 75%. The server notifies A of this reliability score and detailed analysis results. Based on this reliability score, A can determine that the news video is relatively reliable.

[1054] In this way, the present invention provides a system that allows users to easily evaluate the reliability of the content they consume and assists them in making accurate information selections.

[1055] The processing flow will be explained below.

[1056] Step 1:

[1057] The user selects the content and performs the upload operation. The user accesses the web application, selects the file they want to analyze, and clicks the "Upload" button. This operation sends the selected file from the user's device to the server.

[1058] Step 2:

[1059] The user device sends a file to the server via an HTTP POST request, along with metadata, structuring the file in a format that the server can receive.

[1060] Step 3:

[1061] The server receives the file sent from the user terminal, temporarily stores the file, and prepares it for the next analysis process.

[1062] Step 4:

[1063] The server adds the received file to the analysis queue. When it is registered as an analysis job, a job ID is generated and used to track the processing progress.

[1064] Step 5:

[1065] The server initializes the Deep Fake detection algorithm. In this step, the necessary models and libraries are loaded and the system is ready for analysis.

[1066] Step 6:

[1067] The server splits the file. In the case of video, it is split into frames, and in the case of images and audio, it is split into appropriate units (pixels, seconds).

[1068] Step 7:

[1069] The server analyzes each unit (frame, pixel, second) using a Deep Fake detection algorithm, which evaluates each unit individually to determine the presence of Deep Fake.

[1070] Step 8:

[1071] The server aggregates the analysis results for all units and calculates the Deep Fake ratio for the entire content based on the detection results for each unit.

[1072] Step 9:

[1073] The server calculates a trust score based on the Deep Fake ratio. The trust score is expressed between 0 and 100, with higher scores indicating more trustworthy content.

[1074] Step 10:

[1075] The server saves the analysis results and the confidence score in a database. The results are saved as results associated with the job ID, and can be returned in response to a user request.

[1076] Step 11:

[1077] The server notifies the user that the analysis results are ready, and the user can then review the results.

[1078] Step 12:

[1079] The user terminal sends a request to the server to display the analysis results. The request includes the job ID.

[1080] Step 13:

[1081] The server retrieves the analysis results corresponding to the job ID from the database and returns them to the user device along with the reliability score.

[1082] Step 14:

[1083] The user's device displays the received reliability score and analysis results, and the user evaluates the reliability of the content based on these results.

[1084] Example 1

[1085] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1086] In recent years, the development of deep fake technology has increased the risk of inaccurate information spreading. As a result, it has become more difficult for users to determine trustworthy content. Therefore, there is a need for technology that can quickly and accurately evaluate the reliability of content uploaded by users. Conventional technologies require a lot of time to analyze content, and the results are not always highly reliable. The present invention aims to solve this problem.

[1087] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1088] In this invention, the server includes a means for users to upload content, a means for receiving the uploaded content and adding it to an analysis queue, a means for executing an algorithm to detect Deep Fakes contained in the content, a means for calculating a reliability score of the content based on the detection result, a means for returning the reliability score and the analysis result to the user, and a means for displaying the analysis result, thereby enabling the user to quickly and accurately evaluate the reliability of the content.

[1089] A "user" is a person or entity that uses the system to upload content and receive the results of that analysis.

[1090] "Content" refers to the media data to be analyzed, such as videos, images, and audio files.

[1091] "Uploading means" means the functionality or interface that allows a user to submit content to the system.

[1092] "Server" means a computer system that receives, analyzes, and processes content uploaded by users.

[1093] "Analysis Queue" refers to a queue that registers and manages uploaded content for analysis in order.

[1094] "Deep fake" is fake media content generated using deep learning technology.

[1095] "Means for executing algorithms" refers to the ability to execute programs or processes to detect Deep Fakes in content.

[1096] The "Reliability Score" is a numerical value that evaluates the reliability of content, calculated based on the percentage of detected Deep Fakes.

[1097] "Analysis results" refers to the information and data obtained after the server runs the Deep Fake detection algorithm.

[1098] The "means for returning" is a function that allows the server to provide the analysis results and the reliability score to the user.

[1099] "Display means" refers to an interface or method for visually presenting the analysis results and the reliability score to the user.

[1100] A "database" is an electronic data storage system that stores analysis results and allows for searching and referencing as needed.

[1101] A "request" refers to an operation or action in which a user requests information such as analysis results from a system.

[1102] MODE FOR CARRYING OUT THE INVENTION

[1103] In the system based on this invention, a user uploads content using a web application, the server analyzes the content, detects Deep Fake parts, calculates a reliability score, and returns the result to the user. A specific embodiment of this system will be described below.

[1104] System Overview

[1105] Users can access the web application and select and upload video, image, or audio files they want to analyze. The user device issues an HTTP POST request to send the selected content to the server. The server temporarily stores the received content and then adds it to the analysis queue. The server runs the Deep Fake detection algorithm on the content in the analysis queue and calculates a confidence score based on the detection results. Finally, the server returns the analysis results and confidence score to the user and displays them in the web application.

[1106] Hardware and Software Use

[1107] server:

[1108] The server can be a cloud-based virtual machine with high-performance computing resources, such as an EC2 instance from Amazon Web Services (AWS) or a Compute Engine from Google Cloud Platform (GCP).

[1109] The server uses MySQL or PostgreSQL as a database to efficiently store and manage analysis results.

[1110] Deep Learning Frameworks:

[1111] Deep fake detection uses deep learning frameworks such as TensorFlow and PyTorch, which enable advanced image and video analysis.

[1112] User device:

[1113] A user terminal is a device that can connect to the Internet, such as a PC, smartphone, or tablet, and accesses web applications through a web browser (e.g., Google Chrome or Mozilla Firefox).

[1114] Specific examples

[1115] For example, suppose user "A" wants to check the reliability of a certain news video. A uploads the video through a web application, and the server receives the video and adds it to the analysis queue. The server analyzes each frame of the video to detect whether or not it contains Deep fakes. The detection results indicate that 5% of the frames are Deep fakes, and calculate a reliability score of 75%. The server notifies A of this reliability score and detailed analysis results. Based on this reliability score, A can determine that the news video is relatively reliable.

[1116] Example prompts to be input to the generative AI model

[1117] Enter "Please rate the reliability of the news video. Determine whether this video is a Deep Fake and calculate the reliability score." The generated results include the reliability score and the detection results for each analyzed frame.

[1118] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1119] Step 1:

[1120] A user accesses a web application and selects a media file (video, image, or audio file) to upload. When the user clicks the upload button, the device issues an HTTP POST request to send the selected content to the server. This request includes metadata (file name, file size, file type, etc.) along with the file data. The input is the media file selected by the user, and the output is the HTTP request sent to the server.

[1121] Step 2:

[1122] The server receives the HTTP POST request sent from the user device and extracts the content data. The server temporarily stores the received content in local storage, for example in the " / tmp / uploaded_files" directory. The input is the media file sent in step 1, and the output is the file stored in local storage.

[1123] Step 3:

[1124] The server adds the temporarily saved file to the analysis queue. It registers it as a job in the analysis queue management system (e.g., Redis Queue) and generates a job ID. The registered job waits for subsequent analysis processing. The input is the path of the saved file, and the output is the job ID registered in the analysis queue.

[1125] Step 4:

[1126] The server processes jobs registered in the analysis queue. It uses a deep learning framework (e.g., TensorFlow or PyTorch) to execute the Deep Fake detection algorithm. It performs a detailed analysis of each frame (video / image) or each second (audio) of the content to detect traces of Deep Fake. The input is the job ID and the corresponding file path, and the output is the Deep Fake detection result (the judgment result for each frame or second).

[1127] Step 5:

[1128] The server calculates the percentage of detected Deep fakes and calculates a confidence score. Specifically, it calculates a confidence score in the range of 0 to 100 using the percentage of Deep fakes determined to be Deep fakes for all frames and seconds. For example, if 5% of frames are Deep fakes, the confidence score is calculated as 95%. The input is the Deep fake detection result, and the output is the confidence score.

[1129] Step 6:

[1130] The server returns the reliability score and detailed analysis results to the user. The results are sent to the web application in JSON format and displayed in the user's browser. For example, a message such as "The reliability score of the news video is 75%. Click here for detailed analysis results" is displayed. The input is the reliability score and analysis results, and the output is the results displayed in the user's web browser.

[1131] (Application example 1)

[1132] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1133] In recent years, advances in deepfake technology have increased the risk of content such as video and audio being easily altered. Such alterations could have serious implications for important business decisions and communications. Therefore, verifying the authenticity of content is essential, especially for corporate security departments and journalists. However, traditional methods of manually verifying each piece of content are extremely time-consuming and inefficient. This has created a demand for an automated and reliable deepfake detection system.

[1134] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1135] In this invention, the server includes a means for users to upload content, a means for receiving the uploaded content and adding it to an analysis queue, a means for executing an algorithm to detect Deep Fakes contained in the content, a means for calculating a reliability score for the content based on the detection results, a means for analyzing recordings of corporate audio or video conferences to detect whether Deep Fake technology is used, and a means for returning the reliability score and analysis results to the user. This allows users to automatically verify whether the content they uploaded has been altered using Deep Fake technology and quickly evaluate its reliability.

[1136] "User" means a person or entity that uses the system to upload content and receive the results of that analysis.

[1137] "Content" refers to media files such as video, audio, and images uploaded by users.

[1138] "Upload" is the act of a user sending content they own to the system.

[1139] "Receiving" is the process by which the server receives content sent by the user.

[1140] The "analysis queue" is a collection of content that is waiting to be analyzed.

[1141] "Deep fake" is media content that has been artificially generated or altered using deep learning techniques.

[1142] An "algorithm" is a computational procedure or step for solving a particular problem.

[1143] The "confidence score" is a numerical assessment of the likelihood that content has been altered.

[1144] "Analysis Results" means the analytical output generated by the deepfake detection algorithm.

[1145] "Enterprise" means an entity or organization that conducts business activities.

[1146] An "audio or video conference" is a real-time audio and / or video conference between people in different locations.

[1147] "Deepfake technology" is a technology that uses deep learning technology to alter existing media content.

[1148] "Use" refers to whether a particular technology or method is used.

[1149] "Return" is the process in which the server returns the analysis results and the reliability score to the user.

[1150] To implement this invention, it is necessary to build a system and execute a series of processes for users to upload content and return the analysis results and reliability scores. This process is mainly composed of specific means including a server, a user terminal, and an analysis algorithm.

[1151] First, users use a web application to upload content (video, audio, or images). This web application is built using React and provides a user interface that allows users to intuitively upload content. When the user clicks the upload button, the user's device sends the selected content to the server as an HTTP POST request.

[1152] Next, the server is built using Node.js and Express.js and receives the content sent by the user. The received content is temporarily stored and then added to the analysis queue. This analysis queue keeps the content waiting to be analyzed in order.

[1153] The server then runs a DeepFake detection algorithm on the received content. This algorithm, built using TensorFlow and PyTorch, performs detailed analysis of each frame and each second of data to detect signs of deepfakes. Based on the results of this analysis, a confidence score is calculated. The confidence score is a numerical assessment of the accuracy of the content.

[1154] Finally, the server returns the calculated confidence score and analysis results to the user, displaying the results in the user's web application. For example, a user might upload a recording of a business meeting or presentation and check whether the recording has been altered using deepfake technology. In such a case, the server analyzes each frame of the recording, calculates a confidence score, and notifies the user.

[1155] As a concrete example, suppose a user uploads a recording of a business presentation. The server analyzes the file to detect whether deepfake technology has been used. Based on the detection result, the server calculates a confidence score and returns it to the user, such as "confidence score is 97%."

[1156] An example of a prompt might be:

[1157] Upload and analyze a recording of your business presentation, and use DeepFake technology to detect any alterations and calculate a confidence score.

[1158] This completes the description of the embodiment of the present invention. This system enables users to automatically verify whether uploaded content has been altered using deepfake technology and quickly evaluate its authenticity.

[1159] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1160] Step 1:

[1161] A user uploads content. Using the web application interface, the user selects the video, audio, or image file they want to analyze and clicks the upload button. The input is the media file selected by the user, and the output is the content sent from the user device to the server. Specific actions include selecting a file and clicking the upload button.

[1162] Step 2:

[1163] The server receives the uploaded content. The server receives the content sent as an HTTP POST request and temporarily stores it. The input is the received media file, and the output is that file saved in temporary storage on the server. Specific operations include receiving a file and saving a file.

[1164] Step 3:

[1165] The server adds the received content to the analysis queue. The server registers the received media file as an analysis job and adds it to the analysis queue. The input is a media file stored in temporary storage, and the output is that file added to the analysis queue. Specific operations include registering the file as an analysis job and updating the queue.

[1166] Step 4:

[1167] The server runs the Deep Fake detection algorithm. The server runs the Deep Fake detection algorithm using TensorFlow or PyTorch on the media files registered in the analysis queue. The media files registered in the analysis queue are used as input, and the deep fake detection results are obtained as output. Specific operations include analyzing each frame or each second and detecting traces of deep fakes.

[1168] Step 5:

[1169] The server calculates a confidence score based on the detection results. The server calculates a confidence score for the entire media file based on the percentage of deepfakes detected. The deepfake detection results are the input, and the confidence score is calculated and obtained as the output. Specific operations include calculating the percentage of deepfakes and calculating the confidence score.

[1170] Step 6:

[1171] The server returns the analysis results and confidence score to the user. The server returns the calculated confidence score and detailed analysis results to the user's web application and displays the results. The confidence score and analysis results are input, and this information is displayed to the user as output. Specific operations include returning and displaying the results.

[1172] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1173] To implement this invention, a system is required in which a user uploads content using a web application, a server analyzes the content, detects Deep Fake parts, calculates a reliability score, and returns the result to the user, combined with an emotion engine that recognizes the user's emotions.

[1174] The server and user terminal perform the following main operations:

[1175] 1. User uploads content

[1176] A user accesses the web application, selects the file (video, image, audio file) they want to analyze, and clicks the upload button. This action causes the user's device to make an HTTP POST request to send the selected file to the server.

[1177] 2. The server receives the uploaded content

[1178] The server receives the file sent from the user terminal, temporarily stores the content received by the server, and prepares it for the next analysis process.

[1179] 3. Obtaining user emotion data

[1180] Using a camera and microphone installed on the user's device, emotion data is acquired from the user's facial expressions and voice. The emotion engine analyzes this data and estimates the user's emotional state.

[1181] 4. The server adds it to the analysis queue

[1182] The server adds the received file and the acquired emotion data to the analysis queue and registers it as an analysis job. This operation puts the content in a queue and allows subsequent analysis processing to begin.

[1183] 5. The server runs the Deep Fake detection algorithm

[1184] The server runs the DeepFake detection algorithm on the content queued for analysis. This algorithm performs a detailed analysis of each frame (video / image) or each second (audio) to detect DeepFake signatures.

[1185] 6. The server calculates the reliability score based on the detection result.

[1186] The server calculates the reliability score of the entire content based on the ratio of detected deep fakes, and provides the reliability score to the user as a numerical value for evaluating the accuracy of the content.

[1187] 7. The server stores the analysis results, including emotional data.

[1188] The server stores the completed analysis results, the confidence score, and the acquired emotion data in a database. These are stored in association with the job ID, and can be returned in response to subsequent requests.

[1189] 8. The server returns the analysis results and emotion data to the user.

[1190] The server returns the analysis results, a confidence score, and the user's emotion data to the user, which are then displayed in the user's web application.

[1191] Specific examples

[1192] For example, suppose user "B" wants to check the reliability of a news video. B uploads the video through a web application, and the server receives it and adds it to the analysis queue. At this time, the user's device acquires emotional data from B's facial expressions and voice, and sends this data to the server. The server analyzes each frame of the video to detect whether or not it contains Deep Fakes. The detection results indicate that 10% of the frames are Deep Fakes, with a calculated reliability score of 60%. Furthermore, suppose the emotion engine recognizes that B has expressed concerns. The server returns these analysis results to B, who then evaluates the reliability of the video based on this information. In this way, adding emotional data enables a system that enables more detailed analysis and improves the user experience.

[1193] Through this series of processes, the present invention can provide a system that not only evaluates the reliability of content consumed by a user, but also supports comprehensive judgment, including the user's emotional state.

[1194] The processing flow will be explained below.

[1195] Step 1:

[1196] The user selects the content and performs the upload operation. The user accesses the web application, selects the video, image, or audio file they want to analyze, and clicks the "Upload" button. This operation sends the selected file from the user's device to the server.

[1197] Step 2:

[1198] The user device sends a file to the server via an HTTP POST request, along with metadata, structuring the file in a format that the server can receive.

[1199] Step 3:

[1200] The server receives the file sent from the user terminal, temporarily stores the file, and prepares for the next analysis process.

[1201] Step 4:

[1202] Acquires user emotional data. Emotional data is acquired from the user's facial expressions and voice using a camera and microphone installed on the user's device. The emotion engine analyzes this data and estimates the user's emotional state.

[1203] Step 5:

[1204] The server adds the received file and the acquired emotion data to the analysis queue. At the same time as it registers it as an analysis job, a job ID is generated and used to track the progress of the processing.

[1205] Step 6:

[1206] The server initializes the Deep Fake detection algorithm. In this step, the necessary models and libraries are loaded and the system is ready for analysis.

[1207] Step 7:

[1208] The server splits the file. In the case of video, it is split into frames, and in the case of images and audio, it is split into appropriate units.

[1209] Step 8:

[1210] The server analyzes each unit (frame, pixel, second) using a Deep Fake detection algorithm, which evaluates each unit individually to determine the presence of Deep Fake.

[1211] Step 9:

[1212] The server aggregates the analysis results for all units and calculates the Deep Fake ratio for the entire content based on the detection results for each unit.

[1213] Step 10:

[1214] The server calculates a trust score based on the Deep Fake ratio. The trust score is expressed between 0 and 100, with higher scores indicating more trustworthy content.

[1215] Step 11:

[1216] The server stores the analysis results, confidence score, and acquired emotion data in a database, linked to the job ID, so that the results can be returned in response to subsequent requests.

[1217] Step 12:

[1218] The server notifies the user that the analysis results are ready, and the user can then review the results.

[1219] Step 13:

[1220] The user device sends a request to the server to display the analysis results. The request includes the job ID.

[1221] Step 14:

[1222] The server retrieves the analysis results corresponding to the job ID from the database and returns them to the user device along with the confidence score and emotion data.

[1223] Step 15:

[1224] The user's device displays the received trust score, analysis results, and emotional data. Based on these results, the user evaluates the trustworthiness of the content and their own emotional state. This information helps the user to more accurately judge the trustworthiness of the content.

[1225] Example 2

[1226] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1227] Conventional content analysis systems are specialized in detecting deep fakes and calculating their reliability, and are unable to provide comprehensive analysis results that take into account the user's emotional state. This has resulted in issues such as the inability to fully reflect the emotional impact of users when evaluating the reliability of content. Furthermore, there has been an issue with insufficient storage and management of analysis results, making it difficult for users to easily check past analysis results.

[1228] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1229] In this invention, the server includes means for a user to upload data using an electronic device, means for receiving the uploaded data and adding it to an analysis queue, means for acquiring and analyzing the user's emotional data, means for executing an algorithm for detecting deep fakes contained in the data, means for calculating a reliability score of the data based on the detection results, means for storing the analysis results including the emotional data in a database, and means for returning the reliability score and the analysis results to the user. This enables accurate detection of deep fakes in content uploaded by users and provision of comprehensive analysis results including emotional data. Furthermore, by storing the analysis results and emotional data in a database, users can easily check past analysis results.

[1230] "User" refers to an individual or entity that uses an electronic device to upload data and view analytical results.

[1231] "Electronic Device" refers to a device, such as a computer or smartphone, that a User uses to upload Content.

[1232] "Data" refers to content such as videos, images, and audio files uploaded by users.

[1233] "Means for uploading" refers to the operation performed by a user using an electronic device to send data to a server, and the software that supports that operation.

[1234] "Means for receiving" refers to the hardware and software that the server uses to receive and temporarily store data sent by the user.

[1235] "Analysis queue" refers to a waiting list used to process uploaded data in order as analysis jobs.

[1236] "Emotion data" refers to information that indicates the emotional state of a user analyzed from their facial expressions and voice.

[1237] "Means for acquiring emotional data" refers to hardware and software for collecting and analyzing the user's facial expressions and voice to estimate their emotional state.

[1238] "Deep fake" refers to the technology and products that use deep learning technology to generate realistic-looking fake video and audio.

[1239] "Means for implementing algorithms to detect Deep Fakes" refers to hardware and software that runs deep learning models and other algorithms to identify Deep Fake signatures in data.

[1240] "Reliability Score" refers to a number that indicates the reliability of the data, calculated based on the percentage of deep fakes contained in the data.

[1241] "Means for calculating the reliability score" refers to the hardware and software for calculating the reliability score based on the ratio of Deep fake detection results.

[1242] "Means for storing in a database" refers to a storage system and its management software for permanently storing analysis results and emotion data.

[1243] "Means of return" refers to software and communication means for displaying or notifying the user of the analysis results, reliability score, and emotional data.

[1244] MODE FOR CARRYING OUT THE INVENTION

[1245] This invention relates to a system in which a user uploads data using an electronic device, a server analyzes the data to detect Deep Fakes and calculate a reliability score, and also obtains the user's emotional state and returns a comprehensive analysis result.

[1246] The system includes the following major hardware and software:

[1247] User terminal

[1248] The user terminal is an electronic device such as a computer or smartphone that provides an interface for users to upload data. The user terminal is equipped with a camera and microphone, which can capture the user's facial expression and voice data. This data is analyzed by the emotion engine to identify the user's emotional state.

[1249] server

[1250] The server receives data sent by users and has a storage system for temporarily storing it. The received data is added to an analysis queue and registered as an analysis job.

[1251] Deep fake detection algorithm: The server implements a deep learning model to detect deep fakes in the data. This algorithm uses deep learning libraries such as TensorFlow and PyTorch.

[1252] Data analysis: The server uses ffmpeg to split the video into frames and then runs the Deep Fake detection process on each frame.

[1253] Calculation of confidence score: Based on the Deep Fake detection results, a score is calculated to evaluate the reliability of the entire data. This is provided as a confidence score ranging from 0% to 100%.

[1254] Database

[1255] The server is equipped with a database for permanently storing analysis results and user emotion data, making it possible to easily track past analysis results and display them again upon user request.

[1256] Data return

[1257] Once the analysis is complete, the server returns the confidence score, analysis results, and emotion data to the user, which are then displayed on the screen of the user's electronic device.

[1258] Specific examples

[1259] For example, consider a situation where user "B" wants to check the authenticity of a news video. User B accesses a web application, selects the news video file, and clicks the upload button. This action sends the video file from User B's computer to the server via an HTTP POST request.

[1260] The server receives the video file and temporarily stores it in storage. During this time, Person B's device uses the camera and microphone to capture Person B's facial expressions and voice data. The emotion engine analyzes this data and identifies Person B's emotional state (e.g., concern).

[1261] The server adds the received video file and the acquired emotion data to the analysis queue and runs the Deep Fake detection algorithm. Specifically, it divides the video into frames using "ffmpeg" and analyzes each frame using "TensorFlow." This analysis determines that 10% of the video is Deep Fake, and the confidence score is calculated as 90%.

[1262] The analysis results and emotion data are stored in a database along with the job ID. Finally, the server returns the analysis results, confidence score, and emotion data to Person B. Person B can check the analysis results, including the "concern" emotion state, along with the confidence score, in a web application.

[1263] Prompt Sentence Examples

[1264] "Please check the authenticity of this video and tell us how much Deepfake evidence there is."

[1265] Please send the Deepfake analysis results for the uploaded image along with the emotion data.

[1266] "Evaluate the reliability of the audio file and return the result. Emotional data on whether the user is concerned is also important."

[1267] This system allows users to accurately assess the reliability of the content they consume and provides comprehensive analysis based on emotional data.

[1268] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1269] Step 1:

[1270] Users upload content

[1271] Input: Data such as video, images, and audio files

[1272] How it works: A user logs into the web application, selects the file they want to analyze, and clicks the "Upload" button, which causes the user's device to send an HTTP POST request to the server with the selected file.

[1273] Output: Upload request sent to the server

[1274] Step 2:

[1275] The server receives the uploaded content

[1276] Input: HTTP POST request sent from the user terminal

[1277] Operation: The server receives the HTTP request and temporarily saves the sent file to the local disk. As an example, save the file to " / tmp / uploaded_files / video.mp4".

[1278] Output: Temporarily saved video file

[1279] Step 3:

[1280] Obtain the user's emotional data

[1281] Input: The user's facial expression and voice data

[1282] Operation: Use the camera and microphone of the user terminal to obtain the user's facial expression and voice in real time. Pass these data to the emotion engine, and the emotion engine analyzes and estimates the user's emotional state. For example, determine "suspense" from a "frowning expression".

[1283] Output: Obtained emotional data

[1284] Step 4:

[1285] The server adds to the analysis queue

[1286] Input: Temporarily saved video file, obtained emotional data

[1287] Operation: The server adds the received file and emotion data to the analysis queue and registers it as an analysis job. Specifically, it registers the file path " / tmp / uploaded_files / video.mp4" and the emotion data "concern" in the analysis queue.

[1288] Output: Jobs waiting to be analyzed

[1289] Step 5:

[1290] The server runs the Deep Fake detection algorithm

[1291] Input: Video files queued for analysis

[1292] How it works: The server splits the video into frames using ffmpeg and runs the Deep Fake detection algorithm on each frame. It then uses a deep learning model (e.g. TensorFlow) to determine whether or not there is a Deep Fake in each frame.

[1293] Output: Deep fake detection results for each frame

[1294] Step 6:

[1295] The server calculates the reliability score based on the detection result.

[1296] Input: Deep fake detection results for each frame

[1297] Operation: The detection results are aggregated and the overall confidence score is calculated based on the percentage of Deep fake detections. Specifically, if the Deep fake portion accounts for 10% of the total, a "confidence score of 90%" is calculated.

[1298] Output: Confidence Score

[1299] Step 7:

[1300] The server stores the analysis results, including emotional data.

[1301] Input: confidence score, acquired emotion data

[1302] Operation: The server saves the analysis results and emotion data in the database. It associates them with the job ID and saves "Video.mp4", "Confidence Score: 90%", and "Emotional State: Concern".

[1303] Output: Analysis results and emotion data stored in a database

[1304] Step 8:

[1305] The server returns the analysis results and emotion data to the user.

[1306] Input: Analysis results and emotion data stored in the database

[1307] Behavior: The server generates a report of the analysis results and returns it to the user. The user's web application displays the result "Video Confidence Score: 90%, Emotional State: Concern".

[1308] Output: Analysis results and emotion data displayed on the user's electronic device

[1309] (Application example 2)

[1310] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1311] Traditional deep fake detection systems focus on assessing the trustworthiness of content uploaded by users. However, these systems provide analysis results without considering the user's emotional state, and therefore cannot provide information about how users understand and perceive the trustworthiness of content. In particular, when the results of deep fake detection are ambiguous or have low confidence, this can burden users' understanding and behavior. To solve this problem and improve user experience, it is necessary to incorporate user emotional data into the analysis.

[1312] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[1313] In this invention, the server includes a means for users to upload content, a means for receiving the uploaded content and adding it to an analysis queue, a means for executing an algorithm to detect Deep Fakes contained in the content, a means for calculating a reliability score for the content based on the detection results, a means for executing an engine that acquires and analyzes emotional data from the user terminal, and a means for providing the user with feedback based on the emotional data along with the reliability score. This allows the user to receive feedback based on their own emotional state in addition to the reliability of the content, allowing the user to make decisions based on more comprehensive information.

[1314] The "means by which users upload content" refers to an interface that users use to send data such as video, image, and audio files to be analyzed to a server via the Internet.

[1315] The "means for receiving uploaded content and adding it to the analysis queue" is a function by which the server receives content sent by a user and adds the content to a waiting list for analysis processing.

[1316] "Means for executing an algorithm to detect Deep Fakes contained in content" refers to a program that allows the server to perform detailed analysis of each frame (video / image) or each second (audio) to identify traces of Deep Fakes.

[1317] "Means for calculating the reliability score of content based on the detection results" is a function that allows the server to use the detection results of Deep Fake to calculate a numerical value (reliability score) to evaluate the reliability of the entire content.

[1318] "Means for executing an engine that acquires and analyzes emotional data from a user terminal" refers to a program that acquires data using a camera or microphone and performs emotional analysis in order to estimate the emotional state of the user from their facial expressions and voice.

[1319] "Means for providing users with feedback based on emotional data together with the reliability score" is a function in which the server combines the reliability score with the user's emotional data to provide the analysis results to the user, providing feedback in a form that is easier for the user to understand.

[1320] This invention is a system that not only evaluates the reliability of content but also takes into account the emotional state of the user to support more comprehensive judgment. This system is mainly composed of a server and a user terminal.

[1321] The server and user terminal perform the following main operations:

[1322] First, the user device provides an interface for uploading the content (video, image, audio files) that the user wants to analyze. The user selects the file through a web application or a dedicated app and clicks the upload button. The user device then makes an HTTP POST request to send the selected file to the server.

[1323] The server then receives and temporarily stores the file sent from the user terminal, and adds the received content to the analysis queue, ready for analysis processing.

[1324] Furthermore, the emotion engine uses the camera and microphone installed on the user device to acquire emotion data from the user's facial expressions and voice, and analyzes this data to estimate the user's emotional state.

[1325] The server adds the uploaded file and the acquired emotion data to the analysis queue and registers it as an analysis job, which starts the subsequent analysis process.

[1326] The server runs the DeepFake detection algorithm on the content queued for analysis. This algorithm performs detailed analysis of each frame (video / image) and each second (audio) to detect traces of DeepFake.

[1327] The server calculates the reliability score for the entire content based on the Deep Fake detection results. For example, if 10% of the frames are determined to be Deep Fake, the reliability score is calculated as 60%. Furthermore, if the emotion engine recognizes the user's emotional state (happiness, sadness, anger, anxiety, etc.), that information is also added as feedback.

[1328] The analysis results include the reliability score of the content and the user's emotional data, which are stored in a database. The server returns these results in response to a user request. The results are then displayed, for example, on the screen of a web application.

[1329] A specific example would be analyzing images captured by a user during a meeting or video call. The user uploads the image to be analyzed, and the server analyzes the image. During the analysis, the user's facial expression is captured by a camera, and the emotional data is included in the analysis results. As a result, the emotional data "the user is concerned" is provided along with the confidence score of the image.

[1330] The following software and hardware are used to configure the emotion engine and Deep Fake detection algorithm functionality:

[1331] Open Source Computer Vision Library: OpenCV

[1332] Sentiment analysis engine: EmotionEngine

[1333] Deep fake detection algorithm API

[1334] This allows users to receive detailed information about the reliability of the content and feedback that includes their own emotional state, allowing them to make decisions based on more comprehensive information.

[1335] An example of a prompt to input to a generative AI model is as follows:

[1336] "Please confirm whether this image or video frame could have been generated by AI. Please return the analysis results and a confidence score. Below is the image data for verification."

[1337] "Please estimate the user's emotional state (happiness, sadness, anger, anxiety, etc.) from this image. Please return the analysis result and the emotional state. Below is the image data for testing."

[1338] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1339] Step 1: User uploads content

[1340] The user selects the content (video, image, audio file) they want to analyze through a web application or dedicated app and clicks the upload button. The user's device sends the selected file data to the server as an HTTP POST request. The input is the file data, and the output is the status of completion of transmission to the server.

[1341] Step 2: The server receives the uploaded content

[1342] The server receives files sent from the user terminal and temporarily stores them. The input is the file data sent from the user terminal, and the output is the storage path of the temporarily stored file. The specific operation is to receive the file and store it in the specified directory.

[1343] Step 3: Obtain user emotion data

[1344] The camera and microphone installed on the user device capture the user's facial expressions and voice, and analyzes the data. The emotion engine estimates the user's emotional state (happiness, sadness, anger, anxiety, etc.). The input is the captured image and voice data, and the output is numerical data related to the emotional state. Specific operations include activating the camera and microphone, capturing data, and analyzing the obtained data.

[1345] Step 4: Server adds to analysis queue

[1346] The server adds the received file and the acquired emotion data to the analysis queue and registers it as an analysis job. The input is the file data and emotion data, and the output is the status of completion of adding it to the analysis queue. The specific operation is to add the file and emotion data to the analysis queue.

[1347] Step 5: The server runs the Deep Fake detection algorithm

[1348] The server runs the DeepFake detection algorithm on the content queued for analysis. This algorithm performs detailed analysis of each frame (video / image) and each second (audio) to detect DeepFake signatures. The input is the file data to be analyzed, and the output is the DeepFake detection results and a confidence score. Specific operations include analyzing each frame and each second individually to detect specific patterns and features.

[1349] Step 6: The server calculates the trust score

[1350] The server calculates the reliability score of the entire content based on the detection results of Deep Fake. The input is the detection result data of Deep Fake, and the output is a numerical reliability score. Specific operations include executing an algorithm that sums up the detection results and evaluates the overall reliability.

[1351] Step 7: The server stores the analysis results and emotion data

[1352] The server stores the analysis results, confidence scores, and the acquired emotion data in a database. The input is the analysis result data, confidence score data, and emotion data, and the output is the status of completion of saving to the database. Specifically, each dataset is linked to a job ID and stored in the database.

[1353] Step 8: The server returns the analysis results and emotion data to the user.

[1354] The server returns the analysis results, confidence score, and user emotion data to the user. The input is the analysis result data, confidence score data, and emotion data, and the output is data to be displayed to the user. Specific operations include data format conversion and transmission to display the results on the screen of the web application used by the user.

[1355] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.

[1356] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1357] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.

[1358] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[1359] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.

[1360] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.

[1361] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).

[1362] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.

[1363] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."

[1364] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values ​​indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.

[1365] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).

[1366] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.

[1367] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.

[1368] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[1369] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.

[1370] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.

[1371] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.

[1372] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.

[1373] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.

[1374] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.

[1375] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.

[1376] The following is further disclosed regarding the above embodiment.

[1377] (Claim 1)

[1378] a means for users to upload content;

[1379] a means for receiving the uploaded content and adding it to an analysis queue;

[1380] a means for executing an algorithm to detect Deep Fakes in the content;

[1381] A means for calculating a reliability score of the content based on the detection result;

[1382] A system including a means for returning a reliability score and analysis results to the user.

[1383] (Claim 2)

[1384] The system of claim 1, wherein the reliability score of the entire content is calculated based on the rate of detected deep fakes.

[1385] (Claim 3)

[1386] 2. The system according to claim 1, wherein the analysis results are stored in a database and returned to the user upon request.

[1387] "Example 1"

[1388] (Claim 1)

[1389] a means for users to upload content;

[1390] a means for receiving the uploaded content and adding it to an analysis queue;

[1391] a means for executing an algorithm to detect Deep Fakes in the content;

[1392] A means for calculating a reliability score of the content based on the detection result;

[1393] A means for returning the reliability score and analysis results to the user;

[1394] The system includes a means for displaying the analysis results.

[1395] (Claim 2)

[1396] The system of claim 1, wherein the reliability score of the entire content is calculated based on the rate of detected deep fakes.

[1397] (Claim 3)

[1398] 2. The system according to claim 1, wherein the analysis results are stored in a database and returned to the user upon request.

[1399] "Application Example 1"

[1400] (Claim 1)

[1401] a means for users to upload content;

[1402] a means for receiving the uploaded content and adding it to an analysis queue;

[1403] a means for executing an algorithm to detect Deep Fakes in the content;

[1404] A means for calculating a reliability score of the content based on the detection result;

[1405] A means of analyzing a company's audio or video conference recordings to detect whether deepfake technology is being used; and

[1406] A system including a means for returning a reliability score and analysis results to the user.

[1407] (Claim 2)

[1408] The system of claim 1, wherein the reliability score of the entire content is calculated based on the rate of detected deep fakes.

[1409] (Claim 3)

[1410] 2. The system according to claim 1, wherein the analysis results are stored in a database and returned to the user upon request.

[1411] "Example 2: Combining Emotion Engines"

[1412] (Claim 1)

[1413] means for a user to upload data using an electronic device;

[1414] a means for receiving the uploaded data and adding it to an analysis queue;

[1415] A means for acquiring and analyzing user emotion data;

[1416] A means for running an algorithm to detect deep fakes in the data; and

[1417] A means for calculating a reliability score of the data based on the detection result;

[1418] a means for storing the analysis results including the emotion data in a database;

[1419] A system including a means for returning a reliability score and analysis results to the user.

[1420] (Claim 2)

[1421] The system according to claim 1, wherein the reliability score of the entire data is calculated based on the rate of detected deep fakes.

[1422] (Claim 3)

[1423] 2. The system of claim 1, wherein the analysis results and emotion data are stored in a database and the results are returned in response to a user request.

[1424] "Application example 2 when combining emotion engines"

[1425] (Claim 1)

[1426] a means for users to upload content;

[1427] a means for receiving the uploaded content and adding it to an analysis queue;

[1428] a means for executing an algorithm to detect Deep Fakes in the content;

[1429] A means for calculating a reliability score of the content based on the detection result;

[1430] means for executing an engine that acquires and analyzes emotion data from a user terminal;

[1431] A system including a means for providing a user with feedback based on emotional data along with a confidence score.

[1432] (Claim 2)

[1433] The system of claim 1, wherein the reliability score of the entire content is calculated based on the rate of detected deep fakes.

[1434] (Claim 3)

[1435] 2. The system of claim 1, wherein the analysis results and emotion data are stored in a database and the results are returned in response to a user request. [Explanation of symbols]

[1436] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>

Claims

1. a means for users to upload content; a means for receiving the uploaded content and adding it to an analysis queue; a means for executing an algorithm to detect deepfakes in content; and A means for calculating a reliability score of the content based on the detection result; A system including a means for returning a reliability score and analysis results to the user.

2. The system of claim 1, wherein the reliability score of the entire content is calculated based on the rate of detected Deepfakes.

3. 2. The system according to claim 1, wherein the analysis results are stored in a database and returned to the user upon request.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A