Information verification system, information verification method, and information verification program

The system uses AI to analyze multimedia content's acoustic and prosodic features, comparing it with official sources to detect and quantify manipulation, effectively preventing misinformation by identifying and tracking malicious accounts.

JP7843984B1Active Publication Date: 2026-04-13加藤 健資
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
加藤 健資
Filing Date
2025-12-25
Publication Date
2026-04-13

AI Technical Summary

Technical Problem

Existing information verification technologies struggle to accurately detect sophisticated manipulations in multimedia content, such as videos and audio, which can distort the original context and intent by altering tone, deleting pauses, or emphasizing specific words, making it difficult to prevent the spread of misinformation.

Method used

An information verification system that uses AI and machine learning to analyze prosodic and acoustic characteristics, including pitch, volume, and pauses, compares content with official sources, and calculates a modification score to detect and quantify manipulation.

Benefits of technology

Enables high-accuracy detection of information distortion from multiple angles, allowing for early identification and mitigation of misinformation, and tracks accounts spreading manipulated content.

✦ Generated by Eureka AI based on patent content.
Patent Text Reader

Abstract

This invention relates to an information verification system that detects and evaluates information alteration and contextual distortion by comparing video and audio content disseminated on social media, etc., with official information (correct information). It is characterized by its use of AI speech recognition and natural language processing for transcription, and its analysis of audio features such as pitch, speaking speed, pauses, and emotion. By comprehensively analyzing not only text differences but also non-verbal elements and contextual differences, and scoring the degree of alteration, it can accurately detect sophisticated information distortions such as selective editing and manipulation of impressions, contributing to the realization of healthy information distribution.
Need to check novelty before this filing date? Find Prior Art

Description

[Background technology]

[0001] This invention relates to information processing technology, particularly technology for verifying the reliability of digital content. More specifically, this invention relates to a system, method, and program for automatically detecting and evaluating alterations, editing, and distortions of context in media content such as videos and audio disseminated on platforms such as social networking services (SNS), based on officially published sources (hereinafter referred to as "true information" or "official information"). This invention is characterized by its use of AI (artificial intelligence) technology, particularly speech analysis technology, natural language processing technology, and machine learning technology, to comprehensively analyze not only text information but also non-verbal information inherent in the sound itself, such as speaking style, tone of voice, emotion, degree of emphasis, and pauses between utterances, and to evaluate the authenticity of information and changes in intent from multiple perspectives. As a result, this invention contributes to preventing the spread of fake news, deepfakes, misinformation, and disinformation, improving media literacy, and realizing a fair information distribution environment.

[0002] In recent years, with the explosive spread of social networking services (SNS), an environment has been realized where anyone can easily transmit and share information. As a result, the speed of information circulation has leaped forward, bringing the benefit that diverse opinions and knowledge are spread throughout society. At the same time, however, the new issue of ensuring the reliability of information has become more serious. In particular, multimedia content including videos and audio has a large amount of information and high appeal, so it has a great impact on people's opinion formation and public opinion. However, because of its great influence, there are constantly cases where a malicious third party intentionally cuts out part of the content, edits or processes it, severely distorts the original context and the intention of the speaker, fosters misunderstandings and prejudices against specific individuals or organizations, or causes social chaos. Such "clipped videos" and "edited content" give a strong visual and auditory impression, so once they are spread, their content is easily accepted as if it were a fact, and it has the characteristic that it is extremely difficult to correct later.

[0003] Conventionally, as a technology to address such problems of information distortion, mainly comparison of text information and fact-checking have been central. For example, a method is known in which specific keywords or speech contents are extracted as text and it is collated whether they match official announcements or reliable information sources. However, since these methods target only information expressed as text, there is a limit to detecting sophisticated information manipulation. For example, even if the speech content is the same, by changing the tone of the voice, deleting important expressions of silence or hesitation between speeches, or unnaturally emphasizing specific words, the impression received by the listener changes greatly. Furthermore, by cutting out only a part of the speech and deleting important contexts before and after (for example, preconditions, reservations, opposing opinions, etc.), it is also possible to make it seem as if it is a conclusion completely opposite to the original meaning. These "impression manipulations" and "context cuttings" are extremely difficult to detect by only comparing text information, and contain the risk of overlooking the essential problem of information distortion. Therefore, the establishment of a more advanced and multi-faceted information verification technology has been strongly demanded. [Prior art documents] [Patent Documents]

[0004] Patent No. 7742067 [Overview of the project] [Problems that the invention aims to solve]

[0005] This invention has been made in view of the problems of the prior art described above, and its purpose is to provide an information verification system, information verification method, and information verification program that can detect information distortion in video and audio content disseminated on social media and the like with higher accuracy and from multiple perspectives. Specifically, this invention aims to go beyond simply determining match or mismatch at the text level, and comprehensively analyze the prosodic and acoustic characteristics of the audio itself, as well as nonverbal information such as the speaker's emotions and the context of their speech, thereby effectively detecting sophisticated manipulation of impressions and intentional contextual omissions that were difficult to detect with the prior art, and to evaluate the reliability of content based on objective indicators. Furthermore, this invention also aims to identify the source of false information and contribute to the healthy flow of information by identifying accounts that repeatedly post and disseminate content that has been altered or edited in this way and monitoring their activities. [Means for solving the problem]

[0006] To solve the above problems, an information verification system according to one aspect of the present invention comprises at least one processor and at least one memory connected to the processor. The processor is characterized in that it functions as a content acquisition unit that acquires target content to be verified from a predetermined platform including a social networking service; an audio extraction unit that extracts audio data from the target content; a transcription unit that generates text data indicating the content of speech and time information corresponding to the text data from the audio data using an AI speech recognition model; an audio feature analysis unit that extracts at least one or more audio features from the audio data, including pitch, volume, speaking speed, and pauses; a positive information database that stores positive information content derived from official sources; a difference analysis unit that compares the target content with one or more positive information contents stored in the positive information database and analyzes the difference between the two; and a score calculation unit that calculates a modification score indicating the degree of modification of the target content based on the analysis results by the difference analysis unit. The aforementioned positive information database pre-stores audio data, text data, and audio features of the positive information content, and the difference analysis unit can be configured to perform analysis that includes at least differences between text data, differences between audio features, and contextual differences. [Effects of the Invention]

[0007] According to the present invention, it becomes possible to detect information distortion in video and audio content with high accuracy and from multiple angles. In particular, by comparing and analyzing the characteristic quantities of the audio itself (pitch, volume, speaking speed, pauses, emotional tendencies, etc.) with correct information, it becomes possible to quantitatively evaluate, based on objective indicators, sophisticated editing that manipulates impressions even if the text content is the same, and malicious cutouts that disregard context, which were difficult to do with conventional technology. This makes it possible to detect the spread of misinformation and false information at an early stage and minimize its impact. Furthermore, by tracking and evaluating the activities of accounts that generate and spread altered content, it becomes possible to take measures against the source of information manipulation, resulting in a significant effect that contributes to the realization of a healthier and more reliable information distribution environment. [Modes for carrying out the invention]

[0008] (Physical configuration of the information verification system) An information verification system according to one embodiment of the present invention physically consists of at least one server device, client devices, and a network connecting them. The server device is the core part that performs the main processing of the system and may be built in a cloud computing environment or an on-premises environment. The client device provides an interface for users to use the system and includes, for example, a personal computer, smartphone, or tablet terminal. The network includes wired or wireless communication networks such as the Internet, intranet, LAN (Local Area Network), and WAN (Wide Area Network).

[0009] (Server hardware configuration) A server device is a computer equipped with at least one processor (CPU), main memory (RAM), secondary storage (e.g., SSD or HDD), and a communication interface. The processor controls the operation of the entire server device by reading the operating system and application programs stored in the secondary storage into the main memory and executing them. In particular, to perform computationally intensive processes such as speech analysis, natural language processing, and machine learning model inference in this invention, it is highly preferable that the server device be equipped with one or more dedicated AI accelerators such as a GPU (Graphics Processing Unit), TPU (Tensor Processing Unit), or FPGA (Field-Programmable Gate Array). The main memory functions as the processor's working memory and preferably has a capacity of tens of gigabytes to several terabytes. The secondary storage is for permanently storing programs and large amounts of data, and uses SSDs conforming to the NVMe (Non-Volatile Memory Express) standard that enable high-speed reading and writing, or HDDs for economically storing large amounts of data. The communication interface is for communicating with other devices via a network, and a high-speed interface of 10 Gigabit Ethernet (10GbE) or higher is desirable.

[0010] (Functional configuration of server equipment) The server system can be functionally separated into web servers, application servers, database servers, analysis servers, etc. These functions may be implemented on a single physical server or distributed across multiple physical servers. The web server receives HTTP requests from client devices, returns static content, and forwards requests to application servers. The application server implements the main business logic of the present invention and handles content retrieval, analysis instructions, and result management. The database server manages the positive information database, analysis results database, user account information, etc. The analysis server specializes in processing that consumes the most computing resources, such as speech recognition, speech feature extraction, and differential analysis. By adopting such a microservices architecture, the independence of each function is increased, making development, deployment, and scaling easier.

[0011] (Client device configuration) The client device provides the system's functions to the user via a web browser or dedicated application software. The client device includes a processor, memory, storage, display, input devices (keyboard, mouse, touch panel, microphone), and a communication interface. The user enters the URL of the content to be verified and views the analysis results via a GUI (Graphical User Interface) displayed on the screen. An interface for inputting voice commands via a microphone may also be included. The client device communicates with the server device to request processing and receive results.

[0012] Embodiments of the present invention will be described in detail below. Note that the following embodiments are not limiting to the present invention, and various modifications are possible within the scope of the technical idea of ​​the present invention. Furthermore, in each embodiment, the same components are denoted by the same reference numerals, and redundant descriptions may be omitted. In at least one embodiment, the information verification system according to the present invention can be realized by a general-purpose computer system, a cloud computing environment, a distributed system, or a combination thereof. The system may include at least one processor, main memory (e.g., RAM), auxiliary storage (e.g., hard disk drive, solid-state drive), input devices (e.g., keyboard, mouse, touch panel, microphone), output devices (e.g., display, speaker), and a communication interface. These components are interconnected via a communication path such as a bus. The processor loads a program stored in the auxiliary storage into the main memory and executes it, thereby operating as one of the functional units described later (e.g., content acquisition unit, audio extraction unit, transcription unit, audio feature analysis unit, difference analysis unit, score calculation unit, etc.). These functional units may be implemented as a single program module executed on a single processor, or as multiple program modules executed in parallel on multiple processors. Furthermore, dedicated hardware (e.g., ASICs, FPGAs) can be used to implement specific functions. In particular, for computationally intensive processes such as speech feature extraction and AI model inference, it is preferable to use accelerators such as GPUs (Graphics Processing Units) and TPUs (Tensor Processing Units).

[0013] (First embodiment) The overall configuration of the information verification system according to the first embodiment of the present invention will now be described. The system of this embodiment is broadly composed of three layers: a "data collection layer" that collects content to be verified from platforms such as SNS, a "data analysis layer" that analyzes the collected content and official information, and an "interface layer" that presents the analysis results to the user and cooperates with external systems. The data collection layer includes a content acquisition unit and continuously collects videos, audio, text, metadata, etc., using various SNS APIs (Application Programming Interfaces) and web scraping technologies. The data analysis layer is the core part of this system and is equipped with multiple functional modules such as an audio extraction unit, a transcription unit, an audio feature analysis unit, a difference analysis unit, a score calculation unit, and an account evaluation unit. This layer is also closely linked to a positive information database for storing positive information content collected from official sources. The positive information database is constructed by selecting the most suitable database management system (DBMS) according to the type of data to be stored and the access pattern, such as a relational database, a NoSQL database, or a vector database. The interface layer comprises a visualization / UI section and an API provision section, providing a dashboard that displays analysis results in a format that humans can intuitively understand, and APIs that make the analysis results available to other systems. Each of these layers and functional sections may be implemented centrally on a single physical server, or they may be distributed across multiple servers depending on the function. For example, when dealing with large-scale data or when high availability and scalability are required, adopting a microservices architecture, containerizing each function as an independent service (e.g., using Docker), and managing and operating it on a container orchestration system (e.g., Kubernetes) is extremely effective.Furthermore, by fully utilizing cloud computing environments that can dynamically allocate and release computing and storage resources (for example, Amazon Web Services (AWS), Microsoft Azure, and Google Cloud Platform (GCP)), it is possible to optimize operational costs while ensuring scalability.

[0014] The hardware configuration in this embodiment will now be described in more detail. This system operates on a server computer equipped with at least one central processing unit (CPU). This server computer has main memory (RAM) with a capacity of several gigabytes to several hundred gigabytes or more, and auxiliary storage (HDD, SSD) capable of storing large amounts of data, from several terabytes to several petabytes or more. In particular, a storage system that is capable of high-speed reading and writing and is highly scalable is required to store unstructured data such as audio data and video data, as well as training data and parameters for AI models. Suitable solutions for this include RAID configurations using multiple NVMe (Non-Volatile Memory Express) connected SSDs, distributed file systems (e.g., HDFS), and object storage (e.g., Amazon S3). Furthermore, in order to speed up the training and inference processing of AI models, this system is equipped with at least one, preferably multiple, GPUs or other AI accelerators. These accelerators excel at efficiently performing large amounts of parallel computation, and in particular, they achieve an order of magnitude improvement in performance compared to using only a CPU, especially in deep neural network calculations. The network configuration should include high-speed network interfaces of Gigabit Ethernet (GbE) or 10 Gigabit Ethernet (10GbE) or higher, connecting to the internet and the internal network. Depending on the scale of the system, it is desirable to implement load balancing using load balancers and security devices such as firewalls, intrusion detection systems (IDS), and intrusion prevention systems (IPS). Furthermore, instead of configuring the entire system with a single physical server, a configuration can be adopted in which multiple physical servers are clustered to ensure high availability and high scalability. In this case, it is common to use virtualization technology (e.g., VMware, KVM) to efficiently divide and allocate physical resources, and to automate application deployment and management using the aforementioned container technology.

[0015] (Second embodiment) Next, as a second embodiment of the present invention, a function for collecting content to be verified from SNS etc. (data collection layer) will be described in detail. The content acquisition unit in this embodiment includes a data collection module that corresponds to major SNS platforms (for example, X (formerly Twitter), YouTube®, Facebook, Instagram, TikTok, Twitch, etc.). Using the official API provided by each platform is the most stable and recommended method. For example, using the YouTube® Data API v3, it is possible to obtain a list of videos from a specific channel, video metadata (title, description, tags, publication date and time, etc.), comments, etc. Also, using the X API, it is possible to stream and obtain posts containing specific keywords, posts from specific accounts, the number of retweets and "likes" of posts, etc. in near real time. However, there may be constraints on the use of APIs, such as limits on the number of requests, usage fees, and the range of data that can be obtained. Therefore, in this embodiment, web scraping technology can also be used in combination as a means to complement APIs or as a means to correspond to platforms for which APIs are not provided. Web scraping is a technique that sends HTTP requests to obtain the HTML content of a web page and extracts the necessary information from it. In this case, libraries such as Beautiful Soup and Scrapy can be used to analyze the HTML structure. Furthermore, to handle modern websites where content is dynamically generated by JavaScript, it is effective to use tools that manipulate headless browsers, such as Selenium, Playwright, and Puppeteer. These tools launch a web browser in the background, render the page, and then retrieve the content, thus ensuring more reliable data collection.The content collected includes a wide range of items, such as video files themselves (e.g., MP4, WebM format), audio files (extracted from videos, or audio-only content), post text, posting date and time, poster information (account name, ID), and engagement figures (views, likes, shares, comments). The collected data is assigned a unique ID for subsequent analysis and systematically stored in a database or storage system along with information such as the source platform and collection date and time.

[0016] (Third embodiment) Next, as a third embodiment of the present invention, an AI speech recognition and transcription function that generates text information from collected audio data will be described in detail. The transcription unit in this embodiment is characterized by a hybrid configuration that selectively or integrally utilizes multiple state-of-the-art AI speech recognition models. This achieves high recognition accuracy and robustness for diverse acoustic environments, speakers, languages, and technical terms. An example of a speech recognition model available in this embodiment is the Whisper model developed by OpenAI. Whisper is a large-scale transformer-based model trained using a vast amount of diverse audio data collected from the internet, and it exhibits high performance in multilingual speech recognition, translation, and language identification tasks without fine-tuning specific to a particular language or task. Its robustness is particularly effective for real-world audio data that includes background noise, accents, and technical terms. Speech recognition services provided by major cloud providers such as Google Cloud Speech-to-Text, Microsoft Azure Speech Services, and Amazon Transcribe can also be used. These services offer the advantage of continuously updated models and easy access to advanced features via APIs, such as speaker diarization, custom vocabulary registration, and automatic punctuation insertion. Furthermore, the system also includes the ability to build or fine-tune custom speech recognition models using open-source speech recognition toolkits (e.g., Kaldi, ESPnet) to pursue speech recognition performance specific to particular domains (e.g., politics, economics, medicine). In this case, the acoustic model could be a hybrid model combining Hidden Markov Models (HMMs) and Deep Neural Networks (DNNs), or an end-to-end model that generates text directly from acoustic signals (e.g., Connectionist Temporal Classification (CTC) based, Attention based, RNN-Transducer based).As a language model, in addition to conventional N-gram models, the use of transformer-based large-scale language models (LLMs) such as BERT and GPT, which have superior contextual understanding capabilities, can dramatically improve the accuracy of homonym discrimination and technical term recognition. As a preprocessing step for speech data, the input quality to the speech recognition model is optimized and recognition errors are reduced by applying processes such as noise reduction (e.g., spectral subtraction method, deep learning-based noise reducer), volume normalization, and sampling rate unification.

[0017] The speech-to-text process is not only about simply converting speech to text, but also a process of adding important information that greatly affects the accuracy of subsequent differential analysis. The speech-to-text part of this embodiment precisely generates timestamp information including the start time and end time corresponding to each word or clause from the output of the speech recognition model. This timestamp information is essential for enabling synchronization between text and speech and for accurately identifying "which utterance corresponds to which part of the video." This enables the calculation of context omission rates and the association with audio feature quantities, which will be described later. Furthermore, for content such as meetings and interviews where there are multiple speakers, speaker diarization technology is applied. Speaker diarization is a technology that detects speaker change points from the audio signal and clusters which speaker each utterance segment belongs to. For this, techniques for vectorizing the characteristics of a speaker's voice (e.g., i-vector, x-vector, d-vector) are used. By performing speaker diarization, it becomes possible to clearly distinguish "who said what," and it becomes possible to track and analyze only the utterances of a specific person. For the generated text data, a punctuation estimation model is used to automatically insert punctuation marks (e.g., comma (,), period (.), question mark (?), exclamation mark (!), etc.) to improve readability. Also, the speech recognition model can output a confidence score for the recognition result. This score is an indicator showing how confident the model is in its recognition result, suggesting that parts with a low score are more likely to be recognition errors. In this system, this confidence score is recorded, and for parts where the score is below a predetermined threshold, processes such as prompting the user for confirmation or re-verifying with another speech recognition model can be activated. The structured text data (text, timestamp, speaker information, confidence score, etc.) generated through these processes is stored in the positive information database and used as basic information for differential analysis.

[0018] (The Fourth Embodiment) Next, as a fourth embodiment of the present invention, a speech feature analysis function for extracting nonverbal information from speech data will be described in detail. In human communication, the way of speaking and the tone of voice (nonverbal information) are just as important, if not more so, than the content of the words themselves (linguistic information). The speech feature analysis unit of this embodiment extracts a wide range of acoustic and prosodic features from speech data in order to quantitatively capture such nonverbal information. These features are later compared with positive information speech and become key to detecting sophisticated editing such as impression management and emotional distortion. One of the extracted features is the fundamental frequency (F0) and its statistics (mean, standard deviation, maximum, minimum, and range of variation) related to the pitch of the voice. Generally, when emotions are heightened, the pitch of the voice tends to rise, and when calm, it tends to fall, and fluctuations in F0 reflect the intonation and pitch of speech. Next, as features related to the volume of the voice, sound pressure level, loudness, and short-time energy are extracted. When emphasizing a particular part of a speech, the volume increases, so capturing this change is important for measuring the degree of emphasis. In addition, features related to speaking speed, such as speech rate (WPM: Words Per Minute or mora / second) and phoneme duration, are calculated. Editing to intentionally speed up or slow down the speech rate can give the listener different impressions, such as anxiety or composure. Furthermore, features related to pauses and pauses in speech, such as the frequency, length, and location of pauses (silent intervals), are detected. Removing pauses before and after important statements can reduce the weight of the statement or alter the context, so detecting them is extremely important. In addition to these basic features, more advanced features that reflect the timbre and quality of the voice are also extracted. For example, Mel-frequency cepstrum coefficients (MFCCs) capture the features of the spectral envelope that take into account the characteristics of human hearing and are widely used in speech recognition and speaker recognition. Furthermore, jitter (frequency fluctuations) and shimmer (amplitude fluctuations), which indicate periodic disturbances in vocal cord vibration, can serve as indicators of vocal tremor and tension. By extracting these diverse features as time-series data in short frame units of a few milliseconds to tens of milliseconds and analyzing their temporal change patterns, it becomes possible to capture the subtle nuances embedded in speech.

[0019] In this embodiment, in addition to the basic acoustic features described above, higher-order prosodic and emotional features are extracted. Prosody refers to the elements of speech that complement and modulate the meaning of sentences and words, such as intonation, accent, rhythm, and prominence. The speech feature analysis unit of this embodiment estimates intonation by analyzing the contour pattern of the fundamental frequency (F0), the position and intensity of accent from the degree of energy concentration, rhythm from the duration patterns of phonemes and syllables, and the degree of emphasis (prominence) of specific words and phrases from the combined changes of F0, energy, and duration. These prosodic features are extremely important for understanding the intent of a statement (e.g., declaration, question, command) and where the speaker is placing emphasis. For example, even the same response "yes" can have vastly different meanings depending on whether the intonation is rising or falling, indicating a question or agreement. If such prosodic features are unnaturally altered in a clipped video, it suggests that the original intent of the statement may have been distorted. Furthermore, this system introduces emotion recognition technology to estimate the speaker's emotional state from the speech. Speech emotion recognition is a technology that classifies basic emotion categories such as joy, anger, sadness, fear, surprise, disgust, and neutrality from the acoustic and prosodic features of speech, or represents the emotional state as coordinates in a two-dimensional space of arousal (excited or calm) and valence (positive or negative). For this purpose, emotion recognition models based on machine learning or deep learning are used. For example, emotions are estimated using models trained with support vector machines (SVMs), convolutional neural networks (CNNs), recurrent neural networks (RNNs), and their derivatives such as LSTM (Long Short-Term Memory), taking extracted speech features (for example, standard feature sets such as the aforementioned GeMAPS (Geneva Minimalistic Acoustic Parameter Set)) as input.If calm statements in official audio clips are altered and edited in edited videos to evoke anger or anxiety, this emotion recognition function can detect this "emotional distortion" and use it as an indicator of impression manipulation. All of these extracted audio features are associated with timestamps and stored as time-series data in the positive information database.

[0020] (Fifth embodiment) Next, as a fifth embodiment of the present invention, an official audio database for managing positive information content that serves as a reference for comparison will be described in detail. This database is a fundamental component for ensuring the accuracy and reliability of differential analysis by this system. Its design requires a high degree of data integrity, consistency, searchability, and scalability. In this embodiment, it is preferable to adopt a polyglot persistence approach, which combines multiple database systems with different characteristics, rather than using a single monolithic database for the official audio database. Specifically, first, the video and audio files themselves related to official announcements (e.g., government press conferences, corporate earnings announcements, official statements by prominent figures, etc.) are stored in scalable and highly durable object storage such as Amazon S3, Google Cloud Storage, or Azure Blob Storage. These object storages are suitable for securely storing large volumes of unstructured data at low cost. Next, structured data related to each piece of content, such as title, presenter, announcement date and time, source URL, and content summary, are stored in a relational database (RDBMS) with excellent transaction processing capabilities, such as MySQL®, PostgreSQL, or Microsoft SQL Server. This enables the efficient execution of complex search queries while maintaining data integrity. Furthermore, the core data of this system—transcribed text data, timestamp information, speaker information, and a wide variety of speech features extracted as time-series data (pitch, volume, speech rate, MFCC, sentiment score, etc.)—is stored in a database specifically designed for handling these data formats. For example, text data and metadata can be indexed in full-text search engines such as Elasticsearch and Apache Solr, enabling high-speed keyword searches and similar document searches. In addition, high-dimensional time-series data such as speech features can be stored in time-series databases such as InfluxDB and TimescaleDB, enabling high-speed data aggregation and analysis within specified time ranges.In particular, the word and sentence vector embeddings used in the calculation of semantic similarity, which will be discussed later, are extremely effective when stored in vector databases such as Pinecone, Milvus, and Weaviate. Vector databases specialize in nearest-neighbor searches in high-dimensional vector spaces, dramatically speeding up the process of instantly searching for semantically similar content.

[0021] To maintain the quality of the official audio database, automation and systematization of the data collection and updating process are essential. In this embodiment, crawler and agent programs operate to periodically monitor pre-registered reliable sources (e.g., official government websites, corporate IR information pages, news organizations' official YouTube® channels, etc.). These programs automatically detect and incorporate new content (videos, audio, press releases, etc.) as it is published. The incorporated content is immediately fed into the analysis pipeline, where a series of processes such as audio extraction, transcription, and audio feature extraction are performed. The generated data (audio files, text data, feature data, etc.) is assigned a unique content ID and version number and stored in the appropriate databases mentioned above. Version control is important to accommodate cases where official information is later corrected or updated. For example, if the minutes of a press conference are revised later, the old version of the data is retained while the new version is added, and the relationship between the two is recorded. This makes it possible to track which version of the official information was used as the basis for verification. Furthermore, to prevent the duplicate collection of identical content from different sources, a mechanism will be implemented to calculate the hash value (e.g., SHA-256) of the content and perform duplicate removal. All data stored in the database will be managed along with detailed provenance information, including its source of truth, collection date and time, processing module, and version history. This ensures data traceability, and in the event of any doubts about the analysis results, the cause can be quickly tracked and identified. By establishing such a rigorous data management system, the official audio database will function as the foundation of the reliability of the entire system.

[0022] (Sixth embodiment) Next, as a sixth embodiment of the present invention, a difference analysis function that comprehensively analyzes the differences between target content and correct information content will be described in detail. The difference analysis unit of this embodiment detects differences from three different aspects: text, audio, and context, and integrates them to evaluate the possibility of information distortion. First, in text difference analysis, a comparison is made between transcribed text data. Rather than simply checking for matches or mismatches of strings, more advanced analysis methods are used. For example, an edit distance algorithm (e.g., Levenshtein distance, Damerau-Levenshtein distance) is used to calculate the number of word insertions, deletions, and substitutions, and to quantify the extent to which the text has been changed. In addition, to consider the rearrangement of word order, the degree of n-gram overlap (e.g., BLEU score) and the similarity of word sets (e.g., Jacquard coefficient, cosine similarity) are also calculated. However, these methods are merely superficial comparisons of strings and cannot capture semantic changes such as substitution with synonyms or paraphrasing. Therefore, in this embodiment, semantic similarity analysis is introduced. This involves using pre-trained language models such as BERT (Bidirectional Encoder Representations from Transformers), Sentence-BERT, and Universal Sentence Encoder. These models can convert words and sentences into high-dimensional vectors (embedding vectors) that reflect their contextual meaning. By vectorizing the sentences of the target content and the positive information content, and calculating the cosine similarity between these vectors, it is possible to evaluate how semantically similar they are, even if the strings are different. For example, "The government will reconsider the plan" and "The administration will review the plan" are different words, but very similar semantically, so they will receive a high similarity score. This semantic similarity analysis makes it possible to detect distortion caused by clever paraphrasing.

[0023] In addition to text difference analysis, the difference analysis unit of this embodiment analyzes the differences in the feature quantities of the audio itself. This allows for the detection of editing that manipulates the listener's impression, even if there are no changes in the text. Various feature quantities (pitch, volume, speaking speed, pauses, etc.) extracted as time-series data by the audio feature quantity analysis unit are compared between the target content and the positive information content. Dynamic Time Warping (DTW) is effective for this comparison. DTW is an algorithm that can appropriately evaluate the similarity of two time-series patterns even if the speaking speed fluctuates naturally, and is widely used in speech recognition and other applications. DTW is applied to the feature quantity time series of the target audio and the positive information audio, and the distance (similarity) between them is calculated. If this distance is large, it means that the audio features are significantly different, and it is highly likely that some kind of editing has been done. Specifically, the time series of pitch (voice height) is compared to detect whether the average pitch has risen or fallen unnaturally, or whether the intonation pattern has changed. Similarly, the time series of volume (loudness) is compared to detect whether specific words have been unnaturally emphasized or, conversely, made quieter. The system compares time-series data of speech rates to detect whether the overall or partial speaking pace has been altered. In particular, pauses (periods of silence) between statements play a crucial role in conveying nuances. If pauses that should be present in the original audio are removed or unnaturally shortened in the target audio, it is highly likely that this is intentional editing intended to make the context rushed or to erase nuances such as the speaker's hesitation or indecision. This system compares the presence, location, and length of pauses to detect such manipulations of "pauses." Furthermore, it is possible to capture the spectrogram of the audio signal as an image and detect "traces of editing," such as discontinuities and noise caused by editing operations like cutting, pasting, and speed changes, using a deep learning-based anomaly detection model.

[0024] Furthermore, the difference analysis unit of this embodiment analyzes contextual differences. This is extremely important for detecting information distortion caused by taking only a portion of a statement out of context. First, it identifies which part of the overall correct information content the target content corresponds to. This can be done quickly using partial matching search of the transcribed text or speech fingerprinting technology. Once the corresponding section is identified, the system retrieves from the correct information database what statements existed before and after the excerpted portion. Then, it evaluates the semantic relationship between the excerpted portion and the context before and after it. For example, even if the excerpted statement is the conclusion "the plan should be promoted," if there is an important conditional clause immediately before it such as "if sufficient budget is secured and public understanding is obtained," then deleting this conditional clause and presenting only the conclusion is a clear distortion of context. In this system, the extent to which such deleted context (conditions, reservations, objections, background explanations, etc.) alters the meaning of the excerpted portion is evaluated using natural language processing techniques such as dependency syntactic analysis and semantic role labeling. Then, the "contextual coverage rate" is calculated as the ratio of the amount of speech used in the target content to the total amount of speech (time or number of characters) in the positive information content, and conversely, the importance of the unused parts is taken into account and quantified as the "contextual gap rate." The results of these analyses from the text, audio, and contextual aspects are passed to the next scoring unit for a comprehensive evaluation.

[0025] (Seventh Embodiment) Next, as a seventh embodiment of the present invention, a modification score calculation function that integrates the results of difference analysis and calculates the degree of modification of the target content as an objective indicator will be described in detail. The various analysis results obtained by the difference analysis unit (for example, text edit distance, semantic similarity, DTW distance of audio features, number of pause deletions, contextual loss rate, etc.) each have different units and scales, so they cannot be directly compared or added together. Therefore, the score calculation unit of this embodiment first normalizes each analysis result to a range of 0 to 1 (or 0 to 100) and converts it into a dimensionless score. For example, for text edit distance, the ratio of edited parts to the original text length is used as the score, and for semantic similarity, the value obtained by subtracting the similarity score from 1 is used as the dissimilarity score. Next, these normalized individual scores are aggregated into meaningful categories. For example, intermediate scores such as "text modification score," "audio modification score," "sentiment manipulation score," and "context loss score" are calculated. The "text modification score" integrates the edit distance, semantic dissimilarity, etc., to indicate the extent to which the content of the speech itself has been changed. The "Audio Modification Score" integrates the results of feature difference detection, such as pitch, volume, and speech rate, as well as the detection of editing traces, to indicate the extent to which the audio itself has been modified. The "Emotion Manipulation Score" indicates the degree of impression manipulation based on the difference in emotion estimation results between the original information audio and the target audio (e.g., a change from calm to angry). The "Context Loss Score" indicates the degree of context distortion based on the aforementioned context loss rate and the evaluation of the importance of the deleted context. After calculating these category-specific scores, they are finally integrated to calculate a single "Overall Modification Score." When integrating, it is effective to weight each category according to its importance. For example, if a complete disregard for context or an emotionally manipulative audio modification is considered to represent a more serious form of information distortion than a simple text modification such as correcting a slip of the tongue, then the context loss score and emotion manipulation score should be given higher weights. While these weighting coefficients can be pre-set based on expert knowledge, it is preferable to optimize them using machine learning techniques.For example, a large number of modified content items that humans judged to be "malicious" and content items that were judged to be "not problematic" could be prepared as training data. The scores for each category could then be used as input to train a weighting coefficient that best reproduces those judgments (for example, using logistic regression or support vector machines). This would make it possible to calculate an objective and reliable overall modification score that is closer to human perception.

[0026] (Eighth embodiment) Next, as an eighth embodiment of the present invention, an account evaluation function for evaluating the reliability of accounts that post and spread modified content will be described in detail. The problem of information distortion depends not only on the nature of individual content, but also largely on the context of the sender, such as who is spreading it, with what intention, and how. The account evaluation unit of this embodiment aims to analyze such sender contexts and identify accounts that habitually engage in malicious information manipulation or accounts that may be involved in organized information manipulation campaigns. This function begins by accumulating the calculated modification score in a database linked to account information. Specifically, for each piece of content posted by a given account, a detailed score history such as the overall modification score, text modification score, audio modification score, sentiment manipulation score, and contextual gap score is recorded along with the posting date and time, engagement count (likes, retweets, comments), etc. Based on this historical data, the account's behavioral patterns are analyzed from multiple angles. For example, if a large number of pieces of content consistently have a high contextual gap score in a given account's posting history, it can be determined that the account is likely a habitual offender that intentionally cuts out context and distorts information. Furthermore, if a pattern is observed in which content with a high emotional manipulation score is rapidly posted and spread immediately after a particularly high-profile event (e.g., election, disaster, large-scale press conference), it suggests that this may be an activity intended to manipulate public opinion with a specific intent. In addition, the behavior of multiple accounts posting the same or similar modified content almost simultaneously, or retweeting each other, may suggest the existence of a bot network or influencer group (including activities similar to so-called "summary sites") that are colluding to manipulate information. This system can use graph analysis techniques to detect such collusive behavior. A social graph is constructed with accounts as nodes (vertices) and relationships such as retweets and mentions as edges, and by applying a community detection algorithm (e.g., Louvain's algorithm), closely linked groups of accounts are extracted as clusters.Then, by evaluating the average modification score and the synchronicity of posting patterns for the entire cluster, the presence or absence of organized information manipulation is determined. These analysis results are integrated, and indicators such as a "trust score" and "alert level" are assigned to each account. This score is an evaluation based on past behavior and is continuously updated. Accounts judged to have a significantly low trust score or an extremely high alert level are added to a monitoring list, and their subsequent activities are monitored intensively.

[0027] (Ninth embodiment) Next, as a ninth embodiment of the present invention, a visualization and UI (user interface) function for intuitively understanding the results of differential analysis and supporting rapid decision-making will be described in detail. The wide range of analysis data generated by this system is difficult to grasp in essence unless one is an expert, simply by listing the numbers. Therefore, the visualization and UI unit of this embodiment converts this complex information into a visual representation that can be understood at a glance and provides it as an interactive dashboard. At the heart of this dashboard is the "comparison view" which compares the original information content and the target content side by side. The original information content (for example, the full video of an official press conference and its full transcript) is displayed on the left side of the screen, and the target content to be verified (for example, a clipped video posted on social media and its transcript) is displayed on the right side. Users can play both videos in sync and visually confirm which parts have been cut and which parts have been used. On the transcript text, parts deleted from the target content are displayed in gray or with a strikethrough, and changed parts are highlighted (for example, in red). There is also a function to jump to the corresponding part of the video when the text is clicked. Furthermore, detailed analysis results for both audio tracks are graphically displayed at the bottom of the screen. For example, the audio waveforms are displayed side-by-side, making the difference in amplitude (volume) immediately obvious. Below that, a time-series graph of pitch (voice pitch) and a time-series graph of emotion (e.g., the trajectory of arousal and pleasure / displeasure) are overlaid, allowing for a visual comparison of when the tone of voice changed and how emotions shifted. Particularly important is the "difference highlighting" function. For example, pauses that exist in the original information but are removed in the target content are clearly indicated as red bars on the timeline, and sections where the speaking speed is unnaturally sped up are highlighted. This allows users, even without specialized knowledge, to intuitively understand the editing intent, such as, "This pause has been cut, making the speech sound rushed," or "This word is unnaturally loud and overemphasized."The overall modification score and scores for each category are displayed in formats such as meters and radar charts, and the risk level of the content is indicated by color coding (e.g., green: safe, yellow: caution, red: dangerous). All of these visualization elements are interactive, and a wealth of features are included to support in-depth analysis, such as highlighting corresponding videos and text when the user selects a part of the graph.

[0028] (Tenth embodiment) Next, as a tenth embodiment of the present invention, the applications and extended functions of this system will be described in detail. This system can be used not only to retrospectively verify content posted in the past, but also as a more proactive information risk management tool. One such application is the "real-time monitoring function." This function acquires the audio stream in real time while important press conferences, speeches, live broadcasts, etc., are in progress and feeds it into the analysis pipeline. Based on a model of "speech patterns that are easily taken out of context" learned from past data, it predicts from the ongoing speech that there is a high risk of it being taken out of context later and used to distort information. For example, in a statement such as "...there are criticisms that..., but overall...", the system evaluates in real time the possibility that only the part "...there are criticisms that..." will be taken out of context, or the possibility that only a part emphasizing a particular aspect will be misused when a complex issue is mentioned from multiple angles, based on the structure of the speech, the words used, and the speaker's prosodic characteristics. Speeches that are judged to be high risk are immediately notified as warnings (alerts) on the dashboards of public relations personnel and risk managers. This warning includes the reasons why the content was deemed high-risk (e.g., "strong assertive language," "words that could incite conflict," "emotional outbursts," etc.). This allows public relations personnel to take swift and effective action, such as proactively disseminating accurate information and contextual clarification before problematic clips spread, or preparing Q&A for anticipated criticisms. To build this predictive model, numerous past instances of clipped and problematic statements are used as training data, and machine learning (e.g., gradient boosting trees and neural networks) is used to learn patterns of features (linguistic features, acoustic features) associated with the likelihood of clipping. Furthermore, each function of this system (e.g., difference analysis, score calculation) can be provided as an API (Application Programming Interface) for easy use by external systems.For example, content management systems (CMS) of news organizations, social listening tools of companies, and verification platforms of fact-checking organizations can incorporate the information verification function of the present invention into their systems by calling this API. The API is provided in REST (Representational State Transfer) and gRPC (gRPC Remote Procedure Calls) formats and is published along with detailed documentation and sample code. This makes it possible to widely provide the technical value of the present invention to society and contribute to the health of a broader information ecosystem.

[0029] (Detailed flow of content acquisition process) The content acquisition method performs processing in multiple stages when acquiring target content to be verified from a designated platform. First, the system refers to a list of platforms registered as targets for monitoring (e.g., social media services, video sharing sites, news sites, etc.) and accesses each platform. Access can be done using the API (Application Programming Interface) provided by each platform, or by directly extracting information from web pages using web scraping technology. When using an API, the system sends a request to the platform's API endpoint using a previously obtained access token or authentication information. The request can include search conditions such as the type of content to be acquired (e.g., video, image, text), time range, keywords, and hashtags. In response to the request, the platform returns metadata for the relevant content (title, author, posting date and time, number of views, number of likes, etc.) and a link to the content itself. The system analyzes the received metadata and determines the verification priority. Factors such as the speed of content dissemination, the influence of the author, and past modification history are considered in determining priority. For content determined to be high priority, the download of the content itself begins immediately.

[0030] (Detailed flow of the audio extraction process) The audio extraction means performs the process of extracting audio data from the acquired target content. If the content is a video file, the system identifies the container format of the video file (e.g., MP4, AVI, MKV, WebM) and selects the appropriate demultiplexer (decontainerizer). The demultiplexer separates the video stream and audio stream from the video file. Since the audio stream may be encoded with various codecs (e.g., AAC, MP3, Opus, Vorbis, FLAC), the system obtains the codec information of the audio stream and uses the corresponding decoder to convert the compressed audio data into uncompressed PCM (Pulse Code Modulation) format. At this time, basic parameters of the audio data such as the sampling rate (e.g., 8kHz, 16kHz, 44.1kHz, 48kHz), bit depth (e.g., 8-bit, 16-bit, 24-bit), and number of channels (mono, stereo, multi-channel) are recorded. If a specific sampling rate or number of channels is required in subsequent speech recognition or speech feature analysis processes, the system will perform resampling or channel conversion. For example, when converting stereo audio to mono, the system will average the audio data from the left and right channels.

[0031] (Detailed flow of the transcription process) The transcription method uses an AI speech recognition model to generate text data representing the spoken content and corresponding time information from the extracted audio data. The speech recognition process consists of three main components: an acoustic model, a language model, and a pronunciation dictionary. The acoustic model estimates phonemes (the smallest units of sound in language) from the acoustic features of speech (e.g., Mel-frequency cepstrum coefficients, MFCCs). The language model is a model that learns the probability of word and sentence occurrences and is used to estimate the most likely word sequence from a sequence of phonemes. The pronunciation dictionary defines the phoneme sequence that makes up each word. In the latest end-to-end speech recognition models (e.g., Whisper, Wav2Vec 2.0, HuBERT), these components are integrated into a single neural network, allowing for direct text generation from audio waveforms. The system divides the audio data into segments of a fixed length (e.g., 30 seconds) and inputs each segment into the speech recognition model. The model outputs the recognized text and the time range (start and end times) in which the text was spoken for each segment. Time information is recorded in milliseconds or frames. The model also outputs a confidence score for the recognition result. Segments with low confidence scores can be handled specially in subsequent processing (for example, prompting manual verification).

[0032] (Detailed flow of speech feature analysis process) The speech feature analysis means extracts at least one speech feature from speech data, including pitch, volume, speech rate, and pauses. Methods such as autocorrelation, cepstrum analysis, and the YIN algorithm are used to extract pitch. These methods estimate the fundamental frequency (F0) by detecting the periodicity of the speech waveform. The fundamental frequency corresponds to the vibration frequency of the vocal cords and represents the pitch of the voice. The system divides the speech data into short time windows (e.g., 20 milliseconds) and calculates the fundamental frequency for each time window. This yields a trajectory of the voice's pitch (pitch curve) that changes over time. Methods for extracting volume (loudness) include measuring the amplitude of the speech waveform or using a loudness model that considers human auditory characteristics (e.g., A-weighting). The system calculates the speech energy or loudness for each time window and records the time change in volume. Methods for extracting speech rate include calculating the number of phonemes or syllables spoken per unit time. The system obtains phoneme sequences and syllable sequences from the transcription results and calculates the speech rate by dividing their number by the speech duration. Pauses (silent intervals) can be detected by searching for intervals where the speech energy falls below a certain threshold, or by using silence labels output by the speech recognition model. The system records the location (start and end times) and length of detected pauses.

[0033] (Structure and management of the correct information database) A legitimate information database is a database that stores legitimate information content originating from official sources. The database structure can adopt various formats, including relational databases, document-oriented databases, and graph databases. In a relational database, the table storing legitimate information content includes columns such as content ID, title, description, publication date and time, source, content type (video, audio, text), file path, and metadata. There are also multiple related tables, such as tables for storing transcription results, audio features, and audio fingerprints. These tables are interconnected using the content ID as the key. Database management includes operational tasks such as regular backups, index optimization, and archiving of old data. The system automatically detects, downloads, transcribes, extracts audio features, and registers new official content in the database as it is published. Registered content is assigned a unique identifier, and an index is created to allow efficient access during subsequent searches and comparisons.

[0034] (Detailed flow of differential analysis process) The difference analysis method compares the target content with one or more positive information content stored in the positive information database and analyzes the differences between the two. Difference analysis is performed at multiple levels. Firstly, in text-level difference analysis, the transcription result of the target content is compared with the transcription result of the positive information content. Indicators such as edit distance (Levenshtein distance), longest common subsequence (LCS), and N-gram similarity are used for the comparison. The system aligns the two texts and identifies matching parts, deleted parts, inserted parts, and replaced parts. Secondly, in semantic-level difference analysis, natural language processing techniques are used to compare the semantic content of the texts. For example, language models such as BERT, RoBERTa, and GPT are used to convert each text into a high-dimensional vector (embedding representation), and the cosine similarity and Euclidean distance between the vectors are calculated. In addition, textual entailment and contradiction detection models are used to determine whether the content of the target content logically matches or contradicts the content of the positive information content. Thirdly, difference analysis of speech levels compares speech features (pitch, volume, speaking speed, pauses, etc.). The system extracts speech features from corresponding time intervals and calculates the difference. For example, if a part of the speech that was delivered in a calm tone in the positive information content changes to an angry or excited tone in the target content, that change will be detected.

[0035] (Detailed flow of the score calculation process) The scoring means calculates a modification score, which indicates the degree of modification of the target content, based on the analysis results of the difference analysis means. The modification score is a comprehensive index calculated by integrating multiple elements. Key elements include the text difference score, semantic difference score, speech feature difference score, and temporal consistency score. The text difference score is calculated based on the degree of agreement of the transcription results. For example, the smaller the edit distance (i.e., the more similar the text is), the lower the score (less modification). The semantic difference score is calculated based on the semantic similarity and logical consistency of the text. For example, the score is low if an implication relationship exists and high if a contradiction is detected. The speech feature difference score is calculated based on the magnitude of the difference in speech features such as pitch, volume, speech rate, and pauses. For example, the score is high if there is a large change in pitch or if the speech rate has been significantly changed. The temporal consistency score is calculated based on the consistency of which parts of the original information content each part of the target content corresponds to. For example, the score is high if part of the original information content has been deleted or its order has been changed. These individual scores are combined using weighted averaging and machine learning models to calculate the final modification score.

[0036] (Data structure details: Content metadata) The content metadata handled by the system is managed in a structured data format. Content metadata includes the following fields: Content ID (unique identifier), Title (title of the content), Description (description of the content), Author ID (identifier of the author), Author Name (display name of the author), Post Date and Time (date and time the content was posted), Platform (name of the platform on which the content was posted), URL (URL to access the content), Content Type (video, audio, image, text, etc.), File Format (MP4, MP3, JPG, etc.), File Size (in bytes), Playback Time (in seconds, for video and audio), Playback Count, Likes, Comments, Shares, Hashtags (list of hashtags attached to the content), Language (language of the content), Region (region the content targets), Category (category of the content), Verification Status (unverified, under verification, verified), Modification Score (calculated modification score), Verification Date and Time (date and time verification was performed), Verifier ID (identifier of the system or user that performed the verification). These fields are serialized in formats such as JSON or XML and stored in a database or file system.

[0037] (Data structure details: Transcription results) The transcription results are managed as structured data containing text representing the spoken content and time information of when that text was spoken. The data structure of the transcription results includes the following fields: Content ID (identifier of the corresponding content), Segment ID (identifier of the segment of the transcription result), Start Time (start time of the segment in milliseconds), End Time (end time of the segment in milliseconds), Text (recognized text), Confidence (confidence score of the recognition result, ranging from 0 to 1), Speaker ID (identifier of the speaker, if speaker separation was performed), Language (recognized language), Phoneme Sequence (phoneme sequence corresponding to the text, optional), Word Boundaries (list of start and end times for each word, optional). These fields hold multiple segments in the form of arrays or lists, and together they represent the complete transcription result of a single piece of content. The data structure is serialized in JSON or XML format and stored in a database. In addition, a full-text search index is created for the text fields to streamline searching and comparison.

[0038] (Data structure details: Speech features) Speech features are managed as numerical data representing various features extracted from speech data. The data structure of speech features includes the following fields: Content ID (identifier of the corresponding content), Frame ID (identifier of the time frame from which the features were extracted), Time (center time of the frame, in milliseconds), Pitch (fundamental frequency, in Hz), Volume (loudness, in dB), Speech rate (number of phonemes or syllables per unit time), Pause flag (a Boolean value indicating whether the frame is a silent interval), MFCC (vector of Mel-frequency cepstrum coefficients), Spectral centroid (frequency of the centroid of the speech spectrum), Zero crossing rate (frequency at which the speech waveform crosses zero), Energy (energy of the speech), Emotion label (estimated emotion, e.g., neutral, joy, anger, sadness), and Emotion score (confidence score for each emotion). These features are stored as time-series data in array format. For example, if features are extracted every 10 milliseconds for 1 second of speech, 100 frames will be generated. Data structures are sometimes stored in binary format (e.g., NumPy's npy format, HDF5 format) to enable efficient storage and fast access.

[0039] (Security mechanism: data encryption) This system implements encryption at multiple levels to protect highly confidential data. First, for encryption of the communication path, encrypted communication using the TLS (Transport Layer Security) protocol is used for communication between client devices and servers, and for communication between servers. TLS is a standard protocol for preventing data eavesdropping, tampering, and impersonation, and the currently widely adopted versions are TLS 1.2 and TLS 1.3. Next, for encryption of stored data, data stored in the database and files stored in the file system are encrypted using symmetric key encryption algorithms such as AES (Advanced Encryption Standard). AES supports key lengths of 128 bits, 192 bits, and 256 bits, achieving both high security and processing speed. Dedicated key management systems such as KMS (Key Management Service) and HSM (Hardware Security Module) are used for managing encryption keys. This minimizes the risk of unauthorized access to encryption keys. In addition, field-level encryption is applied to fields containing personal or confidential information, preventing even database administrators from viewing their contents.

[0040] (Security mechanism: Access control) The system implements either role-based access control (RBAC) or attribute-based access control (ABAC) to ensure that users and system components can only access resources for which they have appropriate permissions. In RBAC, users are assigned roles, and each role is associated with specific permissions (e.g., read, write, delete). For example, the "administrator" role is granted full access to all resources, while the "viewer" role is granted read-only permission. In ABAC, access permissions are dynamically determined by combining user attributes (e.g., department, job title, clearance level), resource attributes (e.g., confidentiality level, owner), and environment attributes (e.g., access time, source IP address). The system performs authentication (verification of the user's identity) and authorization (verification of the user's authority to perform the requested operation) for each request. Authentication methods include passwords, multi-factor authentication (MFA), biometric authentication, and certificate-based authentication. Authorization uses access control lists (ACLs) or policy engines.

[0041] (Security mechanism: audit log) The system records all important operations and events as audit logs. Audit logs include the following information: timestamp (date and time the operation was performed), user ID (identifier of the user or system component that performed the operation), operation type (e.g., login, logout, data read, data write, data delete, configuration change), target resource (identifier of the resource targeted by the operation), operation result (success or failure), error message (error details if it failed), source IP address, and user agent (client software information). Audit logs are stored in append-only storage to prevent tampering and are backed up regularly. Digital signatures and blockchain technology can also be used to ensure log integrity. Audit logs are used for investigating security incidents, verifying compliance, and analyzing system operation. The system can have features to periodically analyze audit logs and detect unusual access patterns or unauthorized operations.

[0042] (Error handling: Network error) The system performs appropriate error handling for various errors that occur during network communication. Network errors include connection timeouts, read timeouts, connection refusals, hostname resolution failures, and SSL certificate errors. When these errors occur, the system first automatically performs retries. The number and interval of retries are determined based on an exponential backoff algorithm. For example, the first retry is after 1 second, the next after 2 seconds, the next after 4 seconds, and so on, with the waiting time increasing exponentially. This allows time to be waited for recovery from temporary network failures or server overload conditions. If the retries fail a specified number of times, the system marks the task as a failure and records it in the error log. It can also send a notification to the administrator prompting manual action. Furthermore, the system can mitigate the impact of a single network point of failure by utilizing alternative communication paths and mirror servers.

[0043] (Error handling: Data processing error) The system performs appropriate error handling for various errors that occur during data processing. Data processing errors include invalid file formats, data corruption, decoding errors, insufficient memory, and insufficient disk space. The system verifies the validity of input data at each processing step. For example, before processing a video file, it checks the file's magic number (a sequence of bytes that identifies the file format) to ensure it matches the expected format. If data corruption is detected, the system attempts to repair the data if possible. If repair is not possible, it skips the data and proceeds to processing the next data. It also records information about the data that caused the error (content ID, error details, etc.) in the error log so that it can be reviewed and reprocessed later. For insufficient memory or insufficient disk space errors, the system frees up resources by clearing temporary files and caches. If the problem persists, the system pauses processing and sends a notification to the administrator.

[0044] (Error handling: Errors in the AI ​​model) The system performs appropriate error handling for various errors that occur during AI model inference. AI model errors include errors in loading the model file, errors in the input data format, inference timeouts, insufficient GPU memory, and errors where the model output is in a format different from what is expected. When loading the model, the system verifies the integrity of the model file. For example, it calculates a checksum for the file and checks if it matches a previously recorded checksum. For errors in the input data format, the system converts the format in the data preprocessing step. For example, if the model expects audio data at a specific sampling rate, it resamples the input audio to that sampling rate. For inference timeouts, the system sets an appropriate timeout period and, if a timeout occurs, interrupts the inference and records it as an error. For insufficient GPU memory, the system takes measures such as reducing the batch size, lowering the model's precision (e.g., converting from FP32 to FP16), or performing inference on the CPU.

[0045] (Optimization method: caching) The system implements caching at various levels to improve processing speed. First, as API response caching, metadata and files of content obtained from external platforms are cached locally for a certain period. This reduces duplicate requests for the same content and lowers the risk of exceeding the usage limits (rate limits) of external APIs. Next, as database query caching, the results of frequently executed queries are cached in memory. For example, in-memory data stores such as Redis or MemCached are used to enable fast retrieval of query results. Furthermore, as AI model inference result caching, inference results for the same input are cached. For example, by reusing transcription results or audio features for the same audio file, duplicate inference processing can be avoided. The cache expiration time (TTL: Time To Live) is set appropriately according to the nature of the data. For example, data that is frequently updated is set to a short TTL, while data that rarely changes is set to a long TTL.

[0046] (Optimization methods: parallel processing and batch processing) The system utilizes parallel and batch processing to efficiently process large amounts of data. Parallel processing reduces overall processing time by executing multiple tasks simultaneously. For example, the transcription of multiple content items can be distributed across multiple worker processes or threads for parallel execution. Technologies such as multiprocessing, multithreading, and asynchronous I / O are used to implement parallel processing. Furthermore, large-scale parallel processing across multiple machines can be achieved using distributed processing frameworks (e.g., Apache Spark, Dask). Batch processing processes multiple data items together at once, rather than processing individual data items one by one. For example, in AI model inference, inputting multiple audio segments as a batch into the model maximizes the parallel computing capabilities of the GPU, improving throughput. The batch size is adjusted according to the available memory capacity and the characteristics of the model.

[0047] (Optimization method: Model optimization) The system applies model lightweighting techniques to improve the inference speed of AI models and reduce resource consumption. Model lightweighting techniques include quantization, pruning, and knowledge distillation. Quantization is a technique that reduces the size and computational complexity of a model by lowering the precision of its parameters (weights). For example, weights represented as 32-bit floating-point numbers (FP32) are converted to 16-bit floating-point numbers (FP16) or 8-bit integers (INT8). This reduces the model size by half or a quarter and also improves inference speed. Pruning is a technique that makes a model sparse by removing less important parameters or connections. This reduces computational complexity and memory usage. Knowledge distillation is a technique that transfers the knowledge of a large teacher model to a smaller student model. The student model is learned to mimic the output of the teacher model, so it has performance close to the teacher model while significantly reducing its size and computational complexity.

[0048] (Optimization techniques: Optimization of indexes and queries) The system implements appropriate index creation and query optimization to improve database access speed. An index is a data structure created for a database table that allows for fast retrieval of rows based on the values ​​of specific columns. For example, indexes are created for columns that are frequently used as search criteria, such as content ID, publication date and time, and modification score. Types of indexes include B-tree indexes, hash indexes, full-text search indexes, and spatial indexes. Query optimization analyzes the execution plan of database queries, reduces unnecessary joins and subqueries, and rewrites queries to use appropriate indexes. It also keeps database statistics up-to-date so that the query optimizer can select the best execution plan. Furthermore, query performance for large datasets can be improved by implementing database partitioning and sharding.

[0049] (User interface details: Dashboard) The system provides a dashboard that allows users to see the system status and verification results at a glance. The dashboard displays information such as: the total number of verified content, the number of content currently under verification, the number of content with high modification scores, the number of content verified in the past 24 hours, a graph showing the trend of the average modification score, content distribution by platform, content distribution by category, a list of recently verified content, and a list of alerts and notifications. The dashboard is implemented as an interactive interface that runs on a web browser, and users can drill down to more detailed information by clicking on graphs and lists. The data on the dashboard is updated in real time or periodically, so that it always displays the latest information. In addition, the dashboard is customized according to the user's role and permissions, so that each user only sees the information they need.

[0050] (User interface details: Detailed display of verification results) If a user wants to see the verification results for specific content in detail, the system provides a detailed verification results screen. This screen displays the following information: basic content information (title, author, posting date and time, platform, etc.), modification score and its breakdown (text difference score, semantic difference score, audio feature difference score, etc.), a comparison view of the target content and the original content (displayed side-by-side), a comparison of transcription results (differences are highlighted), a comparison of audio waveforms (differences can be visually confirmed), a comparison graph of audio features (time changes in pitch, volume, speech rate, etc.), sentiment analysis results (differences in sentiment between the target and original content), details of deleted and added parts, verification history (records of past verifications), and user comments and feedback. From this screen, users can export the verification results (PDF, CSV, etc.), share them with other users, and submit feedback.

[0051] (User interface details: Search and filtering) The system provides powerful search and filtering functions to efficiently find desired content from a large amount of content. Users can search and filter content using the following criteria: keyword search (keywords in title, description, and transcript results), search by author, specifying a range for posting date and time, filtering by platform, filtering by content type, specifying a range for modification score, filtering by verification status, filtering by language, filtering by category, and filtering by hashtag. Search results can be sorted by relevance, posting date and time, modification score, etc. The search results list also displays summary information for each piece of content, such as a thumbnail image, title, modification score, and posting date and time, and users can click on content of interest to view details. The system saves a history of search queries, allowing users to re-execute searches they have performed in the past.

[0052] (System scalability: Plugin architecture) This system can adopt a plugin architecture to facilitate future feature enhancements and the integration of new technologies. In a plugin architecture, the system's core functions and extensions are clearly separated, with extensions implemented as independent modules called plugins. Each plugin is implemented according to a system-defined interface (API), and the system's core can dynamically load plugins and invoke their functions. For example, plugins can be added to support new speech recognition models, retrieve content from new platforms, or implement new differential analysis algorithms. Plugins can be added, removed, and updated without recompiling or restarting the system. Furthermore, plugins can be developed by third-party developers and distributed through a plugin marketplace.

[0053] (System scalability: API provision) This system provides RESTful APIs and GraphQL APIs to enable external applications and services to utilize the system's functions. Through the APIs, external systems can perform operations such as: sending content validation requests, retrieving validation results, searching for content, registering content in the official information database, and retrieving statistical information. The APIs are equipped with authentication and authorization mechanisms to ensure that only authorized clients can access them. Authentication methods include API keys, OAuth 2.0, and JWT (JSON Web Token). API documentation is provided in OpenAPI (Swagger) format, allowing developers to easily understand the API specifications and develop client applications. The APIs also have rate limits (limits on the number of requests per unit of time) to prevent excessive load on the system.

[0054] (System scalability: Multi-tenant compatible) This system can implement multi-tenancy support, allowing multiple organizations and user groups to use the system independently. In a multi-tenant architecture, a single system instance serves multiple tenants (organizations or user groups). Data for each tenant is logically isolated, and users in one tenant cannot access data in other tenants. Data isolation can be achieved through methods such as database-level isolation (each tenant has its own independent database), schema-level isolation (each tenant has its own independent schema), and row-level isolation (all tenants share the same table, but each row is assigned a tenant ID). The system can provide each tenant with independent settings, customizations, and branding. Furthermore, usage data for each tenant (storage usage, API calls, etc.) can be tracked individually and used for billing and resource limits.

[0055] (System operation: monitoring and alerting) To maintain stable system operation, the system includes comprehensive monitoring and alerting capabilities. Monitoring targets include server CPU usage, memory usage, disk usage, network bandwidth usage, service status, database response time, API response time, error rate, and queue length. The system periodically collects these metrics and stores them in a time-series database (e.g., Prometheus, InfluxDB). A monitoring dashboard (e.g., Grafana) visualizes the metrics in graphs and charts, allowing for real-time monitoring of the system's status. The alerting function sends notifications to administrators when metrics exceed pre-configured thresholds or when specific events occur (e.g., service downtime, sudden spikes in error rates). Notifications are sent via email, SMS, Slack, PagerDuty, and other methods.

[0056] (System operation: Log management) The system includes a log management system that centrally collects, stores, and analyzes logs generated by all components. These logs include application logs (records of application behavior and processing), access logs (records of access to APIs and web servers), error logs (records of errors and exceptions), and audit logs (records of security-related operations). Logs are output in a structured format (e.g., JSON), and each log entry includes a timestamp, log level (DEBUG, INFO, WARN, ERROR, FATAL), source (the component that generated the log), message, and context information (e.g., request ID, user ID). Logs are collected by log collection agents (e.g., Fluentd, Logstash) and sent to central log storage (e.g., Elasticsearch, Splunk). Log analysis tools (e.g., Kibana) can be used to search, filter, and visualize logs to quickly identify system problems.

[0057] (System operation: backup and recovery) The system incorporates regular backup and recovery mechanisms to prevent data loss and enable rapid system recovery in the event of a disaster. Backup targets include databases, file storage, configuration files, and AI model files. Backups are performed using full backups (backing up all data), differential backups (backing up only data changed since the last full backup), and incremental backups (backing up only data changed since the last backup). Backups are performed automatically on a regular basis, and backup data is stored in a location physically separate from the system (e.g., another data center, cloud storage). Recovery procedures are documented in advance and regularly tested. The system defines Recovery Point Objectives (RPO) and Recovery Time Objectives (RTO) and designs backup and recovery strategies to achieve these objectives.

[0058] (Details of industrial applicability) The system of the present invention has broad applicability in various industrial fields. Firstly, in the field of news organizations and journalism, the system can be used as a tool to quickly verify the veracity of information disseminated on social media. Reporters and editors can use the system to check whether video and audio sources have been altered, and prevent the spread of misinformation. Secondly, in government and public institutions, the system can be used as a tool to monitor whether the content of official announcements and press conferences is accurately conveyed. In particular, it can prevent the statements of candidates and politicians from being distorted through selective quoting or alteration during election periods or important policy announcements. Thirdly, in corporate public relations and brand management departments, the system can be used as a tool to monitor whether product announcements and statements by management are inaccurately quoted or maliciously altered. Fourthly, in educational institutions, the system can be used as teaching material for media literacy education. Through the system, students can learn techniques for discerning the veracity of information and how the media processes information. Fifth, this system can be used by law enforcement and judicial bodies as a tool to verify the authenticity of audio and video evidence submitted.

[0059] (Modifications and other embodiments) The present invention is not limited to the embodiments described above, and various modifications are possible within the scope of its technical concept. For example, although the above embodiments mainly describe a configuration in which all processing is performed on the server side, it is also possible to adopt an edge computing configuration in which part of the processing is performed on the client side, for example, on the user's smartphone or PC. For example, some functions of this system can be provided as a plugin embedded in an SNS application, and real-time extraction of audio features and simple difference analysis can be performed when a video is played. Then, only when there is a strong suspicion of modification, a detailed analysis is requested from the server. This configuration can reduce the computational load on the server, reduce the amount of communication, and also contribute to protecting user privacy. Furthermore, regarding AI model training, it is also possible to adopt a federated learning approach in which the model is trained in each local environment (edge ​​device) without collecting data held by each organization or individual on a central server, and only the training results (e.g., the amount of parameter updates) are aggregated on the server to update the global model. This makes it possible to build a high-performance model trained from diverse data without disclosing highly private or confidential data to the outside. For example, each news organization can train models locally using its own archived data and share those insights, thereby improving the accuracy of information verification technology across the entire industry. As a further modification, to enhance the reliability of the present invention, blockchain technology could be integrated. For example, when an official information provider publishes content as "authentic information," the hash value (fingerprint) of that content, a timestamp, and provider information are recorded as a transaction on the blockchain. Due to the nature of its distributed ledger, it is extremely difficult to tamper with data once it is recorded on the blockchain. Therefore, this record functions as a highly reliable anchor that proves "this content is authentic, published at this time by this provider." Later, when this system analyzes the content to be verified, it checks whether the content matches any of the authentic information recorded on the blockchain. This guarantees the reliability of the authentic information itself, which serves as the basis for verification, and dramatically improves the reliability of the entire system. This mechanism can be provided as an "official proof" function, a means for information providers themselves to proactively prove the authenticity of their own information. Furthermore, since this record on the blockchain is verifiable by anyone, independent verification by platform operators and third-party organizations becomes easier.

[0060] Another variation involves more actively utilizing audio fingerprinting technology. Audio fingerprinting is a technique that generates a short, robust identifier (fingerprint) unique to an audio file from audio data. In the above embodiment, an example of its use in mapping target content to legitimate information content was described, but this can be further developed. For example, fingerprints could be extracted in advance from all official audio content stored in the legitimate information database, and a dedicated fingerprint database capable of high-speed searching could be constructed. Then, fingerprints could be extracted in real time from a large number of unspecified video and audio content circulating on social media and compared with this database. This makes it possible to instantly identify which official press conference and which part a particular clipped video was quoted from. This technology is widely used in the field of copyright management, but by applying it to the present invention, the identification of information sources can be dramatically sped up and made more accurate. By immediately starting the calculation of context loss rate and audio modification score based on the identified information sources, the throughput of the entire verification process can be greatly improved.

[0061] Modifications to enhance the transparency and explainability of the AI ​​model used in this invention are also conceivable. AI, especially deep learning models, are sometimes criticized as "black boxes" due to the complexity of their internal structure. Even if this system determines that the "modification score is high," if the reasoning behind that conclusion is unclear, users may find it difficult to trust and accept the results. Therefore, we introduce Explainable AI (XAI) technology. For example, we visualize which parts of the speech the AI ​​model focused on to determine that "emotions have changed." Methods such as LIME (Local Interpretable Model-agnostic Explanations) and SHAP (SHapley Additive exPlanations) can be used for this. These methods calculate the extent to which each input feature contributed to the model's prediction. By displaying this contribution as a spectrogram or heatmap in text, users can understand the specific reasoning behind the judgment, such as "the AI ​​emphasized this part where the sudden rise in pitch and the utterance of a specific word overlap as evidence of 'anger'." Providing this kind of explainability improves the reliability of the conclusions reached by the system, enabling users to make more informed decisions. (Industrial applicability)

[0062] This invention has broad applicability in various industrial fields where the reliability of information is extremely important. First, news organizations and media companies can use it as a fact-checking tool for daily news articles and broadcast content. In particular, by quickly verifying whether information is distorted when quoting or referencing reports from other companies or information on social media, it can reduce the risk of misinformation and maintain the quality and reliability of reporting. Furthermore, in election coverage, it is extremely effective in preventing candidates' statements from being taken out of context and reported unfairly, and in providing voters with fair information to make informed decisions. Second, the public relations and investor relations departments of companies and government agencies can use it as a risk management tool for information dissemination about their own organizations. It can detect early on situations where press conferences or official statements are maliciously edited and disseminated in a way that damages corporate value and social credibility, and support crisis communication such as prompt corrections and the issuance of official statements. It can also monitor misinformation and negative campaigns about a company's products and services, contributing to the protection of brand image. Furthermore, in the financial industry, it is expected to be used as an aid in verifying the veracity of market rumors and insider information, and in platform operators, as a monitoring tool to maintain the integrity of content distributed on their services. In the field of education, it is also highly valuable as a teaching material for students to learn media literacy, as it concretely demonstrates how information can be distorted. Thus, this invention contributes to a wide range of industries as a foundational technology that solves fundamental issues related to information distribution in modern society. (Explanation of symbols)

[0063] (Although this section is not strictly necessary as drawings are omitted, examples of components are provided in preparation for future addition of drawings.) 100... Information Verification System 110... Computer System 120... Server 130…Network 140…Database 200...Content Acquisition Unit 300...Audio extraction unit 400...Transcription Department 500...Speech Feature Analysis Unit 600… Correct Information Database 700...Differential analysis section 800... Score calculation unit 900... Account Evaluation Department 1000…Visualization / UI department 1100…API provision department

[0064] (11th embodiment) Next, as an eleventh embodiment of the present invention, the detailed configuration of the AI ​​speech recognition model in this system and a hybrid recognition method that utilizes its diversity will be described. The accuracy of speech recognition depends on many factors, such as the architecture of the model used, the quality and quantity of the training data, the target acoustic environment, and the characteristics of the speaker. It is difficult for a single model to achieve optimal performance in all situations. Therefore, in this embodiment, higher recognition accuracy and robustness are achieved by using multiple AI speech recognition models with different characteristics in parallel and integrating their outputs. Specifically, this system holds a diverse group of models, including transformer-based end-to-end models pre-trained on large-scale multilingual data (e.g., OpenAI Whisper, Google USM (Universal Speech Model), Meta SeamlessM4T), models fine-tuned for specific languages ​​or domains (e.g., a model specialized for Japanese press conference audio, a model corresponding to specialized terminology in the medical field), and conventional hybrid HMM-DNN models (e.g., a model built with the Kaldi framework). These models are run in parallel on the input audio, and transcription results and confidence scores are obtained from each model. The following approaches are used to integrate these multiple results: First, in Voting-based Fusion, a majority vote is taken on the word or phrase-level results output by each model. The result that receives the most agreement from the models is adopted as the final recognition result. Second, in Confidence-weighted Fusion, a weighted average or weighted vote is performed using the confidence scores assigned to the output of each model as weights. The results of the models with higher confidence are reflected more strongly. Third, in integration using the ROVER (Recognizer Output Voting Error Reduction) algorithm, multiple recognition results are aligned on a time axis, and voting is performed at the word level to generate the most plausible transcription result.Furthermore, an Adaptive Model Selection (AMC) method is also effective, which involves pre-analyzing the characteristics of the input audio (e.g., language, noise level, speaker's gender and age) and dynamically selecting the model best suited to those characteristics. For example, if the language of the audio is automatically identified and determined to be Japanese, a Japanese-specific model is given priority. Alternatively, if a high noise level is detected, a model with superior noise tolerance is selected. This hybrid approach enables the system to consistently provide highly accurate transcriptions for diverse acoustic conditions and speakers.

[0065] (12th embodiment) Next, as a twelfth embodiment of the present invention, a detailed feature extraction method in the speech feature analysis unit and the meaning of those features will be described. Speech signals are complex waveforms that fluctuate over time, and appropriate signal processing and feature design are essential to extract useful information from them. In this embodiment, the speech signal is divided into short frames (for example, 20 to 50 milliseconds), a window function (for example, a Hamming window) is applied to each frame, and then a Fast Fourier Transform (FFT) is performed to convert the time-domain waveform into a frequency-domain spectrum. Various acoustic features are calculated from this spectrum. First, to extract the fundamental frequency (F0), a highly accurate pitch estimation algorithm such as the autocorrelation method, cepstrum method, or the YIN algorithm or PYIN algorithm is used. F0 reflects the vibration frequency of the vocal cords and is directly related to the perception of pitch. By calculating the mean, standard deviation, maximum, minimum, and range of F0, the trend and magnitude of fluctuation in pitch of the entire speech are quantified. Next, to capture the characteristics of the spectral envelope, Mel-frequency cepstrum coefficients (MFCCs) are extracted. MFCCs are coefficients obtained by analyzing the spectrum using a Mel-scale filter bank that mimics human auditory characteristics and applying a discrete cosine transform (DCT) to its logarithmic power, effectively representing the timbre of speech. Typically, 12 to 13 low-order MFCCs, along with delta (Δ) coefficients and delta-delta (ΔΔ) coefficients representing their temporal changes, are used to create a feature vector with a total of 36 to 39 dimensions. Furthermore, the energy (power) of the speech is calculated as the sum of the squares of the amplitudes of each frame, reflecting the loudness of the voice. By determining the mean, standard deviation, maximum value, and dynamic range of the energy, the volume characteristics of the speech are understood. In addition, the zero-crossing rate indicates the frequency with which the speech waveform crosses the positive and negative boundary, and is useful for distinguishing between voiced and voiceless sounds and for detecting noise. The spectral centroid represents the position of the center of gravity of the spectrum and is related to the perception of brightness and sharpness of sound. Spectral roll-off indicates the frequency at which a certain percentage (e.g., 85%) of the spectral energy is concentrated and is used to assess the presence or absence of high-frequency components.These features capture the acoustic properties of speech from multiple perspectives, each providing information on different aspects.

[0066] In this embodiment, even more advanced prosodic feature extraction is performed. Prosody refers to features related to the temporal and dynamic structure of speech, including intonation, accent, rhythm, speech rate, and pauses. Intonation can be understood as a temporal change pattern of F0. If F0 rises at the end of a sentence, it is a question; if it falls, it is a declarative sentence, reflecting the intent of the utterance. In this system, prosodic models such as the Fujisaki model and the Tilt model are applied to mathematically model the contour of F0, and their parameters are extracted as features. Accent and prominence (emphasis) are achieved when F0, energy, and duration change more significantly in specific syllables or words than in their surroundings. In this system, the F0, energy, and duration of each syllable are calculated, and local peaks and fluctuations are detected to identify emphasized areas. Rhythm appears as a pattern of syllable and phoneme durations. The rhythmic characteristics of speech are quantified by calculating the degree of isochronism and the variability of duration (standard deviation, coefficient of variation). Speech rate is defined as the number of syllables, phonemes, or words spoken per unit time. This system calculates precise speech rate by combining transcription results with temporal information of the speech. Temporal fluctuations in speech rate (acceleration and deceleration) are also important information, and their rate of change is extracted as a feature. Pauses (silent intervals, pauses) are detected as intervals where speech power falls below a predetermined threshold. The location, length, and frequency of pauses reflect the structure of the speech and the speaker's psychological state. For example, long pauses are often placed before important statements, and tension or hesitation may manifest as the insertion of unnatural pauses. This system calculates pause statistics (average length, standard deviation, percentage of total pause time) and analyzes the relationship between each pause and the surrounding context. These prosodic features are extremely important for understanding the linguistic and emotional meaning of speech, and in the difference analysis of this invention, they are key to detecting edits that give a different impression even if the text content is the same.

[0067] (13th embodiment) Next, as a thirteenth embodiment of the present invention, the details of the emotion recognition function that estimates the speaker's emotional state from speech will be described. Human emotions are reflected in the acoustic and prosodic features of speech. For example, anger is often associated with a high F0, high energy, fast speech rate, and short pauses, while sadness is often associated with a low F0, low energy, slow speech rate, and long pauses. The emotion recognition unit of this embodiment takes the various features extracted by the aforementioned speech feature analysis unit as input and estimates the speaker's emotional state using an emotion recognition model based on machine learning or deep learning. There are two main approaches to expressing emotions: classification by discrete emotion categories and representation by continuous emotion dimensions. In the discrete approach, emotions are classified into one of six or seven basic emotions (joy, sadness, anger, fear, surprise, disgust, neutral) based on Ekman's basic emotion theory. In this system, an emotion classifier is trained using an emotion speech database (e.g., IEMOCAP, RAVDESS, EmoDB, JTES, etc.) that has these emotion categories as labels. For classifiers, conventional machine learning methods such as support vector machines (SVM), random forests, and gradient boosting trees (e.g., XGBoost, LightGBM) can be used, as well as deep learning models such as convolutional neural networks (CNNs), recurrent neural networks (RNNs, LSTMs, GRUs), and transformers. In particular, approaches that treat audio spectrograms as images and apply image recognition techniques with CNNs, and approaches that process audio features as time-series data with LSTMs are effective. On the other hand, in the continuous approach, emotions are represented as coordinates on a continuous space in two dimensions: arousal (excited or calm) and valence (positive or negative), or even three dimensions by adding dominance (dominant or subordinate). In this approach, the emotion recognition model is constructed as a regression model and predicts the values ​​of each dimension. The advantage of continuous representation is that it can represent subtle changes and gradations of emotion that cannot be captured by discrete categories. This system utilizes both discrete and continuous emotion recognition models in parallel, allowing for their appropriate use depending on the situation.

[0068] To improve the accuracy of emotion recognition, this embodiment employs multimodal emotion recognition that uses not only speech features but also text information (transcription results). The content of the utterance itself (for example, the use of words expressing emotions such as "happy" and "sad," negative expressions, emphasis expressions, etc.) is an important clue in estimating emotion. In this system, text-based emotion analysis is performed on the transcribed text using pre-trained language models such as BERT, RoBERTa, and GPT to estimate emotion from the text. The results of speech-based emotion recognition and text-based emotion analysis are then integrated using either late fusion or early fusion methods. In late fusion, emotion is estimated independently from speech and text, and the results are integrated by weighted averaging or voting. In early fusion, an integrated feature vector is created by combining speech features and text features (for example, BERT embedding vectors), and this is used as input to train a single multimodal emotion recognition model. This multimodal approach utilizes information from both speech and text, achieving more robust and accurate emotion recognition. The estimated emotional information is recorded along with time information (timestamp) and stored in the positive information database. Then, in the difference analysis unit, the emotional time series of the target content and the emotional time series of the positive information content are compared to detect any deformation or manipulation of emotions. For example, if a part that was expressed with neutral or positive emotions in the positive information changes to negative emotions such as anger or anxiety in the target content, this strongly suggests that the impression may have been manipulated through audio editing or processing.

[0069] (14th embodiment) Next, as a 14th embodiment of the present invention, an automated workflow for building and updating a positive information database will be described in detail. The quality and comprehensiveness of the positive information database are directly linked to the verification accuracy of this system. Therefore, it is extremely important to build an automated pipeline that collects, analyzes, and stores content in the database in a timely and continuous manner from reliable sources. In this embodiment, first, a source registry is built to manage a list of reliable official sources. This registry registers information such as the URL of each source, type (government agency, company, news organization, etc.), reliability level, monitoring frequency, and acquisition method (API, RSS feed, web scraping). The source registry can be manually edited by the administrator, but it also has the function of automatically discovering and suggesting new reliable sources. For example, it presents sites linked from existing reliable sources or sites evaluated using domain reliability evaluation services (e.g., domain age, HTTPS support, official domain (.go.jp, .gov), etc.) as candidates. Next, a source monitoring agent periodically (e.g., every few minutes, every few hours, or daily) visits each source registered in the registry to check for new content. When new content is detected, its metadata (title, publication date and time, URL, etc.) is retrieved and added to the content retrieval queue. The content retrieval worker retrieves the task from the queue and downloads the video and audio files. The downloaded files are stored in temporary storage and passed on to the next analysis step. The analysis pipeline automatically performs the following series of processes: First, it extracts the audio track from the video file and converts it to a standard audio format (e.g., WAV, FLAC). Second, it inputs the audio file into the transcription unit to generate text data and timestamps. Third, it inputs the audio file into the audio feature analysis unit to extract various acoustic, prosodic, and sentiment features. Fourth, it applies a natural language processing pipeline (e.g., morphological analysis, named entity recognition, syntactic analysis, semantic analysis) to the generated text data to generate structured text information.Fifth, text data is input into a pre-trained language model (e.g., BERT) to generate sentence or paragraph-level semantic vectors (embeddings). All of these analysis results are associated with unique content IDs and stored in appropriate databases (relational databases, object storage, time-series databases, vector databases).

[0070] Furthermore, this embodiment also includes functions for database quality control and maintenance. First, the deduplication function prevents the same content from being collected redundantly from different sources. This is achieved by methods such as calculating the hash value (e.g., MD5, SHA-256) of video and audio files and comparing it with the existing database, or by using audio fingerprinting technology to determine the similarity of audio content. Next, the quality verification function automatically checks the quality of transcription results and feature extraction results. For example, if the confidence score of the transcription is generally low, or if the extraction of audio features fails (e.g., F0 cannot be detected), a warning is issued and manual verification is prompted. In addition, the database version control function handles cases where official information is later corrected or updated. For example, if the minutes of a press conference are revised later, the new version of the data is added and its relationship to the old version is recorded. This allows tracking of which version of the official information was used as the basis for verification. Furthermore, the database backup and restore function ensures resilience against system failures and data corruption. Regular automatic backups and replication to different geographical locations guarantee data persistence and availability. These automated workflows and quality control functions ensure that the positive information database is always up-to-date and of high quality, serving as the foundation for the reliability of this system.

[0071] (15th embodiment) Next, as a 15th embodiment of the present invention, a detailed method for text difference analysis in the difference analysis unit will be described. Text difference analysis is a process that comprehensively evaluates the similarities and differences between transcribed text data. In this embodiment, analysis is performed at multiple levels, from superficial string comparison to deep semantic understanding. First, at the most basic level, exact and partial string matches are detected. To identify which part of the text in the positive information content corresponds to the text in the target content, the Longest Common Subsequence (LCS) algorithm or high-speed substring search using a suffix array is used. Once the corresponding section is identified, the context before and after it is obtained from the positive information database. Next, the edit distance is calculated. Edit distance is the minimum number of editing operations (insertion, deletion, replacement) required to convert one string to another, and the Levenshtein distance is a typical example. The smaller the edit distance, the more similar the two strings are. This system calculates edit distance at the sentence and paragraph levels and normalizes it (by dividing by the string length) to produce a text similarity score ranging from 0 to 1. However, edit distance is sensitive to word order, and even a slight change in word order can result in a large distance. Therefore, to evaluate the similarity of a set of words, the system calculates the repetition rate of N-grams. An N-gram is a sequence of N consecutive words (or tokens), such as a Bi-gram (N=2) or Tri-gram (N=3). Sets of N-grams are generated from the target text and the positive information text, and their repetition rate is evaluated using metrics such as the BLEU score (Bilingual Evaluation Understudy) and the ROUGE score (Recall-Oriented Understudy for Gisting Evaluation). These scores are widely used in the evaluation of machine translation and automatic summarization, and quantify the content similarity of texts.

[0072] As a deeper level of text difference analysis, this embodiment evaluates semantic similarity. This utilizes pre-trained large-scale language models. Specifically, it employs transformer-based models such as BERT (Bidirectional Encoder Representations from Transformers), RoBERTa (A Robustly Optimized BERT Pretraining Approach), ALBERT (A Lite BERT), and ELECTRA, as well as sentence embedding-specific models such as Sentence-BERT (SBERT) and Universal Sentence Encoder (USE). These models have the ability to convert words, sentences, and paragraphs into high-dimensional vectors (typically several hundred to a thousand dimensions) that reflect their meaning. In this system, each sentence in the target content and the corresponding sentence (and surrounding context) in the positive information content are vectorized. Then, the cosine similarity between these vectors is calculated. Cosine similarity is the cosine value of the angle between the two vectors and takes a value in the range of -1 to 1. The closer the value is to 1, the more the vectors point in the same direction and the more semantically similar they are. This semantic similarity analysis shows that sentences with different wording but the same meaning (for example, "The government will review the plan" and "The administration will reconsider the plan") have a high degree of similarity, while sentences with some common words but significantly different meanings (for example, "Promote the plan" and "Cancel the plan") have a low degree of similarity. Furthermore, this system analyzes the syntactic structure and semantic roles of text. Dependency parsing represents the grammatical relationships between words in a sentence (subject, predicate, object, etc.) as a tree structure. Semantic role labeling identifies the semantic role of each element in relation to the predicate (agent, agent, place, time, etc.). By comparing this structural information, it is possible to determine whether the essential meaning is preserved even if the word order changes, or whether the meaning has been altered by the deletion of important elements.For example, in a conditional sentence, "If condition A is met, then conclusion B is true," if the conditional clause "If condition A is met" is deleted and only the conclusion clause "Conclusion B is true" remains, the absence of the conditional clause will be detected through syntactic analysis and evaluated as a significant change in meaning.

[0073] (16th embodiment) Next, as a sixteenth embodiment of the present invention, a detailed method for speech difference analysis in the difference analysis unit will be described. Speech difference analysis is a process that compares time-series data of speech features and detects acoustic and prosodic changes. In this embodiment, an analysis method centered on Dynamic Time Warping (DTW) is adopted. DTW is an algorithm that finds the optimal correspondence between two time-series data and evaluates their similarity. It is widely used in the fields of speech recognition and speaker recognition and has the advantage of being able to robustly calculate similarity even when the speaking speed fluctuates naturally. In this system, time series of various speech features (F0, energy, MFCC, speech rate, etc.) are extracted for corresponding speech sections of the target content and the positive information content. Then, the DTW algorithm is applied to each feature and the DTW distance is calculated. The smaller the DTW distance, the more similar the two time series are, and the larger the distance, the greater the difference. For example, by applying DTW to the time series of F0 (pitch), it is possible to evaluate how well the pitch pattern of the target speech matches the pitch pattern of the positive information speech. If the F0 of the target audio is generally rising or falling, or if unnatural pitch changes have been applied to specific parts, the DTW distance will be large, suggesting the possibility of pitch manipulation. Similarly, applying DTW to a time series of energy (volume) can detect whether specific words or phrases are unnaturally emphasized or, conversely, quieted. Applying DTW to a time series of MFCC can detect changes in timbre, such as degradation or alteration of sound quality due to audio processing or filtering.

[0074] In this embodiment, in addition to DTW, the system also detects and compares specific audio events. Of particular importance is the detection and comparison of pauses (silent intervals). Pauses are an important element that reflects the natural rhythm of speech and the pauses in the speaker's thoughts, and their deletion or shortening can significantly alter the impression of speech. In this system, intervals where the audio power falls below a predetermined threshold are detected as pauses, and their start time, end time, and length are recorded. Then, the number, position, and length of pauses are compared within the corresponding time ranges of the target content and the positive information content. If a pause present in the positive information content is deleted in the target content, or if the length of a pause is significantly shortened, this is recorded as evidence of editing. Conversely, if a pause that does not exist in the positive information content is inserted into the target content, it is also detected as unnatural editing. Furthermore, changes in speech rate are also an important indicator. In this system, the local speech rate is estimated by calculating the number of syllables or phonemes per unit time. Then, the time series of speech rates of the target content and the positive information content are compared. If the overall speaking speed is either faster or slower, it suggests that the audio playback speed may have been altered. Furthermore, if the speaking speed changes unnaturally in only specific sections, it is highly likely that those sections have been edited. Visualizing the spectrogram (time-frequency representation) of the audio signal as an image and detecting discontinuities and noise patterns on that image is also effective. When audio is cut and pasted, discontinuities can occur in the waveform and spectrum at the editing points. This system uses deep learning-based anomaly detection models (e.g., autoencoders, GAN-based detectors) to detect anomalous patterns on the spectrogram. The results of these audio difference analyses are quantified as numerical scores and passed to a subsequent score calculation unit.

[0075] (Embodiment 17) Next, as a 17th embodiment of the present invention, a detailed method for contextual difference analysis in the difference analysis unit will be described. Contextual difference analysis is a process that evaluates which part of the overall positive information content the target content is an excerpt of, and what context is lost as a result of that excerpt. Much of the information distortion occurs when only a part of a statement is excerpted, and important context (premises, reservations, opposing opinions, background explanations, etc.) before and after it is deleted. In this embodiment, first, it is identified which part of the positive information content text the target content text corresponds to. For this, the substring search and speech fingerprinting techniques used in the text difference analysis described above are utilized. Once the corresponding section is identified, the text before and after it (for example, several sentences to several paragraphs before and after) is obtained from the positive information database. Then, the semantic relationship between this surrounding context and the part excerpted in the target content is analyzed. Specifically, the following analysis is performed. First, it is checked whether the deleted surrounding context contains important elements that limit or modify the meaning of the excerpted part, such as conditional clauses, concessive clauses, and adversative clauses. For this, the results of dependency syntactic analysis and semantic role labeling are utilized. For example, if expressions such as "if such and such," "under such and such conditions," "however such and such," or "but such and such" are deleted, this is evaluated as a significant loss of context. Secondly, it is evaluated whether the deleted context contains information that is semantically opposed to or complements the part that was cut out. This is done using semantic similarity between sentences or topic consistency analysis using topic modeling (e.g., LDA: Latent Dirichlet Allocation). For example, in a press conference where both positive and negative statements are made about a certain policy, if only the positive statements are cut out and the negative statements and concerns are deleted, this is evaluated as an unbalanced presentation of information. Thirdly, the structure of the entire positive information content (e.g., logical structure such as introduction, body, and conclusion) is analyzed and it is evaluated which part of that structure the target content extracts. If only the conclusion is cut out and the reasoning and premise arguments are deleted, this is evaluated as a loss of context.

[0076] In this embodiment, these analysis results are integrated to calculate a quantitative metric called the Context Omission Rate. The Context Omission Rate is calculated considering the following factors. First, the ratio of the length of the portion used in the target content to the total length (time or number of characters) of the positive information content is calculated as the "physical coverage rate." A smaller value indicates that more information has been deleted. Next, the "importance" of the deleted portion is evaluated. Importance is calculated for each sentence or paragraph using, for example, TF-IDF (Term Frequency-Inverse Document Frequency) or a BERT-based sentence importance score. If many high-importance portions have been deleted, the Context Omission Rate will be high. Furthermore, the semantic relationship between the deleted portion and the remaining portion is evaluated. Using the results of the semantic similarity and syntactic relevance analysis mentioned above, if the deleted portion has the potential to significantly alter the meaning of the remaining portion, the Context Omission Rate will be high. These factors are weighted and integrated to calculate a Context Omission Rate score ranging from 0 to 1. A higher score indicates that more important context has been lost and the risk of information distortion is high. Furthermore, this system also includes a function to summarize the content of the deleted context in natural language and present it to the user. For example, by generating a specific explanation such as, "In this clipped video, the conditional clause 'if sufficient budget is secured,' which was a prerequisite for the statement, has been removed," users can more easily understand the specific content of the missing context.

[0077] (Embodiment 18) Next, as an 18th embodiment of the present invention, a detailed scoring method in the score calculation unit and a method for integrating multiple scores will be described. The difference analysis unit outputs various analysis results such as text differences, audio differences, and contextual differences. Since these results have different units and scales, they cannot be directly compared or added together. The score calculation unit of this embodiment first normalizes each analysis result and converts it into a dimensionless score in the range of 0 to 1. As a normalization method, minimum-maximum normalization (Min-Max Normalization), Z-score normalization (standardization), and conversion using the sigmoid function can be used. For example, since a larger edit distance indicates a greater difference, to convert this into a similarity score, processing such as subtracting the normalized edit distance from 1 or taking the reciprocal is performed. Similarly, the DTW distance is converted from distance to similarity. On the other hand, since cosine similarity is already in the range of 0 to 1 (or -1 to 1), it is only necessary to adjust the range as needed. The normalized scores are unified as scores indicating the "degree of modification". In other words, the closer the score is to 0, the more it matches the original information and the less it has been altered; conversely, the closer it is to 1, the more it deviates from the original information and the more it has been altered. Next, weight coefficients are set for each score. These weight coefficients reflect the importance and reliability of each score and can be set empirically by the system designer or optimized by machine learning. For example, since changes in audio features directly lead to impression manipulation, a high weight is set for the audio difference score. On the other hand, subtle differences in text expression are considered within an acceptable range, and the weight of the text difference score is set relatively low. Since missing context is an essential cause of information distortion, a high weight is set for the missing context score. These weighted scores are summed up to calculate the Overall Manipulation Score. The Overall Manipulation Score is a single metric that shows how much the target content has been altered from the original information and serves as a primary indicator for users and external systems to judge the reliability of the information.

[0078] In this embodiment, in addition to the overall modification score, detailed scores by category are also provided. This allows users to specifically understand which aspects have been significantly modified. The categories of scores include the following: Firstly, the Text Manipulation Score is a score that integrates the edit distance, N-gram overlap, and semantic similarity of the transcribed text, indicating the extent to which the content of the speech itself has been altered. Secondly, the Audio Manipulation Score is a score that integrates the DTW distance of speech features such as F0, energy, and MFCC, indicating the extent to which the acoustic characteristics of the speech have been altered. Thirdly, the Prosody Manipulation Score is a score that integrates changes in prosodic features such as intonation, rhythm, speech rate, and pauses, indicating the extent to which the impression of the speech has been altered. Fourthly, the Emotion Manipulation Score is a score that shows the change in emotional state estimated by the emotion recognition model, indicating the extent to which the emotional nuances of the utterance have been manipulated. Fifth, the Context Omission Score is the aforementioned context omission rate itself, indicating the extent to which important context has been omitted. These categorized scores are visualized and presented to the user as radar charts or bar graphs. Furthermore, the system also presents the confidence interval and uncertainty of the scores. For example, if the confidence level of speech recognition is low, or if the identification of corresponding sections with correct information content is ambiguous, the reliability of the calculated score will also decrease. By clearly indicating such uncertainties, users can interpret the scores more appropriately and perform additional verification as needed.

[0079] (19th embodiment) Next, as a 19th embodiment of the present invention, a detailed evaluation method in the account evaluation unit and the detection of organized information manipulation will be described. Not only is it possible to detect modifications to individual content, but by analyzing the behavioral patterns of accounts that post and spread that content, it is possible to identify accounts that habitually engage in malicious information manipulation and organized campaigns. The account evaluation unit of this embodiment first collects and stores the posting history of each account. Specifically, for each piece of content posted by a given account, it records the content ID, posting date and time, overall modification score, category score, engagement count (likes, retweets, comments, views), and content type (video, image, text only). Based on this historical data, the following indicators are calculated. Firstly, the Average Manipulation Score is the average value of the overall modification scores of all content posted by that account, and indicates the extent to which the account has posted modified content as a whole. Secondly, the High-Risk Content Ratio is the percentage of content judged as "high-risk" with an overall modification score exceeding a predetermined threshold (e.g., 0.7), indicating whether the account frequently posts dangerous content. Thirdly, Posting Frequency is the number of posts per unit of time; an unusually high posting frequency suggests the possibility of a bot account. Fourthly, Virality is the rate of increase in engagement in a short period after posting (e.g., the first hour or 24 hours), serving as an indicator of organized dissemination activity. Fifthly, Topic Bias analyzes the distribution of topics (e.g., politics, economics, entertainment, etc.) of content posted by the account; an extreme bias towards a particular topic suggests the possibility of activity with a specific intent. These indicators are integrated to assign a "Confidence Score" or "Alert Level" to each account. Accounts with a low Trust Score or a high Alert Level are added to a monitoring list, and their subsequent activities are monitored intensively.

[0080] Furthermore, this embodiment identifies organized information manipulation campaigns by detecting coordinated behavior between multiple accounts. Coordinated behavior refers to actions such as multiple accounts posting identical or similar content almost simultaneously, or retweeting or liking each other's content. Such behavior often exhibits patterns different from naturally occurring human behavior, suggesting the existence of bot networks or organized covert operations. This system visualizes and analyzes the relationships between accounts using social graph analysis techniques. Specifically, a directed graph is constructed with accounts as nodes (vertices) and interactions such as retweets and mentions as edges. By applying community detection algorithms (e.g., Louvain's method, Girvan-Newman's method, Label Propagation method) to this graph, closely linked groups of accounts are extracted as clusters (communities). The extracted communities are then analyzed as follows: First, the average modification score of the content posted by accounts within the community is calculated to evaluate the overall risk level of the community. Second, the synchronization of posting times of accounts within the community is analyzed. For example, if multiple accounts post the same content within minutes, it can be evidence of coordinated dissemination. Thirdly, the profile information of accounts within the community (creation date, number of followers, presence of a profile picture, etc.) is analyzed to assess the proportion of accounts with characteristics of bot accounts (e.g., recently created accounts, extremely low or high number of followers, default profile picture or similar to other accounts). These analysis results are integrated to assess the likelihood that the community is involved in an organized information manipulation campaign. Communities identified as high-risk are reported to system administrators and relevant agencies and are subject to further investigation and countermeasures.

[0081] (20th embodiment) Next, as a 20th embodiment of the present invention, the detailed interface design of the visualization / UI unit and the optimization of the user experience will be described. The wide range of analysis data generated by this system is difficult for general users who are not experts to understand if it is simply a list of numbers. Therefore, the visualization / UI unit of this embodiment converts this complex information into an intuitively understandable visual representation and provides it as an interactive dashboard. The dashboard is implemented as a web application that runs on a web browser and, through responsive design, provides optimal display on various devices such as desktop PCs, tablets, and smartphones. At the heart of the dashboard is the "comparison view." The comparison view divides the screen into left and right halves, displaying the original information content on the left and the target content on the right. A video player, a transcript text display area, and an audio waveform display area are arranged on each side. Users can play the left and right videos in sync. Synchronized playback allows users to visually confirm which parts of the original information are used in the target and which parts have been deleted. In the transcript text display area, the text is displayed along with a timestamp, and the currently playing section is highlighted. Additionally, deleted portions of the target content are displayed in gray or with a strikethrough in the correct information, while changed portions are highlighted in red. When a user clicks on a specific part of the text, the video jumps to that point, allowing them to instantly hear the audio. In the audio waveform display area, the waveforms of the left and right audio are displayed side by side, making the differences in amplitude (volume) immediately obvious. Below the waveforms, a time-series graph of pitch (F0) and a time-series graph of emotion (for example, the trajectory of arousal and pleasure / displeasure) are overlaid. This allows for a visual comparison of when the tone of voice changed and how emotions shifted.

[0082] At the bottom of the comparison view is the "Difference Highlights" panel. This panel lists the major differences detected in chronological order. Each difference item displays the type of difference (e.g., "Pause Removal," "Pitch Change," "Context Missing"), the time it occurred, the magnitude of the difference (score), and a brief description. When the user clicks on a difference item, the video jumps to that time and the relevant section is displayed in detail. For example, clicking on "Pause Removal" will display the pause (silent section) as a red band on the timeline in the original source, clearly indicating that the pause does not exist in the target source. Similarly, clicking on "Pitch Change" will enlarge the pitch graph for that section, showing in detail the difference between the pitch curves of the original source and the target. In the upper right corner of the screen is the "Score Panel." The score panel displays the overall modification score as a large number and meter, and below that, category-specific scores (text modification, audio modification, prosodic modification, sentiment manipulation, context missing) are displayed as radar charts or bar graphs. The scores are color-coded: green indicates "Safe," yellow indicates "Caution," and red indicates "Danger." Users can intuitively grasp the risk level of content at a glance by simply looking at the score panel. Furthermore, the dashboard includes a "Detailed Analysis" tab where users can view more detailed analysis results. The Detailed Analysis tab displays detailed statistics on audio features, detailed alignment results of text differences, the full text of deleted contexts, and detailed time-series data of sentiment recognition. In addition, the "Export" function allows users to download the analysis results as a PDF report or JSON data. This makes it possible for users to save the analysis results, share them with others, or use them in external systems.

[0083] (21st embodiment) Next, as a 21st embodiment of the present invention, the details of the real-time monitoring function will be described. This function acquires and analyzes the audio stream in real time while important press conferences, speeches, live broadcasts, etc., are in progress, predicts speech points that are at high risk of being later excerpted and used to distort information, and issues a warning. This function allows public relations personnel and risk managers to take proactive measures before problematic excerpted videos are disseminated. In this embodiment, first, the URL or streaming source of the live broadcast to be monitored is registered in the system in advance. When the live broadcast starts, the system automatically connects to the stream and acquires the audio data in real time. The acquired audio is divided into short time units (e.g., a few seconds to tens of seconds) and sequentially fed into the analysis pipeline. In the analysis pipeline, the audio is first transcribed in real time by a speech recognition model. At this time, by using a model that supports streaming speech recognition (e.g., online speech recognition, incremental speech recognition), the recognition results are output sequentially from the time of speech without waiting for the speech to finish. Next, the transcribed text and audio features are input into the "tick-out likelihood prediction model." This prediction model is trained using machine learning (e.g., gradient boosting trees, neural networks) with training data from numerous past instances of speeches that were actually picked out and caused problems. The prediction model comprehensively evaluates the content of the speech (words used, sentence structure, topic), audio features (pitch, volume, speaking speed, emotion), and the context of the speech (the overall topic of the conference, the current flow of the discussion), and outputs a score indicating the high risk that the speech will be picked out in the future.

[0084] For statements that the predictive model identifies as high-risk, an alert is immediately generated and notified to the public relations officer's dashboard. The alert includes a risk score, a transcript of the statement, the time of the statement, and the reason why it was deemed high-risk (e.g., "strong assertions," "words that could incite conflict," "emotional outburst," "complex context that is easily misunderstood"). Upon receiving this alert, the public relations officer can quickly take the following actions: First, provide supplementary explanations or clarify the context of the problematic statement at an appropriate time, either after the press conference or during the conference. Second, prepare a Q&A in advance to address anticipated criticisms and misunderstandings and publish it on the official website and social media. Third, promptly release a full transcript or video of the press conference to facilitate access to accurate information. Fourth, strengthen social media monitoring and establish a system to respond immediately if clips containing the problematic statement begin to spread. This real-time monitoring function is a groundbreaking feature that shifts the paradigm of information risk management from reactive to proactive prevention. To continuously improve the accuracy of the prediction model, this system collects prediction results and actual results (whether or not the statement was later edited) as feedback and uses it to retrain the model. As a result, the prediction model evolves over time, enabling more accurate predictions.

[0085] (22nd embodiment) Next, as a 22nd embodiment of the present invention, the details of the subtitle / caption alteration detection function will be described. In video content, not only audio and video, but also text elements such as subtitles and captions are important means of conveying information. However, these subtitles and captions can be easily altered or fabricated and may be misused as a means of giving viewers a false impression. In this embodiment, alterations and fabrications are detected by automatically detecting and extracting subtitles and captions in a video and comparing their content with the audio transcription results and a database of correct information. First, Optical Character Recognition (OCR) technology is used to detect subtitle / caption areas from video frames. Specifically, the video is sampled at regular intervals (for example, one frame or several frames per second), and a deep learning-based text detection model (for example, models such as EAST, DBNet, and CRAFT) is applied to each frame image. These models have the ability to detect text areas in images as rectangular or polygonal bounding boxes. An OCR engine (e.g., Tesseract, Google Cloud Vision API, Amazon Textract, or a deep learning-based OCR model) is applied to the detected text region to extract the text content. The extracted text is recorded along with its appearance time (frame number or timestamp). Next, the extracted subtitle / caption text is compared with the audio transcription result. Ideally, subtitles / captions should accurately reflect the audio content. However, in modified videos, subtitles may display content different from the audio, or statements not present in the audio may be added as subtitles. This system evaluates the similarity between the subtitle text and the transcription text using the aforementioned text difference analysis methods (edit distance, N-gram overlap, semantic similarity). Low similarity, i.e., a large discrepancy between the subtitle and audio content, indicates the possibility of subtitle modification.

[0086] Furthermore, in this embodiment, the extracted subtitle / caption text is compared with official transcripts and official subtitles stored in the official information database. Videos provided by official sources often have accurate subtitles. This system also stores official subtitle data in the official information database. If the subtitles of the target content do not match the official subtitles in the official information database, it indicates that modified or unofficial subtitles have been added. The system also analyzes the visual features of the subtitles / captions (font, color, size, placement, background, animation effects). Official subtitles usually follow consistent style guidelines and use specific fonts and placements. On the other hand, unofficially added subtitles and captions often have different styles. This system extracts the visual features of the detected text area and evaluates the degree of match with the official style. A low degree of match indicates that unofficial subtitles / captions have been added. Furthermore, natural language processing is used to analyze whether the content of the subtitles / captions contains sensational expressions, exaggerations, assertive claims, or emotionally charged language (so-called "clickbait titles" or "hype captions"). If such expressions are included, it is highly likely that the editing is intended to manipulate the viewer's impression. This system evaluates the content of subtitles and captions using a dictionary of sensational expressions and a sensational expression detection model trained through machine learning. These analysis results are integrated to calculate a "subtitle modification score." A high subtitle modification score indicates that the content is highly likely to contain altered or misleading subtitles or captions. Problematic subtitles and captions detected are displayed as warning marks on the video timeline in the visualization / UI section, and users can click to view details.

[0087] (23rd embodiment) Next, as a 23rd embodiment of the present invention, the details of the speech synthesis and deepfake speech detection function will be described. In recent years, advances in AI technology have led to the development of technologies for synthesizing human speech with high accuracy (TTS: Text-to-Speech, speech cloning) and technologies for converting existing speech into different content (speech conversion, voice changer). While these technologies are used for legitimate purposes, they are also used by malicious users to generate "deepfake speech" that makes it appear as if a real person has said something they did not. In this embodiment, the possibility that the speech in the target content is synthesized by AI rather than actual human speech is detected. Deepfake speech detection is an active research field, and various approaches have been proposed. This system uses a combination of the following methods. First, a method for detecting anomalies in the acoustic characteristics of the speech. AI-synthesized speech may show differences in subtle acoustic characteristics compared to natural human speech. For example, the pitch variation pattern may be unnaturally smooth or, conversely, unnaturally discontinuous, the energy distribution in a particular frequency band may be abnormal, or the transitions between phonemes may be unnatural. This system uses a large dataset of real human speech and synthesized speech to construct a deep learning model (e.g., CNN, RNN, Transformer, or a combination thereof) that learns the subtle differences between them. This model analyzes the spectrogram and acoustic features of the input speech and outputs the probability (deepfake score) that the speech is synthesized. Secondly, it employs a method for analyzing phase information. Many speech synthesis systems prioritize amplitude spectra, and sometimes have poor reproducibility of phase spectra. This system analyzes the phase spectrum of the speech signal in detail and detects unnatural patterns. Thirdly, it employs a method for evaluating speech consistency. In long speeches, the speaker's voice quality, pitch, and speaking style fluctuate within a natural range, but maintain a certain consistency. On the other hand, this consistency can be lost in synthesized speech. This system statistically evaluates the consistency of speaker features throughout the entire speech.

[0088] Furthermore, this embodiment also employs speaker verification technology. Speaker verification is a technique for verifying whether a given voice belongs to a specific speaker. In this system, a large number of voice samples from official speakers (e.g., politicians, corporate public relations representatives) are stored in a positive information database. From these samples, the voice features (speaker embedding vectors) of each speaker are extracted, and a speaker model is constructed. The speaker embedding vectors are extracted by a deep learning-based speaker recognition model (e.g., x-vector, d-vector, ECAPA-TDNN). Similarly, speaker embedding vectors are extracted from the audio of the target content and compared with the speaker model in the positive information database. If the speaker embedding vector of the target audio deviates significantly from that of the speaker model in the positive information database, it indicates that the target audio is not the actual speaker's voice, i.e., it is a synthesized voice or the voice of someone else. Conversely, if the speaker embedding vectors match, but the deepfake detection model outputs a high deepfake score, it indicates that advanced voice cloning technology may have been used. The results of these multiple detection methods are integrated to calculate an overall "audio authenticity score." A low audio authenticity score indicates a high probability that the content's audio is synthesized or deepfake, and a warning is displayed to the user. The system also includes a function to estimate the generation method of detected deepfake audio. Different speech synthesis technologies (e.g., WaveNet, Tacotron, FastSpeech, VITS, HiFi-GAN, etc.) may each have unique acoustic characteristics. By learning these characteristics, the system estimates which technology was used and provides this information as reference for countermeasures.

[0089] (24th embodiment) Next, as a 24th embodiment of the present invention, the details of the video modification detection function will be described. While the present invention mainly focuses on the analysis of speech and text, modification of video content is also an important issue. For example, there are methods such as editing the facial expressions and gestures of speakers, modifying the background, synthesizing other videos, and creating deepfake videos (face swapping). In this embodiment, an auxiliary function for detecting video modification is provided. First, edit points (cuts, transitions) in the video are detected. In video editing, it is common to cut and paste multiple scenes, and these edit points may appear as discontinuities in the video. In this system, the image similarity between consecutive frames (e.g., histogram comparison, structural similarity index SSIM, deep learning-based feature comparison) is calculated, and the location where the similarity drops sharply is detected as an edit point. The location and frequency of the detected edit points are compared with the original information content to evaluate whether or not there is unnatural editing. Next, the faces of people in the video are detected, and their facial expressions and gaze are analyzed. For face detection, deep learning-based face detection models (e.g., MTCNN, RetinaFace, YOLO-Face) are used. For detected face regions, facial landmarks (feature points such as eyes, nose, and mouth) are detected, and an expression recognition model (e.g., a model trained on the FER2013 dataset) is applied to estimate facial expressions (joy, sadness, anger, surprise, neutral, etc.). In addition, a gaze estimation model is used to estimate where the person is looking. By comparing this information with the emotion estimated from the audio, the consistency between audio and video is evaluated. For example, if anger is detected in the audio, but a smiling expression is displayed in the video, it may indicate that the video has been tampered with.

[0090] Furthermore, this embodiment also performs deepfake video detection. Deepfake videos are created using deep learning techniques such as GANs (Generative Adversarial Networks) to replace the face of one person with the face of another person or to manipulate facial expressions. There are several approaches to deepfake video detection. First, there is a method to detect subtle unnaturalness in the face region. Even high-resolution deepfake faces may contain unnatural artifacts (artificial traces) at a subtle level. Examples include unnatural synthesis marks at the boundaries of the face, abnormal blinking patterns, unnatural textures of teeth and eyes, and inconsistent lighting. This system uses deep learning models (e.g., XceptionNet-based models, EfficientNet-based models) that learn these subtle features to detect deepfakes. Second, there is a method to evaluate temporal consistency. Deepfakes often generate each frame independently, which can lead to a loss of temporal consistency between frames. For example, features such as unnatural discontinuity in facial orientation and expression changes, and unstable fluctuations in the position of facial landmarks may appear. This system analyzes changes in facial features between consecutive frames and detects unnatural fluctuations. Thirdly, there is a method for analyzing physiological signals. In human faces, physiological signals such as heart rate and respiration appear as subtle color changes (rPPG: remote PhotoPlethysmography). In real human footage, these physiological signals are detected, but in deepfake footage, they may not be detected or may show unnatural patterns. This system analyzes color changes in the facial region over time and evaluates the presence and naturalness of physiological signals. The results of these video modification detections are integrated with the audio and text modification detection results and reflected in the overall content modification score.

[0091] (25th embodiment) Next, as a 25th embodiment of the present invention, the details of the multilingual support function will be described. Information distortion and the spread of misinformation are not limited to specific language areas, but are global problems. Furthermore, when official information disseminated in one language is translated into other languages, its content may be distorted intentionally or unintentionally. In this embodiment, a function is provided to make the system compatible with a multilingual environment. First, the speech recognition unit utilizes a multilingual speech recognition model. As mentioned above, models such as OpenAI Whisper and Google USM support dozens to over a hundred languages, and a single model can recognize speech in multiple languages. In this system, these models are utilized to automatically identify the language of the input speech and then transcribe it in the appropriate language. For language identification, a speech-based language identification model (for example, an x-vector-based model) or the language probability output by the speech recognition model itself is used. Next, the text analysis unit utilizes a multilingual natural language processing model. Models such as mBERT (multilingual BERT), XLM-RoBERTa, and XLM-R, which are multilingual versions of BERT, support over 100 languages ​​and enable the calculation of semantic similarity across languages ​​and zero-shot transfer learning. This system uses these models to perform text comparison and semantic analysis between different languages. For example, when an official statement in Japanese is translated into English and spread on social media, to verify the accuracy of the translation, the Japanese positive information text and the English target text are vectorized using a multilingual embedding model, and their cosine similarity is calculated. A low similarity score indicates that the translation may be inaccurate or intentionally distorted.

[0092] Furthermore, this embodiment also provides a machine translation quality evaluation function. When official information is translated into other languages, the system automatically translates the official text using a reliable machine translation service (e.g., Google Translate, DeepL, Microsoft Translator). This automated translation result is then compared with the translated text used in the target content. If the similarity between the two is high, the target translation is evaluated as reasonable; if the similarity is low, it indicates that the translation may be inaccurate or altered. In addition, automated evaluation metrics such as the BLEU score, METEOR score, and COMET score are used as indicators to evaluate the quality of the translation. Low scores indicate poor translation quality. Furthermore, this system also provides a function to support reviews by human experts with expertise in each language in order to detect cultural and contextual mistranslations and distortions. Specifically, the system automatically lists and prioritizes suspicious translations it has detected and presents them to experts. Experts efficiently review these sections and make a final judgment. In this way, by combining automated analysis and human expertise, highly accurate information verification is achieved even in multilingual environments. Furthermore, the system's user interface will also be localized into multiple languages, allowing users worldwide to access it in their own language. The interface translation will combine professional translation services with ongoing localization management.

[0093] (26th embodiment) Next, as a 26th embodiment of the present invention, the details of privacy protection and data security functions will be described. This system handles a large amount of data that may contain personal information, such as audio, video, and text. Therefore, privacy protection and data security are of paramount importance in system design. In this embodiment, the following multi-layered security measures are implemented. First, data encryption. All data, such as collected audio and video files, transcribed text, and analysis results, are encrypted at rest using a strong encryption algorithm (e.g., AES-256). In transit, the TLS (Transport Layer Security) protocol is used to prevent eavesdropping and tampering during transmission. Second, access control. Access to the database and analysis functions is controlled by a strict authentication and authorization mechanism. Users are granted only permissions according to their role through Role-Based Access Control (RBAC). For example, general users can only view the results of content they have requested to be verified, while administrators can access data for the entire system. Third, data minimization and retention period limitations. This system collects only the minimum data necessary to achieve its objectives and does not collect unnecessary data. Furthermore, collected data is automatically deleted after a predetermined retention period (e.g., a period mandated by law or a period required for business purposes). Users have the right to request the deletion of their data, and the system will respond to such requests promptly.

[0094] Fourth, anonymization and pseudonymization. When processing data containing personally identifiable information (PII) in this system, anonymization or pseudonymization is performed as much as possible. For example, information that could identify the speaker is removed from audio data, and proper nouns are replaced with common nouns in transcribed text. However, for the purpose of this system, when verifying official statements by public figures, the speaker's identity information is necessary, and in such cases, processing is performed based on appropriate legal grounds (e.g., public interest, legitimate interest). Fifth, audit log recording is performed. All important operations performed within the system (data access, analysis execution, configuration changes, etc.) are recorded as detailed audit logs. The audit logs record the user who performed the operation, the date and time of the operation, the content of the operation, and the data targeted, and are cryptographically protected to prevent tampering. Audit logs are used for investigating security incidents and compliance audits. Sixth, security assessment and vulnerability management are performed. This system undergoes regular security assessments (penetration testing, vulnerability scans), and any vulnerabilities found are promptly fixed. Furthermore, the software libraries and frameworks used are always kept up-to-date with the latest security patches. Seventh, there is Privacy Impact Assessment (PIA). When introducing new features to this system or changing data processing methods, a Privacy Impact Assessment is conducted in advance to identify and evaluate privacy risks and take appropriate measures. Through these multi-layered security measures, this system provides advanced information verification capabilities while protecting user privacy.

[0095] (27th embodiment) Next, as a 27th embodiment of the present invention, the details of a privacy-preserving model learning function using Federated Learning will be described. To improve the accuracy of various AI models of this system (speech recognition, emotion recognition, deepfake detection, etc.), a large amount of training data is required. However, this data often contains information that should be kept private, making it difficult or inappropriate to aggregate it on a central server. In this embodiment, federated learning technology is used to provide a mechanism in which multiple distributed data holders (e.g., servers from different organizations or different regions) cooperate to learn a model without centralizing the data. The basic flow of federated learning is as follows: First, the central server distributes an initial model to each data holder (client). Each client learns (fine-tunes) the distributed model using the local data it holds. At this time, the data itself is not sent externally but remains within the client. Once learning is complete, each client sends only the updated model parameters (weights) or the amount of parameter update (gradient) to the central server. The central server aggregates (e.g., averages) the parameters received from each client and generates a new global model. This new global model is then distributed to each client again, and the above process is repeated. By repeating this process multiple times, a highly accurate model can be trained overall without directly sharing data from each client.

[0096] In this embodiment, federative learning is combined with further privacy protection technologies. First, there is differential privacy. By intentionally adding noise to the model parameters that each client sends to the central server, the risk of information leakage from individual data samples is mathematically limited. Differential privacy allows the strength of privacy protection to be adjusted with parameters (ε and δ), enabling control over the trade-off between privacy and model accuracy. Second, there is secure aggregation. This technology encrypts the parameters sent by each client, and the central server can calculate the aggregate (sum or average) without decrypting the parameters of individual clients, while they remain encrypted. This enhances privacy as even the central server cannot know the model updates of individual clients. Third, there is client selection and sampling. Instead of all clients participating in learning every time, only a randomly selected subset of clients participate in each round. This reduces communication costs and decreases dependence on specific clients. The federative learning function of this embodiment enables the system to build more accurate and general-purpose AI models while protecting privacy, through collaboration among diverse data holders around the world. For example, one possible application is for news organizations and government agencies in different countries to jointly train a deepfake detection model without disclosing their respective official information data to the public.

[0097] (28th embodiment) Next, as a 28th embodiment of the present invention, the details of the Explainable AI (XAI) function will be described. Many AI models used by this system, especially deep learning models, have high accuracy, but their internal operation is a black box, making it difficult for humans to understand why they made such decisions. However, the modification scores and warnings output by this system influence important decision-making by users and stakeholders, so it is important to be able to clearly explain the basis for them. In this embodiment, the technology of Explainable AI is used to provide a function that visualizes and explains the basis for the decisions of the AI ​​model. Explainable AI methods can be broadly divided into model-agnostic methods and model-specific methods. This system uses the following methods. First, there is LIME (Local Interpretable Model-agnostic Explanations). LIME is a method that explains which input features contributed to a prediction by approximating a specific prediction of a complex black-box model with a simple model (e.g., a linear model) that can be interpreted in its neighborhood. In this system, for example, if a deepfake detection model determines that a certain audio is "synthesized speech," LIME is applied to identify which time intervals and frequency bands of the audio strongly influenced that determination, and this is presented to the user. Secondly, there is SHAP (SHapley Additive exPlanations). SHAP is a method that fairly evaluates how much each input feature contributed to the prediction, based on the Shapley value from game theory. SHAP provides a theoretically sound explanation and can be applied to various models. In this system, the SHAP value of each feature (F0, energy, MFCC, etc.) is calculated in the emotion recognition model and modification score calculation, and the most important features are visualized.

[0098] Thirdly, there is the visualization of the attention mechanism. Transformer-based models (such as BERT and Whisper) have an internal attention mechanism and calculate attention weights that indicate which parts of the input are being focused on during processing. This system visualizes these attention weights to show which parts of the text or audio the model prioritized when making its decisions. For example, when calculating the semantic similarity of text, a heatmap is displayed showing which word relationships BERT focused on. Fourthly, there are counterfactual explanations. Counterfactual explanations are a method of providing explanations in the form of "If the input had been changed in this way, the prediction would have been changed in this way." This system generates explanations such as, "If this pause had not been removed, the modification score would have been 0.3 lower." This allows users to specifically understand which editing operations had a significant impact on the modification score. Fifthly, there are prototype-based explanations. By presenting the most similar example (prototype) from the training data to the current input, the system provides an explanation in the form of "This content is similar to similar past examples." This system searches past examples similar to the detected modification pattern from the original information database and the modification example database and presents them to the user. Through these explainable AI functions, the system not only presents a score but also clearly explains the reasoning behind it, thereby gaining the user's understanding and trust and supporting appropriate decision-making.

[0099] (29th embodiment) Next, as a 29th embodiment of the present invention, the details of the Continual Learning function will be described. Information manipulation techniques and synthetic content generation techniques using AI technology are evolving daily. If this system continues to use a model that has been trained once, it may not be able to adapt to new techniques, and the detection accuracy may decrease. In this embodiment, the system uses continuous learning technology to provide a function that adapts to new data and new attack methods while the system is in operation and continuously updates the model. One of the challenges of continuous learning is "catastrophic forgetting." This is a phenomenon in which knowledge previously learned is forgotten when a model is retrained with new data. In this system, the following methods are employed to prevent catastrophic forgetting. First, there is the rehearsal method. When training with new data, past knowledge is retained by mixing in a portion of past training data (representative samples). In this system, important past cases are stored in a memory buffer and reused when training new data. Second, there is a regularization-based method. For example, techniques such as EWC (Elastic Weight Consolidation) and LwF (Learning without Forgetting) prevent forgetting by imposing constraints on model parameters so that they do not change significantly in the direction that was important in past tasks. Thirdly, there are dynamic architecture techniques. By dynamically extending the model architecture for new tasks and new data (for example, by adding new neurons or layers), new knowledge is acquired while retaining past knowledge. In this system, these techniques are used in combination, and the model is periodically updated using new positive information data collected during operation, new modification examples, and user feedback.

[0100] Furthermore, this embodiment incorporates an active learning method. Active learning is a method in which a model actively selects the data sample with the highest learning effect, obtains labels (correct answers) for that sample, and learns from them. In this system, from among the many content encountered during operation, cases where the model was unsure of what to do (for example, cases where the modification score showed an intermediate value, or cases where the judgments of multiple models did not match) are prioritized and reviewed by human experts. Once the experts assign correct labels (whether or not it has been modified, and what type of modification it is), the model is retrained using that data. This process allows for efficient improvement of model accuracy with limited human resources. In addition, this system also has a mechanism to quickly share information about new types of attacks or modification methods across the entire system and take countermeasures when they are detected. For example, if an organization detects an attack using a new deepfake generation technique, that information (attack characteristics, detection method) is shared with other organizations using this system. The shared information is used to update each organization's models, improving overall security. Through this continuous learning and community-based information sharing, the system maintains up-to-date defensive capabilities against ever-evolving threats.

[0101] (30th embodiment) Next, as a 30th embodiment of the present invention, the details of the external system integration function will be described. This system not only operates independently but also functions within a broader information ecosystem by integrating with other information systems and platforms. In this embodiment, the following external integration functions are provided. First, there is integration with SNS platforms. This system uses the official APIs of major SNS platforms (Twitter / X, Facebook, Instagram, YouTube®, TikTok, etc.) to collect content, post analysis results, display warnings, etc. For example, integration could involve requesting the SNS platform to display a warning label for content determined to be high-risk by this system, or posting a link to the system's verification results page as a comment. However, these integrations are carried out in accordance with the policies and terms of use of each platform. Second, there is integration with fact-checking organizations. There are many organizations around the world that perform professional fact-checking (for example, organizations that are members of the International Fact-Checking Network IFCN). This system takes in fact-checking results published by these organizations (for example, structured data using the ClaimReview schema) and integrates them into a positive information database. Furthermore, the system provides a function to report suspicious content it detects to fact-checking organizations and request detailed verification. Thirdly, it facilitates collaboration with news organizations. News organizations can use the system to verify the accuracy of information they report and investigate the veracity of information circulating on social media. The system provides a dedicated dashboard and API for news organizations to support rapid information verification.

[0102] Fourth, there is collaboration with government and public institutions. As disseminators of official information, government and public institutions will register official content in this system, enriching the database of correct information. Furthermore, in situations where the spread of misinformation has a significant impact on society, such as during election periods or public health emergencies, this system will be used for the early detection and response to misinformation. This system will provide government institutions with a real-time monitoring dashboard, periodic report generation, and emergency alert functions. Fifth, there is collaboration with academic research institutions. This system will provide academic research institutions with anonymized datasets and analysis results to support the research and development of misinformation detection technology. In addition, new algorithms and models developed by research institutions will be integrated into this system to improve performance. Through such industry-academia collaboration, this system can always incorporate cutting-edge technology. Sixth, there is integration with browser extensions. This system will work with client applications provided as extensions for web browsers. When users are browsing social media or news sites, the extension will automatically send content to this system and display verification results in real time. For example, it will provide an experience where a warning icon appears next to a suspicious video, and clicking it displays detailed analysis results. Seventh, the system integrates with the company's internal systems. Companies can integrate this system with their public relations monitoring systems, brand protection systems, and compliance systems to detect and respond to misinformation and reputational damage related to the company at an early stage. The system provides companies with customizable APIs, webhooks, dashboard embedding capabilities, and more. Through these diverse external integration functions, the system serves as a foundation for improving the reliability and transparency of the entire information ecosystem.

[0103] (31st Embodiment) Next, as a 31st embodiment of the present invention, the details of the official audio authentication function using audio fingerprinting technology will be described. An audio fingerprint is a unique identifier that represents the content of audio content, and is a technology widely used in music identification services (e.g., Shazam). In this embodiment, this technology is applied to the authentication and tracking of official information. Specifically, an audio fingerprint is generated for audio such as press conferences, speeches, and interviews released by official organizations, and registered in the official information database. The generation of the audio fingerprint uses a method that extracts characteristic landmarks (peaks) from the audio spectrogram and encodes their temporal and frequency arrangement patterns as a hash value. Representative algorithms include Philips' Robust Hashing and the Constellation algorithm used in Shazam. In this system, these algorithms are implemented to generate a robust fingerprint that can identify the original official audio even from short audio fragments of a few seconds to tens of seconds. By extracting a fingerprint from the audio of clipped or remixed videos circulating on social media and comparing it with fingerprints registered in a database of official information, it is possible to instantly identify which official content and which part of the audio was extracted from. This matching function works with high accuracy even if the audio is somewhat compressed, contains noise, or has altered pitch or speed. Information about the identified official content (URL of the original video, title, publication date and time, and relevant time range) is presented to the user, and specific information such as "This clipped video uses the portion from □□ minutes □□ seconds to □□ minutes □□ seconds of the press conference held by the Ministry of XX on △△ year △ month △ day" is displayed.

[0104] Furthermore, this embodiment also provides a function to verify the authenticity of official audio using audio fingerprints. When an official institution publishes audio, it records the audio's fingerprint on a blockchain or public timestamp service, creating proof that it cannot be tampered with later. Users can calculate the fingerprint of any audio file and compare it with the record on the blockchain to verify that the audio was officially published and has not been tampered with. This mechanism cryptographically guarantees the reliability of official information. In addition, this system allows for the distributed management of the audio fingerprint database among multiple organizations using the aforementioned federated learning mechanism. Each organization locally maintains the fingerprints of the official audio it manages, and when a matching query is received, it searches only its own database and returns the result. The central server only aggregates the results from each organization and does not need to access the audio data or fingerprint details of individual organizations. This makes it possible to build a global audio verification network while maintaining the privacy and autonomy of each organization. Audio fingerprinting technology has low computational costs and can search large databases quickly, making it suitable for applications that require real-time processing. This system efficiently identifies content, including official audio, by rapidly performing fingerprint matching on a large volume of videos collected by the SNS monitoring department using parallel processing.

[0105] (32nd embodiment) Next, as a 32nd embodiment of the present invention, the details of the audio edit point detection function using time-series anomaly detection technology will be described. When audio is edited by cutting and pasting, unnatural discontinuities or anomalies may occur in the time series of audio features at the edit points. In this embodiment, such edit points are automatically detected using time-series anomaly detection technology. Time-series anomaly detection is a technology that detects abnormal locations in time-series data that deviate from the normal pattern. In this system, an anomaly detection model is applied to the time series of audio features (F0, energy, MFCC, spectral centroid, etc.) as input. The following methods can be used as anomaly detection models. Firstly, statistical methods. The time series of audio features is modeled using an autoregressive model (AR), a moving average model (MA), an autoregressive moving average model (ARMA), or a state-space model (e.g., a Kalman filter), and locations where the deviation between the model's predicted value and the actual observed value is large are detected as anomalies. Secondly, machine learning-based methods. For example, anomaly detection algorithms such as Isolation Forest, One-Class SVM, and Local Outlier Factor (LOF) are used to detect outliers in the feature space. Thirdly, there are deep learning-based methods. Autoencoders (AEs) or variational autoencoders (VAEs) are used to learn time series of normal speech features and detect locations with large reconstruction errors as anomalies. Additionally, methods using LSTM (Long Short-Term Memory) networks to predict the next value in the time series and detect locations with large prediction errors as anomalies are also effective.

[0106] In this embodiment, these anomaly detection methods are combined to improve detection accuracy. Specifically, multiple anomaly detection models are executed in parallel, and the anomaly scores output by each are integrated. Possible integration methods include averaging the scores, using the maximum value, or using voting (only locations judged as anomaly by multiple models are considered anomalies). Further detailed analysis is performed on the detected anomaly locations (candidate edit points). For example, the audio waveform for several frames before and after the anomaly location is observed in detail to detect discontinuities in the waveform (phase jumps, abrupt changes in amplitude). It is also evaluated whether the speaker's voice quality or background noise characteristics have changed before and after the anomaly location. These additional analyses reduce false positives (detecting natural audio changes as anomalies instead of edits). Detected edit points are recorded along with a timestamp and displayed as markers on the audio waveform timeline in the visualization / UI section. Users can click on the marker to play the audio at the corresponding location and check whether edits have been made. Furthermore, the number and location of edit points are compared with the correct information content to quantitatively evaluate the extent of editing. If numerous edit points are detected, it indicates that the audio may have been significantly altered, and this is reflected in the modification score. This time-series anomaly detection technique can be applied not only to audio but also to video edit point detection. In the case of video, the time series of features such as image similarity between frames, color histograms, and optical flow are analyzed to detect anomalous discontinuities.

[0107] (33rd embodiment) Next, as a 33rd embodiment of the present invention, the details of the integrated analysis function using multimodal deep learning will be described. In previous embodiments, we have described an approach in which each modality (format of information), such as speech, text, and video, is analyzed individually and the results are integrated afterward. In this embodiment, integrated analysis is performed using a multimodal deep learning model that simultaneously receives multiple modalities as input and learns the interactions and consistency between them. Multimodal deep learning has seen significant progress in recent research and has various applications, such as integrated understanding of images and text (e.g., image caption generation, visual question answering), integrated understanding of speech and text (e.g., utilization of text information in speech emotion recognition), and integrated understanding of speech and video (e.g., verification of the consistency between the speaker's mouth movements and speech). In this system, the following multimodal model is constructed. First, there is a speech-text integrated model. This model simultaneously receives acoustic features of speech (spectrogram, speech feature time series) and embedding vectors of transcribed text as input, fuses the information from both, and determines whether or not modifications have been made and what type of modifications they are. For example, if the audio expresses anger, but the text content is neutral, this discrepancy between audio and text can be detected. Secondly, there is the audio-video integration model. This model takes audio and video of the speaker's face (facial landmarks, expressions, mouth movements) as input simultaneously and evaluates the synchronization and consistency between the audio and video. For example, if the audio and mouth movements do not match (lip-sync mismatch), it may indicate the possibility of deepfake video or dubbing. Thirdly, there is the text-video integration model. This model simultaneously analyzes the text of subtitles / captions in the video and the content of the video (scene, object, person) and evaluates the consistency between the two. For example, if the subtitles say "meeting place," but the video is clearly an outdoor scene, this will be detected as a discrepancy.

[0108] Possible architectures for these multimodal models include the following: First, a dedicated encoder (feature extractor) is prepared for each modality. For speech, a CNN or RNN-based encoder is used; for text, a language model such as BERT or GPT is used; and for video, an image encoder such as CNN or Vision Transformer (ViT) is used. The feature vectors output from each encoder are integrated in a fusion layer. Fusion methods include simple concatenation, element-wise multiplication or summation, weighted fusion using an attention mechanism, and a cross-modal attention mechanism (using information from one modality to decide which part of the other modality to focus on). The fused feature vector is input into a classifier or regressor, and the final judgment (whether or not it was modified, modification score, type of modification, etc.) is output. Such multimodal models are trained end-to-end using a large amount of multimodal data (data containing speech, text, and video). During training, a multi-task learning framework is used, in which each encoder learns a better feature representation by simultaneously learning a judgment task for each modality individually and a judgment task for multimodal integration. Through the multimodal integration analysis of this embodiment, the system can detect complex and sophisticated information manipulations that are difficult to detect with a single modality.

[0109] (34th embodiment) Next, as a 34th embodiment of the present invention, the details of the system improvement function utilizing user feedback will be described. The analysis results of this system are not always perfectly accurate, and false positives (judging unaltered content as altered) and omissions (overlooking altered content) may occur. In this embodiment, a mechanism is provided to actively collect feedback from users and utilize it to improve the system. Specifically, users can provide the following feedback on the analysis results presented by this system: Firstly, feedback on the accuracy of the judgment. Users can provide evaluations such as "This judgment is correct" or "This judgment is incorrect" with a simple button operation. Secondly, detailed comments. Users can input detailed comments in text, such as why they think the judgment is correct or incorrect, and what points were overlooked. Thirdly, provision of additional information. Users can provide URLs of relevant official information that the system has not referenced, or information on similar cases. This feedback is stored in a database and used for the following purposes. Firstly, weaknesses in the system are identified by analyzing cases of false positives and omissions. For example, tendencies may become apparent, such as a tendency to miss certain types of modifications (e.g., subtle pitch changes) or low accuracy with certain speakers or languages. To address these weaknesses, improvement measures such as retraining the model, tuning parameters, and adding new features are implemented. Next, the examples that users rated as correct and incorrect are used as positive and negative examples, respectively, to retrain the model. This brings the model closer to the user's judgment criteria.

[0110] Furthermore, this embodiment also considers a mechanism to provide incentives to users who provide feedback. For example, users who provide a lot of useful feedback will be rewarded with access to advanced features of the system or evaluation points within the community. This will encourage active user participation. The system also includes a mechanism to evaluate the reliability of feedback. Not all users' feedback is equally accurate, and some may contain misinformation or malicious feedback (intentionally providing false evaluations to confuse the system). The system tracks the accuracy of each user's past feedback (for example, the percentage of their feedback that was later confirmed as correct by experts) and calculates a confidence score for each user. Feedback from users with high confidence scores is reflected in model learning with greater weight, while feedback from users with low confidence scores is handled with caution. Furthermore, if similar feedback is received from multiple users, the reliability of that feedback is judged to be high. Through this crowdsourcing-like approach, the system aggregates the knowledge of diverse users and is continuously improved.

[0111] (35th embodiment) Next, as a 35th embodiment of the present invention, the details of the specialization function for specific fields using domain adaptation technology will be described. Although this system is designed for general information verification, it may be possible to achieve higher accuracy by specializing in specific fields or domains (e.g., politics, medicine, finance, science and technology, entertainment, etc.). Each field has its own unique terminology, expression styles, information structures, and typical patterns of misinformation. In this embodiment, a function is provided to adapt a general-purpose model to a specific field using domain adaptation technology. Domain adaptation is a technique that adjusts a model trained in one domain (source domain) to another domain (target domain) using data from the target domain. In this system, first, a general-purpose model pre-trained on a large general-purpose dataset is constructed. Next, a dataset from a specific field (e.g., press conferences in the medical field, medical papers, medical-related SNS posts, etc.) is collected. This dataset is used to fine-tune the general-purpose model. When fine-tuning, there are two methods: updating all parameters of the model, and updating some parameters (e.g., only the final layer). The latter approach aims to prevent overfitting, maintain generality, and adapt to specific fields.

[0112] Furthermore, this embodiment also employs a method for explicitly incorporating domain-specific knowledge into the model. For example, in the medical field, structured knowledge such as medical terminology dictionaries, disease ontologs, and drug interaction databases are available. By integrating this knowledge into the model as a knowledge graph, it becomes possible to understand domain-specific contexts and relationships. Specifically, Knowledge Graph Embedding technology is used to embed entities (e.g., disease names, drug names) and relationships (e.g., "treat" "have side effects") into a vector space, and this is integrated with text and audio embeddings. In addition, in order to learn domain-specific misinformation patterns, examples of misinformation that have spread in that domain in the past are collected, and these are used as negative examples (examples of misinformation), while correct information is used as a positive example to train the classification model. For example, in the medical field, there are typical misinformation patterns such as the promotion of unfounded treatments such as "eating XX will cure cancer" and conspiracy theories such as "vaccines contain harmful substances." By learning these patterns, similar misinformation can be efficiently detected. This system manages multiple domain-specific models in parallel and provides a function that automatically selects the appropriate model according to the content area the user wants to verify. The content area is automatically determined using a text classification model. This domain adaptation function allows the system to achieve both versatility and specialization.

[0113] (36th embodiment) Next, as a 36th embodiment of the present invention, the details of the real-time collaborative filtering function will be described. Information spreads very quickly on social media, and viral content in particular can reach millions of people in a few hours to a few days. In this embodiment, a mechanism is provided to monitor content that is spreading in real time and for multiple users and organizations to collaboratively evaluate the reliability of that content. Specifically, when multiple users and organizations using this system verify the same content (for example, the same video URL), their verification results are shared and integrated. Since each user and organization may have a different database of positive information or use different analysis methods, each verification result provides information from a different perspective. This system aggregates these multiple verification results and calculates an overall reliability score. There are the following approaches to aggregation. First, there is the average or weighted average of the scores. The overall score is calculated by averaging the modification scores of each verification result. In the case of a weighted average, weights are set according to the reliability of each user and organization (past verification accuracy, expertise, etc.). Second, there is voting-based aggregation. Each verification result is categorized as "reliable," "doubtful," or "risky," and an overall judgment is made by majority vote or weighted voting. Thirdly, aggregation is performed using Bayesian estimation. Each verification result is used as observed data to estimate the true reliability of the content as a posterior probability. This method can also take into account the uncertainty of each verification result.

[0114] Furthermore, in this embodiment, the urgency is evaluated by combining the speed of content dissemination and verification results. Content that disseminates very quickly and has a low reliability score is designated as "high-urgency misinformation" and should be prioritized for action. When this system detects such high-urgency content, it automatically sends alerts to relevant parties (fact-checking organizations, news organizations, platform operators, and public institutions). The system also provides forums and chat functions within its user community to facilitate information sharing and discussion about the content. Users share their verification results, additional evidence, and expert opinions in the forums, and collaborate to uncover the truth. This collaborative approach is also called crowdsourced fact-checking, and by utilizing the knowledge of not only experts but also the general public, it enables faster and more multifaceted verification. However, collaborative filtering also carries the risk of groupthink and echo chambers (a phenomenon in which only people with the same opinion gather and differing opinions are excluded). This system introduces mechanisms to respect diverse opinions and appropriately consider minority opinions (for example, displaying opinion diversity indicators and making minority opinions visible).

[0115] (37th embodiment) Next, as a 37th embodiment of the present invention, the details of the educational and awareness-raising functions will be described. The purpose of this system is not merely to detect misinformation, but also to help users improve their media literacy and acquire the ability to critically evaluate information. In this embodiment, the following educational and awareness-raising functions are provided. First, a detailed explanation of the analysis results. This system not only presents a score, but also explains in easy-to-understand language why that score was obtained, what kind of alterations were detected, and how that affects the interpretation of the information. For example, it provides an explanation such as, "In this video, the context before and after the statement has been removed. In the original press conference, there was a preface saying, 'This is purely hypothetical,' before this statement. By removing this preface, it gives the impression that the statement is a definitive fact." Second, an information verification tutorial. This system provides an interactive tutorial that allows users to learn how to verify information themselves. The tutorial explains, using examples, how to find official sources, how to identify alterations in audio and video, how to check the context, and the importance of cross-checking multiple sources. Third, an introduction to typical patterns of misinformation. This system classifies and organizes past cases of misinformation that have spread, and introduces typical patterns (e.g., "fragrance," "emotional manipulation," "appeal to authority," "conspiracy theories," etc.). By learning these patterns, users will be able to judge for themselves whether new information they encounter is likely to be misinformation.

[0116] Fourth, it offers learning content in the form of quizzes and games. The system provides quizzes and games that allow users to learn media literacy while having fun. For example, there are quizzes that present multiple videos and ask users to identify which are genuine and which have been altered, and simulation games that teach users how to spot misinformation. By introducing gamification elements (points, badges, rankings), the system encourages users to continue learning. Fifth, it offers community-based learning. Within the system's user community, it provides a space where users can share knowledge and experiences and learn from each other. For example, there are forums where users can post their verification experiences and other users can provide comments and advice, and webinars and Q&A sessions by experts. Sixth, it provides educational materials for educational institutions. To support media literacy education in schools and universities, the system provides lesson plans for teachers, worksheets for students, and demo videos that can be used in class. This helps the next generation of citizens acquire the ability to critically evaluate information. Through these educational and awareness-raising functions, the system goes beyond being a mere technical tool and contributes to improving information literacy throughout society.

[0117] (38th embodiment) Next, as the 38th embodiment of the present invention, the details of the legal and ethical considerations function will be described. This system provides a function that has a significant social impact, which is to determine the truthfulness of information and, in some cases, to assign warning labels to specific content or accounts. Therefore, legal and ethical considerations are extremely important. In this embodiment, the following considerations are taken into account. First, respect for freedom of expression. This system does not have the authority to delete content or prohibit its transmission. It merely provides an evaluation of the reliability of the information and additional information, and the final decision is left to the user. The purpose of this system is not censorship, but to improve the transparency of information. Second, the risk of misjudgment is clearly indicated. The judgment of this system is based on AI and automated analysis and is not necessarily completely accurate. This system clearly indicates the reliability and uncertainty of all judgment results and warns the user that "this judgment is for reference only, and you should make your own final decision." Third, there is an objection mechanism. Creators of content incorrectly identified as "misinformation" by this system, and owners of accounts unfairly rated as having a high alert level, have the right to appeal the decision. This system will establish a contact point for appeals and conduct a re-evaluation by human experts. If the re-evaluation reveals that the decision was incorrect, it will be corrected promptly, and apologies and compensation will be provided as necessary. Fourthly, transparency and accountability are paramount. The system's judgment logic, the type of AI model used, the source of the training data, and the judgment criteria will be made public as much as possible to ensure transparency. Furthermore, the system's operator, responsible parties, and contact information will be clearly identified to clarify who is responsible in the event of a problem.

[0118] Fifth, bias monitoring and mitigation. AI models may learn biases (prejudices) contained in the training data. For example, there is a risk of making unfair judgments regarding specific political stances or specific races, genders, or religions. This system periodically audits the model's judgment results and evaluates whether or not bias is present. If bias is detected, measures such as reviewing the training data, retraining the model, and applying bias mitigation techniques (e.g., adversarial debiasing, reweighting) will be taken. Sixth, protection of privacy and reputation. This system takes care not to infringe on the privacy or reputation of individuals. For example, when analyzing statements or videos of ordinary citizens (individuals who are not public figures), consent will be obtained from the individual, or personally identifiable information will be anonymized. Furthermore, if the system's judgment results may damage the reputation of a particular individual or organization, careful review will be conducted, and legal advice will be sought as necessary. Seventh, compliance with international laws and regulations. This system complies with the laws of the countries and regions in which it operates (e.g., the EU's General Data Protection Regulation (GDPR), Section 230 of the US Communications Decency Act, national defamation laws, copyright laws, etc.). In particular, when transferring data across borders or operating in different jurisdictions, appropriate measures will be taken to consider the differences in legal regulations of each country. Through these legal and ethical considerations, this system will be operated in a socially responsible manner.

[0119] (39th embodiment) Next, as a 39th embodiment of the present invention, the details of the scalability and performance optimization functions will be described. This system requires high scalability and performance because it needs to process a large amount of SNS content in real time. In this embodiment, these requirements are met by the following technical innovations. First, there is a distributed processing architecture. This system adopts a microservices architecture and implements each function (SNS browsing, speech recognition, text analysis, speech feature extraction, difference analysis, score calculation, etc.) as an independent service. Since each service can scale out independently (increase the number of servers), resources can be allocated more effectively to services with high loads. Asynchronous communication using message queues (e.g., Apache Kafka, RabbitMQ) is used for communication between services to achieve loose coupling between services. Second, there is parallel processing and batch processing. When processing a large amount of content, throughput (amount of processing per unit time) is improved by processing multiple content in parallel. In this system, large-scale parallel processing is achieved using a distributed processing framework (e.g., Apache Spark, Apache Flink). Furthermore, batch processing that does not require real-time processing (for example, re-analysis of past content, periodic database maintenance) is executed at night or during periods of low load to efficiently utilize resources.

[0120] Thirdly, there is caching and result reuse. When the same content is requested for verification by multiple users, it is inefficient to repeat the same analysis each time. This system caches (temporarily stores) the analysis results, and for subsequent requests for the same content, the cached results are returned immediately. The cache expiration date is set according to the nature of the content. For example, official information is unlikely to change, so it can be cached for a long period, but SNS posts may be updated or deleted, so they are cached for a short period. Fourthly, there is model optimization and speed optimization. Deep learning models are highly accurate, but computationally expensive. This system uses model optimization techniques (e.g., knowledge distillation, pruning, quantization) to reduce the size and computational cost of the model without significantly compromising accuracy. In addition, inference speed optimization techniques (e.g., utilization of GPUs and TPUs, use of model optimization compilers (ONNX Runtime, TensorRT), batch inference) are used to improve processing speed. Fifthly, there is load balancing and autoscaling. This system uses a load balancer to distribute the load across multiple servers. Furthermore, the system utilizes the auto-scaling features of cloud platforms (e.g., AWS, Google Cloud, Azure) to automatically adjust the number of servers in response to increases or decreases in load. This allows it to handle sudden increases in access (for example, when a major news event occurs). Sixth, there is database optimization. Appropriate database technologies are selected to efficiently store and search large amounts of data. Relational databases (e.g., PostgreSQL, MySQL®), NoSQL databases (e.g., MongoDB, Cassandra), time-series databases (e.g., InfluxDB, TimescaleDB), and vector databases (e.g., Pinecone, Milvus) are used depending on the nature of the data. Search performance is also improved through appropriate index design, query optimization, and data partitioning. Through these scalability and performance optimizations, the system provides a fast and stable service even in large-scale operations.

[0121] (40th embodiment) Next, as a 40th embodiment of the present invention, the details of the function for responding to future technological advancements will be described. AI technology, information manipulation technology, and SNS platform specifications are evolving rapidly. This system needs to be designed to flexibly respond to these changes. In this embodiment, the following countermeasures are taken. First, modular design. Each function of this system is implemented as an independent module, and the interfaces between modules are clearly defined. This makes it easy to update a specific module (e.g., a speech recognition model) when replacing it with new technology, without affecting other modules. Second, a plug-in architecture. This system provides a mechanism that allows external developers to add new functions and analysis methods as plug-ins. Plug-ins interact with the system through a standardized API, enabling functional extensions without changing the core functions of the system. Third, model version control. The AI ​​models used in this system are continuously updated. This system manages the version of each model and provides a mechanism that allows multiple versions to be operated in parallel. This allows new models to be introduced in stages, their performance to be evaluated, and then applied to the production environment. Also, if a problem occurs, it is possible to quickly roll back to a previous version. Fourth, support for new data formats. In the future, new media formats that do not currently exist (e.g., 3D audio, holographic images, EEG interfaces) may emerge. This system has an extensible data pipeline to incorporate new data formats. New data formats will be handled by adding dedicated preprocessing modules. Fifth, standardization and open source. The data formats, API specifications, and analysis methods used by this system will be standardized as much as possible and released as open source. This will allow the entire community to share the technology and collaborate to evolve it. Sixth, continuous research and development. The development team of this system will continuously monitor the latest academic research, technological trends, and changes in threats and reflect them in the system. They will also conduct their own research and development and propose new methods and algorithms.These measures ensure that the system maintains its usefulness over the long term and adapts to the ever-evolving information environment.

[0122] (Effects and Benefits) The embodiments of the present invention described above provide the following remarkable effects. Firstly, it enables the automatic implementation of objective and multifaceted verification of clipped videos and modified content circulating on social media, based on official information. This significantly streamlines the process, which previously required time-consuming manual verification, allowing for the processing of large amounts of content at near real-time speed. Secondly, by comprehensively analyzing information from various modalities, including not only text comparisons but also acoustic characteristics, prosodic characteristics, emotional nuances, and visual characteristics of videos, it can detect sophisticated information manipulation that was difficult to detect with conventional methods. In particular, the ability to detect subtle alterations, such as "the text is the same, but the impression of the voice is different," is a unique strength of the present invention. Thirdly, by quantitatively evaluating the absence of context, it can clearly point out information distortion caused by selectively quoting only a portion of a statement. This allows users to recognize that the presented information is only a part of the whole picture and to make more careful judgments. Fourthly, by analyzing account behavior patterns and detecting organized information manipulation campaigns, it can address not only isolated instances of misinformation but also organized threats. Fifth, the real-time monitoring function allows for the prediction of problematic statements before they are extracted, enabling proactive measures. This represents a paradigm shift from the traditional reactive approach to a proactive, preventative one. Sixth, explainable AI technology clearly presents the rationale behind judgments, gaining user understanding and trust, and supporting appropriate decision-making. Seventh, privacy protection technologies (federated learning, differential privacy) allow for distributed model learning without centralizing data, minimizing privacy risks while building highly accurate models. Eighth, continuous learning and the use of user feedback enable the system to evolve as it operates, continuously adapting to new threats and methods. Ninth, multilingual support enables information verification across language barriers in a global information environment. Tenth, educational and awareness-raising functions contribute not only to providing technical solutions but also to improving information literacy throughout society. Through these effects, the present invention significantly contributes to improving the reliability and transparency of information and building a healthy information ecosystem.

[0123] (Specific example 1: Verification of clipped videos from a politician's press conference) As a specific embodiment of the present invention, we will describe a case of verifying a clipped video from a politician's press conference. A politician held a press conference regarding economic policy, and the entire conference was released as a video on the government's official website. In this press conference, the politician stated, "Considering the current economic situation, we are in a situation where we have no choice but to consider raising the consumption tax as one of the options." Before and after this statement, there was context such as, "However, this is merely one of several options, and the final decision will be made after fully listening to the voices of the people," and "At this point, we have not decided to raise the tax, and we are in the stage of carefully proceeding with discussions." However, on social media, a video clipped only of the part that said, "We are in a situation where we have no choice but to consider raising the consumption tax as one of the options," was spread and posted with a misleading title such as "Government decides to raise consumption tax." The system detected this clipped video by its social media monitoring unit and began analysis. First, the audio was extracted from the video and transcribed by the speech recognition unit. The transcription result was the text, "Considering the current economic situation, we are in a situation where we have no choice but to consider raising the consumption tax as one of the options." Next, this text was compared with the transcript data of the official press conference stored in the official information database. The text difference analysis unit identified, through substring search, that this text corresponds to a specific point in the official press conference (23 minutes and 45 seconds from the start).

[0124] The contextual difference analysis unit retrieved the context before and after the identified section from the positive information database and detected that important preconditions and reservations had been removed from the clipped video. Specifically, it was found that important modifying phrases such as "this is merely one of several options," "the final decision will be made after fully listening to the voices of the people," and "we have not yet decided to raise the price" had been removed. The contextual loss score was calculated to be 0.85 (high risk). Next, the audio difference analysis unit compared the audio of the clipped video with the corresponding section of the official press conference. By comparing audio features using the DTW algorithm, it was confirmed that the audio of both showed a high degree of similarity and that the audio itself had not been altered. However, in the official press conference, there was a pause of about 2 seconds (appearing to be thinking) before the problematic statement, and after the statement, there were questions from reporters and supplementary explanations in response, but all of these were removed from the clipped video. It was analyzed that the removal of the pause made the statement give a more definitive impression. The emotion recognition unit estimated the speaker's emotion in the relevant section of the official press conference as "cautious, somewhat anxious," and similar emotions were detected in the clipped video. However, it was evaluated that the impression received by the viewer differed significantly due to the removal of the surrounding context. The score calculation unit integrated these analysis results and calculated an overall modification score of 0.78 (high risk). By category, the text modification score was low (audio and text matched), but the contextual loss score was very high, which boosted the overall score. The visualization / UI unit presented these results to the user. In the comparison view, the relevant section of the official press conference (including the surrounding context) was displayed on the left, and the clipped video was displayed on the right, with the deleted portion highlighted in red. A warning message was also displayed stating, "This clipped video has removed important assumptions and reservations about the statement, and the intent of the statement may be distorted."

[0125] (Specific Example 2: Detection of deepfake audio in interviews with corporate public relations representatives) As a concrete example, we will describe the detection of deepfake audio in an interview with a corporate public relations representative. A public relations representative of a major company was interviewed regarding a new product announcement, and the audio was officially released. A few days later, an audio recording of the same public relations representative's voice saying, "Our new product has a serious defect, and we are considering a recall," spread on social media. This audio contained content that the public relations representative had not actually spoken, and was a deepfake created using AI speech synthesis technology. This system detected this audio and began analysis. First, the speech recognition unit transcribed the audio and obtained the text, "Our new product has a serious defect, and we are considering a recall." This text was compared with the official interview transcript in the official information database, but no matches were found. This suggests that this statement does not officially exist. Next, the speaker authentication unit extracted a speaker embedding vector from the suspicious audio and compared it with the public relations representative's speaker model stored in the official information database. The speaker embedding vector showed a high degree of similarity, and the voice quality matched that of the public relations representative. However, the deepfake voice detection unit analyzed the detailed acoustic characteristics of the speech and detected the following anomalies: Firstly, the pitch variation pattern was unnaturally smooth, lacking the subtle fluctuations found in natural human speech. Secondly, the energy distribution in certain frequency bands (especially the high-frequency band) was abnormal compared to genuine human speech. Thirdly, unnatural discontinuities were detected in the transitions between phonemes. These features were typical traces of AI speech synthesis. The deepfake detection model outputted a deepfake score of 0.92 (highly likely to be synthesized speech) for this speech.

[0126] Furthermore, the system estimated the type of speech synthesis technology used. Based on the acoustic feature patterns, it was determined that the audio was highly likely to have been generated using a GAN-based speech synthesis model (specifically, HiFi-GAN or a similar model). The scoring unit integrated these analysis results and calculated an overall modification score of 0.95 (extremely high risk). The audio authenticity score was extremely low (0.05), indicating that there was almost no possibility that the audio was genuine. The visualization / UI unit presented these results to the user. A clear warning message was displayed stating, "This audio is highly likely to be a deepfake audio synthesized by AI. No evidence was found that our public relations representative actually made this statement." Details of the detected acoustic anomalies (smoothness of pitch fluctuations, anomalies in frequency distribution, discontinuities in phoneme transitions) were also presented along with graphs. In response to these verification results, the company promptly issued an official statement clarifying that "the audio in question was not spoken by our public relations representative, but is a forged audio created by a malicious third party." They also stated that they would consider legal action. The rapid detection capabilities of this system allowed us to quickly stop the spread of misinformation and minimize damage to the company's reputation.

[0127] (Specific Example 3: Detection of Translation Distortion in a Multilingual Environment) As a third specific example, we will describe an instance of detecting translation distortion in a multilingual environment. At an international conference, a high-ranking Japanese government official delivered a speech in Japanese, and its contents were officially released as a Japanese transcript. In the speech, the official stated, "While valuing international cooperation, our country reserves the right to take any measures necessary to protect its own interests." This speech was later translated into English and disseminated on social media, but the translation was "Our country will prioritize national interests over international cooperation and reserves the right to take any measures necessary." This English translation distorted the intent of the original Japanese, changing the important premise of "while valuing international cooperation" into the confrontational expression of "prioritizing national interests over international cooperation." This system detected this English translation and began analysis. First, the text analysis unit vectorized the English text using a multilingual language model (mBERT). At the same time, the official Japanese transcript stored in the positive information database was also vectorized using the same model. mBERT can map texts from different languages ​​into a common semantic space, enabling the calculation of semantic similarity across languages. The cosine similarity of two vectors was calculated to be 0.65, indicating a degree of relevance, although not a perfect match. However, this similarity was lower than what would be expected for an accurate translation (typically 0.8 or higher).

[0128] Next, the system automatically translated the official Japanese transcript into English using reliable machine translation services (Google Translate and DeepL). The Google Translate translation was "Our country, while valuing international cooperation, reserves the right to take necessary measures to protect our own interests," and the DeepL translation was "While emphasizing international cooperation, our country reserves the right to take measures necessary to protect national interests." When these automatic translation results were compared with the English translation circulating on social media, a clear difference was detected. In particular, the concessive expression "while valuing / emphasizing international cooperation" was replaced in the social media version with the confrontational expression "prioritize national interests over international cooperation." The semantic similarity analysis unit detected that this change resulted in a significant transformation of meaning. Specifically, the original Japanese text expressed a stance of striving to balance "international cooperation" and "national interest," while the SNS version of the English translation portrayed the two as opposing forces, prioritizing national interest. This could lead to significantly different perceptions in the international community. The score calculation unit calculated the translation modification score to be 0.72 (high risk). The visualization / UI unit presented this result to the user. In the comparison view, the official Japanese transcript was displayed at the top, the result of a reliable machine translation in the middle, and the English translation circulating on SNS at the bottom, with differences highlighted in red. A warning message was also displayed stating, "This English translation may not accurately reflect the intent of the original Japanese. In particular, the important premise of 'while valuing international cooperation' has been removed and replaced with an opposing expression."These verification results were shared with international relations experts and the media, reinforcing the importance of accurate translation.

[0129] (Specific Example 4: Predicting the Risk of Clipping Through Real-Time Monitoring of Press Conferences) As a fourth specific example, we will explain an example of predicting the risk of selective editing through real-time monitoring of press conferences. A government agency held a live-streamed press conference regarding an important policy announcement. This system connected to a pre-registered live-stream URL and acquired and analyzed the audio in real time. During the press conference, the minister in charge made the following statement: "Under ideal circumstances, this policy will be highly effective. However, in reality, there are various constraints, and there is a good chance that the expected results will not be achieved. We are preparing not only with optimistic outlooks in mind, but also with the worst-case scenario in mind." The real-time monitoring function of this system transcribed this statement sequentially and input it into a model that predicts the likelihood of selective editing. The prediction model detected the following characteristics: Firstly, the statement contained both a positive part, "This policy will be highly effective," and a negative part, "There is a good chance that the expected results will not be achieved." Such statements that present both sides of an argument are easily taken out of context. Secondly, the statement included conditional clauses and adversative clauses such as "Under ideal circumstances" and "However, in reality," and the meaning changes significantly if these are removed. Thirdly, prosodic analysis of the audio revealed that the phrase "it will have a significant effect" was spoken in a relatively bright tone, while the phrase "there is a possibility that the expected results will not be achieved" was spoken in a lower tone. Such changes in voice are also factors that make the impression easily manipulated through selective editing.

[0130] The predictive model comprehensively evaluated these features and calculated the likelihood of the statement being taken out of context to be 0.88 (very high risk). The system immediately sent an alert to the public relations officer's dashboard. The alert included a warning that "a high-risk statement has been detected. There is a possibility that only the part 'this policy will be very effective' may be taken out of context, with conditions and reservations removed," along with a transcript of the statement, the time it was made (15 minutes and 32 seconds from the start of the press conference), and details of the risk factors. In response to this alert, the public relations officer took the following swift actions: Firstly, during the Q&A session after the press conference, when a reporter asked a related question, the officer urged the minister to clarify the intent of his statement. The minister added, "As I mentioned earlier, the effectiveness of this policy depends on various conditions. I would like you to understand the whole context so that it is not taken out of context and interpreted optimistically or pessimistically." Secondly, immediately after the press conference ended, the full transcript of the press conference was published on the official website with annotations added to the relevant section. Thirdly, the official social media account posted a tweet stating, "For the full text of the remarks regarding the policy's effectiveness at today's press conference, please see here," encouraging access to accurate information. Fourthly, the system's social media monitoring function was enhanced to immediately detect when clipped videos containing the relevant remarks began to be posted. Indeed, as predicted, a few hours after the press conference, videos containing only the part stating, "This policy will have a significant effect," began to be posted on social media. The system immediately detected these and assigned a high modification score. The public relations team responded to these posts by posting replies from the official account stating, "This video only includes a portion of the remarks, and important conditions and reservations have been omitted. Please see here for the full text," thus curbing the spread of misinformation. This real-time monitoring and swift response successfully prevented the spread of misinformation at an early stage.

[0131] (Industrial applicability) This invention has broad applicability in the following industrial fields. First, information verification in news organizations. News organizations such as newspapers, television stations, and online media can use this system to quickly verify the veracity of information spreading on social media and to provide accurate reporting. In particular, the real-time analysis function of this system is of great value in news reporting where speed is required. Second, public relations and crisis management in government and public institutions. Governments, local governments, and public institutions can use this system to monitor whether the official information they disseminate is being distorted and spread, and to quickly issue corrective information as needed. Furthermore, in situations where misinformation has a significant impact on society, such as during election periods or public health emergencies, this system supports the early detection and response to misinformation. Third, reputation management and brand protection in companies. Companies can use this system to quickly detect misinformation and malicious rumors about themselves and their products and take appropriate action. In particular, for listed companies, rapid verification and response are important because misinformation can affect stock prices. Fourth, content moderation for social media platform operators. Social media platform operators such as Twitter / X, Facebook, YouTube®, and TikTok can use this system to automatically detect misinformation and harmful content spreading on their platforms and take action such as displaying warning labels, limiting reach, and removing the content. Fifthly, it improves efficiency for fact-checking organizations. Organizations that perform professional fact-checking can use this system to automate the prioritization of content to be verified, preliminary analysis, and evidence collection, allowing human experts to focus on more complex judgments.

[0132] Sixth, media literacy education in educational institutions. Schools, universities, and lifelong learning institutions can use this system as teaching material to educate students and citizens on how to critically evaluate information, how to identify misinformation, and the importance of media literacy. Seventh, misinformation research in research institutions. Universities and research institutions can use the data collected and analyzed by this system to study the mechanisms of misinformation dissemination, social impact, and countermeasures. This system provides researchers with anonymized datasets and API access. Eighth, criminal investigations in law enforcement. Law enforcement agencies such as the police and prosecutors can use this system to detect forged audio and video used in crimes such as fraud, defamation, and election interference, and use them as evidence. Ninth, information warfare countermeasures in international organizations. International organizations such as the United Nations and NATO can use this system to detect and analyze misinformation campaigns in information warfare, propaganda, and hybrid warfare between nations. Tenth, everyday information verification by ordinary citizens. This system will be made publicly available as a web service and smartphone application, allowing anyone to verify information they are interested in, thereby contributing to the improvement of information literacy throughout society. In these diverse industrial fields, the present invention will play an important role in improving the reliability of information, suppressing misinformation, and ensuring transparency, contributing to the realization of a healthy information society.

[0133] (summary) The embodiments of the present invention have been described in detail above. The present invention is a groundbreaking system that enables multifaceted and automated verification of clipped videos and modified content circulating on social media, based on official information. By integrally utilizing a variety of AI technologies such as speech recognition, speech feature analysis, emotion recognition, text analysis, and video analysis, it can detect sophisticated information manipulation that was difficult to detect with conventional methods. Furthermore, it is equipped with advanced functions such as real-time monitoring, account evaluation, deepfake detection, multilingual support, privacy protection, explainability, and continuous learning, and is adaptable to the ever-evolving information environment. The present invention can be used by a wide range of stakeholders, including news organizations, government agencies, corporations, social media platforms, educational institutions, research institutions, and the general public, and will greatly contribute to improving the reliability and transparency of information and building a healthy information ecosystem. It should be noted that the present invention is not limited to the embodiments described above, and various modifications and applications are possible without departing from the spirit of the invention. For example, it is possible to implement the technical elements described in each embodiment in different combinations. In addition, the basic principles of the present invention can be applied not only to audio and video content, but also to text-only content, image content, and new media formats that will emerge in the future. The technical scope of the present invention should be determined based on the claims, and the invention described therein and its equivalents are included within the technical scope of the present invention.

[0134] (41st Embodiment) Next, as a 41st embodiment of the present invention, the details of the temporal consistency verification function will be described. In video content, the temporal synchronization of video and audio is an important indicator of authenticity. In this embodiment, traces of editing or modification are detected by verifying the temporal consistency of video frames and audio waveforms in detail. Specifically, lip-sync verification is performed to analyze the synchronization of the speaker's mouth movements and the audio. When a person speaks naturally, there is a strict temporal correspondence between the mouth movements and the sound produced. For example, when pronouncing the sound "ma," the lips are first closed and then opened. In this system, the speaker's face region is detected from the video, and facial landmarks (especially feature points around the mouth) are tracked. For face detection, Haar Cascade, HOG (Histogram of Oriented Gradients), or deep learning-based detectors (e.g., MTCNN, RetinaFace) are used. For landmark tracking, libraries such as Dlib, OpenFace, or MediaPipe are used. Features such as the degree of mouth opening and closing, lip shape, and jaw position are extracted in a time series and mapped to the phoneme sequence of the speech. From the speech, a phoneme recognition model (e.g., an HMM-based or DNN-based phoneme recognizer) is used to estimate the phonemes at each time point. The correlation between the mouth movement features obtained from the video and the phoneme sequence obtained from the speech is calculated, and if the two show a high correlation, it is determined that the video and audio are synchronized. Conversely, if the correlation is low, or if a time lag is detected, it indicates the possibility that the video and audio were recorded separately, or that it is a deepfake video.

[0135] Furthermore, this embodiment also verifies the correlation between audio energy fluctuations and video motion. When a person speaks, when the audio energy (volume) is high, the mouth movements are also large, and often accompanied by subtle movements of the head and body. This system extracts the audio energy envelope and compares it with the magnitude of the video's optical flow (pixel movement between frames). If there is a high correlation between the two, it is determined to be a natural video. If the correlation is low, it is possible that the audio was added later (dubbed) or the video was synthesized. This system also verifies the consistency between ambient sound and video. For example, it detects inconsistencies such as when the video shows an outdoor landscape but the audio has reverberation characteristic of an indoor setting. The characteristics of ambient sound are estimated using Acoustic Scene Classification (ASC) technology. The type of scene (indoor / outdoor, conference room / street, etc.) is also estimated from the visual features of the video, and the degree of agreement between the two is evaluated. Through these temporal consistency verifications, this system can detect complex editing and synthesis that are difficult to detect with a single modality.

[0136] (42nd embodiment) Next, as a 42nd embodiment of the present invention, the details of the metadata analysis function will be described. Digital content is usually accompanied by metadata (data about the data). In the case of video files, information such as the shooting date and time, shooting equipment, editing software, GPS location information, file creation date and time, and modification date and time are embedded in the form of EXIF ​​(Exchangeable Image File Format) or XMP (Extensible Metadata Platform). In this embodiment, the authenticity of the content and traces of modification are verified by extracting and analyzing this metadata. First, the existence and completeness of the metadata are confirmed. Officially shot and published content usually contains detailed metadata. On the other hand, if the metadata is completely deleted or unusually small, there is a possibility of intentional concealment. Next, the consistency of the metadata content is verified. For example, if the file creation date and time are earlier than the date and time the content is supposedly shot, it is a contradiction. Also, if the shooting equipment information differs from the equipment officially used, it is suspicious. Furthermore, it is possible to estimate what kind of editing was done from the editing software information. For example, if professional editing software such as Adobe Premiere Pro or Final Cut Pro is used, it's highly likely that a certain level of editing has been done. On the other hand, if simpler editing software or smartphone apps are used, the editing may be limited to minor changes.

[0137] Furthermore, this embodiment also performs metadata tampering detection. Metadata itself can be tampered with using specialized tools. This system verifies the internal integrity of the metadata (e.g., correlations between different metadata fields) and detects traces of tampering. In addition, it calculates the hash value of the file (MD5, SHA-256, etc.) and compares it with the hash value of officially published content to determine whether the file is completely identical or has been modified even slightly. If the hash values ​​match, it is guaranteed that the file has not been altered at all. If the hash values ​​are different, some kind of change has been made. Furthermore, this system also considers guaranteeing the reliability of metadata using blockchain technology. When an official institution publishes content, it records the metadata and hash value on the blockchain to create proof that it cannot be tampered with later. Users can verify that the content is officially published and has not been tampered with by comparing the metadata and hash value of any content with the record on the blockchain. This metadata analysis function, when combined with audio and video content analysis, achieves a stronger authenticity verification.

[0138] (43rd embodiment) Next, as a 43rd embodiment of the present invention, the details of the social graph analysis function will be described. The spread of misinformation and altered content largely depends on the relationships (social graph) between users on social networking services (SNS). In this embodiment, organized information manipulation campaigns and influential accounts are identified by tracking the content's dissemination path and analyzing the social graph. Specifically, all accounts that posted or shared specific content (for example, the URL of a specific clipped video) are collected, and follow relationships, retweet relationships, and mention relationships between these accounts are obtained. This allows the content dissemination network to be represented as a graph structure. Each node in the graph represents an account, and each edge represents a relationship (follow, retweet, etc.). Network analysis methods are applied to this graph. First, the centrality index is calculated. Degree centrality represents the number of edges connected to a given node and identifies hub-like accounts that are connected to many other accounts. Betweenness centrality represents the frequency of being located on the shortest path within the network and identifies accounts that act as bridges for information. Eigenvector centrality highly values ​​nodes connected to influential nodes, measuring their overall influence within the network. Secondly, there is community detection. This divides the graph into closely connected subgroups (communities). Community detection algorithms include Louvain's method, Girvan-Newman's method, and Label Propagation. Detected communities may be groups of accounts with similar ideologies or objectives. In particular, if multiple communities are coordinating to spread the same content, it may indicate a coordinated campaign.

[0139] Thirdly, there is time-series analysis. This involves tracking the spread of content over time, analyzing which account posted first (the starting point), how it spread, and the speed of spread. If multiple accounts start posting simultaneously in the early stages of spread, it may indicate coordinated inauthentic behavior. Fourthly, there is bot detection. Many automated accounts (bots) exist on social media and are sometimes used to spread misinformation. This system analyzes the behavior patterns of accounts (posting frequency, posting time, diversity of posted content, follower-to-following ratio, etc.) to estimate the likelihood of them being bots. Existing tools such as Botometer (formerly BotOrNot) are used as reference models for bot detection. If a group of detected bot accounts are spreading the same content, it serves as evidence of automated information manipulation. Through these social graph analyses, this system reveals not only the authenticity of individual content but also the structure of the organized information manipulation behind it.

[0140] (44th embodiment) Next, as a 44th embodiment of the present invention, the details of the emotion manipulation detection function will be described. One of the purposes of information manipulation is to manipulate the emotions of the recipient and induce specific actions (for example, to incite anger to encourage aggressive posts, or to incite fear to get people to buy specific products). In this embodiment, we detect whether the content is intentionally trying to manipulate emotions. First, we estimate emotions from the text, audio, and video of the content. From the text, we use an emotion analysis model (e.g., BERT, RoBERTa-based emotion classifier) ​​to estimate basic emotions such as joy, sadness, anger, fear, disgust, and surprise, and their intensity. From the audio, we use the aforementioned speech emotion recognition technology. From the video, we analyze the speaker's facial expressions and estimate emotions using a facial expression recognition model (e.g., FER, DeepFace). Next, we compare this emotion information with official information. If strong emotions (especially negative emotions such as anger and fear) are emphasized through cropping or editing, even though the official information showed neutral or mild emotions, it indicates the possibility of emotion manipulation.

[0141] Furthermore, this embodiment detects typical techniques of emotion manipulation. Firstly, the addition of emotional words. It detects cases where emotionally charged adjectives or adverbs (e.g., "shocking," "terrifying," "unforgivable") that were not included in the original statement have been added to subtitles or titles. Secondly, the addition of music and sound effects. It detects cases where tension-building music or sound effects emphasizing shock, which were not included in the original video, have been added. Audio analysis is used to detect the presence of background music and sound effects and compare them with the official video. Thirdly, the adjustment of the video's color tone and contrast. It detects operations that create a gloomy impression by darkening the video or lowering the saturation, or conversely, create a flashy impression by increasing the saturation. Image processing techniques are used to analyze the video's color histogram, brightness, and saturation and compare them with the official video. Fourthly, close-ups of faces. It detects techniques that emphasize facial expressions and amplify emotional impressions by unnaturally zooming in on the speaker's face. The framing (composition) of the video is analyzed and compared with the official video. By detecting these emotional manipulations, this system reveals information manipulation tactics that appeal to the recipient's emotions, thereby hindering rational judgment.

[0142] (45th embodiment) Next, as a 45th embodiment of the present invention, the details of the context completion support function will be described. This system not only detects alterations but also completes missing context to help users understand the overall picture. Specifically, it automatically extracts the deleted portion from the official information in the clipped video and presents it to the user. For example, if the clipped video uses a portion of an official press conference (e.g., from 23 minutes 45 seconds to 24 minutes 10 seconds), the system automatically retrieves the portions before and after that (e.g., from 23 minutes 00 seconds to 23 minutes 45 seconds, and from 24 minutes 10 seconds to 25 minutes 00 seconds) and displays them as "context before and after the deletion". The user can watch the complete video, including the deleted portion, with a single click. The system also automatically generates a summary of the deleted portion. For summary generation, extractive summarization (extracting important sentences) or generative summarization (generating new sentences that summarize the content) techniques are used. For generative summarization, pre-trained language models such as BART, T5, and GPT are used. The summary should concisely convey the key points of the context, such as, "The deleted portion clearly indicated that this statement was a 'hypothetical' one," or "The deleted portion explained that parliamentary approval was required for the implementation of this policy."

[0143] Furthermore, this embodiment also provides links to relevant official information. For example, it automatically collects and presents links to full transcripts of press conferences, relevant government announcements, similar past statements, and expert commentary articles. This allows users to comprehensively refer to related information, not just a single clipped video, and gain a deeper understanding. The system also provides interactive features that allow users to explore the context themselves. For example, when a user clicks on a point of interest on the clipped video's timeline, the corresponding portion of the official video is displayed, allowing the user to freely view the surrounding context. Through this context-enhancing support function, the system not only warns that "this video has been altered," but also provides constructive information such as "the correct information is here."

[0144] (46th embodiment) Next, as a 46th embodiment of the present invention, the details of the cross-platform tracking function will be described. Misinformation and altered content often spread not only across a single SNS platform but also across multiple platforms (such as Twitter / X, Facebook, YouTube®, TikTok, Instagram, LINE, and Telegram). In this embodiment, the spread of the same content across different platforms is tracked. First, a content fingerprint (such as the aforementioned audio fingerprint, image hash, and video hash) is calculated for the content collected from each platform. The spread of the same content is tracked by matching content with the same or similar fingerprints across different platforms. Next, the spread status on each platform (number of posts, views, shares, comments, spread speed, etc.) is aggregated to evaluate the overall impact across platforms. For example, it is possible to identify trends such as a certain content spreading relatively little on Twitter but exploding on TikTok. The spread paths between platforms are also analyzed. For example, the path is tracked if the content was first posted on Twitter, then reposted on YouTube®, and then spread on Facebook. This type of path analysis allows us to understand the strategy of an information manipulation campaign (which platform it starts from and in what order it spreads).

[0145] Furthermore, this embodiment also performs analysis tailored to the characteristics of each platform. For example, short videos are prevalent on TikTok, and clips tend to be edited to be even shorter. On Instagram, visually appealing thumbnails and captions are highly valued. On Telegram, dissemination is often confined to closed groups. This system performs analysis that takes these platform-specific characteristics into account. It also refers to each platform's policies (content moderation rules, deletion criteria, etc.) and tracks situations where certain content is deleted on one platform but remains on another. Through this cross-platform tracking, this system grasps the dynamics of misinformation across the entire information ecosystem.

[0146] (Embodiment 47) Next, as the 47th embodiment of the present invention, the details of the predictive modeling function will be described. This system predicts the future spread of misinformation by learning from past data. Specifically, it makes the following predictions: First, it predicts what kind of misinformation is likely to occur in the future regarding a particular topic or event. For example, when an election is approaching, misinformation about candidates tends to increase. This system learns misinformation patterns from past election periods and predicts the possibility of similar misinformation occurring in the current election. Second, it predicts how far a particular piece of content will spread in the future. A spread prediction model is constructed using features such as the initial spread of the content (number of posts, views, and shares in the first few hours), the influence of the posting account, and the emotional appeal of the content. The model uses time-series forecasting methods (e.g., ARIMA, LSTM, Prophet) to predict the future spread curve. Third, it predicts the likelihood that a particular account will post misinformation in the future. A risk prediction model is constructed using features such as the account's past posting history, the average modification score, and its location on the network. Accounts predicted to be high risk are monitored intensively.

[0147] These prediction results will be used for proactive warning and preventative measures. For example, before important events (elections, international conferences, product launches, etc.), predicted patterns of misinformation will be shared with stakeholders, and countermeasures (e.g., proactive information dissemination, strengthened monitoring systems) will be taken in advance. In addition, content predicted to spread will be verified early, and if it is found to be misinformation, a warning will be issued before widespread dissemination occurs. Through this predictive approach, this system will realize a paradigm shift from reactive response to proactive prevention.

[0148] (Embodiment 48) Next, as the 48th embodiment of the present invention, the details of the voice cloning countermeasure function will be described. In recent years, voice cloning technology (a technology that generates voices that mimic the voice of a specific person from a small amount of sample voices) has been rapidly evolving. This embodiment provides an advanced technology for detecting forged voices generated by voice cloning. The key to voice cloning detection is to capture the subtle differences between real human voices and AI-generated voices. This system analyzes the following characteristics: Firstly, subtle fluctuations in voice. Human voices contain unpredictable subtle fluctuations caused by breathing, vocal cord vibrations, and subtle movements of the vocal organs. AI-generated voices cannot perfectly reproduce these fluctuations and may be unnaturally smooth or, conversely, contain unnatural noise. This system uses higher-order statistics (e.g., higher-order moments, higher-order cumulative quantities) and nonlinear analysis methods (e.g., Lyapunov exponents, fractal dimensions) to analyze the fine structure of voice. Secondly, phase information of voice. Conventional voice analysis has focused on amplitude spectra, but phase spectra also contain important information. AI-generated speech often has incomplete phase reproduction. This system extracts features from the phase spectrum and compares them to real speech. Thirdly, there are the long-term dependencies of speech. Human speech has long-term prosodic patterns (intonation, rhythm, and pause placement) that depend on context and intent. AI-generated speech may sound natural in the short term, but it may lack long-term consistency. This system analyzes prosodic features on long time scales.

[0149] Furthermore, in this embodiment, the detection model is enhanced using adversarial training. Speech cloning technology is constantly evolving, and sophisticated forged speech may emerge that cannot be detected by conventional detection methods. In this system, forged speech is generated using the latest speech cloning model and added to the training data of the detection model, thereby continuously strengthening the model. This is similar to the relationship between the red team (attackers) and the blue team (defenders) in the field of security. In addition, this system also provides proof verification of the source of the speech. When official organizations publish speech, they can embed digital signatures or watermarks to prove that the speech is officially published. Watermarking is a technology that embeds signals into speech that are inaudible to the human ear but can be detected by specialized software. This system provides the function to detect watermarks from speech and verify their authenticity.

[0150] (49th embodiment) Next, as the 49th embodiment of the present invention, the details of the context-aware analysis function will be described. Even the same statement or video can have vastly different meanings and importance depending on the situation (context) in which it was made. In this embodiment, the situation in which the content was made is understood and reflected in the analysis. Specifically, the following situational information is collected and analyzed. Firstly, there is the temporal context. The date and time the content was made, the social and political situation at that time, and related events are considered. For example, the intent and impact of a statement made by a politician will differ depending on whether it was made immediately before an election or during peacetime. This system refers to news databases and event calendars to grasp the temporal context of the content. Secondly, there is the spatial context. The location (country, city, specific venue, etc.) in which the content was made is identified, and the characteristics of that location are considered. For example, the formality of a statement made at an official press conference differs from that made in an informal setting. This system identifies the location through background analysis of the video, GPS information, and extraction of place names in the text. Thirdly, there is the social context. The system aims to understand the social atmosphere, public opinion trends, and the state of related discussions when content is disseminated. This system estimates the social context through topic modeling, sentiment analysis, and trend analysis of discussions on social media. Fourthly, it considers the context of the content creator. This includes the creator's position, role, past statements, and credibility. For example, the statements of an official government spokesperson carry different weight than those of an ordinary citizen. This system analyzes the creator's profile information, past posting history, and network location.

[0151] By integrating this contextual information into content analysis, more accurate and context-sensitive judgments can be made. For example, it can determine that a statement that is normally acceptable may be misleading in a specific tense situation. Considering context also reduces false positives. For instance, it prevents misinterpreting sarcastic or humorous remarks literally and classifying them as misinformation. This system estimates the intent of a statement (serious, joking, sarcastic, etc.) from information such as the tone of the statement, the speaker's facial expressions, and the audience's reactions (laughter, etc.).

[0152] (50th embodiment) Next, as the 50th embodiment of the present invention, the details of the long-term impact assessment function will be described. The impact of misinformation and altered content is important not only in terms of short-term dissemination but also in terms of long-term effects. In this embodiment, the long-term impact of content on society is evaluated. Specifically, the following indicators are tracked. First, changes in beliefs. The extent to which the beliefs and attitudes of people who have been exposed to misinformation have changed is evaluated. In this system, changes in the content of posts on social media, survey data, public opinion poll data, etc. are analyzed to estimate the impact of misinformation. Second, changes in behavior. The extent to which misinformation has affected people's specific behaviors (e.g., voting behavior, purchasing behavior, health behavior). For example, an analysis is conducted to determine whether misinformation about medical care led to a decrease in vaccination rates. Third, deepening of social division. The extent to which misinformation has deepened social division (political polarization, intergroup conflict) is evaluated. In this system, the degree of division is measured by network analysis of discussions on social media, emotional polarity analysis, etc. Fourth, decline in trust. This study assesses the extent to which the spread of misinformation has damaged trust in the media, governments, experts, and science. Trust levels will be measured using methods such as surveys and social media mentions.

[0153] The results of these long-term impact assessments will be used to prioritize measures against misinformation. Misinformation that spreads on a small scale in the short term but has serious long-term effects should be prioritized for action. Furthermore, analyzing the long-term impacts of past misinformation will improve future strategies for combating it. For example, it will help accumulate knowledge about what types of misinformation are most harmful and what countermeasures are most effective.

[0154] (51st Embodiment) Next, as the 51st embodiment of the present invention, the details of the multi-stakeholder collaboration function will be described. Misinformation countermeasures cannot be solved by a single organization or technology alone; collaboration among diverse stakeholders (government, media, platform operators, fact-checking organizations, research institutions, and civil society) is necessary. This embodiment provides a function to support information sharing and collaboration among these stakeholders. Specifically, the system provides a platform for sharing information verified by each stakeholder. Each organization registers information on the content it has verified (content URL, verification results, reasons for judgment, and evidence) in the system in a standardized format. Other organizations can refer to this information and use it in their own verification work. This eliminates the waste of multiple organizations duplicating the same content and improves efficiency. Furthermore, the system supports the division of roles that leverages the expertise of each organization. For example, it coordinates the division of labor so that medical misinformation is verified by medical experts, and legal misinformation is verified by legal experts. The system provides a function to automatically classify the content and request verification from appropriate experts. In addition, the system also supports the building of trust among stakeholders. The accuracy of each organization's verification results is evaluated against each other, and a confidence score is calculated. Verification results from organizations with higher confidence scores are given greater weight. In addition, to ensure transparency, this system also provides a function to disclose information such as each organization's verification process, the methodology used, and whether or not there are any conflicts of interest.

[0155] (52nd embodiment) Next, as a 52nd embodiment of the present invention, the details of the adaptive learning function will be described. Information manipulation techniques are constantly evolving, and this system also needs to continuously learn and adapt. In this embodiment, a mechanism is provided that allows the system to automatically learn and improve its performance while in operation. Specifically, the following adaptive learning is performed. First, there is online learning. The model is updated sequentially each time new data arrives. This allows for rapid response to the latest information manipulation techniques. Stochastic gradient descent (SGD) and methods with adaptive learning rates (Adam, AdaGrad) are used as algorithms for online learning. Second, there is active learning. Content that the system is not confident in judging (for example, those with judgment scores near a threshold) is selectively presented to human experts, who teach the correct labels. This high-quality labeled data is used to efficiently improve the model. Third, there is transfer learning. A model learned in one field or task is applied to another field or task. For example, a model trained on detecting misinformation in English can be transferred to detecting misinformation in Japanese. Transfer learning allows for the construction of highly accurate models even in fields with limited data. Fourthly, there is meta-learning. In the process of learning multiple tasks, the model learns the "how to learn" itself. This allows for the construction of models that can quickly adapt to new tasks with small amounts of data.

[0156] Furthermore, this embodiment also provides a mechanism to continuously monitor the model's performance and automatically retrain it if a performance degradation is detected. Performance monitoring uses indicators such as accuracy on the validation dataset, user feedback, and false positive / missed positive rates. Retraining is triggered if performance falls below a threshold or if a new type of misinformation increases. Retraining can be done by either retraining the model from scratch using the latest data or by fine-tuning the existing model.

[0157] (Embodiment 53) Next, as a 53rd embodiment of the present invention, the details of the function that takes diversity and inclusion into consideration will be described. This system is intended to be used by people with diverse cultures, languages, and values, and it is necessary to take care not to be biased towards any particular culture or value. In this embodiment, the following considerations are taken into account. First, support for diverse languages. In addition to the multilingual support function described above, it also supports minority languages ​​and dialects. Furthermore, analysis is performed that takes into account the cultural nuances and differences in expression in each language. Second, understanding of cultural context. Even the same expression can have different meanings and be received differently depending on the culture. For example, a joke that is acceptable in one culture may be considered inappropriate in another. This system performs analysis that takes into account the cultural background of the content. Third, monitoring and mitigating bias. As mentioned above, AI models may reflect bias in the training data. This system periodically audits biases related to gender, race, religion, political stance, etc., and takes measures to mitigate any detected biases. Fourth, ensuring accessibility. This system is made available to people with visual impairments, hearing impairments, and other disabilities. For example, it will implement features such as text-to-speech, subtitle display, keyboard navigation, and color schemes that are considerate of color blindness. Fifthly, it will respect diverse perspectives. This system will not impose a particular political stance or values, but will present diverse viewpoints and enable users to make their own judgments. For example, it will present official information and expert opinions from different perspectives side by side on a given issue.

[0158] (Embodiment 54) Next, as a 54th embodiment of the present invention, the details of the resilience function will be described. This system is designed to have high resilience against attacks and failures. Specifically, the following measures are taken: First, a distributed architecture. Each component of the system is located in multiple geographically distributed data centers. Even if some data centers fail or are subjected to cyberattacks, other data centers will continue to function. Second, redundancy. Critical data and services are replicated to multiple locations. The database is redundant in a master-slave configuration or a multi-master configuration. Third, automatic failover. If a failure is detected, it automatically switches to a backup system. Failover is designed to be completed within a few seconds to a few minutes. Fourth, cybersecurity measures. Due to the nature of this system, which detects misinformation, it may be a target for attacks from those who manipulate information. This system implements measures such as firewalls, intrusion detection systems (IDS), intrusion prevention systems (IPS), DDoS attack countermeasures, encrypted communication, access control, and regular security audits. Fifth, data backup. Important data is regularly backed up and stored in multiple locations. Recovery procedures from backups are also regularly tested. Sixth, there is an incident response plan. Procedures for responding to failures and attacks are documented in advance and communicated to relevant parties. Regular training is conducted to maintain response capabilities.

[0159] (Embodiment 55) Next, as a 55th embodiment of the present invention, the details of energy efficiency and environmental consideration functions will be described. Operating a large-scale AI system requires enormous computing resources and energy. In this embodiment, considerations are made to minimize the impact on the environment. Specifically, the following measures are taken. First, energy-efficient hardware is selected. The latest low-power processors, GPUs, and dedicated AI chips (e.g., Google TPU, NVIDIA Tensor Core) are used. Second, the model is made lightweight. As mentioned above, the size and computational load of the model are reduced using techniques such as knowledge distillation, pruning, and quantization, thereby reducing the energy required. Third, an efficient data center is used. An energy-efficient data center with a low PUE (Power Usage Effectiveness) is selected. Energy consumption is also reduced by optimizing the cooling system and reusing waste heat. Fourth, renewable energy is used. Priority is given to supplying power to the data center with renewable energy such as solar, wind, and hydroelectric power. Fifth, computing resources are optimized. Computing resources are used only when needed, and resources are released when idle. We will leverage the cloud's auto-scaling capabilities to dynamically adjust resources according to demand. Sixth, we will measure and report our carbon footprint. We will measure and publish the amount of carbon dioxide emitted by the operation of this system. We will also engage in carbon offsetting (activities to offset the carbon dioxide emissions).

[0160] (Embodiment 56) Next, as the 56th embodiment of the present invention, the details of the regulatory compliance function will be described. Regulations concerning AI, data protection, and online content are being strengthened in countries around the world. This embodiment provides a function to comply with these regulations. Specifically, it complies with the following regulations: Firstly, the EU AI Act. The EU classifies AI systems according to their risk level and imposes strict requirements on high-risk AI. Because this system has the socially important function of misinformation detection, it may be classified as a high-risk AI. This system is designed to meet the requirements of transparency, explainability, human oversight, risk management, and data governance required by the AI ​​Act. Secondly, the General Data Protection Regulation (GDPR). When processing users' personal data, this system complies with the requirements of the GDPR (consent acquisition, data minimization, purpose limitation, retention period limitation, protection of data subject rights, etc.). Thirdly, the Digital Services Act (DSA). The EU DSA obligates online platforms to deal with illegal and harmful content. This system provides a function to help platform operators meet the requirements of the DSA. Fourth, there are the defamation, privacy, and copyright laws of each country. This system will take appropriate precautions in collecting, analyzing, and publishing content to avoid violating these laws. Fifth, there are election-related regulations. Many countries have special regulations against misinformation during election periods. This system will operate with particular caution and cooperate with regulatory authorities during election periods.

[0161] (Embodiment 57) Next, as the 57th embodiment of the present invention, the details of the user engagement function will be described. In order to maximize the effectiveness of this system, it is important that users actively use the system and provide feedback. In this embodiment, functions are provided to increase user engagement. Specifically, the following measures are taken. First, personalization. The information presented is customized according to the user's interests. For example, if a user is interested in politics, the results of misinformation verification related to politics will be displayed preferentially. Personalization involves analyzing the user's past browsing history, search history, feedback history, etc. Second, a notification function. Push notifications and email notifications are sent when new verification results are released or important misinformation is detected regarding topics or accounts that the user follows. Third, social functions. Functions are provided that allow users to follow other users, comment on verification results, and participate in discussions. By forming a user community, continued use is promoted. Fourth, gamification. As mentioned above, rewards such as points, badges, and rankings are provided for user contributions (providing feedback, participating in verification, etc.). Fifth, educational content. We will enhance the aforementioned educational and awareness-raising functions and provide content that allows users to learn while having fun. Sixthly, we will create a user-friendly interface. We will design an intuitive and easy-to-understand user interface so that even users who are not tech-savvy can easily use it.

[0162] (Embodiment 58) Next, as the 58th embodiment of the present invention, the details of the business model and sustainability functions will be described. In order to operate this system over the long term and to continuously improve it, a sustainable business model is necessary. In this embodiment, the following business models are considered. First, operation with public funds. The system is operated as a public service with grants and subsidies from governments, international organizations, foundations, etc. In this case, the system is made publicly available free of charge and widely used. Second, a subscription model. Basic functions are provided free of charge, and advanced functions (e.g., detailed analysis reports, API access, priority support) are provided through paid subscriptions. It is assumed that companies, media, research institutions, etc. will be the paying users. Third, a licensing model. The technology of this system is licensed to SNS platform operators, media companies, etc., and they integrate it into their own services. License fees will be the source of revenue. Fourth, an advertising model. Advertisements are displayed on the system's website or app, and advertising revenue is obtained. However, in order to prevent advertisements from becoming a breeding ground for misinformation, the content and placement standards of advertisements will be strictly reviewed. Fifth, a data sales model. The data collected and analyzed by this system (anonymized and with privacy protected) will be sold to research institutions and companies. However, the sale of data will be conducted carefully with ethical and legal considerations in mind. Sixthly, it is a donation model. Donations from the general public and organizations will be solicited to fund operations. This is a non-profit model, similar to Wikipedia.

[0163] (Embodiment 59) Next, as the 59th embodiment of the present invention, the details of the international cooperation function will be described. International cooperation is essential because misinformation spreads across national borders. This embodiment provides a function to promote international cooperation. Specifically, the following initiatives will be undertaken. First, international data sharing. Information verified by fact-checking organizations, research institutions, and government agencies in each country will be shared in an international database. This system will function as the platform for this database. Second, the development of international standards. The system will participate in efforts to internationally standardize methods for verifying misinformation, data formats, evaluation criteria, etc. Standardization will improve interoperability between different organizations and systems. Third, international research cooperation. The system will collaborate with research institutions around the world to promote research on misinformation. Knowledge will be shared through joint research projects, international conferences, and publications. Fourth, support for developing countries. The system's technology and know-how will be provided to developing countries that lack resources for combating misinformation, and capacity building support will be provided. Fifth, contribution to international legal frameworks. Technical expertise will be provided in the development of international legal frameworks and treaties concerning combating misinformation.

[0164] (60th embodiment) Next, as the 60th embodiment of the present invention, the details of the scalability function for future technologies will be described. This system is designed to be compatible not only with current technologies but also with new technologies that will emerge in the future. Specifically, the following future technologies are envisioned: Firstly, quantum computing. If quantum computers become practical, current encryption technologies may be broken, but high-speed analysis using quantum algorithms will also be possible. This system will consider introducing quantum-resistant cryptography and utilizing quantum machine learning algorithms. Secondly, brain-computer interfaces (BCIs). If technologies emerge in the future that allow direct information exchange with the human brain, new forms of information manipulation (for example, implanting false memories directly into the brain) may become possible. This system will continue research to be able to respond to such new threats. Thirdly, augmented reality (AR) and virtual reality (VR). If AR and VR become widespread, immersive misinformation (for example, false events in virtual space) may emerge. This system will develop verification technologies for AR and VR content. Fourthly, next-generation AI. If more advanced artificial general intelligence (AGI) that surpasses current AI emerges, methods of information manipulation will also become more sophisticated. This system will continuously evolve by incorporating the latest AI technologies. Fifthly, it is a new communication platform. In the future, new forms of communication platforms different from current SNS may emerge (for example, decentralized social networks, communication within the metaverse). This system will maintain a flexible architecture to accommodate these new platforms.

[0165] (61st Embodiment) Next, as the 61st embodiment of the present invention, the details of the modification detection function at the audio waveform level will be described. In previous embodiments, analysis at the audio feature quantity and phoneme level has been mainly described, but in this embodiment, subtle modifications are detected by directly analyzing the audio waveform itself at a lower level. An audio waveform is represented as a change in amplitude on the time axis. In this system, the following analysis is performed on the audio waveform. First, the continuity of the waveform is verified. In natural speech, the waveform is smoothly continuous. When audio is cut and pasted through editing, discontinuities in the waveform (phase jumps, abrupt changes in amplitude) may occur at the editing points. In this system, the first and second derivatives of the waveform are calculated to detect abnormal discontinuities. Second, the noise floor is analyzed. Background noise (noise floor) always exists in the recording environment. Audio recorded in the same recording session has the same noise floor. When audio recorded in different environments is cut and pasted, the characteristics of the noise floor change. In this system, the noise spectrum of the silent parts of the audio is analyzed to detect changes in the noise floor. Third, quantization noise is analyzed. Digital audio is obtained by quantizing an analog signal into discrete values. Quantization noise is generated during this quantization process. When audio is re-encoded or converted between different formats, the pattern of quantization noise changes. This system analyzes the statistical characteristics of quantization noise to detect traces of re-encoding. Fourthly, there is clipping detection. If the audio volume is too high, the waveform is cut off at the upper or lower limit (clipping). Clipping not only degrades sound quality but also indicates that the audio may have been manipulated. This system analyzes the amplitude distribution of the waveform to detect the presence or absence of clipping.

[0166] Fifth, there is the analysis of echo and reverberation. The acoustic characteristics of the recording environment (room size, wall material, etc.) are reflected in the sound as echo and reverberation. Sound recorded in the same environment will have the same echo and reverberation pattern. When sound recorded in different environments is combined, the echo and reverberation patterns will be inconsistent. This system uses impulse response estimation technology to extract the echo and reverberation pattern from the sound and verify its consistency. Sixth, there is the detection of power supply hum. Electromagnetic interference from the power supply can introduce hum noise of a specific frequency (e.g., 50Hz or 60Hz) into the sound. The frequency of the hum noise corresponds to the power supply frequency of the area where the recording was made. This system analyzes the presence and frequency of hum noise to estimate the recording area and detect the mixing of sound recorded in different areas. Through these detailed waveform-level analyses, this system can detect even extremely subtle modifications that are difficult to detect at the feature level.

[0167] (62nd Embodiment) Next, as a 62nd embodiment of the present invention, the details of the speech rate manipulation detection function will be described. By changing the speech rate of a voice, it is possible to manipulate the impression of what is being said. For example, by increasing the speech rate, it is possible to give the impression that the speaker is anxious or insincere, while conversely, by decreasing the speech rate, it is possible to give the impression that the speaker is deep in thought or lacking confidence. In this embodiment, the manipulation of speech rate is detected. First, the speech rate is estimated from the voice. Speech rate is defined as the number of phonemes, syllables, or words per unit time. In this system, the speech rate is calculated using phoneme recognition or syllable segmentation techniques. Next, the estimated speech rate is compared with the speaker's normal speech rate. The speaker's normal speech rate is statistically estimated from past official voice recordings. If the estimated speech rate deviates significantly from the normal range, it indicates the possibility of speed manipulation. Furthermore, this system also estimates the method of changing the speech rate. If the speed is changed by simple time stretching, the pitch of the voice does not change, but distortion occurs in the subtle temporal structure of the voice. This system detects traces of time stretching algorithms such as Phase Vocoder and WSOLA (Waveform Similarity Overlap-Add). It also analyzes whether changes in speech rate are applied uniformly to the entire speech or only to specific parts. If the rate is changed only in specific parts, it is highly likely that this was an intentional manipulation.

[0168] (63rd embodiment) Next, as a 63rd embodiment of the present invention, the details of the background audio / ambient sound consistency verification function will be described. Videos and audio include not only the speaker's voice but also background audio and ambient sounds (e.g., background noise, traffic noise, bird songs, wind sounds, etc.). These background sounds are important clues for verifying the authenticity of the content. In this embodiment, the consistency of background audio / ambient sounds is verified. First, the speaker's voice and background sounds are separated from the audio. Mixed audio is separated into multiple sound sources using source separation techniques, particularly deep learning-based methods (e.g., Conv-TasNet, Demucs). Acoustic scene classification is performed on the separated background sounds to estimate the type of scene (e.g., "indoor / conference room," "outdoor / street," "inside a car," etc.). The estimated scene is compared with the visual content of the video, metadata, and text information to verify consistency. For example, if the video shows an outdoor park, but the background sound has indoor reverberation characteristics, this is inconsistent. Furthermore, the temporal changes in background sound are also analyzed. In natural recordings, background sound changes over time (for example, human speech may be heard and then fade away, or cars may pass by). If the background sound is unnaturally constant, it may indicate that it was artificially added. In addition, the acoustic relationship between the background sound and the speaker's voice is also examined. When recorded simultaneously in the same environment, the speaker's voice and background sound share the same reverberation characteristics and the same noise floor. If these do not match, it is possible that the speaker's voice and background sound were recorded separately and later combined.

[0169] (64th embodiment) Next, as the 64th embodiment of the present invention, the details of the automatic verification function for subtitles and captions will be described. Subtitles and captions are often displayed in videos. It is also important to verify whether these subtitles and captions match the actual audio content. In this embodiment, automatic verification of subtitles and captions is performed. First, the text of the subtitles and captions is extracted from the video. Optical character recognition (OCR) technology is used to read characters from the video frames. As the OCR engine, Tesseract, Google Cloud Vision API, Amazon Textract, etc. are used. Non-Latin characters such as Japanese and Chinese are also supported. The extracted subtitle and caption text is compared with the transcript text obtained by speech recognition. If the two match, the subtitles are accurate. If there is a discrepancy, the following possibilities are considered. First, there is an error in the subtitle. The subtitle creator may have entered it incorrectly or intentionally altered it. Second, there is an error in speech recognition. The speech recognition may have been inaccurate. In this case, the results of multiple speech recognition engines are compared to confirm consistency. Third, it is a summary or paraphrase. Subtitles may summarize or paraphrase audio content due to space constraints or readability. In such cases, the degree of semantic similarity is evaluated. The semantic similarity calculation (cosine similarity of sentence embedding vectors) described above is used to evaluate the degree of semantic similarity. If the meaning differs significantly, it is a problematic subtitle.

[0170] Furthermore, in this embodiment, the timing of subtitle and caption display is also verified. Subtitles should be displayed in sync with the corresponding audio. The timing of subtitle display is compared with the timing of the corresponding part of the audio, and if there is a significant discrepancy, it may be a sign of editing. The visual characteristics of subtitles and captions (font, color, position, background) are also analyzed. Official videos often use subtitles with a consistent style. If the style of the subtitles changes midway through, it may indicate that footage from different sources has been combined. In addition, it is verified whether emotional expressions or exaggerations have been added to the subtitles. For example, if the audio says "may increase," but the subtitle displays "will increase dramatically!", that is an exaggeration. Through this automated verification of subtitles and captions, this system can also detect visual information manipulation.

[0171] (65th embodiment) Next, as a 65th embodiment of the present invention, the details of the function for estimating the physiological and psychological state of the speaker will be described. Speech contains information about the speaker's physiological and psychological state (e.g., stress, fatigue, emotion, health). In this embodiment, these states are estimated and used to understand the context of the content. First, features that serve as indicators of the physiological and psychological state are extracted from the speech. Stress and tension often manifest as an increase in pitch (F0), an increase in speech rate, and an increase in voice jitter. Fatigue manifests as a decrease in voice energy, a decrease in speech rate, and an increase in pauses. Emotions are estimated using the emotion recognition technology described above. Health conditions (e.g., a cold, a sore throat) manifest as changes in voice quality (e.g., hoarseness, nasal voice). These features are input into a machine learning model to estimate the speaker's state. The estimated state is compared with the content and situation of the speech to verify consistency. For example, if a speaker is unusually relaxed while making an important announcement, it may be unnatural. Conversely, if a speaker appears to be under extreme stress during a routine press conference, it may indicate that something unusual is happening. Also, if the estimated state of the speaker differs significantly between official videos and clipped videos, it's possible the audio has been altered.

[0172] Furthermore, this embodiment also estimates the speaker's cognitive load. Cognitive load is the degree of burden on the brain when performing a task. Cognitive load increases when lying or processing complex information. Cognitive load manifests as decreased speech fluency (increased hesitation and repetition), increased pauses, and fluctuations in speech rate. This system estimates cognitive load from these characteristics and uses it to evaluate the reliability of speech. However, the estimation of these physiological and psychological states has limited certainty and varies greatly from person to person, so it should be treated only as reference information and not as a decisive factor in judgment.

[0173] (66th Embodiment) Next, as the 66th embodiment of the present invention, the details of the function for responding to the evolution of speech synthesis technology will be described. Speech synthesis technology, in particular neural text-to-speech (Neural TTS), is evolving rapidly and can now generate synthesized speech that is so natural that it is indistinguishable from human speech. Representative technologies include Tacotron, WaveNet, FastSpeech, VITS, and more recently, Diffusion model-based speech synthesis. In this embodiment, continuous research and development is conducted to detect speech generated by these latest speech synthesis technologies. Specifically, the following approaches are taken. First, the latest speech synthesis models are continuously investigated and the characteristics of the speech they generate are analyzed. Each model may have its own unique "fingerprint" (e.g., characteristics of a specific frequency band, phase pattern, temporal structure). These fingerprints are learned and incorporated into the detection model. Second, the detection model is strengthened using adversarial examples. Developers of speech synthesis technology may generate adversarial examples to avoid detection. In this system, a detection model that is robust to adversarial examples is constructed using an adversarial learning framework. Thirdly, there is multimodal verification. By integrating and verifying multiple sources of information, including not only audio but also video (lip-sync), text (semantic consistency), and metadata, it can detect sophisticated forgery that is difficult to detect with a single modality. Fourthly, there is collaboration with the community. We collaborate with the speech synthesis technology research community and the security research community to share the latest threat intelligence and jointly develop countermeasures.

[0174] (Embodiment 67) Next, as the 67th embodiment of the present invention, the details of the real-time audio streaming analysis function will be described. In previous embodiments, the analysis of recorded video and audio files has been mainly described, but in this embodiment, audio streaming delivered in real time (e.g., live broadcasts, radio broadcasts, podcasts) is analyzed. Real-time analysis presents the following technical challenges. First, low latency is required. The analysis results need to be obtained almost simultaneously with the delivery. In this system, the streaming audio is divided into small chunks (e.g., in units of a few seconds), and each chunk is analyzed sequentially. Processes such as speech recognition, feature extraction, and sentiment recognition are accelerated using algorithms and hardware (GPU, dedicated chip) optimized for real-time processing. Second, context preservation is required. When analyzing in chunk units, the context between chunks may be lost. In this system, information from past chunks is preserved and used in the analysis of the current chunk. For example, time-series models such as RNN and LSTM are used to consider the long-term context. Third, addressing network instability is required. Streaming delivery may experience delays, packet loss, and quality degradation depending on the network conditions. This system performs robust analysis to address these issues. For example, even if some audio is lost due to packet loss, it can fill in the gaps using surrounding information.

[0175] Applications of real-time analysis include the following: Firstly, monitoring live streams. Live streams of political speeches, corporate earnings announcements, press conferences, etc., are monitored to detect problematic statements or statements likely to be edited out later in real time. Secondly, immediate fact-checking. Statements are compared with database information in real time to immediately point out factual errors or inconsistencies. Thirdly, real-time subtitle verification. Subtitles displayed on live streams are verified in real time to ensure they match the actual audio. Fourthly, emergency alerts. If significant misinformation or dangerous statements are detected in real time, alerts are immediately sent to the relevant parties.

[0176] (Embodiment 68) Next, as the 68th embodiment of the present invention, the details of the voice data anonymization and privacy protection functions will be described. Since this system collects and analyzes a large amount of voice data, privacy protection is extremely important. In this embodiment, a voice data anonymization technology is provided. Voice may contain not only the speaker's personal identification information (voiceprint), but also information about gender, age, health status, emotional state, and even socioeconomic background. Voice anonymization means removing this personal identification information while retaining the content of the voice (linguistic information). The following are some methods for voice anonymization. First, there is pitch shifting. By changing the pitch of the voice, the original speaker's voiceprint is changed. However, excessive pitch shifting impairs the naturalness of the voice. Second, there is voice conversion. The voice quality of one speaker is converted to that of another speaker. Deep learning-based voice conversion technology (e.g., CycleGAN-VC, StarGAN-VC) is used. Third, there is regeneration by speech synthesis. The content is converted to text using speech recognition, and that text is synthesized into speech using a different voice. This completely removes the original speaker's voiceprint. Fourthly, there is the perturbation of the speaker embedding vector. By adding noise to the speaker embedding vector within the framework of differential privacy, it becomes difficult to identify individuals while still allowing statistical analysis.

[0177] This system applies these anonymization techniques according to the purpose of data use. For example, when data is made public for research purposes, advanced anonymization is performed. When used for internal analysis, privacy is protected through access control and encryption. Furthermore, this system minimizes the retention period of audio data. Once analysis is complete, only the necessary information (e.g., transcribed text, features, and judgment results) is retained, and the original audio data is deleted. Long-term storage or secondary use of data is only performed with the user's consent.

[0178] (69th embodiment) Next, as the 69th embodiment of the present invention, the details of the voice quality evaluation function will be described. Voice quality greatly affects the accuracy of analysis. With low-quality voice (for example, with a lot of noise, low volume, or severe degradation due to compression), the accuracy of speech recognition and feature extraction decreases. In this embodiment, the voice quality is automatically evaluated, and an analysis strategy is selected according to the quality. The following are some of the evaluation indices for voice quali...

Claims

1. A content acquisition means for acquiring target content to be verified from a predetermined platform, Audio extraction means for extracting audio data from the target content, A transcription means that generates text data representing the content of the utterance and time information corresponding to the text data from the aforementioned audio data using a speech recognition model, A speech feature analysis means for extracting speech features including pitch, volume, speaking speed, and pauses from the aforementioned speech data, A positive information database that stores audio data, text data, time information, and audio features of positive information content derived from official sources, **A difference analysis means that compares the target content with one or more positive information contents stored in the positive information database and analyzes the difference between the two, A text difference analysis process is performed which aligns the text data generated by the transcription means with the text data of the positive information content using the time information, and analyzes the semantic similarity between the two using a pre-trained language model. The audio difference analysis process involves aligning the time-series data of the audio features extracted by the audio feature analysis means with the time-series data of the audio features of the positive information content using dynamic time stretching (DTW), and analyzing the similarity for each corresponding interval. A contextual difference analysis process is performed to analyze the degree of contextual gaps based on the coverage rate of the target content relative to the entire positive information content. A differential analysis means** that performs the following: A score calculation means that calculates a modification score indicating the degree of modification of the target content based on the analysis results of the difference analysis means, An information verification system characterized by having the following features.

2. Computers The steps include obtaining the target content to be verified from a designated platform, The steps include extracting audio data from the target content, The steps include generating text data representing the spoken content and time information corresponding to the text data from the aforementioned audio data using a speech recognition model, The steps include extracting speech features from the aforementioned speech data, including pitch, volume, speaking speed, and pauses. The steps include: accessing a positive information database that stores audio data, text data, time information, and audio features of positive information content derived from official sources, and identifying positive information content to compare with the target content; **A step of comparing the target content with the identified positive information content and analyzing the difference between the two, A text difference analysis process is performed which aligns the generated text data and the text data of the positive information content using the time information, and analyzes the semantic similarity between the two using a pre-trained language model. The extracted audio feature time series data and the audio feature time series data of the positive information content are aligned using dynamic time stretching (DTW), and the similarity for each corresponding interval is analyzed in the audio difference analysis process. A contextual difference analysis process is performed to analyze the degree of contextual gaps based on the coverage rate of the target content relative to the entire positive information content. The steps to perform** and The steps include: calculating a modification score indicating the degree of modification of the target content based on the results of the difference analysis; An information verification method characterized by performing the following.

3. Computers, A content acquisition means for acquiring target content to be verified from a predetermined platform. Audio extraction means for extracting audio data from the target content, A transcription means that generates text data representing the content of the utterance and time information corresponding to the text data from the aforementioned audio data using a speech recognition model. A voice feature analysis means for extracting voice features including pitch, volume, speaking speed, and pauses from the aforementioned voice data. A means for accessing a positive information database that stores audio data, text data, time information, and audio features of positive information content derived from official sources, and for identifying positive information content to be compared with the target content. **A difference analysis means for comparing the target content and the identified positive information content and analyzing the difference between the two, A text difference analysis process is performed which aligns the text data generated by the transcription means with the text data of the positive information content using the time information, and analyzes the semantic similarity between the two using a pre-trained language model. The audio difference analysis process involves aligning the time-series data of the audio features extracted by the audio feature analysis means with the time-series data of the audio features of the positive information content using dynamic time stretching (DTW), and analyzing the similarity for each corresponding interval. A contextual difference analysis process is performed to analyze the degree of contextual gaps based on the coverage rate of the target content relative to the entire positive information content. Differential analysis means for performing the analysis** A score calculation means that calculates a modification score indicating the degree of modification of the target content based on the analysis results of the difference analysis means. An information verification program designed to function as such.

4. The aforementioned speech feature analysis means further extracts emotional features from the speech data that indicate the speaker's emotional state, The difference analysis means further analyzes the difference between the emotion features extracted from the audio data of the target content and the emotion features extracted from the audio data of the positive information content. The score calculation means further calculates an emotion manipulation score indicating the degree of emotion transformation based on the difference in the emotion features. The information verification system according to feature 1.

5. The score calculation means is A normalization process that normalizes multiple different difference analysis results obtained by the difference analysis means into a predetermined range, The process involves weighting each of the normalized difference analysis results by multiplying them by predetermined weight coefficients, The process involves summing up the weighted difference analysis results to calculate an overall modification score. The information verification system according to claim 1, characterized by performing the following:

Citation Information

Patent Citations

  • Ambient sound feature-based authenticity analysis method, apparatus and device, and medium

    CN120612960A

  • Audio signal authenticity verification method and device, equipment and medium

    CN120673780A

  • Information processing system

    JP2023135373A