LLM-based multi-modal false intelligence analysis system and method
By building a multimodal false intelligence analysis system based on LLM and using multimodal information for comprehensive analysis, the limitations of false intelligence analysis in the existing technology are solved, and efficient identification and analysis of audio, video and text are realized, forming a closed-loop mechanism for knowledge utilization and updating.
Patent Information
- Application Number
- CN202510447162.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-10
- Publication Date
- 2025-07-11
AI Technical Summary
Existing large language models have limitations in false intelligence analysis, are susceptible to training data bias, and have limited understanding of complex scenarios, making it difficult to effectively identify multimodal false intelligence such as audio, video and text.
A multimodal false intelligence analysis system based on LLM is built, combined with GPU servers, storage servers and cloud platforms, through the AIGC detection module and the intelligent sample data collection and annotation module, multimodal information is used for comprehensive analysis, including LLM large-scale language model and multi-type business model, data forgery analysis, sensitive speech recognition, sensitive vocabulary analysis and sensitive video recognition, and traceability and annotation of multimodal data through vector search and intelligent information search.
It improves the accuracy and efficiency of false information analysis, can effectively identify multi-modal false information such as audio, video and text, form a closed-loop mechanism for knowledge utilization, discovery and update, and realizes intelligent processing of multi-modal false information.
Smart Images

Figure CN120296228A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of open-source intelligence analysis, and particularly to a multi-modal false intelligence analysis system and method based on LLM. Background Art
[0002] The statements in this section only provide background information related to the present disclosure and may not constitute prior art.
[0003] Large Language Models (LLMs) are one of the major breakthroughs in the field of artificial intelligence in recent years. By training on a vast amount of data, they can generate high-quality natural language texts and complete various complex tasks such as translation, question answering, and text generation. Due to the progress of deep learning technology, especially the emergence and popularization of models based on the Transformer architecture, the LLM technology has been changing rapidly.
[0004] The false intelligence analysis technology based on LLM mainly utilizes the capabilities of large language models in understanding and generation, and judges the authenticity of information through methods such as comparison and analysis. Its technical principle is to input the information to be verified into the large language model. The large language model understands and analyzes the input information, extracts the key information and features therein, and verifies the authenticity of the input text by comparing the extracted key information and features with known true information. According to the comparison and verification results, false analysis conclusions or relevant suggestions are output.
[0005] Currently, some researchers have tried to apply large language models to the field of false intelligence analysis. For example, some studies have evaluated the performance of large language models such as ChatGPT in detecting the authenticity of news and found that these models can identify false information to a certain extent. However, these studies still have certain limitations and challenges, being easily affected by biases in the training data and having limited ability to understand complex scenarios, etc.
[0006] With the continuous development and improvement of the current large language model technology, the false intelligence analysis technology based on large language models is expected to play a more important role in the future. By continuously optimizing the model structure, improving the model performance, introducing multi-modal information and other means, the accuracy and efficiency of false intelligence analysis can be further improved.
[0007] The present invention makes full use of the language understanding and generation capabilities of LLM in the field of artificial intelligence, and comprehensively analyzes multi-modal data through deep learning technology. For text forgery, the grammar and semantic understanding capabilities of LLM are used to identify and exclude forged texts; in terms of audio and video, sound and image analysis technologies are adopted to conduct false analysis on sensitive audio-visual content, especially in the field of open-source intelligence analysis, which can effectively identify multi-modal false intelligence such as audio, video, and text. Summary of the Invention
[0008] The object of the present invention is to effectively identify multi-modal false intelligence such as audio, video, and text for false intelligence analysis applications. The overall architecture of the method consists of two parts: intelligent sample data collection, annotation, and fine screening, and the AIGC detection module for multi-modal information, relying on the infrastructure (IAAS) composed of GPU servers, storage servers, and cloud platforms, the size of which is planned according to specific business scenarios. By accessing multi-modal data from external sources, the analysis and disposal of multi-modal false intelligence are realized.
[0009] The technical solution of the present invention is as follows:
[0010] A multi-modal false intelligence analysis system based on LLM consists of an AIGC detection module for multi-modal information and an intelligent sample data collection, annotation, and screening module, relying on the infrastructure composed of GPU servers, storage servers, and cloud platforms. By accessing multi-modal data from external sources, the analysis and disposal of multi-modal false intelligence are realized.
[0011] Further, the AIGC detection module provides a false intelligence analysis model and false intelligence data analysis function, providing capacity support for the false data identification and analysis function;
[0012] Among them, the false intelligence analysis model includes: LLM large language model and multi-type business models;
[0013] The false intelligence data analysis function constructs a sensitive word library based on the false intelligence analysis model to realize various data detection tasks, including: data forgery analysis, sensitive speech recognition, sensitive vocabulary analysis, and sensitive video recognition.
[0014] Further, the multi-type business models include: optical character recognition model, speech transcription model, face recognition model, scene recognition model, text vectorization model, image vectorization model, graphic and text vectorization model, audio vectorization model, video vectorization model.
[0015] Further, the intelligent sample data collection, annotation, and screening module includes: data annotation management function, data retrieval management function, and data crawling management function;
[0016] The data annotation management function includes: text management function, image management function, audio management function, and video management function;
[0017] The data retrieval management function includes: text retrieval function, image retrieval function, audio retrieval function, video retrieval function, and advanced retrieval function;
[0018] The data crawling management function includes: receiving the detection results input by the AIGC detection module, performing intelligent information search through the LLM large language model, planning automated collection tasks, and invoking the web crawler management tool to crawl information of multi-modal data, that is, traceability.
[0019] Furthermore, the text management function includes: managing the forged text data after traceability and analysis by the LLM large language model, and displaying the text information after false analysis and keyword extraction in the form of a list;
[0020] The image management function includes: managing the forged picture data after traceability and analysis by the LLM large language model, and displaying the picture information after false analysis and keyword extraction in the form of a list, including sensitive person information, sensitive text information, and sensitive scene information;
[0021] The audio management function includes: managing the forged and sensitive audio data after traceability, speech transcription, and analysis by the LLM large language model, and displaying the audio information after false analysis and keyword extraction in the form of a list, including sensitive person information and sensitive text information;
[0022] The video management function includes: managing the forged video data after traceability and analysis by the LLM large language model, and displaying the video information after false analysis and keyword extraction in the form of a list, including sensitive person information, sensitive text information, and sensitive scene information.
[0023] Furthermore, the text retrieval function includes: using NLP natural language technology to find the corresponding information in multi-modal data in the form of natural language questions, and presenting the retrieved results after sorting and summarizing;
[0024] The image retrieval function includes: image tag retrieval and image vector retrieval; among them, image tag retrieval intelligently identifies the image through multiple dimensions, returns the image tags contained in the image, and then precisely retrieves the content in the bottom library through the tags to return the required multi-modal data materials; image vector retrieval uses the image vectorization model to convert the image into a vector, and then performs vector retrieval on the content in the bottom library through vector similarity comparison;
[0025] The audio retrieval function includes: audio tag retrieval and audio vector retrieval; audio tag retrieval parses the content in the audio through the speech transcription model, then extracts the keywords in the audio through NLP, and retrieves the multi-modal content in the bottom library through the keywords; audio vector retrieval uses the audio vectorization model to convert the audio into a vector, and then performs vector retrieval on the content in the bottom library through vector similarity comparison;
[0026] Video retrieval function, including: video tag retrieval and video vector retrieval; video tag retrieval intelligently identifies videos through multiple dimensions, returns the tags contained in the videos, and then precisely retrieves the content in the bottom library through the tags to return the required multi-modal data materials; video vector retrieval converts videos into vectors through a video vectorization model, and then performs vector retrieval on the content in the bottom library through vector similarity comparison;
[0027] Advanced retrieval function, including: full-text retrieval, precise retrieval, tag retrieval, fuzzy retrieval and secondary retrieval.
[0028] Furthermore, the video tag extraction process includes:
[0029] First, perform language detection on the audio, and then send tasks to 6 atomic AI algorithms in parallel after decoding the video once, including: audio extraction algorithm, text extraction algorithm, face feature extraction algorithm, scene recognition algorithm, landmark recognition algorithm, object recognition algorithm; if there is no manuscript information in the audio-visual material, the text extracted from the audio will be used as the input for text analysis; if it is in English, the text translation needs to be called to translate the text into Chinese.
[0030] Furthermore, the manuscript tag extraction process includes:
[0031] If the audio-visual material has manuscript information, the manuscript will be used as the input for text tag extraction. If there is no manuscript, the text generated by speech recognition will be used instead.
[0032] Furthermore, the sensitive tag extraction process includes:
[0033] Sensitive information filtering comprehensively detects various media types, including people, scenes, flags, emblems, and text; sensitive information filtering relies on the metadata generated by manuscript tag extraction and video tag extraction, and combines with the sensitive information library it carries to conduct comparison and analysis, so as to identify sensitive information.
[0034] The present invention also proposes a multi-modal false intelligence analysis method based on LLM. Based on the above-mentioned multi-modal false intelligence analysis system based on LLM, it includes:
[0035] Step S1: The AIGC detection module receives the data accessed from the outside, schedules multiple types of business models based on the LLM large language model, and realizes the recognition of sensitive people in the sensitive person library, sensitive words in the sensitive vocabulary, and sensitive scenes; through the large model, various Internet search engines and internal vector retrieval engines are called to trace the origin of multi-modal data, and the logical analysis ability of the large model is used to analyze the common sense logical errors in the text content and the text source after tracing;
[0036] Step S2: Based on the detection results of the AIGC detection module and the task scheduling ability of the LLM large language model, the intelligent sample data collection annotation screening module automatically invokes the network search engine and the open-source information crawling tool to realize the automated and intelligent collection of data in the security industry. Then, the collected data is called to the AIGC detection module for pre-annotation, and after fine screening, it forms the ability to automatically expand the large model dataset of the security industry, forming a closed loop of knowledge utilization + knowledge discovery + knowledge update.
[0037] Compared with the existing technologies, the beneficial effects of the present invention are as follows:
[0038] The present invention utilizes the language understanding and generation ability of the LLM, and aims at multi-modal open-source intelligence data such as audio, video, and text. Relying on key technologies such as vector retrieval, process orchestration and intelligent agent integration, and multi-modal artificial intelligence label extraction, through intelligent sample data collection annotation fine screening and multi-modal information AIGC detection, it realizes the analysis and processing of multi-modal false intelligence. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] Figure 1 It is a block diagram of a multi-modal false intelligence analysis system based on the LLM;
[0040] Figure 2 It is a business process flowchart of multi-modal false intelligence analysis based on the LLM;
[0041] Figure 3 It is a flowchart of multi-modal artificial intelligence label processing of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0042] It should be noted that relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprising", "including" or any other variation thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising a..." does not exclude the existence of additional identical elements in the process, method, article or device comprising the element.
[0043] The features and performance of the present invention will be further described in detail below in conjunction with the embodiments.
[0044] Embodiment 1
[0045] Based on this, this embodiment proposes a multi-modal false intelligence analysis system based on LLM. Please refer to Figures 1 - 3 , which consists of an AIGC detection module for various modal information and an intelligent sample data collection, annotation, and screening module, relying on an infrastructure (IAAS) composed of GPU servers, storage servers, and cloud platforms. Its size is planned according to specific business scenarios. By accessing multi-modal data from external sources, it realizes the analysis and disposal of multi-modal false intelligence.
[0046] In this embodiment, specifically, the AIGC detection module provides a false intelligence analysis model and false intelligence data analysis function, providing capacity support for the false data identification and analysis function;
[0047] Among them, the false intelligence analysis model includes: LLM large language model and multi-type business models;
[0048] The false intelligence data analysis function constructs a sensitive word library based on the false intelligence analysis model to implement various data detection tasks, including: data forgery analysis, sensitive speech recognition, sensitive vocabulary analysis, and sensitive video recognition, etc.
[0049] That is, the AIGC detection module mainly conducts data forgery analysis and management for sensitive figures and sensitive vocabulary, and realizes text forgery analysis, picture forgery analysis, audio forgery analysis, and video forgery analysis relying on the process orchestration and intelligent agent fusion architecture. Among them:
[0050] 1) Text forgery analysis:
[0051] The text forgery analysis is based on the LLM large language model. Through the logical thinking and task scheduling capabilities of the large model, it realizes the analysis of text forgery. When the text forgery analysis receives text information, first, based on the method of large model prompt engineering, it uses the logical thinking ability of the large model to find obvious logical factual defects in the text content as one of the criteria for forgery judgment; at the same time, based on the task scheduling ability of the large model, it automatically calls the internal material search engine and the Internet search engine to trace the origin and compare the text information, and finally uses the tracing result as one of the criteria for judging the AIGC generation of the text.
[0052] 2) Picture forgery analysis
[0053] Image forgery analysis is based on the task scheduling ability of the large model. It automatically invokes various visual models such as the OCR model, face recognition model, and scene recognition model for the accessed images, extracts text, person, and scene information in the images, conducts text forgery analysis on the extracted text information, and simultaneously issues early warnings for sensitive information based on face and scene recognition tags; it will automatically invoke the image vectorization model and the text-image vectorization model to trace the source of the images. The types of traceability include image and video content, and it conducts text-image vector matching on the images based on the words and sentences in the sensitive word library, and issues early warnings for the sensitive image content with a threshold higher than the set value.
[0054] 3) Audio forgery analysis
[0055] Audio forgery analysis is based on the task scheduling ability of the large model. It conducts ASR speech transcription on the received audio files, and then, based on the large model prompting engineering method for the transcribed text content, utilizes the logical thinking ability of the large model to find obvious logical factual defects in the text content; at the same time, based on the tool invocation ability of the model, it automatically invokes the internal material search engine and the Internet search engine to trace the source and compare for this text information, and finally uses the traceability result as one of the bases for judging the AIGC generation of the text.
[0056] 4) Video forgery analysis
[0057] Video forgery analysis separates the audio and video of the received video files, conducts audio forgery analysis on the audio files, the separated video is framed based on the frame extraction model, and the framed images are subjected to image forgery analysis. Finally, the analysis results of the voice and images are combined as the early warning basis for video forgery analysis.
[0058] In this embodiment, specifically, the multi-type business models include: various models such as optical character recognition model, speech transcription model, face recognition model, scene recognition model, text vectorization model, image vectorization model, text-image vectorization model, audio vectorization model, and video vectorization model.
[0059] In this embodiment, it should be noted that the intelligent sample data collection, annotation, and screening module relies on the content-based vector retrieval engine and the key technologies of the multi-modal artificial intelligence label extraction process to achieve data annotation management, data retrieval management, and data crawling management. Among them, data annotation management provides text management, image management, audio management, and video management; data retrieval management realizes text retrieval, image retrieval, audio retrieval, video retrieval, and advanced retrieval; data crawling management realizes acquisition task planning and web crawler management;
[0060] That is, the intelligent sample data collection, annotation, and screening module includes: a data annotation management function, a data retrieval management function, and a data crawling management function; that is, the data knowledge base classifies and summarizes the result data detected and analyzed by the AIGC detection module to form multi-type data knowledge bases such as text management, image management, audio management, and video management, and supports data screening and searching based on various tags.
[0061] In this embodiment, specifically, the data annotation management function includes: a text management function, an image management function, an audio management function, and a video management function;
[0062] The data retrieval management function includes: a text retrieval function, an image retrieval function, an audio retrieval function, a video retrieval function, and an advanced retrieval function;
[0063] The data crawling management function includes: receiving the detection results input by the AIGC detection module, performing intelligent information search through the LLM large language model, planning automated collection tasks, and invoking the web crawler management tool to crawl the information of multi-modal data, that is, tracing the source.
[0064] In this embodiment, specifically, the text management function includes: managing forged text data after tracing the source and analysis by the LLM large language model, including various types such as txt, word, and pdf, and displaying the text information after false analysis and keyword extraction in the form of a list;
[0065] The image management function includes: managing forged picture data after tracing the source and analysis by the LLM large language model, including various types such as jpg, JPEG, and PNG, and displaying the picture information after false analysis and keyword extraction in the form of a list, including sensitive person information, sensitive text information, and sensitive scene information;
[0066] The audio management function includes: managing forged and sensitive audio data after tracing the source, speech transcription, and analysis by the LLM large language model, including various types such as wav, m4a, and mp3, and displaying the audio information after false analysis and keyword extraction in the form of a list, including sensitive person information and sensitive text information;
[0067] The video management function includes: managing forged video data after tracing the source and analysis by the LLM large language model, including various types such as avi, mp4, and rmvb, and displaying the video information after false analysis and keyword extraction in the form of a list, including sensitive person information, sensitive text information, and sensitive scene information.
[0068] In this embodiment, specifically, the text retrieval function includes: using NLP natural language technology to find corresponding answers, people, sensitive scenarios, etc. in multi-modal intelligence data such as text, video, audio, pictures, etc. in the form of natural language questions, and presenting the retrieved results after sorting and summarizing, including the summary of the best answers, the context of the event, public opinion sentiment, knowledge graph, media resource graph, etc. When searching for an input statement, based on the capabilities of the search engine itself, the entire statement will be divided into individual words and recombined according to certain specifications, so as to judge the true search intention of the input search condition and display the content that meets the search condition;
[0069] The image retrieval function includes: image tag retrieval and image vector retrieval; among them, image tag retrieval intelligently identifies images through multiple dimensions such as sensitive people, sensitive scenarios, and sensitive words, returns the image tags contained in the image, and then precisely retrieves the content in the base library through the tags, returning the required multi-modal data materials; image vector retrieval converts the image into a vector through an image vectorization model, and then performs vector retrieval on the content in the base library through vector similarity comparison; image vector retrieval can retrieve image materials and video materials for retrieval in cases where the tags of the content to be searched cannot be confirmed;
[0070] The audio retrieval function includes: audio tag retrieval and audio vector retrieval; audio tag retrieval parses the content in the audio through a speech-to-text model, then extracts the keywords in the audio through NLP, and retrieves the multi-modal content in the base library through the keywords; audio vector retrieval converts the audio into a vector through an audio vectorization model, and then performs vector retrieval on the content in the base library through vector similarity comparison; audio vector retrieval can retrieve audio materials and video materials for retrieval in cases where the tags of the content to be searched cannot be confirmed;
[0071] The video retrieval function includes: video tag retrieval and video vector retrieval; video tag retrieval intelligently identifies videos through multiple dimensions, returns the tags contained in the video, and then precisely retrieves the content in the base library through the tags, returning the required multi-modal data materials; video vector retrieval converts the video into a vector through a video vectorization model, and then performs vector retrieval on the content in the base library through vector similarity comparison; video vector retrieval can retrieve video materials for retrieval in cases where the tags of the content to be searched cannot be confirmed and for homologous video retrieval;
[0072] The advanced retrieval function includes: full-text retrieval, exact retrieval, tag retrieval, fuzzy retrieval, and secondary retrieval; supporting special retrieval methods such as OR, AND, and NOT, which is convenient for screening out the most accurate materials through various retrievals.
[0073] In this embodiment, it should also be noted that the data crawling management function receives the key information input by the AIGC detection module of multi-modal information, conducts intelligent information search through the LLM large language model, plans automated collection tasks, and calls the network crawler management tool to crawl the information of multi-modal data.
[0074] The collection task planning is responsible for creating collection tasks, setting collection strategies, and managing and scheduling collection information sources, including:
[0075] a) Collection task planning. Decompose the collection task according to the intention and form an execution collection plan task in the form of a collection plan. At the same time, the planned task can be associated with the collection information source;
[0076] b) Collection resource control and strategy setting. Allocate collection nodes, collection IP pools, collection temporary throughput control planning, collection transmission links, etc. Equipment control and planning include adjusting collection priorities, setting collection task types (such as scheduled collection), setting collection time periods and collection frequencies;
[0077] c) Collection information source monitoring and management. Register and manage various Internet site information sources for collection, and detect, analyze, and evaluate the availability of each registered information source site.
[0078] The network crawler management is responsible for developing various collection crawler algorithms and maintaining crawling rules, including text crawlers such as think tanks, social media crawlers (such as Twitter, Facebook, WeChat official accounts, etc.), video crawlers, crowdsourced data crawlers (such as Wikipedia, Baidu, etc.), web map crawlers (such as Google Earth, Bing Maps, etc.). The main functions include:
[0079] a) Open-source intelligence collection such as think tanks, etc., to achieve targeted collection of information sources such as think tank sites, official websites, and special topic websites;
[0080] b) Social media information source collection, to achieve targeted collection of social media information such as Weibo and WeChat official accounts;
[0081] c) Video information collection, to achieve collection and download of video files from mainstream video sites;
[0082] d) Crowdsourced data collection, to achieve collection of Internet-related crowdsourced data such as Wikipedia and Baidu;
[0083] e) Web map data collection, to achieve collection of oblique photography data, street view data, time-series image data, etc.
[0084] For this embodiment, please refer to Figure 3, also in response to the characteristics of audio-visual materials and the need for label systemization in the field of open-source intelligence analysis, a multi-modal label processing process with 4 major branches, 9 atomic algorithms, and 4 types of outputs has been formed, and it can be combined with text translation to achieve multi-language label capabilities. Combining with the attached Figure 3 Explanation, the four major branches of the video label extraction process, manuscript label extraction process, sensitive information filtering process, and multi-modal comprehensive analysis process are described.
[0085] The video label extraction process is as follows:
[0086] The video files in the audio-visual materials are input into the video label extraction process. For Chinese and English, first, the language of the audio is judged, and then after the video is decoded once, tasks are sent to 6 atomic AI algorithms to work in parallel, including:
[0087] Audio extraction (speech recognition ASR) algorithm: Convert the speech in the video into text, and call different recognition engines according to different languages.
[0088] Text extraction (optical character recognition OCR) algorithm: Extract the text information in the video frame, such as subtitles, and nameplates of people in news scenes.
[0089] Face feature extraction: Extract the feature information of faces appearing in the video frame, including names, depth of field, face aspect ratio, etc.
[0090] Scene recognition: Recognize important news scenes appearing in the video frame.
[0091] Landmark recognition: Recognize well-known landmark buildings appearing in the frame.
[0092] Object recognition: Recognize various objects appearing in the frame.
[0093] If the audio-visual materials do not carry manuscript information, the text extracted from the audio will be used as the input for text analysis. If it is in English, text translation will also be called to translate the text into Chinese.
[0094] The manuscript label extraction process is as follows:
[0095] If the audio-visual materials carry manuscript information, the manuscript will be used as the input for text label extraction. If there is no manuscript, the text generated by speech recognition will be used instead. The manuscript label extraction process is divided into three types of atomic algorithms to execute concurrently, namely text label, text feature extraction (theme), and text feature extraction (keywords).
[0096] Text label: Input the text content of the news manuscript, and perform text entity extraction on it. The extracted content includes 4 types of information: names, organizational structures, place names, and brands.
[0097] Text feature extraction (theme): Extract theme features based on the content of the manuscript and classify them into corresponding themes.
[0098] Text feature extraction (keywords): Extract keywords based on the content of the manuscript. The keyword extraction is customized based on the characteristics of the manuscript format, with greater weights in the title and lead parts than in the body text and simultaneous interpretation.
[0099] After the manuscript tags are extracted, they will be input into the information filtering process and the multi-modal tag comprehensive analysis process together with the video tags to achieve comprehensive analysis based on the application scenario.
[0100] The sensitive tag extraction process is as follows:
[0101] Sensitive information filtering conducts a comprehensive detection of various media types, including people, scenes, flags, emblems, text, etc. Sensitive information filtering relies on the metadata generated by manuscript tag extraction and video tag extraction, and compares and analyzes them in combination with the sensitive information library it carries to identify sensitive information. Sensitive information filtering supports the following information filtering processes:
[0102] Sensitive people: Support the review of key political, economic, scientific and technological figures, etc. The review scope supports the human faces in the video footage, the names of news anchors, the names of news subtitles, and the nameplates of people in the video footage. Support expanding the review of more sensitive people through face feature training and registration of sensitive person names.
[0103] Sensitive emblems and flags: Support filtering by paying attention to the special emblems and flags of organizations, mainly for picture filtering.
[0104] Sensitive slogans: Review the text that appears in the video footage, including subtitle information, text banners that appear in the video, etc.
[0105] Violation language: Filter the text that appears in the news manuscript and video footage to ensure the compliance of the text.
[0106] Based on the above comprehensive identification of sensitive information, it can uniformly output the sensitive information identification results externally, and then the user conducts a secondary review and evaluation.
[0107] In this embodiment, it should be noted that the core of the multi-modal tag comprehensive analysis is to use the cross-verification between multi-modal tags to correct the accuracy of the tags. Taking the person tag as an example, three technologies of face feature detection, ASR, and OCR are used for associated cross-verification to improve the accuracy of the person tag. The cross-analysis logic is as follows:
[0108] Obtain the tag information of people, including names and corresponding timestamps / time periods, through three atomic algorithms of face recognition, OCR, and ASR respectively.
[0109] Based on the face recognition results, compare the timestamps of the names appearing in OCR and ASR one by one, and check the name information appearing in the three pieces of information to enhance the recognition information.
[0110] In this embodiment, it should be noted that vector retrieval is a content-based retrieval method. By converting text into vector representations and using the similarity between vectors to match and rank search results, the basic principle is to represent each word or phrase as a vector, and these vectors are intertwined in a multi-dimensional space. When searching, the search keyword is converted into a vector, and then the document closest to this vector is found in the document collection.
[0111] Vector retrieval is mainly implemented based on two technologies: TF-IDF and Word2Vec.
[0112] TF-IDF is a statistical method used to evaluate the importance of a word in a document. TF (Term Frequency) is the frequency of a word in a document, that is, the number of occurrences of a word in a document divided by the total number of words in the document. The formula is expressed as:
[0113]
[0114] IDF (Inverse Document Frequency) is a measure of the general importance of a word. The IDF of a specific word can be obtained by dividing the total number of documents in the corpus by the number of documents containing the word, and then taking the logarithm of the resulting quotient. The main idea of IDF is that if the number of documents containing a term is less, the stronger the information-providing ability of the term, and the greater the weight of the term. The formula is expressed as:
[0115]
[0116] Adding 1 here is to avoid the case where the denominator is 0.
[0117] TF-IDF is the product of TF and IDF, used to reflect the importance of a word in a document. The calculation formula is as follows:
[0118] TF-IDF(t,d,D) = TF(t,d) · IDF(t,D)
[0119] Word2Vec is a neural network model that can convert words into vectors through training, making words with similar semantics closer in the vector space. There are mainly two model architectures in Word2Vec: the Continuous Bag-of-Words (CBOW) model and the skip-gram model. In the CBOW model, the goal is to predict the center word based on the context words. Suppose we have a window size of c, then the model tries to predict the middle word based on the c words within the window. The objective function of CBOW can be simplified to minimize the following loss function:
[0120]
[0121] where \(w_t\) is the target word, c is the window size, \(w_{t - c},..., w_{t - 1}, w_{t + 1},..., w_{t + c}\) are the surrounding context words, and T is the number of training samples.
[0122] The goal of the skip-gram model is to predict the context words around a center word. This means that for each center word, the model will try to predict all the words that appear around that word. This makes the skip-gram more sensitive to rare words. The objective function of skip-gram can be simplified to minimize the following loss function:
[0123]
[0124] where \(w_t\) is the target word, c is the window size, and \(w_{t + j}\) represents the context words around the target word \(w_t\).
[0125] The steps for vector retrieval are as follows:
[0126] 1) Select a suitable vector model and parameters
[0127] Selecting a suitable vector model and parameters is the key to improving the effect of vector retrieval. TF-IDF is suitable for traditional text retrieval tasks, while for semantic-level retrieval, the Word2Vec model is more appropriate. When selecting a model, it is necessary to make a choice according to factors such as the dataset and application scenario.
[0128] 2) Use similarity queries
[0129] Similarity query is the core of vector retrieval. When querying, calculate the similarity score between the document and the search keyword. Using similarity queries can make the search engine more accurately match the user's search intent.
[0130] 3) Combine other retrieval methods
[0131] In the case where the storage path is clear and the index construction is imperfect, traditional text retrieval still has advantages. Combining traditional text retrieval with vector retrieval forms a hybrid retrieval method to make full use of the advantages of various retrieval methods.
[0132] 4) Aggregate the results
[0133] By aggregating the search results, it can better meet the application needs of users. Using the cross-aggregation function can improve the quality and relevance of the results to a certain extent. When aggregating the results, the diversity and accuracy of the search results need to be considered.
[0134] In this embodiment, it should also be noted that the system first realizes the visual construction of the thought chain in Langchain based on langflow, realizes the visual process choreography of the disinformation analysis task, and constructs the basic framework for the disinformation analysis business process. And based on AutoGen, task-processing agents are constructed, including playing various roles and promoting the completion of tasks through conversations. Users can construct multiple agents such as programmer agents, commander agents, and user input agents. The task is submitted to the commander agent through the user input agent. The commander agent specifies the plan to complete the task and calls various task agents including the programmer agent to realize the chain execution of related tasks, and hands the final result to the commander for inspection. The commander agent evaluates according to the completed result and puts forward opinions to continue to enter the task process for circulation until the commander agent believes that the task reaches the completion index, and returns the result to the user agent. The user agent can put forward opinions on the result through manual input and hand it over to the commander again for task planning and execution.
[0135] By manually constructing the thought chain framework of the actual business and automatically completing the sub-tasks by constructing agents, a processing architecture integrating manual choreography and agents is realized, fully exploiting the analysis and processing capabilities of the LLM large model, and achieving the rapid intelligent implementation of the multi-modal disinformation analysis business application.
[0136] The present invention also proposes a multi-modal disinformation analysis method based on LLM. Based on the above-mentioned multi-modal disinformation analysis system based on LLM, it includes:
[0137] Step S1: The AIGC detection module receives the data accessed from the outside, schedules multi-type business models based on the LLM large language model, and realizes the identification of sensitive persons in the sensitive person library, sensitive words in the sensitive vocabulary, and sensitive scenarios; through the large model, various Internet search engines and internal vector retrieval engines are called to trace the multi-modal data, and the logical analysis ability of the large model is used to analyze the common sense logical errors in the text content and the traced text sources;
[0138] Step S2: Based on the detection results of the AIGC detection module and the task scheduling ability of the LLM large language model, the intelligent sample data collection, annotation, and screening module automatically invokes the network search engine and the open-source information crawling tool to achieve the automated and intelligent collection of security industry data. Then, the collected data is called to the AIGC detection module for pre-annotation, and after fine screening, it forms the ability to automatically expand the large model dataset of the security industry, forming a closed loop of knowledge utilization + knowledge discovery + knowledge update.
[0139] In summary, this embodiment proposes a multi-modal false intelligence comprehensive analysis method. For text forgery, the syntax and semantic understanding ability of the LLM is used to identify and exclude forged texts; in terms of audio and video, voice and image analysis technologies are adopted to identify sensitive audio, video, and view content.
[0140] This embodiment constructs a content-based vector retrieval engine. By converting text into vector representations, the similarity between vectors is used to match and rank search results. Each word or phrase is represented as a vector, and the vectors are intertwined in a multi-dimensional space. When searching, the search keywords are converted into vectors, and then the documents closest to this vector are searched for in the document collection.
[0141] This embodiment designs a process choreography and agent fusion architecture. By manually constructing a thinking chain framework for open-source intelligence analysis services, the sub-tasks are automatically completed by constructing agents, realizing a processing architecture that combines manual choreography and agent fusion, fully exploiting the analysis and processing capabilities of the LLM large model, and achieving the rapid and intelligent implementation of false intelligence analysis applications.
[0142] This embodiment forms a closed-loop mechanism of "knowledge utilization + knowledge discovery + knowledge update". By using historical data and business knowledge, the AIGC detection is performed on the accessed audio, video, text, and image data. For the sensitive characters, scenes, voices, and vocabulary discovered through detection and identification, further intelligent sample data collection, annotation, and screening are carried out, and they are saved as new text, image, audio, and video annotation datasets, forming a virtuous cycle of knowledge use, discovery, and update.
[0143] This embodiment creates a multi-modal artificial intelligence label extraction process. In response to the characteristics of audio-visual materials and the need for a systematic false intelligence label system, a "4 + 9 + 4" multi-modal label processing process is formed, which can be combined with text translation to achieve multi-language label capabilities. Among them, "4 + 9 + 4" refers to 4 major branches, 9 atomic algorithms, and 4 types of outputs. The 4 major branches are: video label extraction process, manuscript label extraction process, sensitive information filtering process, and multi-modal comprehensive analysis process; the 9 atomic algorithms are: audio extraction (automatic speech recognition ASR) algorithm, text extraction (optical character recognition OCR) algorithm, face feature extraction algorithm, scene recognition algorithm, landmark recognition algorithm, object recognition algorithm, text feature extraction (keyword) algorithm, text feature extraction (theme) algorithm, and text label (person name, place name) algorithm; the 4 types of output information are: general information output, sensitive information output, news label output, and relationship analysis output (correlation calculation).
[0144] The above-described embodiments merely represent the specific implementation manners of the present application, and their descriptions are relatively specific and detailed, but they should not be construed as limiting the protection scope of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the technical solution of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application.
[0145] This background art section is provided to generally present the context of the present invention. The work of the currently named inventors, to the extent described in this background art section, and aspects of the work that are not prior art as of the time of filing this application are neither expressly nor impliedly admitted to be prior art of the present invention.
Claims
1. A multi-modal disinformation analysis system based on LLM, characterized in that, It consists of an AIGC detection module for multiple-modal information and an intelligent sample data collection, annotation, screening and selection module, relying on the infrastructure composed of GPU servers, storage servers and cloud platforms. By accessing multi-modal data from external sources, it realizes the analysis and disposal of multi-modal false intelligence.
2. The multimodal disinformation analysis system based on LLM according to claim 1, wherein, The AIGC detection module provides a false intelligence analysis model and false intelligence data analysis functions, providing capacity support for the false data identification and analysis functions; Among them, the false intelligence analysis model includes: LLM large language model and multi-type business models; The false intelligence data analysis function, based on the false intelligence analysis model, constructs a sensitive word library and realizes various types of data detection tasks, including: data forgery analysis, sensitive speech recognition, sensitive vocabulary analysis and sensitive video recognition.
3. The multimodal disinformation analysis system based on LLM according to claim 2, wherein The multi-type business models include: optical character recognition model, speech transcription model, face recognition model, scene recognition model, text vectorization model, image vectorization model, graphic and text vectorization model, audio vectorization model, video vectorization model.
4. A multimodal disinformation analysis system based on LLM according to claim 3, characterized in that The intelligent sample data collection, annotation, screening and selection module includes: data annotation management function, data retrieval management function and data crawling management function; The data annotation management function includes: text management function, image management function, audio management function and video management function; The data retrieval management function includes: text retrieval function, image retrieval function, audio retrieval function, video retrieval function and advanced retrieval function; The data crawling management function includes: receiving the detection results input by the AIGC detection module, performing intelligent information search through the LLM large language model, planning automated collection tasks, and invoking the web crawler management tool to crawl the information of multi-modal data, that is, tracing the source.
5. The multimodal disinformation analysis system based on LLM according to claim 4, characterized in that, The text management function includes: managing the forged text data after tracing the source and analyzing by the LLM large language model, and presenting the text information after false analysis and keyword extraction in the form of a list; The image management function includes: managing the forged picture data after tracing the source and analyzing by the LLM large language model, and presenting the picture information after false analysis and keyword extraction in the form of a list, including sensitive person information, sensitive text information and sensitive scene information; The audio management function includes: managing the forged and sensitive audio data after tracing the source, speech transcription and analyzing by the LLM large language model, and presenting the audio information after false analysis and keyword extraction in the form of a list, including sensitive person information and sensitive text information; The video management function includes: managing the forged video data after tracing the source and analyzing by the LLM large language model, and presenting the video information after false analysis and keyword extraction in the form of a list, including sensitive person information, sensitive text information and sensitive scene information.
6. The multimodal disinformation analysis system based on LLM according to claim 4, characterized in that, The text retrieval function includes: using NLP natural language technology to find the corresponding information in multi-modal data in the form of natural language questions, and presenting the retrieved results after sorting and summarizing; The image retrieval function includes: image tag retrieval and image vector retrieval. Among them, image tag retrieval intelligently identifies images in multiple dimensions, returns the image tags contained in the images, and then precisely retrieves the content in the base library through the tags, returning the required multi-modal data materials. Image vector retrieval uses an image vectorization model to convert images into vectors, and then performs vector retrieval on the content in the base library through vector similarity comparison. The audio retrieval function includes: audio tag retrieval and audio vector retrieval. Audio tag retrieval parses the content in the audio through a speech-to-text model, then extracts the keywords in the audio through NLP, and retrieves the multi-modal content in the base library through the keywords. Audio vector retrieval uses an audio vectorization model to convert audio into vectors, and then performs vector retrieval on the content in the base library through vector similarity comparison. The video retrieval function includes: video tag retrieval and video vector retrieval. Video tag retrieval intelligently identifies videos in multiple dimensions, returns the tags contained in the videos, and then precisely retrieves the content in the base library through the tags, returning the required multi-modal data materials. Video vector retrieval uses a video vectorization model to convert videos into vectors, and then performs vector retrieval on the content in the base library through vector similarity comparison. The advanced retrieval function includes: full-text retrieval, precise retrieval, tag retrieval, fuzzy retrieval, and secondary retrieval.
7. An LLM-based multimodal disinformation analysis system according to claim 6, characterized in that, The video tag extraction process includes: First, perform language detection on the audio, and then send tasks to 6 atomic AI algorithms in parallel after decoding the video once, including: audio extraction algorithm, text extraction algorithm, face feature extraction algorithm, scene recognition algorithm, landmark recognition algorithm, object recognition algorithm. If there is no manuscript information in the audio-visual material, the text extracted from the audio will be used as the input for text analysis. If it is in English, the text translation needs to be called to translate the text into Chinese.
8. The multimodal disinformation analysis system based on LLM according to claim 7, wherein, The manuscript tag extraction process includes: If there is manuscript information in the audio-visual material, the manuscript will be used as the input for text tag extraction. If there is no manuscript, the text generated by speech recognition will be used instead.
9. The multimodal disinformation analysis system based on LLM according to claim 8, characterized in that, The sensitive tag extraction process includes: Sensitive information filtering comprehensively detects various media types, including people, scenes, flags, emblems, and text. Sensitive information filtering relies on the metadata generated by manuscript tag extraction and video tag extraction, and combines its own sensitive information library for comparison and analysis to identify sensitive information.
10. A multi-modal disinformation analysis method based on LLM, characterized in that, A multi-modal false intelligence analysis system based on LLM according to any one of claims 1-9, including: Step S1: The AIGC detection module receives the data accessed from the outside, schedules multiple types of business models based on the LLM large language model, and realizes the identification of sensitive persons in the sensitive person library, sensitive words in the sensitive vocabulary list, and sensitive scenes. Through the large model, various Internet search engines and internal vector retrieval engines are called to trace the origin of multi-modal data, and the logical analysis ability of the large model is used to analyze the common sense logical errors in the text content and the text source after tracing. Step S2: Based on the detection results of the AIGC detection module and the task scheduling ability of the LLM large language model, the intelligent sample data collection, annotation, and screening module automatically invokes the network search engine and the open-source information crawling tool to achieve the automated and intelligent collection of data in the security industry. Then, the collected data is called the AIGC detection module for pre-annotation, and after fine screening, it forms the ability to automatically expand the large model dataset in the security industry, forming a closed loop of knowledge utilization + knowledge discovery + knowledge update.