Method, device, storage medium and program product for identifying authenticity of fused data
By using pre-trained features to process the fusion data of the network model and large language model to recognize the fusion data of images and text, the problem of low accuracy of multimodal fusion data detection results in the prior art is solved, and high-accuracy authenticity recognition and explanation generation are achieved.
Patent Information
- Application Number
- CN202510187483.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-20
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-02-20
AI Technical Summary
In the prior art, the detection results of multimodal fusion data are relatively low, and it is impossible to effectively identify the authenticity of fusion data between images and text.
The pre-trained feature processing network model is used to extract the features of the data to be identified, and input data carrying real or non-real graphic and text fusion data is input into the pre-trained large language model. By learning the feature information of the real data, the authenticity of the data is distinguished.
The accuracy of the detection results of multimodal fusion data is improved, and the authenticity of the fusion data between images and text can be accurately identified, and corresponding explanations can be generated.
Smart Images

Figure CN119671595B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of data processing, and in particular, to a method, device, storage medium and program product for identifying the authenticity of fused data. Background Art
[0002] With the development of information technology, the amount of information data on the network is becoming increasingly large and diverse in form. Among them, there is a certain amount of forged data. For users, how to identify the authenticity of information is an important means to achieve information security. Related information authenticity identification methods are usually implemented based on deep learning models.
[0003] The method for identifying the authenticity of information based on a deep learning model usually has the following characteristics: it depends on a large amount of training data for training, and cannot be applied to different types of information and complex scenarios. Therefore, the accuracy of the detection results for multi-modal fused data is relatively low. Summary of the Invention
[0004] The present application provides a method, device, storage medium and program product for identifying the authenticity of fused data, so as to at least solve the problem of low accuracy of detection results for multi-modal fused data in related technologies.
[0005] The present application provides a method for identifying the authenticity of fused data, including:
[0006] Collecting first input data and second input data; wherein, the second input data includes image data and text data; the first input data includes at least one first true fused data and / or at least one false fused data;
[0007] Inputting the second input data into a pre-trained feature processing network model, and using the pre-trained feature processing network model to extract the feature vector of the second input data to obtain third input data;
[0008] Inputting the first input data and the third input data into a pre-trained large language model, and using the pre-trained large language model to identify the authenticity of the second input data to obtain the detection result of the second input data; wherein, the detection result includes a target recognition result and target recognition description data corresponding to the target recognition result, and the target recognition result is used to indicate the authenticity of the second input data; the pre-trained feature processing network model is connected to the pre-trained large language model.
[0009] The present application also provides a device for identifying the authenticity of fused data, including:
[0010] A data acquisition module, configured to collect first input data and second input data; wherein, the second input data includes image data and text data; the first input data includes at least one first true fused data and / or at least one false fused data;
[0011] A feature extraction module, configured to input the second input data into a pre-trained feature processing network model, and extract a feature vector of the second input data by using the pre-trained feature processing network model to obtain the third input data;
[0012] An identification module, configured to input the first input data and the third input data into a pre-trained large language model, and identify the authenticity of the second input data by using the pre-trained large language model to obtain a detection result of the second input data; wherein, the detection result includes a target identification result and target identification description data corresponding to the target identification result, and the target identification result is used to indicate the authenticity of the second input data; the pre-trained feature processing network model is connected to the pre-trained large language model.
[0013] The present application also provides an electronic device, including: a memory, configured to store a computer program; a processor, configured to implement the steps of any one of the above-mentioned authenticity identification methods for fused data when executing the computer program.
[0014] The present application also provides a computer-readable storage medium, in which a computer program is stored, and wherein the computer program, when executed by a processor, implements the steps of any one of the above-mentioned authenticity identification methods for fused data.
[0015] The present application also provides a computer program product, including a computer program, and the computer program, when executed by a processor, implements the steps of any one of the above-mentioned authenticity identification methods for fused data.
[0016] Through the present application, by using a pre-trained feature processing network model, the features of the second input data to be identified are extracted, and at the same time, the first input data carrying real (or non-real) graphic-text fused data and the second input data to be identified carrying graphic-text fused data are input into the pre-trained large language model, so that the pre-trained large language model can learn the feature information of the real (or non-real) fused first input data to distinguish the authenticity of the second input data to be identified and explain the identification result. Therefore, the technical problem of low accuracy of the detection result of multi-modal fused data in the related art can be solved, and the authenticity of the fused data of images and texts can be accurately identified and an explanation of the authenticity identification result can be generated. Description of the Drawings
[0017] In order to more clearly illustrate the embodiments of the present application, the following will briefly introduce the drawings required in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0018] Figure 1 It is an application scenario diagram of a method for identifying the authenticity of fused data provided by an embodiment of the present application;
[0019] Figure 2 It is one of the schematic flowcharts of the method for identifying the authenticity of fused data provided by an embodiment of the present application;
[0020] Figure 3 It is a schematic structural diagram of a pre-trained feature processing network model provided by an embodiment of the present application;
[0021] Figure 4 It is the second of the schematic flowcharts of the method for identifying the authenticity of fused data provided by an embodiment of the present application;
[0022] Figure 5 It is an application schematic diagram of the authenticity identification of fused data provided by an embodiment of the present application;
[0023] Figure 6 It is the third of the schematic flowcharts of the method for identifying the authenticity of fused data provided by an embodiment of the present application;
[0024] Figure 7 It is a schematic structural diagram of an initial feature processing network model provided by an embodiment of the present application;
[0025] Figure 8 It is a schematic structural diagram of a device for identifying the authenticity of fused data provided by an embodiment of the present application;
[0026] Figure 9 It is a schematic structural diagram of an electronic device provided by an embodiment of the present application. Detailed implementation manners
[0027] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the protection scope of the present application.
[0028] It should be noted that in the description of the present application, the terms "include", "comprise" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements but also includes other elements not explicitly listed, or further includes elements inherent to such a process, method, article or device. The terms "first", "second", etc. in the present application are used to distinguish similar objects, rather than to describe a specific order or sequence.
[0029] Deep Learning (DL) refers to machine learning based on deep neural network models and methods. The feature of deep learning is the ability to automatically extract features, and the deep neural network is the model basis for deep learning to achieve automatic feature extraction.
[0030] DeepFake (DF) refers to an artificial intelligence technology in which a machine learning model of a generative adversarial network (GAN) combines and superimposes a picture or video onto the original picture or video, conducts large-sample learning through neural network technology, and splices and synthesizes data such as voices, facial expressions, and body movements into fake data.
[0031] The Fully Connected layer (FC) refers to a layer in which each node is connected to all nodes in the previous layer and is used to synthesize the features extracted previously.
[0032] Convolution (Conv), also known as convolution, refers to a mathematical operator that generates a third function from two functions.
[0033] A feature map refers to the result generated after an input image is processed by the convolutional layer of a neural network, representing a feature in neural space. The above features can be edges, textures, shapes, etc., and are the intermediate representation of the image in the convolutional neural network.
[0034] Batch size refers to the unified resizing of a group of images when processing images in batches.
[0035] A Large Language Model (LLM) refers to a deep learning model trained with a large amount of text data, used to understand and generate human language, and usually based on a deep neural network (such as the transformer architecture) to handle the complexity and diversity of language.
[0036] A regularization term refers to an additional term added to the objective function in a convex optimization problem, used to prevent the model from overfitting, improve the generalization ability of the model, control the complexity of the model, and promote the emergence of certain specific properties.
[0037] Pre-training refers to a strategy for training a deep learning model. Its core lies in using a large-scale dataset to preliminarily train the model so that the model learns general feature representations. This process is similar to the basic learning stage of humans before learning new knowledge, accumulating experience through extensive reading and observation.
[0038] Pre-trained large language models generally refer to designing large language model training tasks based on large-scale corpora (including language training materials such as sentences, paragraphs, etc.), training large-scale neural network algorithm structures to learn and implement. The finally obtained large-scale neural network algorithm structure and parameters are the pre-trained large language models. For subsequent other tasks, feature extraction or task fine-tuning can be performed on the basis of this model to achieve specific task purposes. The idea of pre-training is to first train a task to obtain a set of model parameters, then use this set of model parameters to initialize the network model parameters, and then use the initialized network model to train other tasks to obtain models adapted to other tasks. By pre-training on a large-scale corpus, neural language representation models can learn powerful language representation capabilities and can extract rich syntactic and semantic information from text. Pre-trained large language models can provide word elements (tokens) containing rich semantic information and sentence-level features for downstream tasks, or directly perform fine-tuning for downstream tasks on the pre-trained model to conveniently and quickly obtain downstream-exclusive models.
[0039] The neural network algorithm structure trained by the pre-trained large language model can be a large language model (Large Language Model, LLM), or CNN, RNN, LSTM, etc., or it can be a model constructed by an attention network, such as transformer, bert, GPT, Clip, etc. This application does not make a limitation here. The attention network refers to a network model trained using the attention mechanism. This model assigns different weights to each part of the input sequence, thereby extracting more important feature information in the input sequence and enabling the model to finally obtain a more accurate output.
[0040] With the development of information technology, the amount of information data on the network is becoming increasingly large and the forms are becoming more diverse. Among them, there is a certain amount of forged data. For users, how to identify the authenticity of information is an important means to achieve information security. Related information authenticity identification methods are usually implemented based on deep learning models.
[0041] The methods for identifying information authenticity based on deep learning models usually have the following characteristics: relying on a large amount of training data for training, being unable to be applied to different types of information and complex scenarios, and having a relatively low accuracy of detection results for multi-modal fusion data.
[0042] To solve the above technical problems, the present application provides a method, device, storage medium, and program product for authenticating the authenticity of fused data. By using a pre-trained feature processing network model, the features of the second input data to be authenticated are extracted. At the same time, the first input data carrying real graphic-text fused data and the second input data to be authenticated carrying graphic-text fused data are input into the pre-trained large language model, so that the pre-trained large language model can learn the feature information of the real first input data to distinguish the authenticity of the second input data to be authenticated and explain the recognition result. Therefore, the authenticity of the fused data of images and text can be accurately recognized and the corresponding explanation can be generated.
[0043] To enable those skilled in the art of the present technology to better understand the solution of the present application, the following further elaborates on the present application in conjunction with the accompanying drawings and specific embodiments.
[0044] In combination with the specific application environment architecture on which the execution of the method for authenticating the authenticity of fused data depends, the specific application environment architecture is described herein.
[0045] Refer to Figure 1 , Figure 1 FIG. is an application scenario diagram of a method for authenticating the authenticity of fused data provided by an embodiment of the present application.
[0046] The method for authenticating the authenticity of fused data provided by the embodiments of the present application can be applied to an electronic device or a device for authenticating the authenticity of fused data disposed inside the electronic device. The embodiments of the present application do not show specific details here.
[0047] Taking the execution subject as an electronic device as an example, refer to Figure 1 , in the electronic device 200, there is a pre-trained feature processing network model 101 and a pre-trained large language model 102, and the pre-trained feature processing network model 101 is connected to the pre-trained large language model 102.
[0048] The user inputs the first input data including real graphic-text fused data (such as Figure 1 shown as 1011 in Figure 1As shown by 1012 in [reference], the electronic device 200 inputs the second input data 1012 into the pre-trained feature processing network model 101, extracts the feature information (or feature vector) of the second input data 1012, and the pre-trained feature processing network model 101 inputs the feature information of the second input data 1012 into the pre-trained large language model 102. At the same time, the electronic device 200 inputs the first input data 1011 into the pre-trained large language model 102, and the pre-trained large language model 102 outputs the detection result 1021 corresponding to the second input data 1012. Among them, the detection result 1021 includes the target recognition result "The second input data is false graphic-text fusion data" for indicating the authenticity of the second input data 1012 and the corresponding explanatory data "The image content in the second input data is a residence, which does not match the text content 'This is the teaching building of University B'", as shown in the interface 1022.
[0049] Figure 2 It is one of the flow diagrams of the authenticity recognition method for fusion data provided by the embodiments of the present application. As Figure 2 shown, the embodiments of the present application provide an authenticity recognition method for fusion data. The method is described in detail as follows:
[0050] S201: Collect the first input data and the second input data; among them, the second input data includes image data and text data; the first input data includes at least one first real fusion data and / or at least one non-real fusion data.
[0051] Specifically, the first input data includes at least one first real fusion data and / or at least one non-real fusion data. The first real fusion data is used to indicate the graphic-text fusion data that has been verified as real. The non-real fusion data is used to indicate the graphic-text fusion data that has been verified as false, and the second input data is used to indicate the fusion data including image data and text data whose authenticity needs to be recognized.
[0052] For example, the first input data includes two real fusion data; or, the first input data includes three non-real fusion data; or, the first input data includes one real fusion data and two non-real fusion data.
[0053] Among them, the first real fusion data includes the first real image data and the first real text data. At least one of the image and text in the non-real fusion data is false (or non-real).
[0054] For example, the non-real fusion data includes non-real image data and non-real text data; or, the non-real fusion data includes non-real image data and real text data; or, the non-real fusion data includes real image data and non-real text data.
[0055] It can be understood that when the image and text in the non-genuine fusion data do not match, the image in the non-genuine fusion data is determined as non-genuine image data, and the text is determined as non-genuine text data.
[0056] S202: Input the second input data into the pre-trained feature processing network model, and use the pre-trained feature processing network model to extract the feature vector of the second input data to obtain the third input data.
[0057] Specifically, the pre-trained feature processing network model refers to the feature processing network model that has undergone the pre-training process. It is used to extract the feature information of the text-image fusion data to be recognized (i.e., the second input data) and represent it in vector form. The feature vector of the second input data is determined as the third input data.
[0058] S203: Input the first input data and the third input data into the pre-trained large language model, and use the pre-trained large language model to identify the authenticity of the second input data to obtain the detection result of the second input data; wherein, the detection result includes the target recognition result and the target recognition description data corresponding to the target recognition result, and the target recognition result is used to indicate the authenticity of the second input data; the pre-trained feature processing network model is connected to the pre-trained large language model.
[0059] Specifically, input the first input data and the feature vector of the second data extracted by the pre-trained feature processing network model into the pre-trained large language model, so that the pre-trained large language model can learn the feature information in the first input data, facilitate identifying whether there is forged feature information in the feature vector of the second input data, thereby identifying the authenticity of the second input data, and output the detection result of the second data. The detection result includes the target recognition result and the target recognition description data corresponding to the target recognition result. The target recognition result is used to indicate the authenticity of the second input data. The target recognition description data refers to the reason for describing the output of the target detection result by the pre-trained large language model, which is convenient for users to understand and improves the accuracy of the authenticity recognition result of the second input data.
[0060] Optionally, the target recognition result includes "genuine" or "false"; or, the target recognition result includes "genuine" or "error".
[0061] In some embodiments, the pre-trained feature processing network model includes an image feature extraction network, a text feature extraction network, and a feature fusion network. The image feature extraction network and the text feature extraction network are both connected to the feature fusion network.
[0062] Among them, the image feature extraction network is used to extract the first image feature vector of the image data in the second input data, and the first text feature extraction network is used to extract the first text feature vector of the text data in the second input data.
[0063] Among them, the feature fusion network is used to fuse each feature vector extracted by the image feature extraction network and the text feature extraction network.
[0064] Optionally, the pre-trained large language model includes a deep learning network that can perform classification processing, such as the vicuna-7b network model.
[0065] Figure 3 It is a schematic structural diagram of the pre-trained feature processing network model provided by the embodiments of this application.
[0066] See Figure 3 , the pre-trained feature processing network model 301 includes an image feature extraction network 3011, a text feature extraction network 3012, and a feature fusion network 3013. The output ends of the image feature extraction network 3011 and the text feature extraction network 3012 are both connected to the input end of the feature fusion network 3013.
[0067] In some embodiments, inputting the second input data into the pre-trained feature processing network model and using the pre-trained feature processing network model to extract the feature vector of the second input data to obtain the third input data includes:
[0068] Input the image data into the image feature extraction network, and use the image feature extraction network to extract the first image feature vector of the image data;
[0069] Input the text data into the text feature extraction network, and use the text feature extraction network to extract the first text feature vector of the text data;
[0070] Input the first text feature vector and the first image feature vector into the feature fusion network, and use the feature fusion network to fuse the first text feature vector and the first image feature vector to obtain the third input data.
[0071] Specifically, extract the image data and text data in the second input data, input the image data into the image feature extraction network, use the image feature extraction network to extract the first image feature vector of the image data, input the text data into the text feature extraction network, use the text feature extraction network to extract the first text feature vector of the text data, and use the feature fusion network in the pre-trained feature processing network to fuse the first text feature vector and the first image feature vector to obtain the feature information of the second input data, that is, the third input data.
[0072] Optionally, the image extraction network may include a deep learning network for extracting image features, such as a Vision Transformer (VIT), an imagebind, etc.
[0073] Optionally, the text extraction network may include a deep learning network for extracting text features, such as a Bidirectional Encoder Representations from Transformers (Bert) model, an imagebind, etc.
[0074] Figure 4 This is the second flowchart of the authenticity recognition method for the fused data provided by the embodiments of this application. As Figure 4 shown, step S203 of the authenticity recognition method for the fused data includes:
[0075] S2031: Extract the first vector of the first input data, and splice the first vector and the third input data to obtain the target input data;
[0076] S2032: Input the target input data into a pre-trained large language model, and use the pre-trained large language model to identify the authenticity of the second input data to obtain the detection result of the second input data.
[0077] Optionally, extracting the first vector of the first input data facilitates the pre-trained large language model to learn the feature information of the first input data, and splicing the first vector of the first input data and the third input data to obtain the target input data. Based on the target input data, the pre-trained large language model identifies the authenticity of the feature information in the second input data based on the feature information of the real fused first input data to obtain the detection result of the second input data.
[0078] It can be understood that the third input data includes each feature information in the second input data. The pre-trained large language model performs classification processing based on each feature information in the second input data to obtain the corresponding classification result, and the above classification result is the target detection result corresponding to the authenticity of the second input data. Converting information such as the third input data and the target detection result into corresponding text data can obtain the target recognition description data corresponding to the above target detection result.
[0079] In some embodiments, the first input data (fewshot) further includes at least one authenticity recognition result and the authenticity recognition description data corresponding to each authenticity recognition result.
[0080] Among them, the authenticity recognition result is used to indicate the authenticity of the data in the first input data (such as the first true fusion data or the non-true fusion data). For example, if the first input data includes a first true fusion data and two non-true fusion data, the authenticity recognition result includes a first authenticity recognition result for indicating that the first true fusion data is true, and two second authenticity recognition results for indicating that each non-true fusion data is false.
[0081] Optionally, the authenticity recognition result and the authenticity recognition description data corresponding to each authenticity recognition result may be in the form of labels.
[0082] Exemplarily, the authenticity recognition result is: a text label for explaining that the first true fusion data is a true text-image fusion data, and the corresponding authenticity recognition description data includes: a text label for explaining the reason why the first true fusion data is true.
[0083] For example, the authenticity recognition result of the first true fusion data is: the first true fusion data is true, and the corresponding authenticity recognition description data is: the edge of the image data in the first true fusion data is clear, the image content is the head portrait of person A, which matches the text data "the first true fusion data is a photo of person A", so the first true fusion data is true.
[0084] In some embodiments, the target input data is input into a pre-trained large language model, and the pre-trained large language model is used to identify the authenticity of the second input data, and the detection result of the second input data is obtained, including:
[0085] The target input data is input into a pre-trained large language model. Based on the authenticity recognition result, the authenticity recognition description data, and the first true fusion data and / or non-true fusion data, the pre-trained large language model is optimized, and based on the optimized large language model, the authenticity of the second input data is identified to obtain the detection result of the second input data.
[0086] Optionally, the first true fusion data and / or non-true fusion data, the authenticity recognition result, and the authenticity recognition description data are converted into corresponding vectors and spliced with the third input data to obtain the target input data.
[0087] Specifically, based on the feature vectors of the first true fusion data and / or false fusion data, the true / false recognition result, and the true / false recognition description data, the pre-trained large language model is optimized so that the pre-trained large language model learns the feature information of the true graphic-text fusion data and / or false graphic-text fusion data, the recognition results of the true graphic-text fusion data and / or false graphic-text fusion data, the reasons for determining the true graphic-text fusion data as true, and / or the reasons for determining the false graphic-text fusion data as false, thereby obtaining an optimized large language model. And based on the optimized large language model, the true / false of the second input data is recognized to obtain the target recognition result of the second input data and the corresponding target recognition description data.
[0088] Optionally, the first input data further includes guiding vocabulary for guiding the pre-trained large language model to learn the process of inferring the true / false of the second input data.
[0089] For example, the guiding vocabulary includes: first learn the feature information of the image data, and then learn the feature information of the text data. Or, the guiding vocabulary includes: think and learn the feature information one by one. To improve the accuracy of the recognition result of the pre-trained large language model.
[0090] Figure 5 It is a schematic application diagram for the true / false recognition of the fusion data provided by the embodiments of the present application.
[0091] See Figure 5 , the second input data, the first input data, and the guiding vocabulary are input into the true / false recognition model of the fusion data. The true / false recognition model of the fusion data includes a pre-trained feature processing network model and a pre-trained large language model. The feature information of the second input data is extracted through the pre-trained feature processing network model, and is concatenated with the first input data and the guiding vocabulary, and then input into the pre-trained large language model, thereby obtaining the detection result of the second input data.
[0092] Figure 6 It is the third schematic flowchart of the true / false recognition method of the fusion data provided by the embodiments of the present application. As Figure 6 shown, the true / false recognition method of the fusion data includes:
[0093] S301: Obtain a true fusion sample data set; the true fusion sample data set includes multiple second true fusion data; the second true fusion data includes second true image data, second true text data, and the first sample label of the second true text data;
[0094] S302: Preprocess the multiple second true image data to obtain multiple image training data;
[0095] S303: Adjust multiple second real text data to obtain multiple text training data, and add corresponding second sample labels to each text training data;
[0096] S304: Perform splicing processing on multiple image training data and corresponding text training data respectively to obtain an untrue sample data set; among them, the untrue sample data set includes multiple untrue sample data; each untrue sample data carries a second sample label corresponding to each text training data;
[0097] S305: Based on the sample data set, pre-train the initial feature processing network model and the initial large language model to obtain a pre-trained feature processing network model and a pre-trained large language model; among them, the sample data set includes a real fusion sample data set and an untrue sample data set; the initial feature processing network model is connected to the initial large language model.
[0098] Specifically, obtain a real fusion sample data set including multiple second real fusion data. The second real fusion data refers to the text-image fusion data that has been verified as real and is composed of second real image data and second real text data. Add corresponding first sample labels to the second real text data.
[0099] Among them, the first sample label includes reason text describing the authenticity of the second real fusion data.
[0100] Specifically, preprocess the second real image data in the real fusion sample data set to obtain corresponding multiple image training data. The preprocessing includes, but is not limited to, processing methods for changing feature information such as the edges, textures, and colors of the second real image data.
[0101] For example, if the second real image data includes a face, the face in the second real image data can be replaced, or the facial expression of the face in the second real image data can be adjusted, the sizes of the facial features can be adjusted, and accessories such as masks and beards can be added.
[0102] Specifically, adjust the second real text data in the real fusion sample data set, change the content or semantics of the second real text data, and add corresponding second sample labels to the second real text data based on the above image preprocessing operations or text adjustment processing operations to obtain multiple text training data. The second sample label is used to describe the changes in the second real text data, facilitating the initial large language model to learn the false feature information of the image training data and text training data, and reasoning about the reasons or explanations for the authenticity of the image training data and text training data.
[0103] For example, replace the location in the second real text data, or modify the semantic emotion "happy" in the second real text data to "angry".
[0104] For example, if the face in the second real image data corresponding to the second real text data is replaced, a label can be added to the second real text data: There is a forgery operation of face replacement in the image data, so the content of the image data does not match the text data.
[0105] Specifically, multiple image training data are respectively spliced with the corresponding text training data (each text training data carries the corresponding second sample label) to obtain an untrue sample data set. Based on this, a sample data set including a true fusion sample data set and an untrue sample data set is obtained.
[0106] Specifically, based on the sample data set including the true fusion sample data set and the untrue sample data set, the initial feature processing network model and the initial large language model are pre-trained to obtain a pre-trained feature processing network model and a pre-trained large language model.
[0107] Forgery is performed according to the features of the image data and the text data, and an explanatory label is added to each sample data, so that the sample data set covers various forgery methods, such as image synthesis, the text description does not match the image content, etc., so that the pre-trained large prediction model can learn the essential patterns and features of the text-image fusion data, as well as the reasons for judging authenticity, has generalization ability, increases the adaptability and robustness of the model, and improves the accuracy of the authenticity recognition result of the text-image fusion data.
[0108] In some embodiments, based on the sample data set, pre-training the initial feature processing network model and the initial large language model to obtain a pre-trained feature processing network model and a pre-trained large language model, including:
[0109] Input the sample data in the sample data set into the initial feature processing network model for pre-training, so that the initial feature processing network model extracts the training feature vectors of the sample data to obtain a pre-trained feature processing network model; wherein, the sample data includes second real fusion data and untrue sample data;
[0110] Input the training feature vectors into the initial large language model for pre-training to obtain a pre-trained large language model.
[0111] Optionally, input the sample data in the sample dataset into the initial feature processing network model, and pre-train the initial feature processing network model based on each sample data, so that the initial feature processing network model learns to extract the training feature vectors of the sample data during the training process, thereby obtaining a pre-trained feature processing network model. And input the training feature vectors extracted during the training process into the initial large language model, and pre-train the initial large language model based on each training feature vector, so that the initial large language model learns to identify the authenticity of the sample data during the training process, thereby obtaining a pre-trained large language model.
[0112] In some embodiments, the initial feature processing network model includes an initial image feature extraction network, an initial text feature extraction network, and an initial feature fusion network. The initial text feature extraction network includes a first text feature extraction network and a second text feature extraction network. The initial feature fusion network includes a first sub-fusion network and a second sub-fusion network. The input end of the first sub-fusion network is connected to the output ends of the first text feature extraction network and the initial image feature extraction network. The input end of the second sub-fusion network is connected to the output end of the first sub-fusion network and the output end of the second text feature extraction network. The output end of the second sub-fusion network is connected to the input end of the initial large language model.
[0113] Figure 7 It is a schematic structural diagram of the initial feature processing network model provided by the embodiments of the present application.
[0114] See Figure 7 , the output ends of the initial image feature extraction network 7011 and the initial text feature extraction network 7012 of the initial feature processing network model 701 are both connected to the input end of the initial feature fusion network 7013. The initial text feature extraction network 7012 includes a first text feature extraction network 70121 and a second text feature extraction network 70122. The initial feature fusion network 7013 includes a first sub-fusion network 70131 and a second sub-fusion network 70132.
[0115] Among them, the output end of the initial image feature extraction network 7011, the output end of the first text feature extraction network 70121 are connected to the input end of the first sub-fusion network 70131. The output end of the second text feature extraction network 70122, the output end of the first sub-fusion network 70131 are connected to the input end of the second sub-fusion network 70132. The output end of the second sub-fusion network 70132 is connected to the input end of the initial large language model 702.
[0116] It can be understood that the pre-trained feature processing network model also includes a first text feature extraction network, a second text feature extraction network, a first sub-fusion network, and a second sub-fusion network.
[0117] Optionally, the second input data does not carry literal information (such as literal data in the form of a label) for indicating the authenticity of the second input data, and the pre-trained feature processing network model may not perform the operation of feature extraction using the second text feature extraction network.
[0118] In some embodiments, inputting the sample data in the sample dataset into the initial feature processing network model for pre-training, so that the initial feature processing network model extracts the training feature vectors of the sample data, to obtain the pre-trained feature processing network model, includes:
[0119] Inputting the literal samples in the sample data into the first text feature extraction network, and using the first text feature extraction network to extract the second text feature vectors of the literal samples; the literal samples include second true literal data and literal training data;
[0120] Inputting the sample labels in the sample data into the second text feature extraction network, and using the second text feature extraction network to extract the third text feature vectors of the sample labels; the sample labels include first sample labels and second sample labels;
[0121] Inputting the image samples in the sample data into the initial image feature extraction network, and using the initial image feature extraction network to extract the second image feature vectors of the image samples; the image samples include second true image data and image training data;
[0122] Inputting the second text feature vectors and the second image feature vectors into the first sub-fusion network, and using the first sub-fusion network to perform fusion processing on the second text feature vectors and the second image feature vectors, to obtain first fusion data;
[0123] Inputting the first fusion data and the third text feature vectors into the second sub-fusion network, and using the second sub-fusion network to perform fusion processing on the first fusion data and the third text feature vectors, to obtain the training feature vectors of the sample data, and the pre-trained feature processing network model.
[0124] Optionally, the sample data in the sample set includes second true fusion data and non-true sample data, the literal samples in the sample data include second true literal data and literal training data; the image samples in the sample data include second true image data and image training data; the sample labels in the sample data include first sample labels and second sample labels.
[0125] Optionally, input the text samples in the sample data into the first text feature extraction network, input the sample labels in the sample data into the second text feature extraction network, input the image samples in the sample data into the initial image feature extraction network, and use the first text feature extraction network to extract the second text feature vector of the text samples, use the second text feature extraction network to extract the third text feature vector of the sample labels, and use the initial image feature extraction network to extract the second image feature vector of the image samples.
[0126] Optionally, perform feature fusion processing on the feature vectors of the sample data in a staged fusion manner: use the first sub-fusion network to fuse the second text feature vector and the second image feature vector, so that the initial large language model can understand the correlation between the text data and the image data in the sample data, and use the second sub-fusion network to fuse the first fusion data and the third text feature vector, so that the initial large language model can understand the correlation between the information in the sample label and the text samples and image samples, so as to determine the corresponding explanatory data when inferring the authenticity of the sample data.
[0127] Using a feature fusion network to fuse the feature information of text samples and image samples can more comprehensively capture the features and patterns of forged information in the sample data, improve the accuracy of the recognition results, and reduce the probability of misjudgment and missed judgment.
[0128] Optionally, use the cross-attention mechanism calculation method to determine the correlation between the feature vectors.
[0129] For example, use the first sub-fusion network to calculate the third cross-attention vector between the second image feature vector and the second text feature vector, and use the third cross-attention direction as the first fusion data.
[0130] Optionally, the third cross-attention vector can be represented by the following formula:
[0131] F1 = CA(Q = img, K = txt, V = txt);
[0132] Where CA represents cross-attention, the second text feature vector includes the key (K = txt) and the value (V = txt), and Q = img represents the second image feature vector.
[0133] Among them, the third cross-attention vector is used to indicate the correlation between the image sample and the text sample in the sample data, or the influence of the image sample on the text sample, or the image sample guiding the initial large language model to learn the feature information in the text sample.
[0134] In some embodiments, a second sub-fusion network is used to fuse the first fusion data and the third text feature vector to obtain a training feature vector of the sample data, including:
[0135] Using the second sub-fusion network, based on the first fusion data and the third text feature vector, determine a first cross-attention vector and a second cross-attention vector; wherein, the first cross-attention vector is used to indicate the attention value of the sample label to the sample data; the second cross-attention vector is used to indicate the attention value of the sample data to the sample label;
[0136] Concatenate the first cross-attention vector and the second cross-attention vector to obtain the training feature vector of the sample data.
[0137] Specifically, the first sub-fusion network inputs the first fusion data into the second sub-fusion network, and the second text feature extraction network inputs the third text feature vector of the sample label into the second sub-fusion network. The second sub-fusion network is used to determine the first cross-attention vector and the second cross-attention vector, and concatenate the first cross-attention vector and the second cross-attention vector to obtain the training feature vector.
[0138] Optionally, a cross-attention mechanism calculation method is used to determine the first cross-attention vector and the second cross-attention vector.
[0139] Optionally, the first cross-attention vector can be represented by the following formula:
[0140] F2 = CA(Q = rsn, K = v1, V = v1);
[0141] Wherein, Q = rsn is used to represent the third text feature vector, and the first fusion data includes a key (K = v1) and a value (V = v1). The first cross-attention vector can be used to indicate the attention value of the sample label to the sample data, or the influence of the sample label on the sample data, or the sample label guiding the initial large language model to learn the feature information in the sample data.
[0142] Optionally, the second cross-attention vector can be represented by the following formula:
[0143] F3 = (Q = v1, K = rsn, V = rsn);
[0144] Wherein, Q = v1 represents the first fusion data, and the third text feature vector includes a key (K = rsn) and a value (V = rsn). The second cross-attention vector is used to indicate the attention value of the sample data to the sample label, or the influence of the sample data on the sample label, or the sample data guiding the initial large language model to infer reasonable explanatory data.
[0145] In some embodiments, based on a sample data set, pre-training is performed on an initial feature processing network model and an initial large language model to obtain a pre-trained feature processing network model and a pre-trained large language model, including:
[0146] Based on the sample data set, the initial feature processing network model and the initial large language model are pre-trained using a loss function to obtain a pre-trained feature processing network model and a pre-trained large language model. The loss function is used to determine the negative correlation degree between the detection result of the initial large language model and each sample data in the sample data set; the loss function includes a first correlation degree, a second correlation degree, and a third correlation degree;
[0147] Among them, the first correlation degree is used to determine the negative correlation degree between each sample label in the sample data set and the predicted label of the initial large language model corresponding to each sample label; the second correlation degree is used to determine the negative correlation degree between each sample label in the sample data set and the recognition result output by the initial large language model corresponding to each sample label; the third correlation degree is used to determine the negative correlation degree between each sample label in the sample data set and the recognition description data output by the initial large language model corresponding to each sample label.
[0148] It can be understood that the initial feature processing network model and the initial large language model can be pre-trained separately. The initial large language model and the initial feature processing network model can also be used as an overall model for pre-training.
[0149] In this application, the pre-training process of taking the initial large language model and the initial feature processing network model as an overall model is used as an example for illustration, which does not limit the pre-training process of the initial large language model and the initial feature processing network model.
[0150] Optionally, the sample data set includes a real fusion sample data set and a non-real sample data set, and the sample data includes multiple second real fusion data and multiple non-real sample data. The sample labels include a first sample label and a second sample label.
[0151] Specifically, each sample data in the sample data set is input into the initial feature processing network and the initial large language model for pre-training, and the corresponding loss function value is determined. Based on the loss function value, reverse transmission training is performed on the initial feature processing network and the initial large language model. When the loss function value reaches a preset threshold, it is determined that the pre-training process ends, and a pre-trained feature processing network model and a pre-trained large language model are obtained.
[0152] Among them, the loss function includes a first correlation degree, a second correlation degree, and a third correlation degree. The loss function determined based on the first correlation degree, the second correlation degree, and the third correlation degree can be used to represent the negative correlation degree between the detection result output by the initial large language model and each sample data in the sample data set.
[0153] Among them, the preset threshold can be specifically set according to the actual situation. It can be understood that the smaller the loss function value, the greater the correlation between the output detection result and each sample data in the sample dataset, that is, the higher the authenticity of the detection result, or the higher the accuracy of the detection result.
[0154] It can be understood that the pre-trained large language model can convert the feature information of the second input data into corresponding text data to obtain the target recognition description data for the authenticity of the second input data.
[0155] In some embodiments, the first correlation degree is determined based on each sample label and the difference degree between the prediction labels output by the initial large language model and corresponding to each sample label during the pre-training process of the initial large language model.
[0156] Specifically, the first correlation degree is used to indicate the negative correlation degree between each true sample label (including the first sample label and the second sample label) of each sample data in the sample dataset and the prediction label output by the initial large language model and corresponding to each of the above true sample labels.
[0157] It can be understood that the smaller the above first correlation degree, the higher the matching degree between the prediction label output by the initial large language model and the corresponding true sample label.
[0158] Optionally, the first correlation degree can be expressed by the following formula:
[0159] ;
[0160] where N represents the number of sample data, represents the true sample label value of the i-th sample data; when the i-th sample data is true graphic-text fusion data, it is marked as 1, otherwise, when the i-th sample data is false graphic-text fusion data, it is marked as 0. represents the prediction label of the i-th sample data output by the initial large language model.
[0161] In some embodiments, the second correlation degree is determined based on each sample label and the authenticity probability of the recognition result output by the initial large language model and corresponding to each sample label during the pre-training process of the initial large language model.
[0162] Specifically, the second correlation degree is used to indicate the negative correlation degree between each true sample label of each sample data in the sample dataset and the authenticity probability of the recognition result output by the initial large language model and corresponding to each sample label.
[0163] It is understandable that the smaller the second relevance, the higher the probability of the authenticity of the predicted label output by the initial large language model.
[0164] Optionally, the second relevance can be expressed by the following formula:
[0165] ;
[0166] ;
[0167] where Q represents the image training data, rsn represents the sample label; Avgpool() represents average pooling processing; MLP() represents a multi-layer perceptron; sigmoid represents an activation function, represents the predicted probability of the sample label; represents the first fusion data, K represents the key value of the first fusion data; V represents the value corresponding to the key value K. represents the authenticity of the recognition result output by the initial large language model. If the recognition result output by the initial large language model is correct, it is marked as 1, otherwise, if the recognition result output by the initial large language model is incorrect, it is marked as 0. That is, if the label indicates that the sample data is real image-text fusion data and the recognition result output by the initial large language model is also true, it is marked as 1.
[0168] In some embodiments, the third relevance is determined based on each sample label and the similarity probability between the recognition description data output by the initial large language model and corresponding to each sample label during the pre-training process of the initial large language model.
[0169] Specifically, the third relevance is used to indicate the negative correlation between the similarity probabilities between each true sample label of each sample data in the sample dataset and the recognition description data output by the initial large language model and corresponding to the above-mentioned each sample label. The smaller the third relevance, the higher the similarity probability between the recognition description data output by the initial large language model and the corresponding true sample label.
[0170] Optionally, the third relevance can be expressed by the following formula:
[0171] ;
[0172] where, represents the predicted probability of the recognition description data output by the initial large language model, is the correctness label of the recognition description data output by the initial large language model. If the recognition description data output by the initial large language model is consistent with the sample label, it is marked as 1, and if the recognition description data output by the initial large language model is inconsistent with the sample label, it is marked as 0.
[0173] Correspondingly, the loss function can be expressed by the following formula:
[0174] ;
[0175] Wherein, represents a hyperparameter, represents a hyperparameter.
[0176] The loss function includes measuring the difference between the detection result and the true sample label, and introducing the association between the description data and the fused feature information, so that when the pre-trained large prediction model accurately identifies the authenticity of the text-image fused data, corresponding explanatory data can be generated.
[0177] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases the former is a better implementation method.
[0178] Figure 8 This is a schematic structural diagram of the authenticity recognition device for fused data provided by the embodiments of the present application. As Figure 8 shown, the embodiments of the present application also provide an authenticity recognition device for fused data, including:
[0179] A data acquisition module 801, configured to acquire first input data and second input data; wherein, the second input data includes image data and text data; the first input data includes at least one first true fused data and / or at least one non-true fused data;
[0180] A feature extraction module 802, configured to input the second input data into a pre-trained feature processing network model, and extract the feature vector of the second input data by using the pre-trained feature processing network model to obtain third input data;
[0181] An identification module 803, configured to input the first input data and the third input data into a pre-trained large language model, and identify the authenticity of the second input data by using the pre-trained large language model to obtain the detection result of the second input data; wherein, the detection result includes a target recognition result and target recognition explanatory data corresponding to the target recognition result, and the target recognition result is used to indicate the authenticity of the second input data; the pre-trained feature processing network model is connected to the pre-trained large language model.
[0182] In some embodiments, the pre-trained feature processing network model includes an image feature extraction network, a text feature extraction network, and a feature fusion network, and the image feature extraction network and the text feature extraction network are both connected to the feature fusion network; the feature extraction module includes:
[0183] The first extraction sub-module is used to input image data into an image feature extraction network and extract a first image feature vector of the image data using the image feature extraction network;
[0184] The second extraction sub-module is used to input text data into a text feature extraction network and extract a first text feature vector of the text data using the text feature extraction network;
[0185] The fusion processing sub-module is used to input the first text feature vector and the first image feature vector into a feature fusion network and perform fusion processing on the first text feature vector and the first image feature vector using the feature fusion network to obtain third input data.
[0186] In some embodiments, the recognition module includes:
[0187] The vector extraction sub-module is used to extract a first vector of the first input data and splice the first vector and the third input data to obtain target input data;
[0188] The recognition sub-module is used to input the target input data into a pre-trained large language model and use the pre-trained large language model to identify the authenticity of the second input data to obtain a detection result of the second input data.
[0189] In some embodiments, the first input data further includes at least one authenticity recognition result and authenticity recognition description data corresponding to each authenticity recognition result; the recognition unit is specifically used for:
[0190] Input the target input data into a pre-trained large language model, optimize the pre-trained large language model based on the authenticity recognition result, the authenticity recognition description data, and the first real fusion data and / or non-real fusion data, and identify the authenticity of the second input data based on the optimized large language model to obtain a detection result of the second input data.
[0191] In some embodiments, the device further includes:
[0192] The first acquisition module is used to acquire a real fusion sample data set; the real fusion sample data set includes a plurality of second real fusion data; the second real fusion data includes second real image data, second real text data, and a first sample label of the second real text data;
[0193] The first preprocessing module is used to preprocess a plurality of second real image data to obtain a plurality of image training data;
[0194] The second preprocessing module is used to adjust a plurality of second real text data to obtain a plurality of text training data and add corresponding second sample labels to each text training data;
[0195] A third preprocessing module, configured to splice multiple image training data with corresponding text training data respectively to obtain an untrue sample data set; wherein, the untrue sample data set includes multiple untrue sample data; each untrue sample data carries a second sample label corresponding to each text training data.
[0196] A pre-training module, configured to pre-train an initial feature processing network model and an initial large language model based on the sample data set to obtain a pre-trained feature processing network model and a pre-trained large language model; wherein, the sample data set includes a true fusion sample data set and an untrue sample data set; the initial feature processing network model is connected to the initial large language model.
[0197] In some embodiments, the pre-training module includes:
[0198] A first training sub-module, configured to input the sample data in the sample data set into the initial feature processing network model for pre-training, so that the initial feature processing network model extracts the training feature vectors of the sample data to obtain a pre-trained feature processing network model; wherein, the sample data includes second true fusion data and untrue sample data.
[0199] A second training sub-module, configured to input the training feature vectors into the initial large language model for pre-training to obtain a pre-trained large language model.
[0200] In some embodiments, the initial feature processing network model includes an initial image feature extraction network, an initial text feature extraction network, and an initial feature fusion network. The initial text feature extraction network includes a first text feature extraction network and a second text feature extraction network. The initial feature fusion network includes a first sub-fusion network and a second sub-fusion network. The input end of the first sub-fusion network is connected to the output end of the first text feature extraction network and the output end of the initial image feature extraction network. The input end of the second sub-fusion network is connected to the output end of the first sub-fusion network and the output end of the second text feature extraction network. The output end of the second sub-fusion network is connected to the input end of the initial large language model. The second training sub-module includes:
[0201] A first extraction unit, configured to input the text samples in the sample data into the first text feature extraction network and use the first text feature extraction network to extract the second text feature vectors of the text samples; the text samples include second true text data and text training data.
[0202] A second extraction unit, configured to input the sample labels in the sample data into the second text feature extraction network and use the second text feature extraction network to extract the third text feature vectors of the sample labels; the sample labels include first sample labels and second sample labels.
[0203] A third extraction unit for inputting the image samples in the sample data into an initial image feature extraction network and using the initial image feature extraction network to extract a second image feature vector of the image samples; the image samples include second real image data and image training data;
[0204] A first fusion unit for inputting the second text feature vector and the second image feature vector into a first sub-fusion network and using the first sub-fusion network to perform a fusion process on the second text feature vector and the second image feature vector to obtain first fusion data;
[0205] A second fusion unit for inputting the first fusion data and the third text feature vector into a second sub-fusion network and using the second sub-fusion network to perform a fusion process on the first fusion data and the third text feature vector to obtain a training feature vector of the sample data and a pre-trained feature processing network model.
[0206] In some embodiments, the second fusion unit includes:
[0207] A first determination subunit for using the second sub-fusion network to determine a first cross-attention vector and a second cross-attention vector based on the first fusion data and the third text feature vector; wherein, the first cross-attention vector is used to indicate the attention value of the sample label to the sample data; the second cross-attention vector is used to indicate the attention value of the sample data to the sample label;
[0208] A second determination subunit for concatenating the first cross-attention vector and the second cross-attention vector to obtain a training feature vector of the sample data.
[0209] In some embodiments, the pre-training module is specifically configured to:
[0210] Based on the sample data set, pre-train the initial feature processing network model and the initial large language model using a loss function to obtain a pre-trained feature processing network model and a pre-trained large language model, and the loss function is used to determine the negative correlation degree between the detection result of the initial large language model and each sample data in the sample data set; the loss function includes a first correlation degree, a second correlation degree, and a third correlation degree;
[0211] Among them, the sample data includes second true fusion data and untrue sample data; the first correlation degree is used to determine the negative correlation degree between each sample label in the sample data set and the predicted label of the initial large language model corresponding to each sample label. The sample labels include the first sample label and the second sample label; the second correlation degree is used to determine the negative correlation degree between each sample label in the sample data set and the recognition result output by the initial large language model corresponding to each sample label; the third correlation degree is used to determine the negative correlation degree between each sample label in the sample data set and the recognition description data output by the initial large language model corresponding to each sample label.
[0212] In some embodiments, the first correlation degree is determined based on each sample label and the difference degree between the predicted label output by the initial large language model and corresponding to each sample label during the pre-training process of the initial large language model.
[0213] In some embodiments, the second correlation degree is determined based on each sample label and the authenticity probability of the recognition result output by the initial large language model and corresponding to each sample label during the pre-training process of the initial large language model.
[0214] In some embodiments, the third correlation degree is determined based on each sample label and the similarity probability between the recognition description data output by the initial large language model and corresponding to each sample label during the pre-training process of the initial large language model.
[0215] For the description of the features in the corresponding embodiments of the authenticity recognition device for fusion data, reference can be made to the relevant descriptions in the corresponding embodiments of the authenticity recognition method for fusion data, which will not be elaborated here one by one.
[0216] Figure 9 The following is a schematic structural diagram of the electronic device provided by this application. As Figure 9 shown, the electronic device 90 provided in this embodiment includes: at least one processor 901 and a memory 902. Optionally, the device 90 further includes a communication component 903. Among them, the processor 901, the memory 902, and the communication component 903 are connected through a bus.
[0217] In a specific implementation process, at least one processor 901 executes the computer execution instructions stored in the memory 902, so that at least one processor 901 executes the above-mentioned embodiment of the authenticity recognition method for fusion data.
[0218] For the specific implementation process of the processor 901, reference can be made to the above method embodiment, and its implementation principle and technical effect are similar, which will not be elaborated here in this embodiment.
[0219] In the above embodiments, it should be understood that the processor may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), etc. The general-purpose processor may be a microprocessor or any conventional processor, etc. The steps of the method disclosed in combination with the application can be directly implemented by the execution of the hardware processor, or by the combination of the hardware and software modules in the processor.
[0220] The memory may include a random access memory (RAM), and may also include a non-volatile memory (NVM), such as at least one disk memory.
[0221] The bus may be an industry standard architecture (ISA) bus, a peripheral component interconnect (PCI) bus, an extended industry standard architecture (EISA) bus, etc. The bus may be divided into an address bus, a data bus, a control bus, etc. For the sake of convenience of representation, the bus in the drawings of this application is not limited to only one bus or one type of bus.
[0222] The embodiments of the present application further provide a computer-readable storage medium, in which a computer program is stored, and the computer program is configured to execute the steps in any of the above embodiments of the method for identifying the authenticity of fused data when running.
[0223] In an exemplary embodiment, the above computer-readable storage medium may include, but is not limited to: USB flash drives, read-only memories (ROMs), random access memories (RAMs), mobile hard disks, magnetic disks, or optical discs and other media that can store computer programs.
[0224] The embodiments of the present application further provide a computer program product, the above computer program product includes a computer program, and when the computer program is executed by a processor, the steps in any of the above embodiments of the method for identifying the authenticity of fused data are implemented.
[0225] Embodiments of the present application also provide another computer program product, including a non-volatile computer-readable storage medium storing a computer program, where the computer program, when executed by a processor, implements the steps in any of the above-described embodiments of the method for authenticating the authenticity of fused data.
[0226] Those skilled in the art can further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the components and steps of each example have been generally described according to their functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Skilled professionals can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.
[0227] The above has introduced in detail a method, device, storage medium, and program product for authenticating the authenticity of fused data provided by the present application. Specific examples are used herein to elaborate on the principle and implementation manner of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application. It should be noted that for those of ordinary skill in the art in the technical field, without departing from the principle of the present application, several improvements and modifications can be made to the present application, and these improvements and modifications also fall within the protection scope of the claims of the present application.
Claims
1. A method for identifying the authenticity of fused data, characterized in that: include: Collecting first input data and second input data; wherein the first input data includes image data and text data, and the second input data includes image data and text data; the first input data includes at least one first real fusion data and / or at least one non-real fusion data; Inputting the second input data into a pre-trained feature processing network model, and using the pre-trained feature processing network model to extract a feature vector of the second input data to obtain third input data; Inputting the first input data and the third input data into a pre-trained large language model, and using the pre-trained large language model to identify the authenticity of the second input data, to obtain a detection result of the second input data; wherein the detection result includes a target recognition result and target recognition description data corresponding to the target recognition result, and the target recognition result is used to indicate the authenticity of the second input data; the pre-trained feature processing network model is connected to the pre-trained large language model; The pre-trained feature processing network model includes an image feature extraction network, a text feature extraction network and a feature fusion network, wherein the image feature extraction network and the text feature extraction network are both connected to the feature fusion network, wherein the image feature extraction network is used to extract a first image feature vector of the image data, the text feature extraction network is used to extract a first text feature vector of the text data, and the feature fusion network is used to fuse the first text feature vector and the first image feature vector.
2. The authenticity identification method of fused data according to claim 1 is characterized in that: The step of inputting the second input data into a pre-trained feature processing network model, and extracting a feature vector of the second input data using the pre-trained feature processing network model to obtain third input data includes: Inputting the image data into the image feature extraction network, and using the image feature extraction network to extract a first image feature vector of the image data; Inputting the text data into the text feature extraction network, and using the text feature extraction network to extract a first text feature vector of the text data; The first text feature vector and the first image feature vector are input into the feature fusion network, and the first text feature vector and the first image feature vector are fused by the feature fusion network to obtain the third input data.
3. The authenticity identification method of fused data according to claim 1 is characterized in that: The step of inputting the first input data and the third input data into a pre-trained large language model, and using the pre-trained large language model to identify the authenticity of the second input data to obtain a detection result of the second input data includes: Extracting a first vector of the first input data, and concatenating the first vector with the third input data to obtain target input data; The target input data is input into the pre-trained large language model, and the pre-trained large language model is used to identify the authenticity of the second input data to obtain a detection result of the second input data.
4. The authenticity identification method of fused data according to claim 3 is characterized in that: The first input data also includes at least one authenticity recognition result and authenticity recognition description data corresponding to each authenticity recognition result; the step of inputting the target input data into the pre-trained large language model, using the pre-trained large language model to recognize the authenticity of the second input data, and obtaining a detection result of the second input data includes: The target input data is input into the pre-trained large language model, and based on the authenticity recognition result, the authenticity recognition explanation data, and the first real fusion data and / or the non-real fusion data, the pre-trained large language model is optimized, and the authenticity of the second input data is identified based on the optimized large language model to obtain a detection result of the second input data.
5. The method for identifying the authenticity of fused data according to any one of claims 1 to 4, characterized in that: The method further comprises: Acquire a real fusion sample data set; the real fusion sample data set includes a plurality of second real fusion data; the second real fusion data includes second real image data, second real text data and a first sample label of the second real text data; Preprocessing the plurality of second real image data to obtain a plurality of image training data; Adjusting the plurality of second real text data to obtain a plurality of text training data, and adding a corresponding second sample label to each of the text training data; The plurality of image training data are respectively concatenated with the corresponding text training data to obtain a non-real sample data set; wherein the non-real sample data set includes a plurality of non-real sample data; each of the non-real sample data carries a second sample label corresponding to each text training data; Based on the sample data set, the initial feature processing network model and the initial large language model are pre-trained to obtain the pre-trained feature processing network model and the pre-trained large language model; wherein the sample data set includes the real fusion sample data set and the non-real sample data set; the initial feature processing network model is connected to the initial large language model.
6. The authenticity identification method of fused data according to claim 5 is characterized in that: The pre-training of the initial feature processing network model and the initial large language model based on the sample data set to obtain the pre-trained feature processing network model and the pre-trained large language model includes: Inputting the sample data in the sample data set into the initial feature processing network model for pre-training, so that the initial feature processing network model extracts the training feature vector of the sample data to obtain a pre-trained feature processing network model; wherein the sample data includes the second real fusion data and the non-real sample data; The training feature vector is input into the initial large language model for pre-training to obtain the pre-trained large language model.
7. The authenticity identification method of fused data according to claim 6 is characterized in that: The initial feature processing network model includes an initial image feature extraction network, an initial text feature extraction network and an initial feature fusion network. The initial text feature extraction network includes a first text feature extraction network and a second text feature extraction network. The initial feature fusion network includes a first sub-fusion network and a second sub-fusion network. The input end of the first sub-fusion network is connected to the output end of the first text feature extraction network and the output end of the initial image feature extraction network. The input end of the second sub-fusion network is connected to the output end of the first sub-fusion network and the output end of the second text feature extraction network. The output end of the second sub-fusion network is connected to the input end of the initial large language model. The sample data in the sample data set is input into the initial feature processing network model for pre-training so that the initial feature processing network model extracts the training feature vector of the sample data to obtain the pre-trained feature processing network model, including: Inputting a text sample in the sample data into the first text feature extraction network, and using the first text feature extraction network to extract a second text feature vector of the text sample; the text sample includes the second real text data and the text training data; Inputting the sample labels in the sample data into the second text feature extraction network, and using the second text feature extraction network to extract a third text feature vector of the sample labels; the sample labels include the first sample labels and the second sample labels; Inputting an image sample in the sample data into the initial image feature extraction network, and using the initial image feature extraction network to extract a second image feature vector of the image sample; the image sample includes the second real image data and the image training data; Inputting the second text feature vector and the second image feature vector into the first sub-fusion network, and using the first sub-fusion network to fuse the second text feature vector and the second image feature vector to obtain first fusion data; The first fused data and the third text feature vector are input into the second sub-fusion network, and the second sub-fusion network is used to fuse the first fused data and the third text feature vector to obtain the training feature vector of the sample data and the pre-trained feature processing network model.
8. The authenticity identification method of fused data according to claim 7 is characterized in that: The step of using the second sub-fusion network to fuse the first fused data and the third text feature vector to obtain the training feature vector of the sample data includes: The second sub-fusion network is used to determine a first cross-attention vector and a second cross-attention vector based on the first fusion data and the third text feature vector; wherein the first cross-attention vector is used to indicate the attention value of the sample label to the sample data; and the second cross-attention vector is used to indicate the attention value of the sample data to the sample label; The first cross-attention vector and the second cross-attention vector are concatenated to obtain a training feature vector of the sample data.
9. The authenticity identification method of fused data according to claim 5, characterized in that: The pre-training of the initial feature processing network model and the initial large language model based on the sample data set to obtain the pre-trained feature processing network model and the pre-trained large language model includes: Based on the sample data set, an initial feature processing network model and an initial large language model are pre-trained using a loss function to obtain the pre-trained feature processing network model and the pre-trained large language model, wherein the loss function is used to determine a negative correlation between a detection result of the initial large language model and each sample data in the sample data set; the loss function includes a first correlation, a second correlation, and a third correlation; Among them, the first correlation is used to determine the negative correlation between each sample label in the sample data set and the predicted label of the initial large language model corresponding to each sample label; the second correlation is used to determine the negative correlation between each sample label in the sample data set and the recognition result output by the initial large language model corresponding to each sample label; the third correlation is used to determine the negative correlation between each sample label in the sample data set and the recognition description data output by the initial large language model corresponding to each sample label.
10. The authenticity identification method of fused data according to claim 9, characterized in that: The first relevance is determined based on the sample labels and the difference between the predicted labels corresponding to the sample labels output by the initial large language model during pre-training of the initial large language model.
11. The authenticity identification method of fused data according to claim 9, characterized in that: The second relevance is determined based on each of the sample labels and the authenticity probability of the recognition results corresponding to each of the sample labels output by the initial large language model during pre-training of the initial large language model.
12. The authenticity identification method of fused data according to claim 9, characterized in that: The third correlation is determined based on the sample labels and the similarity probability between the recognition description data corresponding to the sample labels output by the initial large language model during the pre-training of the initial large language model.
13. An electronic device, characterized in that: include: Memory for storing computer programs; A processor, configured to implement the steps of the method for identifying the authenticity of fused data as claimed in any one of claims 1 to 12 when executing the computer program.
14. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, wherein the computer program, when executed by a processor, implements the steps of the method for identifying the authenticity of fused data according to any one of claims 1 to 12.
15. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method for identifying the authenticity of fused data as claimed in any one of claims 1 to 12 are implemented.
Citation Information
Patent Citations
Bimodal forged information detection method based on large language model
CN118171235A
Social network false message detection method based on large model and multi-modal fusion
CN119475066A