Unified discrimination method, device and computer readable medium for synthetic data
Patent Information
- Application Number
- CN202410449686.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-04-15
- Publication Date
- 2026-09-04
- Estimated Expiration
- 2044-04-15
AI Technical Summary
[0004]本申请的一个目的是提供一种合成数据的统一辨别方法、设备及计算机可读介质,用以解决现有方案适用范围不够广泛、灵活性不足的问题
[0015] Compared to existing technologies, this application provides a unified identification scheme for synthetic data. This scheme requires constructing model input information based on target data. The target data includes at least one type of modal data. The model input information includes an overall identifier character, modal data from the target data, and modal identifier characters. The model input information is input into a large language model. Features of each modal data are calculated and fused to obtain model output information. The model output information includes overall feature information that fuses all modal data features and modal feature information representing the features of each modal data. The modal features representing the features of each modal data are then processed. Feature information is input to the encoder of the converter model. The features of each modal data are fused using a self-attention mechanism to obtain modality-encoded feature information, which incorporates attention information between the features of each modality data. The overall feature information, which incorporates all modality data features, is then concatenated with the modality-encoded feature information and input to the decoder of the converter model. The overall feature information and the modality-encoded feature information are fused to obtain discrimination feature information. A preset activation function is then used to calculate the discrimination feature information to obtain a discrimination score, and the discrimination score is used to determine whether the target data is synthetic data. Since this scheme does not require limiting the modality of the target data to be identified, regardless of the types of modality data included in the target data, a unified model input information can be automatically constructed. After extracting and fusing feature information from different modal data through a large language model and a converter model, the discrimination feature information used for judgment is finally obtained, thus completing the process of determining whether the target data is synthetic data. The entire process is not limited by the modality of the target data, therefore it has a wide range of applications and high flexibility.
Smart Images

Figure CN118349945B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of information technology, and in particular to a unified method, apparatus and computer-readable medium for identifying synthetic data. Background Technology
[0002] Generative Artificial Intelligence (CAI) technology refers to the use of complex algorithms, models, and rules to learn from large-scale datasets and generate synthetic data that resembles real data, including images, text, audio, and video. As CAI technology continues to develop, the synthetic data it generates is becoming increasingly similar to real data, making it difficult for ordinary users to distinguish between synthetic and real data.
[0003] While some schemes for identifying synthetic data have emerged, their application is often limited. Most schemes only target a single modality of data; for example, they can only identify whether a single image or audio segment is synthetic. They cannot employ a unified approach to identify data containing arbitrary modalities. Therefore, their applicability and flexibility are insufficient. Summary of the Invention
[0004] One objective of this application is to provide a unified method, apparatus, and computer-readable medium for identifying synthetic data, in order to address the problems of insufficient applicability and flexibility of existing solutions.
[0005] To achieve the above objectives, embodiments of this application provide a unified identification method for synthetic data, the method comprising: Model input information is constructed based on target data, wherein the target data includes at least one type of modal data, and the model input information includes overall identifier characters, modal data in the target data, and modal identifier characters; The model input information is input into a large language model, and the features of each modal data are calculated and fused to obtain model output information. The model output information includes overall feature information that fuses the features of all modal data and modal feature information that represents the features of each modal data respectively. Modal feature information representing the features of each modal data is input into the encoder of the converter model. The features of each modal data are fused and calculated using a self-attention mechanism to obtain modal coding feature information. The modal coding feature information is fused with the attention information between the features of each modal data. The overall feature information, which integrates all modal data features, is concatenated with the modal coding feature information and then input into the decoder of the converter model. The overall feature information and the modal coding feature information are then fused to obtain the discrimination feature information. The discrimination feature information is calculated using a preset activation function to obtain a discrimination score, and the discrimination score is used to determine whether the target data is synthetic data.
[0006] Furthermore, the target data includes at least one modal data selected from audio data, image data, video data, and text data; The model input information includes overall identifier characters, modal data in the target data, as well as audio data identifier characters, image data identifier characters, video data identifier characters, and text data identifier characters.
[0007] Furthermore, the format of the model input information is as follows: <flag>Audio data <audio>Image data Video data <video>Text data <text>,in, <flag>Indicates the overall identifier character, <audio>Characters indicating audio data identification. Characters representing image data identifiers. <video>Characters indicating video data identification. <text>This indicates a text data identifier character. If the target data does not contain a certain modal data, then the corresponding modal data in the model input information is empty.
[0008] Further, the model input information is input into a large language model, and the features of each modality data are calculated and fused to obtain model output information. The model output information includes overall feature information that fuses the features of all modality data and modal feature information representing the features of each modality data, including: The model input information is input into a large language model, and the features of each modality data are calculated and fused to obtain model output information with the same sequence length as the model input information. The model output information and the model input information are compared. <flag>The feature information corresponding to the sequence position is determined as the overall feature information that integrates the features of all modal data, and is then compared with the model input information in the model output information. <audio> 、 、 <video>and <text>The feature information corresponding to the sequence position is determined as modal feature information representing the features of the corresponding modal data.
[0009] Furthermore, the overall feature information, which integrates all modal data features, is concatenated with the modal coding feature information and then input into the decoder of the converter model. The overall feature information and the modal coding feature information are further fused to obtain discriminative feature information, including: The first feature information is obtained by concatenating the overall feature information that integrates all modal data features with the modal coding feature information; The first feature information is input into the decoder of the converter model, and the overall feature information and the modal coding feature information are fused to obtain the second feature information with the same sequence length as the first feature information. The feature information in the second feature information that corresponds to the sequence position of the overall feature information in the first feature information is determined as the discriminative feature information that integrates the overall feature information and the modal coding feature information.
[0010] Further, a preset activation function is used to calculate the discrimination feature information to obtain a discrimination score, and the discrimination score is used to determine whether the target data is synthetic data, including: The discrimination score is obtained by calculating the discrimination feature information using a preset activation function; The discrimination score is compared with a preset judgment threshold. If the score is greater than or equal to the judgment threshold, the target data is determined to be synthetic data.
[0011] Furthermore, the preset activation function includes the Sigmoid function.
[0012] Furthermore, the target data is a type of modal data; The method involves calculating a discrimination score by using a preset activation function to identify feature information, and determining whether the target data is synthetic data based on the discrimination score. The feature information is calculated using a preset activation function to obtain a discrimination score, and the type of modal data of the target data is determined as synthetic data based on the discrimination score.
[0013] Some embodiments of this application also provide a unified identification device for synthetic data, wherein the device includes a memory for storing computer program instructions and a processor for executing the computer program instructions, wherein when the computer program instructions are executed by the processor, the device is triggered to execute the aforementioned unified identification method for synthetic data.
[0014] Other embodiments of this application also provide a computer-readable medium having computer program instructions stored thereon, which can be executed by a processor to implement the unified identification method for the synthetic data.
[0015] Compared to existing technologies, this application provides a unified identification scheme for synthetic data. This scheme requires constructing model input information based on target data. The target data includes at least one type of modal data. The model input information includes an overall identifier character, modal data from the target data, and modal identifier characters. The model input information is input into a large language model. Features of each modal data are calculated and fused to obtain model output information. The model output information includes overall feature information that fuses all modal data features and modal feature information representing the features of each modal data. The modal features representing the features of each modal data are then processed. Feature information is input to the encoder of the converter model. The features of each modal data are fused using a self-attention mechanism to obtain modality-encoded feature information, which incorporates attention information between the features of each modality data. The overall feature information, which incorporates all modality data features, is then concatenated with the modality-encoded feature information and input to the decoder of the converter model. The overall feature information and the modality-encoded feature information are fused to obtain discrimination feature information. A preset activation function is then used to calculate the discrimination feature information to obtain a discrimination score, and the discrimination score is used to determine whether the target data is synthetic data. Since this scheme does not require limiting the modality of the target data to be identified, regardless of the types of modality data included in the target data, a unified model input information can be automatically constructed. After extracting and fusing feature information from different modal data through a large language model and a converter model, the discrimination feature information used for judgment is finally obtained, thus completing the process of determining whether the target data is synthetic data. The entire process is not limited by the modality of the target data, therefore it has a wide range of applications and high flexibility. Attached Figure Description
[0016] Other features, objects, and advantages of this application will become more apparent from the following detailed description of non-limiting embodiments with reference to the accompanying drawings: Figure 1 A flowchart illustrating a unified identification method for synthetic data provided in an embodiment of this application; Figure 2 A schematic diagram illustrating the information processing procedure for synthetic data identification in the embodiments of this application; The same or similar reference numerals in the accompanying drawings represent the same or similar parts. Detailed Implementation
[0017] The present application will now be described in further detail with reference to the accompanying drawings.
[0018] In a typical configuration of this application, the terminal and the service network devices each include one or more processors (CPUs), input / output interfaces, network interfaces, and memory.
[0019] Memory may include non-persistent storage in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.
[0020] Computer-readable media include permanent and non-permanent, removable and non-removable media, which can store information by any method or technology. Information can be computer program instructions, data structures, program devices, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, read-only optical disc (CD-ROM), digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transfer medium that can be used to store information accessible by a computing device.
[0021] This application provides a unified method for identifying synthetic data. This method does not require limiting the modality of the target data to be identified. Regardless of how many types of modal data are included in the target data, a unified model input information can be automatically constructed. After extracting and fusing the feature information of different modal data through a large language model and a converter model, the identification feature information used for judgment is finally obtained, thereby completing the identification process of whether the target data is synthetic data. The entire process is not limited by the modality of the target data, so it has a wide range of applications and high flexibility.
[0022] In practical scenarios, the execution subject of this method can be a user device, a network device, or a device composed of user devices and network devices integrated through a network, or it can be an application running on the aforementioned devices. The user device includes, but is not limited to, various terminal devices such as computers, mobile phones, and tablets; the network device includes, but is not limited to, network hosts, single network servers, multiple network server sets, or cloud computing-based computer sets. Here, the cloud consists of a large number of hosts or network servers based on cloud computing, where cloud computing is a type of distributed computing, consisting of a virtual computer composed of a group of loosely coupled computer sets.
[0023] Figure 1 This application provides a unified identification method for synthetic data, which includes at least the following steps: Step S101: Construct model input information based on target data.
[0024] The target data includes at least one type of modal data. Specifically, the modal types in this scheme can include audio data, image data, video data, and text data. Therefore, the target data can include at least one of these four modal data types: audio data, image data, video data, and text data. The audio data refers to the audio content in the target data, the image data refers to the image content in the target data, the video data refers to the video content in the target data, and the text data refers to the text content in the target data.
[0025] For example, for target data that needs to be identified as synthetic data, data content of different modalities can be extracted to construct the model input information in this embodiment. Taking an animated short film as an example, if it is necessary to identify whether the animated short film is synthetic data, the cover image of the animated short film can be used as image data, the background music of the animated short film as audio data, the continuous video frames of the animated short film as video data, and the subtitle text of the animated short film as text data, thereby constructing the model input information.
[0026] When constructing the model input information, at least the following should be included: overall identifier characters, modal data in the target data, and identifier characters for audio data, image data, video data, and text data. The overall identifier characters, audio data identifier characters, image data identifier characters, video data identifier characters, and text data identifier characters are used to identify the sequence positions of overall information and corresponding type of modal information in the model input information. Combined with the trained Large Language Model (LLM), the LLM can output overall feature information that integrates all modal data features, as well as modal feature information representing each modal data feature, at the corresponding sequence positions of the aforementioned identifier characters when outputting the calculation results, making it convenient for users to obtain the above feature information.
[0027] In practical scenarios, since the target data includes at least one modal of data—audio data, image data, video data, and text data—the absence of any one or more modal information is permissible. Therefore, the target data can contain all four types of modal data (audio, image, video, and text), or it can contain only three, two, or one type of modal data. Taking the aforementioned animated short film as an example, if the animated short film does not contain subtitles, it may omit text data and only contain the other three types of modal data.
[0028] In some embodiments of this application, the overall identifier character, audio data identifier character, image data identifier character, video data identifier character, and text data identifier character can be represented by the following special characters. Among them, <flag>Indicates the overall identifier character, <audio>Characters indicating audio data identification. Characters representing image data identifiers. <video>Characters indicating video data identification. <text>This represents the text data identifier character. Based on this, the input information for the model constructed based on the target data can be in the following format: <flag>Audio data <audio>Image data Video data <video>Text data <text>If the target data does not contain a certain modality, then the corresponding modality data in the model input information is empty. For example, when the target data only contains text data, and the other three types of modality data are not input, then the model input information is: <flag> , <audio> , , <video>Text data <text>For modal data that has not been input, it is represented as "unknown" in the figure. Based on the above method of using specific identifier characters to mark each type of modal data, the input in the subsequent processing can always be controlled within these five identifier characters. This eliminates the need to limit the duration of audio data, the size of image data, the duration of visual data, and the number of characters in text data when inputting target data, thereby improving the applicability of the solution and constructing more standardized model input information, thus improving the efficiency of subsequent processing.
[0029] Step S102: Input the model input information into the large language model, calculate and fuse the features of each modality data, and obtain the model output information. The model output information includes overall feature information that fuses the features of all modality data, and modal feature information representing the features of each modality data separately.
[0030] The large language model used in this embodiment is pre-trained. Inputting the model input information into the large language model allows for the calculation and fusion of features from each modality, resulting in model output information with the same sequence length as the model input information. Furthermore, in the model output information, the information corresponding to the sequence position of each identifier character contains specific fusion feature information, wherein the fusion feature information is related to the overall identifier character... <flag>The feature information corresponding to the sequence position is summarized and integrated with the features contained in all modal data, along with the audio data identifier characters. <audio>The feature information corresponding to the sequence position is summarized and integrated with the features contained in the audio data and the character identifiers in the image data. The feature information corresponding to the sequence position is summarized and fused with the features contained in the image data and the video data identifier characters. <video>The feature information corresponding to the sequence position is summarized and fused with the features contained in the video data and the text data identifier characters. <text>The feature information corresponding to the sequence position is summarized and fused with the features contained in the text data. Therefore, the features in the model output information and the features in the model input information can be compared. <flag>The feature information corresponding to the sequence position is determined as the overall feature information that integrates the features of all modal data, and is then compared with the model input information in the model output information. <audio> 、 、 <video>and <text>The feature information corresponding to the sequence position is determined as modal feature information representing the features of the corresponding modal data.
[0031] For example, when constructing the model input information, its shape is (2256, 512), where 2256 is the sequence length and 512 is the feature embedding dimension. Thus, the aforementioned model input information is represented by a feature matrix formed by 2256 feature vectors of dimension 512. <flag>Audio data <audio>Image data Video data <video>Text data <text>After inputting it into a large language model, the model output information with the same sequence length of 2256 can be obtained. The embedding dimension of each sequence element can be set according to the needs of the actual scenario, thus obtaining model output information with a shape of (2256, N), where N is the embedding dimension of the features in the model output information. The model output information contains various identifier characters. <flag> 、 <audio> 、 、 <video> 、 <text>The corresponding sequence position is as follows Figure 2 As shown, the above sequence position can be determined by the overall feature information f_Flag that integrates all modal data features and the modal feature information f_Audio, f_Image, f_Video and f_Text that represent the corresponding modal data features.
[0032] Step S103 involves inputting the modal feature information representing the features of each modal data into the encoder of the Transformer model. The features of each modal data are then fused using a self-attention mechanism to obtain modal coding feature information. Thus, the modal coding feature information incorporates the attention information between the features of each modal data, enabling it to focus on the correlation between different modal data and improving the accuracy of the identification results.
[0033] Taking the aforementioned modal feature information f_Audio, f_Image, f_Video, and f_Text, representing the features of corresponding modal data, as an example, after being input into the encoder of the converter model, the encoder will use a self-attention mechanism to fuse and calculate the features of each modal data to obtain their respective feature information, which is the modal coding feature information. For example, after the modal feature information f_Audio (audio type), f_Image (image type), f_Video (video type), and f_Text (text type) are input into the encoder of the converter model, the corresponding modal coding feature information, including F_Audio, F_Image, F_Video, and F_Text, can be obtained.
[0034] Step S104: The overall feature information that integrates all modal data features is concatenated with the modal coding feature information and then input to the decoder of the converter model. The overall feature information and the modal coding feature information are fused to obtain the discrimination feature information.
[0035] In the scheme of this application embodiment, based on the overall feature information f_Flag that has already integrated all modal data features, the information fusion between the feature information of each modality and the overall feature information is further enhanced by the encoder and decoder of the converter model. This makes the obtained discrimination feature information better reflect the characteristics of the target data, thereby obtaining a more accurate discrimination result and improving the accuracy of the scheme.
[0036] Specifically, when obtaining the discriminative feature information through the decoder of the converter model, the overall feature information that integrates all modal data features can be concatenated with the modal coding feature information to obtain the first feature information. Thus, the first feature information is the result of concatenating the overall feature information f_Flag output by the large language model with the modal coding feature information F_Audio, F_Image, F_Video, and F_Text output by the encoder of the converter model.
[0037] After determining the first feature information, it is input to the decoder of the converter model. The overall feature information and the modal coding feature information are fused to obtain second feature information with the same sequence length as the first feature information. The decoder of the pre-trained converter model also outputs output information with the same sequence length as the input information. Therefore, the feature information in the second feature information corresponding to the sequence position of the overall feature information in the first feature information can be identified as the discriminative feature information that fuses the overall feature information and the modal coding feature information. Figure 2 As shown, the sequence position of the discriminative feature information in the second feature information corresponds to the sequence position of the overall feature information f_Flag in the first feature information.
[0038] Step S105: Calculate the discrimination feature information using a preset activation function to obtain a discrimination score, and determine whether the target data is synthetic data based on the discrimination score.
[0039] After calculating the discrimination feature information using a preset activation function, a discrimination score within a preset numerical range can be obtained. Depending on the activation function used, the calculated discrimination score will have a corresponding numerical range. When determining whether the target data is synthetic data, a reasonable judgment threshold can be determined within the numerical range. The discrimination score is compared with the preset judgment threshold; if it is greater than or equal to the judgment threshold, the target data is determined to be synthetic data.
[0040] In some embodiments of this application, different activation functions can be used depending on the application scenario, and corresponding judgment thresholds can be set to determine whether the target function is synthetic data. For example, the preset activation function in this embodiment can be the Sigmoid function, which can map the discriminative feature information within the numerical range of (0, 1) and calculate a discrimination score greater than 0 and less than 1. If the judgment threshold is set to 0.5, then when the discrimination score is ≥0.5, the target data is considered to be synthetic data; otherwise, it is real data.
[0041] Furthermore, to ensure that the discrimination score calculated by the activation function can be used to determine whether the target data is synthetic data, a binary cross-entropy loss function can be used when training the model. A large number of labeled synthetic data and real data are used as training samples to train the model, so that the discrimination score of synthetic data after processing by the model in this scheme approaches 1, while the discrimination score of real data after processing by the model in this scheme approaches 0.
[0042] In real-world scenarios, for a given target data set, some modalities may be real data, while others may be synthetic data. In such cases, the solution described in this application can be used to identify whether any type of modal data within the target data is synthetic, thereby broadening the application scope of the solution. Specifically, to identify whether a specific type of modal data within the target data is synthetic, when inputting the target data, only one type of modal data can be obtained to identify that type of modal data.
[0043] Since the unified identification method for synthetic data provided in this application embodiment only requires that the target data include at least one type of modal data, model input information can be normally constructed based on one type of modal data, and the identification score can be finally calculated based on the above steps. Because the other three types of modal data are missing in this scenario, the identification score can be used to determine whether the modal data of that type in the target data is synthetic data. For example, when the target data only contains text data, the constructed model input information is: <flag> , <audio> , , <video>Text data <text>After calculating the discrimination score through the above steps, if the discrimination score is ≥0.5, then the text data of the target data can be considered as synthetic data.
[0044] Therefore, this solution can not only uniformly identify target data, but also separately identify one type of modality data. It does not need to limit the modality of the target data to be identified. Regardless of how many types of modality data are included in the target data, it can automatically construct unified model input information. After extracting and fusing the feature information of different modal data through a large language model and a converter model, it finally obtains the discriminative feature information used for judgment, thereby completing the process of identifying whether it is synthetic data. The entire process is not limited by the modality of the target data, so it has a wide range of applications and high flexibility.
[0045] Based on another aspect of this application, embodiments of this application also provide a unified identification device for synthetic data, the device including a memory for storing computer program instructions and a processor for executing the computer program instructions, wherein when the computer program instructions are executed by the processor, the device is triggered to execute the aforementioned unified identification method for synthetic data.
[0046] In particular, the methods and / or embodiments in this application can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowchart. When the computer program is executed by a processing unit, it performs the functions defined in the methods of this application.
[0047] It should be noted that the computer-readable medium described in this application can be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this application, a computer-readable medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.
[0048] In this application, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. The computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium, capable of transmitting, propagating, or transmitting a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium may be transmitted using any suitable medium, including but not limited to: wireless, wireline, optical fiber, RF, etc., or any suitable combination thereof.
[0049] Computer program code for performing the operations of this application can be written in one or more programming languages or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, and C++, and conventional procedural programming languages such as "C" or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0050] The flowcharts or block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of devices, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-specific system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0051] In another aspect, this application also provides a computer-readable medium, which may be included in the device described in the above embodiments; or it may exist independently and not assembled into the device. The aforementioned computer-readable medium carries one or more computer program instructions, which may be executed by a processor to implement the methods and / or technical solutions of the various embodiments of this application.
[0052] It should be noted that this application can be implemented in software and / or a combination of software and hardware, for example, using an application-specific integrated circuit (ASIC), a general-purpose computer, or any other similar hardware device. In some embodiments, the software program of this application can be executed by a processor to implement the steps or functions described above. Similarly, the software program of this application (including related data structures) can be stored in a computer-readable recording medium, such as RAM memory, a magnetic or optical drive, a floppy disk, or similar devices. Furthermore, some steps or functions of this application can be implemented in hardware, for example, as circuitry that cooperates with a processor to perform the various steps or functions.
[0053] It will be apparent to those skilled in the art that this application is not limited to the details of the exemplary embodiments described above, and that this application can be implemented in other specific forms without departing from the spirit or essential characteristics of this application. Therefore, the embodiments should be considered exemplary and non-limiting in all respects, and the scope of this application is defined by the appended claims rather than the foregoing description. Thus, all variations falling within the meaning and scope of equivalents of the claims are intended to be embraced within this application. No reference numerals in the claims should be construed as limiting the scope of the claims. Furthermore, it is clear that the word "comprising" does not exclude other units or steps, and the singular does not exclude the plural. Multiple units or devices recited in the apparatus claims may also be implemented by a single unit or device in software or hardware. The terms "first," "second," etc., are used to indicate names and do not indicate any particular order.< / text> < / video> < / audio> < / flag> < / text> < / video> < / audio> < / flag> < / text> < / video> < / audio> < / flag> < / text> < / video> < / audio> < / flag> < / text> < / video> < / audio> < / flag> < / text> < / video> < / audio> < / flag> < / text> < / video> < / audio> < / flag> < / text> < / video> < / audio> < / flag> < / text> < / video> < / audio> < / flag> < / text> < / video> < / audio> < / flag> < / text> < / video> < / audio> < / flag>
Claims
1. A unified identification method for synthetic data, characterized in that, The method includes: Model input information is constructed based on target data, wherein the target data includes at least one type of modal data, and the model input information includes overall identifier characters, modal data in the target data, and modal identifier characters; The model input information is input into a large language model, and the features of each modal data are calculated and fused to obtain model output information. The model output information includes overall feature information that fuses the features of all modal data and modal feature information that represents the features of each modal data respectively. Modal feature information representing the features of each modal data is input into the encoder of the converter model. The features of each modal data are fused and calculated using a self-attention mechanism to obtain modal coding feature information. The modal coding feature information is fused with the attention information between the features of each modal data. The first feature information is obtained by concatenating the overall feature information that integrates all modal data features with the modal coding feature information; The first feature information is input into the decoder of the converter model, and the overall feature information and the modal coding feature information are fused to obtain the second feature information with the same sequence length as the first feature information. The feature information in the second feature information that corresponds to the sequence position of the overall feature information in the first feature information is determined as the discriminative feature information that integrates the overall feature information and the modality coding feature information; The discrimination feature information is calculated using a preset activation function to obtain a discrimination score, and the discrimination score is used to determine whether the target data is synthetic data.
2. The method according to claim 1, characterized in that, The target data includes at least one modal data selected from audio data, image data, video data, and text data; The model input information includes overall identifier characters, modal data in the target data, as well as audio data identifier characters, image data identifier characters, video data identifier characters, and text data identifier characters.
3. The method according to claim 2, characterized in that, The format of the model input information is as follows: <flag>Audio data <audio>Image data Video data <video>Text data <text>,in, <flag>Indicates the overall identifier character, <audio>Characters indicating audio data identification. Characters representing image data identifiers. <video>Characters indicating video data identification. <text> This indicates a text data identifier character. If the target data does not contain a certain modal data, then the corresponding modal data in the model input information is empty.< / text> < / video> < / audio> < / flag> < / text> < / video> < / audio> < / flag> 4. The method according to claim 3, characterized in that, The model input information is input into a large language model, and the features of each modality data are calculated and fused to obtain model output information. The model output information includes overall feature information that fuses the features of all modality data and modal feature information representing the features of each modality data, including: The model input information is input into a large language model, and the features of each modality data are calculated and fused to obtain model output information with the same sequence length as the model input information. The model output information and the model input information are compared. <flag>The feature information corresponding to the sequence position is determined as the overall feature information that integrates the features of all modal data, and is then compared with the model input information in the model output information. <audio> 、 、 <video>and <text> The feature information corresponding to the sequence position is determined as modal feature information representing the features of the corresponding modal data.< / text> < / video> < / audio> < / flag> 5. The method according to claim 1, characterized in that, The method involves calculating a discrimination score by using a preset activation function to identify feature information, and determining whether the target data is synthetic data based on the discrimination score. The discrimination score is obtained by calculating the discrimination feature information using a preset activation function; The discrimination score is compared with a preset judgment threshold. If the score is greater than or equal to the judgment threshold, the target data is determined to be synthetic data.
6. The method according to claim 5, characterized in that, The preset activation function includes the Sigmoid function.
7. The method according to claim 1, characterized in that, The target data is a type of modal data; The method involves calculating a discrimination score using a preset activation function to identify feature information, and determining whether the target data is synthetic data based on the discrimination score. The discrimination feature information is calculated using a preset activation function to obtain a discrimination score, and the type of modal data of the target data is determined as synthetic data based on the discrimination score.
8. A unified identification device for synthetic data, wherein, The device includes a memory for storing computer program instructions and a processor for executing the computer program instructions, wherein when the computer program instructions are executed by the processor, the device is triggered to perform the method of any one of claims 1 to 7.
9. A computer-readable medium having stored thereon computer program instructions that can be executed by a processor to implement the method as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Method for predicting post interaction behavior state, computing equipment and storage medium
CN113837457A
Model training method, classification method and related device
CN117079046A