Authentic identification method and device, electronic equipment and storage medium

Through multimodal data analysis and text correlation prediction model, the problem of low accuracy in detecting deep learning in the prior art is solved, and higher detection accuracy and robustness are achieved, and suitable for the recognition and real-time detection of complex forged content.

CN120032137APending Publication Date: 2025-05-23ACADEMY OF BROADCASTING SCI STATE ADMINISTATION OF PRESS PUBLICATION RADIO FILM & TELEVISION
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411246380.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-09-05
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

The prior art detects high-fidelity forged content generated by deep learning technology, and it is difficult to effectively identify complex forged content.

Method used

Through the trained image feature extraction model, text feature extraction model and audio feature extraction model, multimodal data text information of images, text and audio is obtained respectively, and the text correlation prediction model is used to analyze the semantic relationship between different modal data, and the text correlation value is generated to determine the false detection results of the object to be detected.

Benefits of technology

Improve the accuracy and robustness of detection, enhance the recognition ability of complex forged content, enable real-time detection, and adapt to large-scale data flows.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120032137A_ABST
    Figure CN120032137A_ABST
Patent Text Reader

Abstract

The invention relates to an authentic identification method and device, electronic equipment and a storage medium, and the method comprises the steps: obtaining first text information according to a to-be-detected object and a trained image feature extraction model, obtaining second text information according to the to-be-detected object and a trained text feature extraction model, obtaining third text information according to the to-be-detected object and the trained audio feature extraction model; inputting the first text information, the second text information and the third text information into a trained text association degree prediction model to obtain a text association degree value; according to the text correlation degree value, a first authentic identification result of the object to be detected is obtained, and the first authentic identification result is real or synthetic.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to counterfeit detection technology, and more specifically, to a counterfeit detection method, device, electronic device and storage medium. Background Art

[0002] With the rapid development of artificial intelligence technology, deep synthesis technology has become a powerful tool that can generate realistic images and videos. This technology has shown great potential in the fields of image synthesis, video production, game design, virtual reality, etc., but it also brings unprecedented challenges. Due to the high-quality synthesis ability of deep synthesis technology, the content generated by deep synthesis technology has a certain ability to be forged, which has formed a certain deception to the human visual system. Some criminals may use deep generation technology to generate false content, spread false content, create false events, etc., which will have an adverse impact on the normal and orderly operation of society. Therefore, it has become an urgent issue to develop effective detection methods to identify these forged contents.

[0003] Traditional identification methods mainly rely on manually extracted features, such as pixel statistics and noise features, but these methods have low detection accuracy when faced with high-fidelity forged content generated by deep learning technology. In recent years, deep learning-based detection technology has gradually become a research hotspot. Its core lies in using deep neural networks to automatically learn discriminative features from real content and forged content. However, these methods often focus on low-level visual features of images and fail to capture sufficient discriminative information, resulting in poor identification accuracy.

[0004] Therefore, it is necessary to provide a new technical solution to improve the accuracy of identification. Summary of the invention

[0005] One purpose of the present disclosure is to provide a new technical solution for a counterfeit identification method.

[0006] According to a first aspect of the present disclosure, a counterfeit identification method is provided, comprising:

[0007] Obtaining first text information according to the object to be detected and the trained image feature extraction model, obtaining second text information according to the object to be detected and the trained text feature extraction model, and obtaining third text information according to the object to be detected and the trained audio feature extraction model;

[0008] Inputting the first text information, the second text information and the third text information into a trained text relevance prediction model to obtain a text relevance value;

[0009] A first authentication result of the object to be detected is obtained according to the text association value, wherein the first authentication result is real or synthetic.

[0010] Optionally, obtaining the first text information according to the object to be detected and a trained image feature extraction model includes:

[0011] Obtaining an image corresponding to the object to be detected;

[0012] According to the image corresponding to the object to be detected and the trained image feature extraction model, text description information of each element in the image, and / or text description information of the relationship between elements in the image, and / or text description information of the plot content are obtained as the first text information.

[0013] Optionally, obtaining the second text information according to the object to be detected and a trained text feature extraction model includes:

[0014] Obtaining text description information and / or subtitle information corresponding to the object to be detected;

[0015] According to the text description information and / or subtitle information corresponding to the object to be detected and the trained text feature extraction model, context association information is obtained as the second text information.

[0016] Optionally, obtaining the third text information according to the object to be detected and the trained audio feature extraction model includes:

[0017] Obtaining audio corresponding to the object to be detected;

[0018] According to the audio corresponding to the object to be detected and the trained audio feature extraction model, text description information of the emotion is obtained as the third text information.

[0019] Optionally, the method further includes:

[0020] Obtaining an image corresponding to the object to be detected;

[0021] Inputting the image corresponding to the object to be detected into a trained image detection model to obtain a second authentication result, wherein the trained image detection model is a model that performs detection based on pixel statistical features, texture features, and color features of the image;

[0022] A final authentication result is obtained according to the first authentication result and the second authentication result.

[0023] Optionally, the method further includes:

[0024] Obtaining audio corresponding to the object to be detected;

[0025] Input the audio corresponding to the object to be detected into a trained audio detection model to obtain a third authentication result, wherein the trained audio detection model is a model that performs detection based on time domain features and frequency domain features of the audio;

[0026] A final authentication result is obtained according to the first authentication result and the third authentication result.

[0027] Optionally, the method further includes:

[0028] Inputting a training sample set into the trained image feature extraction model to obtain a first output result, inputting the training sample set into the trained text feature extraction model to obtain a second output result, and inputting the training sample set into the trained audio feature extraction model to obtain a third output result, wherein the training sample set includes a plurality of training samples and corresponding manually annotated authentication results;

[0029] Inputting the first output result, the second output result and the third output result into a text relevance prediction model to be trained to obtain a text relevance value corresponding to each training sample;

[0030] Obtaining a predicted authentication result for each training sample according to the text association value corresponding to each training sample;

[0031] When the number of consistent results between the predicted authentication results of each training sample and the manually annotated authentication results of each training sample is less than a preset threshold, the parameters of the text relevance prediction model to be trained are adjusted, and the training is continued until the number of consistent results between the predicted authentication results of each training sample and the manually annotated authentication results of each training sample is greater than or equal to the preset threshold.

[0032] According to a second aspect of the present disclosure, there is provided a counterfeit detection device, comprising:

[0033] A text information determination module, used to obtain first text information according to the object to be detected and the trained image feature extraction model, obtain second text information according to the object to be detected and the trained text feature extraction model, and obtain third text information according to the object to be detected and the trained audio feature extraction model;

[0034] A text relevance value determination module, used for inputting the first text information, the second text information and the third text information into a trained text relevance prediction model to obtain a text relevance value;

[0035] The authentication result determination module is used to obtain a first authentication result of the object to be detected according to the text association value, wherein the first authentication result is real or synthetic.

[0036] According to a third aspect of the present disclosure, an electronic device is provided, comprising a memory and a processor, wherein the memory stores a computer program, and the computer program is used to control the processor to operate so as to execute the counterfeit identification method according to any one of the first aspects of the present disclosure.

[0037] According to a fourth aspect of the present disclosure, a non-volatile computer-readable storage medium is provided, on which computer program instructions are stored. When the computer program instructions are executed by a processor, the counterfeit identification method described in any one of the first aspects is implemented.

[0038] The authentication method provided by the embodiment of the present disclosure obtains first text information, second text information and third text information respectively through a trained image feature extraction model, a trained text feature extraction model and a trained audio feature extraction model. The three types of text information are multimodal data. The trained text relevance prediction model is used to analyze the semantic associations between different modal data to obtain text relevance values. The first authentication result of the object to be detected is determined based on the text relevance values, thereby improving the accuracy and robustness of detection and enhancing the ability to identify complex forged content. At the same time, real-time detection can be achieved and large-scale data streams can be coped with.

[0039] Features and advantages of the embodiments of the present specification will become apparent from the following detailed description of exemplary embodiments of the present specification with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] The accompanying drawings, which are incorporated in and constitute a part of the specification, illustrate embodiments of the specification and, together with the description, serve to explain the principles of the embodiments of the specification.

[0041] Figure 1 A processing flow chart of an image authentication method according to an embodiment of the present disclosure is shown.

[0042] Figure 2 A principle block diagram of an image authentication device according to an embodiment of the present disclosure is shown.

[0043] Figure 3 A structural block diagram of an electronic device according to an embodiment of the present disclosure is shown. DETAILED DESCRIPTION

[0044] Various exemplary embodiments of the present specification will now be described in detail with reference to the accompanying drawings.

[0045] The following description of at least one exemplary embodiment is merely illustrative in nature and is in no way intended to limit the embodiments of the present specification and its application or uses.

[0046] It should be noted that like reference numerals and letters refer to similar items in the following figures, and therefore, once an item is defined in one figure, it need not be further discussed in subsequent figures.

[0047] Figure 1 FIG. 1 is a flow chart of a method for detecting counterfeit according to an embodiment of the present disclosure. Figure 1 As shown, the method includes steps S110 to S130.

[0048] Step S110, obtaining first text information according to the object to be detected and the trained image feature extraction model, obtaining second text information according to the object to be detected and the trained text feature extraction model, and obtaining third text information according to the object to be detected and the trained audio feature extraction model.

[0049] The object to be detected can be a video or an image.

[0050] When the object to be detected is not accompanied by text description information and / or subtitle information, the operation of obtaining the second text information based on the object to be detected and the trained text feature extraction model is no longer performed. When the object to be detected is not accompanied by audio, the operation of obtaining the third text information based on the object to be detected and the trained audio feature extraction model is no longer performed.

[0051] Step S120: input the first text information, the second text information and the third text information into a trained text relevance prediction model to obtain a text relevance value.

[0052] The text relevance prediction model is a deep learning model.

[0053] The text association value indicates the degree of matching between the first text information, the second text information, and the third text information. The higher the degree of matching between the first text information, the second text information, and the third text information, the greater the text association value. The lower the degree of matching between the first text information, the second text information, and the third text information, the smaller the text association value.

[0054] Step S130, obtaining a first authentication result of the object to be detected according to the text association value, wherein the first authentication result is authentic or synthetic.

[0055] Specifically, the text association value is compared with a preset threshold to obtain a comparison result. If the comparison result is that the text association value is greater than or equal to the preset threshold, the first authentication result is true. If the comparison result is that the text association value is less than the preset threshold, the first authentication result is synthetic.

[0056] The authentication method provided by the embodiment of the present invention obtains first text information, second text information and third text information respectively through a trained image feature extraction model, a trained text feature extraction model and a trained audio feature extraction model. The three types of text information are multimodal data. The trained text relevance prediction model is used to analyze the semantic association between different modal data to obtain a text relevance value. The first authentication result of the object to be detected is determined based on the text relevance value, thereby improving the accuracy and robustness of detection and enhancing the ability to identify complex forged content. At the same time, real-time detection can be achieved and large-scale data streams can be coped with.

[0057] In one embodiment, obtaining the first text information according to the object to be detected and the trained image feature extraction model in step S110 specifically includes: obtaining an image corresponding to the object to be detected; obtaining text description information of each element in the image, and / or text description information of the relationship between elements in the image, and / or text description information of the plot content, as the first text information according to the image corresponding to the object to be detected and the trained image feature extraction model.

[0058] When the object to be detected is a video, the image corresponding to the object to be detected may be all frames constituting the video, or may be a partial frame constituting the video.

[0059] The trained image feature extraction model is a model in Computer Vision (CV) technology, which can be a Convolutional Neural Network (CNN) model or an image description generation model (such as a Show and Tell model).

[0060] Elements in the image include people, objects, and environmental factors. The text description information of each element in the image includes the name and quantity of the element. The text description information of the relationship between the elements in the image is a sentence that associates the elements in the image in a fixed pattern. The text description information of the plot content is the corresponding description information when the elements in the image and / or the relationship between the elements change. When the object to be detected is an image, the text description information of the plot content is not output. When the object to be detected is a video, the text description information of the plot content can be output.

[0061] In one embodiment, the trained image feature extraction model can be trained according to the following method: obtain a first training set, the first training set includes multiple training samples and corresponding manually annotated text description information; input the first training set into the image feature extraction model to be trained for training, and obtain the text description information corresponding to each training sample; based on each training sample, determine whether the text description information obtained based on the model and the corresponding manually annotated text description information are consistent, and obtain the detection results corresponding to each training sample; when the number of inconsistencies in the detection result is greater than or equal to a preset number, adjust the parameters of the image feature extraction model, continue training, until the number of inconsistencies in the detection result is less than a preset number, stop training, and obtain a trained image feature extraction model. The training samples in the first training set can all be images, or all be videos, or some can be images and some can be videos.

[0062] In one embodiment, obtaining the second text information according to the object to be detected and the trained text feature extraction model in step S110 specifically includes: obtaining text description information and / or subtitle information corresponding to the object to be detected; obtaining context-related information as the second text information according to the text description information and / or subtitle information corresponding to the object to be detected and the trained text feature extraction model.

[0063] The text description information corresponding to the object to be detected is the text information attached to the object to be detected itself, for example, a brief introduction to the content.

[0064] The trained text feature extraction model is a model in the natural language processing (NLP) technology, which can be a BERT (Bidirectional Encoder Representations from Transformers) model or a GPT (Generative Pre-trained Transformer) model.

[0065] The second text information may be a sentence summarized based on the text description information and / or subtitle information corresponding to the object to be detected.

[0066] In one embodiment, the trained text feature extraction model can be trained according to the following method: obtain a second training set, the second training set includes multiple training samples and corresponding manually annotated text description information, each training sample is a content introduction and / or subtitle information obtained from an image or video; input the second training set to the text feature extraction model to be trained for training, and obtain the text description information corresponding to each training sample; based on each training sample, determine whether the text description information obtained based on the model and the corresponding manually annotated text description information are consistent, and obtain the detection result corresponding to each training sample; when the detection result is that the number of inconsistencies is greater than or equal to a preset number, adjust the parameters of the text feature extraction model, and continue training until the detection result is that the number of inconsistencies is less than a preset number, stop training, and obtain a trained text feature extraction model.

[0067] In one embodiment, obtaining the third text information according to the object to be detected and the trained audio feature extraction model in step S110 specifically includes: obtaining the audio corresponding to the object to be detected; obtaining text description information of the emotion as the third text information according to the audio corresponding to the object to be detected and the trained audio feature extraction model.

[0068] The trained audio feature extraction model is a model of speech processing (SP) technology, and may be a long short-term memory network (LSTM) model.

[0069] The text description information of emotion can be any one of happy, sad, angry, and calm.

[0070] In one embodiment, the trained audio feature extraction model can be trained according to the following method: obtain a third training set, the third training set includes multiple training samples and corresponding manually annotated text description information, each training sample is an audio obtained from an image or video; input the third training set to the audio feature extraction model to be trained for training, and obtain the text description information corresponding to each training sample; based on each training sample, determine whether the text description information obtained based on the model and the corresponding manually annotated text description information are consistent, and obtain the detection result corresponding to each training sample; when the detection result is that the number of inconsistencies is greater than or equal to a preset number, adjust the parameters of the audio feature extraction model, and continue training until the detection result is that the number of inconsistencies is less than a preset number, stop training, and obtain a trained audio feature extraction model.

[0071] In one embodiment, the method also includes: obtaining an image corresponding to the object to be detected; inputting the image corresponding to the object to be detected into a trained image detection model to obtain a second authentication result, wherein the trained image detection model is a model that performs detection based on pixel statistical features, texture features, and color features of the image; and obtaining a final authentication result based on the first authentication result and the second authentication result.

[0072] The trained image detection model is a deep learning model, for example, a Convolutional Neural Networks (CNN) model.

[0073] After the image is input into the trained image detection model, the pixel statistical features, texture features and color features of the image can be extracted. The first detection result is obtained by analyzing the pixel statistical features of the image to detect whether the image has traces of tampering. The second detection result is obtained by detecting whether the image has a forged area through the texture features of the image. The authenticity of the image is determined by analyzing the image's verbal features (such as color distribution, color consistency) to obtain a third detection result. When at least one of the first detection result, the second detection result and the third detection result is a synthesis, the second authentication result is a synthesis.

[0074] When at least one of the first authentication result and the second authentication result is a composite, the final authentication result of the object to be detected is a composite. By using the first authentication result and the second authentication result to determine the final authentication result of the object to be detected, the object to be detected can be detected from multiple aspects and the detection accuracy can be improved.

[0075] In one embodiment, the method also includes: obtaining audio corresponding to the object to be detected; inputting the audio corresponding to the object to be detected into a trained audio detection model to obtain a third authentication result, wherein the trained audio detection model is a model that performs detection based on time domain features and frequency domain features of the audio; and obtaining a final authentication result based on the first authentication result and the third authentication result.

[0076] The trained audio detection model is a deep learning model, such as the Detect-2B model.

[0077] After the audio is input into the trained audio detection model, the time domain features and frequency domain features of the audio can be extracted. By analyzing the time domain features and frequency domain features of the audio, it is detected whether the audio is synthesized and a detection result is obtained. When the detection result is that synthesis exists, the third authentication result is synthesis.

[0078] When at least one of the first authentication result and the third authentication result is a composite, the final authentication result of the object to be detected is a composite. By using the first authentication result and the third authentication result to determine the final authentication result of the object to be detected, the object to be detected can be detected from multiple aspects and the detection accuracy can be improved.

[0079] In one embodiment, a final authentication result is obtained according to the first authentication result, the second authentication result and the third authentication result. When at least one of the first authentication result, the second authentication result and the third authentication result is a synthesis, the final authentication result of the object to be detected is a synthesis, which can realize the detection of the object to be detected from multiple aspects and improve the accuracy of the detection.

[0080] In one embodiment, the method further includes: inputting the training sample set into a trained image feature extraction model to obtain a first output result, inputting the training sample set into a trained text feature extraction model to obtain a second output result, and inputting the training sample set into a trained audio feature extraction model to obtain a third output result, wherein the training sample set includes multiple training samples and corresponding manually annotated authentication results; inputting the first output result, the second output result, and the third output result into a text relevance prediction model to be trained to obtain a text relevance value corresponding to each training sample; obtaining a predicted authentication result of each training sample based on the text relevance value corresponding to each training sample; and when the number of consistent predicted authentication results of each training sample and the manually annotated authentication results of each training sample is less than a preset threshold, adjusting the parameters of the text relevance prediction model to be trained, and continuing training until the number of consistent predicted authentication results of each training sample and the manually annotated authentication results of each training sample is greater than or equal to the preset threshold.

[0081] The training samples in the training sample set may all be images, may all be videos, or may be partly images and partly videos. The training samples in the training sample set include both real samples and synthetic samples.

[0082] In one embodiment, a confusion matrix may also be used to determine whether the text relevance prediction model to be trained is trained. Specifically, based on the confusion matrix, at least one indicator among accuracy, recall, and F1 score is used to evaluate the performance of the text relevance prediction model to be trained.

[0083] The present disclosure also provides a counterfeit detection device for implementing any of the above method embodiments. Figure 2 FIG. 4 shows a structural block diagram of a counterfeit detection device according to some embodiments. Figure 2 As shown, the authentication device 200 may include a text information determination module 210 , a text relevance value determination module 220 and an authentication result determination module 230 .

[0084] The text information determination module 210 is used to obtain first text information based on the object to be detected and the trained image feature extraction model, obtain second text information based on the object to be detected and the trained text feature extraction model, and obtain third text information based on the object to be detected and the trained audio feature extraction model.

[0085] The text relevance value determination module 220 is used to input the first text information, the second text information and the third text information into a trained text relevance prediction model to obtain a text relevance value.

[0086] The authentication result determination module 230 is used to obtain a first authentication result of the object to be detected according to the text association value, wherein the first authentication result is real or synthetic.

[0087] In one embodiment, the text information determination module 210 is also used to obtain an image corresponding to the object to be detected; based on the image corresponding to the object to be detected and the trained image feature extraction model, obtain text description information of each element in the image, and / or text description information of the relationship between elements in the image, and / or text description information of the plot content as the first text information.

[0088] In one embodiment, the text information determination module 210 is also used to obtain text description information and / or subtitle information corresponding to the object to be detected; based on the text description information and / or subtitle information corresponding to the object to be detected and the trained text feature extraction model, context association information is obtained as the second text information.

[0089] In one embodiment, the text information determination module 210 is also used to obtain the audio corresponding to the object to be detected; based on the audio corresponding to the object to be detected and the trained audio feature extraction model, obtain the text description information of the emotion as the third text information.

[0090] In one embodiment, the device further includes a final authentication result determination module. The final authentication result determination module is further used to obtain an image corresponding to the object to be detected; input the image corresponding to the object to be detected into a trained image detection model to obtain a second authentication result, wherein the trained image detection model is a model that performs detection based on pixel statistical features, texture features, and color features of the image; and obtain a final authentication result based on the first authentication result and the second authentication result.

[0091] In one embodiment, the device further includes a final authentication result determination module. The final authentication result determination module is further used to obtain the audio corresponding to the object to be detected; input the audio corresponding to the object to be detected into a trained audio detection model to obtain a third authentication result, wherein the trained audio detection model is a model that performs detection based on the time domain features and frequency domain features of the audio; and obtain the final authentication result according to the first authentication result and the third authentication result.

[0092] In one embodiment, the device further includes a model training module. The model training module is further used to input the training sample set into the trained image feature extraction model to obtain a first output result, input the training sample set into the trained text feature extraction model to obtain a second output result, and input the training sample set into the trained audio feature extraction model to obtain a third output result, wherein the training sample set includes multiple training samples and corresponding manually annotated authentication results; input the first output result, the second output result, and the third output result into the text relevance prediction model to be trained to obtain the text relevance value corresponding to each training sample; obtain the predicted authentication result of each training sample according to the text relevance value corresponding to each training sample; when the number of consistent predicted authentication results of each training sample and the manually annotated authentication results of each training sample is less than a preset threshold, adjust the parameters of the text relevance prediction model to be trained, and continue training until the number of consistent predicted authentication results of each training sample and the manually annotated authentication results of each training sample is greater than or equal to the preset threshold.

[0093] The present disclosure also provides an electronic device for implementing any of the above method embodiments. Figure 3 The structure block diagram of the electronic device 3 according to some embodiments is shown. The electronic device 3 can be a PC, a workstation, a notebook computer, a server, etc., which is not limited here.

[0094] like Figure 3 As shown, the electronic device 3 includes a processor 310 and a memory 320 for storing instructions executable by the processor 310. The processor 310 is configured to implement the image authentication method according to any embodiment of the present disclosure when executing the instructions stored in the memory 320.

[0095] The processor 310 is used to execute computer instructions, which can be written in an instruction set of an architecture such as x86, Arm, RISC, MIPS, SSE, etc. The memory 320 includes, for example, ROM (read-only memory), RAM (random access memory), a non-volatile memory such as a hard disk, etc., which are not limited here.

[0096] The present disclosure also provides a non-volatile computer-readable storage medium on which computer program instructions are stored. When the computer program instructions are executed by a processor, the counterfeit identification method provided in any of the above embodiments is implemented.

[0097] Each embodiment in this specification is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. For the device embodiment, its related parts can be referred to the partial description of the method embodiment.

[0098] The above is a description of a specific embodiment of the specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recorded in the claims can be performed in an order different from that in the embodiments and still achieve the desired results. In addition, the processes depicted in the drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0099] The embodiments of the present specification may be systems, methods and / or computer program products. The computer program product may include a computer-readable storage medium carrying computer instructions for causing a processor to implement various aspects of the embodiments of the present specification.

[0100] A computer-readable storage medium may be a tangible device that can hold and store computer instructions used by a computer instruction execution device. A computer-readable storage medium may be, for example, but not limited to, an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. More specific examples of computer-readable storage media (a non-exhaustive list) include: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a static random access memory (SRAM), a portable compact disk read-only memory (CD-ROM), a digital versatile disk (DVD), a memory stick, a floppy disk, a mechanical encoding device, such as a punch card or a raised structure in a groove on which computer instructions are stored, and any suitable combination of the foregoing. As used herein, a computer-readable storage medium is not to be interpreted as a transient signal per se, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagating through a waveguide or other transmission medium (e.g., a light pulse through a fiber optic cable), or an electrical signal transmitted through a wire.

[0101] The computer instructions described herein can be downloaded from a computer-readable storage medium to each computing / processing device, or downloaded to an external computer or external storage device through a network layer, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network layer may include copper transmission cables, optical fiber transmission, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. The network layer adapter card or network layer interface in each computing / processing device receives the computer instructions from the network layer and forwards the computer instructions for storage in a computer-readable storage medium in each computing / processing device.

[0102] The flowcharts and block diagrams in the accompanying drawings show the possible architecture, functions and operations of the systems, methods and computer program products according to the multiple embodiments of this specification. In this regard, each box in the flowchart or block diagram can represent a module, a program segment or a part of a computer instruction, and a part of a module, a program segment or a computer instruction contains one or more executable computer instructions for realizing the specified logical function. In some alternative implementations, the functions marked in the box can also occur in a different order from the order marked in the accompanying drawings. For example, two consecutive boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or the flowchart, and the combination of the boxes in the block diagram and / or the flowchart can be implemented by a dedicated hardware-based system that performs the specified function or action, or can be implemented by a combination of dedicated hardware and computer instructions. It is well known to those skilled in the art that it is equivalent to implement it by hardware, implement it by software, and implement it by combining software and hardware.

[0103] The embodiments of the present specification have been described above, and the above description is exemplary, not exhaustive, and is not limited to the disclosed embodiments. Many modifications and variations are obvious to those of ordinary skill in the art without departing from the scope of the described embodiments. The selection of terms used herein is intended to best explain the principles of the embodiments, practical applications, or improvements to the technology in the market, or to enable other persons of ordinary skill in the art to understand the embodiments disclosed herein.

Claims

1. A method for detecting counterfeit goods, characterized in that: include: Obtaining first text information according to the object to be detected and the trained image feature extraction model, obtaining second text information according to the object to be detected and the trained text feature extraction model, and obtaining third text information according to the object to be detected and the trained audio feature extraction model; Inputting the first text information, the second text information and the third text information into a trained text relevance prediction model to obtain a text relevance value; A first authentication result of the object to be detected is obtained according to the text association value, wherein the first authentication result is real or synthetic.

2. The method according to claim 1, characterized in that The step of obtaining the first text information according to the object to be detected and the trained image feature extraction model includes: Obtaining an image corresponding to the object to be detected; According to the image corresponding to the object to be detected and the trained image feature extraction model, text description information of each element in the image, and / or text description information of the relationship between elements in the image, and / or text description information of the plot content are obtained as the first text information.

3. The method according to claim 1, characterized in that The step of obtaining the second text information according to the object to be detected and the trained text feature extraction model includes: Obtaining text description information and / or subtitle information corresponding to the object to be detected; According to the text description information and / or subtitle information corresponding to the object to be detected and the trained text feature extraction model, context association information is obtained as the second text information.

4. The method according to claim 1, characterized in that: The step of obtaining the third text information according to the object to be detected and the trained audio feature extraction model includes: Obtaining audio corresponding to the object to be detected; According to the audio corresponding to the object to be detected and the trained audio feature extraction model, text description information of the emotion is obtained as the third text information.

5. The method according to claim 1, characterized in that The method further comprises: Obtaining an image corresponding to the object to be detected; Inputting the image corresponding to the object to be detected into a trained image detection model to obtain a second authentication result, wherein the trained image detection model is a model that performs detection based on pixel statistical features, texture features, and color features of the image; A final authentication result is obtained according to the first authentication result and the second authentication result.

6. The method according to claim 1, characterized in that The method further comprises: Obtaining audio corresponding to the object to be detected; Input the audio corresponding to the object to be detected into a trained audio detection model to obtain a third authentication result, wherein the trained audio detection model is a model that performs detection based on time domain features and frequency domain features of the audio; A final authentication result is obtained according to the first authentication result and the third authentication result.

7. The method according to any one of claims 1 to 6, characterized in that: The method further comprises: Inputting a training sample set into the trained image feature extraction model to obtain a first output result, inputting the training sample set into the trained text feature extraction model to obtain a second output result, and inputting the training sample set into the trained audio feature extraction model to obtain a third output result, wherein the training sample set includes a plurality of training samples and corresponding manually annotated authentication results; Inputting the first output result, the second output result and the third output result into a text relevance prediction model to be trained to obtain a text relevance value corresponding to each training sample; Obtaining a predicted authentication result for each training sample according to the text association value corresponding to each training sample; When the number of consistent results between the predicted authentication results of each training sample and the manually annotated authentication results of each training sample is less than a preset threshold, the parameters of the text relevance prediction model to be trained are adjusted, and the training is continued until the number of consistent results between the predicted authentication results of each training sample and the manually annotated authentication results of each training sample is greater than or equal to the preset threshold.

8. A counterfeit detection device, characterized in that: include: A text information determination module, used to obtain first text information according to the object to be detected and the trained image feature extraction model, obtain second text information according to the object to be detected and the trained text feature extraction model, and obtain third text information according to the object to be detected and the trained audio feature extraction model; A text relevance value determination module, used for inputting the first text information, the second text information and the third text information into a trained text relevance prediction model to obtain a text relevance value; The authentication result determination module is used to obtain a first authentication result of the object to be detected according to the text association value, wherein the first authentication result is real or synthetic.

9. An electronic device, characterized in that: The method comprises a memory and a processor, wherein the memory stores a computer program, and the computer program is used to control the processor to operate so as to execute the counterfeit identification method according to any one of claims 1 to 7.

10. A non-volatile computer-readable storage medium having computer program instructions stored thereon, characterized in that: When the computer program instructions are executed by a processor, the counterfeit identification method according to any one of claims 1 to 7 is implemented.