Image detection method and device, training method and device, electronic equipment, medium and program product

Through the visual-language model, the image features and prompt features are fused and aggregated by the cross attention mechanism, and the judgment criteria are dynamically constructed, which solves the problems of fixed decision boundary limitations and insufficient feature representation in the existing technology, and realizes efficient detection of new forgery technologies.

CN120279324APending Publication Date: 2025-07-08BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510389770.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-28
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

Existing image detection technology is difficult to effectively deal with the rapidly evolving deep forgery technology, and there are problems of limited boundaries of fixed decision making and insufficient diversity of feature representations.

Method used

The visual-language model is used to fusion image features and prompt features, and through the dynamic similarity comparison mechanism, an adaptive authenticity evaluation mechanism is built, and the cross-attention mechanism is used to fusion and aggregate features are used to dynamically build customized judgment standards.

Benefits of technology

It improves the accuracy of image detection, can effectively deal with new forgery technologies, and provides a more effective paradigm for image authenticity verification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120279324A_ABST
    Figure CN120279324A_ABST
Patent Text Reader

Abstract

The invention provides an image detection method and device, a training method and device, electronic equipment, a medium and a program product, and relates to the field of artificial intelligence, in particular to the fields of deep learning, large language models, image recognition, deep forgery detection and the like. According to the implementation scheme, image features of a to-be-detected image including a target object and a first prompt feature and a second prompt feature used for detecting the to-be-detected image are obtained; fusing the first prompt feature and the image feature to obtain a first reference feature, and fusing the second prompt feature and the image feature to obtain a second reference feature; determining a first similarity between the first reference feature and the image feature and a second similarity between the second reference feature and the image feature; based on the first similarity and the second similarity, an image detection result is determined, and the image detection result indicates whether the to-be-detected image is real or not.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of artificial intelligence technology, and particularly to fields such as deep learning, large language models, image recognition, and deepfake detection. Specifically, it relates to an image detection method, an image detection training method, a device, an electronic device, a computer-readable storage medium, and a computer program product. Background Art

[0002] Artificial intelligence is a discipline that studies enabling computers to simulate certain human thinking processes and intelligent behaviors (such as learning, reasoning, thinking, planning, etc.), and it has both hardware-level technologies and software-level technologies. Artificial intelligence hardware technologies generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, and big data processing; artificial intelligence software technologies mainly include several major directions such as computer vision technology, speech recognition technology, natural language processing technology, and machine learning / deep learning, big data processing technology, and knowledge graph technology.

[0003] With the progress of artificial intelligence technology, the image processing effects obtained by applying artificial intelligence technology have been further improved. Correspondingly, the occasions where the authenticity of images needs to be detected are increasing day by day, prompting the industry to continuously seek more effective image detection methods, especially for dealing with certain deepfake situations.

[0004] The methods described in this section are not necessarily methods that have been previously conceived or adopted. Unless otherwise specified, no method described in this section should be considered prior art solely because it is included in this section. Similarly, unless otherwise specified, the problems mentioned in this section should not be considered to have been recognized in any prior art. Summary of the Invention

[0005] The present disclosure provides an image detection method, an image detection training method, a device, an electronic device, a computer-readable storage medium, and a computer program product.

[0006] According to one aspect of the present disclosure, there is provided an image detection method, including: obtaining image features of a to-be-detected image including a target object, as well as a first hint feature and a second hint feature for detecting the to-be-detected image, where the first hint feature is associated with the true attribute of the target object, and the second hint feature is associated with the non-true attribute of the target object; fusing the first hint feature and the image features to obtain a first reference feature, and fusing the second hint feature and the image features to obtain a second reference feature; determining a first similarity between the first reference feature and the image features, and a second similarity between the second reference feature and the image features; and based on the first similarity and the second similarity, determining an image detection result, where the image detection result indicates whether the to-be-detected image is real.

[0007] According to another aspect of the present disclosure, there is provided an image detection training method, including: obtaining a sample image for training, a first prompt feature to be trained, and a second prompt feature to be trained, where the sample image includes a real sample target object or an unreal sample target object, the first prompt feature to be trained is associated with the real attribute of the sample target object, and the second prompt feature to be trained is associated with the unreal attribute of the sample target object; performing image detection on the sample image to obtain a sample image detection result; and based on the sample image and the sample image detection result, training through a preset loss function to optimize the first prompt feature to be trained and the second prompt feature to be trained.

[0008] According to another aspect of the present disclosure, there is provided an image detection device, including: a feature acquisition module configured to acquire an image feature of a to-be-detected image including a target object, and a first prompt feature and a second prompt feature for detecting the to-be-detected image, where the first prompt feature is associated with the real attribute of the target object, and the second prompt feature is associated with the unreal attribute of the target object; a feature processing module configured to fuse the first prompt feature and the image feature to obtain a first reference feature, and fuse the second prompt feature and the image feature to obtain a second reference feature; a similarity determination module configured to determine a first similarity between the first reference feature and the image feature, and a second similarity between the second reference feature and the image feature; and a detection result determination module configured to determine an image detection result based on the first similarity and the second similarity, where the image detection result indicates whether the to-be-detected image is real.

[0009] According to another aspect of the present disclosure, there is provided an image detection training device, including: a data acquisition module configured to acquire a sample image for training, a first prompt feature to be trained, and a second prompt feature to be trained, where the sample image includes a real sample target object or an unreal sample target object, the first prompt feature to be trained is associated with the real attribute of the sample target object, and the second prompt feature to be trained is associated with the unreal attribute of the sample target object; an image detection module configured to perform image detection on the sample image to obtain a sample image detection result; and a training execution module configured to train through a preset loss function based on the sample image and the sample image detection result to optimize the first prompt feature to be trained and the second prompt feature to be trained.

[0010] According to another aspect of the present disclosure, there is provided an electronic device, including at least one processor; and a memory communicatively connected to the at least one processor; the memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to execute the method as described above in the present disclosure.

[0011] According to another aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium storing computer instructions for causing a computer to execute the method as described above in the present disclosure.

[0012] According to another aspect of the present disclosure, there is provided a computer program product including a computer program which, when executed by a processor, implements the method as described above in the present disclosure.

[0013] According to one or more embodiments of the present disclosure, the accuracy of image detection can be improved.

[0014] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] The drawings exemplarily show embodiments and form a part of the description, and are used together with the written description of the description to explain the exemplary implementation manners of the embodiments. The shown embodiments are only for illustrative purposes and do not limit the scope of the claims. In all the drawings, the same reference numerals refer to similar but not necessarily identical elements.

[0016] Figure 1 A schematic diagram showing an exemplary system in which various methods described herein can be implemented according to an embodiment of the present disclosure;

[0017] Figure 2 A flowchart showing an image detection method according to an embodiment of the present disclosure;

[0018] Figure 3 A schematic diagram showing obtaining image features using an image encoder in a vision-language model according to an embodiment of the present disclosure;

[0019] Figure 4 A schematic diagram showing obtaining a first prompt feature and a second prompt feature using a text encoder in a vision-language model according to an embodiment of the present disclosure;

[0020] Figure 5 A schematic diagram showing visual condition injection and first and second prompt feature sequences constructed under visual condition injection according to an embodiment of the present disclosure;

[0021] Figure 6 A schematic diagram showing feature aggregation according to an embodiment of the present disclosure;

[0022] Figure 7 A schematic diagram showing determining an image detection result according to an embodiment of the present disclosure;

[0023] Figure 8 The flowchart of an image detection training method according to an embodiment of the present disclosure is shown;

[0024] Figure 9 The schematic diagram of an image detection method based on a vision - language model according to an embodiment of the present disclosure is shown;

[0025] Figure 10 The structural block diagram of an image detection device according to an embodiment of the present disclosure is shown;

[0026] Figure 11 The structural block diagram of an image detection device according to another embodiment of the present disclosure is shown;

[0027] Figure 12 The structural block diagram of an image detection training device according to an embodiment of the present disclosure is shown;

[0028] Figure 13 The structural block diagram of an exemplary electronic device that can be used to implement the embodiments of the present disclosure is shown. Detailed implementation manners

[0029] The following describes exemplary embodiments of the present disclosure with reference to the accompanying drawings. Various details of the embodiments of the present disclosure are included to assist understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope of the present disclosure. Similarly, for clarity and conciseness, descriptions of well - known functions and structures are omitted in the following description.

[0030] In the present disclosure, unless otherwise specified, the terms "first", "second", etc. are used to describe various elements and are not intended to limit the positional relationship, temporal relationship, or importance relationship of these elements. Such terms are only used to distinguish one element from another. In some examples, the first element and the second element may refer to the same instance of the element, and in certain cases, based on the context description, they may also refer to different instances.

[0031] In the description of various examples in the present disclosure, the terms used are only for the purpose of describing specific examples and are not intended to be restrictive. Unless the context clearly indicates otherwise, if the number of elements is not specifically limited, the element may be one or more. In addition, the term "and / or" used in the present disclosure covers any one of the listed items and all possible combinations.

[0032] In the related art, traditional image detection techniques generally rely on supervised binary classification training, and classify by learning to map input features to a fixed decision boundary. This method may have the following problems: First, the fixed decision boundary has limitations and cannot evolve with the emergence of new forgery techniques; second, the diversity and adaptability of feature representations are insufficient, making it difficult to comprehensively capture the diverse features of forged content.

[0033] To this end, embodiments of the present disclosure provide a more effective image detection method and an image detection training method.

[0034] Embodiments of the present disclosure will be described in detail below with reference to the accompanying drawings.

[0035] Figure 1 FIG. shows a schematic diagram of an exemplary system 100 in which the various methods and apparatuses described herein can be implemented according to an embodiment of the present disclosure. Referring to Figure 1 , the system 100 includes one or more client devices 101, 102, 103, 104, 105, and 106, a server 120, and one or more communication networks 110 that couple the one or more client devices to the server 120. The client devices 101, 102, 103, 104, 105, and 106 can be configured to execute one or more applications.

[0036] In an embodiment of the present disclosure, the server 120 can run one or more services or software applications that enable the execution of the methods described in the embodiments of the present disclosure.

[0037] In certain embodiments, the server 120 can also provide other services or software applications, and these services or software applications can include non-virtual environments and virtual environments. In certain embodiments, these services can be provided as web-based services or cloud services, for example, provided to users of the client devices 101, 102, 103, 104, 105, and / or 106 under a software as a service (SaaS) model.

[0038] In Figure 1 the configuration shown, the server 120 can include one or more components that implement the functions performed by the server 120. These components can include software components, hardware components, or a combination thereof that can be executed by one or more processors. Users operating the client devices 101, 102, 103, 104, 105, and / or 106 can in turn use one or more client applications to interact with the server 120 to utilize the services provided by these components. It should be understood that various different system configurations are possible, which can be different from the system 100. Therefore, Figure 1 is an example of a system for implementing the various methods described herein and is not intended to be limiting.

[0039] Users may use client devices 101, 102, 103, 104, 105, and / or 106 to provide images, prompt texts, etc. The client devices may provide an interface that enables the users of the client devices to interact with the client devices. The client devices may also output information to the users via the interface. Although Figure 1 only six client devices are depicted, those skilled in the art will be able to understand that the present disclosure may support any number of client devices.

[0040] Client devices 101, 102, 103, 104, 105, and / or 106 may include various types of computer devices, such as portable handheld devices, general-purpose computers (such as personal computers and laptop computers), workstation computers, wearable devices, smart screen devices, self-service terminal devices, service robots, gaming systems, thin clients, various messaging devices, sensors, or other sensing devices, etc. These computer devices may run various types and versions of software applications and operating systems, such as MICROSOFT Windows, APPLE iOS, UNIX-like operating systems, Linux or Linux-like operating systems (such as GOOGLE Chrome OS); or include various mobile operating systems, such as MICROSOFT WindowsMobile OS, iOS, Windows Phone, Android. Portable handheld devices may include cellular phones, smartphones, tablets, personal digital assistants (PDAs), etc. Wearable devices may include head-mounted displays (such as smart glasses) and other devices. Gaming systems may include various handheld gaming devices, Internet-enabled gaming devices, etc. The client devices are capable of executing various different applications, such as various Internet-related applications, communication applications (such as email applications), short message service (SMS) applications, and may use various communication protocols.

[0041] Network 110 may be any type of network well-known to those skilled in the art, which may support data communication using any one of a variety of available protocols (including but not limited to TCP / IP, SNA, IPX, etc.). By way of example only, one or more networks 110 may be a local area network (LAN), an Ethernet-based network, token ring, wide area network (WAN), the Internet, virtual network, virtual private network (VPN), intranet, extranet, blockchain network, public switched telephone network (PSTN), infrared network, wireless network (such as Bluetooth, WIFI), and / or any combination of these and / or other networks.

[0042] Server 120 may include one or more general-purpose computers, dedicated server computers (such as PC (Personal Computer) servers, UNIX servers, midrange servers), blade servers, mainframes, server clusters, or any other suitable arrangement and / or combination. Server 120 may include one or more virtual machines running a virtual operating system, or other computing architectures involving virtualization (such as one or more flexible pools of logical storage devices that can be virtualized to maintain virtual storage devices of the server). In various embodiments, server 120 may run one or more services or software applications that provide the functions described below.

[0043] The computing units in server 120 may run one or more operating systems including any of the above operating systems as well as any commercially available server operating systems. Server 120 may also run any one of a variety of additional server applications and / or middleware applications, including HTTP servers, FTP servers, CGI servers, JAVA servers, database servers, and the like.

[0044] In some embodiments, server 120 may include one or more applications to analyze and merge data feeds and / or event updates received from users of client devices 101, 102, 103, 104, 105, and / or 106. Server 120 may also include one or more applications to display data feeds and / or real-time events via one or more display devices of client devices 101, 102, 103, 104, 105, and / or 106.

[0045] In some embodiments, server 120 may be a server of a distributed system, or a server incorporating blockchain. Server 120 may also be a cloud server, or an intelligent cloud computing server or intelligent cloud host with artificial intelligence technology. A cloud server is a host product in the cloud computing service system, which solves the defects of high management difficulty and weak business scalability existing in traditional physical hosts and virtual private server (VPS, Virtual Private Server) services.

[0046] System 100 may also include one or more databases 130. In certain embodiments, these databases may be used to store data and other information. For example, one or more of the databases 130 may be used to store information such as audio files and video files. The databases 130 may reside in various locations. For example, a database used by the server 120 may be local to the server 120, or may be remote from the server 120 and may communicate with the server 120 via a network-based or dedicated connection. The databases 130 may be of different types. In certain embodiments, a database used by the server 120 may be, for example, a relational database. One or more of these databases may store, update, and retrieve data to and from the database in response to commands.

[0047] In certain embodiments, one or more of the databases 130 may also be used by an application to store application data. The database used by the application may be a different type of database, such as a key-value store, an object store, or a conventional store supported by a file system.

[0048] Figure 1 System 100 can be configured and operated in various ways to enable the application of the various methods and apparatuses described according to the present disclosure.

[0049] Aspects of an image detection method according to embodiments of the present disclosure will be described in detail below.

[0050] Figure 2 A flowchart of an image detection method 200 according to an embodiment of the present disclosure is shown.

[0051] As Figure 2 shown, the image detection method 200 includes steps S201, S202, S203, and S204.

[0052] In step S201, image features of a to-be-detected image including a target object, as well as a first hint feature and a second hint feature for detecting the to-be-detected image, are obtained. The first hint feature is associated with a true attribute of the target object. The second hint feature is associated with a non-true attribute of the target object.

[0053] In the example, in some application scenarios, with the development of generative artificial intelligence technology, it has gradually become possible to generate realistic images of specific things through corresponding artificial intelligence models without relying on real image acquisition or shooting. In this case, a need for verifying the authenticity of images has also emerged. In this method 200, the authenticity of the image to be detected is unknown, so it will be determined through this method 200. The target object in the image to be detected may include the main thing in the image to be detected, such as a person, an animal, a building, etc. In addition, the image to be detected may only include the target object, or may include other content in addition to the target object. Preprocessing for standardization, such as cropping, resizing (such as 224×224), etc., can be performed on the image to be detected to make it more conducive to the execution of image detection.

[0054] In the example, image features can be used to characterize the image information of the image to be detected. The image features can be global image features of the image to be detected or local image features of the target object in the image to be detected. The image features can be represented in the form of a feature sequence. For this purpose, the image to be measured can be divided into K image blocks (K is a natural number), so the image feature V can be represented as a feature sequence of v1, v2, v3…v k In addition, the image features can be d-dimensional (d is a natural number), so the image feature V can also be represented as

[0055] In the example, the true attributes of the target object can refer to the characteristics or properties that the real target object should possess, and these characteristics or properties conform to the natural laws, physical laws, etc. of the real world. For example, for a person's face, the eyebrows should be above the two eyes, and the nose should be between the two eyes, etc. On the contrary, the non-true attributes of the target object can refer to the characteristics or properties that the target object should not possess. For example, if a person in the image to be detected has an eye on the forehead, obviously this does not conform to the appearance characteristics or properties of a person in the real world, so it can be regarded as a non-true attribute. In the application scenarios of traditional image modification via image processing technology or in the above-mentioned generative artificial intelligence application scenarios, there may be cases where there are some forgery traces on the person, and these can all be regarded as examples of non-true attributes.

[0056] In the example, the first hint feature being associated with the true attribute of the target object may mean that the first hint feature is directly used to characterize the true attribute of the target object, or it may mean that the first hint feature is used to characterize the true attribute of the target object after specific processing. That is, the first hint feature may contain information for hinting at the true attribute of the target object. In addition, the first hint feature may also contain information in other aspects, such as fusing the image information of the image to be detected into the first hint feature using specific processing. Similarly, the second hint feature is the same. Therefore, the second hint feature may contain information for hinting at the non-true attribute of the target object. In addition, the second hint feature may also contain information in other aspects, such as fusing the image information of the image to be detected into the second hint feature using specific processing.

[0057] In the example, both the first hint feature and the second hint feature may include features in multiple aspects. For example, they may include M (M is a natural number) sub-features, and each sub-feature may be d-dimensional (d is a natural number). Therefore, both the first hint feature and the second hint feature can be represented in the form of a feature sequence. For example, the first hint feature R can be represented as a feature sequence of M sub-features r1, r2, r3…r m After splicing, it can be represented as Similarly, the second hint feature F can be represented as a feature sequence of M sub-features f1, f2, f3…f m After splicing, it can be represented as

[0058] In step S202, the first hint feature and the image feature are fused to obtain a first reference feature, and the second hint feature and the image feature are fused to obtain a second reference feature.

[0059] In the example, information of different modalities can be fused. For example, information of one modality can be fused with information of another modality. In this step S202, the image information represented by the image feature can be fused with the information for hinting at the true attribute represented by the first hint feature to obtain a first reference feature representing this fused information. Similarly, the image information represented by the image feature can be fused with the information for hinting at the non-true attribute represented by the second hint feature to obtain a second reference feature representing this fused information.

[0060] In the example, since both the first hint feature and the second hint feature may include features in multiple aspects, dynamic aggregation of multi-aspect features can also be achieved through information fusion. Thus, the multi-aspect features can be adaptively combined according to the image feature to form comprehensive first and second reference features. Correspondingly, the first reference feature can be represented as r aggregated, the second reference feature can be represented as f aggregated .

[0061] In step S203, determine a first similarity between the first reference feature and the image feature, and a second similarity between the second reference feature and the image feature.

[0062] In the example, since the first reference feature obtained through step S202 combines the image information represented by the image feature and the information represented by the first hint feature for hinting at the true attribute, by comparing the similarity between the first reference feature and the image feature, the proximity of the image to be detected to the true attribute can be determined. Similarly, since the second reference feature combines the image information represented by the image feature and the information represented by the second hint feature for hinting at the non-true attribute, by comparing the similarity between the second reference feature and the image feature, the proximity of the image to be detected to the non-true attribute can be determined. That is, by determining the first similarity between the first reference feature and the image feature, and the second similarity between the second reference feature and the image feature, it can be determined whether the image to be detected is closer to the true attribute or the non-true attribute. Both the first similarity and the second similarity can be scored by calculating the cosine similarity. The first similarity can be represented as s real , and the second similarity can be represented as s fake .

[0063] In step S204, based on the first similarity and the second similarity, determine an image detection result, which indicates whether the image to be detected is real.

[0064] In the example, the magnitudes of the first similarity and the second similarity can be compared to determine the image detection result. Additionally, the first similarity and the second similarity can also be normalized to probability values or confidence scores to determine the image detection result.

[0065] According to an embodiment of the present disclosure, an image detection method based on dynamic criteria is proposed. This method constructs an adaptive authenticity evaluation mechanism to dynamically construct customized judgment criteria for each different image to be detected. This is achieved through the fusion of the image information represented by the image features and the information represented by the first hint feature for hinting at the real attribute, as well as the fusion of the image information represented by the image features and the information represented by the second hint feature for hinting at the non-real attribute. Different from the traditional detection methods that rely on static decision boundaries, this method transforms the static binary classification decision mechanism into a dynamic similarity comparison mechanism by constructing an adaptive real first reference feature and non-real second reference feature, changing the technical path of static comparison in traditional detection methods, greatly improving the accuracy of image detection, and being able to effectively cope with the rapidly evolving new forgery technologies, providing a transformation of the technical paradigm for digital media authenticity verification.

[0066] In this embodiment, the image to be detected of the target object can be from a public dataset, or the acquisition of the image to be detected is authorized by the user corresponding to the target object.

[0067] In the technical solution of the present disclosure, the collection, storage, use, processing, transmission, provision, and disclosure of the user's personal information involved in the processing all comply with the provisions of relevant laws and regulations and do not violate public order and good customs.

[0068] In some embodiments, as combined with Figure 2 obtaining the image features of the image to be detected including the target object in step S201 may include: providing the image to be detected to an image encoder in a Vision-Language Model (VLM) that supports at least visual and language modality tasks for encoding to obtain the image features.

[0069] Figure 3 FIG. shows a schematic diagram of obtaining image features by using an image encoder in a vision-language model according to an embodiment of the present disclosure.

[0070] As Figure 3 shown, the image to be detected 301 can be provided to the image encoder 310 to obtain the image feature V. The image feature V can be represented as v1, v2, v3... v kFeature sequence. In the example, the image encoder 310 can be an inherent component in a vision-language model. The image encoder 310 can include a Residual Network (ResNet) or a Vision Transformer (ViT), which is used to map the input image into a feature vector. A vision-language model is a deep learning model that combines at least visual information (such as images, videos) and language information (such as text, speech) to support at least visual modality and language modality tasks. Common vision-language models can include Contrastive Language-Image Pre-training (CLIP) models, Aligned Language and Image Generation (ALIGN) models, Florence models, etc.

[0071] Therefore, the method of the embodiments of the present disclosure can be executed while keeping the infrastructure of the vision-language model unchanged, so as to utilize the vision-language model to construct an adaptive authenticity evaluation mechanism, thereby improving the executability and operability of implementing this method.

[0072] In some embodiments, as combined with Figure 2 The step S201 of obtaining the first prompt feature and the second prompt feature for detecting the image to be detected may include: constructing a first prompt feature sequence and a second prompt feature sequence; and fusing the features in the first prompt feature sequence to obtain the first prompt feature, and fusing the features in the second prompt feature sequence to obtain the second prompt feature. The first prompt feature sequence may include a first initial prompt feature for characterizing the true attributes of the target object and a third prompt feature associated with the inherent attributes of the target object. The second prompt feature sequence may include a second initial prompt feature for characterizing the non-true attributes of the target object and the third prompt feature.

[0073] In the example, compared with the first / second prompt feature, as the name implies, the first / second initial prompt feature can be understood as a direct or original information expression for prompting the true / non-true attributes of the target object, so it is called the "initial" prompt feature. Similar to the first / second prompt feature, the first initial prompt feature R' may include M sub-features r'1, r'2, r'3... r' m , and the second initial prompt feature F' may include M sub-features f'1, f'2, f'3... f' m .

[0074] In the example, similar to the first / second hint features, the third hint feature being associated with the inherent attribute of the target object may mean that the third hint feature is directly used to represent the inherent attribute of the target object, or it may mean that the third hint feature is used to represent the inherent attribute of the target object after undergoing specific processing. The inherent attribute can reflect the characteristics or features that both the real target object and the non-real target object possess, and thus can also be referred to as general visual attributes. The third hint feature C may include N (N is a natural number) sub-features c1, c2, c3... c n .

[0075] Since in the fusion process, it is allowed that when a certain element in the sequence is processed, other elements in the sequence are considered simultaneously, elements expected to be associated can be constructed into the same sequence for processing. That is, in this method, it is desired to make the first initial hint feature representing the real attribute associated with the third hint feature representing the inherent attribute, and also make the second initial hint feature representing the non-real attribute associated with the third hint feature representing the inherent attribute. Thus, it can be promoted that the image detection process does not focus on the characteristics or features that both the real target object and the non-real target object possess, and further can help improve the effect and efficiency of image detection.

[0076] In some embodiments, the steps of fusing the features in the first hint feature sequence to obtain the first hint feature, and fusing the features in the second hint feature sequence to obtain the second hint feature may include: providing the first hint feature sequence to the text encoder in the vision-language model for encoding to obtain the first hint feature; and providing the second hint feature sequence to the text encoder in the vision-language model for encoding to obtain the second hint feature. As described above, this vision-language model can be used to support at least vision modality and language modality tasks.

[0077] Figure 4 Shows a schematic diagram of obtaining the first hint feature and the second hint feature using the text encoder in the vision-language model according to an embodiment of the present disclosure.

[0078] As Figure 4 shown, the first hint feature sequence R_seq can be provided to the text encoder 410 to obtain the first hint feature R. Similarly, the first hint feature sequence F_seq can be provided to the text encoder 410 to obtain the second hint feature F.

[0079] In addition, as described above, since the first hint feature R may include M sub-features r1, r2, r3... r m , the second hint feature F may include M sub-features f1, f2, f3... f m, corresponding hint feature sequences can be set one by one corresponding to these sub - features. For example, the first hint feature sequence R_seq can also include M subsequences, where Figure 4 schematically shows two subsequences. The second hint feature sequence F_seq can also include M subsequences, where Figure 4 also schematically shows two subsequences. As Figure 4 shown, the first subsequence in the first hint feature sequence R_seq can include the third hint features c1, c2, c3... c n and the first sub - feature r’1 in the first initial hint feature R’. The second subsequence in the first hint feature sequence R_seq can include the third hint features c1, c2, c3... c n and the second sub - feature r’2 in the first initial hint feature R’. Similarly, the first subsequence in the second hint feature sequence F_seq can include the third hint features c1, c2, c3... c n and the first sub - feature f’1 in the second initial hint feature F’. The second subsequence in the second hint feature sequence F_seq can include the third hint features c1, c2, c3... c n and the second sub - feature f’2 in the second initial hint feature F’.

[0080] In the example, similar to the image encoder 310 as Figure 3 shown, the text encoder 410 as Figure 4 shown can also be an inherent component in a vision - language model. The text encoder 410 can include a Transformer architecture, which is used to convert the input text into a feature vector. This Transformer architecture can provide a self - attention mechanism. In the case where the input text itself is expressed in the form of a feature sequence, the self - attention mechanism provided by this Transformer architecture can also be used to process the input feature sequence so that the elements in the feature sequence have the desired associations with each other.

[0081] Thus, similar to using the image encoder in a vision-language model, the method of the embodiments of the present disclosure can be executed while keeping the infrastructure of the existing vision-language model unchanged, so as to utilize the vision-language model to construct an adaptive authenticity evaluation mechanism, thereby enhancing the executability and operability of implementing this method. At the same time, since the text encoder in the vision-language model can provide the self-attention mechanism required in the image detection architecture of this method, the third prompt feature representing the inherent attribute can be respectively associated with the first initial prompt feature representing the real attribute and the second initial prompt feature representing the non-real attribute by means of this text encoder, thereby prompting the image detection process not to focus on the characteristics or features that both the real target object and the non-real target object possess, so as to improve the effect and efficiency of image detection.

[0082] Figure 5 A schematic diagram showing the first prompt feature sequence and the second prompt feature sequence constructed under visual condition injection and visual condition injection according to an embodiment of the present disclosure is shown.

[0083] In this article, visual condition injection may refer to associating the image features of the image to be detected with specific features. In this method, selectively only the third prompt feature related to the inherent attribute is related to this visual condition injection, so that the third prompt feature can not only contain the prompt information related to the inherent attribute, but also contain the image information related to the image to be detected. In other words, this visual condition injection does not involve the first initial prompt feature representing the real attribute and the second initial prompt feature representing the non-real attribute, which can maintain the semantic purity and semantic role of these two, and ensure the structural integrity of the information used as a prompt.

[0084] In some embodiments, the above steps of constructing the first prompt feature sequence and the second prompt feature sequence may include: processing the image features and the third initial prompt feature based on the cross-attention mechanism to obtain the third prompt feature, where the third initial prompt feature is used to characterize the inherent attribute of the target object; and splicing the third prompt feature and the first initial prompt feature into the first prompt feature sequence, and splicing the third prompt feature and the second initial prompt feature into the second prompt feature sequence.

[0085] In the example, the step of processing the image features and the third initial prompt feature based on the cross-attention mechanism to obtain the third prompt feature, i.e., visual condition injection, can mean associating the image features of the image to be detected with the third initial prompt feature, so that the generated third prompt feature can contain both the prompt information related to the inherent attributes and the image information related to the image to be detected. It can be understood that the third initial prompt feature mentioned here, similar to the aforementioned first / second initial prompt features, can be understood as a direct or original information expression that prompts the inherent attributes of the target object, so it is called the "initial" prompt feature. The third initial prompt feature C' can include N sub-features c'1, c'2, c'3... c' n .

[0086] In the example, as Figure 5 shown, the image features v1, v2, v3... v k and the third initial prompt features c'1, c'2, c'3... c' n can be processed based on the cross-attention mechanism to obtain the third prompt features c1, c2, c3... c n . Then, the first sub-sequence in the first prompt feature sequence R_seq can be formed by concatenating the third prompt features c1, c2, c3... c n and the first sub-feature r'1 in the first initial prompt feature R'. The second sub-sequence in the first prompt feature sequence R_seq can be formed by concatenating the third prompt features c1, c2, c3... c n and the second sub-feature r'2 in the first initial prompt feature R' ( Figure 5 only two sub-sequences are shown schematically). Similarly, the first sub-sequence in the second prompt feature sequence F_seq can be formed by concatenating the third prompt features c1, c2, c3... c n and the first sub-feature f'1 in the second initial prompt feature F'. The second sub-sequence in the second prompt feature sequence F_seq can be formed by concatenating the third prompt features c1, c2, c3... c n and the second sub-feature f'2 in the second initial prompt feature F'.

[0087] Through the above operations, visual condition injection can be achieved and the first prompt feature sequence and the second prompt feature sequence can be constructed under visual condition injection. Thus, the image information related to the image to be detected can be injected into the prompt information related to the inherent attributes, so as to enable this image detection method to dynamically adapt to the personalized image information of different images to be detected, which helps to dynamically construct customized judgment criteria for each different image to be detected to achieve a dynamic similarity comparison mechanism.

[0088] In some embodiments, the step of processing the image features and the third initial prompt feature based on the cross-attention mechanism to obtain the third prompt feature may include: performing the cross-attention mechanism by setting the third initial prompt feature as the query vector and setting the image features as the key vector and the value vector to obtain the attention calculation feature; and fusing the third initial prompt feature and the attention calculation feature to obtain the third prompt feature.

[0089] In an example, this process may correspond to the process of visual condition injection. As Figure 5 shown, the third initial prompt features c’1, c’2, c’3…c’ n can be set as the query vector (Query, also denoted as Q), and the image features v1, v2, v3…v k can be set as the key vector (Key, also denoted as K) and the value vector (Value, also denoted as V), and the cross-attention mechanism is performed to obtain the attention calculation feature Result Attention . Then, the third initial prompt features c’1, c’2, c’3…c’ n and the attention calculation feature Result Attention can be fused to obtain the third prompt features c1, c2, c3…c n .

[0090] Therefore, the visual condition injection implemented based on the cross-attention mechanism can bring about fine-grained visual and text interaction, enabling the prompt information to capture specific visual information of the input, thereby contributing to dynamically constructing customized judgment criteria for each different image to be detected.

[0091] In some embodiments, the cross-attention mechanism may include a multi-head cross-attention mechanism such that the attention calculation feature is obtained based on multiple attention heads.

[0092] In an example, the number of attention heads can be selected according to the type of the image encoder in the vision-language model, such as 64, 96, or 128 attention heads. Additionally, each attention head can be used to process features such as 8-dimensional features. Correspondingly, the multi-head cross-attention can be calculated by the expression MultiheadAttention(Q, K, V) = Concat(head1, head2, …, head h )W O where the Concat function represents concatenating the output results of the multiple attention heads head1, head2…head h respectively, and W O represents the output weight, which can be learned during the training process. Each attention head can be calculated by the expression head i= Attention(QW i Q , KW i K , VW i V ) calculation, where W i Q , W i K , W i V represent the weights of the query vector Q, the key vector K, and the value vector V respectively. The Attention function can be calculated by the standard scaled dot-product attention calculation, where d k represents the dimension of the key vector K.

[0093] Correspondingly, the above visual condition injection can be calculated by the expression C = C'+ MultiheadAttention(C', V, V), where C represents the third hint feature, C' represents the third initial hint feature, and it should be noted that V in this expression represents the image feature. Therefore, MultiheadAttention(C', V, V) is the above attention calculation feature.

[0094] The multi-head cross-attention can allow capturing the influence of different visual aspects on the hint information, and thus can achieve visual condition injection in a more fine-grained manner, thereby dynamically constructing customized judgment criteria for each different image to be detected.

[0095] In some embodiments, the first hint feature sequence and the second hint feature sequence may further include a first additional feature for indicating the start of the sequence, a second additional feature for indicating the end of the sequence, and a third additional feature for indicating the target object.

[0096] In the example, as Figure 5 shown, the first hint feature sequence R_seq may include the sequence start marker [SOS], that is, the first additional feature, and the sequence end marker [EOS], that is, the second additional feature. In addition, the first hint feature sequence R_seq may further include the third additional feature "face." for indicating that the face is the target object. It can be understood that Figure 5 the face is used as an example of the target object, but the scope of the present disclosure is not limited thereto.

[0097] In the example, in the case of implementing the present image detection method based on the CLIP model, the above three additional features can all be obtained by vector retrieval from the original embedding table E original of the CLIP model itself. The original embedding table E original can be derived from the original knowledge of the CLIP model.

[0098] In the example, compared with the original embedding table E original Similarly, a learnable embedding table E that can be optimized during training can be maintained separately learnable The learnable embedding table E learnable can include M first initial prompt features r’1, r’2, r’3…r’ m , M second initial prompt features f’1, f’2, f’3…f’ m , and N third initial prompt features c’1, c’2, c’3…c’ n , each of which can be d-dimensional. Accordingly, the learnable embedding table E learnable can be represented as

[0099] Therefore, by adding special tokens for indicating the start and end and special tokens for indicating the target object to the sequence, more auxiliary information can be introduced into the sequence to be processed by the self-attention mechanism, so as to improve the processing efficiency. At the same time, since the vectors corresponding to these tokens can be retrieved from the original knowledge of a vision-language model such as the CLIP model, there is no need to change the existing basic architecture of the vision-language model too much, which is beneficial to the executability and operability of the present image detection method.

[0100] In some embodiments, as Figure 5 shown, in the first prompt feature sequence R_seq, the first additional feature [SOS], the third prompt features c1, c2, c3…c n , the first initial prompt features r’1 / r’2, the third additional feature "face.", and the second additional feature [EOS] can be concatenated. Similarly, in the second prompt feature sequence F_seq, the first additional feature [SOS], the third prompt features c1, c2, c3…c n , the second initial prompt features f’1 / f’2, the third additional feature "face.", and the second additional feature [EOS] can be concatenated.

[0101] In the example, based on a specific concatenation order, the input text format of the prompt information can be set. For example, the following prompt template can be constructed: [SOS]+Context tokens +Real aspect_token / Fake aspect_token +"face."+[EOS], where Context tokens can correspond to the third initial prompt feature, Real aspect_token can correspond to the first initial prompt feature vector, Fake aspect_tokenIt may correspond to the second initial prompt feature vector. The input text conforming to the above template can be tokenized by a tokenizer to obtain multiple tokens. Among them, [SOS], [EOS], and "face." in the multiple tokens can be obtained by vector retrieval in the original embedding table E original and the Context tokens , Real aspect_token , Fake aspect_token in the multiple tokens can be obtained by vector retrieval in the learnable embedding table E learnable .

[0102] In this way, the standardized preset format on the sequence can be used to simplify the subsequent understanding of the sequence in the self-attention mechanism, so as to improve the processing efficiency.

[0103] In some embodiments, as described in combination with Figure 2 step S202 of fusing the first prompt feature and the image feature to obtain the first reference feature, and fusing the second prompt feature and the image feature to obtain the second reference feature, it may include: processing the first prompt feature and the image feature based on the cross-attention mechanism to obtain the first reference feature, and processing the second prompt feature and the image feature based on the cross-attention mechanism to obtain the second reference feature.

[0104] In the example, the cross-attention mechanism can be used to fuse the image information represented by the image feature and the information for prompting the real attribute represented by the first prompt feature to obtain the first reference feature representing the fused information. Similarly, the cross-attention mechanism can be used to fuse the image information represented by the image feature and the information for prompting the non-real attribute represented by the second prompt feature to obtain the second reference feature representing the fused information.

[0105] In this way, the inherent fusion mechanism provided by the cross-attention can be effectively used to fuse the image information and the corresponding prompt information, which is beneficial to the execution and implementation of the overall method.

[0106] Figure 6 Shows a schematic diagram of feature aggregation according to an embodiment of the present disclosure.

[0107] In this article, feature aggregation may involve the process of converting the first / second prompt feature into the first / second reference feature based on the cross-attention mechanism. When each type of prompt feature contains information in multiple aspects, the feature aggregation can be used to dynamically combine the information in the multiple aspects, and the dynamic combination process combines the characteristics or features of the image to be detected itself to achieve adaptive calculation.

[0108] In some embodiments, the steps of processing the first prompt feature and the image feature based on the cross-attention mechanism to obtain the first reference feature, and processing the second prompt feature and the image feature based on the cross-attention mechanism to obtain the second reference feature may include: performing the cross-attention mechanism by setting the image feature as the query vector and setting the first prompt feature as the key vector and the value vector to obtain the first reference feature; and performing the cross-attention mechanism by setting the image feature as the query vector and setting the second prompt feature as the key vector and the value vector to obtain the second reference feature.

[0109] In an example, this process may correspond to the process of feature aggregation. As Figure 6 shown, the cross-attention mechanism can be performed by setting the image feature V as the query vector Query, setting the first prompt feature R as the key vector Key and the value vector Value, to obtain the first reference feature r aggregated . Similarly, the cross-attention mechanism can be performed by setting the image feature V as the query vector Query, setting the second prompt feature F as the key vector Key and the value vector Value, to obtain the second reference feature f aggregated .

[0110] Therefore, the feature aggregation implemented based on the cross-attention mechanism can dynamically combine the prompt information from multiple aspects to perform adaptive calculation for the image to be detected itself, thereby helping to dynamically construct customized judgment criteria for each different image to be detected.

[0111] In some embodiments, the cross-attention mechanism may include a multi-head cross-attention mechanism, such that both the first reference feature and the second reference feature are obtained based on multiple attention heads.

[0112] In an example, the first reference feature r aggregated can be calculated by the expression r aggregated =MultiheadAttention(V,R,R), where V represents the image feature and R represents the first prompt feature. Similarly, the second reference feature f aggregated can be calculated by the expression f aggregated =MultiheadAttention(V,F,F), where V represents the image feature and F represents the second prompt feature.

[0113] Feature aggregation can be more finely grained through multi-head cross-attention, thereby dynamically constructing customized judgment criteria for each different image to be detected.

[0114] Figure 7 FIG. shows a schematic diagram of determining an image detection result according to an embodiment of the present disclosure.

[0115] In some embodiments, as combined with Figure 2 step S204 described above to determine the image detection result based on the first similarity and the second similarity, it may include: in response to the score of the first similarity being greater than the score of the second similarity, determining that the image to be detected is real; and in response to the score of the second similarity being greater than the score of the first similarity, determining that the image to be detected is not real.

[0116] In an example, as Figure 7 shown, the score s of the first similarity real can be calculated by the expression s real = cos(V, r aggregated ) / τ, where the cos function represents cosine similarity, V represents the image feature, r aggregated represents the first reference feature, and τ represents the temperature parameter, which can be inherited from a vision-language model such as CLIP. Similarly, the score s of the second similarity fake can be calculated by the expression s fake = cos(V, f aggregated ) / τ, where the cos function represents cosine similarity, V represents the image feature, and f aggregated represents the second reference feature, and τ represents the temperature parameter.

[0117] Since the first similarity represents the degree of proximity of the image to be detected to the real attribute, and the second similarity represents the degree of proximity of the image to be detected to the non-real attribute, by comparing the first similarity and the second similarity, the traditional static binary classification decision mechanism is transformed into a dynamic similarity comparison mechanism, thereby overcoming the traditional detection method that relies on static decision boundaries.

[0118] In some embodiments, the target object in the image to be detected may include a person's face, and the non-real attribute of the target object may include forged traces on the face.

[0119] In an example, this image detection method is particularly beneficial for detecting the authenticity of face images. Generally speaking, non-real face images may include various forged traces, such as unnatural face edges, inconsistent skin colors, inconsistent lighting, etc. These forged traces can all be used as hint information for the non-reality of the face.

[0120] Since in practical applications, mature experience may have been formed on the characteristics of forged face images, the use of such forged trace information can provide basic data for training and learning for this image detection method, so that the image detection method obtained based on this training and learning can more effectively cope with evolving new forgery technologies.

[0121] It can be understood that although the description in this text takes the human face as an example, the scope of the present disclosure is not limited thereto.

[0122] According to an embodiment of the present disclosure, an image detection training method is also provided.

[0123] Figure 8 The flowchart of an image detection training method 800 according to an embodiment of the present disclosure is shown.

[0124] As Figure 8 shown, the method 800 includes steps S801, S802, and S803.

[0125] In step S801, a sample image for training, a first to-be-trained prompt feature, and a second to-be-trained prompt feature are obtained. The sample image includes a real sample target object or a non-real sample target object. The first to-be-trained prompt feature is associated with the real attribute of the sample target object. The second to-be-trained prompt feature is associated with the non-real attribute of the sample target object.

[0126] In an example, the sample image may include a real image and a forged image, and the sample numbers of the two may be balanced. The real image may include a real sample target object, and the forged image may include a non-real sample target object.

[0127] In an example, the first to-be-trained prompt feature may be obtained based on a first initial to-be-trained prompt feature, and the first initial to-be-trained prompt feature may be a direct or original information expression that prompts the real attribute of the sample target object. Similarly, the second to-be-trained prompt feature may be obtained based on a second initial to-be-trained prompt feature, and the second initial to-be-trained prompt feature may be a direct or original information expression that prompts the non-real attribute of the sample target object. In the initialization stage of training, the first initial to-be-trained prompt feature and the second initial to-be-trained prompt feature may be randomly initialized, for example, initialized as random vectors with a Gaussian distribution.

[0128] In step S802, the sample image is subjected to image detection to obtain a sample image detection result.

[0129] In an example, the sample image may be subjected to image detection according to the image detection method described above (for example, combined with Figure 2 the method 200 described).

[0130] In step S803, based on the sample image and the sample image detection result, training is performed through a preset loss function to optimize the first to-be-trained prompt feature and the second to-be-trained prompt feature.

[0131] In the example, as described above, the image detection method of the embodiments of the present disclosure can be implemented based on a vision-language model such as the CLIP model, and the basic architecture of the CLIP model remains unchanged. Therefore, during the training process, the model parameters of the CLIP model can be frozen, and only the first and second prompt features to be trained are updated.

[0132] In the example, the construction of the loss function can take multiple aspects into account. For example, while ensuring accurate authenticity assessment, it can promote the learning of diverse and complementary real / fake feature representations.

[0133] In the image detection training method 800, since the first and second prompt features to be trained respectively represent the prompt information related to the real attributes of the target object and the prompt information related to the non-real attributes of the target object, the purpose of optimization through training is to enable the model to have the ability to detect the authenticity of images by identifying forged visual cues based on semantic understanding. Furthermore, it can provide more optimized prompt information features for the corresponding image detection method to improve the accuracy of image detection.

[0134] In some embodiments, the above preset loss function may include an image-prompt alignment loss. This image-prompt alignment loss can be used to represent whether the first and second prompt features to be trained are aligned with the authenticity of the sample image.

[0135] In the example, if the sample image is real, it is expected that the image detection result should show a higher first similarity corresponding to the first prompt feature to be trained. Conversely, if the sample image included is not real, it is expected that the image detection result should show a higher second similarity corresponding to the second prompt feature to be trained. That is, it is expected to align the first and second prompt features to be trained obtained through training with the authenticity of the sample image.

[0136] In the example, the image-prompt alignment loss can be calculated by the following formula 1:

[0137]

[0138] where L align represents the image-prompt alignment loss, y represents the label, y ∈ {0, 1} (0 represents real, 1 represents forged), p r and p f can be normalized through the softmax function. For example, p r can be calculated by the expression p r = e Sreal / (e Sreal + e Sfake ), and p fIt can be calculated through the expression p f = e Sfake / (e Sreal + e Sfake ), where s real and s fake respectively represent the first similarity and the second similarity obtained during training.

[0139] Therefore, the image - prompt alignment loss constructed in the above - mentioned manner can prompt the learned prompt features to correctly align with real / fake images, thus ensuring accurate authenticity assessment.

[0140] In some embodiments, the above - mentioned preset loss function may further include a prompt diversity loss. The prompt diversity loss may include an inter - class prompt diversity loss and an intra - class prompt diversity loss. The inter - class prompt diversity loss can be used to represent whether there are differences between the first prompt feature to be trained and the second prompt feature to be trained. The intra - class prompt diversity loss can be used to represent whether there is redundancy between the first prompt features to be trained of the same class and whether there is redundancy between the second prompt features to be trained of the same class.

[0141] In an example, the inter - class prompt diversity loss can be calculated by the following formula 2:

[0142]

[0143] where L between represents the inter - class prompt diversity loss, r i and f j respectively represent the first prompt feature to be trained and the second prompt feature to be trained, and M is the number of features. The cosine similarity can be normalized to the range of [0, 1] and squared to reduce the similarity of cross - category features.

[0144] In an example, the intra - class feature diversity loss can be calculated by the following formula 3:

[0145]

[0146] where L within represents the intra - class prompt diversity loss, r i and f j respectively represent the first prompt feature to be trained and the second prompt feature to be trained, and M is the number of features.

[0147] Therefore, the cross-class prompt diversity loss constructed in the above manner can prompt the difference between real and forged prompt features, while the intra-class prompt diversity loss can prompt the diversity between intra-class prompt features and avoid redundancy. Thus, diversity regularization can be achieved through the prompt diversity loss including cross-class prompt diversity loss and intra-class prompt diversity loss, so as to simultaneously consider cross-class and intra-class feature relationships and achieve more comprehensive feature learning.

[0148] In some embodiments, the preset loss function can be expressed as a weighted sum of the image-prompt alignment loss and the prompt diversity loss.

[0149] In an example, the preset loss function can be calculated by the following formula 4:

[0150]

[0151] where L represents the total loss function, L align represents the image-prompt alignment loss, and L between and L within respectively represent the cross-class prompt diversity loss and the intra-class prompt diversity loss in the prompt diversity loss, and λ represents the balance weight, which can be set to 2.0 according to experience, for example.

[0152] Therefore, by constructing the loss function for training from different dimensions, not only can accurate authenticity assessment be ensured, but also the learning of diverse and complementary real / fake feature representations can be promoted.

[0153] Figure 9 FIG. shows a schematic diagram of an image detection method based on a vision-language model according to an embodiment of the present disclosure.

[0154] In some embodiments, as Figure 9 shown, after inputting the image 901 to be detected, it can first be processed by the image encoder 910 of the vision-language model to obtain the global feature 903 of the image, where the global feature 903 of the image is represented in the form of a vector sequence.

[0155] Then, the global features 903 of the image can be incorporated into the prompt embedding sequences 902a and 902b representing real and forged respectively through visual condition injection, where the visual condition injection process can utilize the cross-attention mechanism. Then, the prompt embedding sequences 902a' and 902b' after visual condition injection can be provided to the text encoder 920 of the vision-language model to obtain the prompt feature vectors 904a and 904b representing real and forged respectively. After that, the global features 903 of the image can be respectively aggregated with the prompt feature vector 904a representing real and the prompt feature vector 904b representing forged to obtain the reference features 906a representing real and the reference features 906b representing forged, where the feature aggregation can be performed through the cross-attention mechanism.

[0156] Finally, the reference features 906a representing real and the reference features 906b representing forged can be respectively calculated with the global features 903 of the image to obtain the similarity scores representing real and the similarity scores representing forged. For example, if the similarity score representing real is less than the similarity score representing forged, it is determined that the input image is a forged image; if the similarity score representing real is greater than the similarity score representing forged, it is determined that the input image is a real image.

[0157] In some embodiments, the image detection method of the present disclosure can be applied to a social media content review system. By detecting the faces in social media videos or images, suspicious content can be quickly identified, accelerating the speed of manual review. The image detection method of the present disclosure can also be applied to a digital media authentication and news authenticity verification system. By detecting the faces of public figures in the media, the authenticity of images and videos can be quickly verified. In addition, the image detection method of the present disclosure can also be applied to many scenarios such as financial security identity verification systems, intelligent video surveillance systems, etc.

[0158] Figure 10 The structural block diagram of an image detection device 1000 according to an embodiment of the present disclosure is shown.

[0159] As Figure 10As shown, the apparatus 1000 includes a feature acquisition module 1001, a feature processing module 1002, a similarity determination module 1003, and a detection result determination module 1004. The feature acquisition module 1001 may be configured to acquire image features of a to-be-detected image including a target object, as well as a first hint feature and a second hint feature for detecting the to-be-detected image. The first hint feature is associated with the true attributes of the target object, and the second hint feature is associated with the non-true attributes of the target object. The feature processing module 1002 may be configured to fuse the first hint feature and the image features to obtain a first reference feature, and fuse the second hint feature and the image features to obtain a second reference feature. The similarity determination module 1003 may be configured to determine a first similarity between the first reference feature and the image features, and a second similarity between the second reference feature and the image features. The detection result determination module 1004 may be configured to determine an image detection result based on the first similarity and the second similarity, where the image detection result indicates whether the to-be-detected image is real.

[0160] The operations of the above-mentioned feature acquisition module 1001, feature processing module 1002, similarity determination module 1003, and detection result determination module 1004 may respectively correspond to the operations of steps S201, S202, S203, and S204 as shown in Figure 2 Therefore, the details of each aspect thereof will not be elaborated here.

[0161] Figure 11 The structural block diagram of an image detection apparatus 1100 according to another embodiment of the present disclosure is shown.

[0162] As shown in Figure 11 the apparatus 1100 includes a feature acquisition module 1101, a feature processing module 1102, a similarity determination module 1103, and a detection result determination module 1104. The operations of the above-mentioned modules may be the same as those of the feature acquisition module 1001, feature processing module 1002, similarity determination module 1003, and detection result determination module 1004 as shown in Figure 10 In addition, the above-mentioned modules may further include further sub-modules.

[0163] In some embodiments, the feature acquisition module 1101 may include an image feature acquisition module 1101a. The image feature acquisition module 1101a may be configured to provide the to-be-detected image to an image encoder in a vision-language model that supports at least vision-modal and language-modal tasks for encoding to obtain image features.

[0164] In some embodiments, the feature acquisition module 1101 may further include a prompt feature sequence construction module 1101b and a prompt feature sequence processing module 1101c. The prompt feature sequence construction module 1101b may be configured to construct a first prompt feature sequence and a second prompt feature sequence, where the first prompt feature sequence includes a first initial prompt feature for characterizing the true attributes of the target object and a third prompt feature associated with the inherent attributes of the target object, and the second prompt feature sequence includes a second initial prompt feature for characterizing the non-true attributes of the target object and the third prompt feature. The prompt feature sequence processing module 1101c may be configured to fuse the features in the first prompt feature sequence to obtain a first prompt feature and fuse the features in the second prompt feature sequence to obtain a second prompt feature.

[0165] In some embodiments, the prompt feature sequence processing module 1101c may include a first prompt feature acquisition module 1101c-1 and a second prompt feature acquisition module 1101c-2. The first prompt feature acquisition module 1101c-1 may be configured to provide the first prompt feature sequence to a text encoder in a vision-language model that supports at least vision-modal and language-modal tasks for encoding to obtain a first prompt feature. The second prompt feature acquisition module 1101c-2 may be configured to provide the second prompt feature sequence to the text encoder in the vision-language model for encoding to obtain a second prompt feature.

[0166] In some embodiments, the prompt feature sequence construction module 1101b may include a third prompt feature acquisition module 1101b-1 and a prompt feature concatenation module 1101b-2. The third prompt feature acquisition module 1101b-1 may be configured to process the image features and a third initial prompt feature based on a cross-attention mechanism to obtain a third prompt feature, where the third initial prompt feature is used to characterize the inherent attributes of the target object. The prompt feature concatenation module 1101b-2 may be configured to concatenate the third prompt feature and the first initial prompt feature into a first prompt feature sequence, and concatenate the third prompt feature and the second initial prompt feature into a second prompt feature sequence.

[0167] In some embodiments, the third prompt feature acquisition module 1101b-1 may include a cross-attention mechanism execution module 1101b-11 and a feature summation module 1101b-12. The cross-attention mechanism execution module 1101b-11 may be configured to execute the cross-attention mechanism by setting the third initial prompt feature as the query vector and setting the image features as the key vector and value vector to obtain an attention calculation feature. The feature summation module 1101b-12 may be configured to fuse the third initial prompt feature and the attention calculation feature to obtain a third prompt feature.

[0168] In some embodiments, the feature processing module 1102 may include a cross-attention processing module 1102a. The cross-attention processing module 1102a may be configured to process the first prompt feature and the image feature based on a cross-attention mechanism to obtain a first reference feature, and process the second prompt feature and the image feature based on the cross-attention mechanism to obtain a second reference feature.

[0169] In some embodiments, the feature processing module 1102a may include a first reference feature acquisition module 1102a-1 and a second reference feature acquisition module 1102a-2. The first reference feature acquisition module 1102a-1 may be configured to perform a cross-attention mechanism by setting the image feature as a query vector and setting the first prompt feature as a key vector and a value vector to obtain a first reference feature. The second reference feature acquisition module 1102a-2 may be configured to perform a cross-attention mechanism by setting the image feature as a query vector and setting the second prompt feature as a key vector and a value vector to obtain a second reference feature.

[0170] In some embodiments, the detection result determination module 1104 may include a real judgment module 1104a and an unreal judgment module 1104b. The real judgment module 1104a may be configured to determine that the image to be detected is real in response to the score of the first similarity being greater than the score of the second similarity. The unreal judgment module 1104b may be configured to determine that the image to be detected is unreal in response to the score of the second similarity being greater than the score of the first similarity.

[0171] Figure 12 The structural block diagram of an image detection training device 1200 according to an embodiment of the present disclosure is shown.

[0172] As Figure 12 shown, the device 1200 includes a data acquisition module 1201, an image detection module 1202, and a training execution module 1203. The data acquisition module 1201 is configured to acquire sample images for training, a first training prompt feature, and a second training prompt feature, where the sample images include real sample target objects or unreal sample target objects, the first training prompt feature is associated with the real attribute of the sample target object, and the second training prompt feature is associated with the unreal attribute of the sample target object. The image detection module 1202 is configured to perform image detection on the sample images to obtain sample image detection results. The training execution module 1203 may be configured to perform training based on the sample images and the sample image detection results through a preset loss function to optimize the first training prompt feature and the second training prompt feature.

[0173] In an example, the image detection module 1202 may be configured to be based on the above image detection device (such asFigure 10 the device 1000 shown or as Figure 11 shown, the device 1100) performs image detection on the sample image.

[0174] The operations of the above-mentioned feature data acquisition module 1201, image detection module 1202, and training execution module 1203 can respectively correspond to the operations of steps S801, S802, and S803 as Figure 8 shown. Therefore, details of each aspect thereof will not be elaborated here.

[0175] According to an embodiment of the present disclosure, there is also provided an electronic device, including at least one processor; and a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method described above.

[0176] According to an embodiment of the present disclosure, there is also provided a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause a computer to execute the method described above.

[0177] According to an embodiment of the present disclosure, there is also provided a computer program product, including a computer program, wherein the computer program, when executed by a processor, implements the method described above.

[0178] Referring to Figure 13 , the block diagram of the electronic device 1300 that can be used as a server or a client of the present disclosure will now be described. It is an example of a hardware device that can be applied to various aspects of the present disclosure. The electronic device is intended to represent various forms of digital electronic computer devices, such as, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, a personal digital processor, a cellular phone, a smart phone, a wearable device, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0179] As Figure 13As shown, the electronic device 1300 includes a computing unit 1301, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 1302 or a computer program loaded from a storage unit 1308 into a random access memory (RAM) 1303. In the RAM 1303, various programs and data required for the operation of the electronic device 1300 can also be stored. The computing unit 1301, the ROM 1302, and the RAM 1303 are connected to each other through a bus 1304. An input / output (I / O) interface 1305 is also connected to the bus 1304.

[0180] Multiple components in the electronic device 1300 are connected to the I / O interface 1305, including: an input unit 1306, an output unit 1307, a storage unit 1308, and a communication unit 1309. The input unit 1306 can be any type of device capable of inputting information into the electronic device 1300. The input unit 1306 can receive input digital or character information and generate key signal inputs related to user settings and / or function controls of the electronic device, and can include, but is not limited to, a mouse, a keyboard, a touch screen, a trackpad, a trackball, a joystick, a microphone, and / or a remote control. The output unit 1307 can be any type of device capable of presenting information and can include, but is not limited to, a display, a speaker, a video / audio output terminal, a vibrator, and / or a printer. The storage unit 1308 can include, but is not limited to, magnetic disks and optical discs. The communication unit 1309 allows the electronic device 1300 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks, and can include, but is not limited to, a modem, a network card, an infrared communication device, a wireless communication transceiver, and / or a chipset, such as a Bluetooth device, an 802.11 device, a WiFi device, a WiMax device, a cellular communication device, and / or the like.

[0181] The computing unit 1301 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 1301 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 1301 executes the various methods and processes described above. For example, in some embodiments, the method can be implemented as a computer software program tangibly embodied in a machine-readable medium, such as the storage unit 1308. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 1300 via the ROM 1302 and / or the communication unit 1309. When the computer program is loaded into the RAM 1303 and executed by the computing unit 1301, one or more steps of the method described above can be executed. Alternatively, in other embodiments, the computing unit 1301 can be configured to execute the method by any other suitable means (e.g., by means of firmware).

[0182] Various embodiments of the systems and techniques described above in this document can be implemented in digital electronic circuitry, integrated circuit systems, field-programmable gate arrays (FPGA), application-specific integrated circuits (ASIC), application-specific standard products (ASSP), systems-on-a-chip (SOC), complex programmable logic devices (CPLD), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a dedicated or general-purpose programmable processor that receives data and instructions from a storage system, at least one input device, and at least one output device, and transmits the data and instructions to the storage system, the at least one input device, and the at least one output device.

[0183] The program code for implementing the methods of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to the processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the program codes are executed by the processor or controller, the functions / operations specified in the flowchart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, as an independent software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0184] In the context of this disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. The machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0185] In order to provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0186] The systems and techniques described herein can be implemented in a computing system that includes backend components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes frontend components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system that includes any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), the Internet, and a blockchain network.

[0187] A computer system may include a client and a server. The client and the server are generally far from each other and usually interact via a communication network. The relationship between the client and the server is generated by computer programs running on respective computers and having a client-server relationship with each other. The server may be a cloud server, a server of a distributed system, or a server incorporating a blockchain.

[0188] It should be understood that the various forms of processes shown above can be used, steps can be reordered, added, or deleted. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and this is not limited herein.

[0189] Although the embodiments or examples of the present disclosure have been described with reference to the accompanying drawings, it should be understood that the above methods, systems, and devices are merely exemplary embodiments or examples, and the scope of the present invention is not limited by these embodiments or examples, but is only defined by the authorized claims and their equivalent scope. Various elements in the embodiments or examples can be omitted or replaced by their equivalent elements. In addition, the steps can be executed in an order different from that described in this disclosure. Further, the various elements in the embodiments or examples can be combined in various ways. Importantly, with the evolution of technology, many of the elements described herein can be replaced by equivalent elements that emerge after this disclosure.

Claims

1. An image detection method, comprising: Obtaining image features of a to-be-detected image including a target object, as well as a first hint feature and a second hint feature for detecting the to-be-detected image, where the first hint feature is associated with the true attribute of the target object, and the second hint feature is associated with the non-true attribute of the target object; Fusing the first hint feature and the image features to obtain a first reference feature, and fusing the second hint feature and the image features to obtain a second reference feature; Determining a first similarity between the first reference feature and the image features, and a second similarity between the second reference feature and the image features; And Based on the first similarity and the second similarity, determining an image detection result, where the image detection result indicates whether the to-be-detected image is true.

2. The method according to claim 1, wherein The obtaining image features of the to-be-detected image including the target object includes: Providing the to-be-detected image to an image encoder in a vision-language model that supports at least vision-modal and language-modal tasks for encoding to obtain the image features.

3. The method according to claim 1 or 2, wherein The obtaining the first hint feature and the second hint feature for detecting the to-be-detected image includes: Constructing a first hint feature sequence and a second hint feature sequence, where the first hint feature sequence includes a first initial hint feature for characterizing the true attribute of the target object, and a third hint feature associated with the inherent attribute of the target object, and the second hint feature sequence includes a second initial hint feature for characterizing the non-true attribute of the target object, and the third hint feature; and Fusing the features in the first hint feature sequence to obtain the first hint feature, and fusing the features in the second hint feature sequence to obtain the second hint feature.

4. The method according to claim 3, wherein The fusing the features in the first hint feature sequence to obtain the first hint feature, and fusing the features in the second hint feature sequence to obtain the second hint feature includes: Providing the first hint feature sequence to a text encoder in a vision-language model that supports at least vision-modal and language-modal tasks for encoding to obtain the first hint feature; and Providing the second hint feature sequence to the text encoder in the vision-language model for encoding to obtain the second hint feature.

5. The method according to claim 3 or 4, wherein The constructing the first hint feature sequence and the second hint feature sequence includes: Processing the image features and a third initial hint feature based on a cross-attention mechanism to obtain the third hint feature, where the third initial hint feature is used to characterize the inherent attribute of the target object; and Concatenating the third hint feature and the first initial hint feature into the first hint feature sequence, and concatenating the third hint feature and the second initial hint feature into the second hint feature sequence.

6. The method according to claim 5, wherein, The processing the image features and the third initial hint feature based on the cross-attention mechanism to obtain the third hint feature includes: The cross-attention mechanism is performed by setting the third initial prompt feature as the query vector and setting the image feature as the key vector and value vector to obtain an attention calculation feature; and The third initial prompt feature and the attention calculation feature are fused to obtain the third prompt feature.

7. The method according to claim 6, wherein The cross-attention mechanism includes a multi-head cross-attention mechanism such that the attention calculation feature is obtained based on multiple attention heads.

8. The method according to any one of claims 3 to 7, wherein, The first prompt feature sequence and the second prompt feature sequence further include a first additional feature for indicating the start of the sequence, a second additional feature for indicating the end of the sequence, and a third additional feature for indicating the target object.

9. The method according to claim 8, wherein, In the first prompt feature sequence, the first additional feature, the third prompt feature, the first initial prompt feature, the third additional feature, and the second additional feature are concatenated; In the second prompt feature sequence, the first additional feature, the third prompt feature, the second initial prompt feature, the third additional feature, and the second additional feature are concatenated.

10. The method according to any one of claims 1 to 9, wherein, The fusing the first prompt feature and the image feature to obtain a first reference feature, and fusing the second prompt feature and the image feature to obtain a second reference feature includes: Processing the first prompt feature and the image feature based on the cross-attention mechanism to obtain the first reference feature, and processing the second prompt feature and the image feature based on the cross-attention mechanism to obtain the second reference feature.

11. The method according to claim 10, wherein, The processing the first prompt feature and the image feature based on the cross-attention mechanism to obtain the first reference feature, and processing the second prompt feature and the image feature based on the cross-attention mechanism to obtain the second reference feature includes: Performing the cross-attention mechanism by setting the image feature as the query vector and setting the first prompt feature as the key vector and value vector to obtain the first reference feature; and Performing the cross-attention mechanism by setting the image feature as the query vector and setting the second prompt feature as the key vector and value vector to obtain the second reference feature.

12. The method according to claim 11, wherein The cross-attention mechanism includes a multi-head cross-attention mechanism such that both the first reference feature and the second reference feature are obtained based on multiple attention heads.

13. The method according to any one of claims 1 to 12, wherein, Determining an image detection result based on the first similarity and the second similarity, where the image detection result indicates whether the image to be detected is real, includes: In response to the score of the first similarity being greater than the score of the second similarity, determining that the image to be detected is real; and In response to the score of the second similarity being greater than the score of the first similarity, determining that the image to be detected is not real.

14. The method according to any one of claims 1 to 13, wherein The target object includes a person's face, and the non-real attribute of the target object includes forgery marks on the face.

15. An image detection training method, including: Obtain sample images for training, a first prompt feature to be trained, and a second prompt feature to be trained, where the sample images include real sample target objects or non-real sample target objects, the first prompt feature to be trained is associated with the real attributes of the sample target objects, and the second prompt feature to be trained is associated with the non-real attributes of the sample target objects; Perform image detection on the sample images to obtain sample image detection results; and Based on the sample images and the sample image detection results, perform training through a preset loss function to optimize the first prompt feature to be trained and the second prompt feature to be trained.

16. The method according to claim 15, wherein, The preset loss function includes an image-prompt alignment loss, and the image-prompt alignment loss is used to represent whether the first prompt feature to be trained and the second prompt feature to be trained are aligned with the authenticity of the sample images.

17. The method according to claim 16, wherein The preset loss function further includes a prompt diversity loss, and the prompt diversity loss includes a cross-class prompt diversity loss and an intra-class prompt diversity loss. Among them, the cross-class prompt diversity loss is used to represent whether there is a difference between the first prompt feature to be trained and the second prompt feature to be trained, and the intra-class prompt diversity loss is used to represent whether there is redundancy among the first prompt features to be trained of the same class and whether there is redundancy among the second prompt features to be trained of the same class.

18. The method according to claim 17, wherein, The preset loss function is expressed as a weighted sum of the image-prompt alignment loss and the prompt diversity loss.

19. An image detection device, comprising: A feature acquisition module, configured to acquire image features of a to-be-detected image including a target object, and a first prompt feature and a second prompt feature for detecting the to-be-detected image, where the first prompt feature is associated with the real attributes of the target object, and the second prompt feature is associated with the non-real attributes of the target object; A feature processing module, configured to fuse the first prompt feature and the image features to obtain a first reference feature, and fuse the second prompt feature and the image features to obtain a second reference feature; A similarity determination module, configured to determine a first similarity between the first reference feature and the image features, and a second similarity between the second reference feature and the image features; And A detection result determination module, configured to determine an image detection result based on the first similarity and the second similarity, where the image detection result indicates whether the to-be-detected image is real.

20. An image detection training device, comprising: A data acquisition module, configured to acquire sample images for training, a first prompt feature to be trained, and a second prompt feature to be trained, where the sample images include real sample target objects or non-real sample target objects, the first prompt feature to be trained is associated with the real attributes of the sample target objects, and the second prompt feature to be trained is associated with the non-real attributes of the sample target objects; An image detection module, configured to perform image detection on the sample image to obtain a sample image detection result; and A training execution module, configured to optimize the first to-be-trained prompt feature and the second to-be-trained prompt feature through a preset loss function based on the sample image and the sample image detection result.

21. An electronic device, comprising: At least one processor; And A memory communicatively connected to the at least one processor; Wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method according to any one of claims 1-18.

22. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause a computer to execute the method according to any one of claims 1-18.

23. A computer program product, comprising a computer program, wherein, The computer program, when executed by a processor, implements the method according to any one of claims 1-18.