Forgery detection method and device, storage medium and electronic equipment

By extracting feature maps of multiple object classes through a forgery detection model and performing feature fusion with texture enhancement and attention feature enhancement, the accuracy and adaptability issues of existing models under high-quality forgery data are solved, achieving more accurate and robust forgery detection.

CN115830487BActive Publication Date: 2026-02-17ALIPAY (HANGZHOU) INFORMATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211226699.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-10-09
Publication Date
2026-02-17
Estimated Expiration
2042-10-09

AI Technical Summary

Technical Problem

Existing forgery detection models are inaccurate and lack adaptability when faced with high-quality forged data, making it difficult to effectively detect forged data in complex application scenarios.

Method used

The forgery detection model extracts feature maps of multiple object types from the target object data, performs feature texture enhancement and attention feature enhancement, and finally performs feature attention fusion to output the forgery detection result.

Benefits of technology

It improves the accuracy and robustness of forgery detection, can adaptively focus on regions with differences in texture information, resists forgery detection reverse resolution mechanisms, adapts to complex application scenarios, and improves the overall detection effect.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115830487B_ABST
    Figure CN115830487B_ABST
Patent Text Reader

Abstract

The specification discloses a forgery detection method and device, a storage medium and an electronic device, wherein the method comprises: extracting a multi-class object feature map of target object data through the forgery detection model, and performing feature texture enhancement based on the multi-class object feature map to obtain a texture-enhanced feature map; performing attention feature enhancement on the object feature map through the forgery detection model to obtain an attention feature map; then performing feature attention fusion based on the object feature map, the texture-enhanced feature map and the attention feature map to obtain object fusion features; and performing forgery detection based on the object fusion features to output an object forgery detection result.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present specification relates to the technical field of computer technology, and particularly relates to a forgery detection method and device, a storage medium and an electronic device. BACKGROUND

[0002] With the rapid development of computer technology, the cost of data forgery such as image, video and the like is lower and lower, the quality of generated forgery data is higher and higher, and the forgery data is randomly visible; in daily applications, it is necessary to accurately determine whether object data such as face image, face video, user image and the like is forged, and data forgery detection has become an important part in daily application scenarios such as security privacy, user verification and identity authentication. SUMMARY

[0003] The present specification provides a forgery detection method, device, storage medium and electronic device, and the technical solution is as follows:

[0004] In a first aspect, the present specification provides a forgery detection method, and the method comprises:

[0005] inputting target object data into a forgery detection model, extracting a multi-class object feature map of the target object data through the forgery detection model, and performing feature texture enhancement based on the multi-class object feature map to obtain a texture enhanced feature map;

[0006] performing attention feature enhancement on the object feature map through the forgery detection model to obtain an attention feature map;

[0007] performing feature attention fusion on the object feature map, the texture enhanced feature map and the attention feature map through the forgery detection model to obtain object fusion features, and outputting object forgery detection results based on the object fusion features.

[0008] In a second aspect, the present specification provides a forgery detection device, and the device comprises:

[0009] a texture processing module configured to input target object data into a forgery detection model, extract a multi-class object feature map of the target object data through the forgery detection model, and perform feature texture enhancement based on the multi-class object feature map to obtain a texture enhanced feature map;

[0010] an attention enhancement module configured to perform attention feature enhancement on the object feature map through the forgery detection model to obtain an attention feature map;

[0011] The fusion processing module is configured to perform feature attention fusion on the object feature map, the texture enhanced feature map, and the attention feature map to obtain an object fusion feature, and output an object forgery detection result based on the object fusion feature.

[0012] In a third aspect, the present specification provides a computer storage medium, which stores a plurality of instructions, and the instructions are suitable for being loaded and executed by a processor to perform the method steps described above.

[0013] In a fourth aspect, the present specification provides an electronic device, which can include a processor and a memory, wherein the memory stores a computer program, and the computer program is suitable for being loaded and executed by the processor to perform the method steps described above.

[0014] The technical solutions provided by some embodiments of the present specification have at least the following beneficial effects:

[0015] In one or more embodiments of the present specification, the electronic device inputs target object data into a forgery detection model, extracts a plurality of object feature maps of the target object data through the forgery detection model, and performs feature texture enhancement based on the plurality of object feature maps to obtain a texture enhanced feature map. Then, the electronic device performs attention feature enhancement on the object feature map through the forgery detection model to obtain an attention feature map, and performs feature attention fusion on the object feature map, the texture enhanced feature map, and the attention feature map to obtain an object fusion feature. Based on the object fusion feature, accurate forgery detection can be performed, thereby avoiding the phenomenon that the forgery detection result is wrong. The texture information difference area can be adaptively focused, thereby resisting the forgery detection anti-decoding mechanism, better adapting to complex application scenarios to achieve effective detection. In addition, the global detection effect of the model in complex scenarios can be improved, and the robustness and universality of the forgery detection are ensured. BRIEF DESCRIPTION OF DRAWINGS

[0016] In order to more clearly illustrate the technical solutions in the present specification or the prior art, the drawings needed in the embodiments or prior art description will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present specification, and other drawings can be obtained by those skilled in the art without creative labor.

[0017] Figure 1 is a scene schematic diagram of a forgery detection system provided by the present specification;

[0018] Figure 2 is a flowchart of a forgery detection method provided by the present specification;

[0019] Figure 3 is a scene schematic diagram of feature processing provided by the present specification;

[0020] Figure 4 is a flowchart of a forgery detection method provided by the present specification;

[0021] Figure 5 is a scene diagram of feature map extraction involved in the present specification;

[0022] Figure 6 is a scene diagram of feature map determination involved in the present specification;

[0023] Figure 7 is a scene diagram of a texture enhancement module involved in the present specification;

[0024] Figure 8 is a flowchart of a forgery detection method provided by the present specification;

[0025] Figure 9 is a scene diagram of model training of a forgery detection model provided by the present specification;

[0026] Figure 10 is a flowchart of a forgery detection device provided by the present specification;

[0027] Figure 11 is a module diagram of a texture processing module provided by the present specification;

[0028] Figure 12 is a structural diagram of an electronic device provided by the present specification;

[0029] Figure 13 is a structural diagram of an operating system and user space provided by the present specification;

[0030] Figure 14 is Figure 13 an architecture diagram of an Android operating system;

[0031] Figure 15 is Figure 13 an architecture diagram of an IOS operating system. DETAILED DESCRIPTION

[0032] The technical solutions in the present specification will be described clearly and completely below in conjunction with the drawings in the present specification. Obviously, the described embodiments are only some of the embodiments of the present specification, rather than all the embodiments. Based on the embodiments in the present specification, all other embodiments obtained by those of ordinary skill in the art without creative labor fall within the scope of protection of the present specification.

[0033] In the description of the specification, it should be understood that the terms "first", "second" and the like are used only for descriptive purposes and are not intended to indicate or imply relative importance. In the description of the specification, it should be noted that, unless otherwise expressly specified and limited, "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units is not limited to the listed steps or units, but can optionally include steps or units not listed, or can optionally include other steps or units inherent to the process, method, product or device. The specific meaning of the above terms in the specification can be understood by the person skilled in the art. In addition, in the description of the specification, "a plurality of" means two or more, unless otherwise specified. The association relationship between the associated objects is described as "and / or", which means that there can be three relationships, for example, A and / or B can represent the following three cases: A exists alone, A and B exist together, and B exists alone. The character " / " generally represents an "or" relationship between the associated objects before and after it.

[0034] In the related art, a machine learning model for forgery detection is usually trained to perform forgery detection on corresponding target data, and most of the target data with forgery phenomenon is generated based on a deep convolutional neural network model, the similarity of the model basic network structure is high, and there is a great probability of resistance mechanism of forgery detection, that is, forgery detection anti-decoding mechanism, so that the detection result of the machine learning model for forgery detection after actual online is inaccurate, and the adaptation ability in the actual application scene is not strong.

[0035] The specification will be described in detail below in conjunction with specific embodiments.

[0036] Please refer to Figure 1 , a scene schematic diagram of a forgery detection system provided in the specification. As Figure 1 indicated, the forgery detection system can at least include a client cluster and a service platform 100.

[0037] The client cluster can include at least one client, such as Figure 1 indicated, specifically including a client 1 corresponding to a user 1, a client 2 corresponding to a user 2,..., and a client n corresponding to a user n, n is an integer greater than 0.

[0038] The clients in the client cluster can be electronic devices with communication functions, including but not limited to wearable devices, handheld devices, personal computers, tablet computers, vehicle-mounted devices, smart phones, computing devices, or other processing devices connected to wireless modems, etc. Electronic devices can be called by different names in different networks, such as user equipment, access terminal, user unit, user station, mobile station, mobile station, remote station, remote terminal, mobile device, user terminal, terminal, wireless communication device, user agent or user device, cellular phone, cordless phone, personal digital assistant (PDA), electronic device in 5G network or future evolution network, etc.

[0039] The service platform 100 can be a separate server device, such as a rack-mounted, blade, tower, or cabinet server device, or a hardware device with strong computing power such as a workstation, a mainframe computer, etc. It can also be a server cluster composed of multiple servers. The servers in the service cluster can be composed in a symmetrical manner, where each server is functionally and positionally equivalent in the transaction link, and each server can independently provide services. The independent service can be understood as not requiring the assistance of another server.

[0040] In one or more embodiments of the present specification, the service platform 100 can establish a communication connection with at least one client in the client cluster, and complete the interaction of data in the forgery detection process based on the communication connection, such as online transaction data interaction. For example, the service platform 100 can deploy the forgery detection model obtained based on the forgery detection method of the present specification to several clients, so as to assist the clients in collecting target object data and performing forgery detection on the target object based on the forgery detection model in an actual forgery detection scenario. For another example, the service platform 100 can obtain the target object data collected by the client from the client, and perform forgery detection on the target object by using the forgery detection model.

[0041] It should be noted that the service platform 100 and at least one client in the client cluster establish a communication connection through a network for interactive communication, where the network can be a wireless network or a wired network. The wireless network includes but is not limited to a cellular network, a wireless local area network, an infrared network or a Bluetooth network. The wired network includes but is not limited to an Ethernet, a universal serial bus (USB) or a controller area network. In one or more embodiments of the specification, technologies and / or formats including Hyper Text Mark-up Language (HTML), Extensible Markup Language (XML) and the like are used to represent data (such as the target compressed package) exchanged through the network. In addition, all or some links can be encrypted using conventional encryption technologies such as Secure Socket Layer (SSL), Transport Layer Security (TLS), Virtual Private Network (VPN), Internet Protocol Security (IPsec) and the like. In other embodiments, custom and / or dedicated data communication technologies can be used instead of or in addition to the above data communication technologies.

[0042] The forgery detection system embodiments provided by the specification belong to the same concept as the forgery detection method in one or more embodiments. The execution subject of the forgery detection method involved in one or more embodiments of the specification can be an electronic device corresponding to the service platform 100 described above. The execution subject of the forgery detection method involved in one or more embodiments of the specification can also be an electronic device corresponding to the client, which is determined based on the actual application environment. The implementation process of the forgery detection system embodiment can be seen from the method embodiment described below, which will not be described here.

[0043] Based on Figure 1 The scene diagram shown below provides a detailed description of the forgery detection method provided by one or more embodiments of the specification.

[0044] Please refer to Figure 2 A flowchart of a forgery detection method is provided for one or more embodiments of the specification. The method can be implemented by relying on a computer program and can be run on a forgery detection device based on the von Neumann architecture. The computer program can be integrated in an application or run as an independent tool application. The forgery detection device can be an electronic device.

[0045] Specifically, the forgery detection method includes:

[0046] S102: input the target object data into the forgery detection model, extract a multi-class object feature map of the target object data through the forgery detection model, and perform feature texture enhancement based on the multi-class object feature map to obtain a texture-enhanced feature map;

[0047] It can be understood that with the development of artificial intelligence technology, deep learning and other artificial intelligence technologies can be used to forge corresponding object data, such as replacing the face object region in certain media data (image or video) or tampering with the leg object region in certain media data (image or video). After forgery, the user is difficult to directly distinguish whether the object data is image or video.

[0048] In actual transaction scenarios, it is necessary to perform forgery detection on target object data in corresponding forgery detection scenarios, such as performing forgery detection processing on target object data such as user object images and user object videos for corresponding object parts (such as the face and limb parts of the object).

[0049] Further, the target object data can be detected for forgery based on a pre-trained forgery detection model to verify whether the input target object data has been forged to obtain an object forgery detection result for the target object data.

[0050] The forgery detection model can be obtained after model training of an initial forgery detection model based on a machine learning model. The machine learning model can be one or more of a convolutional neural network (CNN) model, a deep neural network (DNN) model, a recurrent neural network (RNN) model, an embedding model, a gradient boosting decision tree (GBDT) model, a logistic regression (LR) model, and the like. The initial forgery detection model is trained offline using sample object data, and when the model end condition is met, a trained forgery detection model is obtained. The forgery detection model can then be deployed online to actual forgery detection scenarios to perform forgery detection on target object data.

[0051] Further, the forgery detection scene can be a scene in actual application, such as finance, insurance, security, and the like, which needs to detect forgery of an object. For example, when a user performs object verification, the user generally needs to submit object data such as an ID card, a student card, and current face object data on an application program or a website. The target object data can be in the form of visual data such as a video, an image, a short video, and an animation carrying object information.

[0052] In one or more embodiments of the present specification, by inputting the target object data into the forgery detection model, the forgery detection model is used to extract multi-class object feature maps of the target object data layer by layer.

[0053] The multi-class object feature map can be understood as a feature map of the target object extracted at different feature extraction stages of the same target object data. For the feature extraction part of the forgery detection model, the feature extraction part is usually a multi-layer structure, and different levels of the feature extraction part correspond to different classes of object feature maps.

[0054] Illustratively, the feature extraction part can be composed of multiple backbone networks for feature extraction (such as multiple backbone networks stacked together), and different backbone networks can extract different classes of object feature maps. The backbone network can be fitted based on one or more of the above machine learning models. Further, different classes of object feature maps have different feature resolution and semantic expression.

[0055] In one or more embodiments of the present specification, the multi-class object feature map of the target object data is extracted by the forgery detection model, and a texture enhanced feature map is obtained based on all or part of the multi-class object feature map.

[0056] For example, during the processing of the forgery detection model, the feature texture of the object feature map is enhanced after each class of object feature map is extracted by the backbone network of different levels, and a texture enhanced map corresponding to each class of object feature map is obtained.

[0057] For another example, during the processing of the forgery detection model, the feature texture of part of the target object feature map is enhanced after each class of object feature map is extracted by the backbone network of different levels, and a texture enhanced map corresponding to each class of object feature map is obtained. The target object feature map can be pre-set, for example, the first class of object feature map corresponding to the first backbone network and the second class of object feature map corresponding to the second backbone network can be set as the target object feature map.

[0058] Illustratively, the object feature map is enhanced in feature texture, and the forgery detection model mainly focuses on the texture feature information of the object feature map, and enhances the texture feature information to magnify the forgery details corresponding to the object forgery (such as face forgery and body forgery).

[0059] In a feasible implementation, the forgery detection model can at least include a plurality of backbone networks and a texture enhancement module. The multi-class object feature map of the target object data is extracted by the forgery detection model, and the feature texture of the multi-class object feature map is enhanced to obtain a texture-enhanced feature map, which can be:

[0060] The multi-class object feature map of the target object data is extracted by each backbone network, at least one class of target object feature map is determined from the multi-class object feature map, each class of target object feature map is input into the texture enhancement module, and the texture information is enhanced from each class of target object feature map by the texture enhancement module to output the texture-enhanced feature map corresponding to each class of target object feature map.

[0061] Illustratively, as shown in Figure 3 , Figure 3 is a scene diagram of a feature processing involved in the present specification, in Figure 3 , part of the model structure and data flow state of the forgery detection model are shown, Figure 3 The forgery detection model shown can include a plurality of backbone networks, such as backbone network 1, backbone network 2,..., and backbone network n (n is a positive integer). The number of backbone networks can be set based on the actual application scenario. The plurality of backbone networks can be in the form of network stacking structure. Different backbone networks can extract different classes of object feature maps from the target object data, such as object feature Figure 1 extracted by backbone network 1, object feature Figure 2 extracted by backbone network 2, and object feature map n extracted by backbone network n. Further, different classes of object feature maps have different feature resolutions and semantic expression degrees. Generally, the feature resolutions of object feature Figure 1 , object feature Figure 2 ... object feature map n decrease in turn, and generally, the semantic expression degrees of object feature Figure 1 , object feature Figure 2 ... object feature map n increase in turn.

[0062] Further, for the object feature map corresponding to each backbone network, all or part of the target object feature maps in a plurality of object feature maps can be selected and input into the texture enhancement module for texture enhancement to obtain the texture-enhanced feature map corresponding to the target object feature map.

[0063] In one or more embodiments of the present specification, the object feature map with generally high feature resolution retains more texture information of the forgery trace, forgery artifact, etc. In some embodiments, such object feature map is often referred to as a shallow object feature map, for example, the aforementioned object feature Figure 1 Generally, the present specification is not limited to only texture enhancement of the shallow object feature map, but also combines actual conditions to perform texture enhancement on other object feature maps to excavate potential texture information of the forgery trace, forgery artifact, etc. hidden in depth.

[0064] In a consistent and feasible implementation, the texture enhancement module can use convolution operation, pooling operation, etc. to process the input object feature map to enhance the texture information and obtain a texture-enhanced feature map.

[0065] S104: performing attention feature enhancement on the object feature map by using the forgery detection model to obtain an attention feature map.

[0066] In one or more embodiments of the present specification, the attention mechanism part of the forgery detection model can extract features from at least one of the object feature maps to obtain an attention feature map.

[0067] Illustratively, the attention mechanism part can be at least one layer of attention enhancement module Attention module. Based on the attention feature enhancement of the object feature map by using the at least one layer of Attention module, the attention feature map can be obtained.

[0068] Illustratively, the attention feature enhancement on the object feature map can be visual attention enhancement processing and feature extraction on the visual region from the visual dimension, which can be understood as the object easy-to-forgery region from the visual dimension, and the feature extraction of the visual attention layer is enhanced to obtain the attention feature map. The attention feature enhancement on the object feature map can also be semantic attention enhancement processing and feature extraction on the image / video semantic region from the semantic dimension to obtain the attention feature map, and the attention enhancement thereby assists focusing on the forgery difference sensitive region.

[0069] S106: performing feature attention fusion on the object feature map, the texture-enhanced feature map, and the attention feature map by using the forgery detection model to obtain an object fusion feature, and outputting an object forgery detection result based on the object fusion feature.

[0070] In one or more embodiments of the present specification, the object feature map, the texture enhanced feature map, and the attention feature map can be fused by a feature attention fusion model, and the object feature map, the texture enhanced feature map, and the attention feature map can be subjected to an attention pooling operation to perform multi-layer attention fusion to achieve adaptive focusing of texture sensitive regions at the feature level, and finally the object fusion feature after a series of enhancements is subjected to forgery detection classification, and an object forgery detection result is obtained.

[0071] Illustratively, considering the differences between different regional ranges, if global average fusion is used for the object feature map, the texture enhanced feature map, and the attention feature map, the fused object fusion feature will be affected by the attention map intensity, which is not conducive to the purpose of paying attention to texture information. Therefore, an attention pooling operation is used for optimization.

[0072] Illustratively, the object forgery detection result is usually an object forgery result or an object authenticity result for target object data.

[0073] In one or more embodiments of the present specification, an electronic device inputs target object data into a forgery detection model, extracts a multi-class object feature map of the target object data based on the multi-class object feature map, and performs feature texture enhancement to obtain a texture enhanced feature map, then performs attention feature enhancement on the object feature map to obtain an attention feature map, and then performs feature attention fusion on the object feature map, the texture enhanced feature map, and the attention feature map to obtain an object fusion feature. Based on the object fusion feature, accurate forgery detection can be performed to avoid errors in the forgery detection result, adaptively focus on texture information difference regions, resist forgery detection reverse mechanism, better adapt to complex application scenarios for effective detection, and improve the global detection effect of the model in complex scenarios to ensure the robustness and universality of the forgery detection.

[0074] Please refer to Figure 4 , Figure 4 is a flowchart of another embodiment of a forgery detection method according to one or more embodiments of the present specification. Specifically:

[0075] S202: input target object data into a forgery detection model, and extract a multi-class object feature map of the target object data by using each backbone network of the forgery detection model;

[0076] In one or more embodiments of the present specification, the several backbone networks of the forgery detection model can be a backbone network structure composed of multiple layers of Backbone Layers.

[0077] In some embodiments, the object data can often be of a video type, such as a segment of facial video capturing a user's face, and the network structure in the forgery detection model is usually a three-dimensional network structure, and the backbone network can often be referred to as 3D BackboneLayers. When the 3D backbone network extracts features from the target object data, it usually also focuses on the feature information of the target object data in the time dimension. The object feature map corresponding to the 3D backbone network can be in the form of "C*H*W*T", where C is the channel parameter, H represents the length or height parameter, W represents the width parameter, and T represents the time parameter. It can be understood that the forgery detection model trained based on video type object data can be applicable to image type object data, i.e., image type object data can be subjected to forgery detection.

[0078] In a feasible implementation, taking three backbone networks as an example, the plurality of backbone networks can include a first backbone network, a second backbone network, and a third backbone network. The first backbone network can be 3D BackboneLayers1, the second backbone network can be 3D BackboneLayers2, and the third backbone network can be 3D BackboneLayers3. The plurality of backbone networks can be in the form of a stack.

[0079] Further, the plurality of types of object feature maps of the target object data extracted by the respective backbone networks can be as shown in Figure 5 Figure 5 is a scene diagram of feature map extraction involved in the present specification;

[0080] The first type of object feature map of the target object data is extracted by the first backbone network 3D BackboneLayers1, the second type of object feature map of the target object data is extracted by the second backbone network 3D BackboneLayers2, and the third type of object feature map of the target object data is extracted by the third backbone network 3D BackboneLayers3.

[0081] Among them, the feature resolution of the second type of object feature map is smaller than that of the first type of object feature map and larger than that of the third type of object feature map, and the semantic expression degree of the second type of object feature map is greater than that of the first type of object feature map and smaller than that of the third type of object feature map.

[0082] ​Illustratively, the features corresponding to the first type of object feature map are generally closer to the input data, containing more pixel point information, such as color, texture, edge, and other feature information, that is, the feature resolution of the first type of object feature map is generally higher than that of other object feature maps, but since it has passed through fewer backbone networks, its semantic expression degree is smaller than that of other object feature maps. Illustratively, the third type of object feature map has passed through multiple backbone networks, and the semantic expression degree of the third type of object feature map is generally higher than that of other object feature maps.

[0083] Illustratively, the backbone network can be one or more fitted in the machine learning model in one or more embodiments of the present specification. The backbone network can be a 3D backbone network, and the object feature map extracted by the 3D backbone network can be in the form of “C*H*W*T”, for example. The 3D backbone network can be compatible with image type and video type object data for feature map extraction.

[0084] S204: determining at least one target object feature map from the multiple types of object feature maps;

[0085] It can be understood that each backbone network will correspond to an extracted object feature map;

[0086] Illustratively, as shown in Figure 6 Figure 6 is a schematic diagram of a feature map determination scenario according to the present specification, in which Figure 6 , multiple types of object feature maps of the target object data are extracted by each of the backbone networks, for example, object feature Figure 1 , object feature Figure 2 ,... object feature map n, and then at least one target object feature map is determined from the multiple types of object feature maps, as shown in Figure 6 , the determined target object feature map is object feature Figure 1 , object feature Figure 2 , object feature Figure 1 , object feature Figure 2 is input into a texture enhancement module, and each type of target object feature map corresponding texture enhanced feature map is output;

[0087] ​In one or more embodiments of the present specification, considering that the ability of data forgery such as images, videos and the like using artificial intelligence technologies such as deep learning is constantly improving as technology develops, the forged object data can resist a large number of machine learning models in related technologies, and the dimensions of forgery details, global receptive fields and the like are getting higher and higher. Based on this, the forgery detection method involved in the present specification considers the foregoing multi-dimensional factors to select target object feature maps from multiple types of object feature maps for texture enhancement, rather than focusing only on object feature maps of shallow dimensions with high feature resolution (such as the first type of object feature map), so as to mine texture information in object feature maps with different feature resolutions and different semantic expressions from multiple types of object feature maps and further enhance them.

[0088] In a feasible implementation, a default object feature map can be set based on several backbone networks of the forgery detection model, and in the actual model application stage, at least one type of target object feature map can be determined from multiple types of object feature maps;

[0089] For example, assuming that the multiple types of object feature maps can be the first type of object feature map, the second type of object feature map and the third type of object feature map, the default target object feature map can be the first type of object feature map + the second type of object feature map, or the first type of object feature map + the third type of object feature map, or the first type of object feature map + the second type of object feature map + the third type of object feature map, or the first type of object feature map.

[0090] Illustratively, the first type of object feature map and the second type of object feature map are determined from the multiple types of object feature maps, the first type of object feature map and the second type of object feature map are respectively input into the texture enhancement module, and a first texture enhanced feature map corresponding to the first type of object feature map and a second texture enhanced feature map corresponding to the second type of object feature map are output; or,

[0091] Illustratively, the first type of object feature map and the third type of object feature map are determined from the multiple types of object feature maps, the first type of object feature map and the third type of object feature map are respectively input into the texture enhancement module, and a first texture enhanced feature map corresponding to the first type of object feature map and a third texture enhanced feature map corresponding to the third type of object feature map are output;

[0092] Illustratively, the first type of object feature map, the second type of object feature map and the third type of object feature map are determined from the multiple types of object feature maps, the first type of object feature map, the second type of object feature map and the third type of object feature map are respectively input into the texture enhancement module, and a first texture enhanced feature map corresponding to the first type of object feature map, a second texture enhanced feature map corresponding to the second type of object feature map and a third texture enhanced feature map corresponding to the third type of object feature map are output; or,

[0093] Optionally, the first-class object feature map is determined from the plurality of object feature maps, and the first-class object feature map is input into the texture enhancement module to output a first texture enhanced feature map corresponding to the first-class object feature map.

[0094] Optionally, a default target object feature map can be set from all object feature maps based on the transaction scenario type to which the forgery detection model is applied.

[0095] In a feasible implementation, the forgery detection model can be controlled to perform texture enhanced map prediction processing on the plurality of object feature maps to obtain at least one target object feature map.

[0096] The texture enhanced map prediction processing can be understood as predicting a target object feature map that needs to be texture enhanced from the plurality of object feature maps.

[0097] Optionally, the texture enhanced map prediction processing can be forgery feature identification on the object feature map, that is, identifying forgery features from object feature maps that have not been texture enhanced. It can be understood that at this time, the forgery feature identification is a preliminary suspected forgery feature scanning from the object feature map. The suspected forgery features may not be forgeries. After forgery feature identification, identification result data can be obtained, which includes but is not limited to suspected forgery feature quantity, suspected forgery feature distribution, suspected forgery pixel area, and other types of data. When the identification result data indicates that the suspected forgery factors account for a large proportion, a small number of object feature maps can be selected for texture enhancement. It can be understood that a large proportion of suspected forgery factors indicates that the object feature map with high feature resolution has more forgery features. Texture enhancement can be performed only on the object feature map with high feature resolution, which is sufficient for subsequent forgery classification judgment. When the identification result data indicates that the suspected forgery factors account for a small proportion or do not exist, a large number of object feature maps can be selected for texture enhancement. It can be understood that a small proportion of suspected forgery factors or their absence indicates that the object feature map with high feature resolution has no obvious or no forgery features. At this time, there may be a high level of forgery resistance mechanism to resist forgery detection. In order to improve the effect of forgery detection, a large number of types of object feature maps need to be further texture enhanced to achieve accurate judgment based on multiple texture enhanced feature maps. On the other hand, the number of feature maps that need to be texture enhanced at each model processing stage can be intelligently optimized to optimize model processing efficiency and save model computing resources.

[0098] Further, the residual prediction probability can be evaluated based on the identification result data, and data quantization can be realized based on the residual prediction probability. It can be understood that the more types of data such as suspected forgery feature quantity, suspected forgery feature distribution, and suspected forgery pixel area in the identification result data, the larger the value of the residual prediction probability, and vice versa.

[0099] Illustratively, a probability mapping relationship between the data quantity of the types of data such as the number of suspected counterfeit features, the distribution of suspected counterfeit features, and the area of suspected counterfeit pixels in the identification result data and the residual prediction probability can be established, and based on the probability mapping relationship, the residual prediction probability can be quickly obtained in combination with the current data quantity of each type.

[0100] Optionally, an object feature map with high feature resolution can be used for texture enhanced map prediction processing to obtain a residual prediction probability. For example, the object feature map with high feature resolution is usually the object feature map corresponding to the first or the first preset number (such as 2, 3, etc.) of backbone networks. It can be understood that the object feature maps obtained by the remaining backbone networks are more difficult to find suspected counterfeit features as the feature resolution is lower than that of the first or the first preset number of backbone networks.

[0101] Optionally, a texture prediction mapping relationship between a plurality of reference prediction probability ranges and reference object feature maps can be established, and different reference prediction probability ranges correspond to different numbers and types of reference object feature maps. Based on this, in the actual model processing stage, after obtaining the residual prediction probability of a certain object feature map, the target range in the plurality of reference prediction probability ranges into which the residual prediction probability falls is determined, and then a plurality of target object feature maps indicated by the target range in the texture prediction mapping relationship are obtained.

[0102] In a feasible implementation, the texture enhanced map prediction processing on the plurality of object feature maps to obtain at least one type of target object feature map can be: performing artifact residual detection on the object feature map, obtaining identification result data such as the number of suspected counterfeit features, the distribution of suspected counterfeit features, and the area of suspected counterfeit pixels through the artifact residual detection, obtaining the residual prediction probability corresponding to the object feature map based on the identification result data, and then determining at least one type of target object feature map from the plurality of object feature maps based on the residual prediction probability.

[0103] Optionally, at least one type of target object feature map usually includes the first object feature map, such as the first type of object feature map described above.

[0104] Optionally, in the counterfeit detection model, a map prediction module can be configured before the texture enhancement module, and the map prediction module is used for texture enhanced map prediction processing on the plurality of object feature maps to obtain at least one type of target object feature map. That is, the input of the map prediction module is the object feature map, and the output of the map prediction module is at least one type of target object feature map, so as to input each type of target object feature map to the texture enhancement module.

[0105] S206: input each type of target object feature map into the texture enhancement module, and output a texture-enhanced feature map corresponding to each type of target object feature map respectively;

[0106] Illustratively, the texture enhancement module can include a pooling layer and a dense convolution layer based on a residual structure, as shown in Figure 7 Figure 7 For a scene diagram of a texture enhancement module involved in the present specification, the input of the texture enhancement module is an (target) object feature map, and the output is a texture-enhanced feature map corresponding to each (target) object feature map. The pooling layer can be Figure 7 Figure 7 As shown in the 3D Denseblock, it can be understood that the dense convolution layer is a 3D dense convolution layer based on a residual structure, which can be used for texture enhancement of related object feature maps in object data of video type and image type. When the 3D Denseblock performs texture enhancement on the target object feature map, it usually also focuses on the texture time characteristics of the target object data in the time dimension.

[0107] In a feasible implementation, the input of each type of target object feature map into the texture enhancement module and the output of a texture-enhanced feature map corresponding to each type of target object feature map can be: input each type of target object feature map into the texture enhancement module, perform feature average pooling on the target object feature map through the pooling layer to obtain a pooled feature map, and perform fitting processing on the target object feature map and the pooled feature map through the dense convolution layer based on the residual structure to obtain a texture-enhanced feature map corresponding to the target object feature map.

[0108] Illustratively, the feature average pooling operation on the target object feature map through the pooling layer can obtain a pooled feature map, and then the pooled feature map and the high-frequency part of the original target object feature map are associated and input into the dense convolution layer based on the residual structure for processing, so as to obtain a texture-enhanced feature map corresponding to the target object feature map.

[0109] Illustratively, by focusing on the texture feature information of the object feature map, the texture feature information is enhanced to magnify the details of the object forgery (such as face forgery and body forgery), so as to better perform forgery classification in the subsequent process.

[0110] S208: performing attention feature enhancement on the object feature map through the forgery detection model to obtain an attention feature map;

[0111] ​​In one or more embodiments of the present specification, the forgery detection model can comprise an attention enhancement module, specifically, at least one reference object feature map can be selected from the multi-class object feature map, and the texture-sensitive difference features in the at least one reference object feature map are enhanced by the attention mechanism part in the attention enhancement module to obtain an attention feature map.

[0112] Optionally, the reference object feature map can be an object feature map selected by default from the multi-class object feature map, and the texture-sensitive difference features in the at least one reference object feature map are enhanced by the attention mechanism part in the attention enhancement module to obtain an attention feature map.

[0113] Illustratively, the attention feature enhancement of the object feature map can be visual attention enhancement processing and feature extraction of the visual area from the visual level, which can be understood as the object easy-to-forge area from the visual dimension, and the feature extraction of the visual attention level is enhanced to obtain the attention feature map. The attention feature enhancement of the object feature map can also be semantic attention enhancement processing and feature extraction of the image / video semantic area from the semantic level to obtain the attention feature map.

[0114] Optionally, the attention enhancement module can be a 3D Attention module, and the 3D Attention module can be a single-layer attention mechanism Attention module. When the attention enhancement module enhances the texture-sensitive difference features in the reference object feature map, it also focuses on the feature information of the reference object feature map in the time dimension. The texture-sensitive difference features in the at least one reference object feature map are enhanced by the attention mechanism part in the attention enhancement module to obtain an attention feature map.

[0115] S210: The object fusion feature is obtained by performing feature attention fusion on the object feature map, the texture-enhanced feature map, and the attention feature map using the forgery detection model, and the object forgery detection result is output based on the object fusion feature.

[0116] In one or more embodiments of the present specification, the object feature map, the texture-enhanced feature map, and the attention feature map can be fused by the forgery detection model. The object feature map, the texture-enhanced feature map, and the attention feature map can be subjected to an attention pooling operation to perform multi-layer attention fusion to realize adaptive focusing of the texture-sensitive region at the feature level, and the object fusion feature after a series of enhancements is subjected to forgery detection classification to obtain the object forgery detection result.

[0117] Illustratively, considering the difference between different regional ranges, if global average fusion is used for the "object feature map, texture enhanced feature map and attention feature map", the fused object fusion feature will be affected by the attention map intensity, which is not conducive to the purpose of focusing on texture information. Based on this, an attention pooling operation is used for optimization.

[0118] Illustratively, the forgery detection model can include an attention pooling module. The object feature map, the texture enhanced feature map and the attention feature map are subjected to a bilinear attention pooling operation through the attention pooling module to perform feature attention fusion. In the feature attention fusion process, the object feature map, the texture enhanced feature map and the attention feature map are mapped to the same image resolution and then fused to obtain object fusion features, which are then input into a classifier to classify the input object based on the object fusion features.

[0119] Illustratively, the object forgery detection result is usually an object forgery result or an object real result for target object data. For example, when the target object data is a user face video, the object forgery detection result is a real face video or a fake face video.

[0120] In one or more embodiments of the present specification, an electronic device inputs target object data into a forgery detection model, extracts a multi-class object feature map of the target object data based on the multi-class object feature map through the forgery detection model, and performs feature texture enhancement to obtain a texture enhanced feature map. Then, the object feature map is subjected to attention feature enhancement through the forgery detection model to obtain an attention feature map. Then, the object feature map, the texture enhanced feature map and the attention feature map are subjected to feature attention fusion to obtain object fusion features. Based on the object fusion features, accurate forgery detection can be performed, thereby avoiding the phenomenon of forgery detection result error. The texture information difference area can be adaptively focused, thereby resisting the forgery detection anti-decoding mechanism, better adapting to complex application scenarios to achieve effective detection. In addition, the global detection effect of the model in complex scenarios can be improved, and the robustness and universality of the forgery detection can be ensured.

[0121] Please refer to Figure 8 , Figure 8 is a flowchart of another embodiment of a forgery detection method proposed in one or more embodiments of the present specification. Specifically:

[0122] S302: Create an initial forgery detection model containing an object data decoder, input sample object data into the initial forgery detection model for model training, determine the forgery classification loss of the initial forgery detection model and the object restoration prediction loss determined by the object data decoder;

[0123] The object data decoder is configured to generate object restoration data for sample object data based on the object fusion feature during model training. For example, the sample object data is a sample object image of an object, and the object restoration data is an object restoration image of the sample object image generated based on the object fusion feature. The object restoration image is an image generated by re-decoding the object fusion feature by the object data decoder, and the object restoration image corresponds to the original sample object image.

[0124] In one or more embodiments of the present specification, an initial fake detection model including an object data decoder is pre-created. The initial fake detection model can be obtained after model training of an initial fake detection model created based on a machine learning model. The machine learning model can be one or more of a convolutional neural network (CNN) model, a deep neural network (DNN) model, a recurrent neural network (RNN), a model, an embedding model, a gradient boosting decision tree (GBDT) model, a logistic regression (LR) model, etc.

[0125] Illustratively, a large amount of sample object data can be pre-acquired. The sample object data can be sample object images of image type or sample object videos of video type. The sample object data can be labeled with true or fake labels, i.e., labeled as real object data type or fake data type.

[0126] Illustratively, the initial fake detection model is trained offline using sample object data, i.e., each sample object data is input into the initial fake detection model for model training. During model training, a fake classification loss of the initial fake detection model is determined, and an object restoration prediction loss is determined by the object data encryptor. The fake classification loss and the object restoration prediction loss can be combined to obtain a total model loss. Based on the total model loss, the initial fake detection model is adjusted using a backpropagation algorithm until the initial fake detection model satisfies a model end condition. A trained fake detection model is obtained, which can then be deployed online in an actual fake detection scenario to perform fake detection on target object data.

[0127] In one or more embodiments of the present disclosure, the model end training condition can include, for example, a value of a loss function being less than or equal to a preset loss function threshold, a number of iterations reaching a preset number threshold, and the like. The specific model end training condition can be determined based on actual conditions, which will not be described here.

[0128] In one or more embodiments of the present disclosure, after the initial forgery detection model satisfies the model end condition, the object data decoder in the initial forgery detection model can be discarded, and the remaining model network structure is retained, thereby generating a trained forgery detection model.

[0129] In one possible implementation, the object restoration prediction loss determined by the object data decoder can be:

[0130] The initial forgery detection model generally includes multiple rounds of model training processes, and each round of model training process is based on at least one or a batch of sample object data for model training.

[0131] A2: In each round of model training process, the initial forgery detection model is controlled to perform object restoration processing on the sample object data by the object data decoder to obtain object restoration data.

[0132] Illustratively, after the initial forgery detection model receives the input sample object data, it will perform the steps of: “extracting a multi-class object feature map of the sample object data by the initial forgery detection model, and performing feature texture enhancement based on the multi-class object feature map to obtain a texture enhanced feature map; performing attention feature enhancement on the object feature map by the initial forgery detection model to obtain an attention feature map; performing feature attention fusion on the object feature map, the texture enhanced feature map, and the attention feature map by the initial forgery detection model to obtain object fusion features, and performing forgery detection based on the object fusion features to obtain a forgery identification result for the sample object data”. After the object fusion features are generated, the initial forgery detection model performs object restoration processing on the sample object data by the object data decoder to obtain object restoration data.

[0133] A4: The initial forgery detection model is controlled to perform loss calculation based on the object restoration data and the sample object data using a comparison loss function to obtain an object restoration prediction loss L2.

[0134] The comparison loss function can be a Margin Loss function. The specific form of the Margin Loss function can refer to the form of the Margin Loss function in related technologies.

[0135] The comparison loss function can also be a decoding loss function based on the Euclidean distance.

[0136] A6: The initial forgery detection model is controlled based on a forgery identification result for the sample object data and a forgery reference result for the sample object data, and a forgery classification loss L1 can be calculated.

[0137] In a possible implementation, the determination of the forgery classification loss of the initial forgery detection model can include the following steps: determining a forgery identification result of the initial forgery detection model for the sample object data, obtaining a forgery reference result for the sample object data, and calculating a loss based on the forgery identification result and the forgery reference result by using a cross-entropy loss function to obtain a forgery classification loss L1.

[0138] The forgery reference result is a true or false label previously annotated for the sample object data.

[0139] The forgery identification result is an object forgery detection result output by the initial forgery detection model for the sample object data in the current round.

[0140] The input of the cross-entropy loss function is the forgery reference result and the forgery identification result, and the output is the forgery classification loss.

[0141] S304: Model parameter adjustment is performed on the initial forgery detection model based on the forgery classification loss and the object restoration prediction loss until a model training end condition is met, and a forgery detection model without the object data decoder is obtained.

[0142] In a possible implementation, the model parameter adjustment performed on the initial forgery detection model based on the forgery classification loss and the object restoration prediction loss can include the following steps: inputting the forgery classification loss and the object restoration prediction loss into a model comprehensive calculation formula, outputting a model comprehensive loss, and adjusting the model parameters of the initial forgery detection model by using the model comprehensive loss.

[0143] The model comprehensive calculation formula satisfies the following formula:

[0144] L total = a * L1 + b * L2

[0145] wherein, the L total is the model comprehensive loss, the L1 is the forgery classification loss, the a is a first hyperparameter of the forgery classification loss, the L2 is the object restoration prediction loss, and the b is a second hyperparameter of the object restoration prediction loss.

[0146] Illustratively, in each round of model training process for the initial forgery detection model: by calculating the model comprehensive loss in the model training process, adjusting the parameters of the model network based on the model comprehensive loss, for example, the parameters include the weight values and threshold values of the model network, until the entire initial forgery detection model reaches the end training condition to complete the training, the model network converges, the object data decoder in the initial forgery detection model is discarded and the remaining network structure is retained to obtain the trained forgery detection model for forgery detection. It can be understood that the object data decoder can generate object restoration data in each round, and the object restoration data can intuitively feedback the model processing ability of the initial forgery detection model through visual means. Based on this, the forgery detection ability of the initial forgery detection model in each round of training process can be grasped in real time. The object restoration data can be regarded as a kind of visual model supervision training signal, so as to accelerate the model convergence and improve the model parameter adjustment effect.

[0147] In order to better understand the forgery detection method involved in the present specification, a model training process of a forgery detection model is shown as follows:

[0148] Please refer to Figure 9 , Figure 9 is a scene diagram of model training of a forgery detection model involved in the present specification. An initial forgery detection model can be created in advance. The initial forgery detection model can include several backbone networks, texture enhancement modules, attention enhancement modules, attention pooling modules, classifiers, and data decoders.

[0149] Illustratively, in Figure 9In some embodiments, the plurality of backbone networks can include three, the plurality of backbone networks can include a first backbone network, a second backbone network, and a third backbone network, the first backbone network can be 3D Backbone Layers1, the second backbone network can be 3D Backbone Layers2, and the third backbone network can be 3D Backbone Layers3, and the plurality of backbone networks can be in a form of stacking; the texture enhancement module can include a pooling layer average pooling and a dense convolution layer 3D Dense block based on a residual structure; the attention enhancement module can be a 3D Attention module, the 3D Attention module can be an Attention module of a single-layer attention mechanism; the attention pooling module can be a Bilinear attention pooling module, which is used to perform a Bilinear attention pooling operation for feature attention fusion; the classifier, i.e., Classifier, is used to perform forgery classification according to sample object fusion features; and the data decoder can be an SD decoder, which is used to generate object restoration data for sample object data based on object fusion features in a model training process.

[0150] In some embodiments, a large amount of sample object data is obtained in advance, the sample object data can be sample object videos of a video type, each sample object data is input to an initial forgery detection model such as Figure 9 In the model training process, the plurality of backbone networks (the first backbone network, the second backbone network, and the third backbone network) respectively extract a plurality of object feature maps of the sample object data, i.e., a first object feature map corresponding to the first backbone network, a second object feature map corresponding to the second backbone network, and a third object feature map corresponding to the third backbone network.

[0151] Further, in each round of the model training process of the sample object data, at least one target object feature map is determined from the first object feature map, the second object feature map, and the third object feature map, such as Figure 9 As shown in FIG. 6, Figure 9 The target object feature map determined in the first object feature map and the second object feature map;

[0152] Then the target object feature map is input into the texture enhancement module, and the texture enhancement module respectively outputs the texture enhanced feature map corresponding to the target object feature map. Illustratively, the target object feature map is subjected to feature average pooling through the pooling layer of the texture enhancement module to obtain a pooled feature map; the target object feature map and the pooled feature map are subjected to fitting processing through the dense convolution layer of the texture enhancement module to obtain the texture enhanced feature map corresponding to the target object feature map.

[0153] Illustratively, in each round of model training process on the sample object data, the object feature map output by one or more backbone networks is also subjected to attention feature enhancement to obtain an attention feature map, as shown in Figure 9 Figure 9 It is shown that the attention feature map is obtained by performing attention feature enhancement on the second type of object feature map output by the second backbone network.

[0154] Optionally, at least one type of reference object feature map can be selected from the plurality of types of object feature maps, and the texture sensitive difference features in the at least one type of reference object feature map are subjected to attention enhancement through the attention mechanism part in the attention enhancement module to obtain an attention feature map.

[0155] Illustratively, then the object feature map, the texture enhanced feature map and the attention feature map in each round of model training process are subjected to feature attention fusion to obtain a sample object fusion feature by using an attention pooling module, and further, the sample object fusion feature is input into a classifier for fake classification of the input object. The object fake detection result is usually an object fake result or an object real result for the target object data.

[0156] It should be noted that in each round of model training process, the object data decoder generates object restoration data for the sample object data based on the aforementioned (sample) object fusion feature, for example, the sample object data is a sample object image of an object, and the object restoration data is an object restoration image for the sample object image generated based on the object fusion feature.

[0157] Further, the fake classification loss L1 of each round of initial fake detection model and the object restoration prediction loss L2 determined by the object restoration image are input into a model comprehensive calculation formula to output a model comprehensive loss. The model comprehensive loss is used to adjust the model parameters of the initial fake detection model until the entire initial fake detection model reaches the end training condition to complete the training, the model network converges, the object data decoder in the initial fake detection model is discarded, and the remaining network structure is retained to obtain the trained fake detection model for fake detection.

[0158] ​In one or more embodiments of the present specification, the electronic device inputs target object data into a forgery detection model, extracts a multi-class object feature map of the target object data through the forgery detection model and performs feature texture enhancement based on the multi-class object feature map to obtain a texture-enhanced feature map, then performs attention feature enhancement on the object feature map through the forgery detection model to obtain an attention feature map, and then performs feature attention fusion on the object feature map, the texture-enhanced feature map and the attention feature map to obtain object fusion features, so that accurate forgery detection can be performed based on the object fusion features, thereby avoiding the phenomenon that the forgery detection result is wrong, adaptively focusing on the texture information difference area, thereby resisting the forgery detection anti-decoding mechanism, better adapting to complex application scenarios to achieve effective detection, and improving the global detection effect of the model in complex scenarios to ensure the robustness and universality of the forgery detection.

[0159] The following will be described in combination with Figure 10 The forgery detection device provided in the present specification will be described in detail. It should be noted that Figure 10 The forgery detection device shown in the present specification is used to execute the method of the present specification Figures 1-9 The method of the embodiment shown in the present specification is only shown in relation to the present specification for the purpose of illustration, and specific technical details not disclosed are described with reference to the present specification Figures 1-9 The embodiment shown in the present specification.

[0160] Please refer to Figure 10 which shows the structure diagram of the forgery detection device of the present specification. The forgery detection device 1 can be realized by software, hardware or a combination of the two to become all or part of the user terminal. According to some embodiments, the forgery detection device 1 includes a forgery detection module 11, a forgery detection module 12 and a forgery detection module 13, which are specifically used for:

[0161] The texture processing module 11 is configured to input target object data into a forgery detection model, extract a multi-class object feature map of the target object data through the forgery detection model, and perform feature texture enhancement based on the multi-class object feature map to obtain a texture-enhanced feature map.

[0162] The attention enhancement module 12 is configured to perform attention feature enhancement on the object feature map through the forgery detection model to obtain an attention feature map.

[0163] The fusion processing module 13 is configured to perform feature attention fusion on the object feature map, the texture-enhanced feature map and the attention feature map through the forgery detection model to obtain object fusion features, and output an object forgery detection result based on the object fusion features.

[0164] Optionally, the forgery detection model includes a plurality of backbone networks and a texture enhancement module, such as Figure 11As shown, the texture processing module 11 comprises:

[0165] The feature map extraction unit 111 is configured to extract a plurality of object feature maps of the target object data respectively through the plurality of backbone networks.

[0166] The texture enhancement unit 112 is configured to determine at least one target object feature map from the plurality of object feature maps, input each of the target object feature maps into the texture enhancement module, and output a texture enhanced feature map corresponding to each of the target object feature maps respectively.

[0167] Optionally, the feature map extraction unit 111 is configured to:

[0168] determine a default at least one target object feature map from the plurality of object feature maps; or,

[0169] perform texture enhanced map prediction processing on the plurality of object feature maps to obtain at least one target object feature map.

[0170] Optionally, the texture enhancement unit 112 is configured to:

[0171] perform artifact residual detection on the object feature map to obtain a residual prediction probability corresponding to the object feature map, and determine at least one target object feature map from the plurality of object feature maps based on the residual prediction probability.

[0172] Optionally, the texture enhancement module comprises a pooling layer and a dense convolution layer based on a residual structure, and the texture enhancement unit 112 is configured to:

[0173] input each of the target object feature maps into the texture enhancement module, perform feature average pooling on the target object feature map through the pooling layer to obtain a pooled feature map,

[0174] perform fitting processing on the target object feature map and the pooled feature map through the dense convolution layer to obtain a texture enhanced feature map corresponding to the target object feature map.

[0175] Optionally, the plurality of backbone networks comprises a first backbone network, a second backbone network and a third backbone network, and the feature map extraction unit 111 is configured to:

[0176] extract a first type of object feature map of the target object data through the first backbone network, extract a second type of object feature map of the target object data through the second backbone network, and extract a third type of object feature map of the target object data through the third backbone network,

[0177] The feature resolution of the second type of object feature map is smaller than that of the first type of object feature map and larger than that of the third type of object feature map, and the semantic expression degree of the second type of object feature map is larger than that of the first type of object feature map and smaller than that of the third type of object feature map.

[0178] Optionally, the texture enhancement unit 112 is configured to:

[0179] The first type of object feature map and the second type of object feature map are determined from the multi-type object feature map, the first type of object feature map and the second type of object feature map are input into the texture enhancement module respectively, and the first texture enhancement feature map corresponding to the first type of object feature map and the second texture enhancement feature map corresponding to the second type of object feature map are output.

[0180] The first type of object feature map, the second type of object feature map and the third type of object feature map are determined from the multi-type object feature map, the first type of object feature map, the second type of object feature map and the third type of object feature map are input into the texture enhancement module respectively, and the first texture enhancement feature map corresponding to the first type of object feature map, the second texture enhancement feature map corresponding to the second type of object feature map and the third texture enhancement feature map corresponding to the third type of object feature map are output.

[0181] The first type of object feature map is determined from the multi-type object feature map, and the first type of object feature map is input into the texture enhancement module, and the first texture enhancement feature map corresponding to the first type of object feature map is output.

[0182] Optionally, the forgery detection model comprises an attention enhancement module, and the attention enhancement module is configured to:

[0183] At least one type of reference object feature map is selected from the multi-type object feature map, and the at least one type of reference object feature map is subjected to attention enhancement through the attention enhancement module to obtain an attention feature map.

[0184] Optionally, the device 1 is further configured to:

[0185] An initial forgery detection model comprising an object data decoder is created, sample object data is input into the initial forgery detection model for model training, a forgery classification loss of the initial forgery detection model is determined, and an object restoration prediction loss is determined through the object data decoder;

[0186] and the initial forgery detection model is subjected to model parameter adjustment based on the forgery classification loss and the object restoration prediction loss until a model training end condition is met, and a forgery detection model without the object data decoder is obtained.

[0187] Optionally, the device 1 is further used for:

[0188] In each round of model training, the sample object data is subjected to object restoration processing by the object data decoder to obtain object restoration data;

[0189] Based on the object restoration data and the sample object data, a loss calculation is performed using a comparison loss function to obtain an object restoration prediction loss.

[0190] Optionally, the device 1 is further used for:

[0191] The initial forgery detection model is determined to obtain a forgery identification result for the sample object data, and a forgery reference result for the sample object data is obtained;

[0192] Based on the forgery identification result and the forgery reference result, a loss calculation is performed using a cross-entropy loss function to obtain a forgery classification loss.

[0193] Optionally, the device 1 is further used for:

[0194] The forgery classification loss and the object restoration prediction loss are input into a model comprehensive calculation formula to output a model comprehensive loss, and the model parameter of the initial forgery detection model is adjusted using the model comprehensive loss;

[0195] The model comprehensive calculation formula satisfies the following formula:

[0196] L total =a*L1+b*L2

[0197] Wherein, the L total is a model comprehensive loss, the L1 is the forgery classification loss, the a is a first hyperparameter of the forgery classification loss, the L2 is the object restoration prediction loss, and the b is a second hyperparameter of the object restoration prediction loss.

[0198] It should be noted that the above embodiment provides a forgery detection device for executing a forgery detection method, and only the division of the above functional modules is exemplified, and in actual application, the above functions can be completed by different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the above described functions. In addition, the forgery detection device and the forgery detection method provided by the above embodiment belong to the same concept, and the implementation process is described in detail in the method embodiment, which will not be repeated here.

[0199] The above serial numbers in the specification are only for description, and do not represent the advantages or disadvantages of the embodiments.

[0200] In one or more embodiments of the present specification, the electronic device inputs target object data into a forgery detection model, extracts multi-class object feature maps of the target object data through the forgery detection model and performs feature texture enhancement based on the multi-class object feature maps to obtain texture-enhanced feature maps, then performs attention feature enhancement on the object feature maps through the forgery detection model to obtain attention feature maps, and then performs feature attention fusion on the object feature maps, the texture-enhanced feature maps and the attention feature maps to obtain object fusion features. Based on the object fusion features, accurate forgery detection can be performed, thereby avoiding the phenomenon that the forgery detection result is wrong, focusing on the texture information difference area adaptively, thereby resisting the forgery detection anti-decoding mechanism, better adapting to complex application scenarios to achieve effective detection; and the global detection effect of the model in complex scenarios can be improved, and the robustness and universality of the forgery detection are ensured.

[0201] The present specification also provides a computer storage medium, which can store a plurality of instructions, the instructions being adapted to be loaded and executed by a processor to perform the forgery detection method of the above-mentioned Figures 1-9 embodiment. The specific implementation process can be referred to the specific description of the above-mentioned Figures 1-9 embodiment, which will not be repeated here.

[0202] The present specification also provides a computer program product, which stores at least one instruction, the at least one instruction being loaded and executed by the processor to perform the forgery detection method of the above-mentioned Figures 1-9 embodiment. The specific implementation process can be referred to the specific description of the above-mentioned Figures 1-9 embodiment, which will not be repeated here.

[0203] Please refer to Figure 12 , which shows the structural block diagram of the electronic device provided by an exemplary embodiment of the present specification. The electronic device in the present specification can include one or more of the following components: a processor 110, a memory 120, an input device 130, an output device 140 and a bus 150. The processor 110, the memory 120, the input device 130 and the output device 140 can be connected through the bus 150.

[0204] The processor 110 can include one or more processing cores. The processor 110 connects various parts within the entire electronic device with various interfaces and lines, performs various functions of the electronic device 100 and processes data by running or executing instructions, programs, code sets or instruction sets stored in the memory 120, and calling data stored in the memory 120. Alternatively, the processor 110 can be implemented in at least one of a hardware form of a digital signal processing (DSP), a field-programmable gate array (FPGA), a programmable logic array (PLA). The processor 110 can integrate a combination of one or several of a central processing unit (CPU), a graphics processing unit (GPU), and a modem, etc. Among them, the CPU mainly processes an operating system, a user interface, and an application program, etc.; the GPU is responsible for rendering and drawing display content; and the modem is used to process wireless communication. It can be understood that the above-mentioned modem can also not be integrated into the processor 110, but be implemented by a separate communication chip.

[0205] The memory 120 can include a random access memory (RAM) and can also include a read-only memory (ROM). Alternatively, the memory 120 includes a non-transitory computer-readable storage medium. The memory 120 can be used to store instructions, programs, codes, code sets or instruction sets. The memory 120 can include a program storage area and a data storage area, wherein the program storage area can store instructions for implementing an operating system, instructions for implementing at least one function (such as a touch function, a sound playing function, an image playing function, etc.), instructions for implementing each of the following method embodiments, etc., the operating system can be an Android system, an IOS system developed by Apple Inc., a system developed based on the Android system or other systems. The data storage area can also store data created by the electronic device in use, such as a phone book, audio and video data, chat record data, etc.

[0206] Referring to Figure 13As shown, the memory 120 can be divided into an operating system space and a user space, the operating system runs in the operating system space, and native and third-party applications run in the user space. In order to ensure that different third-party applications can achieve good running effect, the operating system allocates corresponding system resources for different third-party applications. However, there are also differences in the demand for system resources in different application scenarios in the same third-party application. For example, in the local resource loading scenario, the third-party application has a higher requirement for the disk reading speed; in the animation rendering scenario, the third-party application has a higher requirement for the GPU performance. However, the operating system and the third-party application are independent of each other, and the operating system often cannot timely perceive the current application scenario of the third-party application, resulting in that the operating system cannot perform targeted system resource adaptation according to the specific application scenario of the third-party application.

[0207] In order to enable the operating system to distinguish the specific application scenario of the third-party application, it is necessary to open up the data communication between the third-party application and the operating system, so that the operating system can obtain the current scenario information of the third-party application at any time, and then perform targeted system resource adaptation based on the current scenario.

[0208] Taking the Android system as an example, the programs and data stored in the memory 120 are as follows Figure 14As shown, the memory 120 can store a Linux kernel layer 320, a system runtime library layer 340, an application framework layer 360, and an application layer 380, wherein the Linux kernel layer 320, the system runtime library layer 340, and the application framework layer 360 belong to an operating system space, and the application layer 380 belongs to a user space. The Linux kernel layer 320 provides underlying drivers for various hardware of the electronic device, such as display drivers, audio drivers, camera drivers, Bluetooth drivers, Wi-Fi drivers, power management, and the like. The system runtime library layer 340 provides main feature support for the Android system through some C / C++ libraries. For example, an SQLite library provides database support, an OpenGL / ES library provides 3D drawing support, a Webkit library provides browser kernel support, and the like. An Android runtime is also provided in the system runtime library layer 340, which mainly provides some core libraries to allow developers to use the Java language to write Android applications. The application framework layer 360 provides various APIs that can be used when building an application, and developers can also build their own applications by using these APIs, such as activity management, window management, view management, notification management, content provider, package management, call management, resource management, and location management. At least one application program is running in the application layer 380, which can be native applications provided by the operating system, such as a contact program, a message program, a clock program, a camera application, and the like, or third-party applications developed by third-party developers, such as game applications, instant messaging programs, photo beautification programs, and the like.

[0209] For example, taking an IOS system as the operating system, the programs and data stored in the memory 120 can include an IOS kernel 310, an IOS runtime library 330, an application framework 350, and an application 370. Figure 15As shown, the IOS system includes: a core operating system layer 420, a core service layer 440, a media layer 460, and a Cocoa Touch layer 480. The core operating system layer 420 includes an operating system kernel, drivers, and low-level hardware abstractions that provide more hardware-specific functionality to program frameworks in the core service layer 440. The core service layer 440 provides system services and / or program frameworks that applications need, such as a Foundation framework, an account framework, an advertisement framework, a data storage framework, a network connection framework, a geographic location framework, a motion framework, and the like. The media layer 460 provides interfaces for applications related to audio and video, such as interfaces related to graphics images, interfaces related to audio technology, interfaces related to video technology, an AirPlay interface for wireless audio and video transmission technology, and the like. The Cocoa Touch layer 480 provides various commonly used interface-related frameworks for application development, and is responsible for user touch interaction on the electronic device. For example, a local notification service, a remote push service, an advertisement framework, a game tool framework, a message user interface (UI) framework, a user interface UIKit framework, a map framework, and the like.

[0210] In Figure 15 In the framework shown, the frameworks related to most applications include, but are not limited to, the Foundation framework in the core service layer 440 and the UIKit framework in the Cocoa Touch layer 480. The Foundation framework provides many basic object classes and data types, and provides the most basic system services for all applications, and is UI-independent. The UIKit framework provides basic UI class libraries for creating touch-based user interfaces, and iOS applications can provide UIs based on the UIKit framework, so it provides the basic framework of the application for building user interfaces, drawing, processing and user interaction events, responding to gestures, and the like.

[0211] In the IOS system, the manner and principle of implementing data communication between a third-party application and an operating system can refer to the Android system, and the present specification will not be repeated here.

[0212] The input device 130 is configured to receive input instructions or data, and the input device 130 includes but is not limited to a keyboard, a mouse, a camera, a microphone, or a touch device. The output device 140 is configured to output instructions or data, and the output device 140 includes but is not limited to a display device and a speaker. In an example, the input device 130 and the output device 140 can be combined, and the input device 130 and the output device 140 are a touch display screen configured to receive a touch operation of a user using a finger, a stylus, or any suitable object on or near the touch display screen, and display a user interface of each application. The touch display screen is usually arranged on a front panel of the electronic device. The touch display screen can be designed as a full screen, a curved screen, or a special-shaped screen. The touch display screen can also be designed as a combination of a full screen and a curved screen, a combination of a special-shaped screen and a curved screen, which is not limited in the present specification.

[0213] In addition, those skilled in the art can understand that the structure of the electronic device shown in the above-described drawings does not constitute a limitation on the electronic device, and the electronic device can include more or fewer components than those shown in the drawings, or combine certain components, or different component arrangements. For example, the electronic device further includes a radio frequency circuit, an input unit, a sensor, an audio circuit, a wireless fidelity (WiFi) module, a power supply, a Bluetooth module, and the like, which are not described herein.

[0214] In the present specification, the execution subject of each step can be the electronic device described above. Alternatively, the execution subject of each step is an operating system of the electronic device. The operating system can be an Android system, an IOS system, or other operating systems, which are not limited in the present specification.

[0215] The electronic device of the present specification can further have a display device installed thereon. The display device can be various devices capable of realizing a display function, such as a cathode ray tube display (CR), a light-emitting diode display (LED), an electronic ink screen, a liquid crystal display (LCD), a plasma display panel (PDP), and the like. A user can use the display device on the electronic device 101 to view displayed text, images, videos, and the like. The electronic device can be a smart phone, a tablet computer, a game device, an AR (Augmented Reality) device, a car, a data storage device, an audio playback device, a video playback device, a notebook computer, a desktop computing device, a wearable device such as an electronic watch, electronic glasses, an electronic helmet, an electronic bracelet, an electronic necklace, an electronic clothing, and the like.

[0216] In Figure 12 In the electronic device shown, the processor 110 can be configured to invoke a network optimization application stored in the memory 120, and specifically perform the following operations:

[0217] inputting the target object data into the forgery detection model, extracting a multi-class object feature map of the target object data through the forgery detection model, and performing feature texture enhancement based on the multi-class object feature map to obtain a texture-enhanced feature map;

[0218] performing attention feature enhancement on the object feature map through the forgery detection model to obtain an attention feature map;

[0219] performing feature attention fusion on the object feature map, the texture-enhanced feature map, and the attention feature map through the forgery detection model to obtain an object fusion feature, and outputting an object forgery detection result based on the object fusion feature.

[0220] In one embodiment, the forgery detection model includes a plurality of backbone networks and a texture enhancement module, and the processor 110 performs the following steps when performing the extracting a multi-class object feature map through the forgery detection model, and performing feature texture enhancement based on the multi-class object feature map to obtain a texture-enhanced feature map:

[0221] extracting a multi-class object feature map of the target object data through each of the backbone networks;

[0222] determining at least one type of target object feature map from the multi-class object feature map, inputting each type of target object feature map into the texture enhancement module, and outputting a texture-enhanced feature map corresponding to each type of target object feature map.

[0223] In one embodiment, the processor 110 performs the following steps when performing the determining at least one type of target object feature map from the multi-class object feature map:

[0224] determining a default at least one type of target object feature map from the multi-class object feature map; or

[0225] performing texture-enhanced map prediction processing on the multi-class object feature map to obtain at least one type of target object feature map.

[0226] In one embodiment, the processor 110 performs the following steps when performing the performing texture-enhanced map prediction processing on the multi-class object feature map to obtain at least one type of target object feature map:

[0227] performing artifact residual detection on the object feature map to obtain a residual prediction probability corresponding to the object feature map, and determining at least one target object feature map from the multiple object feature maps based on the residual prediction probability.

[0228] In one embodiment, the texture enhancement module includes a pooling layer and a dense convolution layer based on a residual structure, and the processor 110, when performing the inputting of each of the target object feature maps into the texture enhancement module to respectively output the texture-enhanced feature maps corresponding to each of the target object feature maps, performs the following steps:

[0229] inputting each of the target object feature maps into the texture enhancement module, performing feature average pooling on the target object feature map through the pooling layer to obtain a pooled feature map;

[0230] performing fitting processing on the target object feature map and the pooled feature map through the dense convolution layer to obtain the texture-enhanced feature map corresponding to the target object feature map.

[0231] In one embodiment, the multiple backbone networks include a first backbone network, a second backbone network, and a third backbone network, and the processor 110, when performing the extracting of the multiple object feature maps of the target object data by each of the backbone networks, performs the following steps:

[0232] extracting a first object feature map of the target object data through the first backbone network, extracting a second object feature map of the target object data through the second backbone network, and extracting a third object feature map of the target object data through the third backbone network,

[0233] wherein the feature resolution of the second object feature map is smaller than that of the first object feature map and larger than that of the third object feature map, and the semantic expression degree of the second object feature map is greater than that of the first object feature map and smaller than that of the third object feature map.

[0234] In one embodiment, the processor 110, when performing the determining of at least one target object feature map from the multiple object feature maps, inputs each of the target object feature maps into the texture enhancement module to respectively output the texture-enhanced feature maps corresponding to each of the target object feature maps, performs the following steps:

[0235] determines a first object feature map and a second object feature map from the multiple object feature maps, inputs the first object feature map and the second object feature map into the texture enhancement module respectively, and outputs a first texture-enhanced feature map corresponding to the first object feature map and a second texture-enhanced feature map corresponding to the second object feature map; or,

[0236] determining a first-class object feature map, a second-class object feature map and a third-class object feature map from the plurality of object feature maps, inputting the first-class object feature map, the second-class object feature map and the third-class object feature map into the texture enhancement module respectively, and outputting a first texture enhanced feature map corresponding to the first-class object feature map, a second texture enhanced feature map corresponding to the second-class object feature map and a third texture enhanced feature map corresponding to the third-class object feature map; or,

[0237] determining a first-class object feature map from the plurality of object feature maps, inputting the first-class object feature map into the texture enhancement module, and outputting a first texture enhanced feature map corresponding to the first-class object feature map.

[0238] In one embodiment, the processor 110, in performing the attention feature enhancement on the object feature map through the forgery detection model, performs the following steps:

[0239] selecting at least one reference object feature map from the plurality of object feature maps, and performing attention enhancement on the at least one reference object feature map through the attention enhancement module to obtain an attention feature map.

[0240] In one embodiment, the processor 110, before performing the inputting of the target object data into the forgery detection model, further performs the following steps:

[0241] creating an initial forgery detection model containing an object data decoder, inputting sample object data into the initial forgery detection model for model training, determining a forgery classification loss of the initial forgery detection model and an object restoration prediction loss determined through the object data decoder;

[0242] and adjusting the model parameters of the initial forgery detection model based on the forgery classification loss and the object restoration prediction loss until a model training end condition is met, to obtain a forgery detection model without the object data decoder.

[0243] In one embodiment, the processor 110, in performing the determination of the object restoration prediction loss through the object data decoder, performs the following steps:

[0244] In each round of model training process, performing object restoration processing on the sample object data through the object data decoder to obtain object restoration data;

[0245] calculating the loss based on the object restoration data and the sample object data using a comparison loss function to obtain an object restoration prediction loss.

[0246] In one embodiment, the processor 110, in performing the determining of the forgery classification loss of the initial forgery detection model, performs the following steps:

[0247] determining a forgery identification result of the initial forgery detection model for the sample object data, obtaining a forgery reference result for the sample object data;

[0248] based on the forgery identification result and the forgery reference result, performing loss calculation using a cross-entropy loss function to obtain a forgery classification loss.

[0249] In one embodiment, the processor 110, in performing the model parameter adjustment of the initial forgery detection model based on the forgery classification loss and the object restoration prediction loss, performs the following steps:

[0250] inputting the forgery classification loss and the object restoration prediction loss into a model comprehensive calculation formula, outputting a model comprehensive loss, and using the model comprehensive loss to adjust the model parameters of the initial forgery detection model;

[0251] The model comprehensive calculation formula satisfies the following formula:

[0252] L total = a*L1 + b*L2

[0253] wherein, the L total is a model comprehensive loss, the L1 is the forgery classification loss, the a is a first hyperparameter of the forgery classification loss, the L2 is the object restoration prediction loss, and the b is a second hyperparameter of the object restoration prediction loss.

[0254] In one or more embodiments of the present specification, the electronic device inputs target object data into a forgery detection model, extracts multi-class object feature maps of the target object data through the forgery detection model and performs feature texture enhancement based on the multi-class object feature maps to obtain texture-enhanced feature maps, then performs attention feature enhancement on the object feature maps through the forgery detection model to obtain attention feature maps, and finally performs feature attention fusion on the object feature maps, the texture-enhanced feature maps, and the attention feature maps to obtain object fusion features. Based on the object fusion features, accurate forgery detection can be performed, thereby avoiding the phenomenon of incorrect forgery detection results, adaptively focusing on texture information difference areas, thereby resisting forgery detection reverse mechanism, better adapting to complex application scenarios to achieve effective detection; and the global detection effect of the model in complex scenarios can be improved, and the robustness and universality of forgery detection are ensured.

[0255] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by a computer program instructing relevant hardware, and the program can be stored in a computer readable storage medium. When the program is executed, it can include the processes of the above-mentioned embodiment methods. The storage medium can be a magnetic disc, an optical disc, a read-only memory or a random access memory, etc.

[0256] It should be noted that the information (including but not limited to user equipment information, user personal information, etc.), data (including but not limited to data for analysis, stored data, displayed data, etc.) and signals involved in the embodiments of the present specification are all authorized by the user or fully authorized by all parties, and the collection, use and processing of related data need to comply with relevant laws, regulations and standards of relevant countries and regions. For example, the target object data, feature processing and forgery detection involved in the present specification are all carried out under full authorization.

[0257] The above only discloses the preferred embodiments of the present specification, and of course cannot limit the scope of the rights of the present specification, so equivalent changes made according to the claims of the present specification still fall within the scope covered by the present specification.

Claims

1. A method for forgery detection, the method comprising: inputting target object data into a forgery detection model, extracting multi-class object feature maps of the target object data by the forgery detection model, and determining at least one class of target object feature maps from the multi-class object feature maps, inputting each class of the target object feature maps into a texture enhancement module of the forgery detection model, and respectively outputting texture enhanced feature maps corresponding to each class of the target object feature maps; performing attention feature enhancement on the object feature maps by the forgery detection model to obtain attention feature maps; performing feature attention fusion on the object feature maps, the texture enhanced feature maps, and the attention feature maps by the forgery detection model to obtain object fusion features, and outputting an object forgery detection result based on the object fusion features; wherein the determining at least one class of target object feature maps from the multi-class object feature maps comprises: setting a default target object feature map from all object feature maps based on a transaction scene type to which the forgery detection model is applied; or performing texture enhanced map prediction processing on the multi-class object feature maps to obtain at least one class of target object feature maps, the texture enhanced map prediction processing being prediction of target object feature maps that need to be texture enhanced from the multi-class object feature maps.

2. The method of claim 1, wherein the forgery detection model comprises a plurality of backbone networks and a texture enhancement module, the extracting multi-class object feature maps by the forgery detection model comprises: extracting multi-class object feature maps of the target object data by each of the backbone networks.

3. The method of claim 1, wherein the performing texture enhanced map prediction processing on the multi-class object feature maps to obtain at least one class of target object feature maps comprises: performing artifact residual detection on the object feature maps to obtain residual prediction probabilities corresponding to the object feature maps, and determining at least one class of target object feature maps from the multi-class object feature maps based on the residual prediction probabilities.

4. The method of claim 1, wherein the texture enhancement module comprises a pooling layer and a dense convolution layer based on a residual structure, the inputting each class of the target object feature maps into the texture enhancement module and respectively outputting texture enhanced feature maps corresponding to each class of the target object feature maps comprises: inputting each class of the target object feature maps into the texture enhancement module, performing feature average pooling on the target object feature maps by the pooling layer to obtain pooled feature maps, and performing fitting processing on the target object feature maps and the pooled feature maps by the dense convolution layer to obtain the texture enhanced feature maps corresponding to the target object feature maps.

5. The method of claim 2, wherein the plurality of backbone networks comprises a first backbone network, a second backbone network, and a third backbone network, the extracting multi-class object feature maps of the target object data by each of the backbone networks comprises: ​ extracting a first type of object feature map of the target object data through the first backbone network, extracting a second type of object feature map of the target object data through the second backbone network, and extracting a third type of object feature map of the target object data through the third backbone network, wherein a feature resolution of the second type of object feature map is smaller than that of the first type of object feature map and larger than that of the third type of object feature map, and a semantic expression degree of the second type of object feature map is larger than that of the first type of object feature map and smaller than that of the third type of object feature map.

6. The method of claim 5, wherein the determining at least one type of target object feature map from the multi-type object feature map, inputting each type of the target object feature map into a texture enhancement module of the forgery detection model, and respectively outputting a texture enhanced feature map corresponding to each type of the target object feature map comprises: determining a first type of object feature map and a second type of object feature map from the multi-type object feature map, inputting the first type of object feature map and the second type of object feature map into the texture enhancement module of the forgery detection model, and outputting a first texture enhanced feature map corresponding to the first type of object feature map and a second texture enhanced feature map corresponding to the second type of object feature map; or, determining a first type of object feature map and a third type of object feature map from the multi-type object feature map, inputting the first type of object feature map and the third type of object feature map into the texture enhancement module of the forgery detection model, and outputting a first texture enhanced feature map corresponding to the first type of object feature map and a third texture enhanced feature map corresponding to the third type of object feature map; or, determining a first type of object feature map, a second type of object feature map, and a third type of object feature map from the multi-type object feature map, inputting the first type of object feature map, the second type of object feature map, and the third type of object feature map into the texture enhancement module of the forgery detection model, and outputting a first texture enhanced feature map corresponding to the first type of object feature map, a second texture enhanced feature map corresponding to the second type of object feature map, and a third texture enhanced feature map corresponding to the third type of object feature map; or, determining a first type of object feature map from the multi-type object feature map, inputting the first type of object feature map into the texture enhancement module of the forgery detection model, and outputting a first texture enhanced feature map corresponding to the first type of object feature map.

7. The method of claim 1, wherein the forgery detection model comprises an attention enhancement module, and the attention feature enhancement of the object feature map through the forgery detection model comprises: selecting at least one type of reference object feature map from the multi-type object feature map, and performing attention enhancement on the at least one type of reference object feature map through the attention enhancement module to obtain an attention feature map.

8. The method of any one of claims 1-7, wherein the inputting of the target object data into the forgery detection model further comprises: creating an initial forgery detection model containing an object data decoder, inputting sample object data into the initial forgery detection model for model training, determining a forgery classification loss of the initial forgery detection model and determining an object restoration prediction loss through the object data decoder; and adjusting model parameters of the initial forgery detection model based on the forgery classification loss and the object restoration prediction loss until a model training end condition is met, obtaining a forgery detection model without the object data decoder.

9. The method of claim 8, wherein the determining an object restoration prediction loss through the object data decoder comprises: in each round of model training, performing object restoration processing on the sample object data through the object data decoder to obtain object restoration data; based on the object restoration data and the sample object data, performing loss calculation using a comparison loss function to obtain an object restoration prediction loss.

10. The method of claim 8, wherein the determining a forgery classification loss of the initial forgery detection model comprises: determining a forgery identification result of the initial forgery detection model for the sample object data, obtaining a forgery reference result for the sample object data; based on the forgery identification result and the forgery reference result, performing loss calculation using a cross-entropy loss function to obtain a forgery classification loss.

11. The method of claim 8, wherein the adjusting model parameters of the initial forgery detection model based on the forgery classification loss and the object restoration prediction loss comprises: inputting the forgery classification loss and the object restoration prediction loss into a model comprehensive calculation formula, outputting a model comprehensive loss, and adjusting model parameters of the initial forgery detection model using the model comprehensive loss; the model comprehensive calculation formula satisfies the following formula: L total =a*L1+b*L2 wherein the L total is a model synthesis loss, the L1 is the fake classification loss, the a is a first hyperparameter of the fake classification loss, the L2 is the object restoration prediction loss, and the b is a second hyperparameter of the object restoration prediction loss.

12. A forgery detection device, the device comprising: a texture processing module configured to input target object data into a forgery detection model, extract a plurality of object feature maps of the target object data from the forgery detection model, determine at least one target object feature map from the plurality of object feature maps, input each of the target object feature maps into a texture enhancement module of the forgery detection model, and output a texture enhanced feature map corresponding to each of the target object feature maps, respectively; an attention enhancement module configured to perform attention feature enhancement on the object feature maps through the forgery detection model to obtain an attention feature map; a fusion processing module configured to perform feature attention fusion on the object feature maps, the texture enhanced feature maps, and the attention feature map through the forgery detection model to obtain an object fusion feature, and output an object forgery detection result based on the object fusion feature; wherein the determining at least one target object feature map from the plurality of object feature maps comprises: setting a default target object feature map from all object feature maps based on a transaction scenario type to which the forgery detection model is applied; or, The texture enhanced map prediction processing is performed on the multi-class object feature map to obtain at least one target object feature map, and the texture enhanced map prediction processing is to predict the target object feature map which needs to be texture enhanced from the multi-class object feature map.

13. A computer storage medium, the computer storage medium storing a plurality of instructions, the instructions being adapted to be loaded and executed by a processor to perform the method steps of any one of claims 1-11.

14. A computer program product, the computer program product storing at least one instruction, the at least one instruction being loaded and executed by a processor to perform the method steps of any one of claims 1-11.

15. An electronic device comprising: A processor and a memory; wherein the memory stores a computer program, and the computer program is adapted to be loaded and executed by the processor to perform the method steps of any one of claims 1-11.

Citation Information

Patent Citations

  • False face video detection method and system based on multi-feature fusion

    CN114898432A

  • Exposing inpainting image forgery under combination attacks with hybrid large feature mining

    US20170091588A1