A Fraudulent Image Recognition Method for Artificial Intelligence Ethics

By extracting the deep visual features and semantic features of fraudulent images and using channel and cross attention mechanisms to fusion of features, the problem of poor recognition efficiency, accuracy and stability of fraudulent images in the prior art is solved, and more efficient and accurate recognition results are achieved.

CN118968183BActive Publication Date: 2025-05-27CENT SOUTH UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411153187.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-08-21
Publication Date
2025-05-27
Estimated Expiration
2044-08-21

AI Technical Summary

Technical Problem

The prior art has poor recognition efficiency, accuracy and stability when processing complex fraud images, especially when processing high-resolution images or large-scale data processing, with high computational complexity and excessive memory usage.

Method used

A fraud image recognition method oriented towards artificial intelligence ethics is adopted. By obtaining fraud image data sets and annotating them, the shallow visual features are extracted using residual modules and U-Net networks, deep visual features and semantic features are extracted in combination with visual state space modules and text state space modules, and feature fusion is adopted for feature fusion by channel attention mechanisms and cross attention mechanisms, and the recognition model with U-Net and Mamba models as the basic structure is trained.

Benefits of technology

It improves the accuracy and stability of the model in fraudulent image recognition, reduces the computational complexity and memory usage, and provides more efficient recognition results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118968183B_ABST
    Figure CN118968183B_ABST
Patent Text Reader

Abstract

In an embodiment of the present disclosure, a fraud image recognition method for artificial intelligence ethics is provided, belonging to the technical field of image recognition, and specifically including: annotating sample fraud images and content description information in a fraud image dataset according to an annotation system; performing different preprocessing on the sample fraud images and content description information; using a residual module and a U-Net network to extract shallow visual features of the sample fraud images; using a visual state space module to extract deep visual features of the sample fraud images; using a text state space module to extract semantic features of the sample fraud images; integrating deep visual features and semantic features by combining a channel attention mechanism and a cross-attention mechanism to obtain fused features; training an identification model by combining the fused features and the labels corresponding to the sample fraud images; and inputting a target fraud image into the trained identification model to obtain an identification result. Through the solution of the present disclosure, the recognition efficiency, accuracy, and stability are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present disclosure relate to the field of image recognition technology, and particularly to a fraud image recognition method for artificial intelligence ethics. Background Art

[0002] At present, with the rapid development of artificial intelligence technology, generative artificial intelligence is widely used in fields such as text, audio, and images, and can create highly realistic content. However, this technological progress has also brought significant ethical risks. Especially in the intelligent generation of forged news and tampered false identity image information, it may be exploited by malicious actors for various fraud activities, which not only damage the interests of individuals and organizations, but also cause certain damage to the social trust mechanism and information ecosystem. Although existing image recognition technologies have made certain progress in some aspects, there are still many limitations in dealing with complex fraud images. The traditional Attention mechanism has a high computational complexity in capturing global image information and local details. Especially when dealing with high-resolution images or large-scale data, it is easy to cause excessive memory occupancy and low operation efficiency.

[0003] It can be seen that there is an urgent need for a precise, efficient, and stable fraud image recognition method for artificial intelligence ethics. Summary of the Invention

[0004] In view of this, embodiments of the present disclosure provide a fraud image recognition method for artificial intelligence ethics, which at least partially solves the problems of poor recognition efficiency, accuracy, and stability in the prior art.

[0005] Embodiments of the present disclosure provide a fraud image recognition method for artificial intelligence ethics, including:

[0006] Step 1: Obtain a fraud image dataset, and label the sample fraud images and content description information in the fraud image dataset according to the annotation system;

[0007] Step 2: Perform different preprocessings on the sample fraud images and content description information;

[0008] Step 3: Use the residual module and the U-Net network to extract the shallow visual features of the sample fraud images;

[0009] Step 4: Based on the shallow visual features, use the visual state space module to extract the deep visual features of the sample fraud images;

[0010] Step 5: Based on the content description information, use the text state space module to extract the semantic features of the sample fraud images;

[0011] Step 6, integrate the channel attention mechanism and the cross-attention mechanism to fuse the deep visual features and semantic features to obtain the fused features;

[0012] Step 7, combine the fused features and the labels corresponding to the sample fraud images to train an identification model based on the U-Net and Mamba models;

[0013] Step 8, input the target fraud image into the trained identification model to obtain the identification result.

[0014] According to a specific implementation manner of the embodiment of the present disclosure, step 1 specifically includes:

[0015] Step 1.1, generate various sample fraud images related to fraud content through a generative artificial intelligence model;

[0016] Step 1.2, according to the definition rules in the annotation system, label the sample fraud images with appropriate fraud type labels, and at the same time add text descriptions corresponding to the image content.

[0017] According to a specific implementation manner of the embodiment of the present disclosure, step 2 specifically includes:

[0018] Step 2.1, perform data augmentation operations on each sample fraud image to generate image pairs with different perspectives and angles, where the data augmentation operations include rotation, scaling, flipping, color jittering, and cropping;

[0019] Step 2.2, perform clause splitting on the content description information, and then perform punctuation removal, case normalization, word segmentation, stop word removal, and part-of-speech filtering on each sentence.

[0020] According to a specific implementation manner of the embodiment of the present disclosure, step 3 specifically includes:

[0021] Step 3.1, for the sample fraud image I f , sequentially pass through multiple residual modules stacked by ResNet for feature extraction to extract the initial feature f 1 ;

[0022] Step 3.2, perform convolution operations and max-pooling downsampling on the initial feature f 1 through the encoder of the U-Net network, where the encoder includes convolution operations and max-pooling downsampling;

[0023] Step 3.3, perform upsampling operations and convolution operations on the initial feature f 1 through the decoder of the U-Net network, splice the upsampled feature map with the feature map of the corresponding encoding path, and obtain the shallow visual feature f 2 .

[0024] According to a specific implementation manner of an embodiment of the present disclosure, step 4 specifically includes:

[0025] Step 4.1, dividing the shallow visual feature f 2 into multiple non-overlapping feature image window blocks, and using a learnable position encoding strategy to represent the position relationship between feature maps;

[0026] Step 4.2, normalizing the feature map after position encoding for f vpos to obtain the normalized image feature f vnorm ;

[0027] Step 4.3, inputting the normalized image feature f vnorm into the SSM state space module. A part of the feature passes through a linear layer and a SiLi activation function, and another part of the feature passes through a linear layer and a one-dimensional convolutional layer, and then enters the SSM state space module to extract features, which are fused through an activation function, and finally output through a linear layer;

[0028] Step 4.4, fusing local and global information through residual connection, performing conversion and combination between visual feature information, and then performing normalization and multi-layer perceptron for feature fusion;

[0029] Step 4.5, after passing through multiple visual state space modules, combining the learned deep visual features to form a new deep visual feature

[0030] f vision = VSSM N (f vpos ), N = 1, 2…, 8

[0031] wherein, H, W, and C respectively represent the height, width of the image, and the number of channels of the image, and N represents the number of visual state space modules.

[0032] According to a specific implementation manner of an embodiment of the present disclosure, step 5 specifically includes:

[0033] Step 5.1, using a BERT text encoder to extract semantic information corresponding to the content description information to obtain a preliminary semantic feature f s , and then performing position encoding on the semantic feature f s to obtain the encoded semantic feature map f pos ;

[0034] Step 5.2, normalizing the encoded semantic feature map f pos to obtain the normalized semantic feature f norm ;

[0035] Step 5.3, input the normalized semantic feature f norm into the SSM state space module. A part of the feature passes through a linear layer and a SiLi activation function, and another part of the feature passes through a linear layer and a one-dimensional convolutional layer, and then enters the SSM state space module to extract features, which are fused through an activation function, and finally output through a linear layer;

[0036] Step 5.4, fuse local and global information through residual connection, perform transformation and combination between semantic feature information, and then perform normalization and feedforward neural network mechanism to fuse features;

[0037] Step 5.5, after passing through multiple text state space modules, combine the learned semantic features to form a new semantic feature f text ,

[0038] f text = TSSM N (f pos ), N = 1, 2…, 8.

[0039] According to a specific implementation manner of an embodiment of the present disclosure, the step 6 specifically includes:

[0040] Step 6.1, perform global average pooling on the obtained visual feature f vision to generate a global average pooling vector f avg describing the global characteristics of each channel;

[0041] Step 6.2, perform global max pooling on the obtained semantic feature f text to generate another global max pooling vector f max describing the global characteristics of each channel;

[0042] Step 6.3, process the global average pooling vector f avg and the global max pooling vector f max through a shared fully connected layer and convolutional layer to obtain a fused feature f catt ;

[0043] Step 6.4, for the fused feature vector f catt , adopt a cross-attention mechanism to calculate visual features and semantic features, where the visual feature f catt is used as the query, and the global max pooling vector f max is used as the key and value, so as to obtain the score f cross after cross-attention, thereby aggregating the global information of the image and the text token information.

[0044] According to a specific implementation manner of an embodiment of the present disclosure, the step 7 specifically includes:

[0045] Step 7.1, input the sample fraud image I f into the recognition model to obtain the predicted fraud result label pre f , and compare it with the true text label pre t to calculate the classification loss;

[0046] Step 7.2, during the training process of the model, adopt a step size decay strategy. After every fixed number of training times, reduce the learning rate by a fixed ratio. Among them, the reduction ratio of the learning rate is

[0047]

[0048] where η 0 represents the initial learning rate, t represents the decay ratio, and s represents the decay step size;

[0049] Step 7.3, adjust the parameters of the recognition model according to the classification loss until they meet the preset requirements.

[0050] The fraud image recognition solution for artificial intelligence ethics in the embodiments of the present disclosure includes: Step 1, obtain a fraud image dataset, and label the sample fraud images and content description information in the fraud image dataset according to the annotation system; Step 2, perform different preprocessings on the sample fraud images and content description information; Step 3, use the residual module and the U-Net network to extract the shallow visual features of the sample fraud images; Step 4, based on the shallow visual features, use the visual state space module to extract the deep visual features of the sample fraud images; Step 5, based on the content description information, use the text state space module to extract the semantic features of the sample fraud images; Step 6, combine the channel attention mechanism and the cross-attention mechanism to fuse the deep visual features and semantic features to obtain the fused features; Step 7, train a recognition model based on the U-Net and Mamba models in combination with the fused features and the labels corresponding to the sample fraud images; Step 8, input the target fraud image into the trained recognition model to obtain the recognition result.

[0051] The beneficial effects of the embodiments of the present disclosure are as follows: Through the solution of the present disclosure, the SSM state space module can not only deeply extract the visual and semantic features of the image, capture the detailed feature information of the fraud image, but also the computational complexity of the SSM model is linear, effectively reducing the memory occupancy rate; through the channel attention mechanism, it is ensured that the model can flexibly adjust the feature weights in different fraud image scenarios and extract the most recognizable features, thereby improving the overall recognition accuracy of the model; through the cross-attention mechanism, the model can better understand and process the complex information interaction in the fraud image, improve the overall recognition performance, and the model can provide accurate and stable recognition results when facing diverse and complex fraud patterns. Brief Description of the Drawings

[0052] In order to more clearly illustrate the technical solutions of the embodiments of the present disclosure, the accompanying drawings required for use in the embodiments will be briefly introduced below. Obviously, the accompanying drawings in the following description are only some embodiments of the present disclosure. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can be obtained based on these drawings.

[0053] Figure 1 It is a schematic flowchart of a fraud image recognition method for artificial intelligence ethics provided by an embodiment of the present disclosure;

[0054] Figure 2 It is a schematic framework diagram of a visual state space module and a text state space module provided by an embodiment of the present disclosure;

[0055] Figure 3 It is a schematic structural diagram of an SSM state space module provided by an embodiment of the present disclosure. Detailed Description of the Embodiments

[0056] The embodiments of the present disclosure will be described in detail below with reference to the accompanying drawings.

[0057] The following illustrates the embodiments of the present disclosure through specific specific examples. Those skilled in the art can easily understand other advantages and effects of the present disclosure from the content disclosed in this specification. Obviously, the described embodiments are only some embodiments of the present disclosure, rather than all embodiments. The present disclosure can also be implemented or applied through other different specific embodiments. Various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present disclosure. It should be noted that, without conflict, the following embodiments and the features in the embodiments can be combined with each other. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present disclosure without creative efforts belong to the scope of protection of the present disclosure.

[0058] It should be noted that the following describes various aspects of the embodiments within the scope of the appended claims. It should be obvious that the aspects described herein can be embodied in a wide variety of forms, and any specific structure and / or function described herein is illustrative only. Based on the present disclosure, those skilled in the art should understand that one aspect described herein can be implemented independently of any other aspect, and two or more of these aspects can be combined in various ways. For example, any number of aspects described herein can be used to implement the device and / or practice the method. Additionally, this device and / or this method can be implemented using other structures and / or functions in addition to one or more of the aspects described herein.

[0059] It should also be noted that the illustrations provided in the following embodiments only schematically illustrate the basic concept of the present disclosure. The components shown in the illustrations only include those related to the present disclosure, rather than being drawn according to the number, shape, and size of the components in actual implementation. The form, quantity, and proportion of each component in actual implementation can be arbitrarily changed, and the layout form of its components may also be more complex.

[0060] In addition, in the following description, specific details are provided to facilitate a thorough understanding of the examples. However, those skilled in the art will understand that the described aspects can be practiced without these specific details.

[0061] An embodiment of the present disclosure provides a fraud image recognition method for artificial intelligence ethics, which can be applied to the process of fraud image recognition in the Internet security scenario.

[0062] See Figure 1 , which is a schematic flowchart of a fraud image recognition method for artificial intelligence ethics provided by an embodiment of the present disclosure. As Figure 1 shown, the method mainly includes the following steps:

[0063] Step 1, obtain a fraud image dataset, and label the sample fraud images and content description information in the fraud image dataset according to the annotation system.

[0064] Further, step 1 specifically includes:

[0065] Step 1.1, generate various sample fraud images related to fraud content through a generative artificial intelligence model.

[0066] Step 1.2, according to the definition rules in the annotation system, label the sample fraud images with appropriate fraud type labels, and at the same time add a text description corresponding to the image content.

[0067] In specific implementation, obtain a fraud image dataset, and label the fraud images and content description information according to the annotation system. The specific process is as follows:

[0068] A. Generate various image data related to fraud through many generative artificial intelligence models such as ChatGPT, Kimichat, Sora, ERNIE Bot, and Pangu.

[0069] B. According to the definition rules in the annotation system, label the image data with appropriate fraud type labels, and at the same time add appropriate text for the description of the image content.

[0070] C. Obtain a fraud dataset containing supervision information such as image labels and image content descriptions.

[0071] Meanwhile, transfer learning and automatic annotation can also be adopted, and the specific process is as follows:

[0072] 1.1 Generating Samples Using Transfer Learning

[0073] Data collection and preprocessing:

[0074] Collect a preliminary dataset containing fraudulent images. If the existing dataset is insufficient, relevant images can be extracted from public datasets.

[0075] Preprocess the images, such as resizing, normalizing, etc., to meet the input requirements of the transfer learning model.

[0076] Select a pre-trained model:

[0077] Select pre-trained models suitable for image classification and detection, such as YOLO, Faster R-CNN, ResNet, etc. These models have usually been trained on large-scale datasets and can identify and classify various image features.

[0078] Transfer learning:

[0079] Fine-tune the pre-trained model using the existing fraudulent image data. This step will adjust the weights of the model to better adapt to specific types of fraud.

[0080] During the training process, data augmentation techniques (such as rotation, cropping, color adjustment) can be used to expand the dataset and improve the generalization ability of the model.

[0081] Generate samples:

[0082] Use the trained model to predict new images, thereby generating sample images of various types of fraud. These images can be real fraud scenarios or synthetic fraud situations.

[0083] 1.2 Automatic Annotation and Manual Correction

[0084] Automatic annotation:

[0085] Input the generated fraudulent images into the trained model to automatically assign preliminary fraud type labels to each image.

[0086] The model can use different classifiers or detectors to label the main fraud types for the images.

[0087] Manual correction:

[0088] Assign the results of automatic annotation to human annotators for review and correction. The annotators will check and correct the results of automatic annotation according to detailed annotation guidelines to ensure accuracy.

[0089] Add a detailed text description to the image, including specific information about the type of fraud and a description of the image content.

[0090] Step 2, perform different preprocessing on the sample fraud images and content description information;

[0091] Based on the above embodiments, step 2 specifically includes:

[0092] Step 2.1, perform data augmentation operations on each sample fraud image to generate image pairs with different perspectives and angles. Among them, the data augmentation operations include rotation, scaling, flipping, color jittering, and cropping;

[0093] Step 2.2, perform sentence splitting on the content description information, and then perform punctuation removal, case normalization, word segmentation, stop word removal, and part-of-speech filtering on each sentence.

[0094] Specifically in implementation, the process of performing different preprocessing on the sample fraud images and content description information can be shown as the following formula:

[0095] A. For fraud images, by performing various transformations on the original images, generate image pairs with different perspectives and variations to improve the generalization ability of the model;

[0096]

[0097] Where f i represents the original image, Transform represents the data augmentation operation, represents the image after data augmentation.

[0098] Rotation: Rotate the image randomly at an angle between 0 and 360 degrees to simulate perspectives at different angles.

[0099] Scaling: Randomly scale the ratio between 0.6 and 1.5 to generate images of different sizes.

[0100] Flipping: Perform random flipping in the horizontal and vertical directions.

[0101] Cropping: Randomly crop different regions of the image to simulate the loss of partial perspectives.

[0102] Color jittering: Modify color parameters such as brightness, contrast, and saturation of the image to simulate different lighting conditions.

[0103] B. For the content description of fraud images, first perform sentence splitting, and then perform punctuation removal, case normalization, word segmentation, stop word removal, and part-of-speech filtering on each sentence. For word segmentation and part-of-speech filtering, use the jieba word segmentation tool, and for stop word removal, use the Harbin Institute of Technology stop word list.

[0104] Step 3: Use the residual module and U-Net network to extract the shallow visual features of the sample fraud images;

[0105] Further, step 3 specifically includes:

[0106] Step 3.1: For the sample fraud image I f , successively pass through multiple stacked ResNet residual modules for feature extraction to extract the initial feature f 1 ;

[0107] Step 3.2: Perform convolution operations and max-pooling downsampling on the initial feature f 1 through the encoder of the U-Net network, where the encoder includes convolution operations and max-pooling downsampling;

[0108] Step 3.3: Perform upsampling operations and convolution operations on the initial feature f 1 through the decoder of the U-Net network, splice the upsampled feature map with the feature map of the corresponding encoding path to obtain the shallow visual feature f 2 .

[0109] Specifically, the process of using ResNet and U-Net networks to extract the shallow visual features of fraud images may include:

[0110] A. For a given fraud image I f ∈R H×W×C , where H, W, and C respectively represent the height, width, and number of channels of the image, for preliminary processing of features.

[0111] f 0 = Init(I f )

[0112] B. Successively pass through multiple stacked residual modules for feature extraction to extract the shallow features of the image I f , where C 0 is the number of channels of the feature map. The residual module contains multiple convolutional layers, normalization layers, and skip connections. The model directly transmits information through skip connections to alleviate the gradient disappearance problem in deep networks.

[0113] f 1 = ResNet(f 0 )

[0114] C. Through a series of encoder parts composed of convolution operations and max-pooling downsampling, the spatial size of the input image gradually decreases, while the depth of the image feature map gradually increases, for gradually extracting deeper features.

[0115] D. The decoding part mainly consists of a series of upsampling operations and convolutional operations, which are used to gradually restore the spatial dimension and resolution of the image, splice the upsampled feature map with the feature map of the corresponding encoding path, and allow the model to use the features extracted in the encoding stage during the decoding stage, enhancing the detailed feature information of the image.

[0116] f 2 = UNet(f 1 ).

[0117] Optionally, a local feature descriptor can also be used to extract the shallow features in the image by calculating the color histogram or texture features of the local area. The specific process is as follows:

[0118] 1. Local area division

[0119] The image is segmented into multiple local areas (such as windows of a fixed size).

[0120] 2. Calculate the color histogram

[0121] Extract color channels: Extract the channels of the RGB or other color spaces from each area.

[0122] Construct the histogram: Calculate the histogram of each color channel to show the color distribution.

[0123] Normalize: Normalize the histogram.

[0124] 3. Calculate the texture features

[0125] Select texture descriptors: Use methods such as gray-level co-occurrence matrix (GLCM), local binary pattern (LBP), or Gabor filter to extract texture features.

[0126] Calculate features: Generate the texture feature vector of the local area and normalize it.

[0127] 4. Feature aggregation and representation

[0128] Generate the feature vector: Summarize the color and texture features of the local area into the overall feature representation of the image.

[0129] Feature selection: Select the most important features for dimensionality reduction.

[0130] These steps help to extract useful shallow features from the local areas of the image for image classification or recognition tasks.

[0131] Step 4: Based on the shallow visual features, use the visual state space module to extract the deep visual features of the sample fraud image;

[0132] Based on the above embodiments, step 4 specifically includes:

[0133] Step 4.1, segment the shallow visual feature f 2 into multiple non-overlapping feature image window blocks, and use a learnable position encoding strategy to represent the positional relationship between feature maps;

[0134] Step 4.2, normalize the feature map after position encoding for f vpos to obtain the normalized image feature f vnorm ;

[0135] Step 4.3, input the normalized image feature f vnorm into the SSM state space module. Part of the features pass through a linear layer and a SiLi activation function, and another part of the features pass through a linear layer and a one-dimensional convolutional layer, and then enter the SSM state space module to extract features, which are fused through an activation function, and finally output through a linear layer;

[0136] Step 4.4, fuse local and global information through residual connection, perform transformation and combination between visual feature information, and then perform normalization and multi-layer perceptron to fuse features;

[0137] Step 4.5, after passing through multiple visual state space modules, combine the learned deep visual features to form a new deep visual feature

[0138] f vision = VSSM N (f vpos ), N = 1, 2…, 8

[0139] where H, W, and C respectively represent the height, width of the image, and the number of channels of the image, and N represents the number of visual state space modules.

[0140] Specifically, the specific frameworks of the visual state space module and the text state space module are as Figure 2 shown, and the process of using the visual state space module to extract the deep features of the image is as follows:

[0141] A. For the extracted image feature map f 2 , segment it into M non-overlapping image window blocks, and the size of each image window block is n×n.

[0142]

[0143] where n represents the width and height of the image window block, and M represents the number of image window blocks.

[0144] B. The model adopts a learnable positional encoding method to better represent the positional relationship between images. During the training process, the positional encoding parameters are updated according to the gradients calculated by the loss function.

[0145] C. For the feature map after positional encoding f vpos a normalization module is adopted, which can effectively improve the model stability and training efficiency, thus obtaining the normalized image feature f vnorm .

[0146] D. The normalized image feature f vnorm is input into the SSM state space module. Part of the features pass through a linear layer and a SiLi activation function, and another part of the features pass through a linear layer and a one-dimensional convolutional layer, and then enter the SSM state space module for feature extraction. Finally, they are fused through an activation function and then output through a linear layer. Among them, the structure of the state space module is as Figure 3 shown.

[0147] In the state space module, the state equation can be obtained by discretizing the linear equation of the continuous-time system:

[0148] h t = Ah t-1 + Bx t

[0149] y t = Ch t-1 + Dx t

[0150] where h t represents the hidden state at the t-th time step, x t represents the input information, y t represents the output information, A is the state transition matrix, describing the dynamic relationship between states; B is the input matrix, describing how the input x affects the state h; C is the output matrix, describing how the state h is transformed into the output y; D represents the influence of the input x on the output y.

[0151] In the Mamba model, the discretization of the state equation uses the zero-order hold rule, and the transformed state transition matrices A and B can be expressed as

[0152]

[0153] where Δ represents the discretization time step, exp is the exponential function, A and B are the discrete state transition matrices, and I is the identity matrix.

[0154] Through the above state equation and the initial hidden state, the change of the hidden state in the Mamba model is calculated iteratively, and the change of the hidden state can be expressed as:

[0155]

[0156] y k = Ch k + Dx k

[0157]

[0158] where is a structured convolutional kernel, M represents the length of the input sequence, and * represents the convolution operation.

[0159] E. Fuse local and global information through residual connections to perform the conversion and combination between visual feature information; then perform normalization and multi-layer perceptron to perform the fusion between features.

[0160] F. After passing through the visual state space modules of multiple images, the model combines the learned deep visual features to form new features

[0161] f vision = VSSM N (f vpos ), N = 1, 2…, 8.

[0162] Optionally, in addition to using the visual state space module, the deep visual features of the sample fraud image can be further extracted based on other methods from the shallow visual features. For example:

[0163] 1. Convolutional Neural Network (CNN)

[0164] Feature extraction: Use a pre-trained convolutional neural network (such as VGG, ResNet, etc.) to extract the deep features of the image. These networks have been trained on large-scale datasets and can automatically learn effective visual features.

[0165] Transfer learning: Fine-tune the pre-trained model to adapt it to a specific fraud image recognition task.

[0166] 2. Deep Feature Fusion

[0167] Fusion of shallow features and deep features: Fuse shallow visual features (such as color histograms and texture features) with deep features (extracted by CNN) to enhance the feature representation ability of the model.

[0168] Feature fusion method: Shallow and deep features can be combined using concatenation, weighted average, or other fusion methods.

[0169] 3. Autoencoder

[0170] Unsupervised learning: Use an autoencoder for unsupervised learning of images, automatically learning deep features from shallow features. The autoencoder can compress shallow features into low-dimensional deep features through the encoder part and then reconstruct the image through the decoder.

[0171] Feature compression: The output of the learned encoder is used as the deep feature representation of the image.

[0172] Step 5: Based on the content description information, use the text state space module to extract the semantic features of the sample fraud image;

[0173] Further, the specific steps of Step 5 are as follows:

[0174] Step 5.1: Use the BERT text encoder to extract the semantic information corresponding to the content description information to obtain the preliminary semantic feature f s , and then perform positional encoding on the semantic feature f s to obtain the encoded semantic feature map f pos ;

[0175] Step 5.2: Normalize the encoded semantic feature map f pos to obtain the normalized semantic feature f norm ;

[0176] Step 5.3: Input the normalized semantic feature f norm into the SSM state space module. Part of the features pass through a linear layer and a SiLi activation function, and another part of the features pass through a linear layer and a one-dimensional convolutional layer, and then enter the SSM state space module to extract features, which are fused through an activation function and finally output through a linear layer;

[0177] Step 5.4: Fuse local and global information through residual connection, perform conversion and combination between semantic feature information, and then perform normalization and feedforward neural network mechanism to fuse features;

[0178] Step 5.5: After passing through multiple text state space modules, combine the learned semantic features to form a new semantic feature f text ,

[0179] f text = TSSM N (f pos ), N = 1, 2…, 8.

[0180] Specifically, the process of using the text state space module to extract the semantic features of the image is as follows:

[0181] A. For the text description information, extract the semantic features of the image through the BERT text encoder to obtain the preliminary semantic feature fs , and then perform positional encoding on the obtained semantic feature f s to obtain the encoded semantic feature map f pos .

[0182] B. Apply a normalization module to the encoded semantic feature map f pos which can effectively improve the model stability and training efficiency, thereby obtaining the normalized semantic feature f norm .

[0183] C. Input the normalized semantic feature f norm into the SSM state space module. Part of the feature passes through a linear layer and a SiLi activation function, and another part of the feature passes through a linear layer and a one-dimensional convolutional layer, and then enters the SSM state space module for feature extraction. Finally, it is fused through an activation function and then output through a linear layer. Through the above state equation and initial hidden state, the change of the hidden state in the SSM model can be calculated iteratively. The change of the hidden state can be expressed as:

[0184]

[0185] y k =Ch k +Dx k

[0186]

[0187] where is the structured convolution kernel, M represents the length of the input sequence, and * represents the convolution operation.

[0188] D. Fuse local and global information through residual connections to perform the conversion and combination between semantic feature information; then perform normalization and feed-forward neural network mechanisms to fuse features.

[0189] E. After passing through multiple text state space modules, the model combines the learned semantic features to form a new semantic feature f text .

[0190] f text =TSSM N (f pos ), N = 1, 2…, 8.

[0191] Optionally, in addition to using the text state space module, the following method can also be adopted:

[0192] 1. Cross-Modal Learning

[0193] Image-Text Alignment: By aligning images with their descriptive texts, cross-modal models are trained to learn the associations between images and texts. For example, an image-text embedding space is used to map images and descriptions so that semantically related images and texts are closer in the embedding space.

[0194] Bidirectional Network: Use a bidirectional network such as CLIP for joint training of images and texts to learn shared semantic features.

[0195] 2. Text Generation Model

[0196] Description Generation: Use text generation models such as GPT, BERT, etc. to generate detailed descriptions related to images, thereby extracting rich semantic features. These descriptions can be used for further analysis of image content.

[0197] Description Completion: If part of the description of an image is missing, a generation model can be used to complete the description to help obtain more complete semantic information.

[0198] Step 6: Incorporate the channel attention mechanism and the cross-attention mechanism to fuse deep visual features and semantic features to obtain fused features;

[0199] Based on the above embodiments, Step 6 specifically includes:

[0200] Step 6.1: Perform global average pooling on the obtained visual feature f vision to generate a global average pooling vector f avg ;

[0201] Step 6.2: Perform global max pooling on the obtained semantic feature f text to generate another global max pooling vector f max ;

[0202] Step 6.3: Process the global average pooling vector f avg and the global max pooling vector f max through a shared fully connected layer and convolutional layer to obtain a fused feature f catt ;

[0203] Step 6.4: For the fused feature vector f catt , use the cross-attention mechanism to calculate visual features and semantic features, where the visual feature f catt is used as the query, and the global max pooling vector f max is used as the key and value, thereby obtaining the score f cross after cross-attention, thus aggregating between image global information and text token information.

[0204] In specific implementation, the set channel and cross-attention mechanism fuse the visual features and semantic features of the image;

[0205] A. Perform global average pooling on the obtained visual feature f vision to generate a vector f avg ∈R C .

[0206]

[0207] B. Perform global max pooling on the obtained semantic feature f text to generate another vector f max ∈R C

[0208]

[0209] C. Process the global average pooling vector f avg and the global max pooling vector f max through a shared fully connected layer and convolutional layer to obtain a fused feature representation f catt .

[0210] f catt =σ(W 2 ·ReLU(W 1 ·(f avg +f max )))

[0211] D. For the fused feature vector f catt , in order to further fuse the feature information between text and image, the image embedding information is passed to the text tokens, and the visual features and semantic features are calculated through the cross-attention mechanism, where the visual feature f catt is used as the query, and the global max pooling vector f max is used as the key and value, so as to obtain the score f cross after cross-attention, and thus aggregate the global information of the image and the information of the text tokens.

[0212]

[0213] Optionally, in addition to the set channel attention mechanism and the cross-attention mechanism, other methods can also be used to effectively fuse deep visual features and semantic features, such as:

[0214] 1. Feature concatenation and fusion

[0215] Feature concatenation: Concatenate visual features and semantic features at the feature level and then further process them through a fully connected layer or a convolutional layer. After concatenation, the model can learn the representation of the combined features.

[0216] Weighted fusion: Apply a weighting mechanism to visual features and semantic features, calculate the weighted sum to obtain the fused features. The weighting coefficients can be automatically learned during the training process.

[0217] 2. Graph Convolutional Networks (GCN for short)

[0218] Graph embedding: Treat visual features and semantic features as nodes in a graph and use a graph convolutional network to fuse the node information. GCN can handle graph-structured data and is suitable for integrating features from different sources.

[0219] Step 7: Train an identification model based on the U-Net and Mamba models by combining the fused features and the labels corresponding to the sample fraud images.

[0220] Based on the above embodiments, step 7 specifically includes:

[0221] Step 7.1: Input the sample fraud image I f into the identification model to obtain the predicted fraud result label pre f , and compare it with the true text label pre t to calculate the classification loss.

[0222] Step 7.2: During the training process of the model, adopt a step decay strategy. After every fixed number of training times, reduce the learning rate by a fixed ratio, where the reduction ratio of the learning rate is

[0223]

[0224] where η 0 represents the initial learning rate, t represents the decay ratio, and s represents the decay step size;

[0225] Step 7.3: Adjust the parameters of the identification model according to the classification loss until the preset requirements are met.

[0226] Specifically in implementation, the fraud image I f obtains the predicted fraud result label pre f after going through model identification, and compares it with the true text label pre t to calculate the classification loss. This loss function can effectively optimize the parameters of the model network and guide the model to optimize the identification performance.

[0227] f rec = 1 / m ∑(pref -pre t ) 2

[0228] Where m represents the number of fraudulent images.

[0229] During the training process of the model, a step decay strategy is adopted. After every fixed number of training times, the learning rate is reduced by a fixed ratio. This method can smooth the training process and avoid premature convergence and oscillation phenomena.

[0230]

[0231] Where η 0 represents the initial learning rate, t represents the decay ratio, and s represents the decay step.

[0232] Evaluate the image recognition effect of the model on the dataset to ensure its good performance in image recognition. And according to the evaluation results, dynamically adjust the model parameters and training strategies until the recognition model is trained to meet the preset requirements. For example, when the classification loss corresponding to the recognition of the sample fraudulent image by the recognition model is less than the preset threshold, the optimal hyperparameter values can be determined and the optimal model parameters can be saved.

[0233] Step 8, input the target fraudulent image into the trained recognition model to obtain the recognition result.

[0234] In specific implementation, when the trained recognition model is obtained, the unlabeled fraudulent images obtained in real time can be input into it to obtain the recognition result corresponding to the fraudulent image. For example, the label corresponding to the fraudulent image can be output. At the same time, during the use of the trained recognition model, the precision, recall rate, and F1 metric can also be calculated in real time as the accuracy metrics for evaluating the model recognition, and the model can be updated and optimized in real time according to the accuracy metrics.

[0235] The fraudulent image recognition method for artificial intelligence ethics provided in this embodiment extracts the visual feature information and semantic features of the fraudulent image through the visual and text state space modules, effectively enhancing the richness of feature extraction in the image. At the same time, by combining channel attention and cross attention, the model can not only dynamically learn the importance of different feature channels, effectively improve the weight of important features, reduce the influence of redundant information, and perform information interaction between visual features and semantic features, effectively improving the recognition ability of the model in fraudulent scenarios and providing more accurate and stable recognition results.

[0236] The following will further illustrate this method with a specific embodiment.

[0237] Step 1. Obtain a fraudulent image dataset and annotate the fraud types and content descriptions of the fraudulent images according to the label system:

[0238] A. The dataset of fraudulent images includes data in 12 major categories such as forged documents, false invoices, false advertisements, false medical reports, forged payment screenshots, false QR codes, false logistics information, pictures of counterfeit goods, forged brand logos, false recruitment pictures, screenshots of fake news, and forged air tickets;

[0239] B. Appropriately label the fraudulent image data according to the label system, and add appropriate text to the content of the fraudulent images for the description of the fraudulent content.

[0240] Step 2. Perform different preprocessings on the fraudulent images and content descriptions obtained in Step 1:

[0241] A. For fraudulent images, a variety of data augmentation techniques such as rotation, scaling, flipping, color jittering, and cropping are adopted in the data processing stage to improve the diversity and richness of training samples and enhance the generalization ability of the model in different environments.

[0242] Suppose there are two fraudulent images d 1 and d 2 , which are pictures of counterfeit goods and contain logo information of counterfeit goods. The original image d 1 is processed by image augmentation such as rotation, scaling, and flipping to generate d° 1 , while the original image d 1 is processed by a combination of image processing such as rotation, color jittering, and flipping to generate d° 2 . Then, (d 1 , d° 1 ) and (d 1 , d° 2 ) are training samples generated through different data augmentation operations, thus improving the diversity of samples.

[0243] B. For content description information, use the jieba word segmentation tool for word segmentation, cut the text description into a sequence of words, then screen the words in the word sequence according to the Harbin Institute of Technology stop word list, delete and filter out the words in the stop word list, and finally compare the part-of-speech tags marked by the jieba tool during word segmentation with the custom reserved part-of-speech list, and delete and filter out the words whose part-of-speech is not in the reserved part-of-speech list;

[0244] Step 3. Use the ResNet and U-Net networks to extract the shallow visual features of fraudulent images;

[0245] A. For a given fraudulent image I f , it is successively passed through multiple residual modules stacked by ResNet for feature extraction to extract the shallow features f 1 of the image;

[0246] B. Higher-level features are captured through an encoder part consisting of a series of convolutional operations and max-pooling downsampling;

[0247] C. It consists of a series of upsampling operations and convolutional operations, which are used to gradually restore the spatial dimension and resolution of the image. The upsampled feature map is concatenated with the feature map of the corresponding encoding path to obtain deeper features f 2 。

[0248] Step 4. Extract deep features of the image using the visual state space module

[0249] A. For the extracted image feature map f 2 , it is divided into multiple non-overlapping image window blocks, and a learnable position encoding strategy is used to better represent the positional relationship between images;

[0250] B. The feature map after position encoding is input into the normalization module and the state space module for in-depth feature extraction;

[0251] C. After passing through multiple visual state space modules, the model combines the learned deep high-frequency features to form new features f vision ;

[0252] Step 5. Extract semantic features of the image using the text state space module

[0253] A. Download the text encoder model and parameter configuration of the pre-trained BERT model;

[0254] B. Use the model parameter values downloaded in sub-step A to initialize the embedding layer parameters of the text encoding model, and freeze the embedding layer parameters so that they do not change during training. In this way, after the pre-processed text is input into the model, it is transformed into a semantic feature vector f s ;

[0255] C. Perform position encoding on the obtained semantic feature f s to obtain the encoded semantic feature map f pos ;

[0256] D. For the encoded semantic feature map f pos , a normalization module is used to obtain f norm , and then the normalized semantic feature f norm is input into the state space module, and local and global information are fused through residual connections to perform the conversion and combination of semantic feature information; then normalization and feed-forward neural network mechanisms are performed for feature fusion.

[0257] E. After passing through multiple text state space modules, the model combines the learned semantic features to form new semantic features ftext ;

[0258] Step Six. Combine the channel and cross-attention mechanism to fuse the visual features and semantic features of the image

[0259] A. Perform global average pooling on the obtained visual feature f vision to generate a vector f that describes the global characteristics of each channel avg ;

[0260] B. Perform global max pooling on the obtained semantic feature f text to generate another vector f that describes the global characteristics of each channel max ;

[0261] C. Process the global average pooling vector f avg and the global max pooling vector f max through a shared fully connected layer and convolutional layer to obtain a fused feature representation f catt ;

[0262] D. For the fused feature vector f catt , in order to further fuse the feature information between text and image, transfer the image embedding information to the tokens, and calculate the visual features and semantic features through the cross-attention mechanism, so as to obtain the score f after cross-attention cross ;

[0263] Step Seven. Combine the supervision information to train a fraud image recognition model based on the U-Net and Mamba models;

[0264] A. The fraud image recognition model uses the U-Net and Mamba models as the main architectures. After generating the visual feature vectors, a Dropout (random inactivation) layer is added to alleviate the overfitting problem during training, making the final trained model have stronger generalization ability. Since the uneven number of samples for each label will affect the training effect, the focal loss cost function is used to assign different weights to the training samples to alleviate the impact caused by the uneven number of samples for different labels;

[0265] B. To evaluate the recognition ability of the model, calculate the loss of the model through the classification loss, so as to optimize the model parameters and structure;

[0266] C. Use the TensorBoard visualization tool to train the curve when adjusting the hyperparameters, determine the optimal hyperparameter values, and save the optimal model parameters;

[0267] Step Eight. Use the trained recognition model to identify fraud images and obtain the identified labels;

[0268] A. Input the result obtained from the preprocessing of sub-step A in step 2 into the fraud recognition model for model recognition;

[0269] B. Calculate the precision rate, recall rate, and F1 metric as the accuracy metrics for evaluating the model recognition.

[0270] The present invention aims to solve the problem of fraud risks existing in artificial intelligence ethics and provides a fraud image recognition method for artificial intelligence ethics. Through the state space module, the present invention can more deeply understand the semantic and visual features of images. Especially for complex fraud image scenarios, it can identify subtle fraud traces. At the same time, by combining the channel attention and cross-attention mechanisms, it can efficiently perform feature fusion and information interaction, providing more accurate, stable, and efficient recognition results, and providing an effective solution for fraud risk recognition in artificial intelligence ethics.

[0271] The method of the present disclosure performs state equation modeling through the SSM module, which can effectively capture the long-range dependence relationships in images, retain the global and local features of the original image information. At the same time, compared with the traditional Attention mechanism, the calculation process of SSM is linear, which can reduce the memory occupancy of the model and improve the operation efficiency of the model. By combining the channel attention and cross-attention modules for the fusion of visual and semantic features, the model can not only dynamically learn the importance of different feature channels, effectively enhance the weights of useful features, and reduce the influence of redundant information. Moreover, through the cross-attention mechanism, information interaction is established between text and visual features, which helps the visual and semantic features to complement each other and improve the overall recognition performance of the model.

[0272] It should be understood that each part of the present disclosure can be implemented by hardware, software, firmware, or a combination thereof.

[0273] As described above, the above is only the specific implementation manner of the present disclosure, but the protection scope of the present disclosure is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed by the present disclosure should be covered by the protection scope of the present disclosure. Therefore, the protection scope of the present disclosure should be subject to the protection scope of the claims.

Claims

1. A fraud image recognition method for artificial intelligence ethics, characterized in that: include: Step 1: Obtain a fraudulent image dataset, and annotate sample fraudulent images and content description information in the fraudulent image dataset according to the annotation system; Step 2, perform different preprocessing on sample fraudulent images and content description information; Step 3: Use the residual module and U-Net network to extract the shallow visual features of the sample fraud image; Step 4, based on the shallow visual features, the visual state space module is used to extract the deep visual features of the sample fraudulent image; Step 5, based on the content description information, using the text state space module to extract the semantic features of the sample fraud image; Step 6: Combine the channel attention mechanism and the cross attention mechanism to fuse the deep visual features and semantic features to obtain the fused features; Step 7: Combine the fusion features and the labels corresponding to the sample fraudulent images to train a recognition model based on the U-Net and Mamba models; Step 8: Input the target fraud image into the trained recognition model to obtain the recognition result.

2. The method according to claim 1, characterized in that , the step 1 specifically includes: Step 1.1, generating various sample fraud images related to fraud content through a generative artificial intelligence model; Step 1.2: According to the definition rules in the annotation system, the sample fraud image is labeled with the appropriate fraud type label, and a text description corresponding to the image content is added.

3. The method according to claim 2, characterized in that , the step 2 specifically includes: Step 2.1, performing data augmentation operations on each sample fraud image to generate image pairs with different viewing angles and perspectives, wherein the data augmentation operations include rotation, scaling, flipping, color jittering, and cropping; Step 2.2, the content description information is divided into sentences, and each sentence is then subjected to punctuation removal, case normalization, word segmentation, stop word removal, and part-of-speech filtering.

4. The method according to claim 3, characterized in that , the step 3 specifically includes: Step 3.1, for the sample fraud image I f , Feature extraction is performed through multiple ResNet stacked residual modules in sequence to extract the initial feature f1 of the image; Step 3.2, performing convolution operation and maximum pooling downsampling on the initial feature f1 through the encoder of the U-Net network, wherein the encoder includes convolution operation and maximum pooling downsampling; In step 3.3, the initial feature f1 is upsampled and convolved through the decoder of the U-Net network, and the upsampled feature map is concatenated with the feature map of the corresponding encoding path to obtain the shallow visual feature f2.

5. The method according to claim 4, characterized in that , the step 4 specifically includes: Step 4.1, divide the shallow visual feature f2 into multiple non-overlapping feature image window blocks, and use a learnable position encoding strategy to represent the position relationship between feature maps; Step 4.2: f in the feature map after position encoding vpos Normalize and get the normalized image feature f vnorm ; Step 4.3, the normalized image feature f vnorm After being input into the SSM state space module, part of the features pass through the linear layer and SiLi activation function, and the other part of the features pass through the linear layer and one-dimensional convolution layer, and then enter the SSM state space module to extract features, fuse them through the activation function, and finally output them through the linear layer; Step 4.4, local and global information are fused through residual connection, the visual feature information is converted and combined, and then normalization and multi-layer perceptron are performed to fuse the features; Step 4.5, after passing through multiple visual state space modules, the learned deep visual features are combined to form new deep visual features f vision =VSSM N (f vpos ),N=1,2…,8 Among them, H, W and C represent the height, width and number of channels of the image respectively, and N represents the number of visual state space modules.

6. The method according to claim 5, characterized in that , the step 5 specifically includes: Step 5.1: Use the BERT text encoder to extract the semantic information corresponding to the content description information and obtain the preliminary semantic features f s , and then the semantic feature f s Perform position encoding to obtain the encoded semantic feature map f pos ; Step 5.2: Encoded semantic feature map f pos Normalize and get the normalized semantic feature f norm ; Step 5.3: normalize the semantic features f norm After being input into the SSM state space module, part of the features pass through the linear layer and SiLi activation function, and the other part of the features pass through the linear layer and one-dimensional convolution layer, and then enter the SSM state space module to extract features, fuse them through the activation function, and finally output them through the linear layer; Step 5.4, local and global information are fused through residual connection to convert and combine semantic feature information, and then normalization and feedforward neural network mechanism are performed to fuse features; Step 5.5, after passing through multiple text state space modules, the learned semantic features are combined to form a new semantic feature f text , f text =TSSM N (f pos ),N=1,2…,8。 7. The method according to claim 6, characterized in that , the step 6 specifically includes: Step 6.1, obtain the visual features f vision Perform global average pooling to generate a global average pooling vector f that describes the global characteristics of each channel avg ; Step 6.2, obtain the semantic features f text Perform global maximum pooling to generate another global maximum pooling vector f that describes the global characteristics of each channel max ; Step 6.3, the global average pooling vector f avg and the global maximum pooling vector f max Through a shared fully connected layer and convolutional layer, a fusion feature f is obtained. catt ; Step 6.4, for the fused feature vector f catt , a cross attention mechanism is used to calculate visual features and semantic features, where the visual feature f catt As a query, the global maximum pooling vector f max As the key and value, we can get the score f after cross attention cross , thereby performing aggregation between image global information and text token information.

8. The method according to claim 7, characterized in that , the step 7 specifically includes: Step 7.1: Take the sample fraud image I f Input the recognition model and get the predicted fraud result label pre f , and compare it with the real text label pre t Compare and calculate the classification loss; Step 7.2: During the training of the model, a step-size decay strategy is adopted. After every fixed number of trainings, the learning rate is reduced by a fixed ratio. The reduction ratio of the learning rate is Where η0 represents the initial learning rate, t represents the decay rate, and s represents the decay step size; Step 7.3, adjust the parameters of the recognition model according to the classification loss until it meets the preset requirements.

Citation Information

Patent Citations

  • Deep forgery detection method based on artifact noise

    CN115100128A

  • Brain tumor image segmentation method based on multi-scale convolution and Mama structure

    CN118447244A