True and false identification method and system for coastal tourism marketing photo generated by AI

Through deep learning models combined with image processing technology, and using multiple eigenvalues ​​to train convolutional neural network models, the problem of difficulty in identifying the difference between AI generated images and real images in the prior art is solved, and higher recognition accuracy and lower time cost are achieved.

CN120164047AInactive Publication Date: 2025-06-17NANJING UNIV OF INFORMATION SCI & TECH
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510631319.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-16
Publication Date
2025-06-17
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing image recognition methods are low in recognition accuracy and recall when processing complex images, especially high-quality AI-generated images, making it difficult to accurately identify the differences between AI-generated images and real images.

Method used

The deep learning model combined with image processing technology is used to acquire and calculate a variety of feature values ​​of the image, such as brightness, color richness and texture features, and train the VGG16 and ResNet50 convolutional neural network models to identify the authenticity of AI-generated seaside tourism marketing photos.

Benefits of technology

It significantly improves the ability to judge the authenticity of images, improves the accuracy of classification, reduces the time cost of manual review, and makes the judgment results more comprehensive and reliable.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120164047A_ABST
    Figure CN120164047A_ABST
Patent Text Reader

Abstract

The invention discloses a true and false identification method and system for a coastal tourism marketing photo generated by AI, and relates to the technical field of computer vision and artificial intelligence, and the method comprises the steps: obtaining a to-be-identified photo with a first feature; obtaining a first photographing numerical value of the to-be-identified photo with the first feature, inputting the first photographing numerical value of the to-be-identified photo with the first feature into a trained deep learning model, and outputting a corresponding photo type by the deep learning model; if the output of the deep learning model is a first photo type, judging that the photo to be identified with the first feature is an AI generated coastal tourism marketing photo; if the output of the deep learning model is a second photo type, judging that the to-be-identified photo with the first feature is not the AI-generated coastal tourism marketing photo; therefore, the user can clearly know the indexes of the uploaded image and the relative difference between the indexes and the AI generated image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of computer vision and artificial intelligence, and specifically, to a method and system for identifying the authenticity of AI-generated seaside tourism marketing photos. Background Art

[0002] With the rapid development of artificial intelligence technology, technologies such as generative adversarial networks (GANs) have been widely applied in the field of image generation, and the quality of the generated images is getting closer and closer to real pictures, and even difficult to distinguish with the naked eye. AI-generated images have been widely used in fields such as entertainment, advertising, and virtual reality, but they also bring potential security risks, such as the spread of false information and copyright infringement. Therefore, developing an efficient method for determining the authenticity of images, which can quickly identify the differences between AI-generated images and real images, is of great significance for ensuring the authenticity and reliability of information.

[0003] Existing image recognition methods mainly rely on traditional image feature extraction technologies and simple classification algorithms. These methods have low recognition accuracy and recall rate when dealing with complex images, especially high-quality AI-generated images, and are difficult to meet actual needs, resulting in the inability to accurately identify the differences between AI-generated images and real images, and thus unable to make accurate distinctions. Summary of the Invention

[0004] To solve the deficiencies mentioned in the above background art, the purpose of the present invention is to provide a method and system for identifying the authenticity of AI-generated seaside tourism marketing photos.

[0005] In a first aspect, the purpose of the present invention can be achieved through the following technical solutions: A method for identifying the authenticity of AI-generated seaside tourism marketing photos, the method comprising the following steps: Obtain a to-be-identified photo with a first feature; Obtain the first photography value of the to-be-identified photo with the first feature, and input the first photography value of the to-be-identified photo with the first feature into a trained deep learning model, and the deep learning model outputs a corresponding photo type; If the deep learning model outputs a first photo type, then determine that the to-be-identified photo with the first feature is an AI-generated seaside tourism marketing photo; If the deep learning model outputs a second photo type, then determine that the to-be-identified photo with the first feature is not an AI-generated seaside tourism marketing photo; wherein, the training of the deep learning model includes: Obtain the first photography values of the input first photo set and the first photography values of the second photo set; The deep learning model learns the first photography values of the first photo set and the first photography values of the second photo set; Output the photo type corresponding to the received first photography value, where the photo type is the first photo type and the second photo type; The first photography value of the first photo set is the photography value of the first photo set with the first feature. The first photo set is a set of first photos, and the first photos are photos generated based on AI with the first feature; the first photography value of the second photo set is the photography value of the second photo set with the first feature. The second photo set is a set of second photos, and the second photos are photos with the first feature shared by users; the first photography value includes the brightness, color richness, saturation, contrast, dynamic range, edge density, granularity, sharpness, gradient, and complexity value of the first photo; the first feature is a coastal tourism feature.

[0006] Combined with the first aspect, in some implementation manners of the first aspect, the method further includes: The preprocessing process of generating the photo with the first feature based on AI includes: Screen the photos generated based on AI with similar hash values by calculating the hash value of the photos, delete one of the photos generated based on AI with similar hash values, and perform secondary screening on the photos generated based on AI after screening through a preset screening criterion; wherein, the preset screening criterion includes: scene relevance, data purity, content authenticity, and format standardization; Extract features from the photos generated based on AI after secondary screening to obtain the first photography value, where the feature extraction process is based on the travel scene exclusive indicators defined in advance, and the travel scene exclusive indicators include natural landscape index, dynamic light and shadow index, urban structure index, and travel atmosphere score; Finally, assign label values to the processed photos based on the images for label distinction, including photos generated based on AI and photos shared by users.

[0007] Combined with the first aspect, in some implementation manners of the first aspect, the method further includes: The generation process of generating the photo with the first feature based on AI includes: Input keywords into the AI, and the AI generates photos with the first feature according to the keywords, where the keywords are one or more of coastal travel, coastal city, coastal tourism, seaside tourism, beach travel, coastline, coastal scenery, seaside resort, coastal city, and coastal scenery.

[0008] Combined with the first aspect, in some implementation manners of the first aspect, the method further includes: The coastal tourism feature is an image feature in the photo, and the image feature includes a beach feature and an ocean feature.

[0009] In combination with the first aspect, in certain implementations of the first aspect, the method further includes: the deep learning model includes a VGG16 convolutional neural network model and a ResNet50 convolutional neural network model.

[0010] In combination with the first aspect, in certain implementations of the first aspect, the method further includes: the VGG16 convolutional neural network model and the ResNet50 convolutional neural network model are preferentially removed from the top fully connected layer, and the convolutional layer is frozen. A fully connected layer, a Dropout layer, and an output layer are added. The model is compiled using the Adam optimizer and callback functions are set. The feature outputs of the VGG16 convolutional neural network model and the ResNet50 convolutional neural network model after the first photographic values are input are flattened, and the processed feature outputs are feature-connected to obtain a comprehensive feature vector.

[0011] In combination with the first aspect, in certain implementations of the first aspect, the method further includes: the process of the feature outputs of the VGG16 convolutional neural network model and the ResNet50 convolutional neural network model: By inputting the first photographic values into the VGG16 convolutional neural network model for learning to generate first training features, and inputting the first photographic values into the ResNet50 convolutional neural network model for learning to generate second training features; the first training features and the second training features are fused to obtain a comprehensive feature vector as the training features; The deep learning model is trained using the training features, and multiple training rounds are set, and finally the trained deep learning model is obtained.

[0012] In a second aspect, in order to achieve the above object, the present invention discloses a true and false identification system for AI-generated seaside tourism marketing photos, including: A data acquisition module for acquiring the to-be-identified photos with the first feature; A model output module for obtaining the first photographic values of the to-be-identified photos with the first feature, inputting the first photographic values of the to-be-identified photos with the first feature into the trained deep learning model, and the deep learning model outputs the corresponding photo types; If the deep learning model outputs the first photo type, it is determined that the to-be-identified photo with the first feature is an AI-generated seaside tourism marketing photo; If the deep learning model outputs the second photo type, it is determined that the to-be-identified photo with the first feature is not an AI-generated seaside tourism marketing photo; A model training module for training the deep learning model, including: Obtaining the first photographic values of the input first photo set and the first photographic values of the second photo set; The deep learning model learns the first photography values of the first photo set and the first photography values of the second photo set; According to the received first photography value, output the photo type corresponding to the first photography value, where the photo type is the first photo type and the second photo type; The first photography value of the first photo set is the photography value of the first photo set with the first feature. The first photo set is a set of first photos, and the first photos are photos generated based on AI with the first feature; the first photography value of the second photo set is the photography value of the second photo set with the first feature. The second photo set is a set of second photos, and the second photos are photos with the first feature shared by users; the first photography value includes the brightness, color richness, saturation, contrast, dynamic range, edge density, granularity, sharpness, gradient, and complexity value of the first photo; the first feature is a seaside tourism feature.

[0013] In another aspect of the present invention, to achieve the above object, a terminal device is disclosed, including a memory, a processor, and a computer program stored in the memory and capable of running on the processor. The memory stores a computer program capable of running on the processor. When the processor loads and executes the computer program, it adopts a method for identifying the authenticity of AI-generated seaside tourism marketing photos as described above.

[0014] In still another aspect of the present invention, to achieve the above object, a computer-readable storage medium is disclosed. The computer-readable storage medium stores a computer program, and when the computer program is loaded and executed by a processor, it adopts a method for identifying the authenticity of AI-generated seaside tourism marketing photos as described above.

[0015] Advantages of the present invention: By combining deep learning and image processing technologies, the present invention significantly improves the ability to judge the authenticity of images. First, using a pre-trained deep learning model, the system can quickly and accurately identify the differences between AI-generated images and real images. This process not only improves the classification accuracy but also reduces the time cost of manual review. Second, the system integrates the calculation of multiple image features, such as brightness, color richness, and contrast. These features provide rich data support for the judgment, making the judgment results more comprehensive and reliable. Through comparative analysis, users can clearly understand the various indicators of the uploaded images and their relative differences from AI-generated images. Description of the Drawings

[0016] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings; Figure 1 It is a schematic diagram of the method flow of the present invention; Figure 2 It is a schematic diagram of the result performance parameters obtained after the operation of the present invention; Figure 3 It is a schematic diagram of the process of the present invention; Figure 4 It is a schematic diagram of the model structure of the present invention; Figure 5 It is a schematic diagram of the working process of the present invention; Figure 6 It is a schematic diagram of the system structure of the present invention. Specific embodiments

[0017] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.

[0018] Embodiment 1: Such as Figure 1 shown, a method for identifying the authenticity of AI-generated seaside tourism marketing photos, the method includes the following steps: Obtain the photo to be identified with the first feature; Specifically, in the data preparation stage, first, it is necessary to obtain image data, and collect images from two main sources: Real images Obtain real photos related to seaside travel notes from online travel notes. These photos are taken and uploaded by users, with natural lighting effects and real scene features, ensuring the authenticity and representativeness of the data. All collected real images are stored in the.jpg format and saved to the local folder real_photo.

[0019] AI-generated images Using the scene features summarized from real photos as keywords, use the drawing function of the AI large model to generate images. The generated image keywords include but are not limited to the following: Sunset coastline, sunset glow coastal scenery, sunrise seaside activities, seaside sunset afterglow, beach sunset afterglow, seaside sunrise impression, sunset beach relaxation, swaying coconut groves, coastal promenade, blue sea and blue sky, coral reef groups, blend of history and nature, underwater world, swimming with sharks, snorkeling adventure, dinner under the stars; These keywords are designed to cover the scene features of real photos to ensure that the AI photo dataset has similar scene features to the real photo dataset, making the training effect of the training model better and the recognition accuracy higher. The generated images will be saved in the.jpg format to the local folder ai_photo. These images will be used for subsequent model training to ensure that the model can recognize AI-generated images in multiple scenarios.

[0020] Through the above method, two folders are formed to store AI-generated images and real images respectively, providing data support for subsequent image preprocessing and classification tasks.

[0021] Obtain the first photographic value of the photo to be recognized for the first feature, and input the first photographic value of the photo to be recognized for the first feature into the trained deep learning model. The deep learning model outputs the corresponding photo type; If the deep learning model outputs the first photo type, it is determined that the photo to be recognized with the first feature is an AI-generated seaside tourism marketing photo; If the deep learning model outputs the second photo type, it is determined that the photo to be recognized with the first feature is not an AI-generated seaside tourism marketing photo; wherein, the training of the deep learning model includes: Obtain the first photographic value of the first photo set and the first photographic value of the second photo set input; The deep learning model learns the first photographic value of the first photo set and the first photographic value of the second photo set; According to the received first photographic value, output the photo type corresponding to the first photographic value, wherein the photo type is the first photo type and the second photo type; The first photographic value of the first photo set is the photographic value possessed by the first photo set with the first feature. The first photo set is a collection of first photos, and the first photos are photos generated by AI with the first feature; the first photographic value of the second photo set is the photographic value possessed by the second photo set with the first feature. The second photo set is a collection of second photos, and the second photos are photos with the first feature shared by users; the first photographic value includes the brightness, color richness, saturation, contrast, dynamic range, edge density, granularity, sharpness, gradient, and complexity value of the first photo; the first feature is a seaside tourism feature.

[0022] Among them, within this embodiment, the photo quality of travel photos mostly depends on their visual features, which play a significant role in enhancing the destination image. The visual features mainly include texture features and color features. In terms of photo color, brighter and more saturated photos tend to arouse stronger participation, and specific colors such as orange and yellow show strong positive effects; in terms of texture, a dry texture evokes feelings of gratitude and sympathy, while an uneven texture causes surprise and excitement. Therefore, the visual features of the photo play a decisive role in photo judgment. Through the calculation of photographic numerical values for the photo dataset, photographic numerical parameters of colors (brightness, color richness, saturation, contrast, dynamic range) and textures (edge density, granularity, sharpness, gradient, complexity value) with a large difference between the average values of AI photos and real photos are selected.

[0023] The preprocessing process for generating photos with the first feature based on AI includes: Screen the AI-generated photos with the first feature that have similar hash values by calculating the hash values of the photos, delete one of the AI-generated photos with the first feature that have similar hash values, and perform a secondary screening on the screened AI-generated photos with the first feature through a preset screening criterion; among them, the preset screening criterion includes: scene relevance, data purity, content authenticity, and format standardization; Extract features from the AI-generated photos with the first feature after the secondary screening to obtain the first photographic numerical values. Among them, the feature extraction process is based on the travel scene exclusive indicators defined in advance, and the travel scene exclusive indicators include natural landscape index, dynamic light and shadow index, urban structure index, and travel atmosphere score; Finally, based on the image, assign label values to distinguish the processed photos, including AI-generated photos and user-shared photos.

[0024] Specifically, it is necessary to perform preprocessing operations on the collected images, which specifically include the following steps: Deletion of similar AI photos Since the AI model may generate two almost identical photos when continuously generating photos, screen the photos by calculating the hash values of the photos, delete the "duplicate" photos among them, achieve the effect of sample denoising, and avoid affecting the model accuracy.

[0025] In the formula: is the output hash value; is the input data; is the th hash sub-function for processing the input data; is the number of hash sub-functions; Ensure that the hash value is within a specific range.

[0026] Resizing Crop all images to a square of 655×655, and then resize them to a fixed size of 64×64 pixels to ensure that the image sizes input into the convolutional neural network are unified and regular images. Use the cv2.resize function in Python to implement the image scaling operation.

[0027] In the formula: is the original image; are the target width and height of the resized image; is the function used to change the image size; is the interpolation method, which determines how the new pixel values are calculated; The output result of the resized image.

[0028] Color space conversion Images are generally stored in the BGR format, while convolutional neural networks usually use the RGB format. Therefore, it is necessary to convert the images from the BGR format to the RGB format, and the cv2.cvtColor function is used to complete the conversion operation.

[0029] Array format conversion Convert the images into array form for subsequent deep learning processing. Each image is represented as a three-dimensional array with a shape of (64, 64, 3), where each dimension represents the pixel height, width, and number of color channels respectively.

[0030] Calculation of photographic values and photographic indicators for travel scenes The colors and textures in travel photos are very rich, and there are significant differences in various photographic values between AI images and real images. Calculating the photographic values of the images and the photographic indicators for travel scenes helps to improve the model performance. Among them, the color features are as follows: Brightness: Brightness quantifies the overall brightness level of the image and is usually represented by the average grayscale value of all pixels.

[0031] Color richness: Color richness measures the saturation of the colors in the image and is calculated through the differences in the red-green and yellow-blue channels in the RGB color matrix.

[0032] In the formula: and They respectively represent the standard deviations of the red-green channel and the yellow-blue channel in the RGB color matrix.

[0033] Saturation: Saturation refers to the intensity of colors in an image.

[0034] In the formula: and are respectively the maximum and minimum values of R, G, and B at the pixel Contrast: Contrast measures the brightness difference between the brightest and darkest parts of an image, usually calculated using Michelson contrast.

[0035] Dynamic range: Dynamic range represents the ratio between the brightest and darkest parts of an image.

[0036] The texture features are as follows: Edge density: Edge density reflects the richness of the texture of an image, that is, the ratio of the edge pixels detected by the Canny algorithm to the total number of pixels.

[0037] In the formula: is 1 if it represents an edge pixel, otherwise it is 0.

[0038] Granularity: Granularity measures the roughness or heterogeneity of the texture of an image.

[0039] In the formula: is the average value of the grayscale matrix G.

[0040] Sharpness: Sharpness measures the clarity of edges and details in an image.

[0041] In the formula: is the Laplacian of the grayscale matrix, is the average value of the Laplacian.

[0042] Gradient: Gradient represents the severity of the brightness change in an image.

[0043] In the formula: and are respectively the horizontal and vertical gradients at the pixel

[0044] Complexity: Complexity measures the average level of local texture changes in an image.

[0045] Where: is the pixel gray value of the Kth neighbor of, if , then , if , then .

[0046] The travel scene photography metrics are as follows: Natural landscape index: Combining the characteristics of high color richness and low graininess ( ), quantifying the visual attractiveness of natural scenery.

[0047] Dynamic light and shadow index: Combining the product of dynamic range and contrast ( ), capturing typical lighting scenes during travel such as sunrise and sunset.

[0048] Urban structure index: Through the geometric mean of edge density and gradient ( ), enhancing the model's ability to recognize building outlines.

[0049] Travel atmosphere score: Balancing the synergistic effect of brightness and color richness ( ), reflecting the overall atmosphere of the photo.

[0050] Construct the photography parameter sub-network This sub-network consists of two fully connected layers and an attention mechanism. The input is 14 standardized photography parameters (including 4 travel scene combination parameters). It learns the physical relationships between the parameters through non-linear mapping and outputs a feature vector, which is concatenated with the dual-path CNN features and then input into the classification layer.

[0051] Sub-network output Feature calculation is as follows: Where: is the weight matrix, is the bias term Label generation For the convenience of training the classification task, a label needs to be assigned to each image: The images generated by AI are assigned a label value of 1 and stored in the corresponding label array.

[0052] The real images are assigned a label value of 0 and also stored in the label array.

[0053] Finally, the image data and its corresponding label array form a complete training dataset, providing a basis for subsequent data partitioning and model training.

[0054] The generation process of generating a photo with the first feature based on AI includes: Input keywords into the AI, and the AI generates a photo with the first feature according to the keywords. Among them, the keywords are one or more of seaside travel, coastal city, coastal tourism, seaside tourism, beach travel, coastline, coastal scenery, seaside resort, coastal city, and seaside scenery.

[0055] The seaside tourism feature is an image feature in the photo, and the image feature includes a beach feature and an ocean feature.

[0056] The deep learning model includes a VGG16 convolutional neural network model and a ResNet50 convolutional neural network model.

[0057] The VGG16 convolutional neural network model and the ResNet50 convolutional neural network model are preferably removed the top fully connected layer, and the convolutional layer is frozen. A fully connected layer, a Dropout layer and an output layer are added. The model is compiled using the Adam optimizer and the callback function is set. And the feature outputs of the VGG16 convolutional neural network model and the ResNet50 convolutional neural network model after being input with the first photographic value are flattened respectively, and the processed feature outputs are feature-connected to obtain a comprehensive feature vector.

[0058] Specifically, in the image dataset division stage: Sample quantity check Before data division, first check whether the number of loaded image samples is sufficient to avoid poor model training effect caused by insufficient data. Use the len function of Python to output the total number of loaded samples to ensure that the dataset size meets the requirements of the classification task.

[0059] In the formula: s is a string or sequence; is the length of the string or sequence s; is the function for calculating the length of the string or sequence.

[0060] Divide the training set, validation set and test set Use the train_test_split function of Python to divide the image data and the label array into a training set, a validation set and a test set, and the division ratio is as follows: Divide the image data into a training set (80%) and a test set (20%).

[0061] Further divide the training set into a training set (60%) and a validation set (20%).

[0062] Where: N is the total number of samples; p is the proportion of the test set; is the number of training set samples, and the calculation formula is ; is the number of test set samples, and the calculation formula is ; D is the data set, and L is the label; is the training set data and label; is the test set data and label is the operation of dividing the data set and label into training set and test set according to the proportion.

[0063] Ensure the balance of samples during division, and each category (AI-generated images and real images) is evenly distributed in different data sets to avoid data bias affecting the model performance.

[0064] Model construction and training stage Load the base model Load the pre-trained VGG16 and ResNet50 convolutional neural network models. The specific operations are as follows: Remove the fully connected layers at the top of the VGG16 and ResNet50 models to avoid the classification results of the original model affecting the training of the new task.

[0065] Set the input shape to (64, 64, 3) to adapt to the preprocessed image data.

[0066] By loading the pre-trained model, the general features learned on the ImageNet data set can be directly utilized, thereby reducing the training time and improving the model performance.

[0067] Model construction and feature merging Based on the pre-trained model, add new layers to complete the binary classification task: Freeze the convolutional layers of VGG16 and ResNet50 to keep their pre-trained features unchanged.

[0068] Flatten the feature outputs of VGG16 and ResNet50 to prepare for subsequent connection operations.

[0069] Connect the two flattened image feature vectors to form a comprehensive feature vector. And connect the image feature vector with the photographic numerical sub-network.

[0070] Add a fully connected layer, a Dropout layer (to prevent overfitting), and an output layer. The output layer uses the Sigmoid activation function to achieve binary classification.

[0071] Where: z is the input value or linear combination; e is the base of the natural logarithm; y is the output value, representing the result after passing through the sigmoid function; is the sigmoid function, and its calculation formula is .

[0072] Model Compilation Compile the model using the Adam optimizer, set the loss function to Binary Cross-Entropy, and monitor the accuracy of the model. The learning rate of the Adam optimizer can be adjusted to achieve faster convergence.

[0073] Where: is the number of samples; is the true label (0 or 1); is the predicted probability (the probability value output by the model).

[0074] Callback Function Settings To prevent overfitting and save the best model, set the following callback functions: Monitor the Validation Loss, Validation Accuracy, and Validation Precision. When the Validation Loss does not decrease for 5 consecutive epochs, stop training and restore to the best weights.

[0075] Where: is the loss function, representing the difference between the true value and the predicted value ; N is the number of samples; is the true label (0 or 1) of the i-th sample; is the predicted probability of the i-th sample; log is the natural logarithm function.

[0076] ModelCheckpoint: Save the model with the minimum Validation Loss to ensure that the best model is retained.

[0077] Train the Model The model is trained using the training set and evaluated on the validation set. The maximum number of training epochs is set to 50. During training, the learning rate is dynamically adjusted according to the performance on the validation set to optimize the training speed and effect. In addition, KFold cross-validation is used to divide the dataset into ten folds to ensure that each sample is fully utilized in training and validation. Each time training is performed, the model is evaluated on different training and validation sets, thereby effectively reducing the risk of overfitting and improving the generalization ability. By calculating the validation accuracy of each fold, the final results are aggregated to obtain the average performance of the model on the overall dataset, providing a more comprehensive basis for robustness evaluation.

[0078] Where: The average performance metric for all K iterations, reflecting the overall performance of the model on the entire dataset; K represents the dataset being divided into K subsets; Indicates that in the i-th iteration, the model is on the validation set The performance metric.

[0079] Model Evaluation Model Performance Evaluation The performance of the model is evaluated on the test set, and the test loss and test accuracy are output as quantitative indicators of the overall performance of the model.

[0080] Model Prediction The trained model is used to perform image classification prediction on the test set to generate prediction labels.

[0081] Performance Metric Calculation Precision and Recall are calculated to comprehensively evaluate the performance of the model on different classes and ensure balanced classification results.

[0082] Where: TP (True Positive) is the correctly identified normal photo; FP (False Positive) is the incorrectly identified normal photo; FN (False Negative) is the incorrectly identified other photo.

[0083] Result Visualization Phase Visualization of the Training Process Use Matplotlib to plot the following graphs: Accuracy curves for training and validation: Observe the learning effect of the model and its performance during the training process.

[0084] Loss curves for training and validation: Understand whether the model has overfitting or underfitting phenomena.

[0085] Bar chart of precision and recall: Intuitively show the model performance.

[0086] The result performance parameters obtained after running are as Figure 2 shown; Comparison of accuracy, precision, and return rate in Table 1 Specifically, the solution of the present invention will be further elaborated through embodiments: the construction stage of running Environment setup Set up a Python development environment and install libraries such as Flask, Keras, OpenCV, and Scikit-image. Through these libraries, image processing, model loading, and web application development are realized.

[0087] File upload module Create a file upload module and use the routing function of the Flask framework to handle the images uploaded by users. Ensure that the uploaded file format is an image, and set the file save path to static / uploads. To prevent file name conflicts, use the secure_filename function to safely process the file name.

[0088] Image processing module In the image processing module, define multiple image feature extraction functions. These functions are implemented using the OpenCV and Scikit-image libraries. By reading the uploaded images, calculate and return the corresponding feature values.

[0089] Model loading Load a pre-trained deep learning model. Use the load_model function provided by Keras to load the model stored as a.keras file into memory. This model has been trained and can effectively distinguish between AI-generated images and real images.

[0090] Result display module Build a result display module and use the rendering function of Flask to pass the calculated image features, model judgment results, and their probabilities to the front-end HTML template. On the result page, display the uploaded images, analysis results, and corresponding statistical data.

[0091] Run startup Start the web application through the app.run method of the Flask framework, and set the debug mode for easy development and testing. Users can access the system through a browser to upload and analyze images.

[0092] Such as Figure 2As shown in the figure, in order to test the performance and stability of the model, the model was continuously run ten times, and its accuracy, precision, and recall data were recorded to generate a visualized bar chart.

[0093] As Figure 3 shown in the figure, in order to build this deep learning model, the following steps were specifically implemented. First, a large number of real photos were collected from online travel notes, and an equal number of AI photos were generated using an AI large model with their scene features as keywords to form a photo dataset; then, data preprocessing work was carried out on the dataset, including photo screening, cropping, etc.; then it was divided into a test set and a training set, and the photographic numerical values of the photos were extracted; in the training set, the VGG16 and RestNet50 models were used to extract features from the photos in a two-way manner, then the features were flattened, and the flattened features were merged with the photographic numerical features to form merged features and input into the test model, and the model was tested in the test set to obtain the output results.

[0094] As Figure 4 shown in the figure, in order to accelerate the model training speed, the present invention divides the preprocessed large number of pictures (size 64×64 pixels) into two steps of calculating photographic numerical values and image feature extraction. When calculating the photographic numerical values, the image first enters the 32RELU layer (32 RELU neurons) for processing, and after dimensionality reduction, it enters the 16RELU layer (16 RELU neurons) to calculate the photographic numerical features of the photo; when extracting image features, a two-way extraction strategy of VGG16 and RestNet50 is adopted. In the VGG16 model, the input image enters the fully connected layer after cyclic iteration through the convolutional layer and pooling layer, and then the output features are flattened. In the RestNet50 model, the input image enters the fully connected layer after cyclic iteration through the convolutional layer, regularization, and non-linear function (RELU), and then the output features are flattened. The flattened features in the two models are collected, merged, and the merged features are connected with the photographic numerical parameters and then enter the fully connected layer, and then enter the dropout layer for processing, and finally the output features are sent into the model, and the model outputs the photo authenticity result.

[0095] As Figure 5As shown in the figure, in order to more conveniently and quickly identify the authenticity of tourist photos, the present invention proposes a photo authenticity discrimination system, which mainly consists of five parts: a user interaction interface, a Flask application server, an AI analysis engine, a visualization rendering layer, and a data storage layer. The user interaction interface mainly includes a browser client part, which can upload photos (index.html) and view detection results (result.html); through this interface, HTTP requests or responses are used to connect to the Flask application server. The server mainly consists of a routing controller (rendering index.html and file processing or result generation) and a file processor (photo upload verification and static resource management); the AI analysis engine performs predictive analysis on the content uploaded by the user to the Flask application server, mainly through a hybrid deep learning model (loading the model and generating prediction probabilities) and image metric analysis (feature parameter calculation and data standardization) to complete; in the visualization rendering layer, the Matplotlib engine is started through model loading for Base64 encoding conversion and generating box plots, and dynamic template rendering, result module assembly, and page optimization are performed through analysis requests; the data storage layer mainly consists of an uploaded photo repository, and the photos uploaded by the user will be saved in the static / uploads folder.

[0096] Specifically, the ability to manually identify the authenticity of tourist photos depends on the experience and ability of individual users. In the experiment designed to manually identify tourist photos, the experimental situation shown in Table 2 below was obtained: Table 2 Comparison of Experimenter Situations The average accuracy of the experimenters was only 67.7%, the precision was only 76.7%, and the recall rate was only 56.8%. Moreover, some participants tended to make more conservative judgments, while others tended to think that most photos were generated by AI.

[0097] In terms of constructing a deep learning model, except for the single-layer CNN model, it mainly includes multi-layer CNN, VGG, ResNet50, VGG+ResNet50, VGG+ResNet50 (with image color and texture features). The final results can be obtained as shown in Table 3 below: Table 3 Final Comparison Results of the Experiment The results show that the addition of explicit color and texture features significantly improves the classification performance of the model. The accuracy rate of only using the multi-layer CNN model is 84.4%, while the accuracy rate of using the VGG ResNet50 model is 89.2%, which is greater than the two used alone. All models are significantly better than the average level of human judgment, which is 67.7%. When color and texture features are combined with VGGResNet50, the model accuracy rate is further increased to 91.5%, and the precision reaches 92.8%. Although the return rate decreases slightly, it performs better than the VGG ResNet50 model in operations that require high precision.

[0098] Example 2: As Figure 6 shown, to achieve the above object, the present invention discloses a true and false recognition system for AI-generated seaside tourism marketing photos, including: A data acquisition module for acquiring the to-be-recognized photos with the first feature; A model output module for obtaining the first photographic value of the to-be-recognized photos with the first feature, inputting the first photographic value of the to-be-recognized photos with the first feature into the trained deep learning model, and the deep learning model outputs the corresponding photo type; If the deep learning model outputs the first photo type, it is determined that the to-be-recognized photos with the first feature are AI-generated seaside tourism marketing photos; If the deep learning model outputs the second photo type, it is determined that the to-be-recognized photos with the first feature are not AI-generated seaside tourism marketing photos; A model training module for training the deep learning model, including: Obtaining the first photographic values of the input first photo set and the first photographic values of the second photo set; The deep learning model learns the first photographic values of the first photo set and the first photographic values of the second photo set; According to the received first photographic value, output the photo type corresponding to the first photographic value, where the photo type is the first photo type and the second photo type; The first photographic values of the first photo set are the photographic values possessed by the first photo set with the first feature. The first photo set is a set of first photos, and the first photos are photos generated based on AI with the first feature; the first photographic values of the second photo set are the photographic values possessed by the second photo set with the first feature. The second photo set is a set of second photos, and the second photos are photos shared by users with the first feature; the first photographic values include the brightness, color richness, saturation, contrast, dynamic range, edge density, granularity, sharpness, gradient, and complexity value of the first photo; the first feature is the seaside tourism feature.

[0099] Based on the same inventive concept, the present invention further provides a computer device, which includes: one or more processors, and a memory for storing one or more computer programs; the program includes program instructions, and the processor is configured to execute the program instructions stored in the memory. The processor may be a Central Processing Unit (CPU), or may also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. It is the computing core and control core of the terminal, and is used to implement one or more instructions. Specifically, it is used to load and execute one or more instructions in the computer storage medium to implement the above method.

[0100] It should be further noted that, based on the same inventive concept, the present invention further provides a computer storage medium, on which a computer program is stored, and the computer program, when run by a processor, executes the above method. The storage medium may adopt any combination of one or more computer-readable media. The computer-readable medium may be a computer-readable signal medium or a computer-readable storage medium. The computer-readable storage medium may, for example, but not be limited to, an electrical, magnetic, optical, electrical, magnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (non-exhaustive list) of the computer-readable storage medium include: an electrical connection having one or more wires, a portable computer disk, a hard disk, a Random Access Memory (RAM), a Read Only Memory (ROM), an Erasable Programmable Read Only Memory (EPROM or flash memory), an optical fiber, a portable compact disk read only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present invention, the computer-readable storage medium may be any tangible medium that contains or stores a program, and the program may be used by or combined with an instruction execution system, apparatus, or device.

[0101] In the description of this specification, the descriptions referring to terms such as "one embodiment", "example", "specific example", etc. mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present disclosure. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described may be combined in any one or more embodiments or examples in a suitable manner.

[0102] The foregoing has shown and described the basic principles, main features and advantages of the present disclosure. Those skilled in the art should understand that the present disclosure is not limited by the above embodiments, and what is described in the above embodiments and the specification is only to illustrate the principles of the present disclosure. Without departing from the spirit and scope of the present disclosure, the present disclosure will have various changes and improvements, and these changes and improvements fall within the scope of the present disclosure claimed.

Claims

1. A method for identifying authenticity of AI-generated seaside tourism marketing photos, characterized in that: The method comprises the following steps: Obtaining a photo to be identified having a first feature; Obtaining a first photographic value of the photo to be identified with a first feature, inputting the first photographic value of the photo to be identified with the first feature into a trained deep learning model, and the deep learning model outputting a corresponding photo type; If the deep learning model outputs the first photo type, it is determined that the photo to be identified with the first feature is an AI-generated seaside tourism marketing photo; If the deep learning model outputs the second photo type, it is determined that the photo to be identified with the first feature is not an AI-generated seaside tourism marketing photo; wherein the training of the deep learning model includes: Obtaining a first photographic value of a first photo set and a first photographic value of a second photo set as input; The deep learning model learns the first photographic value of the first photo set and the first photographic value of the second photo set; Outputting a photo type corresponding to the first photographic value according to the received first photographic value, wherein the photo type is a first photo type and a second photo type; The first photographic value of the first photo set is a photographic value of the first photo set with the first feature, the first photo set is a collection of first photos, and the first photos are photos with the first feature generated based on AI; the first photographic value of the second photo set is a photographic value of the second photo set with the first feature, the second photo set is a collection of second photos, and the second photos are photos with the first feature shared by users; the first photographic value includes brightness, color richness, saturation, contrast, dynamic range, edge density, granularity, clarity, gradient, and complexity values ​​of the first photos; the first feature is a seaside tourism feature.

2. According to claim 1, a method for identifying authenticity of AI-generated seaside tourism marketing photos is characterized in that: The preprocessing process of generating a photo with the first feature based on AI includes: Screening AI-generated photos with similar hash values ​​and having the first feature by calculating the hash values ​​of the photos, deleting one of the AI-generated photos with similar hash values ​​and having the first feature, and performing a secondary screening on the screened AI-generated photos with the first feature by using a preset screening standard; wherein the preset screening standard includes: scene relevance, data purity, content authenticity, and format standardization; Performing feature extraction on the photos with the first feature generated based on AI after the secondary screening to obtain a first photographic value, wherein the feature extraction process is performed based on preset travel scene-specific indicators, and the travel scene-specific indicators include a natural landscape index, a dynamic light and shadow index, an urban structure index, and a travel atmosphere score; Finally, the processed photos are labeled based on the label values ​​assigned to the images, including photos generated by AI and photos shared by users.

3. The method for identifying authenticity of AI-generated seaside tourism marketing photos according to claim 2, characterized in that: The AI-based generation process of the photo with the first feature includes: Input keywords into the AI, and the AI ​​generates photos with the first feature based on the keywords, wherein the keywords are one or more of seaside travel, coastal city, coastal tourism, seaside tourism, beach travel, coastline, coastal scenery, seaside resort, coastal city, and coastal scenery.

4. The method for identifying authenticity of AI-generated seaside tourism marketing photos according to claim 1, characterized in that: The seaside tourism features are image features in the photo, and the image features include beach features and ocean features.

5. The method for identifying authenticity of AI-generated seaside tourism marketing photos according to claim 1, characterized in that: The deep learning models include a VGG16 convolutional neural network model and a ResNet50 convolutional neural network model.

6. The method for identifying authenticity of AI-generated seaside tourism marketing photos according to claim 5, characterized in that: The VGG16 convolutional neural network model and the ResNet50 convolutional neural network model are preferentially removed from the top fully connected layer, and the convolution layer is frozen, and a fully connected layer, a Dropout layer and an output layer are added. The model is compiled using the Adam optimizer and a callback function is set. The feature outputs of the VGG16 convolutional neural network model and the ResNet50 convolutional neural network model after the first photographic value is input are flattened, and the processed feature outputs are feature connected to obtain a comprehensive feature vector.

7. The method for identifying authenticity of AI-generated seaside tourism marketing photos according to claim 6, characterized in that: The feature output process of the VGG16 convolutional neural network model and the ResNet50 convolutional neural network model: The first photographic value is input into a VGG16 convolutional neural network model for learning to generate a first training feature, and the first photographic value is input into a ResNet50 convolutional neural network model for learning to generate a second training feature; The first training feature and the second training feature are combined to obtain a comprehensive feature vector as a training feature; By using the training features to train the deep learning model and setting multiple training rounds, a trained deep learning model is finally obtained.

8. A system for identifying the authenticity of AI-generated seaside tourism marketing photos, which uses a method for identifying the authenticity of AI-generated seaside tourism marketing photos as claimed in any one of claims 1 to 7, characterized in that: include: A data acquisition module, used to acquire a photo to be identified having a first feature; A model output module, used to obtain a first photographic value of the photo to be identified of the first feature, input the first photographic value of the photo to be identified of the first feature into a trained deep learning model, and the deep learning model outputs a corresponding photo type; If the deep learning model outputs the first photo type, it is determined that the photo to be identified with the first feature is an AI-generated seaside tourism marketing photo; If the deep learning model outputs the second photo type, it is determined that the photo to be identified with the first feature is not an AI-generated seaside tourism marketing photo; The model training module is used to train the deep learning model, including: Obtaining a first photographic value of a first photo set and a first photographic value of a second photo set as input; The deep learning model learns the first photographic value of the first photo set and the first photographic value of the second photo set; Outputting a photo type corresponding to the first photographic value according to the received first photographic value, wherein the photo type is a first photo type and a second photo type; The first photographic value of the first photo set is a photographic value of the first photo set with the first feature, the first photo set is a collection of first photos, and the first photos are photos with the first feature generated based on AI; the first photographic value of the second photo set is a photographic value of the second photo set with the first feature, the second photo set is a collection of second photos, and the second photos are photos with the first feature shared by users; the first photographic value includes brightness, color richness, saturation, contrast, dynamic range, edge density, granularity, clarity, gradient, and complexity values ​​of the first photos; the first feature is a seaside tourism feature.

9. A terminal device, comprising a memory, a processor, and a computer program stored in the memory and capable of running on the processor, characterized in that: The memory stores a computer program that can be run on the processor. When the processor loads and executes the computer program, the method for identifying authenticity of AI-generated seaside tourism marketing photos described in any one of claims 1 to 7 is adopted.

10. A computer-readable storage medium having a computer program stored therein, characterized in that: When the computer program is loaded and executed by the processor, the method for identifying authenticity of AI-generated seaside tourism marketing photos as described in any one of claims 1 to 7 is adopted.

Citation Information

Patent Citations

  • False news image detection method based on multi-modal data fusion neural network

    CN114612679A