Method and device for identifying digital content picture with specific style

By building an extended data set and using the MCA-DenseNet model, combining dense connection layers and attention mechanisms, the accuracy and efficiency problems of traditional methods when identifying new digital content-specific style pictures are solved, and accurate and robust recognition of new digital content-specific style pictures are achieved.

CN120336610APending Publication Date: 2025-07-18ZHEJIANG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510304534.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-14
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

When traditional risk image recognition methods identify new types of digital content-specific style pictures, especially 2D and cartoon pixel abstract art pictures, they cannot accurately identify risk elements such as political sensitivity, pornography or violence, resulting in false alarms or missed reports, affecting the accuracy and effectiveness of content supervision.

Method used

New digital content samples are obtained through manual screening and crawling techniques, extended data sets are built, and MCA-DenseNet model is trained, combined with dense connection layer, transition layer and attention mechanism, global average pooling is performed, and model parameters are optimized to improve identification accuracy and robustness.

Benefits of technology

It realizes accurate identification of new digital content-specific pictures, improves computing efficiency and robustness, can handle noise and new data, adapt to new situations, and improves the accuracy and robustness of identifying risk content such as political sensitivity, pornography, violence, etc.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120336610A_ABST
    Figure CN120336610A_ABST
Patent Text Reader

Abstract

The invention discloses a method and a device for identifying a digital content specific style picture, and the method comprises the steps: obtaining a novel digital content picture through manual screening and sorting, automatically obtaining a style migration technology extension data set through a crawler technology, and constructing a data set containing risk content and risk-free content; sequentially inputting the image data into a trunk layer, a first dense connection layer, a transition layer and a plurality of other dense connection layers, executing global average pooling, and converting a feature map into a vector with a fixed length; and converting the output after global average pooling into probability distribution of each category by using a probability type function to obtain a final classification result, continuously optimizing the model, and finally completing the identification of the digital content picture with the specific style. According to the method, on the premise of ensuring accurate recognition, the feature information in the picture is better extracted, the calculation efficiency is improved, and the method has good robustness and generalization ability, has good ability to adapt to new situations and can accurately predict a result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical fields of image recognition and risk recognition, and particularly relates to a method and device for identifying digital content with a specific style of pictures. Background Art

[0002] In recent years, with the development of blockchain technology and the rise of the digital asset market, new digital content based on blockchain technology has become a new form of artistic expression. From surreal anime characters to retro pixel art, and then to complex abstract expressionism, these works have attracted a wide audience with their highly creative and personalized characteristics. However, some lawbreakers have taken the opportunity to create many new digital content with specific styles of pictures containing risky content by taking advantage of the popularity of the products.

[0003] These specific-style pictures with risks not only affect the healthy development of the digital content market but may also lead to the spread of illegal content, causing social problems. For example, politically sensitive content may lead to social instability, pornographic content may have an adverse impact on minors, and violent content may induce criminal behavior.

[0004] These risky specific-style pictures also pose regulatory challenges. Traditional risk image recognition methods are powerless in identifying risk elements such as politically sensitive, pornographic, or violent elements in pictures of abstract art categories such as two-dimensional, cartoon pixel, etc. The main reason is that these digital contents usually adopt exaggerated expression techniques, non-realistic color combinations, and complex visual elements, which are very different from the standard image data sets used in traditional model training. Especially in identifying images involving politically sensitive content, porn, or violence, traditional models may produce false positives or false negatives, challenging the accuracy and effectiveness of content supervision. Summary of the Invention

[0005] Aiming at the deficiencies of the prior art, the present invention provides a method for identifying digital content with a specific style of pictures, which can solve the problems that traditional picture recognition methods cannot accurately identify the risk content in new digital content with specific styles of pictures and the low efficiency of traditional recognition methods. The method proposed by the present invention can better extract the feature information in the pictures on the premise of ensuring accurate recognition, improve the calculation efficiency, and the method has good robustness and generalization ability. The anti-interference ability of this method is very strong, it can handle various noises and errors, and when facing new data, its ability to adapt to new situations is very good, and it can accurately predict the results.

[0006] The object of the present invention is achieved through the following technical solutions:

[0007] The first aspect of the present invention: A method for identifying digital content with a specific style of pictures, comprising the following steps:

[0008] (1) Manually screen and sort out new digital content pictures with relevant risk characteristics and risk-free ones on the new digital content platform, then use web crawler technology to automatically obtain specific digital content samples in these sets, and use style transfer technology to expand the dataset, convert normal pictures into the specified style, and expand the risk dataset; construct a dataset containing risky content and risk-free content and split it into a training set and a test set in the ratio of 8:2 for training;

[0009] (2) Input the image data into the backbone layer, the first dense connection layer, the transition layer, several other dense connection layers and transition layers in sequence, then perform global average pooling to convert the feature map into a vector of fixed length;

[0010] (3) Through step (2), the MCA-DenseNet model can be constructed, randomly load batch data from the training set in step (1), and input it into the MCA-DenseNet model for forward propagation to calculate the prediction result; calculate the loss value according to the difference between the prediction result and the true label, use the backpropagation algorithm to calculate the gradient, and use the Adam optimizer to update the model parameters to minimize the loss;

[0011] (4) At the end of each epoch, use an independent validation set to evaluate the model performance, record the validation loss and accuracy. If the validation performance is better than the previous best performance, save the model weights; cycle to optimize the model and use the optimized model to complete the recognition of digital content pictures with specific styles.

[0012] Further, the step (2) includes the following sub-steps:

[0013] (2.1) First, input the image data into the backbone layer, pass through 3 convolutional layers of 3*3 with strides of 2, 1, 1 respectively, and enter a 3*3 max pooling layer after passing through the convolutional layers for preliminary feature extraction;

[0014] (2.2) Then enter the first dense connection layer. In this layer, there are multiple connection blocks. Each block takes the feature maps of all previous blocks as input and provides its own generated feature maps to all subsequent blocks. The operation sequence inside each block is completed by a combination of two layers; and use a probabilistic function to convert the output after global average pooling into a probability distribution of each category to obtain the final classification result;

[0015] (2.3) Add a permutation attention mechanism between the dense connection layer and the transition layer to calculate the mutual attention weights between features, and capture context-dependent information according to the calculated attention weights;

[0016] (2.4) The data processed by the attention mechanism is input into the transition layer. The transition layer has three channels. The data passes through the three channels respectively, and a downsampling module with the feature of feature fusion is added to fuse the results of the three channels, splice the feature map data, and obtain new feature map information;

[0017] (2.5) Subsequently, subsequent dense connection layers are constructed. The sizes and the number of connection blocks of the feature maps in these connection layers are different from those in the first dense connection layer, ensuring multi-level extraction of features. The operations inside each dense connection layer are the same as those in the first layer, but the sizes and the number of feature maps are different;

[0018] (2.6) Between each dense connection layer, a transition layer is inserted to further reduce the size of the feature map and reduce the number of feature maps through a compression factor; after the last dense connection layer, global average pooling is performed to convert the feature map into a vector with a fixed length;

[0019] (2.7) Finally, a probabilistic function is used to convert the output after global average pooling into the probability distribution of each category to obtain the final classification result.

[0020] Further, the specific combination of the two layers in step (2.2) is as follows: the first layer combination is batch normalization, the Swish activation function, and a 1*1 convolution; the second layer combination is batch normalization, the Swish activation function, and a 3*3 convolutional layer.

[0021] Further, the transition layer structure in step (5) includes three channels, which are respectively:

[0022] (a) Channel 1: The input feature map of the channel undergoes downsampling through a max pooling layer and is normalized. This channel is mainly used to retain global information;

[0023] (b) Channel 2: The input feature map first passes through a 1x1 convolutional layer to adjust the number of channels, and is normalized and activated. Then it passes through two 3*3 convolutional layers. The first has a stride of 1 and the second has a stride of 2. Normalization and activation are performed at each step. This channel is mainly used to extract local information and perform downsampling;

[0024] (c) Channel 3: The input feature map first passes through a 1x1 convolutional layer to adjust the number of channels, and is normalized and activated. Then it passes through a 3*3 convolutional layer with a stride of 2 and is normalized and activated again. This channel is also mainly used to extract local information and perform downsampling.

[0025] Further, step (4) specifically includes the following sub-steps:

[0026] (9.1) Set the total number of training epochs. Each epoch is divided into several batches. The output of each epoch includes: training loss, training accuracy, validation loss, and validation accuracy. The training loss is the average loss value on the training set for the current epoch. The training accuracy is the accuracy on the training set for the current epoch. The validation loss is the average loss value on the validation set for the current epoch. The validation accuracy is the accuracy on the validation set for the current epoch. After training is completed, obtain the model file with the best performance for subsequent inference and testing.

[0027] (9.2) Training process:

[0028] First, load batch data from the training set to ensure the randomness of samples in each batch. Select the number of batches, and then input the data of this number of batches into the MCA-DenseNet model for forward propagation. Calculate the prediction results through the activation functions of each layer to obtain the final model output. Next, calculate the loss value of the current batch based on the difference between the model's prediction results and the true labels to reflect the accuracy of the model output and guide the direction of updating the model parameters in subsequent training. Calculate the gradient of the loss function with respect to the model parameters through the backpropagation algorithm, and use the Adam optimizer to update the model parameters. The model is continuously optimized until the most accurate prediction and recognition of digital content specific style pictures are finally completed.

[0029] The second aspect of the present invention: An apparatus for identifying digital content specific style pictures, including the following modules:

[0030] Dataset construction module: Manually screen and sort out new digital content pictures with relevant risk characteristics and risk-free on the new digital content platform, then use web crawler technology to automatically obtain specific digital content samples in these sets, and use style transfer technology to expand the dataset, convert normal pictures into the specified style, and expand the risk dataset; construct a dataset containing risk content and risk-free content and split it into a training set and a test set in a ratio of 8:2 for training.

[0031] Pooling conversion module: Input the image data into the backbone layer, the first dense connection layer, the transition layer, several other dense connection layers and transition layers in sequence, and then perform global average pooling to convert the feature map into a vector with a fixed length.

[0032] Calculation module: The MCA-DenseNet model can be constructed through the pooling conversion module, randomly load batch data from the training set of the dataset construction module, and input it into the MCA-DenseNet model for forward propagation to calculate the prediction results; calculate the loss value based on the difference between the prediction results and the true labels, calculate the gradient using the backpropagation algorithm, and use the Adam optimizer to update the model parameters to minimize the loss.

[0033] Verification module: At the end of each cycle, the performance of the model is evaluated using an independent validation set, the validation loss and accuracy are recorded, and if the validation performance is better than the previous best performance, the model weights are saved; The loop-optimized model uses the optimized model to complete the recognition of digital content with specific styles of pictures.

[0034] The third aspect of the present invention: An electronic device, comprising:

[0035] One or more processors;

[0036] A memory for storing one or more programs;

[0037] When the one or more programs are executed by the one or more processors, the one or more processors implement the method for identifying digital content with specific styles of pictures.

[0038] The fourth aspect of the present invention: A computer-readable storage medium, on which computer instructions are stored, and when the instructions are executed by a processor, the method steps for identifying digital content with specific styles of pictures are implemented.

[0039] The beneficial effects of the present invention are as follows:

[0040] 1. Promote content supervision and maintain market health:

[0041] The present invention is particularly optimized for the risk content recognition of abstract art pictures in new digital content. By accurately identifying the risk content in new digital content pictures, it provides strong technical support for the supervision of the digital content market, and helps to create a safer and healthier digital content environment.

[0042] 2. Reduce computing costs and improve computing efficiency:

[0043] In view of the fact that new digital content pictures have less detailed information, the present invention replaces the 7×7 convolutional layer in the backbone layer with three 3×3 convolutional layers, effectively reducing the pixels of the input image, greatly reducing the data volume and computational complexity, and thus improving the running efficiency of the model. This improvement not only speeds up the training speed of the model, but also makes real-time recognition in practical applications possible.

[0044] 3. Enhance feature extraction ability and improve recognition accuracy:

[0045] By introducing an attention mechanism between the dense block and the transition layer, the present invention can adjust the weights of the extracted image features in a complex image environment, enabling the model to pay more attention to the key regions in the image. Adding the attention mechanism to the original model significantly improves the accuracy and robustness of the model in identifying risk content such as politically sensitive, pornographic, and violent content, and can maintain a high recognition accuracy even in the case of poor image quality or interference.

[0046] 4. Integrate multi-source information and optimize the decision-making process:

[0047] The present invention creatively replaces the transition layer with a downsampling module with feature fusion characteristics, which can effectively splice and fuse the feature map data from different layers to form a richer and more comprehensive feature representation. This improvement helps the model better understand the overall structure and details of the image, thereby making a more accurate classification decision.

[0048] 5. Strengthen the network structure and enhance the generalization ability:

[0049] The present invention adopts the Swish activation function instead of the traditional ReLU activation function, enabling the model to have better generalization ability and higher learning efficiency when dealing with deep network structures. In addition, by adding a permutation attention mechanism between the dense block and the transition layer, the present invention further enhances the diversity and robustness of the network, enabling it to still maintain good adaptability and prediction accuracy when facing new data. BRIEF DESCRIPTION OF THE DRAWINGS

[0050] The drawings herein are incorporated into the specification and form a part of the specification, showing embodiments consistent with the present application, and are used together with the specification to explain the principles of the present application.

[0051] Figure 1 is the flowchart of the MCA-DenseNet method of the present invention;

[0052] Figure 2 is the structure diagram of the backbone layer of the present invention;

[0053] Figure 3 is the structure diagram of the transition layer of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0054] Here, the exemplary embodiments will be described in detail, and the examples are shown in the drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present application.

[0055] The terms used in this application are for the purpose of describing specific embodiments only and are not intended to limit this application. The singular forms "a", "the", and "said" used in this application and the appended claims are also intended to include the plural forms unless the context clearly dictates otherwise. It should also be understood that the term "and / or" used herein refers to and encompasses any and all possible combinations of one or more of the associated listed items.

[0056] The method identification process of the present invention is as Figure 1 shown, and the specific steps are as follows:

[0057] (1) Manually screen and sort out new digital content pictures with relevant risk characteristics and risk-free ones on the new digital content platform, then use web crawler technology to automatically obtain specific digital content samples in these sets, and use style transfer technology to expand the data set, convert normal pictures into a specified style, and expand the risk data set. Construct a data set containing risky content and risk-free content. The specific risk categories include politics, terror, violence, vulgarity, and crime, etc. The constructed data set is split into a training set and a test set according to 8:2, and trained for 100 epochs.

[0058] (2) First, input the image data into the backbone layer, as Figure 2 shown, through three 3*3 convolutional layers with strides of 2, 1, and 1 respectively. After passing through the convolutional layers, enter a 3*3 max pooling layer for preliminary feature extraction.

[0059] (3) Then enter the first dense connection layer. In this layer, there are multiple connection blocks. Each block takes the feature maps of all the previous blocks as input and provides its own generated feature maps to all subsequent blocks. The operation sequence inside each block is a combination of two layers. The first layer combination is batch normalization, Swish activation function, and 1*1 convolutional layer, and the second layer combination is batch normalization, Swish activation function, and 3*3 convolutional layer.

[0060] (4) A permutation attention mechanism is added between the dense connection layer and the transition layer to calculate the mutual attention weights between features, and capture context-dependent information according to the calculated attention weights.

[0061] (5) The data processed by the attention mechanism is input into the transition layer. The transition layer has three channels. The data passes through the three channels respectively, and a downsampling module with the characteristic of feature fusion is added to fuse the results of the three channels, splice the feature map data, and obtain new feature map information. The structure of the transition layer is as Figure 3 shown. The three channels are respectively:

[0062] (a) Channel 1: The input feature map of the channel is downsampled through a max pooling layer and normalized. This channel is mainly used to retain global information.

[0063] (b) Channel 2: The input feature map first passes through a 1x1 convolutional layer to adjust the number of channels, and then undergoes normalization and activation processing. Then it passes through two 3*3 convolutional layers, with the first having a stride of 1 and the second having a stride of 2. Normalization and activation are performed at each step. This channel is mainly used to extract local information and perform downsampling.

[0064] (c) Channel 3: The input feature map first passes through a 1x1 convolutional layer to adjust the number of channels, and then undergoes normalization and activation processing. Then it passes through a 3*3 convolutional layer with a stride of 2 and is normalized and activated again. This channel is also mainly used to extract local information and perform downsampling.

[0065] (6) Subsequently, subsequent dense connection layers are constructed. These connection layers are similar to the first dense connection layer, but the size of the feature map and the number of connection blocks are different, ensuring multi-level extraction of features. The operations within each dense connection layer are the same as those of the first layer, but the size and number of the feature map are different.

[0066] (7) Between each dense connection layer, a transition layer is inserted to further reduce the size of the feature map and reduce the number of feature maps through a compression factor; after the last dense connection layer, global average pooling is performed to convert the feature map into a fixed-length vector.

[0067] (8) Finally, after global average pooling, a probabilistic function is used to convert the output of the fully connected layer into a probability distribution for each category, obtaining the final classification result.

[0068] (9) Using the training dataset obtained in step (1), start training after inputting the data into the MCA-DenseNet model constructed in steps (2)-(8); specifically, it includes the following steps:

[0069] (9.1) Set the total number of training epochs to 100, divided into 16 batches in each epoch. The output of each epoch mainly includes:

[0070] Training loss: The average loss value on the training set in the current epoch;

[0071] Training accuracy: The accuracy on the training set in the current epoch;

[0072] Validation loss: The average loss value on the validation set in the current epoch;

[0073] Validation accuracy: The accuracy on the validation set in the current epoch;

[0074] After training is completed, the model file with the best performance is obtained for subsequent inference and testing.

[0075] Configuration: NVIDIA GeForce RTX 3090 graphics card.

[0076] Software environment: Using Python 3.8 and PyTorch 1.9.0

[0077] (9.2) Training process:

[0078] First, load batch data from the training set, ensuring the randomness of samples in each batch to enhance the generalization ability of the model. In this method, the number of batches is selected as 16. Subsequently, the loaded batch data is input into the MCA-DenseNet model for forward propagation, and the prediction results are calculated through the activation functions of each layer to obtain the final model output. Next, according to the difference between the model's prediction results and the true labels, calculate the loss value of the current batch to reflect the accuracy of the model output and guide the direction of model parameter update in subsequent training.

[0079] In this process, calculate the gradient of the loss function with respect to the model parameters through the backpropagation algorithm, and use the Adam optimizer to update the model parameters to minimize the loss and improve the model's performance. At the end of each epoch, use an independent validation set to evaluate the model's performance, and record the validation loss and accuracy to monitor the model's performance on unseen data. In addition, if the validation performance of the current model is better than the previous best performance, save the weights of this model, which not only helps to avoid the decline of the model's performance in the subsequent training process, but also facilitates later inference and testing. Through such a cyclic training process, the model can be continuously iteratively optimized and finally achieve good prediction ability.

[0080] The present invention also discloses a device for identifying pictures with specific styles of digital content, including the following modules:

[0081] Dataset construction module: Manually screen and sort out new digital content pictures with relevant risk characteristics and risk-free on the new digital content platform, then use web crawler technology to automatically obtain specific digital content samples in these sets, and use style transfer technology to expand the dataset, convert normal pictures into specified styles, and expand the risk dataset; construct a dataset containing risk content and risk-free content and split it into a training set and a test set according to 8:2 for training;

[0082] Pooling conversion module: After sequentially inputting the image data into the backbone layer, the first dense connection layer, the transition layer, several other dense connection layers and transition layers, perform global average pooling to convert the feature map into a vector with a fixed length;

[0083] Computing module: The MCA-DenseNet model can be constructed through the pooling transformation module, and batch data can be randomly loaded from the training set of the dataset construction module and input into the MCA-DenseNet model for forward propagation to calculate the prediction results; calculate the loss value according to the difference between the prediction results and the true labels, calculate the gradients using the backpropagation algorithm, and update the model parameters using the Adam optimizer to minimize the loss;

[0084] Verification module: At the end of each epoch, evaluate the model performance using an independent validation set, record the validation loss and accuracy, and save the model weights if the validation performance is better than the previous best performance; cycle to optimize the model and use the optimized model to complete the recognition of digital content specific style pictures.

[0085] The present invention also discloses an electronic device, including: one or more processors; a memory for storing one or more programs; when the one or more programs are executed by the one or more processors, enabling the one or more processors to implement the method for identifying digital content specific style pictures as described above; and discloses a computer-readable storage medium, on which computer instructions are stored, and when the instructions are executed by a processor, the method steps for identifying digital content specific style pictures as described above are implemented.

[0086] Those skilled in the art will readily conceive of other embodiments of the present application after considering the specification and practicing the content disclosed herein. The present application is intended to cover any variations, uses, or adaptations of the present application, which follow the general principles of the present application and include common general knowledge or conventional technical means in the technical field not disclosed in the present application.

[0087] It should be understood that the present application is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope.

Claims

1. A method for identifying pictures with specific styles of digital content, characterized in that, It includes the following steps: (1) Manually screen and sort out new digital content pictures with relevant risk characteristics and risk-free ones on the new digital content platform, then use web crawler technology to automatically obtain specific digital content samples in these sets, and use style transfer technology to expand the dataset, convert normal pictures into a specified style, and expand the risk dataset; construct a dataset containing risky content and risk-free content and split it into a training set and a test set according to 8:2 for training; (2) Input the image data into the backbone layer, the first dense connection layer, the transition layer, several other dense connection layers and transition layers in sequence, then perform global average pooling to convert the feature map into a vector of fixed length; (3) Through step (2), an MCA-DenseNet model can be constructed, randomly load batch data from the training set in step (1), and input it into the MCA-DenseNet model for forward propagation to calculate the prediction result; calculate the loss value according to the difference between the prediction result and the true label, use the backpropagation algorithm to calculate the gradient, and use the Adam optimizer to update the model parameters to minimize the loss; (4) At the end of each epoch, use an independent validation set to evaluate the model performance, record the validation loss and accuracy. If the validation performance is better than the previous best performance, save the model weights; loop to optimize the model and use the optimized model to complete the recognition of digital content pictures with specific styles.

2. The method for identifying pictures with specific styles of digital content according to claim 1, wherein The said step (2) includes the following sub-steps: (2.1) First, input the image data into the backbone layer, pass through 3 convolutional layers of 3*3 with strides of 2, 1, 1 respectively, and enter a 3*3 max pooling layer after passing through the convolutional layers for preliminary feature extraction; (2.2) Then enter the first dense connection layer. In this layer, there are multiple connection blocks. Each block takes the feature maps of all the previous blocks as input and provides its own generated feature maps to all subsequent blocks. The operation sequence inside each block is completed by a combination of two layers; and use a probabilistic function to convert the output after global average pooling into a probability distribution of each category to obtain the final classification result; (2.3) A permutation attention mechanism is added between the dense connection layer and the transition layer to calculate the mutual attention weights between features, and capture context-dependent information according to the calculated attention weights; (2.4) The data processed by the attention mechanism is input into the transition layer. The transition layer has three channels. The data passes through the three channels respectively, and a downsampling module with the characteristic of feature fusion is added to fuse the results of the three channels and splice the feature map data to obtain new feature map information; (2.5) Subsequently, construct subsequent dense connection layers. The sizes and the number of connection blocks of the feature maps in these connection layers are different from those in the first dense connection layer, ensuring multi-level extraction of features. The operations inside each dense connection layer are the same as those in the first layer, but the sizes and quantities of the feature maps are different; (2.6) Between each dense connection layer, insert a transition layer to further reduce the size of the feature map and reduce the number of feature maps by a compression factor; after the last dense connection layer, perform global average pooling to convert the feature map into a fixed-length vector; (2.7) Finally, use a probabilistic function to convert the output after global average pooling into a probability distribution for each category to obtain the final classification result.

3. A method for identifying pictures with specific styles of digital content according to claim 2, characterized in that, The specific combination of the two layers in step (2.2) is as follows: the first layer combination is batch normalization, the Swish activation function, and a 1×1 convolution; the second layer combination is batch normalization, the Swish activation function, and a 3×3 convolutional layer.

4. A method for identifying pictures with specific styles of digital content according to claim 1, characterized in that, The transition layer structure in step (5) includes three channels, namely: (a) Channel 1: The input feature map of the channel undergoes downsampling through a max-pooling layer and is normalized. This channel is mainly used to retain global information; (b) Channel 2: The input feature map first passes through a 1×1 convolutional layer to adjust the number of channels, and then undergoes normalization and activation processing. Then it passes through two 3×3 convolutional layers. The first has a stride of 1, and the second has a stride of 2. Normalization and activation processing are performed at each step. This channel is mainly used to extract local information and perform downsampling; (c) Channel 3: The input feature map first passes through a 1×1 convolutional layer to adjust the number of channels, and then undergoes normalization and activation processing. Then it passes through a 3×3 convolutional layer with a stride of 2 and undergoes normalization and activation processing again. This channel is also mainly used to extract local information and perform downsampling.

5. A method for identifying pictures with specific styles of digital content according to claim 1, characterized in that Step (4) specifically includes the following sub-steps: (9.1) Set the total number of training epochs. Each epoch is divided into several batches. The output of each epoch includes: training loss, training accuracy, validation loss, and validation accuracy; the training loss is the average loss value on the training set for the current epoch; the training accuracy is the accuracy on the training set for the current epoch; the validation loss is the average loss value on the validation set for the current epoch; the validation accuracy is the accuracy on the validation set for the current epoch; after training is completed, obtain the model file with the best performance for subsequent inference and testing; (9.2) Training process: First, load batch data from the training set to ensure the randomness of the samples in each batch; select the batch number, and then input the data of the selected batch number into the MCA-DenseNet model for forward propagation. Calculate the prediction result through the activation functions of each layer to obtain the final model output; next, calculate the loss value of the current batch according to the difference between the model's prediction result and the true label to reflect the accuracy of the model output and guide the direction of updating the model parameters in subsequent training. Calculate the gradient of the loss function with respect to the model parameters through the backpropagation algorithm, and use the Adam optimizer to update the model parameters. The model is continuously optimized until the most accurate prediction and recognition of the digital content specific style pictures are finally completed.

6. An apparatus for identifying pictures with specific styles of digital content, characterized in that, Include the following modules: Dataset construction module: Manually screen and organize new digital content pictures with relevant risk characteristics and risk-free ones on the new digital content platform, then use web crawler technology to automatically obtain specific digital content samples in these sets, and use style transfer technology to expand the dataset, convert normal pictures into the specified style, and expand the risk dataset; construct a dataset containing risk content and risk-free content and split it into a training set and a test set according to 8:2 for training; Pooling conversion module: After sequentially inputting the image data into the backbone layer, the first dense connection layer, the transition layer, several other dense connection layers and transition layers, perform global average pooling to convert the feature map into a vector of fixed length; Calculation module: Through the pooling conversion module, an MCA-DenseNet model can be constructed, randomly load batch data from the training set of the dataset construction module, and input it into the MCA-DenseNet model for forward propagation to calculate the prediction result; calculate the loss value according to the difference between the prediction result and the true label, use the backpropagation algorithm to calculate the gradient, and use the Adam optimizer to update the model parameters to minimize the loss; Verification module: At the end of each epoch, use an independent validation set to evaluate the model performance, record the validation loss and accuracy, and if the validation performance is better than the previous best performance, save the model weights; loop to optimize the model and use the optimized model to complete the recognition of digital content pictures with specific styles.

7. An electronic device, characterized in that, Including: One or more processors; A memory for storing one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement a method for identifying digital content pictures with specific styles as described in any one of claims 1-7.

8. A computer-readable storage medium having computer instructions stored thereon, characterized in that, When the instruction is executed by the processor, it implements the method steps of a method for identifying digital content pictures with specific styles as described in any one of claims 1-7.