Object storage content pornographic identification method and device
Through the combination of object storage and deep learning models, the problems of low efficiency and poor accuracy of user-generated content identification in the prior art are solved, efficient and accurate identification of audio and video content are achieved, and the scalability of the system and the healthy operation capabilities of the platform are improved.
Patent Information
- Application Number
- CN202510338188.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-21
- Publication Date
- 2025-07-18
AI Technical Summary
The prior art has low efficiency, high cost, high risk of misjudgment and misjudgment when processing user-generated content, especially insufficient discrimination capabilities for audio and video content, and insufficient performance and scalability of the system during large-scale data processing.
The automated process of object storage module, preprocessing module, large model identification module, result storage module and alarm module is adopted, combined with ResNet50, VGGish and Transformer models, feature extraction and classification identification of pictures, audio and video content is carried out. Through multimodal feature fusion, subtle differences and complex semantics in the content are captured, and multi-threaded parallel processing and real-time alarms are supported.
It significantly improves the identification accuracy and processing efficiency of pornographic content, reduces the risks of misjudgment and misjudgment, supports efficient processing of large-scale data, ensures healthy operation and user experience of the platform, and has good scalability and stability.
Smart Images

Figure CN120336949A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of cloud computing, and specifically provides a method and device for identifying pornographic content in object storage. Background Art
[0002] With the rapid development of the Internet, the quantity of user-generated content (UGC) on online platforms has increased rapidly. While bringing content diversity and richness, it has also caused problems of a large number of bad information transmissions. In particular, the proliferation of pornographic content not only violates laws and regulations but also seriously affects the healthy development of the Internet environment. Therefore, how to efficiently and accurately identify and filter pornographic content has become an urgent problem for major platforms to solve.
[0003] Traditional content review methods mainly rely on manual review, where reviewers check and judge whether the content is pornographic one by one. Although this method is direct, it has significant drawbacks: First, the efficiency of manual review is low. Facing a vast amount of UGC content, manual review cannot process it in a timely manner. Second, the labor cost is high. Reviewers need to concentrate for a long time, which easily leads to a decline in review quality. Finally, manual review is easily affected by subjective factors and there are risks of misjudgment and missed judgment.
[0004] To solve the above problems, automated content review tools have emerged as the times require. Early automated review tools were based on rules and keyword matching, and performed simple text and image analysis through predefined pornographic keyword lists and rules. However, the accuracy and robustness of these methods are limited and cannot cope with the complexity and diversity of content. Especially for audio and video content, they are even more powerless.
[0005] In recent years, with the rapid development of deep learning technology, content identification methods based on deep learning models have gradually become mainstream. Deep learning models, especially convolutional neural networks (CNNs) and self-attention mechanisms, have achieved remarkable results in fields such as image recognition, speech recognition, and natural language processing. This provides a new technical path for the automated identification of pornographic content.
[0006] In the prior art, some deep learning-based pornographic content identification systems have begun to be applied in practice, but there are still some deficiencies. First, a single type of model (such as only using CNN) is prone to problems of low recognition accuracy when dealing with complex and diverse UGC content. Second, existing systems often lack effective processing of audio and video content and can only identify picture or text content. In addition, most systems still need to be improved in terms of performance and scalability when facing large-scale data. Summary of the Invention
[0007] In view of the deficiencies of the above-mentioned prior art, the present invention provides a method for identifying pornographic content in object storage with strong practicability.
[0008] A further technical task of the present invention is to provide an object storage content pornographic identification device with reasonable design, safety and applicability.
[0009] The technical solution adopted by the present invention to solve its technical problems is as follows:
[0010] A method for identifying pornographic content in object storage has the following steps:
[0011] S1. The object storage module receives the user's upload and download requests, converts the data provided by the user into an unstructured data stream and stores it in the disk of the cloud device, or converts the unstructured data stream in the disk of the cloud device into the data format required by the user and sends it to the user;
[0012] S2. The preprocessing module processes the picture, audio and video content uploaded and downloaded by the user;
[0013] S3. The large model identification module identifies pornographic content in the preprocessed content;
[0014] S4. The result storage module stores the identification result and relevant log information;
[0015] S5. The alarm module triggers an alarm mechanism when pornographic content is detected.
[0016] Further, in step S1, during the conversion process, pornographic content identification operations are performed on data in picture, audio and video formats.
[0017] Further, in step S2, for picture preprocessing, first convert all pictures to the JPEG format, scale the pictures to a fixed size of 224x224 pixels, use the interpolation method for scaling, and adopt data augmentation technology to generate diverse training samples;
[0018] For audio preprocessing, first convert all audio files to the WAV format, standardize the sampling rate of all audio files to a unified value of 16 kHz, divide the audio signal into several frames, with each frame having a fixed length of 25 ms and a 10 ms overlap between adjacent frames to capture the temporal characteristics of the audio;
[0019] For video preprocessing, extract key frames from the video file at a fixed interval of 1 frame per second, convert the extracted key frames to the unified JPEG format, scale the key frames to a fixed size of 224x224 pixels, and perform data augmentation processing on the key frames.
[0020] Further, in step S3, it includes:
[0021] (1) Feature extraction, including image feature extraction, audio feature extraction, and video feature extraction;
[0022] (2) Classification and discrimination, which are carried out based on a large Transformer model. The Transformer model processes the input sequence data through the self-attention mechanism to capture local and global information in images, audio, and videos;
[0023] (3) Model training and optimization.
[0024] Further, in step (1), in the image feature extraction, a pre-trained ResNet50 model is used for image feature extraction. The gradient vanishing problem of the deep network is solved through residual connections, enabling the model to be trained more effectively. Specifically as follows:
[0025] (1-1.1) Input layer: The preprocessed image is input into the input layer of ResNet50 for convolution operation;
[0026] (1-1.2) Convolutional layer and residual block: ResNet50 contains multiple convolutional layers and residual blocks. Through multiple convolutions and non-linear transformations of these layers, high-level features of the image are extracted layer by layer;
[0027] (1-1.3) Global average pooling layer: The feature map output by the convolutional layer is subjected to global average pooling to obtain a feature vector of a fixed length;
[0028] (1-1.4) Feature vector output: The finally output feature vector represents the high-level semantic information of the image and is used for subsequent classification tasks;
[0029] In the audio feature extraction, a VGGish model is used for audio feature extraction. By converting the audio signal into a spectrogram and then inputting it into a pre-trained VGG model for feature extraction, specifically as follows:
[0030] (1-2.1) Spectrogram generation: Each frame of the audio signal is converted into a spectrogram. The short-time Fourier transform (STFT) is used to convert the time-domain signal into a frequency-domain signal to generate a two-dimensional spectrogram;
[0031] (1-2.2) VGGish model input: The spectrogram is input into the pre-trained VGGish model for convolution operation and feature extraction;
[0032] (1-2.3) Convolutional layer and pooling layer: The VGGish model contains multiple convolutional layers and pooling layers. Through multiple convolutions and pooling operations of these layers, high-level features of the audio are extracted layer by layer;
[0033] (1-2.4) Feature vector output: The finally output feature vector represents the high-level semantic information of the audio and is used for subsequent classification tasks;
[0034] In the video feature extraction, a pre-trained ResNet50 model is used to extract features from video key frames, and a Transformer model is combined to capture the temporal features of the video, specifically as follows:
[0035] (1-3.1) Input layer: The pre-processed key frames are input into the input layer of ResNet50 for convolution operations;
[0036] (1-3.2) Convolution layer and residual block: ResNet50 extracts high-level features of key frames layer by layer through multiple convolutions and non-linear transformations of multiple convolution layers and residual blocks;
[0037] (1-3.3) Global average pooling layer: The feature map output by the convolution layer is globally average pooled to obtain a feature vector of a fixed length;
[0038] (1-3.4) Feature vector output: The finally output feature vector represents the high-level semantic information of the key frame;
[0039] (1-3.5) Sequence feature combination: The feature vectors of all key frames of the same video are combined in chronological order to form a feature sequence for subsequent classification tasks.
[0040] Furthermore, in step (2), a large Transformer-based model is used for classification and discrimination. The Transformer model processes the input sequence data through the self-attention mechanism and can capture local and global information in pictures, audio, and videos. The specific classification and discrimination process is as follows:
[0041] (2.1) Feature sequence input: The feature vectors of pictures, audio, and videos are input into the Transformer model to form a feature sequence; for pictures and audio, each input feature vector is processed separately; for videos, the feature vectors of all key frames of the same video are combined in chronological order to form a feature sequence;
[0042] (2.2) Self-attention calculation: Transformer calculates the correlation between each position in the feature sequence and other positions through the self-attention mechanism to generate a weighted feature representation;
[0043] (2.3) Multi-layer stacking: Transformer is stacked by multiple layers of self-attention and feed-forward neural networks, and processes the input feature sequence layer by layer to generate high-level feature representations;
[0044] (2.4) Classification layer: Input the output feature representation of the Transformer into the classification layer for classification judgment, and output the probability of whether it involves pornographic content.
[0045] Further, in step (3), it includes:
[0046] (3.1) Data preparation: Collect a large number of pornographic and non-pornographic pictures, audio, and video clips, annotate them, and form a training data set and a validation data set;
[0047] (3.2) Transfer learning: Based on the pre-trained ResNet50, VGGish, and Transformer models, perform fine-tuning through the transfer learning method to train the classification model;
[0048] (3.3) Loss function: Adopt the cross-entropy loss function Cross-Entropy Loss to measure the difference between the model output and the true label, and guide the optimization of model parameters;
[0049] (3.4) Optimization algorithm: Adopt the Adam optimization algorithm, and through backpropagation and gradient descent, iteratively update the model parameters.
[0050] Further, in step S4, it includes:
[0051] Identification result: Two legal values, pornographic or non-pornographic;
[0052] Metadata of the uploaded content: User ID, upload time, or download time;
[0053] Identification time and model version.
[0054] Further, in step S5, it includes:
[0055] a) Send an alarm notification to the management staff;
[0056] b) Generate an alarm record in the management background.
[0057] An apparatus for identifying pornographic content in object storage includes: at least one memory and at least one processor;
[0058] The at least one memory is used to store machine-readable programs;
[0059] The at least one processor is used to call the machine-readable program to execute a method for identifying pornographic content in object storage.
[0060] Compared with the prior art, the method and apparatus for identifying pornographic content in object storage of the present invention have the following prominent beneficial effects:
[0061] The present invention significantly improves the discrimination accuracy and processing efficiency. The system adopts ResNet50, VGGish and Transformer models to extract features and classify and discriminate picture, audio and video content respectively. Through multi-modal feature fusion, it captures the subtle differences and complex semantics in the content, effectively identifies pornographic content, and significantly reduces the risks of misjudgment and missed judgment.
[0062] A set of automated processing processes is designed, including object storage, preprocessing, large model discrimination, result storage and warning modules, ensuring the full automation of the review of UGC content, greatly reducing the workload of manual review and time cost. At the same time, it supports multi-threaded parallel processing and high-concurrency content uploading and processing, with good scalability and stability. Through key frame extraction and analysis, it can accurately capture pornographic segments in videos, enhancing the effect of video content review.
[0063] Data augmentation technology and transfer learning methods further improve the generalization ability and robustness of the model, ensuring that the model maintains high-precision discrimination performance in content of different types and styles. The real-time warning function is triggered immediately when pornographic content is detected, sending notifications to management personnel and generating warning records in the management background, ensuring that pornographic content can be processed and filtered in a timely manner to prevent its spread. The detailed logging and tracing function improves the transparency and controllability of content management.
[0064] In addition, through efficient and accurate discrimination of pornographic content, this method helps the platform operator comply with relevant laws and regulations, prevent the spread of illegal and bad information, create a healthy network environment, improve user experience and satisfaction, and promote the long-term healthy development of the platform.
[0065] In summary, the present invention has significant practical application value and broad promotion prospects in improving the efficiency and accuracy of content review, the scalability and stability of the system, and the guarantee of the healthy operation of the platform. Brief Description of the Drawings
[0066] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0067] Attached Figure 1 is a schematic flow chart of a method for discriminating pornographic content in object storage content. Detailed Embodiments
[0068] To enable those skilled in the art to better understand the solution of the present invention, the present invention will be further described in detail below in conjunction with specific embodiments. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.
[0069] The following gives an optimal embodiment:
[0070] As Figure 1 shown, a method for identifying pornographic content in object storage in this embodiment has the following steps:
[0071] S1. The object storage module receives the upload / download requests of users, converts the data provided by users into unstructured data streams and stores them in the disks of cloud devices, or converts the unstructured data streams in the disks of cloud devices into the data formats required by users and sends them to users; during the conversion process, pornographic content identification operations are performed on data in the formats of pictures, audio, and video.
[0072] S2. The preprocessing module processes the pictures, audio, and video content uploaded / downloaded by users;
[0073] The preprocessing module is used to process the pictures, audio, and video content uploaded / downloaded by users for subsequent identification by the large model. It includes:
[0074] For picture preprocessing, the uploaded pictures may have various formats (such as PNG, BMP, TIFF, etc.). For unified processing, all pictures are first converted into the JPEG format.
[0075] The pictures are scaled to a fixed size of 224x224 pixels to ensure that the sizes of the pictures input into the model are the same. The interpolation method is used for scaling to retain the original information of the pictures as much as possible. To improve the generalization ability of the model, data augmentation techniques (such as random cropping, rotation, flipping, adjusting brightness and contrast, etc.) are adopted to generate diverse training samples.
[0076] For audio preprocessing, the uploaded audio files may have various formats (such as MP3, AAC, FLAC, etc.). First, all audio files are converted into the WAV format.
[0077] The sampling rates of all audio files are standardized to a unified value of 16kHz to ensure that the audio data input into the model is the same.
[0078] The audio signal is divided into several frames, each frame having a fixed length of 25ms, and there is an overlap of 10ms between adjacent frames to capture the temporal characteristics of the audio.
[0079] For video preprocessing, key frames are extracted from the video file at a fixed interval of 1 frame per second to ensure that the selected video frames can represent the main content of the video. The extracted key frames are converted into a unified JPEG format and scaled to a fixed size of 224x224 pixels to ensure that the input images to the model have the same size.
[0080] To improve the generalization ability of the model, data augmentation is performed on the key frames, such as random cropping, rotation, flipping, etc.
[0081] S3. The large model discrimination module discriminates against pornographic content in the preprocessed content;
[0082] The large model discrimination module is used to discriminate against pornographic content in the preprocessed content. The specific steps are as follows:
[0083] (1) Feature extraction, including image feature extraction, audio feature extraction, and video feature extraction;
[0084] The goal of image feature extraction is to convert the original image data into feature vectors that can represent the content of the image. These feature vectors will be used for subsequent classification tasks. A pre-trained ResNet50 model is used for image feature extraction. ResNet50 is a deep convolutional neural network (CNN) that solves the problem of gradient disappearance in deep networks through residual connections, enabling the model to be trained more effectively.
[0085] The specific feature extraction process:
[0086] (1-1.1) Input layer: The preprocessed image is input to the input layer of ResNet50 for convolution operations;
[0087] (1-1.2) Convolutional layer and residual block: ResNet50 contains multiple convolutional layers and residual blocks. Through multiple convolutions and non-linear transformations of these layers, high-level features of the image are extracted layer by layer;
[0088] (1-1.3) Global average pooling layer: The feature map output by the convolutional layer is subjected to global average pooling to obtain a feature vector of a fixed length;
[0089] (1-1.4) Feature vector output: The finally output feature vector represents the high-level semantic information of the image and is used for subsequent classification tasks;
[0090] Audio feature extraction aims to convert the original audio signal into feature vectors that can represent the audio content. In the present invention, a VGGish model is used for audio feature extraction. VGGish is an audio feature extraction tool based on the VGG model. By converting the audio signal into a spectrogram and then inputting it into the pre-trained VGG model to extract features.
[0091] Specific feature extraction process:
[0092] (1-2.1) Spectrogram generation: Convert the audio signal of each frame into a spectrogram, and use the short-time Fourier transform (STFT) to convert the time-domain signal into a frequency-domain signal to generate a two-dimensional spectrogram;
[0093] (1-2.2) Input to the VGGish model: Input the spectrogram into the pre-trained VGGish model for convolution operations and feature extraction;
[0094] (1-2.3) Convolutional layers and pooling layers: The VGGish model contains multiple convolutional layers and pooling layers. Through multiple convolution and pooling operations of these layers, high-level features of the audio are extracted layer by layer;
[0095] (1-2.4) Output of the feature vector: The finally output feature vector represents the high-level semantic information of the audio and is used for subsequent classification tasks;
[0096] The goal of video feature extraction is to convert the original video data into a feature vector that can represent the video content. Use the pre-trained ResNet50 model to extract features from the video key frames, and combine the Transformer model to capture the temporal features of the video.
[0097] Specific feature extraction process:
[0098] (1-3.1) Input layer: The preprocessed key frame is input into the input layer of ResNet50 for convolution operations;
[0099] (1-3.2) Convolutional layers and residual blocks: ResNet50 extracts high-level features of the key frame layer by layer through multiple convolution and non-linear transformation of multiple convolutional layers and residual blocks;
[0100] (1-3.3) Global average pooling layer: Perform global average pooling on the feature map output by the convolutional layer to obtain a feature vector of a fixed length;
[0101] (1-3.4) Output of the feature vector: The finally output feature vector represents the high-level semantic information of the key frame;
[0102] (1-3.5) Sequence feature combination: Combine the feature vectors of all key frames of the same video in chronological order to form a feature sequence for subsequent classification tasks.
[0103] (2) Classification and discrimination: Perform classification and discrimination based on the large model of Transformer. The Transformer model processes the input sequence data through the self-attention mechanism to capture local and global information in pictures, audio, and videos;
[0104] Specific classification and identification process:
[0105] (2.1) Feature sequence input: Input the feature vectors of pictures, audio, and video into the Transformer model to form a feature sequence. For pictures and audio, each input feature vector is processed separately. For video, all key-frame feature vectors of the same video are combined in chronological order to form a feature sequence.
[0106] (2.2) Self-attention calculation: The Transformer calculates the correlation between each position and other positions in the feature sequence through the self-attention mechanism to generate a weighted feature representation.
[0107] (2.3) Multi-layer stacking: The Transformer is composed of multiple layers of self-attention and feed-forward neural networks stacked together. It processes the input feature sequence layer by layer to generate high-level feature representations.
[0108] (2.4) Classification layer: Input the output feature representation of the Transformer into the classification layer for classification judgment and output the probability of being pornographic.
[0109] (3) Model training and optimization;
[0110] (3.1) Data preparation: Collect a large number of pornographic and non-pornographic pictures, audio, and video clips, annotate them to form a training data set and a validation data set.
[0111] (3.2) Transfer learning: Based on the pre-trained ResNet50, VGGish, and Transformer models, use the transfer learning method to fine-tune and train the classification model.
[0112] (3.3) Loss function: Adopt the cross-entropy loss function Cross-Entropy Loss to measure the difference between the model output and the true label, and guide the optimization of model parameters.
[0113] (3.4) Optimization algorithm: Adopt the Adam optimization algorithm. Through backpropagation and gradient descent, iteratively update the model parameters to improve the classification accuracy.
[0114] Through the above processes of feature extraction, classification and identification, and model training and optimization, the present invention can efficiently and accurately identify whether the pictures, audio, and video content uploaded by users is pornographic, effectively improving the efficiency and accuracy of content review.
[0115] S4. The result storage module stores the identification result and relevant log information;
[0116] The specific content includes:
[0117] Identification result: two legal values, pornographic / non-pornographic;
[0118] Metadata of the uploaded content: user ID, upload time or download time;
[0119] Identification time and model version.
[0120] S5. The alarm module triggers the alarm mechanism when detecting pornographic content;
[0121] Specific measures include:
[0122] a) Sending an alarm notification to the management staff (such as email, text message, etc.)
[0123] b) Generating an alarm record in the management background.
[0124] Based on the above method, an apparatus for identifying pornographic content in object storage in this embodiment includes: at least one memory and at least one processor;
[0125] The at least one memory is used to store machine-readable programs;
[0126] The at least one processor is used to call the machine-readable program to execute a method for identifying pornographic content in object storage.
[0127] The above specific implementation manners are only specific cases of the present invention. The patent protection scope of the present invention includes but is not limited to the above specific implementation manners. Any technical solution that conforms to the above specific implementation manners recorded in the present invention and any appropriate changes or substitutions made by those of ordinary skill in the art shall fall within the patent protection scope of the present invention.
[0128] Although the embodiments of the present invention have been shown and described, for those of ordinary skill in the art, it can be understood that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A method for identifying pornographic content in object storage, characterized in that, It has the following steps: S1. The object storage module receives the user's upload and download requests, converts the data provided by the user into an unstructured data stream and stores it in the disk of the cloud device, or converts the unstructured data stream in the disk of the cloud device into the data format required by the user and sends it to the user; S2. The preprocessing module processes the pictures, audio, and video content uploaded and downloaded by the user; S3. The large model discrimination module discriminates against pornographic content in the preprocessed content; S4. The result storage module stores the discrimination results and relevant log information; S5. The alarm module triggers the alarm mechanism when pornographic content is detected.
2. The method for identifying pornographic content in object storage according to claim 1, wherein In step S1, during the conversion process, pornographic content discrimination operations are performed on data in the picture, audio, and video formats.
3. The object storage content pornographic identification method according to claim 2, characterized in that, In step S2, for picture preprocessing, first convert all pictures to the JPEG format, scale the pictures to a fixed size of 224x224 pixels, use the interpolation method for scaling, and adopt data augmentation techniques to generate diverse training samples; For audio preprocessing, first convert all audio files to the WAV format, standardize the sampling rate of all audio files to a unified value of 16kHz, divide the audio signal into several frames, with each frame having a fixed length of 25ms and a 10ms overlap between adjacent frames to capture the temporal characteristics of the audio; For video preprocessing, extract key frames from the video file at a fixed interval of 1 frame per second, convert the extracted key frames to the unified JPEG format, scale the key frames to a fixed size of 224x224 pixels, and perform data augmentation processing on the key frames.
4. The object storage content pornographic identification method according to claim 3, characterized in that, In step S3, it includes: (1) Feature extraction, including picture feature extraction, audio feature extraction, and video feature extraction; (2) Classification and discrimination, using a large model based on Transformer for classification and discrimination. The Transformer model processes the input sequence data through the self-attention mechanism to capture local and global information in pictures, audio, and video; (3) Model training and optimization.
5. The object storage content pornographic identification method according to claim 4, characterized in that, In step (1), in the picture feature extraction, a pre-trained ResNet50 model is used for picture feature extraction. The problem of gradient disappearance in deep networks is solved through residual connections, enabling the model to be trained more effectively. Specifically as follows: (1-1.1) Input layer: The preprocessed picture is input into the input layer of ResNet50 for convolution operations; (1-1.2) Convolutional layer and residual blocks: ResNet50 contains multiple convolutional layers and residual blocks. Through multiple convolutions and non-linear transformations of these layers, high-level features of the picture are extracted layer by layer; (1-1.3) Global average pooling layer: Perform global average pooling on the feature map output by the convolutional layer to obtain a feature vector of a fixed length; (1-1.4) Feature vector output: The finally output feature vector represents the high-level semantic information of the picture and is used for subsequent classification tasks; In the audio feature extraction, a VGGish model is used for audio feature extraction. By converting the audio signal into a spectrogram and then inputting it into a pre-trained VGG model to extract features, specifically as follows: (1-2.1) Spectrogram generation: Convert the audio signal of each frame into a spectrogram. Use the short-time Fourier transform (STFT) to convert the time-domain signal into a frequency-domain signal and generate a two-dimensional spectrogram. (1-2.2) Input to the VGGish model: Input the spectrogram into the pre-trained VGGish model for convolution operations and feature extraction. (1-2.3) Convolutional layers and pooling layers: The VGGish model contains multiple convolutional layers and pooling layers. Through multiple convolution and pooling operations of these layers, high-level features of the audio are extracted layer by layer. (1-2.4) Output of the feature vector: The finally output feature vector represents the high-level semantic information of the audio and is used for subsequent classification tasks. In the video feature extraction, use the pre-trained ResNet50 model to extract features from the video key frames, and combine the Transformer model to capture the temporal features of the video. Specifically as follows: (1-3.1) Input layer: The pre-processed key frames are input into the input layer of ResNet50 for convolution operations. (1-3.2) Convolutional layers and residual blocks: ResNet50 extracts high-level features of the key frames layer by layer through multiple convolution and non-linear transformation operations of multiple convolutional layers and residual blocks. (1-3.3) Global average pooling layer: Perform global average pooling on the feature map output by the convolutional layer to obtain a feature vector of a fixed length. (1-3.4) Output of the feature vector: The finally output feature vector represents the high-level semantic information of the key frames. (1-3.5) Sequence feature combination: Combine the feature vectors of all key frames of the same video in chronological order to form a feature sequence for subsequent classification tasks.
6. The object storage content pornographic identification method according to claim 5, characterized in that In step (2), a large model based on Transformer is used for classification and discrimination. The Transformer model processes the input sequence data through the self-attention mechanism and can capture local and global information in pictures, audio, and video. The specific classification and discrimination process is as follows: (2.1) Input of the feature sequence: Input the feature vectors of pictures, audio, and video into the Transformer model to form a feature sequence; for pictures and audio, each input feature vector is processed separately; for video, combine the feature vectors of all key frames of the same video in chronological order to form a feature sequence. (2.2) Self-attention calculation: Transformer calculates the correlation between each position in the feature sequence and other positions through the self-attention mechanism to generate a weighted feature representation. (2.3) Multi-layer stacking: Transformer is stacked by multiple self-attention and feed-forward neural networks, and processes the input feature sequence layer by layer to generate high-level feature representations. (2.4) Classification layer: Input the output feature representation of Transformer into the classification layer for classification judgment and output the probability of being pornographic.
7. The object storage content pornographic identification method according to claim 6, characterized in that, In step (3), it includes: (3.1) Data preparation: Collect a large number of pornographic and non-pornographic pictures, audio, and video clips, perform annotation to form a training data set and a validation data set. (3.2) Transfer learning: Based on the pre-trained ResNet50, VGGish, and Transformer models, fine-tune through transfer learning to train a classification model; (3.3) Loss function: Use the cross-entropy loss function Cross-Entropy Loss to measure the difference between the model output and the true label, and guide the optimization of model parameters; (3.4) Optimization algorithm: Use the Adam optimization algorithm to iteratively update model parameters through backpropagation and gradient descent.
8. The method for identifying pornographic content in object storage according to claim 7, characterized in that, In step S4, it includes: Identification results: Two legal values, pornographic or non-pornographic; Metadata of the uploaded content: User ID, upload time, or download time; Identification time and model version.
9. The method for identifying pornographic content in object storage according to claim 8, characterized in that, In step S5, it includes: a) Send an alarm notification to the administrator; b) Generate an alarm record in the management background.
10. An object storage content pornographic identification device, characterized in that, It includes: At least one memory and at least one processor; The at least one memory is used to store machine-readable programs; The at least one processor is used to call the machine-readable program and execute the method described in any one of claims 1 to 9.