Content auditing method based on deep learning

Through the content review method based on deep learning, a multimodal data processing model is constructed, and public opinion information is automatically collected and analyzed, which solves the problems of incomplete monitoring, unprofessional response and untimely handling in the existing technology, and achieves efficient and professional public opinion monitoring and handling.

CN120011964APending Publication Date: 2025-05-16SUZHOU LIANRUIKE ELECTRONIC TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510026352.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-08
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

When facing massive information, the existing technology has problems such as incomplete monitoring, unprofessional response and untimely handling.

Method used

Using a content audit method based on deep learning, by selecting appropriate deep learning frameworks (such as TensorFlow, PyTorch), a deep learning model with multimodal data processing capabilities is designed and constructed, public opinion information is automatically collected and preprocessed, features are extracted and reviewed results are generated, and public opinion analysis reports are automatically generated and pushed to relevant management personnel.

Benefits of technology

It has realized comprehensive monitoring of various types of public opinion information such as text, images, and videos, which has improved the comprehensiveness and professionalism of monitoring, shortened the cycle of public opinion analysis, improved the timeliness of disposal, and reduced the situation of false alarms and missed reports.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120011964A_ABST
    Figure CN120011964A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of artificial intelligence, in particular to a content auditing method based on deep learning, and the method comprises the steps: firstly, selecting a proper deep learning framework according to the demands of public opinion analysis, and designing and constructing a deep learning model with a multi-modal data processing capability; then, public opinion information is automatically collected from various social media and news website platforms, and the collected data is preprocessed; feature extraction is carried out on the preprocessed data; inputting the extracted feature vectors into a deep learning model for operation, and outputting an auditing result; and finally, according to an output result of the deep learning model, automatically generating a public opinion analysis report, and pushing the public opinion analysis report to related managers, so that the technical problems that an auditing mode in the prior art looks careful when facing massive information, monitoring is not comprehensive, coping is not professional and disposal is not timely are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a content review method based on deep learning. Background Art

[0002] With the rapid development of Internet technology, social media, news websites, forums, blogs and other online platforms have become an indispensable source of information in people's daily lives. The amount of information on these platforms has exploded, covering multiple fields such as news, entertainment, education, and social interaction. However, with the surge in the amount of information, the Internet is also filled with a large amount of illegal, false, violent, pornographic and other harmful information. The spread of harmful information on the Internet is extremely fast. Once released, it is likely to spread rapidly in a short period of time, causing adverse social impacts. Therefore, it is particularly important to deal with harmful information in a timely manner.

[0003] However, traditional methods of Internet content review seem to be unable to cope with massive amounts of information, and there are problems such as incomplete monitoring, unprofessional response, and untimely disposal. Summary of the invention

[0004] The purpose of the present invention is to provide a content review method based on deep learning, which aims to solve the technical problems that the review methods in the prior art are unable to cope with massive amounts of information, and have the problems of incomplete monitoring, unprofessional response and untimely disposal.

[0005] To achieve the above purpose, the present invention adopts a content review method based on deep learning, which includes the following steps:

[0006] Step 1: According to the needs of public opinion analysis, select a suitable deep learning framework (such as TensorFlow, PyTorch, etc.), and design and build a deep learning model with multimodal data processing capabilities;

[0007] Step 2: Automatically collect public opinion information from major social media and news website platforms, and pre-process the collected data;

[0008] Step 3: Extract features from the preprocessed data;

[0009] Step 4: Input the extracted feature vector into the deep learning model for calculation and output the audit result;

[0010] Step 5: Based on the output results of the deep learning model, a public opinion analysis report is automatically generated and pushed to relevant managers.

[0011] Among them, the constructed deep learning model is built based on the TensorFlow or PyTorch framework and adopts a multimodal data processing architecture.

[0012] Among them, in step 2, when automatically collecting public opinion information from major social media and news website platforms, crawler technology or API interface is used to automatically collect public opinion information from major social media and news website platforms. The collected information includes text comments, picture sharing, and video uploads;

[0013] Specifically, a crawler program is developed using programming languages ​​such as Python to access the target website by simulating user behavior (such as sending HTTP requests) and crawling the required information; the crawler program needs to handle anti-crawler mechanisms such as verification code verification, IP blocking, etc.

[0014] For platforms that provide API interfaces, you can obtain data by calling the API; this usually requires applying for an API key, setting request parameters, and parsing the data returned by the API.

[0015] Finally, the collected data is stored in a local database or cloud storage for subsequent processing and analysis.

[0016] When the collected data is stored in a local database or cloud storage, a relational database (such as MySQL, Oracle, etc.) or a non-relational database (such as MongoDB, Cassandra, etc.) is used to store the data.

[0017] Wherein, in step 2, preprocessing the data includes removing duplicate data, processing missing values, and removing advertisements or junk information;

[0018] Among them, deduplication uses a hash algorithm or a unique constraint of a database to detect and remove duplicate data;

[0019] For missing data, fill-in methods (such as mean filling, mode filling), interpolation methods, or deletion methods are used for processing;

[0020] Removing advertisements or spam information uses regular expressions, keyword matching or machine learning algorithms (such as Naive Bayes, Support Vector Machine, etc.) to identify and remove advertisements or spam information.

[0021] Among them, in step three, feature extraction is performed on the preprocessed data to convert the preprocessed data (including text, pictures, videos, etc.) into feature vectors that can be understood and processed by the deep learning model, where:

[0022] For text data, use word embedding technology (such as Word2Vec, BERT, etc.) to convert text into vector representation;

[0023] For image data, models such as convolutional neural networks (CNN) are used to extract visual features of images;

[0024] For video data, image processing and time series analysis techniques are combined to extract features.

[0025] In step 4, the audit results output include content classification labels, sentiment tendency scores, and sensitive information labels;

[0026] After outputting the audit results, the output results of the model are verified to ensure their accuracy and reliability.

[0027] Among them, the specific implementation method of automatically generating a public opinion analysis report based on the output results of the deep learning model and pushing it to relevant managers is:

[0028] Design templates for public opinion analysis reports;

[0029] Receive output results from deep learning models and parse them to extract key information;

[0030] Based on the designed report template, fill in the extracted key information data into the corresponding positions;

[0031] Use automation tools (e.g., Python scripts, report generation software, etc.) to convert the filled report template into a readable document format (e.g., PDF, Word, etc.);

[0032] The report is then sent to relevant managers via email, internal corporate communication platforms, etc.

[0033] Among them, during the process of data collection, storage and processing, ensure compliance with relevant laws, regulations and privacy policies to protect the security and privacy of user data.

[0034] Among them, after the model design is completed, the labeled data set is used for model training and verification to evaluate the performance and accuracy of the model;

[0035] When using labeled data sets for model training and validation, the validation method adopts the cross-validation method: by comparing the performance on the training set and the validation set, check whether the model has overfitting (performing well on the training set and poorly on the validation set) or underfitting (performing poorly on both the training set and the validation set) problems.

[0036] A content review method based on deep learning of the present invention first selects a suitable deep learning framework according to the needs of public opinion analysis, designs and constructs a deep learning model with multimodal data processing capabilities; then automatically collects public opinion information from major social media and news website platforms, and preprocesses the collected data; then extracts features from the preprocessed data; then inputs the extracted feature vectors into the deep learning model for calculation, and outputs the review results; finally, according to the output results of the deep learning model, automatically generates a public opinion analysis report, and pushes it to relevant management personnel, thereby solving the technical problems that the review method in the prior art is unable to cope with massive amounts of information, and has the problems of incomplete monitoring, unprofessional response and untimely disposal.

[0037] By selecting a suitable deep learning framework and designing a deep learning model with multimodal data processing capabilities, the present invention can achieve comprehensive monitoring of multiple types of public opinion information such as text, images, and videos. This multimodal data processing capability enables the system to more accurately capture and understand complex information in the network, thereby significantly improving the comprehensiveness of monitoring;

[0038] Advanced deep learning algorithms are used in the feature extraction stage to automatically extract features that are valuable for public opinion analysis from raw data. These features include not only keywords and sentiment in the text, but also object recognition and scene understanding in the image. Through the operation of deep learning models, the system can output more accurate and professional audit results, thereby enhancing the professionalism of the response.

[0039] The automatic collection, preprocessing, feature extraction and model calculation of public opinion information are realized, and the whole process is highly automated; the cycle of public opinion analysis is greatly shortened. Once the system detects potential public opinion risks, it can immediately generate a public opinion analysis report and push it to relevant managers; this efficient handling process enables managers to respond quickly and take necessary measures to deal with public opinion events, thereby improving the timeliness of handling.

[0040] Through continuous learning and optimization, the deep learning model can gradually adapt to various complex public opinion scenarios and reduce false alarms and missed alarms. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0042] Figure 1It is a step flow chart of the content review method based on deep learning of the present invention. DETAILED DESCRIPTION

[0043] Embodiments of the present invention are described in detail below. Examples of the embodiments are shown in the accompanying drawings. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to be used to explain the present invention, but should not be construed as limiting the present invention.

[0044] See also Figure 1 , Figure 1 It is a step flow chart of the content review method based on deep learning of the present invention.

[0045] The present invention provides a content review method based on deep learning, comprising the following steps:

[0046] S1. According to the needs of public opinion analysis, select a suitable deep learning framework (such as TensorFlow, PyTorch, etc.), and design and build a deep learning model with multimodal data processing capabilities;

[0047] For this specific implementation, the constructed deep learning model is built based on the TensorFlow or PyTorch framework and adopts a multimodal data processing architecture.

[0048] S2, automatically collect public opinion information from major social media and news website platforms, and pre-process the collected data;

[0049] For this specific implementation, when automatically collecting public opinion information from major social media and news website platforms, crawler technology or API interface is used to automatically collect public opinion information from major social media and news website platforms. The collected information includes text comments, picture sharing, and video uploads;

[0050] Specifically, a crawler program is developed using programming languages ​​such as Python to access the target website by simulating user behavior (such as sending HTTP requests) and crawling the required information; the crawler program needs to handle anti-crawler mechanisms such as verification code verification, IP blocking, etc.

[0051] For platforms that provide API interfaces, you can obtain data by calling the API; this usually requires applying for an API key, setting request parameters, and parsing the data returned by the API.

[0052] Finally, the collected data is stored in a local database or cloud storage for subsequent processing and analysis.

[0053] When the collected data is stored in a local database or cloud storage, a relational database (such as MySQL, Oracle, etc.) or a non-relational database (such as MongoDB, Cassandra, etc.) is used to store the data.

[0054] Preprocessing the data includes removing duplicate data, handling missing values, and removing advertisements or spam information;

[0055] Among them, deduplication uses a hash algorithm or a unique constraint of a database to detect and remove duplicate data;

[0056] For missing data, fill-in methods (such as mean filling, mode filling), interpolation methods, or deletion methods are used for processing;

[0057] Removing advertisements or spam information uses regular expressions, keyword matching or machine learning algorithms (such as Naive Bayes, Support Vector Machine, etc.) to identify and remove advertisements or spam information.

[0058] S3, extracting features from the preprocessed data;

[0059] For this specific implementation, feature extraction is performed on the preprocessed data to convert the preprocessed data (including text, pictures, videos, etc.) into feature vectors that can be understood and processed by the deep learning model, where:

[0060] For text data, use word embedding technology (such as Word2Vec, BERT, etc.) to convert text into vector representation;

[0061] For image data, models such as convolutional neural networks (CNN) are used to extract visual features of images;

[0062] For video data, image processing and time series analysis techniques are combined to extract features.

[0063] S4, input the extracted feature vector into the deep learning model for calculation, and output the audit result;

[0064] For this specific implementation, the audit results output include content classification labels, sentiment tendency scores, and sensitive information identification;

[0065] After outputting the audit results, the output results of the model are verified to ensure their accuracy and reliability.

[0066] S5. Based on the output results of the deep learning model, a public opinion analysis report is automatically generated and pushed to relevant managers.

[0067] Design a template for public opinion analysis report for this specific implementation method;

[0068] Receive output results from deep learning models and parse them to extract key information;

[0069] Based on the designed report template, fill in the extracted key information data into the corresponding positions;

[0070] Use automation tools (e.g., Python scripts, report generation software, etc.) to convert the filled report template into a readable document format (e.g., PDF, Word, etc.);

[0071] The report is then sent to relevant managers via email, internal corporate communication platforms, etc.

[0072] During the process of data collection, storage and processing, ensure compliance with relevant laws, regulations and privacy policies to protect the security and privacy of user data.

[0073] Among them, after the model design is completed, the labeled data set is used for model training and verification to evaluate the performance and accuracy of the model;

[0074] When using labeled data sets for model training and validation, the validation method adopts the cross-validation method: by comparing the performance on the training set and the validation set, check whether the model has overfitting (performing well on the training set and poorly on the validation set) or underfitting (performing poorly on both the training set and the validation set) problems.

[0075] Specifically, model training:

[0076] Data preparation:

[0077] The labeled data set is divided into a training set and a validation set (sometimes a test set is also required). The training set is used to train the model, and the validation set is used to adjust the model parameters and evaluate the model performance.

[0078] The dataset should contain multiple types of data (such as text, images, etc.) to simulate the multimodal data environment in the real world.

[0079] Model initialization:

[0080] Initialize the model's weights and biases randomly, or use a pretrained model to speed up the training process.

[0081] Training process:

[0082] Forward propagation: The input data is calculated through each layer of the model and finally the output result is obtained.

[0083] Loss function: Calculates the difference between the output result and the true label. Commonly used loss functions include mean square error (MSE), cross-entropy loss, etc.

[0084] Backpropagation: Update the weights and biases in the model using the chain rule based on the gradient of the loss function.

[0085] Optimization algorithms: such as Gradient Descent and Adam, which are used to adjust network parameters during training to minimize the loss function.

[0086] Hyperparameter Tuning:

[0087] Adjust hyperparameters such as learning rate, batch size, number of iterations, etc. through grid search, random search, or Bayesian optimization to improve the performance of the model.

[0088] 2. Model Verification:

[0089] Performance evaluation indicators:

[0090] Accuracy: The ratio of the number of correctly predicted samples to the total number of samples.

[0091] Precision: The ratio of the number of correctly predicted positive samples to the number of all predicted positive samples.

[0092] Recall: The ratio of the number of correctly predicted positive samples to the number of all actual positive samples.

[0093] F1 score: The harmonic average of precision and recall, used to comprehensively evaluate the performance of the model.

[0094] ROC curve and AUC value: The ROC curve is a graph of the true positive rate (TPR) against the false positive rate (FPR), and the AUC value (area under the ROC curve) measures the model's ability to distinguish between positive and negative samples.

[0095] Cross Validation:

[0096] The training dataset is divided into multiple subsets, and the model is trained and validated on each subset to reduce the randomness of the validation set and obtain a more accurate estimate of the model performance.

[0097] Overfitting and underfitting check:

[0098] By comparing the performance on the training set and the validation set, check whether the model is overfitting (performing well on the training set and poorly on the validation set) or underfitting (performing poorly on both the training set and the validation set).

[0099] Model selection and tuning:

[0100] Based on the performance evaluation results on the validation set, the best model architecture and hyperparameter configuration are selected.

[0101] Further tuning of the model, such as using data augmentation techniques, regularization methods, learning rate scheduling, etc., can improve the generalization ability of the model.

[0102] A content review method based on deep learning in this embodiment is used. First, according to the needs of public opinion analysis, a suitable deep learning framework is selected, and a deep learning model with multimodal data processing capabilities is designed and constructed; then, public opinion information is automatically collected from major social media and news website platforms, and the collected data is preprocessed; then, feature extraction is performed on the preprocessed data; the extracted feature vector is then input into the deep learning model for calculation, and the review result is output; finally, according to the output result of the deep learning model, a public opinion analysis report is automatically generated and pushed to relevant management personnel, thereby solving the technical problems that the review method in the prior art is unable to cope with massive information, and there are incomplete monitoring, unprofessional response and untimely disposal.

[0103] By selecting a suitable deep learning framework and designing a deep learning model with multimodal data processing capabilities, the present invention can achieve comprehensive monitoring of multiple types of public opinion information such as text, images, and videos. This multimodal data processing capability enables the system to more accurately capture and understand complex information in the network, thereby significantly improving the comprehensiveness of monitoring;

[0104] Advanced deep learning algorithms are used in the feature extraction stage to automatically extract features that are valuable for public opinion analysis from raw data. These features include not only keywords and sentiment in the text, but also object recognition and scene understanding in the image. Through the operation of deep learning models, the system can output more accurate and professional audit results, thereby enhancing the professionalism of the response.

[0105] The automatic collection, preprocessing, feature extraction and model calculation of public opinion information are realized, and the whole process is highly automated; the cycle of public opinion analysis is greatly shortened. Once the system detects potential public opinion risks, it can immediately generate a public opinion analysis report and push it to relevant managers; this efficient handling process enables managers to respond quickly and take necessary measures to deal with public opinion events, thereby improving the timeliness of handling.

[0106] Through continuous learning and optimization, the deep learning model can gradually adapt to various complex public opinion scenarios and reduce false alarms and missed alarms.

[0107] What is disclosed above is only a preferred embodiment of the present invention, and it certainly cannot be used to limit the scope of rights of the present invention. Ordinary technicians in this field can understand that all or part of the processes of the above embodiment and equivalent changes made according to the claims of the present invention still fall within the scope of the invention.

Claims

1. A content review method based on deep learning, characterized in that: The steps include: Step 1: According to the needs of public opinion analysis, select a suitable deep learning framework, design and build a deep learning model with multimodal data processing capabilities; Step 2: Automatically collect public opinion information from major social media and news website platforms, and pre-process the collected data; Step 3: Extract features from the preprocessed data; Step 4: Input the extracted feature vector into the deep learning model for calculation and output the audit result; Step 5: Based on the output results of the deep learning model, a public opinion analysis report is automatically generated and pushed to relevant managers.

2. The content review method based on deep learning as claimed in claim 1, characterized in that: The constructed deep learning model is built based on the TensorFlow or PyTorch framework and adopts a multimodal data processing architecture.

3. The content review method based on deep learning as claimed in claim 2, characterized in that: In step 2, when automatically collecting public opinion information from major social media and news website platforms, crawler technology or API interfaces are used to automatically collect public opinion information from major social media and news website platforms. The collected information includes text comments, picture sharing, and video uploads; Specifically, a crawler program is developed using programming languages ​​such as Python to access the target website by simulating user behavior and crawling the required information; For platforms that provide API interfaces, data can be obtained by calling the API; Finally, the collected data is stored in a local database or cloud storage for subsequent processing and analysis.

4. The content review method based on deep learning as claimed in claim 3, characterized in that: When the collected data is stored in a local database or cloud storage, a relational database or a non-relational database is used to store the data.

5. The content review method based on deep learning as claimed in claim 4, characterized in that: In step 2, data is preprocessed including removing duplicate data, processing missing values, and removing advertisements or junk information; Among them, deduplication uses a hash algorithm or a unique constraint of a database to detect and remove duplicate data; For missing data, fill-in, interpolation or deletion methods are used to handle them; Remove ads or spam Use regular expressions, keyword matching, or machine learning algorithms to identify and remove ads or spam.

6. The content review method based on deep learning as claimed in claim 5, characterized in that: In step three, feature extraction is performed on the preprocessed data to convert the preprocessed data into a feature vector that can be understood and processed by the deep learning model, where: For text data, word embedding technology is used to convert text into vector representation; For image data, models such as convolutional neural networks are used to extract visual features of images; For video data, image processing and time series analysis techniques are combined to extract features.

7. The content audit method based on deep learning as claimed in claim 6, characterized in that: In step 4, the audit results output include content classification labels, sentiment tendency scores, and sensitive information labels; After outputting the audit results, the output results of the model are verified to ensure their accuracy and reliability.

8. The content audit method based on deep learning as claimed in claim 7, characterized in that: The specific implementation method of automatically generating a public opinion analysis report based on the output results of the deep learning model and pushing it to relevant managers is as follows: Design templates for public opinion analysis reports; Receive output results from deep learning models and parse them to extract key information; Based on the designed report template, fill in the extracted key information data into the corresponding positions; Utilize automated tools to convert the populated report templates into a readable document format; The report is then sent to relevant managers via email, internal corporate communication platforms, etc.

9. The content audit method based on deep learning as claimed in claim 8, characterized in that: During the process of data collection, storage and processing, ensure compliance with relevant laws, regulations and privacy policies to protect the security and privacy of user data.

10. The content audit method based on deep learning as claimed in claim 9, characterized in that: After the model design is completed, the labeled data set is used for model training and validation to evaluate the performance and accuracy of the model; When using labeled data sets for model training and validation, the validation method adopts cross-validation: by comparing the performance on the training set and the validation set, check whether the model has overfitting or underfitting problems.