Malicious comment real-time detection system based on deep learning

By designing a real-time malicious comment detection system based on deep learning, the problems of low efficiency and poor accuracy of malicious comment recognition in the existing technology are solved, and efficient and accurate malicious comment recognition and real-time detection capabilities are achieved.

CN120196799APending Publication Date: 2025-06-24EAST CHINA UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510254187.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-05
Publication Date
2025-06-24

AI Technical Summary

Technical Problem

The existing technology is difficult to effectively identify and detect complex and changeable malicious comments in real time. The traditional methods are inefficient and have poor accuracy, and cannot meet the needs of massive data processing.

Method used

A real-time detection system for malicious comments based on deep learning is designed, including data acquisition module, preprocessing module, deep learning model construction module, model training module, real-time detection module and feedback module. By dynamically adapting to the data interfaces of different platforms, deeply cleaning comment data, building convolutional neural network and recurrent neural network models, detecting and feedback malicious comments in real time.

Benefits of technology

It realizes efficient and accurate malicious comment recognition, can complete single comment detection in milliseconds, meet real-time requirements, dynamically optimize model performance, and continuously improve recognition accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120196799A_ABST
    Figure CN120196799A_ABST
Patent Text Reader

Abstract

The invention discloses a malicious comment real-time detection system based on deep learning. The system comprises a data acquisition module, a preprocessing module, a deep learning model construction module, a model training module, a real-time detection module and a feedback module. Efficient and accurate comment acquisition: the data acquisition module has strong adaptation capability for different types of network platforms, such as social media, e-commerce, news information platforms and the like; by using an open AP I, simulating browser behaviors and combining an RSS feed source and a crawler technology, comment data can be captured from various sources comprehensively, continuously and stably, it is ensured that the data source is wide and comprehensive, and any potential malicious comment information is not missed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a real-time detection system, and specifically refers to a real-time malicious comment detection system based on deep learning. Background Art

[0002] With the rapid development of the Internet, online platforms have shown explosive growth. Various online communication venues such as social media, e-commerce platforms, and news information platforms have become an indispensable part of people's lives, work, and social interactions. A large number of users freely post comments on these platforms, sharing opinions, experiences, and feelings, which greatly promotes information circulation and interaction. However, this openness has also given rise to a large number of malicious comment problems.

[0003] On the one hand, the forms of malicious comments have become increasingly diverse and complex. From traditional insults and slander of others to today's malicious expressions disguised with Internet buzzwords, veiled metaphors, professional terms, and even cross-field knowledge, the traditional detection methods relying on simple keyword matching and fixed rule judgment are stretched to the limit. For example, in the electronic product review area, malicious competing merchants may use seemingly professional technical analysis terms to actually imply product defects and mislead consumers; under the discussion of hot topics on social media, some users skillfully use homophonic puns and abbreviations to post inciting or offensive remarks to avoid conventional supervision.

[0004] On the other hand, the rapid iteration and differential development of online platforms have increased the detection difficulty. Different platforms have their own unique interface designs, data interaction rules, and user group characteristics. Taking the API interface as an example, social media platforms frequently update interface permissions and data return formats, most e-commerce platforms restrict external direct access to data, and news information platforms focus on changes in content push logic, which makes it difficult to collect data uniformly and stably. At the same time, the amount of comment data generated by platform users has increased exponentially. For example, thousands of comments can be aggregated under a popular social media post in a short time, posing a severe challenge to the real-time performance and processing ability of the detection system.

[0005] In this context, the traditional method based on manual review is extremely inefficient, unable to meet the needs of processing massive data, and is easily affected by subjective factors, making it difficult to unify the judgment criteria; the automated detection tools based on simple rules lack adaptability and accuracy. To effectively respond to the complex and changing situation of malicious comments and accurately and real-time purify the network environment, an innovative solution is urgently needed. Summary of the Invention

[0006] The technical problem to be solved by the present invention is to overcome the above-mentioned technical defects and provide a real-time malicious comment detection system based on deep learning.

[0007] To solve the above technical problems, the technical solution provided by the present invention is a real-time malicious comment detection system based on deep learning: a data collection module, which is used to collect comment data in real time from multiple network platforms, and the network platforms include but are not limited to social media platforms, e-commerce platforms, and news information platforms. This module has the ability to dynamically adapt to the data interfaces of different platforms and can automatically adjust the collection strategy according to changes in platform rules to ensure the continuous and stable acquisition of comment data;

[0008] A preprocessing module, connected to the data collection module, is used to clean the collected comment data, remove invalid characters, duplicate data, and garbled codes, and uniformly convert the text into a standard format. At the same time, the text is segmented and part-of-speech tagged to prepare for the input of the subsequent deep learning model. This module supports a variety of natural language processing algorithms and can automatically select the optimal algorithm combination according to data characteristics;

[0009] A deep learning model construction module constructs a deep learning model based on a convolutional neural network (CNN), a recurrent neural network (RNN) and its variants, or a combination of the two. The model includes an input layer, a hidden layer, and an output layer, where the hidden layer contains multiple neurons for automatically extracting the features of comment data. The hyperparameters of the model can be dynamically adjusted according to the training effect to optimize the model performance;

[0010] A model training module, connected to the deep learning model construction module, uses labeled malicious comment samples and normal comment samples to train the deep learning model. The training process adopts a batch training method, combined with optimization algorithms such as stochastic gradient descent, to continuously adjust the model weights to improve the recognition accuracy of the model for malicious comments. The sample library is updated regularly to include samples of newly emerging malicious comment types;

[0011] A real-time detection module, connected to the model training module, inputs the preprocessed real-time comment data into the trained deep learning model, and the model outputs a determination result of whether the comment is a malicious comment or a normal comment. This module has the characteristic of low latency and can complete the detection of a single comment within milliseconds to meet the real-time requirement;

[0012] A feedback module, connected to the real-time detection module, when a malicious comment is detected, on the one hand, the result is fed back to the corresponding network platform so that the platform can take measures such as blocking and warning, and on the other hand, the malicious comment sample and its feature information are fed back to the model training module for further optimization of the model.

[0013] As an improvement, the convolutional neural network in the deep learning model construction module adopts a multi-layer convolutional structure, captures local and global features of the text through convolutional kernels of different sizes, and a pooling layer is connected after the convolutional layer to reduce the dimension of the feature map and reduce the amount of calculation.

[0014] As an improvement, the recurrent neural network in the deep learning model construction module adopts a long short-term memory network (LSTM) or a gated recurrent unit (GRU), which uses its memory unit to effectively process the long-term and short-term dependencies in the text sequence, avoiding the problems of gradient disappearance or gradient explosion.

[0015] As an improvement, during the word segmentation process, the preprocessing module combines a domain dictionary and a pre-trained word vector model to improve the accuracy of word segmentation, especially for the processing of special words such as new network words, technical terms, and abbreviations.

[0016] As an improvement, during the training process, the model training module adopts a cross-validation method, divides the samples into multiple subsets, and alternately uses one subset as the validation set and the remaining subsets as the training set to comprehensively evaluate the model performance and prevent overfitting.

[0017] As an improvement, the real-time detection module supports distributed deployment and can process comment data in parallel on multiple server nodes, further improving the real-time detection ability and throughput of the system.

[0018] As an improvement, when the feedback module feeds back the malicious comment results to the network platform, it attaches detailed malicious classification bases, such as abusive words, sensitive topics, and malicious intent categories, to help the platform manage precisely.

[0019] The advantages of the present invention compared with the prior art are as follows: Efficient and accurate comment collection:

[0020] The data collection module has strong adaptability to different types of network platforms, such as social media, e-commerce, news information platforms, etc. By using open APIs, simulating browser behaviors, and combining RSS feeds and crawler technologies, it can comprehensively and continuously and stably capture comment data from various sources, ensuring that the data sources are extensive and comprehensive without missing any potential malicious comment information.

[0021] At the same time, the set preliminary screening rules can quickly remove invalid data and only store valuable comments in a structured format in the temporary cache database, which not only saves storage space but also provides high-quality data input for subsequent modules, greatly improving the data processing efficiency of the entire system and guaranteeing the accuracy basis of detection from the source.

[0022] Optimized text preprocessing:

[0023] The preprocessing module uses regular expressions, character encoding detection libraries, and advanced natural language processing toolkits to deeply clean the comment data, remove invalid characters such as HTML tags, emojis, and garbled codes, standardize the text, effectively reducing data noise and making the text data processed by subsequent models cleaner and more standardized.

[0024] Combined with custom domain dictionaries for word segmentation and part-of-speech tagging, it can accurately identify Internet buzzwords, professional terms, and platform abbreviations, so that the semantic structure of the text can be clearly presented. It provides more targeted and semantically rich feature information for deep learning models, significantly improves the model's ability to understand the semantics of complex texts, and thus improves the accuracy of identifying malicious comments.

[0025] Flexible and powerful model building and training:

[0026] The deep learning model building module can flexibly select or combine model architectures based on the characteristics of the comment data. Whether it is CNN that focuses on local features, RNN and its variants that are good at processing sequence dependencies, or a hybrid model that combines the two, it can accurately adapt to the needs of different scenarios and fully explore malicious feature patterns in the text.

[0027] The model training module relies on a large and diverse source of samples, covering a variety of malicious types. It adopts a scientific sample division strategy and cross-validation method, combined with automatic hyperparameter adjustment technology, to continuously optimize model performance, so that the model has a strong generalization ability, can accurately identify the endless new malicious comments, and effectively respond to the complex and changeable network language environment.

[0028] Excellent real-time detection performance:

[0029] The real-time detection module uses high-performance server clusters, multi-threading, asynchronous I / O technology, model quantization compression and GPU acceleration to achieve a single comment detection speed of milliseconds. Even under the impact of high-concurrency comment streams, it can ensure real-time response of the system, intercept malicious comments in time, and minimize the harm of their spread.

[0030] The distributed deployment solution further expands the system's throughput, linearly improves processing capabilities, and easily manages massive amounts of comment data, allowing real-time detection to run stably and efficiently in large-scale application scenarios, and comprehensively ensuring the healthy ecology of the comment area of ​​the network platform.

[0031] Dynamic optimization and platform collaboration:

[0032] On the one hand, the feedback module works closely with the network platform to provide real-time feedback on the details and classification of malicious comments, helping the platform to quickly take precise control measures such as blocking and warnings to purify the platform's comment environment; on the other hand, it feeds malicious comment samples and features back to the model training module, prompting the model to continuously learn new malicious patterns, dynamically optimize and upgrade, and continuously improve recognition accuracy, forming a virtuous cycle of system self-improvement mechanism, and always maintaining the ability to keenly capture and accurately strike malicious comments. BRIEF DESCRIPTION OF THE DRAWINGS

[0033] Figure 1It is a schematic diagram of a real-time malicious comment detection system based on deep learning according to the present invention. Detailed implementation manners

[0034] To make the objectives, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention. Components of the embodiments of the present invention generally described and illustrated in the drawings herein may be arranged and designed in a variety of different configurations.

[0035] In the description of the embodiments of the present invention, it should be noted that if terms such as "center", "upper", "lower", "left", "right", "vertical", "horizontal", "inner", "outer", etc. are used to indicate the orientation or positional relationship, it is based on the orientation or positional relationship shown in the drawings, or the orientation or positional relationship in which the product of the invention is usually placed during use. This is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and thus should not be construed as a limitation of the present invention. In addition, terms such as "first", "second", "third", etc. are only used for descriptive distinction and cannot be understood as indicating or implying relative importance.

[0036] In addition, if terms such as "horizontal", "vertical", "hanging" are used, it does not mean that the component is required to be absolutely horizontal or hanging, but can be slightly inclined. For example, "horizontal" only means that its direction is more horizontal relative to "vertical", and does not mean that the structure must be completely horizontal, but can be slightly inclined.

[0037] In the description of the embodiments of the present invention, "a plurality of" represents at least two.

[0038] In the description of the embodiments of the present invention, it should also be noted that unless otherwise clearly defined and limited, if terms such as "set", "installed", "connected", "connected" are used, they should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be directly connected, or indirectly connected through an intermediate medium, and can be the communication inside two elements. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific situations.

[0039] Combined with the attached drawings, a real-time malicious comment detection system based on deep learning, a data collection module, is used to collect comment data from multiple network platforms in real time. The network platforms include but are not limited to social media platforms, e-commerce platforms, and news information platforms. This module has the ability to dynamically adapt to the data interfaces of different platforms and can automatically adjust the collection strategy according to the changes in platform rules to ensure the continuous and stable acquisition of comment data;

[0040] A preprocessing module, connected to the data collection module, is used to clean the collected comment data, remove invalid characters, duplicate data, and garbled codes, and uniformly convert the text into a standard format. At the same time, the text is segmented and part-of-speech tagged to prepare for the input of the subsequent deep learning model. This module supports a variety of natural language processing algorithms and can automatically select the optimal algorithm combination according to data characteristics;

[0041] A deep learning model construction module constructs a deep learning model based on convolutional neural networks (CNNs), recurrent neural networks (RNNs) and their variants, or a combination of the two. The model includes an input layer, a hidden layer, and an output layer. The hidden layer contains multiple neurons for automatically extracting the features of comment data. The hyperparameters of the model can be dynamically adjusted according to the training effect to optimize the model performance;

[0042] A model training module, connected to the deep learning model construction module, uses labeled malicious comment samples and normal comment samples to train the deep learning model. The training process adopts a batch training method, combined with optimization algorithms such as stochastic gradient descent, to continuously adjust the model weights to improve the recognition accuracy of the model for malicious comments. The sample library is updated regularly to include samples of newly emerging malicious comment types;

[0043] A real-time detection module, connected to the model training module, inputs the preprocessed real-time comment data into the trained deep learning model. The model outputs a determination result of whether the comment is a malicious comment or a normal comment. This module has the characteristic of low latency and can complete the detection of a single comment within milliseconds to meet the real-time requirements;

[0044] A feedback module, connected to the real-time detection module, when a malicious comment is detected, on the one hand, the result is fed back to the corresponding network platform so that the platform can take measures such as blocking and warning, and on the other hand, the malicious comment sample and its feature information are fed back to the model training module for further optimization of the model.

[0045] The convolutional neural network in the deep learning model construction module adopts a multi-layer convolutional structure, captures local and global features of the text through convolutional kernels of different sizes, and a pooling layer is connected after the convolutional layer to reduce the dimension of the feature map and reduce the amount of calculation.

[0046] The recurrent neural network in the deep learning model construction module adopts a long short-term memory network (LSTM) or a gated recurrent unit (GRU), and uses its memory unit to effectively process the long-term and short-term dependencies in the text sequence, avoiding the problems of gradient disappearance or gradient explosion.

[0047] During the word segmentation process, the preprocessing module combines a domain dictionary and a pre-trained word vector model to improve the accuracy of word segmentation, especially for the processing of special words such as new network words, technical terms, and abbreviations.

[0048] During the training process, the model training module adopts a cross-validation method, divides the samples into multiple subsets, and takes turns using one subset as the validation set and the remaining subsets as the training set to comprehensively evaluate the model performance and prevent overfitting.

[0049] The real-time detection module supports distributed deployment and can process comment data in parallel on multiple server nodes, further improving the real-time detection ability and throughput of the system.

[0050] When the feedback module feeds back the malicious comment results to the network platform, it attaches detailed malicious classification bases, such as abusive words, sensitive topics, and malicious intention categories, to help the platform manage precisely.

[0051] I. Overview of System Architecture:

[0052] The overall architecture of this deep learning-based real-time malicious comment detection system consists of multiple closely collaborating modules, aiming to efficiently and accurately capture comment data from various network platforms and real-time determine whether they are malicious comments, providing strong support for purifying the network environment. Each module includes a data collection module, a preprocessing module, a deep learning model construction module, a model training module, a real-time detection module, and a feedback module. The following will elaborate on the specific implementation details of each module.

[0053] II. Implementation of Data Collection Module:

[0054] Interface Adaptation:

[0055] For social media platforms, use their open API interfaces to obtain access tokens according to the authorization process specified by the platform, thereby establishing a data connection channel. For example, for the Facebook platform, by registering a developer account, applying for corresponding permissions, and using the Graph API to capture comment data posted by users in public pages, groups, and other areas. At the same time, set a scheduled task to check the availability of the API interface and rule changes at regular time intervals (such as every 5 minutes). Once it is found that the interface rules are adjusted, such as changes in data request frequency limits or modifications to the returned data format, immediately start the adaptive program. This program parses the new rules through a predefined rule library and dynamically adjusts the data request parameters and parsing logic to ensure the uninterrupted collection process.

[0056] For e-commerce platforms, since most e-commerce platforms do not directly provide public APIs for large-scale external data collection, the method of simulating browser behavior is adopted. Tools such as Selenium are combined with ChromeDriver or FirefoxDriver to simulate operations such as user login, browsing product detail pages, and viewing comment sections. By parsing the HTML structure of the web page, comment elements are located and the text content is extracted. To cope with the frequent page structure updates of e-commerce platforms, a page template library is established, and the differences between the current page structure and the template are compared regularly. Once a change is detected, the location rules are automatically updated to ensure the stability of data collection.

[0057] In the case of news and information platforms, RSS subscription sources are combined with web crawler technology. For platforms that provide RSS subscriptions, the subscription sources are directly parsed to obtain comment links, and then the comment content is crawled; for platforms that do not support RSS, the Scrapy framework is used to customize crawlers, and XPath or CSS selector rules are written according to the characteristics of the platform's web page layout to accurately locate the comment area. Similarly, a monitoring mechanism is set up to track the changes in the platform's web page structure in real time and update the crawler rules in a timely manner.

[0058] Data screening and storage:

[0059] During the collection process, preliminary screening rules are set to filter out obviously invalid data, such as strings that are too short (less than 3 characters) or too long (more than 500 characters and without punctuation separation), or composed entirely of numbers or special symbols. These data usually do not contain valid comment information. The preliminarily screened data collected is stored in a temporary cache database in a structured format (such as JSON format, including fields such as comment content, publication time, publisher ID, and affiliated platform) in real time, such as Redis cache, so that subsequent modules can quickly read and process it. At the same time, a data expiration policy is set to avoid the accumulation of cached data. For example, data that has not been processed for more than 1 hour is automatically cleared.

[0060] III. Implementation of the preprocessing module:

[0061] Cleaning process:

[0062] After reading the comment data from the temporary cache database, character cleaning is first performed. Regular expressions are used to remove invalid characters such as HTML tags, emoji encodings (such as Unicode emoji codes), extra spaces, and line breaks. For duplicate data, by calculating the text hash value and establishing a hash table for comparison, if the hash values are the same and the text content is exactly the same, it is determined as duplicate data and excluded. For garbled data, the original encoding is identified in combination with a character encoding detection library (such as chardet), and an attempt is made to convert it to the common UTF-8 encoding format. If the conversion fails, it is marked as invalid data and discarded.

[0063] In terms of format standardization, all text is uniformly converted to lowercase to facilitate the consistency of subsequent processing. At the same time, the time formats of different platforms (such as the diverse timestamp formats on social media platforms and the custom date formats on e-commerce platforms, etc.) are uniformly converted to a standard time format (such as ISO 8601 format), which is convenient for subsequent data analysis in time series.

[0064] Word segmentation and part-of-speech tagging:

[0065] Open-source natural language processing toolkits, such as NLTK or Jieba, are adopted, combined with a custom domain dictionary for word segmentation operations. The domain dictionary covers Internet buzzwords, professional terms in specific fields (such as professional parameter vocabulary in electronic product reviews, disease names in the medical and health field, etc.), and platform-specific abbreviations (such as "yxh" on Weibo representing marketing accounts). During the word segmentation process, the words in the domain dictionary are preferentially matched, and for the text that cannot be matched, the default word segmentation strategy of the toolkit is used.

[0066] Part-of-speech tagging also uses the above-mentioned toolkit to tag the part of speech of each word after word segmentation, such as nouns, verbs, adjectives, etc. This helps the subsequent deep learning model better understand the semantic structure of the text. For example, the part-of-speech tagging of common verbs like "abuse" and adjectives like "vulgar" in malicious comments can provide key clues for model feature extraction. The tagged results are stored in a new data structure (such as a nested list containing a word segmentation list and a corresponding part-of-speech list) and await input into the deep learning model.

[0067] IV. Implementation of the deep learning model construction module:

[0068] Model selection and architecture construction:

[0069] According to the actual application scenario requirements and data characteristics, a suitable deep learning model architecture is selected. If the comment data focuses on local feature extraction, such as the malicious tendency judgment of individual words or phrases, an architecture mainly based on a convolutional neural network (CNN) is preferentially adopted. A multi-layer convolutional neural network is constructed, and convolutional kernels of different sizes are set, such as 3x3, 5x5, etc. Small convolutional kernels are used to capture subtle features at the word level, while large convolutional kernels focus on semantic information at the phrase level. After the convolutional layer, a max pooling layer is connected, with a pooling window size of 2x2 and a stride of 2. Through downsampling, the dimension of the feature map is reduced, reducing the computational burden of the subsequent fully connected layer while retaining key feature information.

[0070] When the comment data has strong sequential dependence, that is, there is a close semantic association between the front and back words and sentences (such as logical coherent malicious expressions in long comments), recurrent neural networks (RNNs) and their variants, such as long short-term memory networks (LSTMs) or gated recurrent units (GRUs), are selected. Taking the LSTM as an example, a hidden layer containing multiple LSTM units is constructed. Each unit contains an input gate, a forget gate, an output gate and a memory unit. Through the gating mechanism, the flow and preservation of information are effectively controlled, avoiding the vanishing gradient or exploding gradient problems that occur in traditional RNNs when processing long sequences, so as to accurately capture the long-term and short-term dependence relationships in the text sequence.

[0071] For complex and diverse actual comment data, the advantages of CNN and RNN can also be combined to build a hybrid model. For example, first use CNN to perform preliminary feature extraction on the text, and input the obtained feature maps into the RNN layer in sequence order. The RNN further mines the sequential associations between the features, and finally outputs to the fully connected layer for classification and judgment.

[0072] Hyperparameter tuning:

[0073] The initial setting of the model hyperparameters is based on empirical values and the results of small-scale experiments. For example, the learning rate is initially set to 0.001, the batch size is set to 64, and the number of filters in the convolutional neural network gradually increases between 32 and 128 according to the number of model layers. During the model training process, an automatic hyperparameter tuning strategy is adopted, combined with tools such as Hyperopt or Ray Tune, to define the hyperparameter search space. Taking the accuracy of the model on the validation set as the optimization goal, algorithms such as random search and Bayesian optimization are used to find the optimal hyperparameter combination. Re-evaluate the effect of the hyperparameter configuration every certain training period (such as 10 epochs) to continuously optimize the model performance.

[0074] V. Implementation of the model training module:

[0075] Sample preparation:

[0076] Collect a large number of labeled malicious comment samples and normal comment samples to construct a training set and a validation set. The sample sources include manually labeled historical data, publicly available network comment datasets, and data collected from various network platforms in real time and manually re-reviewed and labeled. Ensure that the malicious comment samples cover various malicious types, such as abuse, slander, advertising spam, etc. The proportions of various samples are as balanced as possible to avoid the model being biased towards a specific type of recognition. Randomly divide the samples into a training set and a validation set according to a ratio of 8:2. At the same time, k-fold cross-validation (such as k = 5) is adopted during the training process to further enhance the generalization ability of the model.

[0077] Training process:

[0078] Based on the selected deep learning model architecture and initialized hyperparameters, the training set data is used for batch training. After each batch of data is input into the model, the prediction results are calculated through forward propagation, and the loss value is obtained by comparing with the true labels. Then, the model weights are updated through backpropagation using Stochastic Gradient Descent (SGD) and its variant algorithms (such as Adagrad, Adadelta, Adam, etc.). During the training process, the training loss and the validation set accuracy are monitored in real time, and visualization tools (such as TensorBoard) are used to plot the loss curve and the accuracy curve to intuitively observe the convergence of the model. If the validation set accuracy does not improve for multiple consecutive training epochs (such as 5 epochs), and the training loss also tends to be stable, the model is determined to reach the current optimal state and the training is stopped; otherwise, the training is continued to be optimized according to the hyperparameter adjustment strategy.

[0079] VI. Implementation of the real-time detection module:

[0080] Low-latency optimization:

[0081] The trained deep learning model is deployed to a high-performance server cluster, and technologies such as multi-threading and asynchronous I / O are used to achieve low-latency detection. Multiple worker threads are started on the server side, and each thread is responsible for reading a comment data from the preprocessing module, loading the model for inference calculation, and returning the results. The asynchronous I / O mode is adopted. When the thread waits for the model inference result, it does not block the data reading and processing operations of other threads, making full use of the server resources to ensure that the system can complete the detection of a single comment within milliseconds.

[0082] To further accelerate the inference process, the model is quantized and compressed. Using tools such as TensorFlow Lite or PyTorch Mobile, the model parameters are converted from high precision (such as 32-bit floating-point numbers) to low precision (such as 8-bit integers) representation. Without significantly losing the model accuracy, the model storage space and the computational amount are reduced, and the inference speed is improved. At the same time, combined with GPU acceleration technology, the model calculation tasks are distributed to the GPU cores for parallel processing, giving full play to the advantages of GPU in large-scale matrix operations. Especially for large-scale deep neural network models, the inference time can be significantly shortened.

[0083] Distributed deployment:

[0084] When the system faces a high concurrent comment stream, a distributed deployment solution is used to improve the overall detection capability and throughput. A distributed real-time detection platform is built based on a distributed framework (such as Apache Flink or Spark Streaming), and the comment data is partitioned according to certain rules (such as based on platform source, geographical distribution, etc.), and distributed to multiple computing nodes for parallel processing. Each computing node runs the real-time detection module independently, equipped with independent model copies and processing resources. The nodes coordinate work through high-speed network communication and summarize the final detection results. Through distributed deployment, the system can linearly expand processing capabilities and easily cope with the real-time detection needs of massive comment data.

[0085] 7. Feedback module implementation:

[0086] Platform feedback mechanism:

[0087] When the real-time detection module determines that a comment is a malicious comment, the feedback module immediately starts the interaction process with the corresponding network platform. Using the management interface provided by the platform (such as the reporting interface of the social media platform, the illegal comment processing interface of the e-commerce platform, etc.), the detailed information of the malicious comments, including the comment content, publisher ID, detection time, and malicious classification basis (such as insulting words belong to personal attack malicious comments, and sensitive topics belong to illegal information, etc.) are sent to the platform in a standardized format (such as XML or JSON). After receiving the feedback, the platform takes corresponding measures against the malicious comments according to its own rules, such as immediately blocking, sending warning notifications to the publisher, restricting some of the publisher's permissions, etc., and at the same time feeds back the processing results to this detection system for recording and statistical analysis.

[0088] Model optimization feedback:

[0089] In addition to feedback to the platform, malicious comment samples and their feature information (such as deep feature vectors extracted from the middle layer of the deep learning model, pre-processed keyword lists, etc.) are sent back to the model training module in real time. The model training module incorporates these new samples into the sample library, regularly updates the training set, triggers a new round of model training process, and optimizes model parameters using feature information in new samples, so that the model can adapt to new malicious comment patterns in a timely manner and continuously improve the accuracy and generalization ability of malicious comment recognition.

[0090] Through the detailed implementation methods of the above modules, this real-time detection system for malicious comments based on deep learning can operate efficiently and accurately identify malicious comments on the network platform, providing a solid guarantee for creating a healthy and harmonious network environment.

[0091] Efficient and accurate review collection:

[0092] The data collection module has strong adaptability for different types of network platforms, such as social media, e-commerce, news and information platforms, etc. By utilizing open APIs, simulating browser behavior, and combining RSS feeds with crawler technology, it can capture comment data from various sources in an all-round, continuous and stable manner, ensuring that the data sources are extensive and comprehensive, without missing any potential malicious comment information.

[0093] At the same time, the preliminary screening rules set up can quickly remove invalid data and only store valuable comments in a structured format in a temporary cache database, which not only saves storage space, but also provides high-quality data input for subsequent modules, greatly improving the data processing efficiency of the entire system and ensuring the accuracy of detection from the source.

[0094] Optimized text preprocessing:

[0095] The preprocessing module uses regular expressions, character encoding detection libraries, and advanced natural language processing toolkits to deeply clean comment data, remove invalid characters such as HTML tags, emoticons, and garbled characters, standardize the text, and effectively reduce data noise, making the text data processed by subsequent models purer and more standardized.

[0096] Combined with custom domain dictionaries for word segmentation and part-of-speech tagging, it can accurately identify Internet buzzwords, professional terms, and platform abbreviations, so that the semantic structure of the text can be clearly presented. It provides more targeted and semantically rich feature information for deep learning models, significantly improves the model's ability to understand the semantics of complex texts, and thus improves the accuracy of identifying malicious comments.

[0097] Flexible and powerful model building and training:

[0098] The deep learning model building module can flexibly select or combine model architectures based on the characteristics of the comment data. Whether it is CNN that focuses on local features, RNN and its variants that are good at processing sequence dependencies, or a hybrid model that combines the two, it can accurately adapt to the needs of different scenarios and fully explore malicious feature patterns in the text.

[0099] The model training module relies on a large and diverse source of samples, covering a variety of malicious types. It adopts a scientific sample division strategy and cross-validation method, combined with automatic hyperparameter adjustment technology, to continuously optimize model performance, so that the model has a strong generalization ability, can accurately identify the endless new malicious comments, and effectively respond to the complex and changeable network language environment.

[0100] Excellent real-time detection performance:

[0101] With the help of a high-performance server cluster, the real-time detection module uses multi-threading, asynchronous I / O technologies, as well as model quantization compression and GPU acceleration means to achieve a single-comment detection speed of milliseconds. Even under the impact of a high-concurrency comment stream, it can ensure the real-time response of the system, intercept malicious comments in a timely manner, and minimize their dissemination hazards to the greatest extent.

[0102] The distributed deployment solution further expands the throughput of the system, linearly improves the processing capacity, easily handles massive comment data, enables the real-time detection to operate stably and efficiently in large-scale application scenarios, and comprehensively guarantees the healthy ecosystem of the comment area on the network platform.

[0103] Dynamic optimization and platform collaboration:

[0104] On the one hand, the feedback module closely collaborates with the network platform, and real-time feedbacks the details of malicious comments and the classification basis, helping the platform quickly take precise control measures such as blocking and warning to purify the comment environment of the platform; on the other hand, it transmits malicious comment samples and features back to the model training module, prompting the model to continuously learn newly emerging malicious patterns, dynamically optimize and upgrade, continuously improve the recognition accuracy, and form a virtuous cycle of system self-improvement mechanism, always maintaining the ability to keenly capture and precisely strike malicious comments.

[0105] The above describes the present invention and its implementation manners. Such description is not restrictive. What is shown in the drawings is only one of the implementation manners of the present invention, and the actual structure is not limited thereto. All in all, if those of ordinary skill in the art are inspired by it and design similar structural manners and embodiments without creative efforts without departing from the spirit of the present invention, they shall fall within the protection scope of the present invention.

Claims

1. A real-time malicious comment detection system based on deep learning, characterized by: A data collection module is used to collect comment data in real time from multiple network platforms, including but not limited to social media platforms, e-commerce platforms, and news information platforms. This module has the ability to dynamically adapt to data interfaces of different platforms and can automatically adjust collection strategies according to changes in platform rules to ensure continuous and stable acquisition of comment data; The preprocessing module is connected to the data collection module and is used to clean the collected comment data, remove invalid characters, duplicate data, and garbled characters, and convert the text into a standard format. At the same time, it performs word segmentation and part-of-speech tagging on the text to prepare for the subsequent deep learning model input. This module supports a variety of natural language processing algorithms and can automatically select the optimal algorithm combination based on data features; Deep learning model building module, which builds a deep learning model based on convolutional neural network (CNN), recurrent neural network (RNN) and its variants or a combination of the two. The model includes an input layer, a hidden layer and an output layer. The hidden layer contains multiple neurons for automatically extracting the features of the review data. The hyperparameters of the model can be dynamically adjusted according to the training effect to optimize the model performance; A model training module is connected to the deep learning model building module, and uses the labeled malicious comment samples and normal comment samples to train the deep learning model. The training process adopts a batch training method, combined with optimization algorithms such as stochastic gradient descent, and continuously adjusts the model weights to improve the model's recognition accuracy for malicious comments. The sample library is regularly updated to include new malicious comment type samples; A real-time detection module is connected to the model training module, and inputs the pre-processed real-time comment data into the trained deep learning model. The model outputs the judgment result of whether the comment is a malicious comment or a normal comment. The module has a low latency feature and can complete the detection of a single comment within milliseconds, meeting the real-time requirements; The feedback module is connected to the real-time detection module. When malicious comments are detected, the results are fed back to the corresponding network platform so that the platform can take measures such as blocking and warning. On the other hand, the malicious comment samples and their feature information are fed back to the model training module for further optimization of the model.

2. According to the deep learning-based real-time malicious comment detection system of claim 1, it is characterized by: The convolutional neural network in the deep learning model building module adopts a multi-layer convolution structure, captures local and global features of text through convolution kernels of different sizes, and connects a pooling layer after the convolution layer to reduce the dimension of the feature map and reduce the amount of calculation.

3. A real-time malicious comment detection system based on deep learning according to claim 2, characterized in that: The recurrent neural network in the deep learning model building module adopts a long short-term memory network (LSTM) or a gated recurrent unit (GRU), and uses its memory unit to effectively process the long-term and short-term dependencies in the text sequence to avoid the gradient disappearance or gradient explosion problem.

4. According to claim 3, a real-time malicious comment detection system based on deep learning is characterized in that: During the word segmentation process, the preprocessing module combines the domain dictionary and the pre-trained word vector model to improve the accuracy of word segmentation, especially for the processing of special words such as new network words, professional terms, abbreviations, etc.

5. A real-time malicious comment detection system based on deep learning according to claim 4, characterized in that: During the training process, the model training module adopts a cross-validation method to divide the samples into multiple subsets, and uses one of the subsets as a validation set and the remaining subsets as training sets in turn to comprehensively evaluate the model performance and prevent overfitting.

6. The real-time detection system for malicious comments based on deep learning according to claim 5 is characterized by: The real-time detection module supports distributed deployment and can process comment data in parallel on multiple server nodes, further improving the real-time detection capability and throughput of the system.

7. The real-time detection system for malicious comments based on deep learning according to claim 6 is characterized by: When the feedback module feeds back malicious comment results to the network platform, it also includes detailed malicious classification criteria, such as abusive words, sensitive topics, and malicious intent categories, to help the platform manage accurately.