Training Method for SMS Content Review Model, Electronic Device, Computer Program Product, and SMS Review Method

By selecting appropriate training methods and models (such as vector database, logistic regression, naive Bayes) for the SMS content audit model, the existing SMS audit methods have high maintenance costs and poor flexibility, and fast and accurate SMS content audits have been achieved.

CN119357671BActive Publication Date: 2025-06-13BEIJING LINGDU TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411526897.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-10-30
Publication Date
2025-06-13
Estimated Expiration
2044-10-30

AI Technical Summary

Technical Problem

The existing SMS audit methods have high maintenance costs, poor flexibility, difficult to handle complex content, and are inefficient and prone to omissions.

Method used

Provide a training method for SMS content auditing model, by obtaining sample data based on historical data, selecting appropriate training models (vector database, logistic regression, naive Bayes) for processing, and generating audit models.

Benefits of technology

This enables fast and accurate model training, reduces maintenance costs, improves flexibility, and can accurately and quickly process complex content and avoid omissions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119357671B_ABST
    Figure CN119357671B_ABST
Patent Text Reader

Abstract

The present invention provides a training method for a short message content review model, an electronic device, a computer program product, and a short message review method. Among them, the training method for the short message content review model includes: obtaining sample data based on historical data; selecting a vector database, a logistic regression training model, or a naive Bayes training model to process the corresponding sample data according to the data volume of the sample data, so as to obtain a first training model, a second training model, and a third training model, and obtaining a review model based on the first training model, the second training model, and the third training model. The purpose of reducing maintenance costs, improving flexibility, accurately and quickly processing complex content, and avoiding omissions is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of SMS review, and in particular, relates to a training method for an SMS content review model, an electronic device, a computer program product, and an SMS review method. Background Art

[0002] With the development of mobile communication technology, SMS has become one of the important communication methods in people's daily lives. However, with the increase in the number of SMS, it also contains some inappropriate content, such as spam messages, fraud messages, etc. Traditional review methods often rely on pre-setting rule keywords in the program for text review or manual review, which have high maintenance costs, poor flexibility, difficulty in handling complex content, low efficiency, and are prone to omissions. Summary of the Invention

[0003] In view of the problems existing in the prior art, the present invention provides a training method for an SMS content review model, an electronic device, a computer program product, and an SMS review method, which at least partially solves the problems of high maintenance costs, poor flexibility, difficulty in handling complex content, low efficiency, and prone to omissions existing in the prior art.

[0004] In a first aspect, an embodiment of the present disclosure provides a training method for an SMS content review model, including:

[0005] Obtaining sample data based on historical data;

[0006] Obtaining a first data volume of the sample data. If the first data volume is less than a first set data volume, the vector database is used to process the sample data to obtain a first training model;

[0007] If the first data volume is not less than the first set data volume, similarity deduplication processing is performed on the sample data to obtain a second data volume after similarity deduplication. If the second data volume is greater than a second set data volume, a logistic regression training model is used to process the sample data after similarity deduplication to obtain a second training model;

[0008] If the second data volume is not greater than the second set data volume, a naive Bayes training model is used to process the sample data before similarity deduplication to obtain a third training model;

[0009] An audit model is obtained based on the first training model, the second training model, and the third training model.

[0010] Optionally, the obtaining sample data based on historical data includes:

[0011] Load historical data within a set time period as training samples. After the loading is completed, judge the ratio of positive and negative samples. If the ratio is unbalanced, continue to load the data before the set time period.

[0012] Optionally, obtaining sample data based on historical data includes: excluding spam messages from the historical data; that is, the negative samples of the historical data include the rejected message data but do not include the spam message data.

[0013] Optionally, obtaining sample data based on historical data includes: excluding the message data with at least two different review results for the same message.

[0014] Optionally, before the step of processing the sample data after deduplication of similarity using a logistic regression training model, it is necessary to judge whether the positive and negative samples of the sample data after deduplication of similarity meet the requirements. If they do not meet the requirements, oversampling or undersampling is selected to process the sample data after deduplication of similarity.

[0015] Optionally, before the step of processing the sample data before deduplication of similarity using a naive Bayes training model, it is necessary to judge whether the positive and negative samples of the sample data before deduplication of similarity meet the requirements. If they do not meet the requirements, undersampling is selected to process the sample data before deduplication of similarity.

[0016] Optionally, after the step of obtaining the review model based on the first training model, the second training model, and the third training model, it is also necessary to update the obtained review model. When updating, back up the review model before the update. During the update time period, when data is received, judge whether the current review model status is available. If it is not available, use the backed-up review model before the update to process the received data.

[0017] In a second aspect, an embodiment of the present disclosure also provides an electronic device, which includes:

[0018] At least one processor; and,

[0019] A memory communicatively connected to the at least one processor; wherein,

[0020] The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the training method for the short message content review model according to any one of the first aspects.

[0021] In a third aspect, an embodiment of the present disclosure also provides a computer program product, including a computer program / instructions, and when the computer program / instructions are executed by a processor, the training method for the short message content review model according to any one of the first aspects is implemented.

[0022] Fourthly, embodiments of the present disclosure also provide a text message auditing method. An auditing model obtained by using any of the training methods in the first aspect is used, and the auditing result of the auditing model is stored in a cache. When new auditing data is received, first, the new auditing data is compared with the auditing results in the cache. If the auditing results in the cache exist in the new auditing data, the auditing result of the new auditing data is directly obtained based on the auditing results in the cache.

[0023] The present invention provides a training method for a text message content auditing model, an electronic device, a computer program product, and a text message auditing method. In the training method for the text message content auditing model, a corresponding training model is selected according to the data volume of sample data, so as to obtain an auditing model, thereby enabling the model training to be fast and accurate. The trained auditing model is used to audit text messages, achieving the purposes of reducing maintenance costs, improving flexibility, accurately and quickly processing complex content, and avoiding omissions. Description of the Drawings

[0024] By describing the exemplary embodiments of the present disclosure in more detail in conjunction with the drawings, the above and other objects, features, and advantages of the present disclosure will become more obvious. Among them, in the exemplary embodiments of the present disclosure, the same reference numerals generally represent the same components.

[0025] Figure 1 It is a flowchart of the training method for the text message content auditing model provided by the embodiments of the present disclosure;

[0026] Figure 2 It is a flowchart of a text message auditing provided by the embodiments of the present disclosure;

[0027] Figure 3 It is a schematic block diagram of an electronic device provided by the embodiments of the present disclosure. Detailed Embodiments

[0028] The following describes the embodiments of the present disclosure in detail in conjunction with the drawings.

[0029] It should be clear that the following illustrates the embodiments of the present disclosure through specific specific examples. Those skilled in the art can easily understand other advantages and effects of the present disclosure from the content disclosed in this specification. Obviously, the described embodiments are only a part of the embodiments of the present disclosure, rather than all of the embodiments. The present disclosure can also be implemented or applied through other different specific embodiments, and various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present disclosure. It should be noted that, without conflict, the following embodiments and the features in the embodiments can be combined with each other. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present disclosure without creative efforts belong to the scope of protection of the present disclosure.

[0030] It should be noted that the following describes various aspects of embodiments within the scope of the appended claims. It will be apparent that the aspects described herein can be embodied in a wide variety of forms, and any specific structure and / or function described herein is illustrative only. Based on this disclosure, those skilled in the art should understand that one aspect described herein can be implemented independently of any other aspect, and two or more of these aspects can be combined in various ways. For example, any number of aspects set forth herein can be used to implement a device and / or practice a method. Additionally, this device and / or method can be implemented using other structures and / or functionality in addition to one or more of the aspects set forth herein.

[0031] It should also be noted that the diagrams provided in the following embodiments only schematically illustrate the basic concept of the present disclosure. Only the components related to the present disclosure are shown in the diagrams, rather than being drawn according to the number, shape, and size of the components in actual implementation. The type, quantity, and ratio of each component in actual implementation can be arbitrarily changed, and the component layout type may also be more complex.

[0032] In addition, in the following description, specific details are provided to facilitate a thorough understanding of the examples. However, those skilled in the art will understand that the aspects described can be practiced without these specific details.

[0033] Faiss is a library developed by Facebook AI Research, specifically for efficient similarity search and dense vector clustering. Faiss is very efficient in dealing with large-scale vector search problems and is a very important tool in the fields of machine learning and big data.

[0034] Main features: Efficiency: Faiss provides a variety of algorithms and data structures to accelerate the process of vector search, including CPU-based and GPU-based implementations.

[0035] Scalability: It can handle vectors in the hundreds of millions and supports distributed search.

[0036] Versatility: Supports a variety of distance calculation methods, such as Euclidean distance, inner product, etc.

[0037] Multi-language interfaces: Although Faiss is written in C++, it provides interfaces for languages such as Python and MATLAB.

[0038] Main scenarios and problems solved:

[0039] Recommendation System: Problem: In a recommendation system, it is necessary to find a set of items that are most similar to a user's historical behavior or interests. Solution: Faiss can quickly find the most similar ones among a large number of item vectors, thus providing personalized recommendations for users.

[0040] Image Recognition and Retrieval: Problem: In an image retrieval system, it is necessary to find similar images in the database according to a query image. Solution: By converting the image into a vector and using Faiss for similarity search, matching images can be quickly found.

[0041] Speech Recognition: Problem: In speech recognition, it is necessary to convert speech signals into text, which usually involves finding matches among a large number of speech feature vectors. Solution: Faiss can be used to accelerate this process and improve the efficiency of speech recognition.

[0042] Natural Language Processing (NLP): Problem: In NLP tasks, such as text similarity analysis or question-answering systems, it is necessary to find a text segment that is most similar to the query text. Solution: By converting the text into a vector representation and using Faiss for similarity search, this problem can be efficiently solved.

[0043] Security Monitoring: Problem: In a video monitoring system, it is necessary to identify and track multiple targets. Solution: Through feature vector matching, Faiss can help the system quickly identify the targets in the monitoring footage.

[0044] Gene Data Analysis: Problem: In gene sequence analysis, it is necessary to find similar sequences among a large number of gene expression vectors. Solution: Faiss can handle high-dimensional data and is suitable for efficient search in gene databases.

[0045] Audio Fingerprint Technology: Problem: In audio recognition, it is necessary to identify and match audio segments. Solution: By converting the audio into a vector and using Faiss for search, matching audio can be quickly found.

[0046] Due to its high efficiency and versatility, Faiss has a wide range of applications in the fields of machine learning and big data. By providing fast and accurate vector search capabilities, Faiss has greatly promoted the development of applications that require large-scale similarity search.

[0047] The Naive Bayes classifier is a simple probabilistic classifier based on Bayes' theorem, which assumes that the features are independent of each other. Although this assumption often does not hold in reality, the Naive Bayes classifier still shows amazing effectiveness and efficiency in many practical applications.

[0048] Basic Principle:

[0049] Bayes' theorem describes the probability of one event occurring given that another event has occurred. For classification problems, the Naive Bayes classifier uses the following formula to calculate the probability of the target variable \(y\) given the input features \(x\):

[0050] \(P(y|x_1,x_2,\ldots,x_n)=\frac{P(y)\prod_{i = 1}^{n}P(x_i|y)}{P(x_1,x_2,\ldots,x_n)}\), \(P(y|x_1,x_2,\ldots,x_n)=\frac{P(x_1,x_2,\ldots,x_n)P(y)\prod_{i = 1}^{n}P(x_i|y)}{}\),

[0051] where: \(P(y)\) is the prior probability of the class label \(y\). \(P(x_i|y)\) is the conditional probability of the feature \(x_i\) given the class label \(y\). \(P(x_1,x_2,\ldots,x_n)\) is the joint probability of all features, which is usually ignored during classification because it is constant for all classes.

[0052] Features: Simple: The model of the Naive Bayes classifier is very simple and easy to understand and implement.

[0053] Efficient: The calculation is very efficient when the assumption of feature independence holds.

[0054] Requires fewer samples: It can perform well even with less data.

[0055] Application scenarios: The Naive Bayes classifier is applicable to fields such as text classification, spam detection, disease prediction, etc.

[0056] Advantages and disadvantages: Advantages: The algorithm is simple and easy to implement. Requires fewer samples and can handle multi-classification problems. The prediction speed is fast and it is suitable for real-time systems.

[0057] Disadvantages: The assumption of feature independence limits its accuracy. It is very sensitive to the representation form of the input data.

[0058] Logistic Regression is a predictive analysis method widely used in statistics and machine learning, especially in binary classification problems. Despite its name "regression", it is actually a classification algorithm used to predict the occurrence or non-occurrence of an event.

[0059] Basic principle:

[0060] Logistic Regression maps the output of a linear regression model to between 0 and 1 using a logistic function (usually the Sigmoid function) to predict the probability of an event occurring. The formula for the Sigmoid function is:

[0061] σ(z)=11+e−zσ(z)=1+e−z1,

[0062] where \(z\) is a linear combination of the model, i.e., \(z = \beta_0+\beta_1x_1+\beta_2x_2+\cdots+\beta_nx_n\), \(\beta\) are coefficients, and \(x\) are features.

[0063] Features: Output Explanation: The output value can be interpreted as the probability of belonging to a certain class. Binary Classification: Mainly used for binary classification problems, but can also be extended to multi-class classification problems. Output Range: The output value ranges between 0 and 1.

[0064] Application Scenarios:

[0065] Logistic regression is applicable to various binary classification problems, such as: Credit Scoring Model: Predict whether an individual has a default risk. Medical Diagnosis: Predict the presence or absence of a disease. Spam Detection: Determine whether an email is spam. Customer Churn Prediction: Predict whether a customer will stop using the service.

[0066] Advantages: The model is simple, easy to implement and interpret. The output is easily interpretable as a probability. It has relatively high computational efficiency.

[0067] Disadvantages: It cannot handle non-linear problems well. When the feature space is large, the performance may decline. For non-linear relationships, feature engineering is required to extract useful features.

[0068] 3. Oversampling and Undersampling:

[0069] Oversampling and Undersampling are two commonly used techniques for dealing with imbalanced datasets.

[0070] Oversampling is a method of balancing the dataset by increasing the number of minority class samples. Common oversampling methods include:

[0071] Random Oversampling: Increase the number of minority class samples by randomly replicating them, but it may lead to overfitting because it does not introduce new information.

[0072] SMOTE (Synthetic Minority Over-sampling Technique): Generate new samples by interpolating between minority class samples, thus reducing the risk of overfitting.

[0073] ADASYN: Generate synthetic samples adaptively according to the classification difficulty of the samples. More new samples are generated for samples with greater classification difficulty.

[0074] The advantages of oversampling include the ability to increase the size of the dataset, improve the performance of the model, address the problem of class imbalance in the dataset, avoid information loss, and enhance the performance of the model. However, it may also lead to overfitting because the generated new samples may introduce noise, cause a shift in the sample distribution, and increase the requirements for computing time and storage space.

[0075] Undersampling is a method to balance the dataset by reducing the number of majority-class samples. Common undersampling methods include:

[0076] Random undersampling: Randomly delete majority-class samples to reduce their number to the same or similar as the minority class, but it may lose a large amount of useful information and lead to a decline in model performance.

[0077] Ensemble undersampling (Cluster Centroids): Cluster the majority-class samples into several clusters through a clustering algorithm and replace the original samples with the cluster centroids to reduce the number of majority-class samples.

[0078] Tomek Links: Reduce the number of majority-class samples by deleting the majority-class sample pairs (Tomek links) closest to the minority-class samples. This method not only reduces the majority-class samples but also removes the noise samples in the dataset.

[0079] The advantages of undersampling are that it can reduce the size of the dataset, thereby reducing the requirements for computing time and storage space, lowering the complexity of the model, avoiding overfitting, balancing the class distribution, and thus improving the accuracy of the model. However, it may also lead to information loss. When the number of minority-class samples in the dataset is too small, undersampling may cause underfitting and may introduce randomness, resulting in different results for different undersampling methods.

[0080] The choice between oversampling and undersampling depends on the specific dataset and problem. If the number of minority-class samples is extremely small, directly deleting majority-class samples will cause excessive information loss, and oversampling is suitable. If the number of majority-class samples is extremely large, randomly replicating minority-class samples will lead to overfitting, and undersampling is suitable. Sometimes, the methods of oversampling and undersampling can also be combined to further improve the model performance.

[0081] Existing SMS auditing systems generally adopt rule-based text auditing schemes and manual auditing. It mainly relies on predefined rules and keyword lists to detect inappropriate content. When encountering a large number of complex SMS contents that are difficult to process, manual auditing is required, and the workload is very large.

[0082] Disadvantages:

[0083] (1) High maintenance cost: With the change and development of the language, the rules and keyword lists need to be updated frequently to adapt to new forms of inappropriate content.

[0084] (2) Poor flexibility: Rules usually target known forms of bad content and may not be able to effectively detect newly emerging forms.

[0085] (3) False positives and false negatives: Limited ability to understand variants, synonyms, and context may lead to false positives (wrongly marking normal content) or false negatives (failing to identify bad content).

[0086] (4) Difficulty in handling complex content: For language content with complex context, metaphors, or irony, rule-based methods may have difficulty making accurate judgments.

[0087] Such as Figure 1 As shown, this embodiment discloses a training method for a short message content review model, including:

[0088] Obtaining sample data based on historical data;

[0089] The historical data sources mainly come from two places:

[0090] 1. The review record table stores all historical manually reviewed short message content and status (review passed, review rejected, etc.) information, which is used as labeled positive and negative samples for training.

[0091] 2. Historically accumulated spam messages such as fraud and violations.

[0092] Because of the large amount of data, each time during training, data is not loaded from MySQL all at once, but is retrieved from the database incrementally at regular intervals and placed on the server. Therefore, each time during training, the text data accumulated on the server (with very fast reading speed) is used. Compared with directly querying from the MySQL database, the speed is increased by dozens or hundreds of times.

[0093] During training, short message content is mainly divided into the following types: industry, marketing, light reminder, heavy reminder, Internet finance, finance, spam messages, etc. It supports expanding more types, and more types can be easily added.

[0094] According to different types, different types of machine learning models will be trained, which can ensure higher accuracy.

[0095] Obtain the first data volume of the sample data. If the first data volume is less than the first set data volume, then use a vector database to process the sample data, thereby obtaining the first training model;

[0096] In a specific application scenario, if the data volume of a certain type of short message is too small, in the order of ten thousand or less, directly use a vector database, such as the Faiss vector database. The types include industry, marketing, light reminder, heavy reminder, Internet finance, finance, spam messages, etc.

[0097] If the first data volume is not less than the first set data volume, perform similarity deduplication on the sample data to obtain the second data volume after similarity deduplication. If the second data volume is greater than the second set data volume, use the logistic regression training model to process the sample data after similarity deduplication to obtain the second training model;

[0098] If the second data volume is not greater than the second set data volume, use the Naive Bayes training model to process the sample data before similarity deduplication to obtain the third training model;

[0099] For the data volume after similarity deduplication, if the data volume of this type of sample is not large enough, in the order of hundreds of thousands, the model will select the Naive Bayes training model, otherwise use the logistic regression training model.

[0100] When using the Naive Bayes training model, use the samples after similarity deduplication for training, and use the logistic regression for the sample data without similarity deduplication.

[0101] The reason for choosing Naive Bayes and logistic regression is that the training speed of these two models is faster, and the model can be trained and updated faster.

[0102] Based on the first training model, the second training model and the third training model, an audit model is obtained.

[0103] Optionally, the obtaining of the sample data based on the historical data includes:

[0104] Load the historical data within a set time period as the training sample. After loading, judge the ratio of positive and negative samples. If the ratio is unbalanced, continue to load the data before the set time period.

[0105] Optionally, the obtaining of the sample data based on the historical data includes: excluding the spam messages in the historical data; that is, the negative samples of the historical data include the rejected message data, but do not include the spam message data.

[0106] When loading the data, do not load all the historical data, because too much historical data has accumulated. Therefore, load the data within a certain time period as the training sample. After loading, judge the ratio of positive and negative samples. If it is unbalanced, load the insufficient samples forward along the time line. For example, if there are not enough audited and rejected messages, continue to load the audited and rejected data before the loaded time range until the earliest allowable loading time limit or there is no data. When loading the data, the models of the existing technologies will incorporate the spam messages as negative samples. In this embodiment, the spam messages are excluded, so that the sample data is more accurate.

[0107] Optionally, obtaining sample data based on historical data includes: excluding text message data with at least two different review results for the same text message. For example, for the same message content, if one reviewer determines that the text message may be in violation and should not be sent, while another reviewer misreviews it and approves it or deems it approvable, such text messages with inconsistent review statuses will be removed from the training samples, thereby making the training more accurate.

[0108] Optionally, before the step of processing the sample data after deduplication of similarity using a logistic regression training model, it is necessary to determine whether the positive and negative samples of the sample data after deduplication of similarity meet the requirements. If they do not meet the requirements, oversampling or undersampling is selected to process the sample data after deduplication of similarity.

[0109] Deduplication of similarity means that text messages similar to this text message will be matched. For example, the main content is the same, but the link or a part of the name is different. These highly similar text messages will also be removed from the training data. At this time, a deduplication calculation is performed based on the similarity of the loaded data, and only one similar text message will be retained instead of the same text message.

[0110] Optionally, before the step of processing the sample data before deduplication of similarity using a naive Bayes training model, it is necessary to determine whether the positive and negative samples of the sample data before deduplication of similarity meet the requirements. If they do not meet the requirements, undersampling is selected to process the sample data before deduplication of similarity.

[0111] After data preprocessing and deduplication of similarity, before selecting the training samples and the model to start training, if a machine learning model is used instead of a vector database, and it is found that the positive and negative samples of the data are unbalanced at this time, another processing will be done:

[0112] If it is logistic regression, oversampling or undersampling is selected according to the configuration.

[0113] If it is naive Bayes, undersampling is used, and through testing and verification, the accuracy rate is the highest in this way.

[0114] For undersampling, the random undersampling method is used, and for oversampling, the SMOTE (Synthetic Minority Over-sampling Technique) method is used.

[0115] Use the trained sample data and the selected training model for training.

[0116] The trained sample data needs to be vectorized first before training and prediction can be performed.

[0117] Vectorization tool for samples:

[0118] Logistic regression uses: CountVectorizer (term frequency (TF) vectorizer). The vector database for Naive Bayes uses: TfidfVectorizer (term frequency-inverse document frequency (TF-IDF) vectorizer). After testing, the accuracy is higher in the current training samples.

[0119] After the model training is completed, the vectorizer used for training, the trained model, and the relevant model information (including information such as the type of short message) will be saved.

[0120] Using the vector database, the process is basically the same, and the constructed Faiss index object and the relevant model information will also be saved.

[0121] Optionally, after the step of obtaining the audit model based on the first training model, the second training model, and the third training model, it further includes updating the obtained audit model. When updating, the audit model before the update is backed up. During the update period, when data is received, it is judged whether the current state of the audit model is available. If it is not available, the backed-up audit model before the update is used to process the received data.

[0122] Data will be loaded regularly every day and a new model will be trained. After the training is completed, the registered prediction service instances will be obtained from the Nacos registry, and the internal network IP and port information exposed by the relevant instances will be obtained. Then, the trained relevant model data (model, vectorizer, model information) will be distributed to the instances of the prediction service.

[0123] After the distribution is completed, all instances will be notified to update the model.

[0124] The update of the model is a smooth process, as follows:

[0125] 1. The prediction service receives a notification of model update.

[0126] 2. The prediction service loads the corresponding model file, vectorizer, and model information from the local. Different types of short messages load different models.

[0127] 3. After the loading is completed, the current model object and vectorization object are set as the previous generation model, and the current model is pointed to the newly loaded model.

[0128] 4. How to ensure that even if any accident occurs during the update and the initialization process of the model will not affect the prediction is because the previous model object will be retained. When predicting, it is first judged whether the current state of the model is available. For example, during the update, if the state is still not available (including that either the model or the vectorizer is not available, it is considered not available), then the previous model object and vectorizer are first used for prediction.

[0129] The system corresponding to the training method includes: a data collection system, a machine review model training system, a prediction system, and a machine review background management system.

[0130] Data collection system: Responsible for obtaining the original SMS data from the existing SMS system and sending the data to the machine review model training system through the message middleware.

[0131] Machine review model training system: After receiving the data processed by the data collection system, it first performs data preprocessing, data similarity deduplication, and model selection on the data, and then trains the model for the selected model and sample data.

[0132] Prediction system: Receives the trained model file, vectorizer, and model information. At the same time, it receives the http request sent by the client. When predicting the SMS content, it will first perform spam SMS prediction. The vector database will give priority to similarity search. If there is a match, it will directly reject the review. If it is not a spam SMS, it will select the corresponding type of model according to the SMS type of the sender to predict whether the SMS is approved or rejected in the review.

[0133] Machine review background management system: Provides a graphical interface for reviewers to operate, including viewing review results, manual review, adjusting review policies, etc.

[0134] The application of this embodiment is as follows: SMS content filtering technology: In order to prevent the spread of spam SMS and bad information, the SMS platform uses algorithms such as text analysis, keyword matching, and machine learning to detect and filter the SMS content in real time. By establishing a large SMS content library and blacklist database, the platform can identify and intercept SMS containing sensitive words, advertising information, illegal content, etc., thus effectively purifying the SMS environment and protecting the legitimate rights and interests of users.

[0135] Machine review technology for SMS content: To ensure that the SMS content complies with the content specified by different industries, the platform has established an SMS content review system. This system combines manual review and automated review to strictly review the sent SMS content. The SMS content is roughly divided into the following industries: industry, marketing, light collection, heavy collection, Internet finance, finance, spam SMS, etc. Since the SMS content required to be sent by different industries is different, different industry models are trained according to different industries to improve the accuracy of automatic review. At the same time, manual review is performed on SMS that is unknown in automatic review.

[0136] Foreign IP identification in SMS content: If there is a foreign IP in the SMS content, the probability of illegal content will be greatly increased, and foreign IPs will be blocked. Use IP address location query services, such as IP databases (such as IP2Location, MaxMind GeoIP, etc.), to determine the geographical location of the IP address, match the IP address in the SMS content with the IP database, obtain the IP address location information, and determine whether the IP address location is foreign. If it is foreign, it will be blocked; if it is domestic, it can be excluded.

[0137] This embodiment has the following advantages: the audit scheme based on machine learning can automatically learn new features and patterns from a large amount of data, so as to better adapt to the ever-changing language environment.

[0138] Machine learning-based methods can adapt to these emerging content through continuous learning.

[0139] Machine learning models can learn the connections between synonyms and variations through training, thereby more accurately identifying bad content and more accurately distinguishing between normal and bad content, thereby reducing false positives and negatives.

[0140] Machine learning-based methods can better understand the context of text, and thus perform better when dealing with content containing metaphors, sarcasm, or other complex language structures. Machine learning models can also better handle ambiguous concepts and polysemous words, thereby improving the accuracy of moderation.

[0141] Automatically learn new features: The machine learning-based review solution can automatically learn new features and patterns from a large amount of reviewed SMS data, thereby better adapting to the ever-changing SMS content.

[0142] Higher accuracy: Different types of machine review models are trained according to different types of SMS content, and the accuracy is higher when the corresponding type of model is used for machine review of a certain type of SMS content.

[0143] Higher efficiency: The audit method based on machine learning can achieve faster processing speed and is suitable for processing large amounts of data and real-time auditing needs.

[0144] Better customizability: Machine learning-based audit solutions can customize different audit strategies by adjusting model parameters or training data to meet the specific content requirements of different types of SMS content audit results.

[0145] This embodiment also discloses a short message review method. An audit model obtained by using all or part of the steps of the training method disclosed in this embodiment is used. The audit result of the audit model is saved in the cache. When new audit data is received, first, the new audit data is compared with the audit results in the cache. If the audit results in the cache exist in the new audit data, the audit result of the new audit data is directly obtained based on the audit results in the cache.

[0146] As Figure 2 shown, when predicting a short message, spam message prediction will be used first, and the vector database will give priority to similarity search. If there is a match, the audit will be directly rejected.

[0147] If it is not a spam message, then according to the type of the short message sent by the sender, a corresponding type of model will be selected for prediction. Although different types of models and vector databases are used in this embodiment, a unified training and prediction interface is encapsulated.

[0148] The result predicted by the model and the vector database is a score. According to the configuration, subsequent processing is performed on the threshold value of the score. Whether the audit passes or is rejected, a small number of short messages will enter the manual audit process at the threshold edge. This improves the accuracy rate, reduces the misjudgment rate, and avoids prediction errors, resulting in short messages that should not be sent being sent.

[0149] The short messages that have been predicted or audited will be cached and can be cached for 48 hours. The cache time is set according to the device. When a new short message comes next time, it will be checked from the cache first. If the short message content is the same, the result will be directly returned. Those not in the cache will go through machine automatic auditing.

[0150] In short message auditing, because a small amount of data will be manually audited, within a few hours when the new model (trained with the latest samples) has not been updated, the incremental data that has been audited and rejected is matched for similarity using the vector database. Before prediction, the spam message model is used for prediction first, and then the model trained by the training method disclosed in this embodiment is used to predict whether to reject. If not, the corresponding training model in the corresponding model is used for prediction.

[0151] The electronic device disclosed in this embodiment includes a memory and a processor. The memory is used to store non-temporary computer-readable instructions. Specifically, the memory may include one or more computer program products, and the computer program products may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. The volatile memory may include, for example, random access memory (RAM) and / or cache memory, etc. The non-volatile memory may include, for example, read-only memory (ROM), hard disk, flash memory, etc.

[0152] The processor may be a central processing unit (CPU) or other forms of processing units with data processing capabilities and / or instruction execution capabilities, and may control other components in the electronic device to perform desired functions. In an embodiment of the present disclosure, the processor is used to run the computer-readable instructions stored in the memory, so that the electronic device executes all or part of the steps of the training method for the short message content review model in the foregoing embodiments of the present disclosure.

[0153] Those skilled in the art should understand that, in order to solve the technical problem of how to obtain good user experience effects, this embodiment may also include well-known structures such as communication buses and interfaces, and these well-known structures should also be included in the protection scope of the present disclosure.

[0154] As Figure 3 FIG. is a schematic structural diagram of an electronic device provided by an embodiment of the present disclosure. It shows a schematic structural diagram of an electronic device suitable for implementing the electronic device in the embodiments of the present disclosure. Figure 3 The illustrated electronic device is only an example and should not impose any limitations on the functions and usage scope of the embodiments of the present disclosure.

[0155] As Figure 3 As shown, the electronic device may include a processing device (such as a central processing unit, a graphics processing unit, etc.), which may perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) or a program loaded from a storage device into a random access memory (RAM). In the RAM, various programs and data required for the operation of the electronic device are also stored. The processing device, ROM, and RAM are connected to each other through a bus. An input / output (I / O) interface is also connected to the bus.

[0156] Generally, the following devices may be connected to the I / O interface: an input device including, for example, a sensor or a visual information acquisition device; an output device including, for example, a display screen; a storage device including, for example, a magnetic tape, a hard disk, etc.; and a communication device. The communication device may allow the electronic device to communicate with other devices (such as edge computing devices) wirelessly or wiredly to exchange data. Although Figure 3 the illustrated electronic device has various devices, it should be understood that it is not required to implement or include all the illustrated devices. Instead, more or fewer devices may be implemented or included.

[0157] In particular, according to an embodiment of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a non-transitory computer-readable medium, and the computer program includes program codes for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network through a communication device, or installed from a storage device, or installed from a ROM. When the computer program is executed by a processing device, all or part of the steps of the method for training a short message content review model according to an embodiment of the present disclosure are executed.

[0158] For a detailed description of this embodiment, reference may be made to the corresponding descriptions in the foregoing embodiments, and details will not be repeated here.

[0159] The computer-readable storage medium disclosed in this embodiment stores non-temporary computer-readable instructions. When the non-temporary computer-readable instructions are run by a processor, all or part of the steps of the method for training a short message content review model according to the foregoing embodiments of the present disclosure are executed.

[0160] The above-mentioned computer-readable storage medium includes but is not limited to: optical storage media (such as CD-ROMs and DVDs), magneto-optical storage media (such as MOs), magnetic storage media (such as magnetic tapes or external hard drives), media with built-in rewritable non-volatile memories (such as memory cards), and media with built-in ROMs (such as ROM cartridges).

[0161] For a detailed description of this embodiment, reference may be made to the corresponding descriptions in the foregoing embodiments, and details will not be repeated here.

[0162] The basic principles of the present disclosure have been described above in conjunction with specific embodiments. However, it should be noted that the advantages, benefits, effects, etc. mentioned in the present disclosure are only examples and not limitations, and it cannot be considered that these advantages, benefits, effects, etc. are essential for each embodiment of the present disclosure. In addition, the above-mentioned specific details are only for illustrative and facilitating understanding purposes, rather than limitations, and the above details do not limit the present disclosure to necessarily adopt the above specific details for implementation.

[0163] In this disclosure, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. The block diagrams of devices, apparatuses, equipment, and systems involved in this disclosure are only illustrative examples and do not intend to require or imply that they must be connected, arranged, and configured in the manner shown in the block diagrams. As those skilled in the art will recognize, these devices, apparatuses, equipment, and systems can be connected, arranged, and configured in any manner. Words such as "including", "comprising", "having", etc. are open-ended terms, meaning "including but not limited to", and can be used interchangeably with each other. The words "or" and "and" used herein refer to the word "and / or", and can be used interchangeably with it, unless the context clearly indicates otherwise. The word "such as" used herein refers to the phrase "such as but not limited to", and can be used interchangeably with it.

[0164] In addition, as used herein, "or" in the listing of items starting with "at least one" indicates a disjunctive listing, so that for example, the listing of "at least one of A, B, or C" means A or B or C, or AB or AC or BC, or ABC (i.e., A and B and C). Furthermore, the term "exemplary" does not mean that the examples described are preferred or better than other examples.

[0165] It should also be noted that in the systems and methods of this disclosure, each component or each step can be decomposed and / or recombined. These decompositions and / or recombinations should be regarded as equivalent solutions of this disclosure.

[0166] Various changes, substitutions, and alterations to the technologies described herein can be made without departing from the teachings defined by the appended claims. In addition, the scope of the claims of this disclosure is not limited to the specific aspects of the processes, machines, manufactures, compositions of events, means, methods, and acts described above. Current or later-developed processes, machines, manufactures, compositions of events, means, methods, or acts that perform substantially the same function or achieve substantially the same result as the corresponding aspects described herein can be utilized. Thus, the appended claims include such processes, machines, manufactures, compositions of events, means, methods, or acts within their scope.

[0167] The above description of the disclosed aspects is provided to enable any person skilled in the art to make or use this disclosure. Various modifications to these aspects will be readily apparent to those skilled in the art, and the general principles defined herein can be applied to other aspects without departing from the scope of this disclosure. Therefore, this disclosure is not intended to be limited to the aspects shown herein, but rather to the broadest scope consistent with the principles and novel features disclosed herein.

[0168] The foregoing description has been presented for purposes of illustration and description. In addition, the description is not intended to limit embodiments of the disclosure to the form disclosed herein. Although several example aspects and embodiments have been discussed above, those skilled in the art will recognize some variations, modifications, alterations, additions, and sub-combinations thereof.

Claims

1. A training method for a text message content review model, characterized in that: include: Get sample data based on historical data; Acquire a first data volume of sample data, and if the first data volume is less than a first set data volume, process the sample data using a Faiss vector database to obtain a first training model; If the first data volume is not less than the first set data volume, the sample data is subjected to similarity deduplication processing to obtain a second data volume after similarity deduplication; if the second data volume is greater than the second set data volume, the sample data after similarity deduplication is processed using a logistic regression training model to obtain a second training model; If the second data volume is not greater than the second set data volume, the naive Bayes training model is used to process the sample data before similarity deduplication, thereby obtaining a third training model; An audit model is obtained based on the first training model, the second training model and the third training model.

2. The training method for a text message content review model according to claim 1, characterized in that: The acquiring of sample data based on historical data includes: Load the historical data within the set time period as training samples. After loading, determine the ratio of positive and negative samples. If the ratio is unbalanced, continue to load the data before the set time period.

3. The training method for a text message content review model according to claim 2, characterized in that: The obtaining of sample data based on historical data includes: excluding junk text messages in the historical data; that is, the negative samples of the historical data include rejected text message data, but do not include junk text message data.

4. The training method for a text message content review model according to claim 1, characterized in that: The obtaining of sample data based on historical data includes: excluding SMS data with at least two different review results for the same SMS.

5. The training method for a text message content review model according to claim 1, characterized in that: The step of using a logistic regression training model to process the sample data after similarity deduplication includes determining whether the positive and negative samples of the sample data after similarity deduplication meet the requirements. If not, oversampling or undersampling is selected to process the sample data after similarity deduplication.

6. The training method for a text message content review model according to claim 1, characterized in that: The step of using the naive Bayes training model to process the sample data before similarity deduplication includes: judging whether the positive and negative samples of the sample data before similarity deduplication meet the requirements, and if not, selecting undersampling to process the sample data before similarity deduplication.

7. The training method for a text message content review model according to claim 1, characterized in that: After the step of obtaining the audit model based on the first training model, the second training model and the third training model, the method also includes updating the obtained audit model, backing up the audit model before the update during the update, and determining whether the current audit model status is available when data is received during the update time period. If not, use the backed-up audit model before the update to process the received data.

8. An electronic device, characterized in that: The electronic device comprises: at least one processor; and, a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the training method for the SMS content review model described in any one of claims 1-7.

9. A computer program product comprising a computer program / instructions, characterized in that When the computer program / instruction is executed by a processor, the training method for a text message content review model described in any one of claims 1-7 is implemented.

10. A text message audit method, using the audit model obtained by the training method according to any one of claims 1 to 7, characterized in that: The audit results of the audit model are saved in the cache. When new audit data is received, the new audit data is first compared with the audit results in the cache. If the audit results in the cache are the same as the new audit data, the audit results of the new audit data are directly obtained based on the audit results in the cache.

Citation Information

Patent Citations

  • Invalid report audio filtering method and device, equipment, storage medium and product

    CN118173096A

  • Sample label information verification method and device

    WO2020253636A1