An information determination method, device and computer readable storage medium
By acquiring user attribute information and message data, performing annotation processing and model training, a harassment information identification model is generated, solving the problem of difficulty in obtaining interactive message content and achieving accurate harassment information identification and privacy protection.
Patent Information
- Application Number
- CN202110514775.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-05-10
- Publication Date
- 2026-01-23
- Estimated Expiration
- 2041-05-10
AI Technical Summary
In existing technologies, obtaining the content of interactive messages is difficult, especially when users are unwilling to disclose the content, which increases the difficulty of identifying harassing messages and may infringe on user privacy.
By acquiring user attribute information and message data from sample users, performing annotation processing and model training, a harassment information identification model is generated. By utilizing the correspondence between user attributes and message data, the reliance on interactive message content is reduced, thereby achieving the identification of harassment information.
It effectively protects user privacy, reduces the difficulty of obtaining interactive message content, and improves the accuracy and efficiency of identifying harassing information.
Smart Images

Figure CN115329165B_ABST
Abstract
Description
Technical Field
[0001] This application relates to information determination technology in the field of information processing, and more particularly to an information determination method, apparatus and computer-readable storage medium. Background Technology
[0002] Currently, information has become an important means of communication between people. While information strengthens connections, it has also led to the emergence of nuisance messages such as prize-winning offers and advertisements, affecting users' normal work and life. Related technologies use keyword matching to determine whether user interaction messages are nuisance; however, some users are unwilling to disclose the content of their interaction messages, and obtaining the content would violate their privacy, making it difficult to obtain the content of interaction messages and increasing the complexity of implementation. Summary of the Invention
[0003] To address the aforementioned technical problems, embodiments of this application aim to provide an information determination method, apparatus, and computer-readable storage medium, which solves the problem of difficulty in obtaining the content of interactive messages and reduces the implementation difficulty.
[0004] The technical solution of this application is implemented as follows:
[0005] An information determination method, the method comprising:
[0006] Obtain user attribute information and message data information corresponding to the first user identifier of the sample user; wherein, the message data information characterizes the degree of interaction of the sample user's messages;
[0007] The user attribute information and the message data information are labeled to obtain tagged user attribute information and tagged message data information; wherein, the tagged user attribute information and tagged message data information characterize whether the interactive message corresponding to the user attribute information and the message data information is harassing information;
[0008] Based on the first user identifier, a first correspondence is determined between the user attribute information and the data information of the message;
[0009] Based on the user attribute information, the message data information, the user attribute information with tags, the message data information with tags, and the first correspondence, a harassment information identification model is obtained through model training.
[0010] In the above scheme, obtaining the user attribute information and message data information corresponding to the first user identifier of the sample user includes:
[0011] Based on the first communication interface, the pending message distribution traffic and pending message distribution frequency corresponding to the first user identifier are obtained from the first target information platform, and the pending message distribution traffic and pending message distribution frequency are standardized to obtain message distribution traffic and message distribution frequency; wherein, the data information of the message includes message distribution traffic and message distribution frequency;
[0012] Based on the second communication interface, the user attribute information to be processed corresponding to the first user identifier is obtained from the second target information platform, and the user attribute information to be processed is standardized to obtain the user attribute information.
[0013] In the above scheme, the step of labeling the user attribute information and the message data information to obtain tagged user attribute information and tagged message data information includes:
[0014] Obtain the negative credit record information corresponding to the first user identifier;
[0015] Based on the negative credit record information, the user attribute information and message data information are labeled to obtain the tagged user attribute information and the tagged message data information.
[0016] In the above scheme, obtaining the negative credit record information corresponding to the first user identifier includes:
[0017] Based on the third communication interface, obtain the information on the unprocessed bad credit records corresponding to the first user identifier from the third target information platform;
[0018] The negative credit record information to be processed is obtained by standardizing the negative credit record information.
[0019] In the above scheme, the step of labeling the user attribute information and message data based on the negative credit record information to obtain the tagged user attribute information and the tagged message data includes:
[0020] Obtain keywords obtained during the standardization process of the negative credit record information to be processed;
[0021] Based on the keywords and target harassment information determination rules, the negative credit record information is labeled to determine a first tag for the negative credit record information; wherein, the first tag indicates whether the negative credit record information is harassment information.
[0022] Based on the first user identifier, a second correspondence is determined between the user attribute information and the message data information and the bad credit record information;
[0023] Based on the second correspondence and the first tag, the user attribute information and message data information are labeled to obtain the tagged user attribute information and the tagged message data information.
[0024] In the above scheme, the step of labeling the user attribute information and the message data information to obtain tagged user attribute information and tagged message data information includes:
[0025] Receive a second tag for the user attribute information and the message data information; wherein, the second tag is obtained after analyzing the interactive message corresponding to the user attribute information and the message data information; the second tag indicates whether the interactive message corresponding to the user attribute information and the message data information is harassing information;
[0026] The user attribute information and the message data information are labeled based on the second tag to obtain the tagged user attribute information and the tagged message data information.
[0027] In the above scheme, the step of training a model based on the user attribute information, the message data information, the tagged user attribute information, the tagged message data information, and the first correspondence to obtain a harassment information identification model includes:
[0028] The user attribute information, the message data information, the tagged user attribute information, and the tagged message data information are divided into two parts to obtain a training dataset and a test dataset; wherein, the training dataset includes first user attribute information, first message data information, first tagged user attribute information, and first tagged message data information, and the test dataset includes second user attribute information, second message data information, second tagged user attribute information, and second tagged message data information;
[0029] Based on the training dataset, the test dataset, and the first correspondence, the harassment information identification model is obtained by training the model using a neural network algorithm.
[0030] In the above scheme, the step of training the harassment information identification model using a neural network algorithm based on the training dataset, the test dataset, and the first correspondence to obtain the model includes:
[0031] A first initial harassment information identification model is obtained by algorithm modeling based on the training dataset and the first correspondence, and the accuracy of the first initial harassment information identification model is evaluated based on the test dataset and the first correspondence.
[0032] Based on the evaluation results and the first initial harassment information identification model, the harassment information identification model is determined.
[0033] In the above scheme, determining the harassment information identification model based on the evaluation results and the first initial harassment information identification model includes:
[0034] If the evaluation results indicate that the accuracy of the first initial harassment information identification model meets the target accuracy threshold, then the harassment information identification model is determined to be the first initial harassment information identification model.
[0035] If the evaluation results indicate that the accuracy of the first initial harassment information identification model does not meet the target accuracy threshold, the training dataset is grouped according to the message distribution traffic and the distribution frequency at a first grouping interval.
[0036] For each set of sub-training datasets, an algorithm model is performed based on each set of sub-training datasets and the first correspondence to determine a second initial harassment information identification model, and the accuracy of the second initial harassment information identification model is evaluated based on the test dataset and the first correspondence.
[0037] If the evaluation results indicate that the accuracy of the second initial harassment information identification model meets the target accuracy threshold, the harassment information identification model is determined based on the second initial harassment information identification model whose accuracy meets the target accuracy threshold.
[0038] The method in the above scheme further includes:
[0039] If the evaluation results indicate that the accuracy of the second initial harassment information identification model does not meet the target accuracy threshold, the training dataset is grouped according to the message distribution traffic and distribution frequency based on the second grouping interval, until the accuracy of the Nth initial harassment information identification model determined based on the grouped sub-training dataset and the first correspondence meets the target accuracy threshold, and the harassment information identification model is determined based on the Nth initial harassment information identification model whose accuracy meets the target accuracy threshold.
[0040] The method in the above scheme further includes:
[0041] Obtain the attribute information of the user to be monitored and the data information of the message to be monitored corresponding to the second user identifier of the object to be monitored;
[0042] Based on the user attribute information to be monitored, the data information of the message to be monitored, and the harassment information identification model, it is determined whether the interaction message of the object to be monitored is harassment information.
[0043] An information determining device, the device comprising: a processor, a memory, and a communication bus;
[0044] The communication bus is used to realize the communication connection between the processor and the memory;
[0045] The processor is used to execute an information determination program in memory to perform the following steps:
[0046] Obtain user attribute information and message data information corresponding to the first user identifier of the sample user; wherein, the message data information characterizes the degree of interaction of the sample user's messages;
[0047] The user attribute information and the message data information are labeled to obtain tagged user attribute information and tagged message data information; wherein, the tagged user attribute information and tagged message data information characterize whether the interactive message corresponding to the user attribute information and the message data information is harassing information;
[0048] Based on the first user identifier, a first correspondence is determined between the user attribute information and the data information of the message;
[0049] Based on the user attribute information, the message data information, the user attribute information with tags, the message data information with tags, and the first correspondence, a harassment information identification model is obtained through model training.
[0050] A computer-readable storage medium storing one or more programs that can be executed by one or more processors to implement the steps of the information determination method described above.
[0051] The information determination method, device, and computer-readable storage medium provided in the embodiments of this application acquire user attribute information and message data information corresponding to a first user identifier of a sample user; wherein, the message data information characterizes the interaction level of the sample user's messages; the user attribute information and message data information are labeled to obtain labeled user attribute information and labeled message data information; wherein, the labeled user attribute information and labeled message data information characterize whether the interactive message corresponding to the user attribute information and message data information is harassing information; based on the first user identifier, a first correspondence relationship is determined between the user attribute information and message data information; a harassing information identification model is obtained by training a model based on the user attribute information, message data information, labeled user attribute information, labeled message data information, and the first correspondence relationship; thus, only the user attribute information and message data information need to be acquired, that is, a harassing information identification model for determining whether interactive information is harassing information is generated based on the acquired user attribute information and message data information, without relying on the content of the interactive message to determine whether the interactive message is harassing information, reducing the dependence on the content of the interactive message, effectively protecting user privacy, and reducing the difficulty of information acquisition. Attached Figure Description
[0052] Figure 1 A flowchart illustrating an information determination method provided in an embodiment of this application;
[0053] Figure 2 A flowchart illustrating another information determination method provided in an embodiment of this application;
[0054] Figure 3 A flowchart illustrating the standardization and annotation processes of information in the information determination method provided in this application embodiment;
[0055] Figure 4 A flowchart illustrating the process of determining a harassment information identification model in an information determination method provided in this application embodiment;
[0056] Figure 5 A flowchart illustrating the process of determining a harassment information identification model in another information determination method provided in this application embodiment;
[0057] Figure 6 Another flowchart illustrating the determination of a harassment information identification model in a further information determination method provided in an embodiment of this application;
[0058] Figure 7 This is a schematic diagram of the structure of an information determination device provided in an embodiment of this application. Detailed Implementation
[0059] The technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings.
[0060] It should be understood that the phrases "embodiments of this application" or "foreign embodiments" throughout the specification mean that a specific feature, structure, or characteristic related to an embodiment is included in at least one embodiment of this application. Therefore, "embodiments of this application" or "in the foreign embodiments" appearing throughout the specification do not necessarily refer to the same embodiment. Furthermore, these specific features, structures, or characteristics can be combined in any suitable manner in one or more embodiments. In the various embodiments of this application, the sequence numbers of the above-described processes do not imply a sequential order of execution; the execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application. The sequence numbers of the above-described embodiments are merely descriptive and do not represent the superiority or inferiority of the embodiments.
[0061] Unless otherwise specified, any step in the embodiments of this application performed by the electronic device may be executed by the processor of the electronic device. It is also worth noting that the embodiments of this application do not limit the order in which the electronic device performs the following steps. Furthermore, the methods used to process data in different embodiments may be the same or different methods. It should also be noted that any step in the embodiments of this application can be executed independently by the electronic device; that is, when the electronic device performs any step in the following embodiments, it may not depend on the execution of other steps.
[0062] It should be understood that the specific embodiments described herein are for illustrative purposes only and are not intended to limit the scope of this application.
[0063] This application provides an information determination method, which can be applied to an information determination device, and refers to... Figure 1 As shown, the method includes the following steps:
[0064] S101. Obtain the user attribute information and message data information corresponding to the first user identifier of the sample user.
[0065] Among them, the data information of the message represents the degree of interaction of the sample users' messages.
[0066] In this embodiment, the information determination device can be a device with information processing and communication functions, and the user attribute information can include, but is not limited to, the user's work experience information and social activity information; the first user identifier can be the identity identifier of the sample user, and the identity identifier of the sample user can be used to characterize who the sample user refers to; the data information of the message can be the data information generated when the sample user interacts with other users.
[0067] In one feasible implementation, the user attribute information and message data information can be obtained by the information determining device from multiple data platforms based on the first user identifier.
[0068] S102. Label the user attribute information and message data information to obtain user attribute information with tags and message data information with tags.
[0069] Among them, the user attribute information with tags and the message data information with tags represent whether the interactive message corresponding to the user attribute information and the message data information is a nuisance.
[0070] In this embodiment, the information determination device analyzes user attribute information and message data, and annotates the user attribute information and message data based on the analysis results. If the analysis results indicate that the interaction information corresponding to the user attribute information and message data is harassing information, the information determination device can add a first sub-label to the user attribute information and message data indicating that the interaction information is harassing information, thus obtaining user attribute information with the first sub-label and message data information with the first sub-label. If the analysis results indicate that the interaction information corresponding to the user attribute information and message data is not harassing information, the information determination device can add a second sub-label to the user attribute information and message data indicating that the interaction information is not harassing information, thus obtaining user attribute information with the second sub-label and message data information with the second sub-label.
[0071] S103. Based on the first user identifier, determine the first correspondence between user attribute information and message data information.
[0072] In this embodiment of the application, the information determining device can establish a first correspondence between user attribute information and message data information based on user attribute information and message data information under a first user identifier (that is, associate user attribute information and message data information). The first user identifier includes at least the identity identifiers of multiple users.
[0073] In one feasible implementation, the information determining device can obtain user attribute information and message data information of sample user A based on the identity identifier of sample user A, and then associate the user attribute information and message data information of sample user A to obtain a first correspondence. Similarly, the information determining device can obtain user attribute information and message data information of sample user B based on the identity identifier of sample user B, and then associate the user attribute information and message data information of sample user B to obtain a first correspondence. In this way, a first correspondence can be obtained between the user attribute information and message data information of multiple sample users.
[0074] S104. Based on user attribute information, message data information, user attribute information with tags, message data information with tags, and the first correspondence, a model is trained to obtain a harassment information identification model.
[0075] In this embodiment of the application, the information determination device can first use the first part of the information from user attribute information, message data information, user attribute information with tags, and message data information with tags, and the first correspondence relationship to generate an initial harassment information identification model. Then, it can use the second part of the information from user attribute information, message data information, user attribute information with tags, and message data information with tags to verify the initial harassment information identification, and then obtain the harassment information identification model based on the verification result.
[0076] The information determination method provided in this application only needs to obtain user attribute information and message data information. That is, it generates a harassment information identification model based on the obtained user attribute information and message data information to determine whether the interactive information is harassment information, without relying on the content of the interactive message to determine whether the interactive message is harassment information. This reduces the dependence on the content of the interactive message, effectively protects the user's privacy, and reduces the difficulty of obtaining information.
[0077] Based on the foregoing embodiments, this application provides an information determination method, referring to... Figure 2 As shown, the method includes the following steps:
[0078] S201. The information determination device obtains the message distribution traffic and message distribution frequency corresponding to the first user identifier from the first target information platform based on the first communication interface, and performs standardization processing on the message distribution traffic and message distribution frequency to obtain the message distribution traffic and message distribution frequency.
[0079] In this embodiment, the information determination device can obtain the pending message distribution traffic and pending message distribution frequency corresponding to the first user identifier from the first target information platform based on the first communication interface. Then, it processes the pending message distribution traffic and pending message distribution frequency using at least one of the following algorithms: a preset language model extraction algorithm, the TextRank algorithm, a topic clustering algorithm, and a text summary keyword extraction algorithm, to obtain the message distribution traffic and message distribution frequency. The first communication interface includes at least the 5GM-01 instant messaging interface, abbreviated as 5GM-01 interface; the first target information platform includes at least Messaging as a Platform (Maap).
[0080] In one feasible implementation, the information determination device uses the 5GM-01 instant messaging interface and its interactive functions for various instant messaging methods to obtain the pending message distribution traffic and frequency from the Maap platform based on the first user's ID number, using point-to-point messages, group chats, chatbot messages, and other related signaling and media as carriers.
[0081] S202. The information determination device obtains the user attribute information to be processed corresponding to the first user identifier based on the second communication interface, and performs standardized processing on the user attribute information to be processed to obtain user attribute information.
[0082] User attribute information includes, but is not limited to: the industry in which the user works, the specific services the user performs, the registration entity of the user's company, and the information code of the user's company.
[0083] In this embodiment, the information determining device can obtain the user attribute information to be processed for the first user corresponding to the first user identifier from the second target information platform based on the second communication interface, and then perform standardization processing on the user attribute information to be processed to obtain user attribute information. The algorithm used for standardization processing of the user attribute information to be processed is the same as the algorithm used in S201, and will not be described again in this embodiment.
[0084] The second target information platform and the first target information platform may be the same platform; the second target information platform and the first target information platform may also be different platforms; the number of the first target information platform and / or the second target information platform may be at least one; the purpose of standardizing the user attribute information to be processed is to convert the format of the user attribute information to be processed so that it can be processed subsequently based on the format-converted user attribute information; the second communication interface may include, but is not limited to, the 5GM-03 interface and / or the 5GM-04 interface; the second target information platform may include, but is not limited to, the Maap platform.
[0085] In one feasible implementation, the information determining device can obtain information from the Maap platform based on the 5GM-03 and / or 5GM-04 interfaces. This information includes publicly available data on the user's work history (which does not involve user privacy or security), such as the user's industry, specific services performed, the registered entity of the user's company, and the company's information code. The 5GM-03 and 5GM-04 interfaces are primarily used to connect to the chatbot list server, allowing the querying of chatbot information from the server.
[0086] In this embodiment of the application, when the information determining device can obtain the bad credit record information corresponding to the first user identifier, S102 can be implemented by S203 and S204; when the information determining device cannot obtain the bad credit record information corresponding to the first user identifier, S102 can also be implemented by S205 and S206.
[0087] S203. The information determination device obtains the bad credit record information corresponding to the first user identifier.
[0088] In this embodiment of the application, the information determination device can also obtain the bad credit record information of the first user corresponding to the first user based on the first user identifier.
[0089] It should be noted that for different users, users with good credit scores do not have any bad credit records. In other words, there are two scenarios when obtaining bad credit records for the first user's identifier: if the first user has good credit, no bad credit records can be obtained (i.e., the first user has no bad credit records), and if the first user has poor credit, bad credit records can be obtained.
[0090] In the embodiments of this application, S203 can be implemented by S203a and S203b.
[0091] S203a, The information determination device obtains the information on the unprocessed bad credit records corresponding to the first user identifier from the third target information platform based on the third communication interface.
[0092] The number of third target information platforms can be at least one. A third target information platform can be the same as the first target information platform and / or the second target information platform, or it can be a different platform from the first target information platform and / or the second target information platform. The third communication interface can include, but is not limited to, the IS-01 interface; the third target information platform can include, but is not limited to, the 5GMC-Bad Message Monitoring Platform, the MaaP Platform-Bad Message Monitoring Platform, and the 5GMC-SMS Monitoring Platform.
[0093] In one feasible implementation, taking the example of multiple third target information platforms, where each third target information platform is different from the first target information platform and / or the second target data, the information determination device can obtain the corresponding pending bad credit record information from the 5GMC-Bad Message Monitoring Platform, the MaaP Platform-Bad Message Monitoring Platform, and the 5GMC-SMS Monitoring Platform based on the first user's ID card number through the IS-01 interface.
[0094] S203b: The information determination device standardizes the negative credit record information to be processed to obtain negative credit record information.
[0095] Specifically, the information determination device standardizes the negative credit record information to be processed to obtain negative credit record information. The algorithm used for standardizing the negative credit record information is the same as that used in S201 and S202; the purpose of standardizing the negative credit record information is to convert its format so that subsequent processing can be performed based on the converted user attribute information.
[0096] S204. The information determination device, based on the information of bad credit records, labels the user attribute information and message data information to obtain user attribute information with tags and message data information with tags.
[0097] In this embodiment of the application, the information determination device can analyze the pending bad credit records corresponding to the identification information of the first user based on the target harassment information determination rules, and label the user attribute information and message data information corresponding to the identification information of the first user based on the analysis results.
[0098] In the embodiments of this application, S204 can also be implemented by S204a, S204b, S204c and S204d.
[0099] S204a. Keywords obtained by the information determination device during the standardized processing of information on adverse credit records to be processed.
[0100] In the embodiments of this application, the algorithms used in the standardization process of the negative credit record information to be processed are the same in S201, S202 and S203b, and the keywords corresponding to the negative credit record information to be processed can be obtained in the end.
[0101] S204b: The information determination device, based on keyword and target harassment information judgment rules, labels the negative credit record information and determines the first label of the negative credit record information.
[0102] The first label indicates whether the negative credit record information is harassment information; the target harassment information determination rule is a rule generated based on multiple samples of harassment information and multiple samples of non-harassment information (normal information).
[0103] In this embodiment of the application, the information determination device can analyze and judge the keywords in the negative credit record information based on the target harassment information determination rule, and label the negative credit record information according to the judgment and analysis results; specifically, if the keywords in the negative credit record information that exceed the target proportion threshold are determined to be harassment information based on the target harassment information determination rule, then the first label is determined as a label indicating that the negative credit record information is harassment information; or, conversely, the first label is determined as a label indicating that the negative credit record information is not harassment information.
[0104] S204c. Based on the first user identifier, determine the second correspondence between user attribute information and message data information and bad credit record information.
[0105] In this embodiment of the application, the information determination device can associate the user attribute information and message data information under the first user identifier with the bad credit record information to obtain a second correspondence.
[0106] In one feasible implementation, the information determination device can associate user attribute information and message data information under user A's identity with negative credit record information based on user A's identity identifier to obtain a second correspondence; the information determination device can also associate user attribute information and message data information under user B's identity identifier with negative credit record information based on user B's identity identifier to obtain a second correspondence; in this way, a second correspondence can be obtained between the negative credit record information of multiple users and the message data information of multiple users in the sample users.
[0107] S204d, The information determination device, based on the second correspondence and the first tag, performs annotation processing on the user attribute information and message data information to obtain user attribute information with tags and message data information with tags.
[0108] Among them, the tags in the user attribute information with tags and the data information of the messages with tags can be the same as the first tag; or they can be different from the first tag. When they are different from the first tag, the tags in the user attribute information with tags and the data information of the messages with tags have the same meaning as the first tag. That is to say, when they are different from the first tag, they are only different in form. If the first tag represents harassment information, the tags in the user attribute information with tags and the data information of the messages with tags also represent harassment information.
[0109] In this embodiment of the application, the information determination device can determine whether the bad credit record information corresponding to the first user identifier is harassment information based on the first tag. Then, based on the second correspondence, it obtains the user attribute information and message data information corresponding to the bad credit record information under the first user identifier. Based on the first tag, it performs annotation processing on the user attribute information and the message data information of the first user to obtain the user attribute information with the first tag and the message data information with the first tag.
[0110] It should be noted that, as Figure 3 As shown, the first target information platform, the second target data platform, and the third target data platform are all Maap platforms. Based on the message distribution traffic and frequency, user attribute information, and bad credit record information to be processed obtained from the Maap platform, the message distribution traffic and frequency, user attribute information, and bad credit record information to be processed are then standardized to obtain message distribution traffic and frequency, user attribute information, and bad credit record information. The message distribution traffic and frequency, user attribute information, and bad credit record information are then labeled. The labeling process can be based on the bad credit record information or manually labeled using methods such as telephone follow-up.
[0111] S205, The information determination device receives a second tag for data information related to user attribute information and messages.
[0112] The second label is obtained by analyzing the interactive messages corresponding to the user attribute information and message data information; the second label indicates whether the interactive messages corresponding to the user attribute information and message data information are harassing messages.
[0113] In this embodiment, the information determining device can use a follow-up method to access the recipient of the interactive message to determine whether the recipient has given a valid response to the interactive message, thereby determining a second tag. If the recipient of the interactive message gives a valid response, the second tag indicates that the message is not harassing; if the recipient does not give a valid response, the second tag indicates that the message is harassing. It should be noted that during the follow-up, messages that are intermediary-type or otherwise related to messages that have been implicitly accepted by the user should be excluded.
[0114] S206. The information determination device performs annotation processing on the user attribute information and message data information based on the second tag to obtain user attribute information with tags and message data information with tags.
[0115] The meaning of the second tag is the same as that of the tags in the user attribute information and the data information of the message with tags.
[0116] In this embodiment of the application, if the second tag represents the interactive message corresponding to the user attribute information and the message data information as harassing information, then the tag in the user attribute information with the tag and the tag in the message data information with the tag can represent the user attribute information and the message data information as harassing information.
[0117] In one feasible implementation, such as Figure 4 As shown, the information determination device can obtain the message distribution traffic and frequency to be processed, the user attribute information to be processed, and the bad credit record information to be processed from the Maap platform. It can also perform standardization and annotation processing on the message distribution traffic and frequency to be processed, the user attribute information to be processed, and the bad credit record information to be processed. In the subsequent process, the information obtained after annotation processing is divided into two parts to obtain a training dataset and a test dataset. The harassment information identification model is determined based on the training dataset and the test dataset.
[0118] S207. The information determination device divides the user attribute information, message data information, and tagged user attribute information and tagged message data information into two parts to obtain a training dataset and a test dataset.
[0119] The training dataset includes first user attribute information, first message data information, first labeled user attribute information, and first labeled message data information. The test dataset includes second user attribute information, second message data information, second labeled user attribute information, and second labeled message data information.
[0120] In this embodiment of the application, the information determination device can first use a preset word vector model to convert the user attribute information, message data information, user attribute information with tags, and message data information with tags into a vectorized dataset. Then, according to a preset ratio, the converted user attribute information, message data information, user attribute information with tags, and message data information with tags are divided into two parts to obtain a training dataset and a test dataset.
[0121] In one feasible implementation, such as Figure 5 As shown, the information determination device can use the Word2Vec model (a related model used to generate word vectors) to convert the format of the information before and after annotation (user attribute information, message data information, user attribute information with tags, and message data information with tags) into vector forms of user attribute information, message data information, user attribute information with tags, and message data information with tags. Then, according to a preset ratio, the vector forms of user attribute information, message data information, user attribute information with tags, and message data information with tags are divided into two parts: a training dataset and a test dataset. In subsequent processes, a neural network algorithm is used to train the harassment information identification model based on the training dataset and the test dataset.
[0122] S208. The information determination device uses a neural network algorithm to train a model based on the training dataset, the test dataset, and the first correspondence to obtain a harassment information identification model.
[0123] In this embodiment of the application, the information determination device can train the training dataset and the first correspondence based on the neural network algorithm to generate an initial harassment information identification model. Then, it can determine whether the initial harassment information identification model is effective based on the test dataset, and generate a harassment information identification model based on the determination result.
[0124] In the embodiments of this application, S208 can be implemented by S208a and S208b.
[0125] S208a. The information determination device performs algorithm modeling based on the training dataset and the first correspondence to obtain a first initial harassment information identification model, and evaluates the accuracy of the first initial harassment information identification model based on the test dataset and the first correspondence.
[0126] In this embodiment, the information determination device can perform algorithmic modeling processing on the training dataset and the first correspondence based on a neural network algorithm to obtain a first initial harassment information identification model. The neural network algorithm includes at least one of the following: logistic regression algorithm, support vector machine algorithm, Naive Bayes algorithm, and decision tree algorithm.
[0127] The purpose of evaluating the accuracy of the first initial harassment information identification model based on the test dataset and the first correspondence is to determine whether the obtained first initial harassment information identification model is effective.
[0128] S208b, The information determination device determines the harassment information identification model based on the evaluation results and the first initial harassment information identification model.
[0129] The evaluation results are used to characterize whether the first initial harassment information identification model is effective.
[0130] In this embodiment of the application, S208b can also be implemented by step ae.
[0131] a. If the accuracy of the evaluation result characterizing the first initial harassment information identification model meets the target accuracy threshold, the information determination device determines the harassment information identification model as the first initial harassment information identification model.
[0132] In this embodiment of the application, if the accuracy of the first initial harassment information identification model meets the target accuracy threshold in the evaluation result, the information determination device can determine that the first initial harassment information identification model is valid, and then determine the first initial harassment information identification model as the harassment information identification model.
[0133] b. If the accuracy of the evaluation result characterizing the first initial harassment information identification model does not meet the target accuracy threshold, the information determination device groups the training dataset according to the first grouping interval based on message distribution traffic and distribution frequency.
[0134] The first grouping interval is determined based on the number of message distribution traffic and distribution frequency; in one feasible implementation, the first grouping interval can be the average of the number of message distribution traffic and distribution frequency.
[0135] In this embodiment, if the accuracy of the evaluation result representing the first initial harassment information identification model does not meet the target accuracy threshold, the information determination device can determine that the first initial harassment information identification model is invalid and needs to be re-determined. When the information determination device re-determines the first initial harassment information identification model, it first needs to group the message distribution traffic and distribution frequency according to the first grouping interval to obtain the groups corresponding to the message distribution traffic and distribution frequency to be processed. Then, based on the groups corresponding to the message distribution traffic and distribution frequency to be processed, the training dataset is grouped to obtain multiple sets of sub-training datasets.
[0136] like Figure 6As shown, in one feasible implementation, based on the evaluation results, if the accuracy meets the target accuracy threshold (i.e., the accuracy is valid), the harassment information identification model is determined as the first initial harassment information identification model. If the accuracy does not meet the target accuracy threshold (i.e., the accuracy is invalid), the training dataset can be grouped according to the first grouping interval based on message distribution traffic and distribution frequency, and a neural network algorithm can be used to train the harassment identification model based on the grouped information.
[0137] c. For each set of sub-training datasets, the information determination device performs algorithm modeling based on each set of sub-training datasets and the first correspondence to determine the second initial harassment information identification model, and evaluates the accuracy of the second initial harassment information identification model based on the test dataset and the first correspondence.
[0138] In this embodiment of the application, the information determination device can redetermine the model by performing algorithm modeling based on each set of sub-training datasets and the first correspondence, and obtain a second initial harassment information identification model. Then, the second initial harassment information identification model is evaluated to determine whether the second initial harassment information identification model is effective.
[0139] d. When the evaluation results indicate that the accuracy of the second initial harassment information identification model meets the target accuracy threshold, the information determination device determines the harassment information identification model based on the second initial harassment information identification model whose accuracy meets the target accuracy threshold.
[0140] In this embodiment, if the evaluation result indicates that the accuracy of the second initial harassment information identification model meets the target accuracy threshold, then the information determining device indicates that there is a valid second initial harassment information identification model among the determined multiple second initial harassment information. The harassment information identification model can then be determined based on the second initial harassment information identification model whose accuracy meets the target accuracy threshold. When the information determining device determines that the accuracy of the second initial harassment information identification model meets the target accuracy threshold, it stops determining other second initial harassment information identification models and designates the second initial harassment information identification model whose accuracy meets the target accuracy threshold as the harassment information identification model.
[0141] e. If the accuracy of the second initial harassment information identification model as represented by the evaluation results does not meet the target accuracy threshold, the information determination device groups the training dataset according to the second grouping interval based on message distribution traffic and distribution frequency until the accuracy of the Nth initial harassment information identification model determined based on the grouped sub-training dataset and the first correspondence meets the target accuracy threshold, and determines the harassment information identification model based on the Nth initial harassment information identification model whose accuracy meets the target accuracy threshold.
[0142] In this embodiment of the application, if the evaluation results indicate that the accuracy of the second initial harassment information identification model does not meet the target accuracy threshold, then multiple second initial harassment information identification models are determined to be invalid. In this case, the training dataset needs to be grouped based on message distribution traffic and distribution frequency by the second grouping interval to determine the Nth initial harassment information identification model. The process of determining the Nth initial harassment information identification model is the same as the process of determining the second initial harassment information identification model, and will not be described again in this embodiment of the application.
[0143] It should be noted that when the information determining device determines that the accuracy of the Nth initial harassment information identification model meets the target accuracy threshold, it stops determining other Nth initial harassment information identification models and uses the Nth initial harassment information identification model whose accuracy meets the target accuracy threshold as the harassment information identification model.
[0144] In this embodiment of the application, after obtaining the labeled message distribution traffic, message distribution frequency, user attribute information and bad credit record information, and processing the labeled message distribution traffic and frequency, user attribute information and bad credit record information, and processing the message distribution traffic and frequency in sequence, a harassment information identification model can be obtained.
[0145] Based on the foregoing embodiments, in other embodiments of this application, the method may further include the following steps:
[0146] S209. The information determination device obtains the user attribute information to be monitored and the data information of the message to be monitored corresponding to the second user identifier of the object to be monitored.
[0147] The attribute information of the user to be monitored includes, but is not limited to: the industry in which the user works, the specific services performed, the registered entity of the company, and the information code of the company. In this embodiment, the information determining device can obtain the attribute information and message data of the user to be monitored from the first target information platform based on the second user identifier of the monitored object.
[0148] S210. The information determination device determines whether the interactive messages of the monitored object are harassing messages based on the user attribute information to be monitored, the data information of the messages to be monitored, and the harassment information identification model.
[0149] In this embodiment of the application, the information determination device can input the attribute information of the user to be monitored and the data information of the message of the user to be monitored into the harassment information identification model, analyze the data information of the attribute information of the monitored user and the message of the user to be monitored based on the harassment information identification model, and determine whether the interaction message of the monitored object is harassment information based on the output result of the harassment information identification model.
[0150] It should be noted that the information determination device can also predict the relationship between the receiver and sender of the interactive message based on whether the interactive message of the monitored object is harassing. If the interactive message is harassing, it is determined that there is no relationship between the receiver and sender of the interactive message; if the interactive message is non-harassing, it is determined that there is a relationship between the receiver and sender of the interactive message.
[0151] The information determination method provided in this application embodiment only needs to obtain user attribute information and message data information. That is, it generates a harassment information identification model based on the obtained user attribute information and message data information to determine whether the interactive information is harassment information, without relying on the content of the interactive message to determine whether the interactive message is harassment information. This reduces the dependence on the content of the interactive message, effectively protects the user's privacy, and reduces the difficulty of obtaining information.
[0152] Based on the foregoing embodiments, embodiments of this application provide an information determining device, which can be applied to... Figures 1-2 In the information determination method provided in the corresponding embodiment, refer to Figure 7 As shown, the information determining device 3 may include: a processor 31, a memory 32, and a communication bus 33, wherein:
[0153] Communication bus 33 is used to realize the communication connection between processor 31 and memory 32;
[0154] The processor 31 is used to execute the information determination program in the memory 32 to perform the following steps:
[0155] Obtain the user attribute information and message data information corresponding to the first user identifier of the sample user; among which, the message data information characterizes the degree of interaction of the sample user's messages;
[0156] User attribute information and message data are labeled to obtain tagged user attribute information and tagged message data; the tagged user attribute information and tagged message data indicate whether the interactive message corresponding to the user attribute information and message data is harassing information;
[0157] Based on the first user identifier, determine the first correspondence between user attribute information and message data information;
[0158] The harassment information identification model is obtained by training the model based on user attribute information, message data information, user attribute information with tags, message data information with tags, and the first correspondence.
[0159] In other embodiments of this application, the processor 31 is used to execute the information determination program in the memory 32 to obtain user attribute information and message data information corresponding to the first user identifier of the sample user, in order to implement the following steps:
[0160] Based on the first communication interface, the message distribution traffic and message distribution frequency corresponding to the first user identifier are obtained from the first target information platform, and the message distribution traffic and message distribution frequency are standardized to obtain the message distribution traffic and message distribution frequency.
[0161] The data information of the message includes message distribution traffic and message distribution frequency;
[0162] Based on the second communication interface, the user attribute information corresponding to the first user identifier is obtained from the second target information platform, and the user attribute information is standardized to obtain user attribute information.
[0163] In other embodiments of this application, the processor 31 is used to execute the information determination program in the memory 32 to annotate the user attribute information and message data information to obtain user attribute information with tags and message data information with tags, so as to implement the following steps:
[0164] Obtain negative credit record information corresponding to the first user identifier;
[0165] Based on negative credit records, user attribute information and message data are labeled to obtain tagged user attribute information and tagged message data.
[0166] In other embodiments of this application, the processor 31 is used to execute the information determination program in the memory 32 to obtain the bad credit record information corresponding to the first user identifier, so as to implement the following steps:
[0167] Based on the third communication interface, obtain the information on the unprocessed bad credit records corresponding to the first user identifier from the third target information platform;
[0168] The negative credit record information is obtained by standardizing the processing of negative credit record information.
[0169] In other embodiments of this application, the processor 31 is used to execute the information determination program in the memory 32, based on bad credit record information, to annotate user attribute information and message data information to obtain tagged user attribute information and tagged message data information, in order to implement the following steps:
[0170] Keywords obtained during the standardized processing of negative credit record information to be processed;
[0171] Based on keyword and target harassment information judgment rules, negative credit record information is labeled to determine the first label of negative credit record information; wherein, the first label indicates whether negative credit record information is harassment information;
[0172] Based on the first user identifier, a second correspondence is determined between user attribute information and message data information and negative credit record information;
[0173] Based on the second correspondence and the first label, the user attribute information and message data information are labeled to obtain user attribute information with labels and message data information with labels.
[0174] In other embodiments of this application, the processor 31 is used to execute the information determination program in the memory 32, based on bad credit record information, to annotate user attribute information and message data information to obtain tagged user attribute information and tagged message data information, in order to implement the following steps:
[0175] Receive a second tag for user attribute information and message data information; wherein, the second tag is obtained after analyzing the interactive message corresponding to the user attribute information and message data information; the second tag indicates whether the interactive message corresponding to the user attribute information and message data information is harassing information;
[0176] The user attribute information and message data are labeled based on the second tag to obtain user attribute information with tags and message data with tags.
[0177] In other embodiments of this application, the processor 31 is used to execute the information determination program in the memory 32 to train a model based on user attribute information, message data information, user attribute information with tags, message data information with tags, and a first correspondence, to obtain a harassment information identification model, so as to implement the following steps:
[0178] The user attribute information, message data information, and labeled user attribute information and labeled message data information are divided into two parts to obtain a training dataset and a test dataset. The training dataset includes the first user attribute information, the first message data information, the first labeled user attribute information, and the first labeled message data information. The test dataset includes the second user attribute information, the second message data information, the second labeled user attribute information, and the second labeled message data information.
[0179] Based on the training dataset, the test dataset, and the first correspondence, a harassment information identification model is obtained by training the model using a neural network algorithm.
[0180] In other embodiments of this application, the processor 31 is used to execute the information determination program in the memory 32, which uses a neural network algorithm to train a model based on a training dataset, a test dataset, and a first correspondence to obtain a harassment information identification model, in order to achieve the following steps:
[0181] The first initial harassment information identification model is obtained by algorithm modeling based on the training dataset and the first correspondence, and the accuracy of the first initial harassment information identification model is evaluated based on the test dataset and the first correspondence.
[0182] Based on the evaluation results and the initial harassment information identification model, the harassment information identification model is determined.
[0183] In other embodiments of this application, processor 31 is used to execute the information determination program in memory 32 to determine the harassment information identification model based on the evaluation results and the first initial harassment information identification model, in order to implement the following steps:
[0184] If the accuracy of the first initial harassment information identification model meets the target accuracy threshold as represented by the evaluation results, the harassment information identification model is determined to be the first initial harassment information identification model.
[0185] If the accuracy of the first initial harassment information identification model does not meet the target accuracy threshold, the training dataset is grouped according to the first grouping interval based on message distribution traffic and distribution frequency.
[0186] For each set of sub-training datasets, an algorithm model is performed based on each set of sub-training datasets and the first correspondence to determine the second initial harassment information identification model, and the accuracy of the second initial harassment information identification model is evaluated based on the test dataset and the first correspondence.
[0187] If the evaluation results indicate that the accuracy of the second initial harassment information identification model meets the target accuracy threshold, the harassment information identification model is determined based on the second initial harassment information identification model whose accuracy meets the target accuracy threshold.
[0188] In other embodiments of this application, the processor 31 is used to execute an information determination program in the memory 32 to perform the following steps:
[0189] If the accuracy of the second initial harassment information identification model does not meet the target accuracy threshold in the evaluation results, the training dataset is grouped according to the second grouping interval based on message distribution traffic and distribution frequency until the accuracy of the Nth initial harassment information identification model determined based on the grouped sub-training dataset and the first correspondence meets the target accuracy threshold. The harassment information identification model is then determined based on the Nth initial harassment information identification model whose accuracy meets the target accuracy threshold.
[0190] In other embodiments of this application, the processor 31 is used to execute an information determination program in the memory 32 to perform the following steps:
[0191] Obtain the attribute information of the user to be monitored and the data information of the message to be monitored corresponding to the second user identifier of the object to be monitored;
[0192] Based on the user attribute information to be monitored, the data information of the messages to be monitored, and the harassment information identification model, it is determined whether the interaction messages of the monitored object are harassment information.
[0193] It should be noted that the specific implementation process of the steps executed by the processor in this embodiment can be referred to Figures 1-2 The implementation process of the information determination method provided in the corresponding embodiments will not be described in detail here.
[0194] The information determination device provided in this application embodiment only needs to obtain user attribute information and message data information. That is, it generates a harassment information identification model based on the obtained user attribute information and message data information to determine whether the interactive information is harassment information, without relying on the content of the interactive message to determine whether the interactive message is harassment information. This reduces the dependence on the content of the interactive message, effectively protects the user's privacy, and reduces the difficulty of obtaining information.
[0195] Based on the foregoing embodiments, embodiments of this application provide a computer-readable storage medium storing one or more programs, which can be executed by one or more processors to implement... Figures 1-2 The corresponding embodiments provide the steps of the information determination method.
[0196] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of hardware embodiments, software embodiments, or embodiments combining software and hardware aspects. Furthermore, this application can take the form of a computer program product implemented on one or more computer-usable storage media (including, but not limited to, disk storage and optical storage) containing computer-usable program code.
[0197] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0198] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0199] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0200] The above description is merely a preferred embodiment of this application and is not intended to limit the scope of protection of this application.
Claims
1. A method for determining information, characterized in that, The method includes: Obtain user attribute information and message data information corresponding to the first user identifier of the sample user; wherein, the message data information represents the message of the sample user and includes message distribution traffic and message distribution frequency; the user attribute information includes the user's work experience information and social activity information; The user attribute information and the message data information are labeled to obtain tagged user attribute information and tagged message data information; wherein, the tagged user attribute information and tagged message data information characterize whether the interactive message corresponding to the user attribute information and the message data information is harassing information; Based on the first user identifier, a first correspondence is determined between the user attribute information and the data information of the message; Based on the user attribute information, the message data information, the user attribute information with tags, the message data information with tags, and the first correspondence, a harassment information identification model is obtained through model training.
2. The method according to claim 1, characterized in that, The data information obtained, including the user attribute information and message information corresponding to the first user identifier of the sample user, includes: Based on the first communication interface, the pending message distribution traffic and pending message distribution frequency corresponding to the first user identifier are obtained from the first target information platform, and the pending message distribution traffic and pending message distribution frequency are standardized to obtain message distribution traffic and message distribution frequency. Based on the second communication interface, the user attribute information to be processed corresponding to the first user identifier is obtained from the second target information platform, and the user attribute information to be processed is standardized to obtain the user attribute information.
3. The method according to claim 1 or 2, characterized in that, The step of labeling the user attribute information and the message data information to obtain tagged user attribute information and tagged message data information includes: Obtain the negative credit record information corresponding to the first user identifier; Based on the negative credit record information, the user attribute information and message data information are labeled to obtain the tagged user attribute information and the tagged message data information.
4. The method according to claim 3, characterized in that, The step of obtaining the negative credit record information corresponding to the first user identifier includes: Based on the third communication interface, obtain the information on the unprocessed bad credit records corresponding to the first user identifier from the third target information platform; The negative credit record information to be processed is obtained by standardizing the negative credit record information.
5. The method according to claim 4, characterized in that, The step of labeling the user attribute information and message data based on the negative credit record information to obtain the tagged user attribute information and the tagged message data includes: Obtain keywords obtained during the standardization process of the negative credit record information to be processed; Based on the keywords and target harassment information determination rules, the negative credit record information is labeled to determine a first tag for the negative credit record information; wherein, the first tag indicates whether the negative credit record information is harassment information. Based on the first user identifier, a second correspondence is determined between the user attribute information and the message data information and the bad credit record information; Based on the second correspondence and the first tag, the user attribute information and message data information are labeled to obtain the tagged user attribute information and the tagged message data information.
6. The method according to claim 1 or 2, characterized in that, The step of labeling the user attribute information and the message data information to obtain tagged user attribute information and tagged message data information includes: Receive a second tag for the user attribute information and the message data information; wherein, the second tag is obtained after analyzing the interactive message corresponding to the user attribute information and the message data information; the second tag indicates whether the interactive message corresponding to the user attribute information and the message data information is harassing information; The user attribute information and the message data information are labeled based on the second tag to obtain the tagged user attribute information and the tagged message data information.
7. The method according to claim 2, characterized in that, The process of training a model based on the user attribute information, the message data information, the tagged user attribute information, the tagged message data information, and the first correspondence to obtain a harassment information identification model includes: The user attribute information, the message data information, the tagged user attribute information, and the tagged message data information are divided into two parts to obtain a training dataset and a test dataset; wherein, the training dataset includes first user attribute information, first message data information, first tagged user attribute information, and first tagged message data information, and the test dataset includes second user attribute information, second message data information, second tagged user attribute information, and second tagged message data information; Based on the training dataset, the test dataset, and the first correspondence, the harassment information identification model is obtained by training the model using a neural network algorithm.
8. The method according to claim 7, characterized in that, The step of training the harassment information identification model using a neural network algorithm based on the training dataset, the test dataset, and the first correspondence includes: A first initial harassment information identification model is obtained by algorithm modeling based on the training dataset and the first correspondence, and the accuracy of the first initial harassment information identification model is evaluated based on the test dataset and the first correspondence. Based on the evaluation results and the first initial harassment information identification model, the harassment information identification model is determined.
9. The method according to claim 8, characterized in that, The step of determining the harassment information identification model based on the evaluation results and the first initial harassment information identification model includes: If the evaluation results indicate that the accuracy of the first initial harassment information identification model meets the target accuracy threshold, then the harassment information identification model is determined to be the first initial harassment information identification model. If the evaluation results indicate that the accuracy of the first initial harassment information identification model does not meet the target accuracy threshold, the training dataset is grouped according to the message distribution traffic and the distribution frequency at a first grouping interval. For each set of sub-training datasets, an algorithm model is performed based on each set of sub-training datasets and the first correspondence to determine a second initial harassment information identification model, and the accuracy of the second initial harassment information identification model is evaluated based on the test dataset and the first correspondence. If the evaluation results indicate that the accuracy of the second initial harassment information identification model meets the target accuracy threshold, the harassment information identification model is determined based on the second initial harassment information identification model whose accuracy meets the target accuracy threshold.
10. The method according to claim 9, characterized in that, The method further includes: If the evaluation results indicate that the accuracy of the second initial harassment information identification model does not meet the target accuracy threshold, the training dataset is grouped according to the message distribution traffic and distribution frequency based on the second grouping interval, until the accuracy of the Nth initial harassment information identification model determined based on the grouped sub-training dataset and the first correspondence meets the target accuracy threshold, and the harassment information identification model is determined based on the Nth initial harassment information identification model whose accuracy meets the target accuracy threshold.
11. The method according to claim 1, characterized in that, The method further includes: Obtain the attribute information of the user to be monitored and the data information of the message to be monitored corresponding to the second user identifier of the object to be monitored; Based on the user attribute information to be monitored, the data information of the message to be monitored, and the harassment information identification model, it is determined whether the interaction message of the object to be monitored is harassment information.
12. An information determining device, characterized in that, The device includes: a processor, a memory, and a communication bus; The communication bus is used to realize the communication connection between the processor and the memory; The processor is used to execute an information determination program in memory to perform the following steps: Obtain user attribute information and message data information corresponding to the first user identifier of the sample user; wherein, the message data information characterizes the interaction level of the sample user's messages, and includes message distribution traffic and message distribution frequency; the user attribute information includes the user's work experience information and social activity information; The user attribute information and the message data information are labeled to obtain tagged user attribute information and tagged message data information; wherein, the tagged user attribute information and tagged message data information characterize whether the interactive message corresponding to the user attribute information and the message data information is harassing information; Based on the first user identifier, a first correspondence is determined between the user attribute information and the data information of the message; Based on the user attribute information, the message data information, the user attribute information with tags, the message data information with tags, and the first correspondence, a harassment information identification model is obtained through model training.
13. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores one or more programs, which can be executed by one or more processors to implement the steps of the information determination method as described in any one of claims 1 to 11.
Citation Information
Patent Citations
Method and device for recognizing spam
CN103970832A
Crank call identification method and device, and storage medium
CN109688275A