A method, device, electronic device and medium for model training and comment recognition

By using preset training sets and word-partition importance index values ​​in comment recognition to determine comment marks and training the binary classification model, the problem of insufficient accuracy of comment recognition in the prior art is solved, and higher comment recognition accuracy is achieved.

CN114911936BActive Publication Date: 2025-06-27CHINA TELECOM CORP LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210547432.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-05-18
Publication Date
2025-06-27
Estimated Expiration
2042-05-18

AI Technical Summary

Technical Problem

The prior art is difficult to effectively identify comments related to and unrelated to the target object published by users, resulting in insufficient accuracy of comment recognition.

Method used

By obtaining the preset training set, the importance index value of each participle is calculated, and the mark of the comment is determined based on the mark and index value, and then the binary classification model is trained to obtain the target model used for comment recognition.

Benefits of technology

Improves the accuracy of comment recognition, allowing the model to more accurately identify comments related to and unrelated to the target object.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114911936B_ABST
    Figure CN114911936B_ABST
Patent Text Reader

Abstract

Embodiments of the present disclosure provide a method, apparatus, electronic device, and medium for model training and comment recognition. The solution is as follows: Obtain a preset training set, where the preset training set includes multiple sample comments of sample objects and a first label corresponding to each sample comment; for each sample comment, calculate a first metric value corresponding to each word segment included in the sample comment, where the first metric value is used to indicate the importance of the word segment in the multiple sample comments; based on the first label corresponding to each sample comment and the first metric value corresponding to each word segment, determine a second label corresponding to each sample comment; use the multiple sample comments and the second label corresponding to each sample comment to train a preset binary classification model to obtain a target model for comment recognition. Through the technical solution provided by the embodiments of the present disclosure, a model for comment recognition is provided, thereby improving the accuracy of comment recognition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the technical field of big data processing, and particularly to a method, apparatus, electronic device and medium for model training and comment recognition. Background Art

[0002] In the Internet field, users can freely post comments on a certain target object. For example, users can post corresponding comments on the goods they purchase. For another example, users can post corresponding comments on the topic of a certain event.

[0003] Currently, among the comments posted by users, in addition to including comments related to the target object, there are also a large number of comments unrelated to the target object. Therefore, it is necessary to effectively identify the comments posted by users. Summary of the Invention

[0004] The purpose of the embodiments of the present disclosure is to provide a method, apparatus, electronic device and medium for model training and comment recognition, so as to provide a model for comment recognition, thereby improving the accuracy of comment recognition. The specific technical solutions are as follows:

[0005] The embodiments of the present disclosure provide a model training method, and the method includes:

[0006] Obtain a preset training set, where the preset training set includes multiple sample comments of a sample object and a first label corresponding to each sample comment, and the first label is: a first identifier indicating that the sample comment is related to the sample object, or a second identifier indicating that the sample comment is not related to the sample object;

[0007] For each sample comment, calculate a first index value corresponding to each word segment included in the sample comment, and the first index value is used to indicate the importance degree of the word segment in the multiple sample comments;

[0008] Based on the first label corresponding to each sample comment and the first index value corresponding to each word segment, determine a second label corresponding to each sample comment, and the second label is the first identifier or the second identifier;

[0009] Use the multiple sample comments and the second label corresponding to each sample comment to train a preset binary classification model to obtain a target model for comment recognition.

[0010] Optionally, the step of calculating, for each sample comment, a first index value corresponding to each word segment included in the sample comment includes:

[0011] For each sample comment, calculate the quotient of the number of occurrences of each word segment included in the sample comment in the multiple sample comments and the number of the word segment included in the multiple sample comments, as the word frequency of the word segment;

[0012] Based on the number of the multiple sample comments and the number of sample comments including each word segment, calculate the weight of the word segment;

[0013] Calculate the product of the word frequency corresponding to each word segment and the weight, as the first index value of the word segment.

[0014] Optionally, the step of determining the second label corresponding to each sample comment based on the first label corresponding to each sample comment and the first index value corresponding to each word segment includes:

[0015] Use the Bootstrapping algorithm to extract the word segments included in the multiple sample comments multiple times, and determine the second label corresponding to each sample comment according to the extracted word segments, the first label corresponding to each sample comment, and the first index value corresponding to each word segment.

[0016] Optionally, the step of training a preset binary classification model by using the multiple sample comments and the second label corresponding to each sample comment to obtain a target model for comment recognition includes:

[0017] For each sample comment, input the sample comment into a preset Support Vector Machine (SVM) classification model to obtain a third label corresponding to the sample comment;

[0018] According to the second label and the third label corresponding to each sample comment, calculate the loss value of the preset SVM classification model;

[0019] When the preset SVM classification model does not converge, adjust the parameters of the preset SVM classification model based on the loss value, and return to execute the step of inputting each sample comment into the preset SVM classification model to obtain a third label corresponding to the sample comment until the preset SVM classification model converges, and determine the preset SVM classification model at the current moment as the target model for comment recognition.

[0020] An embodiment of the present disclosure provides a comment recognition method, and the method further includes:

[0021] Obtain at least one comment to be recognized of a target object;

[0022] For each comment to be recognized, input the comment to be recognized into a pre-trained target model to obtain a fourth label of the comment to be recognized, where the target model is a binary classification model for comment recognition trained by the model training method described in any of the above items.

[0023] Optionally, the method further includes:

[0024] For each comment to be recognized with the fourth label being the second identifier, determine the target user who published the comment to be recognized;

[0025] Obtain the comments published by the target user within the first time period before the current time as the comments to be analyzed;

[0026] Based on the comments to be analyzed, calculate a second index value of the target user, where the second index value is used to indicate the probability that the target user maliciously publishes comments;

[0027] When the second index value is greater than a preset threshold, determine the target user as a malicious comment user.

[0028] Optionally, when the target object is a commodity, the step of calculating, based on the comments to be analyzed, a second index value of the target user, where the second index value is used to indicate the probability that the target user maliciously publishes comments, includes:

[0029] Based on the comments to be analyzed, calculate the ratio between the number of comments first published by the target user for different commodities within the first time period and the total number of comments published within the first time period as a first ratio;

[0030] Based on the comments to be analyzed, calculate the ratio between the number of commodities commented by the target user within the first time period and the number of commodities purchased by the target user within the first time period as a second ratio;

[0031] Based on the comments to be analyzed, calculate the ratio between the number of comments published by the target user within a preset time range and the total number of comments published by the target user within the first time period as a third ratio;

[0032] Based on the comments to be analyzed, calculate the ratio between the number of comments published by the target user within the second time period before the current time and the second time period as a fourth ratio, where the second time period is less than or equal to the first time period;

[0033] Calculate the weighted sum of the first ratio, the second ratio, the third ratio, and the fourth ratio as the second index value of the target user.

[0034] Optionally, the method further includes:

[0035] Execute a preset operation on the malicious comment user.

[0036] An embodiment of the present disclosure provides a model training device, the device includes:

[0037] A first acquisition module, configured to acquire a preset training set, the preset training set includes multiple sample comments of a sample object, and a first label corresponding to each sample comment, the first label is: a first identifier indicating that the sample comment is related to the sample object, or a second identifier indicating that the sample comment is not related to the sample object;

[0038] A first calculation module, configured to calculate, for each sample comment, a first index value corresponding to each word segment included in the sample comment, the first index value is used to indicate the importance degree of the word segment in the multiple sample comments;

[0039] A first determination module, configured to determine, based on the first label corresponding to each sample comment and the first index value corresponding to each word segment, a second label corresponding to each sample comment, the second label is the first identifier or the second identifier;

[0040] A training module, configured to use the multiple sample comments and the second label corresponding to each sample comment to train a preset binary classification model to obtain a target model for comment recognition.

[0041] Optionally, the first calculation module is specifically configured to calculate, for each sample comment, the quotient of the number of times the word segment included in the sample comment appears in the multiple sample comments and the number of the word segment included in the multiple sample comments as the word frequency of the word segment; calculate the weight of the word segment based on the number of the multiple sample comments and the number of sample comments including the word segment; calculate the product of the word frequency and the weight corresponding to each word segment as the first index value of the word segment.

[0042] Optionally, the first determination module is specifically configured to use the Bootstrapping algorithm to perform multiple extractions on the word segments included in the multiple sample comments, and determine the second label corresponding to each sample comment according to the extracted word segments, the first label corresponding to each sample comment, and the first index value corresponding to each word segment.

[0043] Optionally, the training module is specifically configured to, for each sample comment, input the sample comment into a preset SVM classification model to obtain a third label corresponding to the sample comment; calculate a loss value of the preset SVM classification model according to the second label and the third label corresponding to each sample comment; when the preset SVM classification model does not converge, adjust parameters of the preset SVM classification model based on the loss value, and return to execute the step of inputting the sample comment into the preset SVM classification model to obtain the third label corresponding to the sample comment for each sample comment, until when the preset SVM classification model converges, determine the preset SVM classification model at the current moment as the target model for comment recognition.

[0044] The embodiments of the present disclosure further provide a comment recognition device, and the device further includes:

[0045] A second acquisition module, configured to acquire at least one comment to be recognized of a target object;

[0046] A recognition module, configured to, for each comment to be recognized, input the comment to be recognized into a preselected and trained target model to obtain a fourth label of the comment to be recognized, where the target model is a binary classification model for comment recognition trained by the model training method described in any one of the above.

[0047] Optionally, the device further includes:

[0048] A second determination module, configured to, for each comment to be recognized with the fourth label being the second identifier, determine a target user who published the comment to be recognized;

[0049] A third acquisition module, configured to acquire comments published by the target user within a first time period before the current time as comments to be analyzed;

[0050] A second calculation module, configured to calculate a second index value of the target user based on the comments to be analyzed, where the second index value is used to indicate the probability that the target user maliciously publishes comments;

[0051] A third determination module, configured to determine the target user as a malicious comment user when the second index value is greater than a preset threshold.

[0052] Optionally, when the target object is a commodity, the second calculation module is specifically configured to calculate, based on the comments to be analyzed, a ratio between the number of comments first published by the target user for different commodities within the first time period and the total number of comments published within the first time period as a first ratio;

[0053] Based on the comment to be analyzed, calculate the ratio between the number of products commented by the target user within the first time period and the number of products purchased by the target user within the first time period, as the second ratio;

[0054] Based on the comment to be analyzed, calculate the ratio between the number of comments published by the target user within a preset time range and the total number of comments published by the target user within the first time period, as the third ratio;

[0055] Based on the comment to be analyzed, calculate the ratio between the number of comments published by the target user within the second time period before the current time and the second time period, as the fourth ratio, where the second time period is less than or equal to the first time period;

[0056] Calculate the weighted sum of the first ratio, the second ratio, the third ratio, and the fourth ratio as the second metric value of the target user.

[0057] Optionally, the device further includes:

[0058] An execution module, configured to perform a preset operation on the malicious comment user.

[0059] An embodiment of the present disclosure further provides an electronic device, including a processor, a communication interface, a memory, and a communication bus, where the processor, the communication interface, and the memory complete communication with each other through the communication bus;

[0060] The memory is used to store a computer program;

[0061] When the processor is configured to execute the program stored in the memory, it implements the steps of any of the above-mentioned model training methods.

[0062] An embodiment of the present disclosure further provides an electronic device, including a processor, a communication interface, a memory, and a communication bus, where the processor, the communication interface, and the memory complete communication with each other through the communication bus;

[0063] The memory is used to store a computer program;

[0064] When the processor is configured to execute the program stored in the memory, it implements the steps of any of the above-mentioned comment recognition methods.

[0065] An embodiment of the present disclosure further provides a computer-readable storage medium, where a computer program is stored in the computer-readable storage medium, and when the computer program is executed by a processor, it implements the steps of any of the above-mentioned model training methods.

[0066] An embodiment of the present disclosure also provides a computer-readable storage medium, in which a computer program is stored, and when the computer program is executed by a processor, the steps of any one of the above-mentioned comment recognition methods are implemented.

[0067] An embodiment of the present disclosure also provides a computer program product containing instructions, which when running on a computer, causes the computer to execute any one of the above-mentioned model training methods.

[0068] An embodiment of the present disclosure also provides a computer program product containing instructions, which when running on a computer, causes the computer to execute any one of the above-mentioned comment recognition methods.

[0069] Advantages of the embodiments of the present disclosure:

[0070] For the technical solution provided by the embodiment of the present disclosure, after obtaining a preset training set, that is, after obtaining a plurality of sample comments and the first label corresponding to each sample comment, for each sample comment, calculate the first index value corresponding to each word segment included in the sample comment, so as to determine the second label of each sample comment based on the first label corresponding to the sample comment and the first index value corresponding to each word segment, and then use each sample comment in the preset training set and the second label of each sample comment to train a preset binary classification model to obtain a target model for comment recognition. Compared with the related art, by using the first label of each sample comment in the preset training set and the first index value corresponding to each word segment included in each sample comment, the second label corresponding to each sample comment is re-determined, effectively improving the accuracy of the determined second label, so that the target model trained based on the second label can accurately identify comments related and unrelated to the target object, which effectively improves the accuracy of the trained target model, and thus improves the accuracy of subsequent comment recognition.

[0071] Of course, implementing any product or method of the present disclosure does not necessarily require achieving all the above-mentioned advantages at the same time. BRIEF DESCRIPTION OF THE DRAWINGS

[0072] In order to more clearly illustrate the technical solutions in the embodiments of the present disclosure or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present disclosure, and those of ordinary skill in the art can also obtain other embodiments according to these drawings.

[0073] Figure 1 It is the first flowchart of the model training method provided by the embodiment of the present disclosure;

[0074] Figure 2The second process schematic diagram of the model training method provided by the embodiments of the present disclosure;

[0075] Figure 3 The third process schematic diagram of the model training method provided by the embodiments of the present disclosure;

[0076] Figure 4 A schematic diagram of the execution process of the Bootstrapping algorithm provided by the embodiments of the present disclosure;

[0077] Figure 5 The fourth process schematic diagram of the model training method provided by the embodiments of the present disclosure;

[0078] Figure 6 The first process schematic diagram of the comment recognition method provided by the embodiments of the present disclosure;

[0079] Figure 7 The second process schematic diagram of the comment recognition method provided by the embodiments of the present disclosure;

[0080] Figure 8 The third process schematic diagram of the comment recognition method provided by the embodiments of the present disclosure;

[0081] Figure 9 The fourth process schematic diagram of the comment recognition method provided by the embodiments of the present disclosure;

[0082] Figure 10 A schematic diagram of the structure of the model training device provided by the embodiments of the present disclosure;

[0083] Figure 11 A schematic diagram of the structure of the comment recognition device provided by the embodiments of the present disclosure;

[0084] Figure 12 The first schematic diagram of the structure of the electronic device provided by the embodiments of the present disclosure;

[0085] Figure 13 The second schematic diagram of the structure of the electronic device provided by the embodiments of the present disclosure. Detailed implementation manners

[0086] Next, the technical solutions in the embodiments of the present disclosure will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present disclosure. Obviously, the described embodiments are only a part of the embodiments of the present disclosure, rather than all the embodiments. Based on the embodiments in the present disclosure, all other embodiments obtained by those of ordinary skill in the art based on the present disclosure belong to the scope of protection of the present disclosure.

[0087] In the related art, when training a model for comment recognition, since the labels corresponding to the sample comments in the training dataset used for training the model are all determined manually, there are certain errors, which will cause the trained model to be unable to accurately recognize comments related to the target object.

[0088] To solve the problems in the related art, an embodiment of the present disclosure provides a model training method, which is applied to an electronic device. The electronic device can be a mobile device or a server, etc. Here, no specific limitation is imposed on the electronic device. As Figure 1 shown, Figure 1 FIG. 1 is a first flowchart of the model training method provided by an embodiment of the present disclosure. The method includes the following steps.

[0089] Step S101, obtain a preset training set. The preset training set includes multiple sample comments of a sample object and a first label corresponding to each sample comment. The first label is: a first identifier indicating that the sample comment is related to the sample object, or a second identifier indicating that the sample comment is not related to the sample object.

[0090] Step S102, for each sample comment, calculate a first index value corresponding to each word segment included in the sample comment. The first index value is used to indicate the importance of the word segment in multiple sample comments.

[0091] Step S103, based on the first label corresponding to each sample comment and the first index value corresponding to each word segment, determine a second label corresponding to each sample comment. The second label is the first identifier or the second identifier.

[0092] Step S104, use multiple sample comments and the second label corresponding to each sample comment to train a preset binary classification model to obtain a target model for comment recognition.

[0093] In the embodiment of the present disclosure, although the above-mentioned electronic device can be various different devices, considering the deployment cost of the device, the above-mentioned electronic device can be a server providing services, such as the above-mentioned server.

[0094] By Figure 1In the method shown, after obtaining a preset training set, that is, after obtaining multiple sample comments and the first label corresponding to each sample comment, for each sample comment, calculate the first index value corresponding to each word segment included in the sample comment. Then, based on the first label corresponding to the sample comment and the first index value corresponding to each word segment, determine the second label of each sample comment. Furthermore, use each sample comment in the preset training set and the second label of each sample comment to train a preset binary classification model to obtain a target model for comment recognition. Compared with the related art, by using the first label of each sample comment in the preset training set and the first index value corresponding to each word segment included in each sample comment, re-determine the second label corresponding to each sample comment, effectively improving the accuracy of the determined second label. As a result, the target model trained based on this second label can accurately identify comments related and unrelated to the target object, which effectively improves the accuracy of the trained target model and thus improves the accuracy of subsequent comment recognition.

[0095] In the following, the embodiments of the present disclosure will be described through specific examples.

[0096] Regarding the above step S101, that is, obtaining a preset training set, the preset training set includes multiple sample comments of a sample object and the first label corresponding to each sample comment. The first label is: the first identifier indicating that the sample comment is related to the sample object, or the second identifier indicating that the sample comment is not related to the sample object.

[0097] In this step, the electronic device can obtain the comments corresponding to the sample object as sample comments for the sample object. For each obtained sample comment, the electronic device can determine whether the sample comment is a comment related to the sample object. If so, use the first identifier to label the sample comment to obtain the first label corresponding to the sample comment; if not, use the second identifier to label the first sample to obtain the first label corresponding to the sample comment.

[0098] In the embodiments of the present disclosure, the above sample object can be an enterprise, a commodity, a star topic, a social topic, etc. Here, no specific limitation is imposed on the above sample object. For the sake of understanding, only the sample object being a commodity will be described below as an example, which does not impose any limitation. In addition, the number of the above sample objects can be one or multiple. Here, no specific limitation is imposed on the number of the above sample objects. For the sake of understanding, only one sample object will be described below as an example, which does not impose any limitation.

[0099] The above sample comments can be the text content published by the user for the above sample object. For example, when the above sample object is a certain commodity, the sample comments can be the comments of the user on this commodity, such as the text content that the cost performance of the commodity is very high and the use process is very convenient. For different sample objects, the comments sent by the user are also different. Here, no specific limitation is made on the above sample comments.

[0100] In some embodiments, the comments related to the sample object in the above sample comments can be expressed as: the sample comments include statements about the attribute information of the sample object. The comments in the above sample comments that are not related to the sample object can be expressed as: the sample comments do not include statements about the attribute information of the sample object.

[0101] For ease of understanding, an example is given with the sample object being a commodity. The attribute information of this commodity includes but is not limited to the color, size, etc. of the commodity. For each sample comment, when the sample comment includes a statement related to the commodity attribute information, such as when the sample comment includes the statement "The color of the baby is really nice", the electronic device can determine that the sample comment is related to the sample comment. When the sample comment does not include a statement related to the commodity attribute information, the electronic device can determine that the sample comment is not related to the sample object.

[0102] In the embodiments of the present disclosure, according to the different sample objects, the corresponding attribute information of the sample object is also different. Therefore, the representation methods of whether the above sample comments are related to the sample object are also different. Here, no specific limitation is made on the representation methods of whether the above sample comments are related to the sample object.

[0103] The above first identifier can be 1, and the above second identifier can be 0. In addition, the above first identifier and second identifier can also be other values, such as 2, 3, 4, 5, etc. Here, no specific limitation is made on the above first identifier and second identifier. Among them, the sample comments marked with the above first identifier as the first identifier can be: positive samples in the binary classification model training set, and the sample comments marked with the above second identifier as the second identifier can be: negative samples in the binary classification model training set. The number of positive and negative samples in the above preset training set can be the same or different. Here, no specific limitation is made on the number of positive and negative samples in the above preset training set.

[0104] In some embodiments, when obtaining the above sample comments, the electronic device can obtain all the comments of the user for the sample object as the sample comments.

[0105] In some other embodiments, when obtaining the above sample comments, the electronic device may obtain all the comments of the user for the sample object, and preprocess the obtained comments to obtain the sample comments. Among them, the preprocessing of the comments may be expressed as: removing meaningless comments, or removing meaningless words in the comments. For example, removing comments that are all filled with modal particles, such as comments with the content "hahaha". Here, the preprocessing process of the comments will not be specifically described.

[0106] Regarding the above step S102, that is, for each sample comment, calculate the first index value corresponding to each word segment included in the sample comment, and the first index value is used to indicate the importance degree of the word segment in multiple sample comments.

[0107] In this step, for each sample comment in the preset training set, the electronic device may perform word segmentation processing on the sample comment to obtain multiple word segments, and for each word segment obtained by the word segmentation processing, calculate the importance degree value of the word segment in all sample comments as the first index value corresponding to the word segment.

[0108] Depending on the different sample comments, the number of word segments obtained by the electronic device for word segmentation processing of each sample comment is also different. Here, the number of word segments included in each sample comment is not specifically limited.

[0109] In some embodiments, for each sample comment, the electronic device may use the Term Frequency-Inverse Document Frequency (TF-IDF) algorithm to calculate the first index value corresponding to each word segment included in the sample comment.

[0110] In some embodiments, according to Figure 1 the method shown, the embodiments of the present disclosure also provide a model training method. As Figure 2 shown, Figure 2 is the second process schematic diagram of the model training method provided by the embodiments of the present disclosure. In Figure 2 the method shown, the above step S102 is refined into the following steps, that is, step S1021-step S1023.

[0111] Step S1021, for each sample comment, calculate the quotient of the number of times the word segment included in the sample comment appears in multiple sample comments and the number of the word segment included in multiple sample comments as the word frequency of the word segment.

[0112] In some embodiments, the electronic device may use the following formula to calculate the word frequency of the word segments included in each sample comment.

[0113]

[0114] Among them, TF i is the word frequency of the i-th word segmentation, and n i is the number of times the i-th word segmentation appears in multiple sample comments, and N is the number of the i-th word segmentation included in multiple sample comments.

[0115] Step S1022: Calculate the weight of each word segmentation based on the number of multiple sample comments and the number of sample comments including each word segmentation.

[0116] In some embodiments, the electronic device can use the following formula to calculate the weight of each word segmentation.

[0117]

[0118] Among them, IDF i is the weight of the i-th word segmentation, log represents the logarithmic operation, K is the number of multiple sample comments, and k i is the number of sample comments including the i-th word segmentation.

[0119] The weight of each of the above word segmentations is used to indicate the commonness of the word segmentation in all sample comments, and the magnitude of the weight of each word segmentation is inversely proportional to the commonness of the word segmentation. That is, for each word segmentation, when the word segmentation is more common in the above multiple sample comments, for example, when the frequency of the word segmentation appearing in the multiple sample comments is higher, the weight corresponding to the word segmentation is smaller. When the word segmentation is less common in the above multiple sample comments, for example, when the frequency of the word segmentation appearing in the multiple sample comments is lower, the weight corresponding to the word segmentation is larger.

[0120] In the embodiments of the present disclosure, the execution order of the above steps S1021 and S1022 is not specifically limited.

[0121] Step S1023: Calculate the product of the word frequency corresponding to each word segmentation and the weight as the first index value of the word segmentation.

[0122] In some embodiments, the electronic device can use the following formula to calculate the first index value of each word segmentation.

[0123] TF-IDF i = TF i * IDF i

[0124] Among them, TF-IDF i is the first index value of the i-th word segmentation.

[0125] Through the above steps S1021 - S1023, for each word segment obtained through word segmentation processing, the electronic device uses the above TF-IDF algorithm to calculate the TF value (i.e., the above word frequency) and IDF value (i.e., the above weight) corresponding to the word segment respectively, and thus determines the product of the TF value and the IDF value as the first index value of the word segment. This makes the first index value corresponding to each word segment highly correlated with the word frequency of the word segment in all sample comments and the weight of the word segment in all sample comments, and can accurately indicate the importance of each word segment in the above multiple sample comments, effectively ensuring the accuracy of the first index value corresponding to each word segment.

[0126] Regarding the above step S103, that is, based on the first label corresponding to each sample comment and the first index value corresponding to each word segment, determine the second label corresponding to each sample comment, and the second label is the first identifier or the second identifier.

[0127] For each sample comment in the above preset training set, the first label corresponding to the sample comment and the second label corresponding to the sample comment may be the same or different.

[0128] In some embodiments, according to the above Figure 1 shown method, the embodiments of the present disclosure also provide a model training method. As Figure 3 shown, Figure 3 This is the third process schematic diagram of the model training method provided by the embodiments of the present disclosure. In Figure 3 the method shown above, the above step S103 can be expressed as the following step, that is, step S1031.

[0129] Step S1031, use the Bootstrapping algorithm to extract the word segments included in multiple sample comments multiple times, and determine the second label corresponding to each sample comment according to the extracted word segments, the first label corresponding to each sample comment, and the first index value corresponding to each word segment.

[0130] The above Bootstrapping algorithm is: using limited sample data, through repeated sampling multiple times, to re-establish a new sample that is sufficient to represent the distribution of the population sample.

[0131] For ease of understanding, in combination with Figure 4 for illustration, Figure 4 This is a schematic diagram of the execution process of the Bootstrapping algorithm provided by the embodiments of the present disclosure.

[0132] Through the above step S102, the electronic device can determine the first index value corresponding to each word segment. When performing step S1031, the electronic device can generate a word feature set according to each word segment included in all sample comments and the first index value corresponding to each word segment. Among them, the first index value corresponding to each word segment can be the feature label of the word segment.

[0133] The electronic device can randomly extract multiple word features from the above word feature set with replacement. Each word feature includes a word segment and the feature label of the word segment. The electronic device can search for the sentence including the word segment in the word feature in the multiple sample comments included in the above preset training set to obtain the sentence where the feature appears. The sentence where the feature appears can be a complete sample comment or a partial sentence in the sample comment, such as the sentence including the word segment in the sample comment. Here, no specific limitation is imposed on the sentence where the feature appears.

[0134] The electronic device can regenerate the feature label corresponding to each word according to the extracted word feature and other word features in the sentence where the extracted word feature appears, and complete the process of feature label replacement.

[0135] The electronic device generates a text sampling pattern according to the replaced feature label and the word segment corresponding to each feature label.

[0136] The electronic device performs text pattern sampling according to the generated text sampling pattern, that is, extracts comments from the above multiple sample comments according to the text sampling pattern and determines the second label corresponding to the extracted comment to obtain a text pattern set. In addition, the electronic device can also generate a new word feature set according to the generated text sampling pattern and return to execute the step of randomly extracting multiple word features from the word feature set with replacement until the second label of each sample comment is determined.

[0137] Through the above step S1031, the electronic device can use the Bootstrapping algorithm to determine the second label corresponding to each sample comment, effectively improving the accuracy of the determined second label.

[0138] Regarding the above step S104, that is, using multiple sample comments and the second label corresponding to each sample comment to train a preset binary classification model to obtain a target model for comment recognition.

[0139] In some embodiments, according to the above Figure 1 shown method, the embodiments of the present disclosure also provide a model training method. As Figure 5 shown, Figure 5 is the fourth schematic diagram of the process of the model training method provided by the embodiments of the present disclosure. In Figure 5In the method shown above, the above step S104 can be refined into the following steps, namely step S1041 - step S1043.

[0140] Step S1041: For each sample comment, input the sample comment into a preset SVM classification model to obtain a third label corresponding to the sample comment.

[0141] In this step, for each sample comment, the electronic device can input the sample comment into a preset SVM classification model. The preset SVM classification model will predict whether the sample comment is relevant to the above sample object based on the feature information of the sample comment, and output a third label indicating whether the sample comment is relevant to the sample object. The electronic device obtains the third label corresponding to the sample comment output by the preset SVM classification model.

[0142] The above third label can be the above first identifier or second identifier.

[0143] In some embodiments, when inputting the above sample comment into the preset SVM classification model, the electronic device can input the feature vector corresponding to the sample comment into the preset SVM classification model. Alternatively, the preset SVM classification model includes a feature extraction module. The electronic device can input the above sample comment into the preset SVM classification model, and the feature extraction module in the preset SVM classification model extracts the feature information of the sample comment to obtain a feature vector, so as to perform comment recognition based on the feature vector. Here, the structure of the preset SVM classification model is not specifically limited.

[0144] Step S1042: Calculate the loss value of the preset SVM classification model according to the second label and the third label corresponding to each sample comment.

[0145] In this step, the electronic device can calculate the loss value of the preset SVM classification model according to the second label and the third label corresponding to each sample comment by using a preset loss function.

[0146] For example, for each sample comment, the electronic device can compare the second label corresponding to the sample comment with the third label. The electronic device can count the number of sample comments with different second labels and third labels as the loss value of the preset SVM classification model. In addition, the electronic device can also use a variety of loss functions, such as the mean squared error loss function (Mean Squared Error, MSE), cross-entropy loss function, etc., to calculate the loss value of the preset SVM classification model. Here, the calculation method of the above loss value is not specifically limited.

[0147] Step S1043: When the preset SVM classification model has not converged, adjust the parameters of the preset SVM classification model based on the loss value, and return to execute the step of inputting each sample comment into the preset SVM classification model to obtain the third label corresponding to the sample comment until the preset SVM classification model converges. Then, determine the preset SVM classification model at the current moment as the target model for comment recognition.

[0148] In the embodiments of the present disclosure, the electronic device can determine whether the preset SVM classification model converges based on the loss value of the preset SVM classification model or the number of training times of the preset SVM classification model. Here, the method for judging the convergence of the preset SVM classification model is not specifically limited.

[0149] When the above preset SVM classification model has not converged, the electronic device can adjust the parameters of the preset SVM classification model based on the above loss value, and return to execute the above step S1041, that is, return to execute the step of inputting each sample comment into the preset SVM classification model to obtain the third label corresponding to the sample comment.

[0150] When the above preset SVM classification model converges, the electronic device can determine that the training of the preset SVM classification model is completed. At this time, the electronic device can determine the preset SVM classification model at the current moment as the target model for comment recognition.

[0151] The parameters of the above preset SVM classification model can be the weights and bias amounts in the preset SVM classification model. The electronic device can use the gradient descent method or the reverse adjustment method to adjust the parameters of the preset SVM classification model. Here, the process of adjusting the parameters of the above preset SVM classification model is not specifically described.

[0152] Through the above steps S1041 - S1043, the electronic device can use the above multiple sample comments and the second label of each sample comment to train the preset SVM classification model to obtain the target model for comment recognition, effectively improving the accuracy of the trained target model, and thus improving the accuracy of later comment recognition using the target model.

[0153] In Figure 5 the illustrated embodiment, only the preset binary classification model is taken as an example of the preset SVM classification model for illustration. In addition, the electronic device can also use other binary classification models as the preset binary classification model. Here, the above preset binary classification model is not specifically limited.

[0154] Based on the same inventive concept, according to the model training method provided in the above embodiments of the present disclosure, the embodiments of the present disclosure also provide a comment recognition method. This method is applied to an electronic device. The electronic device for model training and the electronic device for comment recognition described above may be the same device or different devices. Here, no specific limitation is imposed on these two electronic devices. For ease of understanding, hereinafter, only the case where the electronic device for model training and the electronic device for comment recognition are the same device will be taken as an example for illustration, which does not impose any limitation. As Figure 6 shown, Figure 6 FIG. 1 is a first schematic flowchart of the comment recognition method provided by the embodiments of the present disclosure. This method includes the following steps.

[0155] Step S601, obtain at least one comment to be recognized of a target object.

[0156] In the embodiments of the present disclosure, the above target object may be an enterprise, a commodity, a star topic, a social topic, etc. Here, no specific limitation is imposed on the above target object. Additionally, the number of target objects may be one or multiple. Here, no specific limitation is imposed on the number of the above target objects.

[0157] In some embodiments, the above comment to be recognized may be all comments corresponding to the above target object, or a comment obtained by preprocessing all comments corresponding to the target. The preprocessing method may refer to the preprocessing method in step S101 above, and no specific description will be given here.

[0158] Step S602, for each comment to be recognized, input the comment to be recognized into a preselected and trained target model to obtain a fourth label of the comment to be recognized, where the target model is a binary classification model for comment recognition trained by the above model training method.

[0159] In this step, for each comment to be recognized, the electronic device may input the comment to be recognized into the target model trained in step S104 above to obtain a fourth label of the sample comment. The determination process of the fourth label may refer to the determination process of the third label above, and no specific description will be given here.

[0160] By Figure 6 the method shown, the electronic device can use the target model trained by the model training method provided by the embodiments of the present disclosure to recognize the comment to be recognized, so as to determine whether the comment to be recognized is relevant to the target object, effectively improving the accuracy of comment recognition while realizing comment recognition.

[0161] In some embodiments, according to the above Figure 6 shown method, the embodiments of the present disclosure also provide a comment recognition method. As Figure 7 shown, Figure 7Schematic diagram of the second process of the comment recognition method provided by the embodiments of the present disclosure. The method includes the following steps.

[0162] Step S701, obtain at least one comment to be recognized of the target object.

[0163] Step S702, for each comment to be recognized, input the comment to be recognized into a pre-trained target model to obtain a fourth label of the comment to be recognized, where the target model is a binary classification model for comment recognition trained by the above model training method.

[0164] The above Step S701 - Step S702 are the same as the above Step S601 - Step S602.

[0165] Step S703, for each comment to be recognized with the fourth label being the second identifier, determine the target user who published the comment to be recognized.

[0166] In this step, for each comment to be recognized, through the above Step S702, the electronic device can determine that the fourth label of the comment to be recognized is either the above first identifier or the above second identifier. At this time, for each comment to be recognized with the fourth label being the above second identifier, the electronic device can determine the user who published the comment to be recognized as the target user.

[0167] In the embodiments of the present disclosure, the above target user can be represented by the user's name, account name, identity identifier, etc. Here, the representation method of the above target user is not specifically limited.

[0168] The number of target users determined in the above Step S703 can be one or multiple. Here, the number of the above target users is not specifically limited. For ease of understanding, only one target user is taken as an example for illustration below, which does not play any limiting role.

[0169] Step S704, obtain the comments published by the target user within the first duration before the current time as the comments to be analyzed.

[0170] In this step, the electronic device can obtain all the comments published by the target user within the first duration before the current time according to the account information of the target user as the comments to be analyzed.

[0171] The above first duration can be one week, one month, two months, one year, etc. Here, the above first duration is not specifically limited. For ease of understanding, the first duration is taken as one month for illustration below, which does not play any limiting role.

[0172] Step S705, calculate a second metric value of the target user based on the comments to be analyzed, where the second metric value is used to indicate the probability that the target user maliciously publishes comments.

[0173] The above second index value is directly proportional to the probability that the target user maliciously posts comments. That is, the smaller the second index value of the target user, the smaller the probability that the target user maliciously posts comments; the larger the second index value of the target user, the greater the probability that the target user maliciously posts comments. The calculation process of the above second index value can be seen in the following description and will not be specifically described here.

[0174] Step S706, when the second index value is greater than a preset threshold, determine the target user as a malicious comment user.

[0175] In this step, after determining the second index value of the above target user, the electronic device can compare the second index value with a preset threshold. When the second index value is greater than the preset threshold, the electronic device can determine that the target user maliciously posts comments. At this time, the electronic device can determine the target user as a malicious comment user.

[0176] In some embodiments, when the second index value of the above target user is less than or equal to the preset threshold, the electronic device can determine that the target user does not maliciously post comments. At this time, the electronic device can determine the target user as a normal comment user.

[0177] Through the above steps S703 - S706, the electronic device can determine whether the target user maliciously posts comments according to the second index value of each target user, so as to determine malicious comment users, effectively improving the accuracy of the determined malicious comment users.

[0178] In some embodiments, when the above target object is a commodity, according to the above Figure 7 shown method, the embodiments of the present disclosure further provide a comment recognition method. As Figure 8 shown, Figure 8 is the third process schematic diagram of the comment recognition method provided by the embodiments of the present disclosure. In the Figure 8 shown method, step S705 is refined into the following steps, that is, steps S7051 - S7055.

[0179] Step S7051, based on the comment to be analyzed, calculate the ratio between the number of comments that the target user first posts for different commodities within the first time period and the total number of comments posted within the first time period as the first ratio.

[0180] In this step, the electronic device can count the number of comments that the target user first posts for different commodities within the first time period according to the comment to be analyzed by the target user, that is, the number of commodities associated with the comments posted by the target user, and count the total number of comments posted by the target user within the first time period. The electronic device can calculate the ratio between these two numbers to obtain the first ratio.

[0181] In some embodiments, the electronic device may use the following formula to calculate the above-mentioned first ratio P1.

[0182]

[0183] Where D1 is the number of comments first published by the target user for different products within the first time period, and D2 is the total number of comments published by the target user within the first time period.

[0184] The above-mentioned first ratio is used to indicate the proportion of the target user's first comments, and this first ratio is proportional to the probability of the user maliciously publishing comments.

[0185] Step S7052: Based on the comment to be analyzed, calculate the ratio between the number of products commented by the target user within the first time period and the number of products purchased by the target user within the first time period, as the second ratio.

[0186] In this step, the electronic device may, according to the comment to be analyzed of the target user, count the number of products associated with the comments published by the target user within the first time period to obtain the number of products commented by the target user within the first time period. The electronic device may also, according to the account information of the target user, obtain its purchase records within the first time period and determine the number of products purchased by the target user within the first time period. The electronic device calculates the ratio between these two quantities as the second ratio.

[0187] In some embodiments, the electronic device may use the following formula to calculate the above-mentioned second ratio P2.

[0188]

[0189] Where D3 is the number of products commented by the target user within the first time period, and D2 is the number of products purchased by the target user within the first time period. The above D3 is the same as the above D1.

[0190] The above-mentioned second ratio is used to indicate the proportion of comments made by the target user on the products it has purchased. This second ratio is proportional to the probability of the user maliciously publishing comments.

[0191] Step S7053: Based on the comment to be analyzed, calculate the ratio between the number of comments published by the target user within a preset time range and the total number of comments published by the target user within the first time period, as the third ratio.

[0192] In this step, the electronic device may, according to the comment to be analyzed of the target user, count the number of comments published by the target user within a preset time range and count the total number of comments published by the target user within the first time period. The electronic device may calculate the ratio between these two quantities as the third ratio.

[0193] In some embodiments, the electronic device may use the following formula to calculate the above-mentioned third ratio P3.

[0194]

[0195] Wherein, D5 is the number of comments published by the target user within a preset time range.

[0196] In the embodiments of the present disclosure, the above-mentioned preset time range may be the normal working time range of the user. For example, the preset time range may be: 9:00 - 11:00, 13:00 - 17:00. Here, the above-mentioned preset time range is not specifically limited.

[0197] The above-mentioned third ratio is used to indicate the proportion of comments published by the user during working hours, and this third ratio is proportional to the probability of the user publishing malicious comments.

[0198] Step S7054, based on the comment to be analyzed, calculate the ratio of the number of comments published by the target user within the second time period before the current time to the second time period, as the fourth ratio, where the second time period is less than or equal to the first time period.

[0199] In this step, the electronic device may count the number of comments published by the target user within the second time period before the current time according to the comment to be analyzed by the target user. The electronic device may calculate the ratio of this number of comments to the second time period as the fourth ratio.

[0200] In some embodiments, the electronic device may use the following formula to calculate the above-mentioned fourth ratio P4.

[0201]

[0202] Wherein, D6 is the number of comments published by the target user within the second time period before the current time, and T is the above-mentioned second time period.

[0203] In the embodiments of the present disclosure, the above-mentioned second time period may be less than or equal to the first time period. For ease of understanding, only the case where the second time period is less than the first time period is used for illustration. When the above-mentioned first time period is one month, the above-mentioned second time period may be one week; when the above-mentioned first time period is one week, the above-mentioned second time period may be one day or three days, etc.

[0204] The above-mentioned second time period may be set according to the first time period. Here, the above-mentioned second time period is not specifically limited.

[0205] The above-mentioned fourth ratio is used to indicate the frequency of the user publishing comments within the second time period. This fourth ratio is proportional to the probability of the user publishing malicious comments.

[0206] In the embodiments of the present disclosure, the execution order of the above steps S7051, S7052, S7053, and S7054 is not specifically limited.

[0207] Step S7055: Calculate the weighted sum of the first ratio, the second ratio, the third ratio, and the fourth ratio as the second index value of the target user.

[0208] In this step, the electronic device can determine the weights corresponding to the above first ratio, second ratio, third ratio, and fourth ratio, and calculate the weighted sum of the first ratio, second ratio, third ratio, and fourth ratio as the second index value of the target user.

[0209] In some embodiments, the electronic device can use the following formula to determine the weights corresponding to the first ratio, the second ratio, the third ratio, and the fourth ratio.

[0210]

[0211] Where, is the weight of the i-th ratio, and P i is the i-th ratio.

[0212] In some embodiments, the electronic device can use the following formula to calculate the second index value of the above target user.

[0213]

[0214] Where, P is the second index value of the target user.

[0215] Through the above steps S7051 - S7055, the electronic device can calculate the above first ratio, second ratio, third ratio, and fourth ratio based on the comment to be analyzed of the target user, and thus calculate the second index value of the target user based on the first ratio, second ratio, third ratio, and fourth ratio. While ensuring the accuracy of the calculated second index value, it can make the second index value highly correlated with whether the target user maliciously posts comments, thereby improving the accuracy of determining whether the target user is a malicious comment user based on the second index value.

[0216] In the embodiment shown in the above step S7051 - step S7055, only taking the target object as a commodity and calculating the second index value of the target user using the above first ratio, second ratio, third ratio, and fourth ratio as an example for illustration. When the above target object is other objects, for example, when the target object is a certain star topic, the above malicious comment users may be a large number of online water armies (i.e., hired online writers who post specific information for specific content on the Internet). At this time, considering that online water armies often belong to the same company, the electronic device can also obtain the IP address corresponding to each comment in this topic, and determine the proportion of comments with the same or similar IP addresses among the comments included in this topic, as a ratio required for calculating the above second index value. Therefore, according to the different above target objects and specific application scenarios, the calculation method of the second index value of the above target user is also different. Here, the calculation method of the above second index value is not specifically limited.

[0217] In some embodiments, according to the above Figure 7 shown method, the embodiments of the present disclosure also provide a comment recognition method. As Figure 9 shown, Figure 9 is the fourth process schematic diagram of the comment recognition method provided by the embodiments of the present disclosure. In Figure 9 the shown method, after step S706, step S707 may be included.

[0218] Step S707, perform a preset operation on the malicious comment user.

[0219] In this step, after determining through the above step S706 that the target user is a malicious comment user, the electronic device can perform a preset operation on this malicious comment user. Among them, the preset operation includes but is not limited to account monitoring, account function restriction, etc.

[0220] For example, the electronic device can perform account monitoring on the malicious comment user. When it detects that this malicious comment user posts a large number of malicious comments again, it can prompt that this malicious comment user cannot comment, or issue an alarm, etc.

[0221] In the embodiments of the present disclosure, according to different application scenarios, the electronic device can perform different preset operations on the malicious comment user. Here, the preset operations performed on the malicious comment user are not specifically limited.

[0222] Through the above step S707, the electronic device can perform a preset operation on the malicious comment user, which can realize the processing of the malicious comment user and reduce the probability of users making malicious comments.

[0223] In the embodiments of the present disclosure, in addition to performing a preset operation on malicious comment users, the electronic device can also perform business empowerment based on the comments made by the above-mentioned normal comment users. For example, if most of the normal comment users reflect that the color of the product is single, at this time, the merchant can enrich the color of the product to improve the purchasing power of users.

[0224] Based on the same inventive concept, according to the model training method provided in the above embodiments of the present disclosure, the embodiments of the present disclosure also provide a model training device. As Figure 10 shown, Figure 10 is a schematic structural diagram of a model training device provided by an embodiment of the present disclosure. The device includes the following modules.

[0225] The first acquisition module 1001 is configured to acquire a preset training set, where the preset training set includes multiple sample comments of sample objects and a first label corresponding to each sample comment, and the first label is: a first identifier indicating that the sample comment is related to the sample object, or a second identifier indicating that the sample comment is not related to the sample object;

[0226] The first calculation module 1002 is configured to, for each sample comment, calculate a first index value corresponding to each word segment included in the sample comment, and the first index value is used to indicate the importance degree of the word segment in multiple sample comments;

[0227] The first determination module 1003 is configured to determine a second label corresponding to each sample comment based on the first label corresponding to each sample comment and the first index value corresponding to each word segment, and the second label is the first identifier or the second identifier;

[0228] The training module 1004 is configured to use multiple sample comments and the second label corresponding to each sample comment to train a preset binary classification model to obtain a target model for comment recognition.

[0229] In some embodiments, the above-mentioned first calculation module 1002 may specifically be configured to, for each sample comment, calculate the quotient of the number of times each word segment included in the sample comment appears in multiple sample comments and the number of the word segment included in multiple sample comments as the word frequency of the word segment; calculate the weight of the word segment based on the number of multiple sample comments and the number of sample comments including each word segment; calculate the product of the word frequency corresponding to each word segment and the weight as the first index value of the word segment.

[0230] In some embodiments, the above-mentioned first determination module 1003 may specifically be configured to use the Bootstrapping algorithm to perform multiple extractions on the word segments included in multiple sample comments, and determine the second label corresponding to each sample comment according to the extracted word segments, the first label corresponding to each sample comment, and the first index value corresponding to each word segment.

[0231] In some embodiments, the above training module 1004 may specifically be configured to, for each sample comment, input the sample comment into a preset SVM classification model to obtain a third label corresponding to the sample comment; calculate a loss value of the preset SVM classification model according to the second label and the third label corresponding to each sample comment; when the preset SVM classification model has not converged, adjust the parameters of the preset SVM classification model based on the loss value, and return to execute the step of inputting the sample comment into the preset SVM classification model for each sample comment to obtain a third label corresponding to the sample comment, until when the preset SVM classification model converges, determine the preset SVM classification model at the current moment as the target model for comment recognition.

[0232] Based on the same inventive concept, according to the comment recognition method provided in the above embodiments of the present disclosure, embodiments of the present disclosure further provide a comment recognition device. As Figure 11 shown, Figure 11 is a schematic structural diagram of a comment recognition device provided by an embodiment of the present disclosure. The device includes the following modules.

[0233] A second acquisition module 1101, configured to acquire at least one comment to be recognized of a target object;

[0234] A recognition module 1102, configured to, for each comment to be recognized, input the comment to be recognized into a preselected and trained target model to obtain a fourth label of the comment to be recognized, where the target model is a binary classification model for comment recognition trained by the above model training method.

[0235] In some embodiments, the above comment recognition device may further include:

[0236] A second determination module, configured to determine a target user who publishes each comment to be recognized with a fourth label being a second identifier;

[0237] A third acquisition module, configured to acquire comments published by the target user within a first time period before the current time as comments to be analyzed;

[0238] A second calculation module, configured to calculate a second index value of the target user based on the comments to be analyzed, where the second index value is used to indicate the probability that the target user maliciously publishes comments;

[0239] A third determination module, configured to determine the target user as a malicious comment user when the second index value is greater than a preset threshold.

[0240] In some embodiments, when the target object is a commodity, the second calculation module may specifically be configured to calculate, based on the comment to be analyzed, the ratio between the number of comments first published by the target user for different commodities within the first time period and the total number of comments published by the target user within the first time period, as the first ratio;

[0241] calculate, based on the comment to be analyzed, the ratio between the number of commodities commented by the target user within the first time period and the number of commodities purchased by the target user within the first time period, as the second ratio;

[0242] calculate, based on the comment to be analyzed, the ratio between the number of comments published by the target user within the preset time range and the total number of comments published by the target user within the first time period, as the third ratio;

[0243] calculate, based on the comment to be analyzed, the ratio between the number of comments published by the target user within the second time period before the current time and the second time period, as the fourth ratio, where the second time period is less than or equal to the first time period;

[0244] calculate the weighted sum of the first ratio, the second ratio, the third ratio, and the fourth ratio, as the second index value of the target user.

[0245] In some embodiments, the above comment recognition device may further include:

[0246] an execution module, configured to perform a preset operation on a malicious comment user.

[0247] Through the device provided by the embodiments of the present disclosure, after obtaining a preset training set, that is, after obtaining a plurality of sample comments and the first label corresponding to each sample comment, for each sample comment, calculate the first index value corresponding to each word segment included in the sample comment, so as to determine the second label of each sample comment based on the first label corresponding to the sample comment and the first index value corresponding to each word segment, and then use each sample comment and the second label of each sample comment in the preset training set to train a preset binary classification model to obtain a target model for comment recognition. Compared with the related art, by using the first label of each sample comment in the preset training set and the first index value corresponding to each word segment included in each sample comment, the second label corresponding to each sample comment is re-determined, effectively improving the accuracy of the determined second label, so that the target model trained based on the second label can accurately identify comments related and unrelated to the target object, which effectively improves the accuracy of the trained target model, thereby improving the accuracy of subsequent comment recognition.

[0248] Based on the same inventive concept, according to the model training method provided by the embodiments of the present disclosure above, the embodiments of the present disclosure also provide an electronic device, which is used for model training, as Figure 12As shown, it includes a processor 1201, a communication interface 1202, a memory 1203, and a communication bus 1204. Among them, the processor 1201, the communication interface 1202, and the memory 1203 complete mutual communication through the communication bus 1204.

[0249] The memory 1203 is used to store computer programs.

[0250] When the processor 1201 is used to execute the program stored on the memory 1203, the following steps are implemented:

[0251] Obtain a preset training set, which includes multiple sample comments of a sample object and a first label corresponding to each sample comment. The first label is: a first identifier indicating that the sample comment is related to the sample object, or a second identifier indicating that the sample comment is not related to the sample object.

[0252] For each sample comment, calculate a first index value corresponding to each word segment included in the sample comment. The first index value is used to indicate the importance of the word segment in multiple sample comments.

[0253] Based on the first label corresponding to each sample comment and the first index value corresponding to each word segment, determine a second label corresponding to each sample comment. The second label is the first identifier or the second identifier.

[0254] Use multiple sample comments and the second label corresponding to each sample comment to train a preset binary classification model to obtain a target model for comment recognition.

[0255] Based on the same inventive concept, according to the comment recognition method provided in the above embodiments of the present disclosure, the embodiments of the present disclosure also provide an electronic device for comment recognition, as Figure 13 As shown, it includes a processor 1301, a communication interface 1302, a memory 1303, and a communication bus 1304. Among them, the processor 1301, the communication interface 1302, and the memory 1303 complete mutual communication through the communication bus 1304.

[0256] The memory 1303 is used to store computer programs.

[0257] When the processor 1301 is used to execute the program stored on the memory 1303, the following steps are implemented:

[0258] Obtain at least one comment to be recognized of a target object.

[0259] For each comment to be recognized, input the comment to be recognized into a preselected and trained target model to obtain a fourth label of the comment to be recognized, where the target model is a binary classification model for comment recognition trained by the above model training method.

[0260] Through the electronic device provided by the embodiments of the present disclosure, after obtaining a preset training set, that is, after obtaining a plurality of sample comments and the first label corresponding to each sample comment, for each sample comment, calculate the first index value corresponding to each word segment included in the sample comment, so as to determine the second label of each sample comment based on the first label corresponding to the sample comment and the first index value corresponding to each word segment, and then use each sample comment in the preset training set and the second label of each sample comment to train a preset binary classification model to obtain a target model for comment recognition. Compared with the related art, by using the first label of each sample comment in the preset training set and the first index value corresponding to each word segment included in each sample comment, the second label corresponding to each sample comment is re-determined, effectively improving the accuracy of the determined second label, so that the target model trained based on the second label can accurately identify comments related and unrelated to the target object, which effectively improves the accuracy of the trained target model, and thus improves the accuracy of subsequent comment recognition.

[0261] The communication bus mentioned in the above electronic device may be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This communication bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of simplicity, only a thick line is shown in the figure, but it does not mean that there is only one bus or one type of bus.

[0262] The communication interface is used for communication between the above electronic device and other devices.

[0263] The memory may include a Random Access Memory (RAM), and may also include a Non-Volatile Memory (NVM), such as at least one disk memory. In some embodiments, the memory may also be at least one storage device located far from the aforementioned processor.

[0264] The above-mentioned processor may be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), etc.; it may also be a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field-Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.

[0265] Based on the same inventive concept, according to the model training method provided in the above embodiments of the present disclosure, the embodiments of the present disclosure also provide a computer-readable storage medium, in which a computer program is stored. When the computer program is executed by a processor, the steps of any of the above model training methods are implemented.

[0266] Based on the same inventive concept, according to the comment recognition method provided in the above embodiments of the present disclosure, the embodiments of the present disclosure also provide a computer-readable storage medium, in which a computer program is stored. When the computer program is executed by a processor, the steps of any of the above comment recognition methods are implemented.

[0267] Based on the same inventive concept, according to the model training method provided in the above embodiments of the present disclosure, the embodiments of the present disclosure also provide a computer program product containing instructions. When it runs on a computer, the computer is caused to execute any of the model training methods in the above embodiments.

[0268] Based on the same inventive concept, according to the comment recognition method provided in the above embodiments of the present disclosure, the embodiments of the present disclosure also provide a computer program product containing instructions. When it runs on a computer, the computer is caused to execute any of the comment recognition methods in the above embodiments.

[0269] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present disclosure are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center by wire (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or a data center that includes one or more integrated available media. The available medium can be a magnetic medium (such as a floppy disk, a hard disk, a magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as a solid state disk (SSD)).

[0270] It should be noted that in this document, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including", or any other variation thereof is intended to cover non-exclusive inclusion, so that a process, method, article, or device that includes a series of elements includes not only those elements but also other elements that are not explicitly listed, or also includes elements that are inherent to such process, method, article, or device. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of additional identical elements in the process, method, article, or device that includes the element.

[0271] Each embodiment in this specification is described in a related manner. The same or similar parts between the embodiments can be referred to each other, and the differences between each embodiment and other embodiments are emphasized. In particular, for embodiments such as devices, electronic devices, computer-readable storage media, and computer program products, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiments.

[0272] The above are only the preferred embodiments of the present disclosure and are not intended to limit the protection scope of the present disclosure. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present disclosure are all included in the protection scope of the present disclosure.

Claims

1. A model training method, characterized in that, The method includes: Obtaining a preset training set, where the preset training set includes multiple sample comments of a sample object and a first label corresponding to each sample comment. The first label is: a first identifier indicating that the sample comment is related to the sample object, or a second identifier indicating that the sample comment is not related to the sample object. The sample comment is in text form; For each sample comment, calculating a first index value corresponding to each word segment included in the sample comment. The first index value is used to indicate the importance degree of the word segment in the multiple sample comments, and the first index value is obtained according to the product of the word frequency and the weight corresponding to each word segment; Based on the first label corresponding to each sample comment and the first index value corresponding to each word segment, determining a second label corresponding to each sample comment. The second label is the first identifier or the second identifier; Using the multiple sample comments and the second label corresponding to each sample comment to train a preset binary classification model to obtain a target model for comment recognition.

2. The method according to claim 1, characterized in that, The step of, for each sample comment, calculating a first index value corresponding to each word segment included in the sample comment includes: For each sample comment, calculating the quotient of the number of times a word segment included in the sample comment appears in the multiple sample comments and the number of sample comments including the word segment as the word frequency of the word segment; Based on the number of the multiple sample comments and the number of sample comments including each word segment, calculating the weight of the word segment; Calculating the product of the word frequency and the weight corresponding to each word segment as the first index value of the word segment.

3. The method according to claim 1, wherein The step of, based on the first label corresponding to each sample comment and the first index value corresponding to each word segment, determining a second label corresponding to each sample comment includes: Using the self-expanding Bootstrapping algorithm to extract the word segments included in the multiple sample comments multiple times, and determining the second label corresponding to each sample comment according to the extracted word segments, the first label corresponding to each sample comment, and the first index value corresponding to each word segment.

4. The method according to claim 1, wherein The step of, using the multiple sample comments and the second label corresponding to each sample comment to train a preset binary classification model to obtain a target model for comment recognition includes: For each sample comment, inputting the sample comment into a preset support vector machine (SVM) classification model to obtain a third label corresponding to the sample comment; Calculating a loss value of the preset SVM classification model according to the second label and the third label corresponding to each sample comment; When the preset SVM classification model has not converged, adjusting the parameters of the preset SVM classification model based on the loss value, and returning to execute the step of, for each sample comment, inputting the sample comment into the preset SVM classification model to obtain a third label corresponding to the sample comment until the preset SVM classification model converges, and determining the preset SVM classification model at the current moment as the target model for comment recognition.

5. A comment recognition method, characterized in that, The method further includes: Obtaining at least one comment to be recognized of a target object; For each comment to be recognized, input the comment to be recognized into a pre-trained target model to obtain a fourth label of the comment to be recognized, where the target model is a binary classification model for comment recognition trained by the method described in any one of claims 1-4.

6. The method according to claim 5, characterized in that The method further includes: For each comment to be recognized with the fourth label being the second identifier, determine the target user who published the comment to be recognized; Obtain the comments published by the target user within a first time period before the current time as the comments to be analyzed; Based on the comments to be analyzed, calculate a second metric value of the target user, where the second metric value is used to indicate the probability that the target user maliciously publishes comments; When the second metric value is greater than a preset threshold, determine the target user as a malicious comment user.

7. The method according to claim 6, wherein When the target object is a product, the step of calculating, based on the comments to be analyzed, a second metric value of the target user, where the second metric value is used to indicate the probability that the target user maliciously publishes comments, includes: Based on the comments to be analyzed, calculate the ratio of the number of comments first published by the target user for different products within the first time period to the total number of comments published within the first time period as a first ratio; Based on the comments to be analyzed, calculate the ratio of the number of products commented by the target user within the first time period to the number of products purchased by the target user within the first time period as a second ratio; Based on the comments to be analyzed, calculate the ratio of the number of comments published by the target user within a preset time range to the total number of comments published by the target user within the first time period as a third ratio; Based on the comments to be analyzed, calculate the ratio of the number of comments published by the target user within a second time period before the current time to the second time period as a fourth ratio, where the second time period is less than or equal to the first time period; Calculate the weighted sum of the first ratio, the second ratio, the third ratio, and the fourth ratio as the second metric value of the target user.

8. The method according to claim 7, characterized in that, The method further includes: Perform a preset operation on the malicious comment user.

9. A model training device, characterized in that, The device includes: A first acquisition module, configured to acquire a preset training set, where the preset training set includes multiple sample comments of a sample object and a first label corresponding to each sample comment, and the first label is: a first identifier indicating that the sample comment is related to the sample object, or a second identifier indicating that the sample comment is not related to the sample object, and the sample comment is in text form; A first calculation module, configured to calculate, for each sample comment, a first metric value corresponding to each word segment included in the sample comment, where the first metric value is used to indicate the importance of the word segment in the multiple sample comments, and the first metric value is obtained according to the product of the word frequency and weight corresponding to each word segment; A first determination module, configured to determine a second label corresponding to each sample comment based on the first label corresponding to each sample comment and the first metric value corresponding to each word segment, and the second label is the first identifier or the second identifier; A training module, configured to train a preset binary classification model by using the multiple sample comments and the second label corresponding to each sample comment, so as to obtain a target model for comment recognition.

10. The device according to claim 9, wherein The first calculation module is specifically configured to, for each sample comment, calculate the quotient of the number of occurrences of each word segment included in the sample comment in the multiple sample comments and the number of the word segment included in the multiple sample comments as the word frequency of the word segment; calculate the weight of the word segment based on the number of the multiple sample comments and the number of sample comments including each word segment. Calculate the product of the word frequency and the weight corresponding to each word segment as the first index value of the word segment.

11. The device according to claim 9, characterized in that, The first determination module is specifically configured to use the self-expanding Bootstrapping algorithm to extract the word segments included in the multiple sample comments multiple times, and determine the second label corresponding to each sample comment according to the extracted word segments, the first label corresponding to each sample comment, and the first index value corresponding to each word segment.

12. The device according to claim 9, characterized in that, The training module is specifically configured to, for each sample comment, input the sample comment into a preset support vector machine (SVM) classification model to obtain a third label corresponding to the sample comment; calculate the loss value of the preset SVM classification model according to the second label and the third label corresponding to each sample comment. When the preset SVM classification model does not converge, adjust the parameters of the preset SVM classification model based on the loss value, and return to execute the step of inputting the sample comment into the preset SVM classification model for each sample comment to obtain a third label corresponding to the sample comment, until when the preset SVM classification model converges, determine the preset SVM classification model at the current moment as the target model for comment recognition.

13. A comment recognition device, characterized in that, The device further includes: A second acquisition module, configured to acquire at least one comment to be recognized of a target object. A recognition module, configured to, for each comment to be recognized, input the comment to be recognized into a pre-trained target model to obtain a fourth label corresponding to the comment to be recognized, where the target model is a binary classification model for comment recognition trained by the method according to any one of claims 1-4.

14. The device according to claim 13, wherein The device further includes: A second determination module, configured to determine a target user who publishes each comment to be recognized with the fourth label being the second identifier. A third acquisition module, configured to acquire the comments published by the target user within a first time period before the current time as comments to be analyzed. A second calculation module, configured to calculate a second index value of the target user based on the comments to be analyzed, where the second index value is used to indicate the probability that the target user maliciously publishes comments. A third determination module, configured to determine the target user as a malicious comment user when the second index value is greater than a preset threshold.

15. The device according to claim 14, characterized in that, When the target object is a commodity, the second calculation module is specifically configured to calculate, based on the comments to be analyzed, the ratio of the number of comments first published by the target user for different commodities within the first time period to the total number of comments published within the first time period as a first ratio. Based on the comment to be analyzed, calculate the ratio between the number of products commented by the target user within the first time period and the number of products purchased by the target user within the first time period, as the second ratio; Based on the comment to be analyzed, calculate the ratio between the number of comments published by the target user within a preset time range and the total number of comments published by the target user within the first time period, as the third ratio; Based on the comment to be analyzed, calculate the ratio between the number of comments published by the target user within the second time period before the current time and the second time period, as the fourth ratio, where the second time period is less than or equal to the first time period; Calculate the weighted sum of the first ratio, the second ratio, the third ratio, and the fourth ratio as the second metric value of the target user.

16. The device according to claim 14, characterized in that, The device further includes: An execution module for performing a preset operation on the malicious comment user.

17. An electronic device, characterized in that, It includes a processor, a communication interface, a memory, and a communication bus. Among them, the processor, the communication interface, and the memory complete communication with each other through the communication bus; The memory is used to store computer programs; The processor is configured to implement the method steps described in any one of claims 1-4 or 5-8 when executing the programs stored on the memory.

18. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, and when the computer program is executed by the processor, it implements the method steps described in any one of claims 1-4 or 5-8.

Citation Information

Patent Citations

  • Context-aware approach to detection of short irrelevant texts

    CN105279146A

  • Spam comment training and recognition method and device, equipment and a readable storage medium

    CN109582788A