Content quality scoring model training method, scoring method and electronic device

By training the content quality scoring model, using text sample attribute characteristics and cross-processing, the problem of low scoring accuracy caused by manual evaluation is solved, and more efficient content quality evaluation is achieved.

CN115470933BActive Publication Date: 2025-08-12INST OF AUTOMATION CHINESE ACAD OF SCI
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211086311.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-09-06
Publication Date
2025-08-12
Estimated Expiration
2042-09-06

AI Technical Summary

Technical Problem

In the prior art, book content quality assessment relies on manual assessment, resulting in poor accuracy of scoring results.

Method used

By obtaining multiple text samples and their attribute features, cross-processing and inputting them into multiple initial classification regression trees, the model parameters are updated using the mean square variance loss function and the average loss function to train the content quality scoring model.

Benefits of technology

The accuracy and determination efficiency of content quality scoring results are improved, and more accurate content quality evaluation is achieved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115470933B_ABST
    Figure CN115470933B_ABST
Patent Text Reader

Abstract

The present invention provides a training method, a scoring method and an electronic device for a content quality scoring model. The method comprises: obtaining multiple text samples, and multiple attribute features and content quality scores corresponding to each text sample; cross-processing the multiple attribute features corresponding to each text sample to obtain multiple combined attribute features corresponding to the text sample; inputting the multiple attribute features and the multiple combined attribute features corresponding to each of the multiple text samples into multiple initial classification and regression trees of an initial content quality scoring model to obtain multiple predicted content quality scores corresponding to each text sample; and updating the model parameters of the multiple initial classification and regression trees according to the content quality scores and the multiple predicted content quality scores corresponding to each text sample. In this way, content quality assessment is performed on the content quality scoring model obtained through training, which can improve the accuracy of the content quality scoring results.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data processing technology, and in particular to a training method, a scoring method and an electronic device for a content quality scoring model. Background Art

[0002] Book content quality assessment is a way to proactively explore the value of books and is an important reference for subsequent book recommendations.

[0003] In the prior art, when evaluating the content quality of a book, experienced editors usually manually evaluate the book's topic selection, cultural connotation, creative form, etc. to obtain a content quality score.

[0004] However, the existing manual evaluation method will result in poor accuracy of the content quality rating results. Summary of the Invention

[0005] The present invention provides a training method, a scoring method and an electronic device for a content quality scoring model, which improve the accuracy of content quality scoring results.

[0006] The present invention provides a method for training a content quality scoring model, the method comprising:

[0007] Obtain multiple text samples, as well as multiple attribute features and content quality scores corresponding to each text sample.

[0008] For each of the text samples, a plurality of attribute features corresponding to the text sample are cross-processed to obtain a plurality of combined attribute features corresponding to the text sample.

[0009] The multiple attribute features and the multiple combined attribute features corresponding to each of the multiple text samples are respectively input into multiple initial classification and regression trees of the initial content quality scoring model to obtain multiple predicted content quality scores corresponding to each of the text samples.

[0010] According to the content quality scores corresponding to the text samples and the multiple predicted content quality scores, the model parameters of the multiple initial classification and regression trees are updated to obtain a trained content quality scoring model.

[0011] According to a content quality scoring model training method provided by the present invention, the updating of the model parameters of the multiple initial classification and regression trees according to the content quality scores corresponding to the text samples and the multiple predicted content quality scores includes:

[0012] For each text sample, a mean square error loss function is constructed between the content quality score corresponding to the text sample and each predicted content quality score in multiple predicted content quality scores to obtain multiple mean square error loss functions corresponding to the text sample.

[0013] The model parameters of the multiple initial classification and regression trees are updated according to the multiple mean square error loss functions corresponding to the text samples.

[0014] According to a content quality scoring model training method provided by the present invention, the updating of the model parameters of the multiple initial classification and regression trees according to the multiple mean square error loss functions corresponding to the text samples includes:

[0015] For each of the text samples, determine the weight of the initial classification and regression tree corresponding to the predicted content quality score used when constructing each mean square error loss function in the multiple mean square error loss functions corresponding to the text sample, and obtain the weight of the classification and regression tree corresponding to each mean square error loss function; and determine the first average loss function of the multiple mean square error loss functions based on the multiple mean square error loss functions and the weight of the classification and regression tree corresponding to each mean square error loss function.

[0016] The model parameters of the multiple initial classification and regression trees are updated according to the first average loss function corresponding to each text sample.

[0017] According to a content quality scoring model training method provided by the present invention, updating the model parameters of the multiple initial classification and regression trees according to the first average loss function corresponding to each text sample includes:

[0018] Determine a second average loss function corresponding to the plurality of text samples based on the first average loss function corresponding to each text sample.

[0019] According to the second average loss function, the model parameters of the multiple initial classification and regression trees are updated.

[0020] The present invention also provides a content quality scoring method, which includes:

[0021] Obtain a text to be scored and multiple attribute features corresponding to the text to be scored.

[0022] Cross-processing is performed on the multiple attribute features to obtain multiple combined attribute features corresponding to the text to be scored.

[0023] The multiple attribute features and the multiple combined attribute features are respectively input into multiple classification and regression trees of a content quality scoring model to obtain multiple content quality scores corresponding to the text to be scored; wherein the content quality scoring model is any of the content quality scoring models described above.

[0024] A target content quality score corresponding to the to-be-scored text is determined according to the multiple content quality scores.

[0025] According to a content quality scoring method provided by the present invention, determining a target content quality score corresponding to the to-be-scored text based on the multiple content quality scores includes:

[0026] The target content quality score corresponding to the text to be scored is determined according to the multiple content quality scores and the weights of the classification and regression trees corresponding to the content quality scores.

[0027] The present invention also provides a training device for a content quality scoring model, which may include:

[0028] The first acquisition unit is used to acquire multiple text samples, and multiple attribute features and content quality scores corresponding to each text sample.

[0029] The first processing unit is configured to perform cross-processing on a plurality of attribute features corresponding to each text sample to obtain a plurality of combined attribute features corresponding to the text sample.

[0030] The second processing unit is used to input the multiple attribute features and the multiple combined attribute features corresponding to each of the multiple text samples into multiple initial classification and regression trees of the initial content quality scoring model to obtain multiple predicted content quality scores corresponding to each of the text samples.

[0031] An updating unit is used to update the model parameters of the multiple initial classification and regression trees according to the content quality scores corresponding to the text samples and the multiple predicted content quality scores to obtain a trained content quality scoring model.

[0032] According to a training device for a content quality scoring model provided by the present invention, the updating unit is specifically used to construct, for each text sample, a mean square error loss function between the content quality score corresponding to the text sample and each predicted content quality score in a plurality of predicted content quality scores, thereby obtaining multiple mean square error loss functions corresponding to the text sample; and update the model parameters of the multiple initial classification and regression trees according to the multiple mean square error loss functions corresponding to the text samples.

[0033] According to a training device for a content quality scoring model provided by the present invention, the updating unit is specifically used to determine, for each text sample, the weight of the initial classification and regression tree corresponding to the predicted content quality score adopted when constructing each mean square error loss function in the multiple mean square error loss functions corresponding to the text sample, and obtain the weight of the classification and regression tree corresponding to each mean square error loss function; and determine the first average loss function of the multiple mean square error loss functions based on the multiple mean square error loss functions and the weights of the classification and regression trees corresponding to the each mean square error loss function; and update the model parameters of the multiple initial classification and regression trees based on the first average loss function corresponding to the text samples.

[0034] According to a training device for a content quality scoring model provided by the present invention, the updating unit is specifically used to determine the second average loss function corresponding to the multiple text samples based on the first average loss function corresponding to each text sample; and update the model parameters of the multiple initial classification and regression trees based on the second average loss function.

[0035] The present invention also provides a content quality scoring device, which may include:

[0036] The second acquisition unit is used to acquire the text to be scored and multiple attribute features corresponding to the text to be scored.

[0037] The third processing unit is configured to perform cross-processing on the multiple attribute features to obtain multiple combined attribute features corresponding to the text to be scored.

[0038] The fourth processing unit is configured to input the plurality of attribute features and the plurality of combined attribute features into a plurality of classification and regression trees of a content quality scoring model to obtain a plurality of content quality scores corresponding to the to-be-scored text; wherein the content quality scoring model is the content quality scoring model described above.

[0039] A determination unit is configured to determine a target content quality score corresponding to the to-be-scored text based on the multiple content quality scores.

[0040] According to a training device for a content quality scoring model provided by the present invention, the determination unit is specifically used to determine the target content quality score corresponding to the text to be scored based on the multiple content quality scores and the weights of the classification and regression trees corresponding to each content quality score.

[0041] The present invention also provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and runnable on the processor, wherein when the processor executes the program, the training method of the content quality scoring model as described in any one of the above-mentioned methods is implemented; or, the content quality scoring method as described in any one of the above-mentioned methods is implemented.

[0042] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements a training method for a content quality scoring model as described in any one of the above; or, implements a content quality scoring method as described in any one of the above.

[0043] The present invention also provides a computer program product, including a computer program, which, when executed by a processor, implements the training method of the content quality scoring model as described in any one of the above; or implements the content quality scoring method as described in any one of the above.

[0044] The present invention provides a training method, a scoring method, and an electronic device for a content quality scoring model, which obtains multiple text samples, as well as multiple attribute features and content quality scores corresponding to each text sample; cross-processes the multiple attribute features corresponding to each text sample to obtain multiple combined attribute features corresponding to the text sample; inputs the multiple attribute features and the multiple combined attribute features corresponding to each of the multiple text samples into multiple initial classification and regression trees of the initial content quality scoring model to obtain multiple predicted content quality scores corresponding to each text sample; updates the model parameters of the multiple initial classification and regression trees based on the content quality scores and the multiple predicted content quality scores corresponding to each text sample to obtain a trained content quality scoring model, so that content quality assessment can be performed using the content quality scoring model obtained through training to accurately determine the content quality scoring result, thereby improving the accuracy of the content quality scoring result; in addition, the efficiency of determining the content quality scoring result can also be effectively improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] In order to more clearly illustrate the technical solutions in the present invention or the prior art, a brief introduction is given below to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0046] Figure 1 A flowchart of a method for training a content quality scoring model provided by an embodiment of the present invention;

[0047] Figure 2A flowchart of a content quality scoring method provided by an embodiment of the present invention;

[0048] Figure 3 A schematic diagram of the structure of a training device for a content quality scoring model provided by an embodiment of the present invention;

[0049] Figure 4 A schematic structural diagram of a content quality scoring device provided by an embodiment of the present invention;

[0050] Figure 5 A schematic diagram of the physical structure of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0051] To make the objectives, technical solutions, and advantages of the present invention more clear, the technical solutions of the present invention will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the embodiments described are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the present invention.

[0052] In the embodiments of the present invention, "at least one" refers to one or more, and "more than one" refers to two or more. "And / or" describes the relationship between associated objects, indicating that three possible relationships exist. For example, "A and / or B" can mean: A exists alone, A and B exist simultaneously, or B exists alone. A and B can be singular or plural. In the text descriptions of the present invention, the character " / " generally indicates that the associated objects are in an "or" relationship.

[0053] The technical solutions provided by the embodiments of the present invention can be applied in data processing scenarios, particularly in the context of text content quality assessment. Taking book content quality assessment as an example, in the prior art, experienced editors typically manually assess the book's topic, cultural connotations, and creative style to produce a content quality score. However, the use of existing manual assessment methods results in poor accuracy in the resulting content quality score.

[0054] In order to accurately determine the content quality scoring result and thus improve the accuracy of the content quality scoring result, an embodiment of the present invention provides a training method for a content quality scoring model, which obtains multiple text samples, and multiple attribute features and content quality scores corresponding to each text sample; for each text sample, cross-processes the multiple attribute features corresponding to the text sample to obtain multiple combined attribute features corresponding to the text sample; inputs the multiple attribute features and the multiple combined attribute features corresponding to each of the multiple text samples into multiple initial classification and regression trees of the initial content quality scoring model respectively to obtain multiple predicted content quality scores corresponding to each text sample; updates the model parameters of the multiple initial classification and regression trees according to the content quality score corresponding to each text sample and the multiple predicted content quality scores to obtain a trained content quality scoring model, so that content quality assessment can be performed through the content quality scoring model obtained through training to accurately determine the content quality scoring result, thereby improving the accuracy of the content quality scoring result; in addition, the efficiency of determining the content quality scoring result can also be effectively improved.

[0055] The following specific embodiments will be used to describe in detail the training method of the content quality scoring model provided by the present invention. It is understood that the following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments.

[0056] Figure 1 This is a flow chart of a method for training a content quality scoring model provided by an embodiment of the present invention. The method for training a content quality scoring model can be executed by software and / or hardware devices. For example, see Figure 1 As shown, the training method of the content quality scoring model may include:

[0057] S101: Acquire multiple text samples, and multiple attribute features and content quality scores corresponding to each text sample.

[0058] The content quality score can be understood as a label corresponding to a text sample. For example, multiple attribute features may include author, publisher, category, publication year, number of pages, comment content, number of reads, number of clicks, etc., which can be set according to actual needs.

[0059] For example, when obtaining multiple text samples, multiple initial text samples can be obtained by searching, but not limited to, Douban Reading, JD Reading, Yuewen Reading, Baidu Encyclopedia, external knowledge graphs, etc.; and the multiple initial texts searched are further screened, for example, from the multiple initial text samples, initial text samples with large amounts of plagiarized content are eliminated, and / or initial text samples with a reading number less than a reading threshold, and / or initial text samples with a number of ratings less than a number threshold, etc. The specific settings can be made according to actual needs.

[0060] After obtaining the filtered initial text sample through the above-mentioned screening operation, the training set, test set, and validation set can be divided into a ratio of 7:2:1. From the filtered initial text sample, multiple text samples are determined as training sets for training the content quality scoring model. It is understood that in the embodiment of the present invention, the distribution of the corresponding data sets of the training set, test set, and validation set is consistent. When the distribution of the data sets corresponding to the training set, test set, and validation set is consistent, the correlation between the attribute features corresponding to the data set and the target value can be measured, and the correlation between the attribute features can be measured. Methods include but are not limited to the Pearson correlation coefficient, Spearman rank correlation coefficient, Kendall correlation coefficient, information entropy, etc. For example, for the text in the test set and the training set, the distribution ratio of the attribute feature scores in each interval is consistent; to determine the correlation between the attribute features, the Pearson correlation coefficient between the comment sentiment and the score, the Pearson correlation coefficient between the author's education level and the score, etc. can be calculated.

[0061] When obtaining multiple attribute features corresponding to text samples, considering that different text samples come from different sources, the corresponding attribute features of the obtained text samples are different. For example, for text samples from Douban, the corresponding attribute features obtained usually include author, publisher, category, publication year, number of pages, comment content, etc.; for text samples from Baidu Encyclopedia, the corresponding attribute features obtained usually include author, date of birth, representative works, achievements, position, education, etc.; for text samples from JD shopping websites, the corresponding attribute features obtained usually include comments, sales volume, price, etc. Therefore, in order to make the attribute features corresponding to the obtained text samples consistent, for text samples from different sources, relevant attribute features collected from other different sources can also be used to supplement and verify the attribute features of the text samples, so that the corresponding attribute features of text samples obtained from different sources remain consistent, providing a basis for the subsequent training of the content quality scoring model.

[0062] Exemplarily, after obtaining multiple attribute features corresponding to a text sample, the attribute features can be further optimized. For example, for the attribute feature "comment", invalid words such as "of", "@", "to", etc. can be removed from the comment. Specifically, it can be set according to actual needs. Here, the embodiments of the present invention only take removing invalid words from the comment as an example for illustration, but it does not mean that the embodiments of the present invention are only limited to this.

[0063] After obtaining multiple attribute features corresponding to each text sample, in order to make the attribute features corresponding to the text sample more rich and diverse, the multiple attribute features corresponding to the text sample can be further cross-processed, that is, execute the following S102:

[0064] S102. For each text sample, cross-process the multiple attribute features corresponding to the text sample to obtain multiple combined attribute features corresponding to the text sample.

[0065] Exemplarily, when cross-processing the multiple attribute features corresponding to a text sample, a rule-based method, a statistic-based method, or a neural network-based method can be used to explicitly or implicitly multiply two or more attribute features to perform a non-linear transformation on the sample space of the attribute features and increase the non-linear ability of the content quality scoring model used for subsequent training. For example, the attribute feature category and the attribute feature author can be explicitly combined to obtain a combined feature; the attribute feature reading volume and the attribute feature click volume can be explicitly combined to obtain a combined feature, etc. Specifically, it can be set according to actual needs.

[0066] After respectively obtaining multiple attribute features and multiple combined attribute features corresponding to each of the multiple text samples, the multiple attribute features and multiple combined attribute features corresponding to each of the multiple text samples can be input into multiple initial classification regression trees of the initial content quality scoring model, that is, execute the following S103:

[0067] S103. Input the multiple attribute features and multiple combined attribute features corresponding to each of the multiple text samples into multiple initial classification regression trees of the initial content quality scoring model respectively to obtain multiple predicted content quality scores corresponding to each text sample.

[0068] Exemplarily, when inputting the multiple attribute features and multiple combined attribute features corresponding to each of the multiple text samples into multiple initial classification regression trees of the initial content quality scoring model respectively, the multiple attribute features and multiple combined attribute features corresponding to each of the multiple text samples can be vectorized first to obtain their respective corresponding feature vectors, and the respective corresponding feature vectors are input into multiple initial classification regression trees of the initial content quality scoring model respectively.

[0069] For example, when vectorizing the multiple attribute features and multiple combined attribute features corresponding to each of the multiple text samples, methods including but not limited to one-hot encoding, pre-trained word vectors, self-trained word vectors, and a Bidirectional Encoder Representations from Transformer (BERT) model can be used to vectorize the multiple attribute features and multiple combined attribute features corresponding to each of the multiple text samples. For example, for attribute feature categories, assuming that there are science fiction categories, novel categories, and social categories, and the value corresponding to the novel category is 1 and the values corresponding to other categories are 0, then the encoding corresponding to the novel text is (0, 1, 0), which can be used as the feature vector corresponding to the attribute feature category; alternatively, the BERT model can be directly used to determine the feature vector corresponding to the attribute feature category. The specific setting can be made according to actual needs, and the embodiments of the present invention do not impose specific limitations here.

[0070] Since the initial content quality scoring model includes multiple initial classification and regression trees, for each text sample, multiple attribute features and multiple combined attribute features corresponding to the text sample are input into multiple initial classification and regression trees respectively. Each initial classification and regression tree will output the predicted content quality score corresponding to the text sample, that is, multiple predicted content quality scores corresponding to the text sample can be output through multiple initial classification and regression trees.

[0071] S104 : updating the model parameters of the multiple initial classification and regression trees according to the content quality score corresponding to each text sample and the multiple predicted content quality scores, so as to obtain a trained content quality scoring model.

[0072] For example, in an embodiment of the present invention, when the model parameters of multiple initial classification and regression trees are updated according to the content quality score corresponding to each text sample and multiple predicted content quality scores, for each text sample, a mean square error loss function between the content quality score corresponding to the text sample and each predicted content quality score in the multiple predicted content quality scores can be constructed respectively to obtain multiple mean square error loss functions corresponding to the text sample; then, according to the multiple mean square error loss functions corresponding to each text sample, the model parameters of the multiple initial classification and regression trees are updated, that is, automatic parameter optimization is performed to obtain a trained content quality scoring model.

[0073] For example, when updating the model parameters of multiple initial classification and regression trees based on multiple mean square error loss functions corresponding to each text sample, for each text sample, we can first determine the weight of the initial classification and regression tree corresponding to the predicted content quality score used when constructing each mean square error loss function in the multiple mean square error loss functions corresponding to the text sample, and obtain the weight of the classification and regression tree corresponding to each mean square error loss function; and determine the first average loss function of the multiple mean square error loss functions based on the multiple mean square error loss functions and the weight of the classification and regression tree corresponding to each mean square error loss function; and then update the model parameters of the multiple initial classification and regression trees based on the first average loss function corresponding to each text sample.

[0074] For example, when updating the model parameters of multiple initial classification and regression trees according to the first average loss function corresponding to each text sample, the second average loss function corresponding to the multiple text samples can be first determined according to the first average loss function corresponding to each text sample; and then the model parameters of the multiple initial classification and regression trees are updated according to the second average loss function to obtain a trained content quality scoring model.

[0075] It can be seen that in an embodiment of the present invention, multiple text samples, as well as multiple attribute features and content quality scores corresponding to each text sample are obtained; for each text sample, the multiple attribute features corresponding to the text sample are cross-processed to obtain multiple combined attribute features corresponding to the text sample; the multiple attribute features and multiple combined attribute features corresponding to each of the multiple text samples are respectively input into multiple initial classification and regression trees of the initial content quality scoring model to obtain multiple predicted content quality scores corresponding to each text sample; according to the content quality score corresponding to each text sample and the multiple predicted content quality scores, the model parameters of the multiple initial classification and regression trees are updated to obtain a trained content quality scoring model, so that content quality assessment can be performed through the content quality scoring model obtained through training, and the content quality scoring result can be accurately determined, thereby improving the accuracy of the content quality scoring result; in addition, the efficiency of determining the content quality scoring result can also be effectively improved.

[0076] above Figure 1 The embodiment shown in the figure describes in detail how to train a content quality scoring model in the embodiment of the present invention. Figure 2 The illustrated embodiment describes the application process of the content quality scoring model.

[0077] Figure 2 This is a flow chart of a content quality scoring method provided by an embodiment of the present invention. The content quality scoring method can be executed by software and / or hardware devices. For example, see Figure 2 As shown, the content quality scoring method may include:

[0078] S201: Obtain a text to be scored and a plurality of attribute features corresponding to the text to be scored.

[0079] For example, when obtaining the text to be rated, the text to be rated can be obtained by searching, including but not limited to, Douban Reading, JD Reading, Yuewen Reading, Baidu Encyclopedia, external knowledge graphs, etc., and the specific settings can be made according to actual needs.

[0080] When obtaining multiple attribute features corresponding to the text to be scored, taking into account the different sources of different texts, the corresponding attribute features of the corresponding texts are different. For example, for texts from Douban, the corresponding attribute features obtained usually include author, publisher, category, publication year, number of pages, comment content, etc.; for texts from Baidu Encyclopedia, the corresponding attribute features obtained usually include author, date of birth, representative works, achievements, position, education, etc.; for texts from JD shopping websites, the corresponding attribute features obtained usually include comments, sales volume, price, etc. Therefore, in order to make the attribute features corresponding to the obtained text consistent with the attribute features corresponding to the text samples used in the training of the above-mentioned content quality scoring model, for texts from different sources, relevant attribute features collected from other different sources can also be used to supplement and verify the attribute features of the text, so that the attribute features corresponding to the texts obtained from different sources are consistent with the attribute features corresponding to the text samples used in the training of the above-mentioned content quality scoring model.

[0081] After obtaining multiple attribute features corresponding to the text to be scored, in order to make the attribute features corresponding to the text to be scored more abundant and diversified, the multiple attribute features corresponding to the text to be scored can be further cross-processed, that is, the following S202 is executed:

[0082] S202: Cross-process the multiple attribute features to obtain multiple combined attribute features corresponding to the text to be scored.

[0083] For example, when cross-processing multiple attribute features, a rule-based method, a statistical method, or a neural network-based method can be used to explicitly or implicitly multiply two or more attribute features to perform a nonlinear transformation on the sample space of the attribute features. The specific settings can be made according to actual needs.

[0084] After obtaining the multiple attribute features and the multiple combined attribute features corresponding to the text to be scored, the multiple attribute features and the multiple combined attribute features can be input into the multiple classification and regression trees of the content quality scoring model, that is, the following S103 is executed:

[0085] S203: Input the multiple attribute features and the multiple combined attribute features into multiple classification and regression trees of the content quality scoring model respectively to obtain multiple content quality scores corresponding to the text to be scored.

[0086] Among them, the content quality scoring model is the above Figure 1 The content quality scoring model in the illustrated embodiment.

[0087] For example, when multiple attribute features and multiple combined attribute features are respectively input into multiple classification and regression trees of the content quality scoring model, the multiple attribute features and multiple combined attribute features can be first vectorized to obtain their respective corresponding feature vectors, and then the respective corresponding feature vectors are respectively input into the multiple classification and regression trees of the content quality scoring model.

[0088] S204: Determine a target content quality score corresponding to the text to be scored based on the multiple content quality scores.

[0089] Since the content quality scoring model includes multiple classification and regression trees, the multiple attribute features and multiple combined attribute features corresponding to the text to be scored are respectively input into the multiple classification and regression trees of the content quality scoring model. Each classification and regression tree will output the predicted content quality score corresponding to the text to be scored, that is, multiple classification and regression trees can output multiple predicted content quality scores corresponding to the text to be scored.

[0090] For example, when determining the target content quality score corresponding to the text to be scored based on multiple content quality scores, a weighted average can be performed based on the multiple content quality scores and the weights of the classification regression trees corresponding to each content quality score to determine the target content quality score corresponding to the text to be scored.

[0091] It can be seen that in the embodiment of the present invention, by obtaining the text to be scored and multiple attribute features corresponding to the text to be scored; and cross-processing the multiple attribute features to obtain multiple combined attribute features corresponding to the text to be scored; the multiple attribute features and the multiple combined attribute features are respectively input into multiple classification and regression trees of the content quality scoring model to obtain multiple content quality scores corresponding to the text to be scored; and then, based on the multiple content quality scores, the target content quality score corresponding to the text to be scored is determined. In this way, by using the content quality scoring model obtained through training to perform content quality assessment, the content quality score result can be accurately determined, thereby improving the accuracy of the content quality score result; in addition, the efficiency of determining the content quality score result can be effectively improved.

[0092] The content quality scoring model training device and content quality scoring device provided by the present invention are described below. The content quality scoring model training device described below and the content quality scoring model training method described above can be referenced to each other, and the content quality scoring device and the content quality scoring method described above can be referenced to each other.

[0093] Figure 3 A structural diagram of a training device for a content quality scoring model provided by an embodiment of the present invention, for example, see Figure 3 As shown, the training device 30 of the content quality scoring model may include:

[0094] The first acquisition unit 301 is configured to acquire a plurality of text samples, and a plurality of attribute features and content quality scores corresponding to each text sample.

[0095] The first processing unit 302 is configured to perform cross-processing on multiple attribute features corresponding to each text sample to obtain multiple combined attribute features corresponding to the text sample.

[0096] The second processing unit 303 is used to input the multiple attribute features and the multiple combined attribute features corresponding to the multiple text samples into the multiple initial classification and regression trees of the initial content quality scoring model to obtain multiple predicted content quality scores corresponding to each text sample.

[0097] The updating unit 304 is configured to update the model parameters of the multiple initial classification and regression trees according to the content quality score corresponding to each text sample and the multiple predicted content quality scores, so as to obtain a trained content quality scoring model.

[0098] Optionally, the updating unit 304 is specifically used to construct, for each text sample, a mean square error loss function between the content quality score corresponding to the text sample and each predicted content quality score in multiple predicted content quality scores, to obtain multiple mean square error loss functions corresponding to the text sample; and update the model parameters of multiple initial classification and regression trees according to the multiple mean square error loss functions corresponding to each text sample.

[0099] Optionally, the updating unit 304 is specifically used to determine, for each text sample, the weight of the initial classification and regression tree corresponding to the predicted content quality score used when constructing each mean square error loss function in the multiple mean square error loss functions corresponding to the text sample, and obtain the weight of the classification and regression tree corresponding to each mean square error loss function; and determine the first average loss function of the multiple mean square error loss functions based on the multiple mean square error loss functions and the weights of the classification and regression trees corresponding to each mean square error loss function; and update the model parameters of the multiple initial classification and regression trees based on the first average loss function corresponding to each text sample.

[0100] Optionally, the updating unit 304 is specifically configured to determine a second average loss function corresponding to multiple text samples based on the first average loss function corresponding to each text sample; and update the model parameters of multiple initial classification and regression trees based on the second average loss function.

[0101] The content quality scoring model training device 30 provided in an embodiment of the present invention can execute the technical solution of the content quality scoring model training method in any of the above embodiments. Its implementation principle and beneficial effects are similar to the implementation principle and beneficial effects of the content quality scoring model training method. Please refer to the implementation principle and beneficial effects of the content quality scoring model training method, and no further details will be given here.

[0102] Figure 4 This is a structural diagram of a content quality scoring device provided by an embodiment of the present invention. For example, see Figure 4 As shown, the content quality scoring device 40 may include:

[0103] The second acquisition unit 401 is used to acquire the text to be scored and a plurality of attribute features corresponding to the text to be scored.

[0104] The third processing unit 402 is configured to perform cross-processing on the multiple attribute features to obtain multiple combined attribute features corresponding to the text to be scored.

[0105] The fourth processing unit 403 is used to input multiple attribute features and multiple combined attribute features into multiple classification and regression trees of the content quality scoring model respectively to obtain multiple content quality scores corresponding to the text to be scored; wherein the content quality scoring model is the content quality scoring model shown in the above embodiment.

[0106] The determining unit 404 is configured to determine a target content quality score corresponding to the to-be-scored text according to the multiple content quality scores.

[0107] Optionally, the determining unit 404 is specifically configured to determine a target content quality score corresponding to the text to be scored according to the multiple content quality scores and the weights of the classification and regression trees corresponding to the content quality scores.

[0108] The content quality scoring device 40 provided in an embodiment of the present invention can implement the technical solution of the content quality scoring method in any of the above embodiments. Its implementation principle and beneficial effects are similar to those of the content quality scoring method. Please refer to the implementation principle and beneficial effects of the content quality scoring method, and no further details will be given here.

[0109] Figure 5 A schematic diagram of the physical structure of an electronic device provided by an embodiment of the present invention is shown in FIG. Figure 5As shown, the electronic device may include: a processor 501, a communication interface 502, a memory 503 and a communication bus 504, wherein the processor 501, the communication interface 502 and the memory 503 communicate with each other via the communication bus 504. The processor 501 may call the logic instructions in the memory 503 to execute a training method for a content quality scoring model, the method comprising: obtaining multiple text samples, and multiple attribute features and content quality scores corresponding to each text sample; for each text sample, cross-processing the multiple attribute features corresponding to the text sample to obtain multiple combined attribute features corresponding to the text sample; inputting the multiple attribute features and the multiple combined attribute features corresponding to each of the multiple text samples into multiple initial classification and regression trees of the initial content quality scoring model to obtain multiple predicted content quality scores corresponding to each text sample; and updating the model parameters of the multiple initial classification and regression trees according to the content quality scores corresponding to each text sample and the multiple predicted content quality scores to obtain a trained content quality scoring model.

[0110] Furthermore, the logic instructions in the aforementioned memory 503 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the portion that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as a USB flash drive, a mobile hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.

[0111] On the other hand, the present invention also provides a computer program product, which includes a computer program, which can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute a content quality scoring model training method provided by the above methods, the method including: obtaining multiple text samples, and multiple attribute features and content quality scores corresponding to each text sample; for each text sample, cross-processing the multiple attribute features corresponding to the text sample to obtain multiple combined attribute features corresponding to the text sample; inputting the multiple attribute features and multiple combined attribute features corresponding to each of the multiple text samples into multiple initial classification and regression trees of the initial content quality scoring model, respectively, to obtain multiple predicted content quality scores corresponding to each text sample; updating the model parameters of the multiple initial classification and regression trees according to the content quality score corresponding to each text sample and the multiple predicted content quality scores to obtain a trained content quality scoring model.

[0112] On the other hand, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements a training method for a content quality scoring model provided by the above-mentioned methods, the method comprising: obtaining multiple text samples, and multiple attribute features and content quality scores corresponding to each text sample; for each text sample, cross-processing the multiple attribute features corresponding to the text sample to obtain multiple combined attribute features corresponding to the text sample; inputting the multiple attribute features and multiple combined attribute features corresponding to each of the multiple text samples into multiple initial classification and regression trees of the initial content quality scoring model respectively to obtain multiple predicted content quality scores corresponding to each text sample; updating the model parameters of the multiple initial classification and regression trees according to the content quality score corresponding to each text sample and the multiple predicted content quality scores to obtain a trained content quality scoring model.

[0113] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, i.e., they may be located in one location or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of the present embodiment. Persons of ordinary skill in the art will be able to understand and implement the present invention without inventive effort.

[0114] Through the above description of the embodiments, those skilled in the art will clearly understand that each embodiment can be implemented using software plus a necessary general-purpose hardware platform, or of course, hardware. Based on this understanding, the essence of the above technical solution, or the portion that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a magnetic disk, or an optical disk, and includes a number of instructions for causing a computer device (such as a personal computer, server, or network device) to execute the methods described in each embodiment or certain portions of the embodiments.

[0115] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.

Claims

1. A training method for a content quality scoring model, characterized in that: include: Obtain multiple text samples, as well as multiple attribute features and content quality scores corresponding to each text sample; For each of the text samples, cross-processing is performed on a plurality of attribute features corresponding to the text sample to obtain a plurality of combined attribute features corresponding to the text sample; Inputting the multiple attribute features and the multiple combined attribute features corresponding to each of the multiple text samples into multiple initial classification and regression trees of the initial content quality scoring model to obtain multiple predicted content quality scores corresponding to each of the text samples; updating the model parameters of the multiple initial classification and regression trees according to the content quality scores corresponding to the text samples and the multiple predicted content quality scores to obtain a trained content quality scoring model; The updating of the model parameters of the multiple initial classification and regression trees according to the content quality scores corresponding to the text samples and the multiple predicted content quality scores includes: For each of the text samples, constructing a mean square error loss function between the content quality score corresponding to the text sample and each predicted content quality score in a plurality of predicted content quality scores, to obtain a plurality of mean square error loss functions corresponding to the text sample; The model parameters of the multiple initial classification and regression trees are updated according to the multiple mean square error loss functions corresponding to the text samples.

2. The method for training a content quality scoring model according to claim 1, wherein: The updating of the model parameters of the plurality of initial classification and regression trees according to the plurality of mean square error loss functions corresponding to the respective text samples comprises: For each of the text samples, determining, in a plurality of mean square error loss functions corresponding to the text sample, the weight of an initial classification and regression tree corresponding to the predicted content quality score used when constructing each mean square error loss function, and obtaining the weight of the classification and regression tree corresponding to each mean square error loss function; and determining a first average loss function of the plurality of mean square error loss functions based on the plurality of mean square error loss functions and the weights of the classification and regression trees corresponding to each mean square error loss function; The model parameters of the multiple initial classification and regression trees are updated according to the first average loss function corresponding to each text sample.

3. The method for training a content quality scoring model according to claim 2, wherein: The updating of the model parameters of the plurality of initial classification and regression trees according to the first average loss function corresponding to each text sample includes: Determining a second average loss function corresponding to the plurality of text samples according to the first average loss function corresponding to each text sample; According to the second average loss function, the model parameters of the multiple initial classification and regression trees are updated.

4. A content quality scoring method, characterized in that: include: Obtaining a text to be scored and a plurality of attribute features corresponding to the text to be scored; Cross-processing the multiple attribute features to obtain multiple combined attribute features corresponding to the text to be scored; Inputting the multiple attribute features and the multiple combined attribute features into multiple classification and regression trees of a content quality scoring model respectively to obtain multiple content quality scores corresponding to the text to be scored; wherein the content quality scoring model is the content quality scoring model according to any one of claims 1 to 3 above; A target content quality score corresponding to the to-be-scored text is determined according to the multiple content quality scores.

5. The content quality scoring method according to claim 4, characterized in that: Determining a target content quality score corresponding to the to-be-scored text according to the multiple content quality scores includes: The target content quality score corresponding to the text to be scored is determined according to the multiple content quality scores and the weights of the classification and regression trees corresponding to the content quality scores.

6. A training device for a content quality scoring model, characterized in that: include: A first acquisition unit is used to acquire multiple text samples, and multiple attribute features and content quality scores corresponding to each text sample; A first processing unit is configured to perform cross-processing on a plurality of attribute features corresponding to each text sample to obtain a plurality of combined attribute features corresponding to the text sample; a second processing unit, configured to input the plurality of attribute features and the plurality of combined attribute features corresponding to each of the plurality of text samples into a plurality of initial classification and regression trees of an initial content quality scoring model, to obtain a plurality of predicted content quality scores corresponding to each of the text samples; an updating unit, configured to update the model parameters of the plurality of initial classification and regression trees according to the content quality scores corresponding to the respective text samples and the plurality of predicted content quality scores, so as to obtain a trained content quality scoring model; Among them, the updating unit is specifically used to construct, for each text sample, a mean square error loss function between the content quality score corresponding to the text sample and each predicted content quality score in multiple predicted content quality scores, to obtain multiple mean square error loss functions corresponding to the text sample; and update the model parameters of multiple initial classification and regression trees according to the multiple mean square error loss functions corresponding to each text sample.

7. A content quality scoring device, characterized in that: include: A second acquisition unit is used to acquire the text to be scored and a plurality of attribute features corresponding to the text to be scored; A third processing unit is configured to perform cross-processing on the multiple attribute features to obtain multiple combined attribute features corresponding to the text to be scored; The fourth processing unit is used to input the multiple attribute features and the multiple combined attribute features into multiple classification and regression trees of the content quality scoring model respectively to obtain multiple content quality scores corresponding to the text to be scored; wherein, the content quality scoring model is the content quality scoring model described in any one of claims 1 to 3 above.

8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, it implements the training method of the content quality scoring model as described in any one of claims 1 to 3; or, it implements the content quality scoring method as described in any one of claims 4 to 5.

9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, it implements the training method of the content quality scoring model as described in any one of claims 1 to 3; or, it implements the content quality scoring method as described in any one of claims 4 to 5.

10. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, it implements the training method of the content quality scoring model as described in any one of claims 1 to 3; or, it implements the content quality scoring method as described in any one of claims 4 to 5.

Citation Information

Patent Citations

  • Applying a trained model for predicting quality of a content item along a graduated scale

    US20190095961A1

  • Classifying text to determine a goal type used to select machine learning algorithm outcomes

    US20190318269A1