Method for evaluating article by using language model
By splicing articles U and articles R into two orders, using the language model to generate abstracts and calculate the similarity, the total winning numbers and relative rankings of the articles are obtained, and the appropriate language model is selected for evaluation is solved, which solves the problem of insufficient accuracy and emotional understanding of article evaluation in the existing technology, and objective, accurate and consistent evaluation of the articles is achieved.
Patent Information
- Application Number
- CN202510327288.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-19
- Publication Date
- 2025-06-17
AI Technical Summary
The prior art lacks accuracy and emotional understanding in article evaluation, language models may generate inaccurate information, and lack human intuition and judgment, making it difficult to understand the emotional depth and subtleties in articles.
By splicing article U and article R into articles in two orders, the language model is used to generate an abstract and calculate the similarity, the total number of winning games and relative ranking of the article is obtained, and the appropriate language model is selected for evaluation.
It realizes objective, accurate and consistent evaluation of the article, reduces artificial bias, improves the reliability and standardization of the evaluation, and promotes the research and application of language models and natural language processing.
Smart Images

Figure CN120162448A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of artificial intelligence and natural language processing, and more specifically to a method for evaluating an article using a language model. Background Art
[0002] In today's digital age, with the rapid advancement of artificial intelligence technology, language models are showing an extremely rapid development trend. In recent years, with the strong support of deep learning algorithms and the nourishment of massive data, the number of parameters of language models has grown exponentially, from a relatively small scale in the early days to models with billions or even tens of billions of parameters. This scale expansion has brought about a huge improvement in performance. Language models have gained stronger text generation and semantic understanding capabilities in natural language processing, as well as longer context lengths.
[0003] In the face of the current era of article explosion, the demand for efficiency and accuracy in article evaluation is increasing, and the emergence of language models just fills the current gap. It is not limited by time and energy, and can process a large number of articles in a short period of time, saving a lot of time costs. Secondly, it is not constrained by personal emotions and can objectively give article evaluation results. Finally, it has good repeatability. For the same article input, the language model will give basically consistent evaluation results, so that articles can be evaluated at different times and backgrounds.
[0004] Although language models are now very powerful, there are still many areas that need improvement: 1. Large models may generate inaccurate or wrong information. Although they are trained on a large amount of data, they may still misunderstand the nature of the problem, confuse facts, or give illogical answers. 2. Lack of human intuition and judgment. Language models are trained based on data, and they only reason based on previous data and patterns. It may not be able to truly understand the emotional depth and subtlety contained in the article. Humans can keenly perceive the complex emotions conveyed by the author between the lines with their intuition and life experience, while language models can only match and judge based on the emotional patterns that have appeared in the data, and often ignore those unique emotional expressions that are difficult to cover with existing models.
[0005] Therefore, it is an urgent problem for those skilled in the art to propose a method for evaluating articles using a language model to solve the difficulties existing in the prior art. Summary of the invention
[0006] In view of this, the present invention provides a method for evaluating articles using a language model, which comprehensively and objectively evaluates the quality of articles and is dedicated to replacing the traditional human evaluation method of articles.
[0007] To achieve the above object, the present invention provides the following technical solutions:
[0008] A method for evaluating an article using a language model, comprising the following steps:
[0009] Generate a concatenated article from article U and article R, where article U is in the front and article R is in the back;
[0010] Use the language model to generate an abstract of the concatenated article and generate corresponding embedding vectors based on the abstract;
[0011] Calculate a first similarity and a second similarity between the embedding vector and the embedding vectors corresponding to the abstracts of article U and article R respectively;
[0012] Reverse the order of article R and article U to generate a new concatenated article, and use the language model to generate a new abstract;
[0013] Generate an embedding vector based on the new abstract, and then calculate a third similarity and a fourth similarity between it and the embedding vectors corresponding to the abstracts of article U and article R respectively;
[0014] Add the first similarity and the third similarity and compare it with the sum of the second similarity and the fourth similarity. If the similarity is greater, add one to the winning field, and add up the winning fields of the articles to obtain the total winning field number of the articles;
[0015] Rank all articles according to the total winning field number to obtain the relative ranking of the articles under the evaluation of the language model;
[0016] Select four language models for evaluation, and calculate evaluation metrics such as inversion number, cosine similarity, Manhattan distance, Pearson correlation coefficient, and Krippendorff's α according to the evaluation results to select as the language model for evaluating articles;
[0017] Use the selected language model to evaluate the articles.
[0018] Specifically, the first similarity is the similarity calculated between the article with article U in the front and article R in the back and article U, the third similarity is the similarity calculated between the article with article U in the back and article R in the front and article U, the second similarity is the similarity calculated between the article with article U in the front and article R in the back and article R, and the fourth similarity is the similarity calculated between the article with article U in the back and article R in the front and article R.
[0019] Optionally, the formula for variance is:
[0020]
[0021] where n represents a total of n articles, x1, x2, ……, x nIndicates the ranking order summarized by the language model for the scores of n articles. Indicates the average of x1, x2, ……, x n .
[0022] Optionally, the inversion number is as follows: In two sequences, for each position i, 1 ≤ i ≤ n, if A[i] > A[j], i < j and B[i] < B[j], or A[i] < A[j], i < j and B[i] > B[j], then an inversion pair is formed. The total number of inversion pairs counted is the inversion number of these two sequences.
[0023] Optionally, the calculation formula for cosine similarity is:
[0024] For two vectors and They have components [a1, a2, ……, a n and [b1, b2, ……, b n with n dimensions respectively, and the calculation formula is:
[0025]
[0026] where θ represents the angle between the two vectors.
[0027] Optionally, the calculation formula for Manhattan distance is:
[0028] In a two-dimensional plane, the calculation formula for the Manhattan distance between two points (x1, y1) and (x2, y2) is:
[0029] d = |x1 - x2| + |y1 - y2|
[0030] In a high-dimensional space, for two points P(p1, p2, ……, p n ) and Q(q1, q2, ……, q n ), the calculation formula for Manhattan distance is:
[0031]
[0032] Optionally, the calculation formula for Pearson correlation coefficient is:
[0033] For two variables X and Y, they have observed values {x1, x2, ……, x n} and {y1, y2, ……, y n} respectively, and the calculation formula is:
[0034]
[0035] where represents {x1, x2, ……, x nThe average of { represents the average of {y1, y2, ……, y n}.
[0036] Optionally, the calculation formula of Krippendorff's α is:
[0037]
[0038] where observeddisagreement represents the actual disagreement observed among coders, and expecteddisagreement represents the disagreement expected under random circumstances.
[0039] As can be seen from the above technical solutions, compared with the prior art, the present invention discloses a method for evaluating articles using a language model, and its beneficial effects are as follows:
[0040] 1) Promote research and applications in the fields of language models and natural language processing; by scoring articles using a language model, more accurate and consistent evaluation results can be provided for articles based on objective algorithms and data. Reducing the output of human and material resources, a large number of articles can be evaluated in a short time, meeting the need for quickly screening high-quality articles;
[0041] 2) And it helps to reduce human biases, improve the reliability of article evaluation. At the same time, this systematic evaluation method will promote the evaluation of articles to develop in a more standardized and normalized direction. This method will also stimulate more research and exploration on article evaluation, prompting researchers to continuously optimize the algorithms and parameters of language models to adapt to changing article types and evaluation requirements, thus forming a virtuous cycle, continuously improving the quality and efficiency of article evaluation, and providing strong support for the prosperity and development of the entire cultural and knowledge field. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on the provided drawings.
[0043] Figure 1 It is a flowchart of a method for evaluating articles using a language model provided by the present invention;
[0044] Figure 2The bar chart of the measurement indicators provided by the present invention; among them, 2a is the bar chart of the measurement indicator of the Pearson correlation coefficient, 2b is the bar chart of the measurement indicator of the variance, 2c is the bar chart of the measurement indicator of the cosine similarity, 2d is the bar chart of the measurement indicator of Krippendorff’s α, 2e is the bar chart of the measurement indicator of the inversion number, and 2f is the bar chart of the measurement indicator of the Manhattan distance. Detailed implementation manners
[0045] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0046] See Figure 1 As shown, the present invention discloses a method for evaluating an article using a language model, including the following steps:
[0047] Generate a spliced article from article U and article R, where article U is in the front and article R is in the back;
[0048] Use the language model to generate an abstract of the spliced article, and generate corresponding embedding vectors according to the abstract;
[0049] Calculate the first similarity and the second similarity between the embedding vectors and the embedding vectors corresponding to the abstracts of article U and article R respectively;
[0050] Reverse the order of article R and article U to generate a new spliced article, and use the language model to generate a new abstract;
[0051] Generate embedding vectors according to the new abstract, and then calculate the third similarity and the fourth similarity between the embedding vectors and the embedding vectors corresponding to the abstracts of article U and article R respectively;
[0052] Add the first similarity and the third similarity and compare them with the sum of the second similarity and the fourth similarity. If the similarity is greater, the winning field is incremented by one. Add up the winning fields of the articles to obtain the total winning field number of the articles;
[0053] Rank all the articles according to the total winning field number to obtain the relative ranking of the articles under the evaluation of the language model;
[0054] Select four language models for evaluation, and calculate the evaluation indicators of the inversion number, cosine similarity, Manhattan distance, Pearson correlation coefficient, and Krippendorff’s α according to the evaluation results to select as the language model for evaluating the articles;
[0055] Use the selected language model to evaluate the articles.
[0056] Specifically, when selecting a language model, consider available computing resources (such as GPUs, TPUs, etc.), whether the selected language model is easy to fine-tune, whether it has an active development community and good technical support, etc. In addition to these key factors, there are still many comprehensive factors to be considered. First, the performance and accuracy of the language model on different tasks should be evaluated, including the performance of tasks such as text generation, text classification, and named entity recognition. Second, the diversity and quality of the training data are also important evaluation criteria to ensure that the selected language model uses diverse and high-quality data, especially data related to the target application domain. In addition, the scalability, interpretability, privacy, and security of the language model are also key factors. It should be ensured that the language model is easy to expand and integrate, and the prediction results are easy to understand and comply with relevant laws, regulations, and industry standards. Finally, select a language model with long-term support and regular updates to ensure that it can continue to improve and adapt to the latest technological developments.
[0057] According to the above criteria, the present invention has selected representative language models as shown in Table 1:
[0058] Table 1 Introduction to Language Models
[0059]
[0060] The Llama3 language model demonstrates excellent performance in natural language processing tasks, especially in areas such as complex reasoning, creative writing, and code generation. The Qwen language model incorporates a variety of high-quality language data into its training data, enhances 27 languages other than Chinese and English, and is specifically optimized for common language conversion problems in multilingual scenarios, thus significantly improving the multilingual understanding and generation capabilities of the language model. The Gemma language model has been carefully optimized. In the case of a relatively small parameter scale, the Gemma language model can demonstrate performance comparable to that of models with a larger parameter scale.
[0061] Finally, based on the data in the ret list, sort out the rankings, and then calculate the variance of a single model and the inversion number, cosine similarity, Manhattan distance, Pearson correlation coefficient, and Krippendorff’s α between pairwise models, which are indicators to measure the quality of the language model in completing this task, so as to screen out the most suitable language model for this work among the selected language models.
[0062] Furthermore, the calculation formula for variance is:
[0063]
[0064] Among them, n represents a total of n articles, x1, x2, ……, x nIndicates the ranking order summarized by the language model for the scores of n articles. Indicates the average of x1, x2, ……, x n .
[0065] Specifically, variance is an indicator to measure the degree of data dispersion. The larger the variance, the higher the degree of dispersion.
[0066] Furthermore, the number of inversions is as follows: In two sequences, for each position i, 1 ≤ i ≤ n, if A[i] > A[j], i < j and B[i] < B[j], or A[i] < A[j], i < j and B[i] > B[j], then an inversion pair is formed. The total number of inversion pairs counted is the number of inversions of these two sequences.
[0067] Specifically, in a list, if the front - back position of a pair of numbers is opposite to the size order, that is, the number in front is greater than the number behind, then they are called an inversion pair. The total number of inversion pairs in a sequence is called the number of inversions of this list. The smaller the number of inversions, the lower the degree of data chaos.
[0068] Furthermore, the calculation formula for cosine similarity is:
[0069] For two vectors and They have components [a1, a2, ……, a n and [b1, b2, ……, b n in n dimensions respectively, and the calculation formula is:
[0070]
[0071] where θ represents the angle between the two vectors.
[0072] Specifically, the closer the value of cosine similarity is to 1, the higher the similarity degree of the two vectors. When the two vectors are in exactly the same direction, the cosine similarity is 1; when the two vectors are in exactly the opposite direction, the cosine similarity is - 1; when the two vectors are perpendicular to each other, the cosine similarity is 0.
[0073] Furthermore, the calculation formula for Manhattan distance is:
[0074] In a two - dimensional plane, the calculation formula for the Manhattan distance between two points (x1, y1) and (x2, y2) is:
[0075] d = |x1 - x2| + |y1 - y2|
[0076] In a high - dimensional space, for two points P(p1, p2, ……, p n ) and Q(q1, q2, ……, q n), the Manhattan distance calculation formula is:
[0077]
[0078] Specifically, the smaller the value of the Manhattan distance, the lower the degree of difference between the data.
[0079] Furthermore, the calculation formula for the Pearson correlation coefficient is:
[0080] For two variables X and Y, with observed values {x1, x2, ……, x n} and {y1, y2, ……, y n}, the calculation formula is:
[0081]
[0082] Among them, represents the average of {x1, x2, ……, x n}, represents the average of {y1, y2, ……, y n}.
[0083] Specifically, in the Pearson correlation coefficient, if the Pearson correlation coefficient is close to 1, it indicates that the two sets of data show a strong positive linear correlation. If the Pearson correlation coefficient is close to -1, it indicates that the two sets of data show a strong negative linear correlation. If the Pearson correlation coefficient is close to 0, it indicates that the degree of linear correlation between the two sets of data is very low.
[0084] Furthermore, the calculation formula for Krippendorff’s α is:
[0085]
[0086] Among them, observeddisagreement represents the actual observed disagreement among coders, and expecteddisagreement represents the expected disagreement in a random situation.
[0087] The value range of Krippendorff’s α is between -1 and 1.
[0088] When α = 1, it indicates complete agreement among coders. This means that all coders have the same coding results for each data point, with high reliability.
[0089] When α = 0, it indicates that the agreement among coders is the same as in a random situation. That is, the coding results of coders do not have better agreement than random coding, and the reliability is low.
[0090] When α < 0, it indicates that the consistency among coders is worse than random. This situation may suggest problems in the coding process and requires rechecking the coding standards.
[0091] Finally, the results of data calculation are as Figure 2 shown. Among them, 2a is the bar chart of the measurement index of Pearson correlation coefficient, 2b is the bar chart of the measurement index of variance, 2c is the bar chart of the measurement index of cosine similarity, 2d is the bar chart of the measurement index of Krippendorff’s α, 2e is the bar chart of the measurement index of inversion number, and 2f is the bar chart of the measurement index of Manhattan distance.
[0092] Figure 2 L3-150 (Llama3:8b truncated at 150 characters), L3-300 (Llama3:8b truncated at 300 characters), Q-300 (Qwen:4b truncated at 300 characters), Q-500 (Qwen:4b truncated at 500 characters), G-500 (Gemma:7b truncated at 500 characters and the temperature of the language model is set to 0), L2-500 (Llama2:7b truncated at 500 characters and the temperature of the language model is set to 0). Among them, the calculations of various indicators such as inversion number, cosine similarity, Manhattan distance, Pearson correlation coefficient, and Krippendorff’s α are obtained by calculating and adding pairwise between language models, and the variance is the variance of the language model itself.
[0093] Variance is an indicator to measure the degree of data dispersion. In article scoring, the larger the variance, the stronger the ability of the language model to distinguish between articles of different qualities. The L3-300 language model has a relatively large variance, indicating that it can more clearly distinguish between good and bad articles. In contrast, except for the result of G-500 being relatively close to it, the other language models have significant differences from it.
[0094] The inversion number reflects the degree of disorder in data sorting, and the Manhattan distance measures the degree of difference in data in space. The L3-300 language model obtains the minimum values in these two indicators, indicating that it has a high similarity with other language models in the sorting and spatial distribution of article scoring. This means that the scoring results of the L3-300 language model are more consistent with those of other language models, and its evaluation criteria are more universal and objective, without significant differences from the evaluation results of other language models, thus improving the credibility of its scoring results.
[0095] The Pearson correlation coefficient measures the degree of linear correlation between two variables. The cosine similarity is used to evaluate the similarity in direction between two vectors. Krippendorff's α is an index for measuring the inter-rater agreement. The L3-300 language model achieved the maximum value in these metrics, indicating that it performs best among other language models in terms of the correlation, similarity, and agreement in article scoring. That is, the scoring results of the L3-300 language model are highly correlated and similar to those of other language models. At the same time, it also shows that other language models have a high degree of recognition of the scoring results of the L3-300, further proving the reliability and applicability of the L3-300 language model in article scoring work.
[0096] Considering the variance and the other five metrics, the variance reflects the advantage of the L3-300 language model in distinguishing good and bad articles, while the other five metrics indicate that the L3-300 language model has a high similarity with other language models. Overall, this means that other language models have a high degree of consistency with the L3-300 language model in judging the quality of articles. Therefore, the L3-300 language model is more suitable for this work.
[0097] The various embodiments in this specification are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. For the same or similar parts among the embodiments, reference can be made to each other.
[0098] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present invention. Various modifications to these embodiments will be obvious to those skilled in the art. The general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to these embodiments shown herein, but will conform to the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A method for evaluating an article using a language model, characterized in that: It includes the following steps: Generate a spliced article from article U and article R, where article U is in the front and article R is in the back; Use a language model to generate an abstract of the spliced article and generate corresponding embedding vectors based on the abstract; Calculate the first similarity and the second similarity between the embedding vector and the embedding vectors corresponding to the abstracts of article U and article R respectively; Reverse the order of article R and article U to generate a new spliced article, and use the language model to generate a new abstract; Generate an embedding vector based on the new abstract, and then calculate the third similarity and the fourth similarity between it and the embedding vectors corresponding to the abstracts of article U and article R respectively; Add the first similarity and the third similarity and compare them with the sum of the second similarity and the fourth similarity. If the similarity of the former is greater, the winning field of the article is incremented by one. Add up the winning fields of the articles to obtain the total winning field number of the articles; Rank all articles according to the total winning field number to obtain the relative ranking of the articles under the evaluation of the language model; Select four language models for evaluation, and calculate evaluation metrics such as the number of inversions, cosine similarity, Manhattan distance, Pearson correlation coefficient, and Krippendorff’s α according to the evaluation results to select as the language models for evaluating articles; Use the selected language models to evaluate the articles.
2. A method for evaluating an article using a language model according to claim 1, characterized in that: The calculation formula for variance is: Among them, n means there are n articles in total, x1, x2, ..., x n It represents the ranking order of the language model's results for n articles. represents x1, x2, ..., x n The average of .
3. The method for evaluating an article using a language model according to claim 1, characterized in that: The number of inversions is: In two sequences, for each position i, 1 ≤ i ≤ n, if A[i] > A[j], i < j and B[i] < B[j], or A[i] < A[j], i < j and B[i] > B[j], then an inversion pair is formed. Count the total number of inversion pairs, which is the number of inversions of these two sequences.
4. The method for evaluating an article using a language model according to claim 1, characterized in that: The calculation formula for cosine similarity is: For two vectors and They have n-dimensional components [a1, a2, ..., a n ] and [b1,b 2, ……,b n ], the calculation formula is: where θ represents the angle between two vectors.
5. The method for evaluating an article using a language model according to claim 1, characterized in that: The calculation formula for Manhattan distance is: On a two-dimensional plane, the calculation formula for the Manhattan distance between two points (x1, y1) and (x2, y2) is: d = |x1 - x2| + |y1 - y2| In a high-dimensional space, for two points P(p1, p2, ..., p n ) and Q(q1,q2,……,q n ), the Manhattan distance calculation formula is:
6. The method for evaluating an article using a language model according to claim 1, characterized in that: The calculation formula for Pearson correlation coefficient is: For two variables X and Y, there are observations {x1, x2, …, x n } and {y1,y2,…,y n }, the calculation formula is: in, represents {x1,x2,……,x n }, represents {y1,y2,……,y n }Average.
7. The method for evaluating an article using a language model according to claim 1, characterized in that: The calculation formula for Krippendorff’s α is: where observeddisagreement represents the actual observed disagreement between coders, and expecteddisagreement represents the expected disagreement in the random case.