Language model evaluation method and device and electronic equipment

By obtaining the extended word segmentation set and output probabilities of the language model, the problem of low evaluation accuracy of the language model is solved, and a more accurate evaluation effect is achieved.

CN120996036APending Publication Date: 2025-11-21北京中关村科金技术有限公司
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511109224.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-08
Publication Date
2025-11-21

AI Technical Summary

Technical Problem

Existing language model evaluation methods have low accuracy in open-ended dialogue scenarios, especially in situations where there are no reference answers or where strategic dialogue rounds are required, making it difficult to accurately assess the quality of the content generated by the language model.

Method used

By obtaining the first keyword and its extended word segmentation set, and determining the perplexity of the language model based on the output probability of the segmentation in the extended word segmentation set, the evaluation accuracy is improved.

Benefits of technology

It improves the accuracy of language model evaluation, solves the problems of keyword segmentation misalignment and probability underestimation, and enhances the reliability of model evaluation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120996036A_ABST
    Figure CN120996036A_ABST
Patent Text Reader

Abstract

The invention provides a language model evaluation method and device and electronic device.The method comprises the steps that a first keyword and a word list are obtained, the first keyword is a keyword output by a language model expected by first input information, and the word list comprises a plurality of segmented words obtained by segmenting corpora in a corpus; searching segmented words corresponding to the first keyword in a word list to obtain an extended segmented word set of the first keyword; obtaining an output probability of each segmented word in the word list, wherein the output probability of each segmented word is a segmented word output probability determined by the language model for the first input information; and determining the confusion degree for evaluating the language model based on the output probability of the segmented words in the extended segmented word set. And the model evaluation accuracy is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of natural language processing technology, and in particular to a language model evaluation method, apparatus and electronic device. Background Technology

[0002] With the widespread application of large language models in open-ended dialogue tasks, evaluating the quality and effectiveness of their generated content has become a key research and application issue. Among related technologies, language model evaluation methods mostly rely on: text matching metrics (BLEU, ROUGE, BERTScore), reference-based scoring (reference-based in NLG tasks), and LLM-as-a-Judge scoring, etc.

[0003] However, in dialogue scenarios (e.g., question-and-answer / consultation scenarios), real-time responses to user input are required. Strict reproduction of the reference answer is not necessary, or there may be no reference answer at all. Instead, keywords should be strategically used in appropriate dialogue rounds to advance the interaction objective. Therefore, the language model evaluation methods described above can easily lead to low accuracy in language model evaluation. Summary of the Invention

[0004] This application provides a language model evaluation method, apparatus, and electronic device to address the problem of low accuracy in existing language model evaluations.

[0005] To solve the above-mentioned technical problems, this application is implemented as follows:

[0006] In a first aspect, embodiments of this application provide a language model evaluation method, the method comprising:

[0007] Obtain a first keyword and a vocabulary list. The first keyword is the keyword that the language model is expected to output for the first input information. The vocabulary list includes multiple word segments obtained by segmenting the corpus.

[0008] Search the word segment corresponding to the first keyword in the word list to obtain the extended word segment set of the first keyword;

[0009] Obtain the output probability of each word segment in the vocabulary, wherein the output probability of the word segment is the probability of the language model outputting the word segment based on the first input information;

[0010] Based on the output probabilities of word segmentation in the extended word segmentation set, the perplexity used to evaluate the language model is determined.

[0011] Secondly, embodiments of this application provide a language model evaluation apparatus, the apparatus comprising:

[0012] The first acquisition module is used to acquire a first keyword and a vocabulary. The first keyword is the keyword that the language model is expected to output in response to the first input information. The vocabulary includes multiple word segments obtained by segmenting the corpus.

[0013] The search module is used to search for the word segment corresponding to the first keyword in the word list to obtain an extended word segment set for the first keyword;

[0014] The second acquisition module is used to acquire the output probability of each word segment in the vocabulary, wherein the output probability of the word segment is the probability of the language model to output the word segment based on the first input information, and the plurality of word segments include the first keyword;

[0015] A determination module is used to determine the perplexity of the language model based on the output probabilities of word segmentation in the extended word segmentation set.

[0016] Thirdly, embodiments of this application provide an electronic device, including a transceiver and a processor.

[0017] The processor is used for:

[0018] Obtain a first keyword and a vocabulary list. The first keyword is the keyword that the language model expects to output based on the first input information. The vocabulary list includes multiple word segments obtained by segmenting the corpus.

[0019] Search the word segment corresponding to the first keyword in the word list to obtain the extended word segment set of the first keyword;

[0020] The output probability of each word segment in the vocabulary is obtained. The output probability of each word segment is the probability of the language model outputting the word segment based on the first input information. The multiple word segments include the first keyword.

[0021] Based on the output probabilities of word segmentation in the extended word segmentation set, the perplexity used to evaluate the language model is determined.

[0022] Fourthly, embodiments of this application provide an electronic device, including: a processor, a memory, and a program stored in the memory and executable on the processor, wherein when the program is executed by the processor, it implements the steps of the language model evaluation method described in the first aspect.

[0023] Fifthly, embodiments of this application provide a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps of the language model evaluation method described in the first aspect.

[0024] In a sixth aspect, embodiments of this application provide a computer program product, including computer instructions that, when executed by a processor, implement the steps of the method described in the first aspect above.

[0025] In this embodiment, a first keyword expected to be output by the language model for the first input information can be obtained first. The first keyword can be expanded according to the vocabulary, and an expanded word segmentation set for the first keyword can be determined in the vocabulary. That is, the first keyword has been segmented and expanded to increase the coverage of the keyword. The probability of the language model outputting each word in the vocabulary for the first input information can be obtained. In this way, the output probability of the word segmentation based on the expanded word segmentation set can be used to determine the perplexity for evaluating the language model, thereby improving the accuracy of the perplexity and thus improving the accuracy of the language model evaluation. Attached Figure Description

[0026] To more clearly illustrate the technical solutions of the embodiments of this application, the drawings used in the description of the embodiments of this application will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0027] Figure 1 This is one of the flowcharts of a language model evaluation method provided in the embodiments of this application;

[0028] Figure 2 This is the second flowchart of a language model evaluation method provided in the embodiments of this application;

[0029] Figure 3 This is a schematic diagram of a language model evaluation device provided in an embodiment of this application;

[0030] Figure 4 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation

[0031] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0032] See Figure 1 , Figure 1 This is a flowchart of a language model evaluation method provided in an embodiment of this application, executed by a first electronic device, which can be a terminal or a network device. Figure 1 As shown, the language model evaluation method provided in this embodiment includes the following steps:

[0033] Step 101: Obtain the first keyword and the vocabulary. The first keyword is the keyword that the language model expects to output based on the first input information. The vocabulary includes multiple word segments obtained by segmenting the corpus.

[0034] The first input information can be understood as information entered by the user, which can be, but is not limited to, inquiries, questions, or queries. For example, in a dialogue scenario, the first input information could be a question entered by the user. Different first input information corresponds to different user expectations, leading to different expected responses from the language model. Therefore, the expected keywords output by the language model may also differ. For example, for user input information X, the expected keywords might include "quote," "sign up," and "discount," while for user input information Y, the expected keywords might include "validity period" and "package." As an example, the language model can include, but is not limited to, ChatGPT, Claude, and BERT models. The corpus can be the corpus used in the application scenario of this method, including a large amount of data. Each corpus entry can be pre-segmented (word-segmented) to obtain multiple word segments, forming a vocabulary.

[0035] Step 102: Search for the word segment corresponding to the first keyword in the word list to obtain the extended word segment set of the first keyword.

[0036] The word segmentation in the vocabulary may include the first keyword (i.e., the first keyword can be one or at least two of the multiple word segmentations in the vocabulary, and the multiple word segmentations include the first keyword), or it may not include the first keyword. The word segmentation in the vocabulary may also include other word segmentations, and among these other word segmentations, there may also be word segmentations that are semantically similar to the first keyword (extended word segmentations). Therefore, the word segmentation corresponding to the first keyword can be searched in the vocabulary to obtain the extended word segmentation set of the first keyword. It should be understood that semantic similarity can be understood as a small degree of semantic similarity, for example, less than or equal to a preset similarity threshold (greater than or equal to 0). This preset similarity threshold can be set in advance according to actual needs or historical experience.

[0037] Step 103: Obtain the output probability of each word segmentation in the vocabulary. The output probability of each word segmentation is the probability of the output word segmentation determined by the language model for the first input information.

[0038] It should be understood that during the process of outputting or generating text, the language model determines whether to generate a segment based on the predicted probability of its occurrence given the preceding words (i.e., conditional probability). In this application, the output probability can be a conditional probability. The language model can determine the output probability of each segment in the vocabulary for the first input information (the sum of the output probabilities of each segment in the vocabulary is 1) to output the corresponding content. For example, the segment with the highest output probability can be selected as the output segment for this round.

[0039] Step 104: Determine the perplexity of the language model based on the output probabilities of word segmentation in the extended word segmentation set.

[0040] Perplexity (PPL) represents the degree of uncertainty a language model faces when predicting text (e.g., the first keyword). It measures the language model's ability to predict text and assesses its policy mastery. A higher output probability for the first keyword corresponds to a lower perplexity, indicating a stronger ability to generate the first keyword and a better prediction performance, and vice versa. However, when using perplexity, misalignment can occur between the keyword and the tokenizer's segmentation results (i.e., the word segments generated from the corpus). For example, the keyword may not be independently segmented into individual tokens by the tokenizer but may be embedded in long tokens or combinations of multiple tokens, leading to keyword localization failure or perplexity distortion. This type of problem can be called tokenizer-induced keyword segmentation misalignment. For example, given the keyword "house," if the model outputs "Have you received your house yet?", the tokenizer will segment it into "you," "house," "receive," "house," and "have you received it yet?". If we search for "house" using its index (id) in the tokenizer's vocabulary, it won't match because the ID of "house" is different from "house." If we only search for the probability of "house," the probability is low and doesn't reflect the model's bias, making it impossible to match in the vocabulary. This results in low accuracy of the model's predicted output probability for this keyword, leading to poor evaluation accuracy. The language model generates a high level of perplexity and low accuracy for this keyword, easily resulting in low evaluation accuracy for the language model. However, in this embodiment, in determining the perplexity used to evaluate the language model, the given first keyword is expanded. An expanded word segmentation set for the first keyword is determined in the vocabulary. Even if the first keyword is not favored by the model, the output probability of the words in the expanded word segmentation set corresponding to the first keyword can be combined to determine the perplexity, improving the accuracy of the perplexity and thus improving the accuracy of the language model evaluation.

[0041] In this embodiment, a first keyword expected to be output by the language model for the first input information can be obtained first. The first keyword can be expanded according to the vocabulary, and an expanded word segmentation set for the first keyword can be determined in the vocabulary. That is, the first keyword has been segmented and expanded to increase the coverage of the keyword. The probability of the language model outputting each word in the vocabulary for the first input information can be obtained. In this way, the output probability of the word segmentation based on the expanded word segmentation set can be used to determine the perplexity for evaluating the language model, thereby improving the accuracy of the perplexity and thus improving the accuracy of the language model evaluation.

[0042] In some embodiments, the first keyword consists of a prefix and a suffix, and the step of searching for the word segment corresponding to the first keyword in the word list to obtain an extended word segment set for the first keyword includes at least one of the following:

[0043] The vocabulary containing the first keyword is identified as the first extended word segmentation set;

[0044] The word segment ending with the prefix word in the word list is determined as the second extended word set, and the word segment starting with the suffix word in the word list is determined as the third extended word set. The word segment obtained by concatenating any word in the second extended word set with any word in the third extended word set contains the first keyword but is not the first keyword.

[0045] Wherein, the extended word segmentation set of the first keyword includes at least one of the following:

[0046] The first extended word segmentation set;

[0047] The second extended vocabulary set and the third extended vocabulary set.

[0048] It should be understood that the first extended word set may include the first keyword itself as well as word segments that include the first keyword but are not the first keyword. The word segments in the second extended word set end with the prefix of the first keyword, and the word segments in the third extended word set begin with the suffix of the first keyword. They are concatenated in the order of word segments in the second extended word set first, followed by word segments in the third extended word set. The resulting concatenated word segment contains the first keyword but is not the first keyword, and the length of the concatenated word segment is greater than that of the first keyword.

[0049] By means of the embodiments of this application, word segments that are semantically similar to the first keyword or word segments that are semantically similar to the first keyword after being concatenated can be found in the word list (i.e., a word segment pair includes a word segment in the second extended word set and a word segment in the third extended word set), so as to expand the keyword, improve the keyword coverage, improve the accuracy of perplexity, and thus improve the accuracy of model evaluation.

[0050] In some embodiments, determining the perplexity of the language model based on the output probabilities of word segmentation in the extended word segmentation set includes:

[0051] The perplexity is determined based on the maximum value in the first probability, and the perplexity is inversely correlated with the maximum value; wherein the first probability includes at least one of the following:

[0052] Output probabilities of word segmentation in the first extended word segmentation set;

[0053] There are K second probabilities, where the number of word segments in the second extended word set is M, and the number of word segments in the third extended word set is N, where K = M * N. Each second probability is the average of the output probabilities of a word segment in the second extended word set and the output probabilities of a word segment in the third extended word set. K, M, and N are all positive integers. * indicates multiplication.

[0054] It is understood that in determining the perplexity, for the second and third extended word sets, the average output probability of the two words in the word segmentation pair is used. Each word in the second extended word set and each word in the third extended word set can be concatenated to contain the first keyword, thus forming M*N second probabilities. The first probability includes K second probabilities and / or the output probabilities of words in the first extended word set. The perplexity is determined using the maximum probability among the first probabilities to improve the accuracy of the perplexity, thereby improving the accuracy of model evaluation.

[0055] In some embodiments, determining the perplexity based on the maximum value among the first probabilities includes:

[0056] When the number of the first keyword is one, the reciprocal of the maximum value is determined as the perplexity.

[0057] When the number of first keywords is at least two, the inverse of the mean of the third probabilities of the at least two first keywords is determined as the perplexity, and the third probability of the first keyword is the maximum value among the first probabilities corresponding to the first keyword.

[0058] It should be understood that the perplexity is inversely correlated with the maximum value and can be expressed in various forms, such as the negative of the logarithm of the maximum value or the reciprocal of the maximum value, without specific limitations. Furthermore, it should be noted that for at least two first keywords, each first keyword has a corresponding extended word segmentation set. Each extended word segmentation set of a first keyword includes at least one of the following: a first extended word segmentation set corresponding to the first keyword; a second extended word segmentation set corresponding to the first keyword; and a third extended word segmentation set. Thus, in determining the perplexity of the language model based on the output probabilities of word segmentation in the extended word segmentation sets corresponding to at least two first keywords, the first probability corresponding to each first keyword may include at least one of the following: the output probability of word segmentation in the first extended word segmentation set corresponding to the first keyword; and K second probabilities corresponding to the first keyword. Each second probability corresponding to the first keyword is the average of the output probability of a word segmentation in the second extended word segmentation set corresponding to the first keyword and the output probability of a word segmentation in the third extended word segmentation set corresponding to the first keyword.

[0059] In this embodiment, the method of determining the perplexity level differs depending on the number of first keywords, thus improving the flexibility of perplexity level determination.

[0060] In some embodiments, the vocabulary is obtained in the following ways:

[0061] Call the tokenizer interface of the language model to obtain the index of the words in the vocabulary;

[0062] The word index in the vocabulary is decoded to obtain the vocabulary.

[0063] It is understandable that each word in the vocabulary has its corresponding index, and also a corresponding word vector; the three are in a one-to-one correspondence. In this application, the word segmentation in the vocabulary is obtained by the tokenizer after segmenting the corpus. Therefore, the tokenizer interface can be called to segment the corresponding index, and the vocabulary can be obtained by decoding the index of the word segmentation in the vocabulary.

[0064] like Figure 2 As shown, the above method will be described in detail below with a specific embodiment.

[0065] First, keyword segmentation misalignment can be divided into the following two types:

[0066] Embedded substrings: Set keyword a to be a substring of a token b after being split by the tokenizer (a in b), such as "house" in "house".

[0067] Cross-token inclusion compound: Keyword a spans tokens b and c (a in b+c), for example, "house" in "new house" + "child".

[0068] As mentioned above, misaligned keyword segmentation can lead to inaccurate calculations of keyword evaluation probabilities based on perplexity, thus affecting the reliability of the evaluation.

[0069] This application can solve the following technical problems:

[0070] 1. Keyword segmentation misalignment issue: The set keywords cannot be directly mapped to the token sequence in the model's tokenizer;

[0071] 2. The probability of the generated keyword is underestimated: Because the original token is not matched, it is impossible to accurately extract the predicted probability of the word by the model. If the keyword token is used directly for matching or PPL calculation, it will lead to misjudgment and will not be able to truly reflect the tendency or naturalness of the model when generating the keyword.

[0072] To address the above problems, this invention proposes a keyword perplexity assessment method with a thesaurus expansion mechanism, the process of which is as follows:

[0073] First, obtain the original keyword library K (corresponding to the first keyword): for example, extract a set of core business keywords (such as "house", "quote", "package") from manual or domain terminology.

[0074] Secondly, construct the tokenizer inverse mapping table (obtain the vocabulary V of the language model): call the tokenizer interface of the selected language model to obtain its complete vocabulary index (vocab id) range; decode the id of each token in the vocabulary to obtain its corresponding text representation, and all text representations constitute the vocabulary V;

[0075] Then, an expanded thesaurus is generated (constructing an expanded word segmentation set for each first keyword):

[0076] (1) Obtaining the substring embedded token set (obtaining the first extended word segmentation set): Filter the single token terms (a in b) that contain the first keyword;

[0077] For the first keyword a (e.g., "accepting a house"), iterate through the vocabulary V above to find all tokens that satisfy a in b; and include all tokens b containing the first keyword as a first-level extended term set (i.e., the first extended word segmentation set).

[0078] (2) Obtaining composite token sets across tokens (obtaining extended word segmentation set B and extended word segmentation set C, i.e., the second extended word segmentation set and the third extended word segmentation set): mining combinations formed by splicing two tokens to contain the first keyword (a inb+c);

[0079] For all possible split positions of the first keyword 'a', 'a' is split into a suffix and a prefix.

[0080] Find the set B of all tokens b that end with the prefix part, and the set C of all tokens c that begin with the suffix part;

[0081] Concatenate sets B and C (Cartesian product). If the concatenated string contains the keyword 'a', and b+c itself is not equal to 'a' (i.e., avoid duplicate matching of first-level terms), then the combination (b,c) is considered a valid token concatenation item. This combination (b,c) can also be called a word segmentation pair.

[0082] All token combinations (b, c) that meet the conditions are included in the second-level extended term set.

[0083] (3) Output the expanded vocabulary results

[0084] Finally, the first-level terms (single token hits) and second-level terms (token concatenation hits) mentioned above are returned as extended vocabulary entries for the first keyword 'a', which are used for subsequent identification and localization of matching keyword expressions in the model output content.

[0085] For matching expanded thesaurus items during evaluation:

[0086] In the text generated by the model, the output probability of the token is extracted by position polling. For the b+c type or other cases with multiple consecutive tokens, the geometric mean of the output probabilities of each token is taken as the probability of this type. The output probability of each word in the extended word segmentation set of the first keyword output by the aggregated language model is extracted. The maximum probability of the extended word segmentation in the extended word segmentation set of the first keyword a is taken as the probability of keyword a, which can be used to calculate the perplexity.

[0087] The solution described in the above embodiments of this application can improve keyword matching coverage and evaluation accuracy: effectively solve the problem of inaccurate evaluation caused by keyword substring embedding or cross-token inclusion; and has universality and compatibility: the process is applicable to any tokenizer (including BPE, WordPiece, SentencePiece, etc.); and can adapt to the token granularity differences of different models (ChatGLM, Baichuan, Qwen, LLaMA, etc.).

[0088] like Figure 3 As shown, Figure 3 This is a schematic diagram of a language model evaluation device provided in an embodiment of this application, such as... Figure 3 As shown, the language model evaluation device 300 includes:

[0089] The first acquisition module 301 is used to acquire the first keyword and the vocabulary. The first keyword is the keyword that the language model expects to output for the first input information. The vocabulary includes multiple word segments obtained by segmenting the corpus in the corpus.

[0090] The search module 302 is used to search for the word segment corresponding to the first keyword in the vocabulary and obtain the extended word segment set of the first keyword;

[0091] The second acquisition module 303 is used to acquire the output probability of each word segment in the vocabulary. The output probability of the word segment is the probability of the output word segment determined by the language model for the first input information.

[0092] The determination module 304 is used to determine the perplexity of the language model based on the output probabilities of word segmentation in the extended word segmentation set.

[0093] In some embodiments, the first keyword consists of a prefix and a suffix, and the search module 302 is used for at least one of the following:

[0094] The first determining submodule is used to determine the first extended word segmentation set as the vocabulary containing the first keyword;

[0095] The second determination submodule is used to determine the word segments in the vocabulary that end with the prefix word as the second extended word set, and to determine the word segments in the vocabulary that start with the suffix word as the third extended word set. The word segment obtained by concatenating any word segment in the second extended word set with any word segment in the third extended word set contains the first keyword but is not the first keyword.

[0096] The extended word segmentation set of the first keyword includes at least one of the following:

[0097] First extended word segmentation set;

[0098] The second and third extended vocabulary sets.

[0099] In some embodiments, the determining module 304 includes:

[0100] The third determining submodule is used to determine the perplexity based on the maximum value in the first probability, wherein the perplexity is inversely correlated with the maximum value; wherein the first probability includes at least one of the following:

[0101] Output probabilities of word segmentation in the first extended word segmentation set;

[0102] There are K second probabilities, where the number of word segments in the second extended word set is M, and the number of word segments in the third extended word set is N, and K = M * N. Each second probability is the average of the output probability of a word segment in the second extended word set and the output probability of a word segment in the third extended word set. K, M, and N are all positive integers.

[0103] In some embodiments, the third determining submodule is specifically used for:

[0104] When the number of the first keyword is one, the reciprocal of the maximum value is determined as the perplexity.

[0105] When the number of first keywords is at least two, the inverse of the mean of the third probabilities of the at least two first keywords is determined as the perplexity, and the third probability of the first keyword is the maximum value among the first probabilities corresponding to the first keyword.

[0106] In some embodiments, the first acquisition module 301 includes:

[0107] The index acquisition module is used to call the tokenizer interface of the language model to obtain the index of the words segmented in the vocabulary;

[0108] A decoding module is used to decode the word segmentation indexes in the vocabulary to obtain the vocabulary.

[0109] The device 300 provided in this embodiment can realize each process of each embodiment of the above-described language model evaluation method, with one-to-one correspondence of technical features and the same technical effect. To avoid repetition, it will not be described again here.

[0110] This application also provides an electronic device, including: a processor, a memory, and a program stored in the memory and executable on the processor. When the program is executed by the processor, it implements the various processes of the above-described language model evaluation method embodiments and achieves the same technical effect. To avoid repetition, it will not be described again here.

[0111] For details, see Figure 4 This application also provides an electronic device, including a bus 401, a transceiver 402, an antenna 403, a bus interface 404, a processor 405, and a memory 406.

[0112] The processor 405 is used for:

[0113] Obtain the first keyword and the vocabulary. The first keyword is the keyword that the language model expects to output based on the first input information. The vocabulary includes multiple word segments obtained by segmenting the corpus.

[0114] Find the word segment corresponding to the first keyword in the vocabulary to obtain the extended word segment set of the first keyword;

[0115] Obtain the output probability of each word segmentation in the vocabulary. The output probability of each word segmentation is the probability of the output word segmentation determined by the language model for the first input information.

[0116] Based on the output probabilities of word segmentation in the extended word segmentation set, the perplexity used to evaluate the language model is determined.

[0117] In some embodiments, the first keyword consists of a prefix and a suffix, and the processor 405 is used for at least one of the following:

[0118] The vocabulary containing the first keyword is identified as the first extended word segmentation set;

[0119] The word segments ending with prefix words in the vocabulary are defined as the second extended word set, and the word segments starting with suffix words in the vocabulary are defined as the third extended word set. The word segment obtained by concatenating any word segment in the second extended word set with any word segment in the third extended word set contains the first keyword but is not the first keyword.

[0120] The extended word segmentation set of the first keyword includes at least one of the following:

[0121] First extended word segmentation set;

[0122] The second and third extended vocabulary sets.

[0123] In some embodiments, the processor 405 is configured to:

[0124] The perplexity is determined based on the maximum value in the first probability, and the perplexity is inversely correlated with the maximum value; wherein the first probability includes at least one of the following:

[0125] Output probabilities of word segmentation in the first extended word segmentation set;

[0126] There are K second probabilities, where the number of word segments in the second extended word set is M, and the number of word segments in the third extended word set is N, and K = M * N. Each second probability is the average of the output probability of a word segment in the second extended word set and the output probability of a word segment in the third extended word set. K, M, and N are all positive integers.

[0127] In some embodiments, the processor 405 is configured to:

[0128] When the number of the first keyword is one, the reciprocal of the maximum value is determined as the perplexity.

[0129] When the number of first keywords is at least two, the inverse of the mean of the third probabilities of the at least two first keywords is determined as the perplexity, and the third probability of the first keyword is the maximum value among the first probabilities corresponding to the first keyword.

[0130] In some embodiments, the processor 405 is configured to:

[0131] The word segmentation index in the vocabulary is decoded to obtain the vocabulary decoding module, which is used to decode the index of the vocabulary to obtain the vocabulary.

[0132] exist Figure 4 In this document, a bus architecture (represented by bus 401) is used. Bus 401 may include any number of interconnected buses and bridges, linking various circuits including one or more processors represented by processor 405 and memory represented by memory 406. Bus 401 may also link various other circuits such as peripheral devices, voltage regulators, and power management circuits, which are well known in the art and therefore will not be described further herein. Bus interface 404 provides an interface between bus 401 and transceiver 402. Transceiver 402 may be a single element or multiple elements, such as multiple receivers and transmitters, providing a unit for communicating with various other devices over a transmission medium. Data processed by processor 405 is transmitted over a wireless medium via antenna 403, which further receives data and transmits data to processor 405.

[0133] Processor 405 is responsible for managing bus 401 and general processing, and can also provide various functions, including timing, peripheral interface, voltage regulation, power management, and other control functions. Memory 406 can be used to store data used by processor 405 during operation.

[0134] Optionally, the processor 405 can be a CPU, ASIC, FPGA, or CPLD.

[0135] This application also provides a computer-readable storage medium storing a computer program. When executed by a processor, the computer program implements the various processes of the above-described language model evaluation method embodiments and achieves the same technical effects. To avoid repetition, it will not be described again here. The computer-readable storage medium may be a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.

[0136] This application provides a computer program product, including computer instructions. When the computer instructions are executed by a processor, they implement the various processes of the method described in the embodiment. The technical features are one-to-one and can achieve the same technical effect. To avoid repetition, they will not be described again here.

[0137] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.

[0138] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal (which may be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods of the various embodiments of this application.

[0139] The embodiments of this application have been described above with reference to the accompanying drawings. However, this application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of this application without departing from the spirit and scope of the claims, and all of these forms are within the protection scope of this application.

Claims

1. A language model evaluation method, characterized in that, The method includes: Obtain a first keyword and a vocabulary list. The first keyword is the keyword that the language model is expected to output for the first input information. The vocabulary list includes multiple word segments obtained by segmenting the corpus. Search the word segment corresponding to the first keyword in the word list to obtain the extended word segment set of the first keyword; Obtain the output probability of each word segment in the vocabulary, wherein the output probability of the word segment is the probability of the language model outputting the word segment based on the first input information; Based on the output probabilities of word segmentation in the extended word segmentation set, the perplexity used to evaluate the language model is determined.

2. The method according to claim 1, characterized in that, The first keyword consists of prefixes and suffixes. The step of searching the word segment corresponding to the first keyword in the word list to obtain an extended word segment set for the first keyword includes at least one of the following: The vocabulary containing the first keyword is identified as the first extended word segmentation set; The word segment ending with the prefix word in the word list is determined as the second extended word set, and the word segment starting with the suffix word in the word list is determined as the third extended word set. The word segment obtained by concatenating any word in the second extended word set with any word in the third extended word set contains the first keyword but is not the first keyword. Wherein, the extended word segmentation set of the first keyword includes at least one of the following: The first extended word segmentation set; The second extended vocabulary set and the third extended vocabulary set.

3. The method according to claim 2, characterized in that, The determination of the perplexity of the language model based on the output probabilities of word segmentation in the extended word segmentation set includes: The perplexity is determined based on the maximum value in the first probability, and the perplexity is inversely correlated with the maximum value; wherein the first probability includes at least one of the following: Output probabilities of word segmentation in the first extended word segmentation set; There are K second probabilities, where the number of word segments in the second extended word set is M, and the number of word segments in the third extended word set is N, and K = M * N. Each second probability is the average of the output probability of a word segment in the second extended word set and the output probability of a word segment in the third extended word set. K, M, and N are all positive integers.

4. The method according to claim 3, characterized in that, Determining the perplexity based on the maximum value in the first probability includes: When the number of the first keyword is one, the reciprocal of the maximum value is determined as the perplexity. When the number of first keywords is at least two, the inverse of the mean of the third probabilities of the at least two first keywords is determined as the perplexity, and the third probability of the first keyword is the maximum value among the first probabilities corresponding to the first keyword.

5. The method according to any one of claims 1-4, characterized in that, The vocabulary list was obtained through the following methods: Call the tokenizer interface of the language model to obtain the index of the words in the vocabulary; The word index in the vocabulary is decoded to obtain the vocabulary.

6. A language model evaluation device, characterized in that, The device includes: The first acquisition module is used to acquire a first keyword and a vocabulary. The first keyword is the keyword that the language model is expected to output in response to the first input information. The vocabulary includes multiple word segments obtained by segmenting the corpus. The search module is used to search for the word segment corresponding to the first keyword in the word list to obtain an extended word segment set for the first keyword; The second acquisition module is used to acquire the output probability of each word segment in the vocabulary, wherein the output probability of the word segment is the probability of the language model to output the word segment based on the first input information, and the plurality of word segments include the first keyword; A determination module is used to determine the perplexity of the language model based on the output probabilities of word segmentation in the extended word segmentation set.

7. An electronic device, characterized in that, Including transceivers and processors, The processor is used for: Obtain a first keyword and a vocabulary list. The first keyword is the keyword that the language model expects to output based on the first input information. The vocabulary list includes multiple word segments obtained by segmenting the corpus. Search the word segment corresponding to the first keyword in the word list to obtain the extended word segment set of the first keyword; The output probability of each word segment in the vocabulary is obtained. The output probability of each word segment is the probability of the language model outputting the word segment based on the first input information. The multiple word segments include the first keyword. Based on the output probabilities of word segmentation in the extended word segmentation set, the perplexity used to evaluate the language model is determined.

8. An electronic device, characterized in that, include: A processor, a memory, and a program stored in the memory and executable on the processor, wherein the program, when executed by the processor, implements the steps of the method as described in any one of claims 1 to 5.

9. A computer-readable storage medium having a computer program stored thereon, the computer program, when executed by a processor, implementing the steps of the method of any one of claims 1-5.

10. A computer program product, characterized in that, Includes computer instructions, which, when executed by a processor, implement the steps of the method as described in any one of claims 1-5.