Translation effect evaluation method and device, computer equipment and storage medium

By generating and filtering target prompt words and using the target model to evaluate the evaluation translation, the problem of inaccurate translation evaluation in the prior art is solved, and a more accurate and efficient translation effect evaluation is achieved.

CN120106094APending Publication Date: 2025-06-06GUANGZHOU QUCHUANG NETWORK TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510321023.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-18
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

In the prior art, the evaluation of translation effect is inaccurate, and it relies heavily on the quality of reference translations, and cannot fully consider semantic understanding and context information.

Method used

By obtaining the translation dataset, generating training prompt words and inputting the target big model, candidate prompt words are obtained. Enter the candidate prompt words and translation samples into the target model, obtain the evaluation annotation and quality score, and filter out the target prompt words. Use target prompt words and target big model to evaluate the evaluation translation to obtain the target evaluation results.

Benefits of technology

It realizes the evaluation of translation effectiveness that does not rely on the quality of reference translations, learn evaluation rules through the semantic understanding ability of large language models, improves the evaluation accuracy and reduces labor costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120106094A_ABST
    Figure CN120106094A_ABST
Patent Text Reader

Abstract

The invention provides a translation effect evaluation method and device, computer equipment and a storage medium. The method comprises the steps of firstly obtaining a translation data set, then generating training cue words according to the translation data set, inputting the training cue words into a target large model to obtain a plurality of candidate cue words, inputting the candidate cue words and a translation sample into the target large model, obtaining a second evaluation label, and screening out a target cue word based on the second evaluation label and a corresponding first quality score. And finally, evaluating the to-be-evaluated translation by the target cue word and the target large model to obtain a target evaluation result. The method does not depend on the quality of the reference translation, but enables the reference translation to continuously learn the evaluation rule of the translation effect based on the semantic understanding ability of the large language model, so that a set of reasonable target cue words is generated, the target cue words can be utilized to automatically evaluate the translation effect in a manner close to manual evaluation, the evaluation accuracy is improved, and the evaluation efficiency is improved. And the labor cost is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of artificial intelligence technology, and in particular to a translation effect evaluation method, apparatus, computer equipment and storage medium. Background Art

[0002] In today's era of accelerating globalization, the demand for cross-language communication is growing. As a key means to break down language barriers, text translation is of great importance. Early machine translation was mainly based on rules and statistical methods. However, these methods have great limitations in dealing with complex language structures and semantic understanding. With the rise of deep learning technology, machine translation models based on neural networks, such as recurrent neural networks (RNNs), long short-term memory networks (LSTMs) and their variants, have become mainstream translation technologies. These models can automatically learn the mapping relationship between languages, which improves the accuracy and fluency of translation to a certain extent.

[0003] At present, the evaluation of machine translation effects is mainly based on reference translations. This type of method compares the machine translation results with one or more reference translations and calculates the similarity or difference between them to evaluate the quality of the translation results. However, traditional evaluation schemes rely heavily on reference translations, and the quality of the reference translations directly affects the evaluation results. There is also insufficient consideration of semantics and context, and matching is only performed at the lexical level, which cannot fully consider the semantic understanding and contextual information of the translation results. This makes it impossible to fully and accurately evaluate the quality of the translation results. Summary of the invention

[0004] The purpose of this application is to solve at least one of the above-mentioned technical deficiencies, especially the defect of inaccurate translation effect evaluation in the prior art.

[0005] In a first aspect, the present application provides a translation effect evaluation method, comprising:

[0006] Obtain a translation data set; the translation data set includes a plurality of translation samples and their corresponding first evaluation annotations, the first evaluation annotations including evaluation results and evaluation reasons;

[0007] Generate training prompt words according to the translation data set, input the training prompt words into the target large model, and obtain multiple candidate prompt words; the training prompt words are used to instruct the target large model to learn the evaluation method according to the translation data set, and generate multiple candidate prompt words for instructing the target large model to evaluate the translation effect;

[0008] Input the candidate prompt words and the translation samples into the target large model to obtain second evaluation annotations, obtain the first quality score corresponding to each second evaluation annotation, and obtain the target prompt word according to the first quality score;

[0009] The translation to be evaluated is evaluated using the target prompt words and the target large model to obtain the target evaluation result.

[0010] In one embodiment, obtaining a target prompt word according to the first quality score includes:

[0011] Screening multiple candidate prompt words according to the first quality score to obtain multiple preferred prompt words;

[0012] Input the preferred prompt words and translation samples into the target large model to obtain the third evaluation annotation;

[0013] The second quality scores corresponding to the third evaluation annotations are obtained, and the preferred prompt word with the highest second quality score is selected as the target prompt word.

[0014] In one embodiment, before the preferred prompt words and the translation samples are input into the target large model to obtain the third evaluation annotation, the method further includes:

[0015] For any selected preferred prompt word, the preferred prompt word is used as the root node;

[0016] Traversing from the root node, if the current node does not have the first number of child nodes, performing the first number of mutation operations on the current node, and using the results of the mutation operations as the child nodes of the current node;

[0017] Otherwise, select the target child node according to the exploration weight corresponding to each child node, and update the cumulative reward values ​​of all nodes along the path from the target child node to the root node according to the third quality score corresponding to the target child node, and return to the step of traversal starting from the root node to continue execution until the loop end condition is met;

[0018] The second number of nodes with the highest cumulative reward values ​​are used as newly generated preferred prompt words.

[0019] In one embodiment, the mutation operation includes at least one of sentence reorganization, synonym replacement, and adding format-qualifying phrases.

[0020] In one embodiment, the loop end condition includes: when the depth of any branch reaches a preset maximum exploration depth, or the growth rate of the cumulative reward value in the current iteration cycle is less than the convergence threshold.

[0021] In one embodiment, before evaluating the translation to be evaluated using the target prompt word and the target large model to obtain the target evaluation result, the method further includes:

[0022] Determine the domain label of the translation to be evaluated;

[0023] When the domain label does not match the set domain label, a prompt word update prompt is issued.

[0024] In one embodiment, the translation to be evaluated is evaluated using the target prompt word and the target large model to obtain a target evaluation result, including:

[0025] The target prompt words are solidified in the target macro model, and the translation to be evaluated is input into the target macro model to obtain the target evaluation result.

[0026] In a second aspect, the present application provides a translation effect evaluation device, comprising:

[0027] A data acquisition module is used to acquire a translation data set; the translation data set includes a plurality of translation samples and their corresponding first evaluation annotations, and the first evaluation annotations include evaluation results and evaluation reasons;

[0028] A candidate prompt word generation module is used to generate training prompt words according to the translation data set, input the training prompt words into the target large model, and obtain multiple candidate prompt words; the training prompt words are used to instruct the target large model to learn the evaluation method according to the translation data set, and generate multiple candidate prompt words for instructing the target large model to evaluate the translation effect;

[0029] A target prompt word generation module is used to input the candidate prompt words and the translation sample into the target large model, obtain the second evaluation annotation, obtain the first quality score corresponding to each second evaluation annotation, and obtain the target prompt word according to the first quality score;

[0030] The evaluation module is used to evaluate the translation to be evaluated using the target prompt words and the target large model to obtain a target evaluation result.

[0031] In a third aspect, the present application provides a computer device, comprising one or more processors and a memory, wherein the memory stores computer-readable instructions, and when the computer-readable instructions are executed by the one or more processors, the steps of the translation effect evaluation method in any of the above embodiments are executed.

[0032] In a fourth aspect, the present application provides a storage medium storing computer-readable instructions. When the computer-readable instructions are executed by one or more processors, the one or more processors execute the steps of the translation effect evaluation method in any of the above embodiments.

[0033] It can be seen from the above technical solutions that the embodiments of the present application have the following advantages:

[0034] Based on the translation effect evaluation method in this embodiment, first obtain a translation data set, then generate training prompt words according to the translation data set and input them into the target large model to obtain multiple candidate prompt words, input the candidate prompt words and translation samples into the target large model, obtain a second evaluation annotation, and screen out the target prompt words based on the second evaluation annotation and the corresponding first quality score. Finally, the target prompt words and the target large model evaluate the translation to be evaluated to obtain the target evaluation result. This method does not rely on the quality of the reference translation, but is based on the semantic understanding ability of the large language model to allow it to continuously learn the evaluation rules of the translation effect, thereby generating a set of reasonable target prompt words, and then the target prompt words can be used to automatically evaluate the translation effect in a manner close to manual evaluation, which not only improves the evaluation accuracy, but also reduces labor costs. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative labor.

[0036] Figure 1 A flowchart of a translation effect evaluation method provided in one embodiment of the present application;

[0037] Figure 2 A schematic diagram of a process for obtaining a target prompt word in one embodiment of the present application;

[0038] Figure 3 A schematic diagram of a process of generating a new preferred prompt word similar to an original preferred prompt word in one embodiment of the present application;

[0039] Figure 4 An internal structure diagram of a computer device provided for one embodiment of the present application. DETAILED DESCRIPTION

[0040] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.

[0041] This application provides a translation effect evaluation method, please refer to Figure 1 , including steps S102 to S108.

[0042] S102, obtaining a translation data set. The translation data set includes a plurality of translation samples and their corresponding first evaluation annotations, wherein the first evaluation annotations include evaluation results and evaluation reasons.

[0043] It can be understood that obtaining a translation dataset is to provide basic data for learning and evaluation in subsequent steps. By collecting a rich variety of translation samples and their corresponding evaluation annotations, the models involved in subsequent learning and evaluation can understand the pros and cons of different types of translations and the basis for judgment. Specifically, a translation dataset is a structured data set containing multilingual parallel corpora. Each translation sample includes the original text and the target text. For example, the translation of a sentence from English to Chinese, "The cat is on the mat" is translated as "The cat is on the mat". This pair of sentences is a translation sample. The first evaluation annotation contains the evaluation results of the translation sample and the evaluation reasons for the result. The evaluation results can be qualitative descriptions such as "excellent", "good", "average", "poor", or quantitative scores, such as values ​​between 0 and 10. The evaluation reasons elaborate on the basis for giving the evaluation results, such as considerations of grammatical accuracy, semantic completeness, and appropriate use of terminology.

[0044] The first evaluation annotation is usually done manually on the translation samples. In order to improve the annotation quality, each translation sample can be annotated by at least two different annotators. If the difference in the evaluation results exceeds the set conditions, the relevant multiple first evaluation annotations will be sent to another annotator for evaluation to obtain the final first evaluation annotation.

[0045] S104, generating training prompt words according to the translation data set, inputting the training prompt words into the target large model, and obtaining a plurality of candidate prompt words. The training prompt words are used to instruct the target large model to learn the evaluation method according to the translation data set, and to generate a plurality of candidate prompt words for instructing the target large model to evaluate the translation effect.

[0046] It can be understood that the target large model is an artificial intelligence model that has the ability to understand and process various natural language tasks after pre-training with massive data. The training prompt word is a text instruction specially designed to guide the target large model to learn specific tasks or knowledge. In this step, the training prompt word is used to instruct the target large model to learn the evaluation method based on the translation data set, and generate multiple candidate prompt words for instructing the target large model to evaluate the translation effect.

[0047] After receiving the training prompt words and the translation data set, the neural network structure inside the target large model will extract features and recognize patterns from the data. Guided by the training prompt words, the model attempts to summarize the rules and standards for evaluating translation effects from the translation data set, and then generates a series of candidate prompt words. The training prompt words can be instructions such as "analyze the relationship between translation samples and evaluation annotations in the translation data set, and generate multiple candidate prompt words for evaluating translation effects." The preprocessed translation data set and the training prompt words are input into the input layer of the target large model. The model processes the data in its internal multi-layer neural network through the forward propagation process, and finally generates multiple candidate prompt words at the output layer.

[0048] S106, inputting the candidate prompt words and the translation samples into the target large model to obtain second evaluation annotations, obtaining first quality scores corresponding to the second evaluation annotations, and obtaining target prompt words according to the first quality scores.

[0049] It can be understood that the second evaluation annotation is the annotation result generated by the target large model after receiving the candidate prompt word and the translation sample, and evaluating the translation sample according to the evaluation method indicated by the candidate prompt word. The first quality score is a quantitative indicator used to measure the quality of the second evaluation annotation, which reflects the degree of closeness between the evaluation made by the target large model based on the candidate prompt word and the actual situation.

[0050] After receiving the candidate prompt words and the translation sample, the target large model analyzes and judges the translation sample according to the evaluation rules and standards set by the candidate prompt words, thereby generating a second evaluation annotation. For example, if the candidate prompt words require the model to evaluate the translation sample from aspects such as grammar, semantics, and terminology usage, the model will check these aspects of the translation sample and give corresponding evaluation annotations. Then, by evaluating the results obtained by the model, the quality of the prompt words used can be measured. The first quality score can be obtained by calculating the similarity between the second evaluation annotation and the first evaluation annotation. It can also be that the annotation personnel directly score the second evaluation standard.

[0051] The process of obtaining the target prompt word according to the first quality score is actually a process of screening and optimizing the candidate prompt words, and selecting the target prompt word that can enable the model to generate more accurate evaluation annotations based on multiple candidate prompt words. Specifically, it can be directly selecting the highest score as the target prompt word according to the first quality score. However, it can also be using the candidate prompt words to further generate more approximate samples, and further selecting the best from the best.

[0052] S108, using the target prompt words and the target large model to evaluate the translation to be evaluated, and obtaining a target evaluation result.

[0053] It can be understood that the translation to be evaluated is the text whose translation effect needs to be judged. It can be a newly generated translation work or a translation content to be reviewed obtained from an actual application scenario. The target evaluation result is the final evaluation conclusion output by the target large model after a comprehensive analysis and judgment of the translation to be evaluated under the guidance of the target prompt words. The target prompt words provide clear evaluation criteria and methods for the target large model. When the translation to be evaluated is input into the target large model, the model conducts a complete analysis of the input translation to be evaluated based on the evaluation dimensions set by the target prompt words, such as grammatical accuracy, semantic completeness, style consistency, etc. Moreover, once the target prompt words are confirmed, subsequent translations to be evaluated can use the same set of target prompt words for effect evaluation, without having to repeat the above adjustment and optimization process.

[0054] Based on the translation effect evaluation method in this embodiment, first obtain a translation data set, then generate training prompt words according to the translation data set and input them into the target large model to obtain multiple candidate prompt words, input the candidate prompt words and translation samples into the target large model, obtain a second evaluation annotation, and screen out the target prompt words based on the second evaluation annotation and the corresponding first quality score. Finally, the target prompt words and the target large model evaluate the translation to be evaluated to obtain the target evaluation result. This method does not rely on the quality of the reference translation, but is based on the semantic understanding ability of the large language model to allow it to continuously learn the evaluation rules of the translation effect, thereby generating a set of reasonable target prompt words, and then the target prompt words can be used to automatically evaluate the translation effect in a manner close to manual evaluation, which not only improves the evaluation accuracy, but also reduces labor costs.

[0055] In one embodiment, the target prompt word is obtained according to the first quality score, see Figure 2 , including steps S202 to S206.

[0056] S202: Screen multiple candidate prompt words according to the first quality score to obtain multiple preferred prompt words.

[0057] It can be understood that this step is intended to conduct preliminary screening and identification of numerous candidate prompt words through the quantitative index of the first quality score. Since the first quality score intuitively reflects the effectiveness of the candidate prompt word guiding the model evaluation, a higher score means that the candidate prompt word can enable the model to generate a result closer to the real evaluation. Based on this, a certain number of candidate prompt words can be directly selected as preferred prompt words in the order of the first quality score from high to low. Its core purpose is to narrow the scope of subsequent evaluation, focus on those candidate prompt words that show the potential of guiding the model to accurately translate and evaluate in the preliminary evaluation, lay the foundation for further in-depth evaluation in the future, and improve the efficiency and pertinence of the entire screening process. It can also be to set a scoring threshold, select the candidate prompt words above the threshold, and form a preferred prompt word set. The scoring threshold here can be dynamically adjusted, specifically, it can be the discrete degree of calculating the first quality score, and the higher the discrete degree, the lower the corresponding scoring threshold. In this way, more candidate prompt words with potential value can be included when the scoring distribution is relatively dispersed, and will not be mistakenly omitted due to unreasonable threshold setting.

[0058] S204, inputting the preferred prompt words and the translation samples into the target macro model to obtain a third evaluation annotation.

[0059] It can be understood that the translation sample here can be a sample extracted from the translation data set, and the third evaluation annotation is the result annotation output by the target large model after receiving the preferred prompt word and the translation sample, and re-evaluating the translation sample according to the evaluation rules and standards specified by the preferred prompt word. This step is to further verify and explore the ability of these preferred prompt words to guide the target large model to perform translation evaluation on the basis of screening out the preferred prompt words.

[0060] S206: Obtain the second quality scores corresponding to the third evaluation annotations, and select the preferred prompt word with the highest second quality score as the target prompt word.

[0061] It can be understood that the second quality score is a quantitative indicator introduced to further measure the quality of the third evaluation annotation generated by the target large model based on the preferred prompt word. This score is similar to the first quality score, but focuses on the result evaluation after a round of screening (S202) and re-evaluation (S204). By calculating the second quality score corresponding to each third evaluation annotation, it is possible to clearly understand the actual performance of each preferred prompt word in guiding the target large model to perform translation evaluation. A higher second quality score means that the preferred prompt word can enable the target large model to generate a third evaluation annotation that is highly consistent with the real evaluation annotation, that is, the preferred prompt word has a stronger ability and accuracy in guiding the model to evaluate the translation effect. Selecting the target prompt word with the highest second quality score among many preferred prompt words is based on the consideration of pursuing the best evaluation effect, ensuring that the final target prompt word can guide the target large model to the greatest extent when evaluating the subsequent translation to be evaluated, output the most accurate and reliable evaluation results, and optimize the performance of the entire translation evaluation system.

[0062] In one embodiment, before inputting the preferred prompt word and the translation sample into the target large model to obtain the third evaluation annotation, refer to Figure 3 , also including steps S302 to S308.

[0063] S302: For any selected preferred prompt word, use the preferred prompt word as a root node.

[0064] It can be understood that this embodiment is to use the originally selected preferred prompt words to further generate a batch of similar preferred prompt words to select better target prompt words. For an original preferred prompt word, taking it as the root node means building a tree structure for exploring and optimizing prompt words with it as the starting point, and all subsequent node traversals, mutation operations, etc. are derived from this node.

[0065] S304, traversing from the root node, if the current node does not have a first number of child nodes, performing a first number of mutation operations on the current node, and using the results of the mutation operations as child nodes of the current node.

[0066] It can be understood that the first number is a pre-set key parameter, and its value is determined comprehensively based on the complexity of the actual evaluation task, the diversity of candidate prompt words, and computing resources. Node traversal is a common method in tree structure operation. Starting from the root node, the nodes in the tree are accessed in a certain order (such as breadth first or depth first). If the current result traversed does not have the first number of child nodes, it means that the child node has not yet branched out, and a mutation operation needs to be performed on it to generate a prompt word that is similar to but different from the current node. The mutation operation is carried out on the current node, and the text content of the prompt word is modified by natural language processing means such as adjusting the word order, replacing synonyms, and adding or removing specific words to generate new prompt word forms, which become the child nodes of the current node. This process is similar to performing a local search in the search space. By mutating the existing prompt words, more possible prompt word forms are explored in the hope of finding a better prompt word to guide the evaluation of the target large model.

[0067] S306, otherwise, select the target child node according to the exploration weight corresponding to each child node, and update the cumulative reward values ​​of all nodes along the path from the target child node to the root node according to the third quality score corresponding to the target child node, and return to the step of traversal starting from the root node to continue execution until the loop end condition is met.

[0068] It can be understood that the exploration weight sets a value for each child node, which is used to characterize the possibility of the child node being selected during the exploration process. The target child node is selected from multiple child nodes according to the exploration weight of the child node. The third quality score is used to measure the quality performance of the target child node in the translation evaluation task. This score is similar to the first quality score, but focuses on the performance of the new preferred prompt word generated based on the preferred prompt word. The cumulative reward value has a corresponding record on each node, reflecting the comprehensive performance of each node in the entire exploration process. The cumulative reward value of each child node is updated based on the root node. If the mutation operation causes the third quality score to decrease, the cumulative reward value of the child node will also decrease based on the parent node, and vice versa. The magnitude of the increase and decrease is determined by the difference between the third quality scores of the child node and the parent node. The greater the difference, the greater the magnitude. For other nodes on its path, as the child node is traced back to the root node, the update amplitude of the cumulative reward value will gradually decay for each node passed. The exploration weight can be dynamically changed based on the number of times the node has been traversed, the cumulative reward value of the node, etc. This exploration weight setting method can encourage the exploration of paths with higher reward values. The loop end conditions are pre-set, and may include reaching a preset number of traversals, the cumulative reward value converging to a certain range, the number of new prompt words generated meeting the requirements, etc. Specifically, it may be when the depth of any branch reaches the preset maximum exploration depth, or when the cumulative reward value growth rate in the current iteration cycle is less than the convergence threshold.

[0069] S308: Use the second number of nodes with the highest cumulative reward values ​​as newly generated preferred prompt words.

[0070] It is understandable that the second quantity is a preset value, which can be determined according to the demand for the number of new preferred prompt words in the actual application scenario. The cumulative reward value is constantly updated in the previous steps, reflecting the comprehensive performance of the prompt word form represented by the node in the whole exploration process. A higher cumulative reward value means that the prompt word form represented by the node has better potential in guiding the model to perform translation evaluation. In this step, the second number of nodes with the highest cumulative reward value are selected from all nodes, and the prompt words stored by these nodes are used as the newly generated preferred prompt words. For example, if the second quantity is set to 3, the nodes with the top three cumulative reward values ​​are selected, and the prompt words corresponding to them are used as new preferred prompt words, which are used for subsequent guiding target large models to perform translation evaluation, and it is expected that these new preferred prompt words can enable the model to generate more accurate and high-quality evaluation results.

[0071] In one embodiment, the mutation operation includes at least one of sentence reorganization, synonym replacement, and adding format-qualifying phrases. Sentence reorganization is to rearrange the grammatical structure of the prompt word text, for example, adjusting "evaluate the translation from the aspects of grammar, semantics and terminology accuracy" to "evaluate the translation from the aspects of terminology accuracy, grammar and semantics". Synonym replacement is to replace some words in the prompt word with other words with similar meanings, such as replacing "evaluate" with "evaluate". Adding format-qualifying phrases is to add descriptive phrases with specific format requirements to the prompt word, such as adding format-qualifying content such as "check according to the subject-verb-object structure" to the prompt word, so as to generate new prompt word forms, and these new forms serve as child nodes of the current node.

[0072] In one embodiment, before evaluating the translation to be evaluated using the target prompt word and the target large model to obtain the target evaluation result, the method further includes: determining the domain label of the translation to be evaluated. If the domain label does not match the set domain label, a prompt word update prompt is issued.

[0073] It can be understood that the domain label is used to identify the domain to which the translation to be evaluated belongs, such as "medicine", "law", "technology", etc. It analyzes the content of the translated text and extracts key features to determine the professional field to which the text belongs. The set domain label is a pre-set domain label that is adapted to the translation dataset. Before using the target prompt word and the target large model to evaluate the translation to be evaluated, the domain label of the translation to be evaluated is determined to determine whether the current translation content is consistent with the domain adapted by the target prompt word. Translations in different fields have different characteristics and different evaluation logics are used. For example, the medical field has a large number of professional terms, and the legal field has extremely high requirements for the rigor and logic of the language, and the focus of the evaluation will also be on rigor and logic. If the field of the translation to be evaluated is inconsistent with the field targeted by the target prompt word, then the existing target prompt word may not accurately guide the target large model to evaluate, resulting in inaccurate evaluation results. By comparing the domain labels, once a mismatch is found, a prompt word update prompt is issued to prompt relevant personnel to take measures, such as regenerating the target prompt word applicable to the field, thereby ensuring the quality of translation evaluation.

[0074] In one embodiment, the translation to be evaluated is evaluated using a target prompt word and a target large model to obtain a target evaluation result, including: solidifying the target prompt word in the target large model, and inputting the translation to be evaluated into the target large model to obtain the target evaluation result.

[0075] It can be understood that the target prompt words are fixed in the target large model, which essentially allows the target large model to always perform evaluation operations based on these specific prompt words when encountering problems with translation effect evaluation during operation. Specifically, the system prompt of the target large model can be set to the target prompt words. After the target prompt words are fixed in the target large model, it is equivalent to converting the target large model into an agent that can stably perform translation effect evaluation, providing users with long-term, efficient and accurate translation effect evaluation services.

[0076] The present application provides a translation effect evaluation device, which includes a data acquisition module, a candidate prompt word generation module, a target prompt word generation module and an evaluation module.

[0077] The data acquisition module is used to acquire a translation data set. The translation data set includes multiple translation samples and their corresponding first evaluation annotations, and the first evaluation annotations include evaluation results and evaluation reasons. The candidate prompt word generation module is used to generate training prompt words according to the translation data set, input the training prompt words into the target large model, and obtain multiple candidate prompt words. The training prompt words are used to instruct the target large model to learn the evaluation method according to the translation data set, and generate multiple candidate prompt words for instructing the target large model to evaluate the translation effect. The target prompt word generation module is used to input the candidate prompt words and the translation samples into the target large model, obtain the second evaluation annotation, obtain the first quality score corresponding to each second evaluation annotation, and obtain the target prompt word according to the first quality score. The evaluation module is used to evaluate the translation to be evaluated using the target prompt words and the target large model to obtain the target evaluation result.

[0078] For the specific definition of the melody generation device, please refer to the definition of the translation effect evaluation method above, which will not be repeated here. The various modules in the above-mentioned melody generation device can be implemented in whole or in part by software, hardware and their combination. The above-mentioned modules can be embedded in or independent of the processor in the computer device in the form of hardware, or can be stored in the memory of the computer device in the form of software, so that the processor can call and execute the operations corresponding to the above modules. It should be noted that the division of modules in the embodiment of the present application is schematic and is only a logical function division. There may be other division methods in actual implementation.

[0079] The present application provides a computer device, including one or more processors and a memory, wherein the memory stores computer-readable instructions, and when the computer-readable instructions are executed by the one or more processors, the steps of the translation effect evaluation method in any of the above embodiments are executed.

[0080] Indicatively, Figure 4 As shown, Figure 4 A schematic diagram of the internal structure of a computer device provided in an embodiment of the present application. Figure 4 The computer device 400 includes a processing component 402, which further includes one or more processors, and a memory resource represented by a memory 401, for storing instructions that can be executed by the processing component 402, such as an application. The application stored in the memory 401 may include one or more modules, each corresponding to a set of instructions. In addition, the processing component 402 is configured to execute instructions to perform the steps of the translation effect evaluation method of any of the above embodiments.

[0081] The computer device 400 may further include a power supply component 403 configured to perform power management of the computer device 400 , a wired or wireless model interface 404 configured to connect the computer device 400 to the model, and an input / output (I / O) interface 405 .

[0082] The present application provides a storage medium storing computer-readable instructions. When the computer-readable instructions are executed by one or more processors, the one or more processors execute the steps of the translation effect evaluation method in any of the above embodiments.

[0083] Finally, it should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "comprise a ..." do not exclude the presence of other identical elements in the process, method, article or device including the elements.

[0084] The various embodiments in this specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The various embodiments can be combined as needed, and the same or similar parts can refer to each other.

[0085] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present application. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application will not be limited to the embodiments shown herein, but will conform to the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A translation effect evaluation method, characterized in that: include: Get the translation dataset; The translation data set includes a plurality of translation samples and their corresponding first evaluation annotations, wherein the first evaluation annotations include evaluation results and evaluation reasons; Generate training prompt words according to the translation data set, input the training prompt words into the target large model, and obtain multiple candidate prompt words; The training prompt words are used to instruct the target large model to learn the evaluation method according to the translation data set, and generate a plurality of candidate prompt words for instructing the target large model to evaluate the translation effect; Inputting the candidate prompt words and the translation sample into the target large model to obtain second evaluation annotations, obtaining first quality scores corresponding to each of the second evaluation annotations, and obtaining target prompt words according to the first quality scores; The translation to be evaluated is evaluated using the target prompt word and the target large model to obtain a target evaluation result.

2. The translation effect evaluation method according to claim 1, characterized in that: The obtaining of a target prompt word according to the first quality score includes: Screening the plurality of candidate prompt words according to the first quality score to obtain a plurality of preferred prompt words; Inputting the preferred prompt word and the translation sample into the target macro model to obtain a third evaluation annotation; The second quality score corresponding to each of the third evaluation annotations is obtained, and the preferred prompt word with the highest second quality score is selected as the target prompt word.

3. The translation effect evaluation method according to claim 2, characterized in that: Before inputting the preferred prompt word and the translation sample into the target large model to obtain the third evaluation annotation, the method further includes: For any selected preferred prompt word, use the preferred prompt word as a root node; Traversing from the root node, if the current node does not have a first number of child nodes, performing the first number of mutation operations on the current node, and using the results of the mutation operations as the child nodes of the current node; Otherwise, a target child node is selected according to the exploration weight corresponding to each of the child nodes, and the cumulative reward values ​​of all nodes along the path from the target child node to the root node are updated according to the third quality score corresponding to the target child node, and the step of traversing from the root node is returned to continue execution until the loop end condition is met; The second number of nodes with the highest cumulative reward values ​​are used as the newly generated preferred prompt words.

4. The translation effect evaluation method according to claim 3, characterized in that: The variation operation includes at least one of sentence reorganization, synonym replacement, and adding format-restricted phrases.

5. The translation effect evaluation method according to claim 3, characterized in that: The loop end conditions include: when the depth of any branch reaches a preset maximum exploration depth, or the growth rate of the cumulative reward value in the current iteration cycle is less than the convergence threshold.

6. The translation effect evaluation method according to claim 1, characterized in that: Before evaluating the translation to be evaluated by using the target prompt word and the target large model to obtain the target evaluation result, the method further includes: Determining a domain label of the translation to be evaluated; When the domain label does not match the set domain label, a prompt word update prompt is issued.

7. The translation effect evaluation method according to claim 1, characterized in that: The step of evaluating the translation to be evaluated by using the target prompt word and the target large model to obtain a target evaluation result includes: The target prompt word is solidified in the target macro model, and the translation to be evaluated is input into the target macro model to obtain the target evaluation result.

8. A translation effect evaluation device, characterized in that: include: A data acquisition module, used to acquire translation data sets; The translation data set includes a plurality of translation samples and their corresponding first evaluation annotations, wherein the first evaluation annotations include evaluation results and evaluation reasons; A candidate prompt word generation module is used to generate training prompt words according to the translation data set, input the training prompt words into the target large model, and obtain multiple candidate prompt words; The training prompt words are used to instruct the target large model to learn the evaluation method according to the translation data set, and generate a plurality of candidate prompt words for instructing the target large model to evaluate the translation effect; a target prompt word generation module, configured to input the candidate prompt words and the translation sample into the target large model, obtain a second evaluation annotation, obtain a first quality score corresponding to each of the second evaluation annotations, and obtain a target prompt word according to the first quality score; The evaluation module is used to evaluate the translation to be evaluated using the target prompt word and the target large model to obtain a target evaluation result.

9. A computer device, characterized in that: The method comprises one or more processors and a memory, wherein the memory stores computer-readable instructions, and when the computer-readable instructions are executed by the one or more processors, the steps of the translation effect evaluation method according to any one of claims 1 to 7 are executed.

10. A storage medium, characterized in that: The storage medium stores computer-readable instructions, and when the computer-readable instructions are executed by one or more processors, the one or more processors execute the steps of the translation effect evaluation method according to any one of claims 1 to 7.