Translation difficulty assessment method and device based on back-translated text, equipment and medium

By obtaining the original text and translating and back-translating, semantic similarity and differences are calculated, the problem of relying on manual annotation and feature engineering in the existing technology is solved, and efficient and accurate text translation difficulty evaluation is achieved.

CN120337944APending Publication Date: 2025-07-18PENG CHENG LAB
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510337415.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-20
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The text translation difficulty evaluation method in the prior art relies too much on manual annotation and cumbersome feature engineering, and cannot efficiently evaluate text translation difficulty.

Method used

By obtaining the original text, using the translation channel to translate it from the source language to the target language and back-translated to the source language, the semantic similarity between the original text and the back-translated text is calculated, and the translation difficulty score of the text is determined based on the semantic difference coefficient and the translation mode.

Benefits of technology

It improves the efficiency and accuracy of text translation difficulty evaluation, avoids manual annotation and feature engineering, and realizes more general translation difficulty evaluation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120337944A_ABST
    Figure CN120337944A_ABST
Patent Text Reader

Abstract

The invention provides a translation difficulty assessment method and device based on a back-translated text, equipment and a medium, and belongs to the technical field of text translation.The translation difficulty assessment method comprises the steps that an original text is obtained, the original text is translated into a target language from a source language through a translation channel, and the translated text is obtained; the method comprises the steps of translating a translated text from a target language to a source language through a translation channel to obtain a back-translated text, calculating semantic similarity between an original text and the back-translated text, determining a translation mode of the translation channel, and determining a translation difficulty score from the original text to the target language according to a semantic difference coefficient, the semantic similarity and the translation mode. The efficiency of text translation difficulty assessment can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of text translation, and particularly to a method, apparatus, device and medium for evaluating the translation difficulty based on back-translated text. Background Art

[0002] The translation difficulty of text refers to the degree of difficulty in translating a piece of text from the source language to the target language. In related technologies, features of the source text are extracted from multiple dimensions based on the linguistic characteristics of the text, such as text length, vocabulary, grammar, and theme, and the features are evaluated for difficulty in multiple dimensions manually to comprehensively calculate the translation difficulty of the text. Or a translation difficulty prediction model is trained using a corpus with annotated text translation difficulties, and the translation difficulty of the text is predicted through the trained translation difficulty prediction model. However, these methods rely too much on manual annotation or cumbersome feature engineering and cannot efficiently evaluate the translation difficulty of text. Summary of the Invention

[0003] The main purpose of the embodiments of this application is to propose a method, apparatus, device and medium for evaluating the translation difficulty based on back-translated text, aiming to improve the efficiency of evaluating the translation difficulty of text.

[0004] To achieve the above object, a first aspect of the embodiments of this application proposes a method for evaluating the translation difficulty based on back-translated text, the method including:

[0005] Obtain the original text;

[0006] Translate the original text from the source language to the target language through a translation channel to obtain a translated text; the translation channel has a semantic difference coefficient;

[0007] Translate the translated text from the target language to the source language through the translation channel to obtain a back-translated text;

[0008] Calculate the semantic similarity between the original text and the back-translated text;

[0009] Determine the translation mode of the translation channel;

[0010] Determine the translation difficulty score of the original text to the target language according to the semantic difference coefficient, the semantic similarity and the translation mode.

[0011] In some embodiments, the determining the translation difficulty score of the original text to the target language according to the semantic difference coefficient, the semantic similarity and the translation mode includes:

[0012] Obtain the number of channels of the translation channel;

[0013] If the number of the channels is equal to one, determine the initial translation difficulty of the original text into the target language according to the semantic difference coefficient, the semantic similarity, and the translation mode, and use the initial translation difficulty as the translation difficulty score;

[0014] If the number of the channels is greater than one, determine the initial translation difficulty of the original text into the target language for each translation channel according to the semantic difference coefficient, the corresponding semantic similarity, and the translation mode, obtain the channel weight of each translation channel, and perform a weighted average calculation on the initial translation difficulty of the corresponding translation channel according to the channel weight of each translation channel to obtain the translation difficulty score.

[0015] In some embodiments, determining the initial translation difficulty of the original text into the target language according to the semantic difference coefficient, the semantic similarity, and the translation mode includes:

[0016] If the translation mode is the human translation mode, determine the semantic difference degree according to the semantic similarity;

[0017] Multiply the semantic difference coefficient by the semantic difference degree to obtain the initial translation difficulty of the original text into the target language.

[0018] In some embodiments, determining the initial translation difficulty of the original text into the target language according to the semantic difference coefficient, the semantic similarity, and the translation mode includes:

[0019] If the translation mode is the machine translation mode and the semantic similarity is greater than or equal to a preset similarity threshold, determine the semantic difference degree according to the semantic similarity, multiply the semantic difference coefficient by the semantic difference degree to obtain the initial translation difficulty of the original text into the target language;

[0020] If the translation mode is the machine translation mode and the semantic similarity is less than the preset similarity threshold, use the back-translated text as the original text input to the translation channel, and repeat the steps of translating the original text from the source language into the target language through the translation channel to obtain a translated text, translating the translated text from the target language into the source language through the translation channel to obtain a back-translated text, and calculating the semantic similarity between the original text and the back-translated text until the semantic similarity is greater than or equal to the preset similarity threshold, determine the semantic difference degree according to the target similarity, multiply the semantic difference coefficient by the semantic difference degree to obtain the initial translation difficulty of the original text into the target language; the target similarity is the semantic similarity between the first original text and the last back-translated text.

[0021] In some embodiments, after determining the translation difficulty score of the original text into the target language according to the semantic difference coefficient, the semantic similarity, and the translation mode, the translation difficulty evaluation method based on the back-translated text further includes:

[0022] Performing text screening on the original text according to the translation difficulty score of the original text into the target language to obtain a test text;

[0023] Testing the preset translation model according to the test text to obtain the translation accuracy score of the preset translation model;

[0024] Determining a target text according to the translation accuracy score and the original text;

[0025] Updating the model parameters of the preset translation model based on the target text to obtain a target translation model.

[0026] In some embodiments, the determining the target text according to the translation accuracy score and the original text includes:

[0027] If the translation accuracy score is less than a first preset score threshold, adding adversarial noise to the original text to obtain a noise text, calculating a reference translation difficulty score of the noise text into the target language, and combining the original text and the noise text according to the reference translation difficulty score and the translation difficulty score to obtain the target text;

[0028] If the translation accuracy score is greater than or equal to the first preset score threshold and less than a second preset score threshold, determining the difficulty level of the original text according to the translation difficulty score of the original text into the target language, and screening the original text according to the difficulty level to obtain the target text; wherein, the second preset score threshold is greater than the first preset score threshold.

[0029] In some embodiments, the updating the model parameters of the preset translation model based on the target text to obtain a target translation model includes:

[0030] Translating the target text from the source language into the target language through the preset translation model to obtain a target translation text;

[0031] Translating the target translation text from the target language into the source language through the preset translation model to obtain a target back-translation text;

[0032] Calculating a target loss according to the target text and the target back-translation text;

[0033] Update the model parameters of the preset translation model according to the target loss to obtain the target translation model.

[0034] To achieve the above object, a second aspect of the embodiments of the present application proposes a translation difficulty evaluation device based on back-translated text, and the device includes:

[0035] An acquisition module, configured to acquire the original text;

[0036] A first translation module, configured to translate the original text from the source language into the target language through a translation channel to obtain a translated text; the translation channel has a semantic difference coefficient;

[0037] A second translation module, configured to translate the translated text from the target language into the source language through the translation channel to obtain a back-translated text;

[0038] A calculation module, configured to calculate the semantic similarity between the original text and the back-translated text;

[0039] A translation mode determination module, configured to determine the translation mode of the translation channel;

[0040] A translation difficulty evaluation module, configured to determine a translation difficulty score from the original text to the target language according to the semantic difference coefficient, the semantic similarity, and the translation mode.

[0041] To achieve the above object, a third aspect of the embodiments of the present application proposes an electronic device, which includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the translation difficulty evaluation method based on back-translated text in the first aspect is implemented.

[0042] To achieve the above object, a fourth aspect of the embodiments of the present application proposes a computer-readable storage medium, which stores a computer program, and when the computer program is executed by a processor, the translation difficulty evaluation method based on back-translated text in the first aspect is implemented.

[0043] The translation difficulty evaluation method, translation difficulty evaluation device, electronic device, and computer-readable storage medium based on back-translated text proposed in this application estimate the translation difficulty of the original text by obtaining the original text. To improve the efficiency of text translation difficulty evaluation and avoid over-reliance on manual annotation and cumbersome feature engineering, text translation difficulty evaluation is performed based on the back-translation technique. The original text is translated from the source language to the target language through a translation channel to obtain the translated text. Ideally, the semantic meaning of the translated text should be the same as that of the original text, but in actual situations, there will be semantic differences between the translated text and the original text. The magnitude of the semantic difference is positively correlated with the text translation difficulty of the original text. The greater the semantic difference, the greater the text translation difficulty. Considering that the judgment of semantic differences between texts of the same language type is more accurate than that between texts of different language types, the translated text is translated from the target language to the source language through a translation channel to obtain a back-translated text with the same language type as the original text. Calculate the semantic similarity between the original text and the back-translated text to measure the translation difficulty of the text based on the semantic deviation between the original text and the back-translated text. Considering the influence of different translation methods on translation difficulty evaluation and achieving a more general text translation difficulty evaluation, determine the translation mode of the translation channel, and quantitatively calculate the translation difficulty score from the original text to the target language based on the semantic difference coefficient, semantic similarity, and translation mode, avoiding manual annotation and cumbersome feature engineering processes and improving the efficiency of text translation difficulty estimation. Description of the Drawings

[0044] Figure 1 is a flowchart of the translation difficulty evaluation method based on back-translated text provided by an embodiment of this application;

[0045] Figure 2 is Figure 1 a flowchart of step S160 in

[0046] Figure 3 is Figure 2 a flowchart of step S220 in

[0047] Figure 4 is Figure 2 another flowchart of step S220 in

[0048] Figure 5 is another flowchart of the translation difficulty evaluation method based on back-translated text provided by an embodiment of this application;

[0049] Figure 6 is Figure 5 a flowchart of step S530 in

[0050] Figure 7 is Figure 5 a flowchart of step S540 in

[0051] Figure 8 It is a schematic structural diagram of a translation difficulty evaluation device based on back-translated text provided by an embodiment of the present application;

[0052] Figure 9 It is a schematic hardware structure diagram of an electronic device provided by an embodiment of the present application. Specific embodiments

[0053] In order to make the objectives, technical solutions and advantages of the present application more clear and understandable, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0054] It should be noted that although the functional modules are divided in the device schematic diagram and the logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order from the module division in the device or the order in the flowchart. Terms such as "first" and "second" in the specification, claims and the above-mentioned drawings are used to distinguish similar objects and do not necessarily need to describe a specific order or sequence.

[0055] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which this application belongs. The terms used herein are only for the purpose of describing the embodiments of this application and are not intended to limit this application.

[0056] The translation difficulty of a text refers to the degree of difficulty in translating a text from the source language to the target language. In the related art, features of the source text are extracted from multiple dimensions based on the linguistic characteristics of the text, such as text length, vocabulary, grammar, and theme, and the features are evaluated for difficulty in multiple dimensions manually to comprehensively calculate the translation difficulty of the text. Or a translation difficulty prediction model is trained using a corpus with labeled translation difficulties, and the translation difficulty of the text is predicted through the trained translation difficulty prediction model. However, these methods rely too much on manual annotation or cumbersome feature engineering and cannot efficiently evaluate the translation difficulty of the text.

[0057] Based on this, the embodiments of the present application provide a method for evaluating translation difficulty based on back-translated text, a device for evaluating translation difficulty based on back-translated text, an electronic device, and a computer-readable storage medium, aiming to improve the efficiency of evaluating the translation difficulty of the text.

[0058] The method for evaluating translation difficulty based on back-translated text, the device for evaluating translation difficulty based on back-translated text, the electronic device, and the computer-readable storage medium provided by the embodiments of the present application will be specifically described through the following embodiments. First, the method for evaluating translation difficulty based on back-translated text in the embodiments of the present application will be described.

[0059] The translation difficulty assessment method based on back-translated text provided by the embodiments of the present application relates to the technical field of text translation. The translation difficulty assessment method based on back-translated text provided by the embodiments of the present application can be applied to a terminal, or to a server side, or can be software running on a terminal or a server side. In some embodiments, the terminal can be a smart phone, a tablet computer, a notebook computer, a desktop computer, etc.; the server side can be configured as an independent physical server, or can be configured as a server cluster or a distributed system composed of multiple physical servers, or can be configured as a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms; the software can be an application that implements the translation difficulty assessment method based on back-translated text, etc., but is not limited to the above forms.

[0060] The present application can be used in many general or special computer system environments or configurations. For example: personal computers, server computers, handheld or portable devices, tablet devices, multi-processor systems, microprocessor-based systems, set-top boxes, programmable consumer electronic devices, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, and so on. The present application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. The present application can also be practiced in a distributed computing environment where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including storage devices.

[0061] Figure 1 is an optional flowchart of a translation difficulty assessment method based on back-translated text provided by the embodiments of the present application, Figure 1 The method in can include but is not limited to steps S110 to S160.

[0062] Step S110, obtain the original text;

[0063] Step S120, translate the original text from the source language to the target language through a translation channel to obtain a translated text; the translation channel has a semantic difference coefficient;

[0064] Step S130, translate the translated text from the target language to the source language through the translation channel to obtain a back-translated text;

[0065] Step S140, calculate the semantic similarity between the original text and the back-translated text;

[0066] Step S150, determine the translation mode of the translation channel;

[0067] Step S160, determine the translation difficulty score from the original text to the target language according to the semantic difference coefficient, semantic similarity, and translation mode.

[0068] In step S110 of some embodiments, the original text is collected from a specific data source. The data source can be an open-source dataset, a specific website, a social media platform, etc. The original text is the text to be evaluated for translation difficulty.

[0069] In step S120 of some embodiments, the original text is converted from the source language to the target language through a translation channel, and the text content obtained after the language conversion is used as the translated text. The translation channel is the way and platform to obtain text translation services. The translation channel can be a human translator, a translation engine platform, or a translation model, etc. The translation channel has a semantic difference coefficient, which is used to measure the degree of difference between the original text and the translated text and can be used as an evaluation index for translation difficulty. The source language is the language type used for the original text to be translated, the target language is the translated language type, and the translated text uses the target language.

[0070] In step S130 of some embodiments, the translated text is converted from the target language to the source language through a translation channel, and the text content obtained after the language conversion is used as the back-translated text. Both the back-translated text and the original text use the source language. The translation channel has a translation mode, which is the translation method adopted by the translation channel. It should be noted that the translation methods used in the process of translating from the source language to the target language and the process of back-translating from the target language to the source language can be different, but the translation method used by the translation channel must be kept unchanged in the same translation difficulty evaluation task. For example, when evaluating the translation difficulty of the original text, if the translation method used in the process of translating from the source language to the target language is human translation, then the translation method used in the process of back-translating from the target language to the source language must also be human translation. If a text translation model is used for translation from the source language to the target language, then the back-translation from the target language to the source language also uses this text translation model.

[0071] In step S140 of some embodiments, for a given original text t, it is translated from the source language S (source language) to the target language T (target language) through the translation channel M to obtain the translated text t ST . Ideally, the translated text t STIt should have the same semantics as the original text t, but in practice, there will be semantic differences between the two. And the magnitude of the semantic difference is positively correlated with the translation difficulty of the original text t. The greater the translation difficulty, the greater the semantic difference. Considering that it is easier and more accurate to judge the semantic difference between texts in the same language than between texts in different languages, the embodiments of this application use the translation channel M to translate the translated text t ST from the target language T back to the source language S, obtaining a back-translated text t in the same language as the original text t TS . The magnitude of the semantic difference between the original text t and the back-translated text t TS is also positively correlated with the translation difficulty. The greater the translation difficulty, the greater the semantic difference between the two. Therefore, the embodiments of this application measure the translation difficulty of the original text based on the semantic difference between the back-translated text and the original text.

[0072] Specifically, methods such as text edit distance, cosine similarity, pre-trained model prediction, or self-trained model prediction can be used to calculate the semantic similarity between the original text and the back-translated text, so as to measure the semantic difference degree based on the semantic similarity. The larger the value of the semantic similarity, the smaller the value of the semantic difference degree, and the more similar the original text and the back-translated text are.

[0073] In step S150 of some embodiments, obtain the translation method adopted by the translation channel to obtain the translation mode. The translation mode includes a human translation mode or a machine translation mode. Among them, the human translation mode is based on the translation technology of human translators, and the machine translation mode is based on the translation technology of computer software and algorithms, such as translation engines, translation models, etc.

[0074] Please refer to Figure 2 , in some embodiments, step S160 may include but is not limited to steps S210 to S230:

[0075] Step S210, obtain the number of channels of the translation channel;

[0076] Step S220, if the number of channels is equal to one, determine the initial translation difficulty from the original text to the target language according to the semantic difference coefficient, semantic similarity, and translation mode, and use the initial translation difficulty as the translation difficulty score;

[0077] Step S230, if the number of channels is greater than one, determine the initial translation difficulty from the original text to the target language for each translation channel according to the semantic difference coefficient, corresponding semantic similarity, and translation mode of each translation channel, obtain the channel weight of each translation channel, and perform a weighted average calculation on the initial translation difficulty of the corresponding translation channel according to the channel weight of each translation channel to obtain the translation difficulty score.

[0078] In step S210 of some embodiments, the number of translation channels is obtained to get the channel number. The channel number is an integer greater than or equal to 1.

[0079] In step S220 of some embodiments, if the channel number is equal to one, it indicates that a single translation channel is used to evaluate the text translation difficulty of the original text. Then, the translation difficulty is evaluated according to the semantic difference coefficient, semantic similarity, and translation mode to obtain the initial translation difficulty from the original text to the target language, and the initial translation difficulty is used as the translation difficulty score from the original text to the target language. The initial translation difficulty is used to measure the difficulty level of translating the original text from the source language to the target language by a single translation channel. The initial translation difficulty can be represented by a value within the range of [0, 1]. The greater the initial translation difficulty, the greater the translation difficulty of the original text.

[0080] In step S230 of some embodiments, if the channel number is greater than one, it indicates that multiple translation channels are used to evaluate the text translation difficulty of the original text. There are differences in the translation capabilities of different translation channels. If the translation difficulty of the text is calculated only based on the back-translated text of one translation channel, the obtained translation difficulty score result is only the translation difficulty of the text for that translation channel, and the generality of the translation difficulty evaluation is not strong. To make the text translation difficulty evaluation more general, for each translation channel, refer to step S220 to evaluate the translation difficulty according to the semantic difference coefficient, semantic similarity, and translation mode corresponding to the translation channel, and obtain the initial translation difficulty from the original text to the target language corresponding to the translation channel. The channel weight of each translation channel is obtained. The sum of the channel weights of each translation channel is 1. The channel weight is used to measure the importance of the translation channel in the text translation difficulty evaluation, that is, the level of translation ability. The channel weight of each translation channel can be determined according to the actual difficulty evaluation task. The channel weights of each translation channel can be the same, or different channel weights can be set according to the importance of different translation channels. Higher channel weights can be assigned to professional human translation and machine translation channels with high translation accuracy. The initial translation difficulty of the translation channel is calculated by weighted average according to the channel weight of each translation channel, and the weighted average value obtained by the weighted average calculation is used as the translation difficulty score. The translation difficulty score is used to represent the final comprehensive translation difficulty of the original text from the source language to the target language, and can be represented by a value within the range of [0, 1]. The greater the translation difficulty score, the greater the translation difficulty of the original text. The translation difficulty score is expressed as:

[0081]

[0082] where, t represents the original text; MTTD(t) represents the translation difficulty score from the original text to the target language; n represents the channel number; S represents the source language; T represents the target language; M i represents the i-th translation channel; λi Represents the channel weight of the \(i\)th translation channel; TTD represents the initial translation difficulty of the original text from the source language to the target language based on the translation channel.

[0083] The translation mode of multiple translation channels is flexible. It can use all manual translation modes or machine translation modes, or a mixed use of manual translation modes and machine translation modes. If all machine translation modes are used for multiple translation channels, a fully automatic evaluation of translation difficulty can be achieved. The embodiments of the present application can not only distinguish the translation difficulties of different single texts, but also evaluate the overall translation difficulty of a given data set.

[0084] Specifically, given a data set \(\{t_1,t_2,\cdots,t\) i ,\cdots,t\) m}\), \(t\) i is the \(i\)th original text in the data set. Refer to step S220 or step S230 to calculate the translation difficulty scores \(\{MTTD(t_1),MTTD(t_2),\cdots,MTTD(t\) i ),\cdots,MTTD(t\) m )\) of each original text in the data set to the target language, so that the obtained text translation difficulty can be applied according to the specific usage scenario requirements, such as sorting and annotating task assignment for texts according to the magnitude of translation difficulty, resampling data set samples according to a certain difficulty distribution, etc. Determine the overall translation difficulty score of the data set according to the translation difficulty scores of each original text to the target language. The overall translation difficulty score is expressed as:

[0085]

[0086] where STTD represents the overall translation difficulty score; \(m\) is the number of texts included in the data set; MTTD is the translation difficulty score of the original text to the target language.

[0087] In the above steps S210 to S230, the translation difficulty of the original text is quantified by the semantic loss amount between the back-translated text and the original text, avoiding manual annotation and cumbersome feature engineering, and improving the efficiency of text translation difficulty evaluation. And considering the back-translated texts of multiple different translation channels takes into account the translation ability differences of different translation channels, improving the generality and accuracy of text translation difficulty evaluation.

[0088] Please refer to Figure 3 , in some embodiments, step S220 may include but is not limited to steps S310 to S320:

[0089] Step S310, if the translation mode is the manual translation mode, determine the semantic difference degree according to the semantic similarity;

[0090] Step S320: Multiply the semantic difference coefficient and the semantic difference degree to obtain the initial translation difficulty from the original text to the target language.

[0091] In step S310 of some embodiments, if the translation mode is the manual translation mode, subtract one from the semantic similarity degree to obtain the semantic difference degree. The semantic difference degree is used to measure the semantic loss amount between the original text and the back-translated text. The greater the semantic difference degree, the greater the difference degree between the original text and the back-translated text.

[0092] In step S320 of some embodiments, multiply the semantic difference coefficient and the semantic difference degree to obtain the initial translation difficulty from the original text to the target language. The initial translation difficulty is expressed as:

[0093] TTD M (t, S, T) = α(1 - Sim(t, t TS ))

[0094] wherein, TTD represents the initial translation difficulty of the original text t from the source language S to the target language T based on the translation channel M; α represents the semantic difference coefficient, which can be set by itself; t TS represents the back-translated text obtained by first translating the original text t from the source language S to the target language T through the translation channel M and then translating it back from the target language T to the source language S; Sim represents the semantic similarity between the original text and the back-translated text.

[0095] Through the above steps S310 to S320, the translation difficulty value of the original text can be determined without relying on manual annotation and complex feature engineering, improving the efficiency of text translation difficulty assessment.

[0096] Please refer to Figure 4 , in some embodiments, step S220 may further include but is not limited to steps S410 to S420:

[0097] Step S410: If the translation mode is the machine translation mode and the semantic similarity is greater than or equal to the preset similarity threshold, determine the semantic difference degree according to the semantic similarity, and multiply the semantic difference coefficient and the semantic difference degree to obtain the initial translation difficulty from the original text to the target language;

[0098] Step S420: If the translation mode is the machine translation mode and the semantic similarity is less than the preset similarity threshold, then use the back-translated text as the original text input to the translation channel, and repeatedly execute the steps of translating the original text from the source language to the target language through the translation channel to obtain the translated text, translating the translated text from the target language to the source language through the translation channel to obtain the back-translated text, and calculating the semantic similarity between the original text and the back-translated text until the semantic similarity is greater than or equal to the preset similarity threshold. Determine the semantic difference degree according to the target similarity, and multiply the semantic difference coefficient by the semantic difference degree to obtain the initial translation difficulty from the original text to the target language; the target similarity is the semantic similarity between the first original text and the last back-translated text.

[0099] In step S410 of some embodiments, during the text translation and back-translation processes, the content with a relatively low translation difficulty remains unchanged, while the content with a relatively high translation difficulty experiences content and semantic losses. By translating the text from the source language to the target language through the translation channel and then back from the target language to the source language, after multiple round-trip translations, the content that is easy to translate remains unchanged. After multiple semantic losses or semantic changes in the difficult-to-translate content, it will also become content that the translation channel can translate. Eventually, the back-translated text will tend to converge and the content will be stable and unchanged. This back-translated text with stable content is called the balanced text.

[0100] If the translation mode is the human translation mode, the translation accuracy of the human translation mode is relatively high. To reduce the complexity and cost of manual annotation, the back-translated text obtained from a single round-trip translation is used to calculate the translation difficulty score from the original text to the target language. However, the machine translation mode cannot accurately translate the given text content, and there will be certain semantic losses during the translation process. The back-translated text obtained from multiple round-trip translations will lose some grammar or semantics of the original text. The relatively difficult-to-understand words or deep meanings in the original text are likely to be lost, and the final balanced text will only retain the main information of the original text. Therefore, the multiple round-trip translation process can be regarded as pruning the original text until only the main trunk remains. From this perspective, if the vocabulary, semantics, or grammar of the original text is relatively simple and plain, then the information lost in multiple round-trip translations is less, the semantic retention degree of the balanced text is high, and the translation difficulty of the corresponding original text is low; on the contrary, if the vocabulary, semantics, or grammar of the original text is relatively complex and profound, then the information lost in multiple round-trip translations is more, the semantic retention degree of the balanced text is low, and the translation difficulty of the corresponding original text is high. If the translation mode is the machine translation mode, then the translation difficulty is evaluated based on the semantic retention degree of the balanced text obtained from multiple round-trip translations to accurately and efficiently calculate the comprehensive translation difficulty of the text.

[0101] Specifically, if the translation mode is the machine translation mode and the semantic similarity is greater than or equal to the preset similarity threshold, it indicates that the semantic retention degree of the back-translated text is relatively high and the translation difficulty of the original text is relatively low. Then, referring to steps S310 to S320, subtract the semantic similarity to obtain the semantic difference degree, and multiply the semantic difference coefficient by the semantic difference degree to obtain the initial translation difficulty from the original text to the target language. The preset similarity threshold is the limit value of the semantic similarity and can be set according to actual situations.

[0102] In step S420 of some embodiments, if the translation mode is the machine translation mode and the semantic similarity is less than the preset similarity threshold, it indicates that the semantic retention degree of the back-translated text is relatively low and the translation difficulty of the original text is relatively high. Then, use the stable and unchanging balanced text obtained through multiple round-trip translations as the final back-translated text, and calculate the initial translation difficulty. Specifically, use the back-translated text as the original text input to the translation channel, and repeatedly execute the steps of translating the original text from the source language to the target language through the translation channel to obtain the translated text, and then translating the translated text from the target language to the source language through the translation channel to obtain the back-translated text, and calculate the semantic similarity between the original text input this time and the obtained back-translated text until the semantic similarity is greater than or equal to the preset similarity threshold, indicating that the back-translated text tends to be stable. Then, use the semantic similarity between the initial original text and the balanced text as the target similarity, subtract one from the target similarity to obtain the semantic difference degree, and multiply the semantic difference coefficient by the semantic difference degree to obtain the initial translation difficulty.

[0103] The above steps S410 to S420 evaluate the text translation difficulty based on the stable and unchanging characteristics of the balanced text, avoiding the adverse effects on the translation difficulty evaluation caused by the relatively low translation accuracy of the machine translation mode for complex texts, and can accurately and efficiently evaluate the text translation difficulty.

[0104] For a given original text t, the detailed evaluation steps for its translation difficulty value MTTD(t) from the source language to the target language are as follows:

[0105] S0, Initialize n translation channels that support translation from the source language to the target language and from the target language to the source language. These translation channels are M1, M2, M3,..., M n , Set the channel weight values of each translation channel as λ1, λ2, λ3,..., λ n , and keep Set the semantic difference coefficient values of each translation channel as α1, α2,..., α n ;

[0106] S1, Obtain the original text t;

[0107] S2. For each translation channel, use the original text t as the source text c and input it into the translation channel, i.e., c = t, and perform the following text translation difficulty assessment process:

[0108] S21. Through the translation channel M i Translate the source text c from the source language to the target language to obtain the translated text tt i ;

[0109] S22. Through the translation channel M i Translate the translated text tt i Back from the target language to the source language to obtain the back-translated text btt i ;

[0110] S23. Determine whether the translation channel is manual translation. If the judgment result is yes, execute S25; if the judgment result is no, execute S24;

[0111] S24. Determine whether the content of the back-translated text btt i Is the same as the source text c. If the judgment result is yes, execute S25; if the judgment result is no, use the back-translated text btt i As the source text c, i.e., c = btt i , and execute S21;

[0112] S25. Calculate the translation difficulty value of the original text t based on the translation channel M i Of

[0113] S3. Calculate the weighted average of the translation difficulty values of n translation channels to obtain the comprehensive translation difficulty value MTTD(t).

[0114] To achieve an efficient and accurate assessment of the difficulty of text translation, the embodiments of this application propose a text translation difficulty assessment method based on multi-back-translated texts that can be quantitatively calculated. The method assesses the translation difficulty of text from the source language to the target language based on the semantic loss between the back-translated result and the original text, effectively integrating the characteristics of the text itself, the characteristics between the source language and the target language, and the characteristics of the translation channel. The assessment results have stronger universality and higher comparability, and can support both machine translation-based assessment and human translation-based assessment, enhancing the flexibility of translation difficulty assessment. In the assessment mode based on fully machine translation, the assessment process does not require human intervention, with higher assessment efficiency, lower cost, and high reliability of the assessment results. It can effectively reflect the differences in translation difficulty between different texts, avoiding the problems of traditional translation assessment methods such as inability to perform quantitative calculations, low assessment efficiency, and high assessment cost, achieving an efficient, accurate, low-cost, and automatic assessment of text translation difficulty, which is beneficial for the assessment of the translation difficulty of large-scale texts. It has high application value in the process of constructing parallel data sets and can be used for the efficient and rapid assessment of the translation difficulty of large-scale text data sets, supporting the construction of high-quality machine translation training and evaluation parallel text data sets with a reasonable difficulty distribution.

[0115] Text translation difficulty assessment plays a very important role in scenarios such as translation exams. In order to test the translation level of examinees, it is necessary to provide texts to be translated covering different translation difficulties to more finely distinguish the translation ability levels of examinees. If the difficulty of the texts to be translated is concentrated at a certain level, then the translation exam lacks discrimination and cannot detect the true level of the examinees.

[0116] With the development of artificial intelligence technology, the level of machine translation has also been rapidly improved, especially machine translation models based on deep neural networks and large model technologies. The training and evaluation of deep translation models require the use of large-scale high-quality bilingual parallel text data. Bilingual parallel text data needs to be translated and annotated by professional translators of the corresponding languages, which results in low efficiency and high cost in constructing bilingual parallel data, affecting the scale of parallel data construction.

[0117] The evaluation of translation difficulty can effectively improve the efficiency of manual annotation of bilingual parallel data and enhance the quality of parallel text data. On the one hand, through the evaluation of translation difficulty, annotation tasks can be assigned to translators at different levels according to the translation difficulty level of the text. Translators with a high level are responsible for translating texts with a high annotation difficulty, which can improve the efficiency and quality of parallel text data annotation and reduce the annotation cost. On the other hand, the training and evaluation of translation models require data support with a balanced distribution of difficulty levels. During the training stage, if the translation difficulty of the selected training data is too low, it will affect the training efficiency and performance of the model, and the trained translation model cannot support the translation of texts with high difficulty. If the translation difficulty of the selected training data is too high, the translation model is not easy to converge during the training process, which will also lead to poor translation results of the translation model. During the model testing stage, test data with a balanced difficulty distribution is required to test the translation model in order to accurately reflect the true translation ability of the translation model.

[0118] Please refer to Figure 5 , in some embodiments, after step S160, the translation difficulty evaluation method based on the back-translated text may further include, but is not limited to, steps S510 to S540:

[0119] Step S510, perform text screening on the original text according to the translation difficulty score from the original text to the target language to obtain test texts;

[0120] Step S520, perform model testing on the preset translation model according to the test texts to obtain the translation accuracy score of the preset translation model;

[0121] Step S530, determine the target text according to the translation accuracy score and the original text;

[0122] Step S540, update the model parameters of the preset translation model based on the target text to obtain the target translation model.

[0123] In step S510 of some embodiments, in order to accurately test the translation ability of the preset translation model, text screening is performed on the original text according to the translation difficulty score to understand the difficulty distribution characteristics of the text based on the translation difficulty score, and texts with a balanced difficulty distribution are selected as test texts. Specifically, obtain the translation difficulty score from each original text in the dataset to the target language, divide the original texts with the same translation difficulty score into a cluster to obtain multiple text clusters. Count the number of texts in each text cluster to obtain the first quantity of each text cluster. Obtain the number of original texts in the dataset to obtain the second quantity. For each text cluster, calculate the ratio of the first quantity and the second quantity of the text cluster to obtain the difficulty distribution probability of the text cluster. Perform text screening on each text cluster. When the difficulty distribution probabilities of each text cluster after screening are the same, use the original texts in each text cluster at this time as test texts.

[0124] In step S520 of some embodiments, text annotation is performed on the test text to obtain the true translated text of the test text. The test text is in the source language, and the true translated text is in the target language. The test text is translated from the source language to the target language through a preset translation model to obtain a predicted translated text. The preset translation model is a deep learning model for translating text from one language to another, such as a Transformer model, a long short-term memory network, etc. Calculate the semantic similarity between the predicted translated text and the true translated text, and count the number of test texts whose semantic similarity is greater than or equal to a preset similarity threshold to obtain a third quantity. Obtain the number of test texts to obtain a fourth quantity. Calculate the ratio of the third quantity to the fourth quantity to obtain the translation accuracy score of the preset translation model. The translation accuracy score is used to measure the translation accuracy of the preset translation model. The larger the translation accuracy score, the higher the translation accuracy.

[0125] In step S530 of some embodiments, if the translation accuracy score is low, it indicates that the translation accuracy of the preset translation model is low, and the preset translation model needs to be optimized. In the embodiments of the present application, according to the translation accuracy score and the original text, the text for optimizing the preset translation model is determined to obtain the target text.

[0126] In step S540 of some embodiments, based on the target text, the model parameters of the preset translation model are updated to obtain a model with better translation performance, that is, the target translation model. The target translation model can be used to translate texts to be translated with different difficulties, improving the accuracy of text translation.

[0127] Through the above steps S510 to S540, the difficulty distribution of the text can be determined according to the translation difficulty score, so as to obtain texts with a balanced difficulty distribution for model testing and training, improving the reliability and accuracy of model testing and training, and thus improving the text translation accuracy of the translation model.

[0128] Please refer to Figure 6 , in some embodiments, step S530 may include but is not limited to steps S610 to S620:

[0129] Step S610, if the translation accuracy score is less than the first preset score threshold, add adversarial noise to the original text to obtain a noisy text, calculate the reference translation difficulty score of the noisy text to the target language, and combine the original text and the noisy text according to the reference translation difficulty score and the translation difficulty score to obtain the target text;

[0130] Step S620: If the translation accuracy score is greater than or equal to the first preset score threshold and less than the second preset score threshold, determine the difficulty level of the original text according to the translation difficulty score from the original text to the target language, and screen the original text according to the difficulty level to obtain the target text; wherein, the second preset score threshold is greater than the first preset score threshold.

[0131] In step S610 of some embodiments, the first preset score threshold is greater than 0 and less than 1. If the translation accuracy score is less than the first preset score threshold, it indicates that the translation accuracy of the preset translation model is very low, and the preset translation model needs to be retrained. To improve the translation accuracy of the translation model for texts with different translation difficulties, the training texts need to cover different translation difficulties. Add adversarial noise to the original text through a text noise addition tool, such as repeating some words in the original text, deleting some words in the original text, replacing some words with synonyms, and adjusting the order of words in the original text, to obtain the noise text. Referring to steps S110 to S160, calculate the reference translation difficulty score of the noise text. To increase the number of training samples and make the training samples cover different translation difficulties, if the reference translation difficulty score is different from the translation difficulty score, merge the original text and the noise text to obtain the target text. If the reference translation difficulty score is the same as the translation difficulty score, select the original text or the noise text as the target text.

[0132] In step S620 of some embodiments, the second preset score threshold is greater than the first preset score threshold and less than 1. If the translation accuracy score is greater than or equal to the first preset score threshold and less than the second preset score threshold, it indicates that the translation accuracy of the preset translation model is relatively low, and the preset translation model needs to be fine-tuned. The types of difficulty levels include low difficulty, medium difficulty, and high difficulty. Each type of difficulty level corresponds to a score interval, and the score intervals of different types of difficulty levels do not overlap. If the translation difficulty score from the original text to the target language is within the score interval, obtain the difficulty level corresponding to the score interval. If the translation model can accurately translate high-difficulty texts, then it can also accurately translate low-difficulty and medium-difficulty texts. To improve the efficiency of model training, therefore, select the original text with a difficulty level of high difficulty as the target text.

[0133] If the translation accuracy score is greater than or equal to the second preset score threshold, it indicates that the translation accuracy of the preset translation model is high, and there is no need to optimize the preset translation model. The preset translation model can be used as the target translation model for text translation.

[0134] The above steps S610 to S620 can improve the translation accuracy of the model for texts with different translation difficulties by optimizing the preset translation model.

[0135] Please refer toFigure 7 , in some embodiments, step S540 may include but is not limited to steps S710 to S740:

[0136] Step S710, translating the target text from the source language to the target language through a preset translation model to obtain a target translated text;

[0137] Step S720, translating the target translated text from the target language to the source language through a preset translation model to obtain a target back-translated text;

[0138] Step S730, calculating a loss based on the target text and the target back-translated text to obtain a target loss;

[0139] Step S740, updating the model parameters of the preset translation model according to the target loss to obtain a target translation model.

[0140] In step S710 of some embodiments, the target text is input into a preset translation model, and the target text is translated from the source language to the target language through the preset translation model to obtain a target translated text.

[0141] In step S720 of some embodiments, the embodiments of the present application reduce the annotation cost of the target text based on the back-translation technology. The target translated text is input into a preset translation model, and the target translated text is translated from the target language to the source language through the preset translation model to obtain a target back-translated text.

[0142] In step S730 of some embodiments, based on the mean square error loss function, the loss value between the target text and the target back-translated text is calculated to obtain a first loss. The back-translated texts obtained by multiple round-trip translations should be the same or similar. To ensure the stability of multiple round-trip translations, the target back-translated text is translated from the source language to the target language and then from the target language to the source language for a preset number of round-trips to obtain a reference back-translated text. Based on the mean square error loss function, the loss value between the target back-translated text and the reference back-translated text is calculated to obtain a second loss, and the loss value between the target text and the reference back-translated text is calculated to obtain a third loss. To ensure the semantic similarity between the target text and the target back-translated text, the target text is vectorized to obtain a first text embedding vector, the target back-translated text is vectorized to obtain a second text embedding vector, and the semantic difference degree between the first text embedding vector and the second text embedding vector is calculated. The first loss weight of the first loss, the second loss weight of the second loss, and the third loss weight of the third loss are obtained. The sum of the first loss weight, the second loss weight, and the third loss weight is one. The first loss weight is multiplied by the first loss, the second loss weight is multiplied by the second loss, the third loss weight is multiplied by the third loss, and the sum of the three multiplication results and the semantic difference degree is added to obtain the target loss.

[0143] In step S740 of some embodiments, to minimize the objective loss, the model parameters, the first loss weight, the second loss weight, and the third loss weight of the preset translation model are updated respectively to obtain the target translation model.

[0144] Through the above steps S710 to S740, the objective loss can be obtained to guide the optimization process of the preset translation model based on the objective loss, so as to obtain a model with better translation performance, thereby improving the translation accuracy of the model for texts with different translation difficulties.

[0145] Please refer to Figure 8 , the embodiment of the present application further provides a translation difficulty evaluation device based on back-translated text, which can implement the above translation difficulty evaluation method based on back-translated text. The translation difficulty evaluation device based on back-translated text includes:

[0146] An acquisition module 810, configured to acquire the original text;

[0147] A first translation module 820, configured to translate the original text from the source language into the target language through a translation channel to obtain a translated text; the translation channel has a semantic difference coefficient;

[0148] A second translation module 830, configured to translate the translated text from the target language into the source language through the translation channel to obtain a back-translated text;

[0149] A calculation module 840, configured to calculate the semantic similarity between the original text and the back-translated text;

[0150] A translation mode determination module 850, configured to determine the translation mode of the translation channel;

[0151] A translation difficulty evaluation module 860, configured to determine the translation difficulty score from the original text to the target language according to the semantic difference coefficient, the semantic similarity, and the translation mode.

[0152] The specific implementation manner of the translation difficulty evaluation device based on back-translated text is basically the same as the specific embodiments of the above translation difficulty evaluation method based on back-translated text, and will not be elaborated here.

[0153] The embodiment of the present application further provides an electronic device. The electronic device includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, it implements the above translation difficulty evaluation method based on back-translated text. The electronic device can be any intelligent terminal including a tablet computer, an in-vehicle computer, etc.

[0154] Please refer to Figure 9 , Figure 9 schematically shows the hardware structure of an electronic device in another embodiment. The electronic device includes:

[0155] The processor 910 can be implemented in the form of a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, etc., and is used to execute relevant programs to implement the technical solutions provided by the embodiments of the present application;

[0156] The memory 920 can be implemented in the form of a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM), etc. The memory 920 can store an operating system and other application programs. When implementing the technical solutions provided by the embodiments of this specification through software or firmware, the relevant program codes are stored in the memory 920 and are called by the processor 910 to execute the translation difficulty evaluation method based on the back-translated text of the embodiments of the present application;

[0157] The input / output interface 930 is used to implement information input and output;

[0158] The communication interface 940 is used to implement communication interaction between this device and other devices, and can implement communication through wired means (such as USB, network cable, etc.) or wireless means (such as mobile network, WI-FI, Bluetooth, etc.);

[0159] The bus 950 transmits information between various components of the device (such as the processor 910, the memory 920, the input / output interface 930, and the communication interface 940);

[0160] Among them, the processor 910, the memory 920, the input / output interface 930, and the communication interface 940 are communicatively connected to each other inside the device through the bus 950.

[0161] The embodiments of the present application also provide a computer-readable storage medium, which stores a computer program, and when the computer program is executed by a processor, the above-mentioned translation difficulty evaluation method based on the back-translated text is implemented.

[0162] The memory, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs and non-transitory computer-executable programs. In addition, the memory may include high-speed random access memory, and may also include non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some embodiments, the memory may optionally include memories remotely located relative to the processor, and these remote memories can be connected to the processor through a network. Examples of the above-mentioned network include but are not limited to the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0163] The embodiments described in the embodiments of the present application are for more clearly illustrating the technical solutions of the embodiments of the present application, and do not constitute a limitation on the technical solutions provided by the embodiments of the present application. Those skilled in the art will know that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of the present application are equally applicable to similar technical problems.

[0164] Those skilled in the art can understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of the present application, and may include more or fewer steps than those shown in the figures, or combine certain steps, or different steps.

[0165] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0166] Those of ordinary skill in the art can understand that all or some of the steps in the methods disclosed above, and the functional modules / units in the systems and devices can be implemented as software, firmware, hardware, and appropriate combinations thereof.

[0167] The terms "first", "second", "third", "fourth", etc. (if any) in the specification of the present application and the above-mentioned drawings are used to distinguish similar objects, and do not have to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present application described here can be implemented in an order other than those illustrated or described here. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units does not have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.

[0168] It should be understood that in this application, "at least one (item)" means one or more, and "a plurality" means two or more. "And / or" is used to describe the association relationship of associated objects, indicating that there can be three relationships. For example, "A and / or B" can mean: only A exists, only B exists, and both A and B exist at the same time. Among them, A and B can be singular or plural. The character " / " generally represents an "or" relationship between the associated objects before and after. "At least one (item) of the following" or its similar expression refers to any combination of these items, including any combination of single item (item) or plural items (items). For example, at least one (item) of a, b, or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.

[0169] In several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the above division of units is only a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. The displayed or discussed coupling or direct coupling or communication connection to each other can be through some interfaces. The indirect coupling or communication connection of devices or units can be in electrical, mechanical or other forms.

[0170] The units described above as separate components may or may not be physically separated. The components shown as units may or may not be physical units, that is, they can be located in one place, or they can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0171] In addition, each functional unit in various embodiments of this application can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit.

[0172] When an integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes multiple instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods of various embodiments of this application. The aforementioned storage medium includes: various media that can store programs, such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs.

[0173] The preferred embodiments of the embodiments of this application have been described above with reference to the accompanying drawings. This does not limit the scope of the rights of the embodiments of this application. Any modifications, equivalent replacements, and improvements made by those skilled in the art without departing from the scope and essence of the embodiments of this application shall be within the scope of the rights of the embodiments of this application.

Claims

1. A method for evaluating translation difficulty based on back-translated text, characterized in that, The method includes: Obtaining the original text; Translating the original text from the source language into the target language through a translation channel to obtain a translated text; the translation channel has a semantic difference coefficient; Translating the translated text from the target language into the source language through the translation channel to obtain a back-translated text; Calculating the semantic similarity between the original text and the back-translated text; Determining the translation mode of the translation channel; Determining a translation difficulty score for the original text to the target language according to the semantic difference coefficient, the semantic similarity, and the translation mode.

2. The method according to claim 1, wherein The determining the translation difficulty score for the original text to the target language according to the semantic difference coefficient, the semantic similarity, and the translation mode includes: Obtaining the number of channels of the translation channel; If the number of channels is equal to one, determining an initial translation difficulty for the original text to the target language according to the semantic difference coefficient, the semantic similarity, and the translation mode, and using the initial translation difficulty as the translation difficulty score; If the number of channels is greater than one, determining an initial translation difficulty for the original text to the target language for each translation channel according to the semantic difference coefficient, the corresponding semantic similarity, and the translation mode, obtaining the channel weight of each translation channel, and performing a weighted average calculation on the initial translation difficulty of the corresponding translation channel according to the channel weight of each translation channel to obtain the translation difficulty score.

3. The method according to claim 2, wherein The determining an initial translation difficulty for the original text to the target language according to the semantic difference coefficient, the semantic similarity, and the translation mode includes: If the translation mode is a human translation mode, determining a semantic difference degree according to the semantic similarity; Multiplying the semantic difference coefficient by the semantic difference degree to obtain the initial translation difficulty for the original text to the target language.

4. The method according to claim 2, wherein The determining an initial translation difficulty for the original text to the target language according to the semantic difference coefficient, the semantic similarity, and the translation mode includes: If the translation mode is a machine translation mode and the semantic similarity is greater than or equal to a preset similarity threshold, determining a semantic difference degree according to the semantic similarity, multiplying the semantic difference coefficient by the semantic difference degree, and obtaining the initial translation difficulty for the original text to the target language; If the translation mode is the machine translation mode and the semantic similarity is less than the preset similarity threshold, then use the back-translated text as the original text input to the translation channel, and repeat the steps of translating the original text from the source language to the target language through the translation channel to obtain a translated text, translating the translated text from the target language to the source language through the translation channel to obtain a back-translated text, and calculating the semantic similarity between the original text and the back-translated text until the semantic similarity is greater than or equal to the preset similarity threshold. Determine the semantic difference degree according to the target similarity, and multiply the semantic difference coefficient and the semantic difference degree to obtain the initial translation difficulty from the original text to the target language; the target similarity is the semantic similarity between the first original text and the last back-translated text.

5. The method according to any one of claims 1 to 4, characterized in that, After determining the translation difficulty score from the original text to the target language according to the semantic difference coefficient, the semantic similarity, and the translation mode, the translation difficulty evaluation method based on the back-translated text further includes: Screen the original text according to the translation difficulty score from the original text to the target language to obtain a test text; Test the preset translation model according to the test text to obtain the translation accuracy score of the preset translation model; Determine the target text according to the translation accuracy score and the original text; Update the model parameters of the preset translation model based on the target text to obtain a target translation model.

6. The method according to claim 5, wherein The determining the target text according to the translation accuracy score and the original text includes: If the translation accuracy score is less than the first preset score threshold, add adversarial noise to the original text to obtain a noisy text, calculate the reference translation difficulty score from the noisy text to the target language, and merge the original text and the noisy text according to the reference translation difficulty score and the translation difficulty score to obtain the target text; If the translation accuracy score is greater than or equal to the first preset score threshold and less than the second preset score threshold, determine the difficulty level of the original text according to the translation difficulty score from the original text to the target language, and screen the original text according to the difficulty level to obtain the target text; wherein, the second preset score threshold is greater than the first preset score threshold.

7. The method according to claim 5, characterized in that The updating the model parameters of the preset translation model based on the target text to obtain a target translation model includes: Translate the target text from the source language to the target language through the preset translation model to obtain a target translated text; Translate the target translated text from the target language to the source language through the preset translation model to obtain a target back-translated text; Calculate a target loss according to the target text and the target back-translated text; Update the model parameters of the preset translation model according to the target loss to obtain the target translation model.

8. A translation difficulty assessment device based on back-translated text, characterized in that, The device includes: An acquisition module, configured to acquire an original text; A first translation module, configured to translate the original text from a source language to a target language through a translation channel to obtain a translated text; the translation channel has a semantic difference coefficient; A second translation module, configured to translate the translated text from the target language to the source language through the translation channel to obtain a back-translated text; A calculation module, configured to calculate the semantic similarity between the original text and the back-translated text; A translation mode determination module, configured to determine the translation mode of the translation channel; A translation difficulty evaluation module, configured to determine a translation difficulty score from the original text to the target language according to the semantic difference coefficient, the semantic similarity, and the translation mode.

9. An electronic device, characterized in that, The electronic device includes a memory and a processor, the memory stores a computer program, and when the processor executes the computer program, the method according to any one of claims 1 to 7 is implemented.

10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, the method according to any one of claims 1 to 7 is implemented.

Citation Information

Cited By

  • Translation difficulty assessment method and device for text translation model

    CN121118922A

  • Text translation model translation difficulty evaluation method and device

    CN121118922B