Method and system for generating cross-lingual target text on the basis of prompt information, and medium
By receiving prompts in natural language from users and using a learning model to generate cross-language target text, the interpretability and controllability issues of neural machine translation systems are solved. This enables multi-dimensional text control in cross-language scenarios and generates target language text that better meets user needs.
Patent Information
- Application Number
- PCT/CN2024/126383
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-07-12
- Filing Date
- 2024-10-22
- Publication Date
- 2026-01-15
AI Technical Summary
Existing neural machine translation systems lack interpretability and controllability, cannot be applied to complex text control in cross-language scenarios, and require specific model structures for different control objectives, lacking compatibility and extensibility.
By receiving prompts in natural language from user input, including various types of fine-grained cross-language control information, and using a trained learning model to generate cross-language target text, it supports multi-dimensional text control.
It achieves a user-friendly cross-language text control method, generates more accurate target language text, meets users' complex multi-dimensional requirements, and improves generation efficiency and accuracy.
Smart Images

Figure CN2024126383_15012026_PF_FP_ABST
Abstract
Description
Methods, systems, and media for generating cross-language target text based on prompt information Technical Field
[0001] This application relates to the field of natural language processing technology, and in particular to a method, system, and medium for generating cross-language target text based on prompt information. Background Technology
[0002] Traditional neural machine translation (NMT) systems, relying on end-to-end training with large parallel corpora, have made significant progress in recent years, with translation results approaching human translation levels to some extent. However, these models typically exhibit black-box characteristics; that is, the trained model is a black box, only able to produce output given input. Its internal matrix-based reasoning process cannot be interfered with by humans, resulting in a lack of interpretability and controllability. This makes them unsuitable for situations requiring precise handling of specific terminology and entity translation, or for stylistic constraints on the target text.
[0003] Recent research has attempted to incorporate human control into models. For example, existing controllable generative models integrate lexical constraints into the model's decoding process or achieve this through special post-processing steps. However, style constraints on the target text rely on data synthesis or specific model design. In other words, existing controllable generative models exhibit significant model-specificity, meaning they depend on a specific model structure; different control objectives correspond to different model structures. Furthermore, existing technologies lack applications for fine-grained, complex text control in cross-linguistic scenarios.
[0004] Therefore, existing models lack compatibility with constraint types, as well as compatibility and extensibility with languages and model architectures. They cannot meet the complex intervention needs of users in cross-language scenarios, where they can conveniently use natural language to perform multi-dimensional compounding of target language generation, with a unified model architecture.
[0005] Summary of the Invention
[0006] This application addresses the aforementioned deficiencies in the prior art. There is a need for a method, system, and medium for generating cross-language target text based on prompts, which enables users to more easily submit more complex cross-language text control requirements and generates more accurate cross-language target text that better meets user control requirements for the source language text.
[0007] According to the first aspect of this application, a method for generating cross-language target text based on prompt information is provided, comprising: a processor receiving source language text input by a user as the cross-language text to be generated; the processor receiving prompt information input by the user, wherein the prompt information supports expression in natural language and contains multiple types of cross-language fine-grained control information; and the processor generating cross-language target text for display based on the received source language text and in combination with the prompt information using a trained first learning model.
[0008] According to a second aspect of this application, a system for generating cross-language target text based on prompt information is provided. The system includes an interface configured to acquire source language text of the cross-language text to be generated, input by a user, and prompt information input by the user. The prompt information supports expression in natural language and includes various types of fine-grained cross-language control information. The system further includes a processor configured to perform various operations of the method for generating cross-language target text based on prompt information according to various embodiments of this application. The system also includes a display configured to display the generated cross-language target text.
[0009] According to a third aspect of this application, a non-transitory computer-readable storage medium is provided, the program causing a processor to perform various operations of a method for generating cross-language target text based on prompt information according to various embodiments of this application.
[0010] The methods, systems, and media for generating cross-language target text based on prompt information provided in various embodiments of this application allow users to provide prompt information in natural language to control and constrain target language text, thereby providing users with a more convenient and natural way to control text. Furthermore, the prompt information is not limited to a single dimension or perspective, but can, for example, contain multiple types of cross-language fine-grained control information. Thus, when the trained first learning model combines the prompt information provided by the user to generate target language text for the source language text, it can achieve higher efficiency, make the generated target language text more accurate, and better meet the multi-dimensional and complex control requirements proposed by the user.
[0011] The above description is only an overview of the technical solution of this application. In order to better understand the technical means of this application and to implement it in accordance with the contents of the specification, and to make the above and other objects, features and advantages of this application more obvious and understandable, the following are specific embodiments of this application.
[0012] It should be understood that the foregoing general description and the following detailed description are merely illustrative and explanatory, and are not intended to limit the scope of the claimed invention. Attached Figure Description
[0013] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0014] Figure 1 shows a flowchart illustrating a method for generating cross-language target text based on prompt information according to an embodiment of this application.
[0015] Figure 2 shows a flowchart illustrating the user correction prompt information according to an embodiment of this application.
[0016] Figure 3 shows a schematic diagram of the composition structure of the first learning model according to an embodiment of this application.
[0017] Figure 4 shows a schematic diagram of a training method for a first learning model according to an embodiment of this application.
[0018] Figure 5 shows a schematic diagram of some components of a system for generating cross-language target text based on prompt information according to an embodiment of this application. Detailed Implementation
[0019] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this application. All other embodiments obtained by those skilled in the art based on the described embodiments of this application without creative effort are within the scope of protection of this application.
[0020] Unless otherwise defined, the technical or scientific terms used in this application shall have the ordinary meaning understood by one of ordinary skill in the art to which this application pertains. Words such as "comprising" or "including" mean that the element or object preceding the word covers the element or object listed after the word and its equivalents, without excluding other elements or objects.
[0021] The terms "first," "second," and similar words used in this application do not indicate any order, quantity, or importance, but are merely used for distinction. Words such as "including" or "comprising" mean that the element preceding the word encompasses the elements listed after it, and do not exclude the possibility of encompassing other elements as well. The execution order of the steps in the method described in conjunction with the accompanying drawings in this application is not intended to be limiting. As long as the logical relationship between the steps is not affected, several steps can be integrated into a single step, a single step can be decomposed into multiple steps, and the execution order of the steps can be changed according to specific needs.
[0022] It should also be understood that the term "and / or" in this application is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, or B existing alone. Furthermore, the character " / " in this application generally indicates that the preceding and following related objects have an "or" relationship.
[0023] To keep the following description of the embodiments of this application clear and concise, detailed descriptions of known functions and known components are omitted.
[0024] Figure 1 shows a flowchart illustrating a method for generating cross-language target text based on prompt information according to an embodiment of this application.
[0025] As shown in Figure 1, in step 101, the processor can receive the source language text input by the user for generating cross-language text. It should be noted that the "source language text" here can be a specific language text to be translated into other languages, or a specific language text to be rewritten into other languages with certain constraints. It can be a sentence or a paragraph, and the length of the text and the specific languages of the source language and cross-language are not specifically limited in this application.
[0026] In step 102, the processor may further receive prompt information input by the user, wherein the prompt information supports expression in natural language and includes various types of cross-language fine-grained control information.
[0027] In some embodiments, users, due to requirements for translation accuracy and personalization, or for control over cross-language text rewriting, may wish to receive prompts for the generation of cross-language text. The method according to embodiments of this application enables users to provide these prompts in their native source language, natural language. In this application, "fine-grained" is a relative concept and does not limit the specific level, specifications, or content of the prompts. Rather, it means that users can not only provide overall control requirements for cross-language text generation, but also propose specific constraints and control requirements at a finer granularity, such as vocabulary, phrases, and syntax.
[0028] Furthermore, the prompts provided by the user can include not only single-type fine-grained cross-language control requirements, but also multiple fine-grained cross-language control prompts regarding text translation or rewriting, such as sentiment, style, and syntactic structure, based on the user's needs. This is especially beneficial when the user's text control requirements are very specific, allowing for the provision of complex, multi-dimensional, and multi-faceted control information at once, thus improving the efficiency of cross-language target text generation. In other embodiments, since translation from the source language to the target language is involved, the prompts may also include some cross-language vocabulary, phrases, or sentences.
[0029] Next, in step 103, the trained first learning model can be used to generate cross-language target text for display based on the received source language text and the prompts provided by the user.
[0030] In some embodiments, the first learning model may adopt a network structure similar to an encoder-decoder, or any other applicable neural network architecture. This application does not specifically limit it, as long as it can uniformly process the source language text and the prompt information given by the user that may contain multiple requirements after being trained as described in the embodiments of this application, and generate cross-language target text that meets the text translation or rewriting requirements given by the user in the prompt information.
[0031] The method for generating cross-language target text based on prompt information according to embodiments of this application provides users with a more convenient and natural way to control target language text by allowing users to provide prompt information in their preferred natural language and in a more diverse manner. Furthermore, since the content of the prompt information is not limited to a single control dimension or perspective, but can include various types of fine-grained cross-language control information, when the trained first learning model combines the prompt information provided by the user to generate target language text for the source language text, it can generate more accurate and user-relevant target language text in a more efficient manner.
[0032] In some embodiments, users may wish to further modify or supplement the previously given prompts. Figure 2 shows a flowchart illustrating the process of a user correcting prompts according to an embodiment of this application.
[0033] As shown in Figure 2, in step 201, if the user has provided prompts in previous rounds, the prompts from each previous round and their corresponding cross-language target text can be displayed for comparison. This helps the user clearly understand the previously provided prompts, serving as a basis for modifying and supplementing them. In other embodiments, for ease of comparison, the source language text can always be displayed to the user.
[0034] Next, in step 202, the user's correction operation for the prompt information of a specified round can be received. The correction operation is based at least on the deviation between the cross-language target text of the specified round and the prompt information of that round.
[0035] Since step 201 displays the prompts from previous rounds and the cross-language target text generated accordingly, users can not only analyze potential errors in the target language text, but also easily check the degree to which the given control requirements are met, whether there are mismatches in certain fine-grained controls, etc.; or, they can evaluate whether the generated cross-language target text achieves the desired effect, etc., thereby efficiently supplementing or iteratively optimizing the prompts.
[0036] In some embodiments, the prompts for a specific round can be the last prompts provided by the user. For example, the user can first select the text control dimensions or angles that they care about more and have higher priority, and provide prompts. If the given control dimensions or angles are sufficiently satisfied, then control requirements for other dimensions or angles with lower priority are gradually provided. In other embodiments, the user can also use prompts from other rounds that are not the last round as a basis for supplementation and modification. In this way, for example, if the user has tried different prompt methods for a specific dimension in previous rounds, they can choose the most suitable prompt method in a new round and make further modifications and supplements based on that round.
[0037] Finally, in step 203, based on the source language text and the corrected prompt information, the trained first learning model can be used again to generate updated cross-language target text for display.
[0038] The embodiment shown in Figure 2 allows users to modify and supplement prompts in a progressive manner, and further improves the efficiency of users selecting prompts that best meet their control needs by providing prompts in various rounds for users to choose from.
[0039] As mentioned earlier, the control information provided by the user can encompass multiple dimensions and perspectives. In some embodiments, in addition to cross-language fine-grained control information, the prompts may also include, for example, translation prompts in terms of dimensions.
[0040] Furthermore, to enable users, especially new users of the system, to more conveniently and accurately request the prompts they need and that will be responded to, some embodiments may provide guidance for the prompts. For example, examples of typical prompts in natural language can be provided on the user interface, or all available prompt types can be listed so that users can select and add the corresponding prompt framework to their desired prompt. This allows users to easily access prompts for cross-language target text generation, including at least one type of translation prompt and / or at least one type of cross-language fine-grained control information. In other words, in embodiments according to this application, users can request translation only from the source language text to the target language text, or request cross-language fine-grained control only from the source language text to the target language text, or provide prompts from both dimensions simultaneously.
[0041] In some embodiments, translation prompts and cross-language fine-grained control can each have multiple perspectives that are allowed to coexist. Specifically, the types of translation prompt information may include, for example, target phrase classes, target short phrase classes, and target word order classes; the types of cross-language fine-grained control information may include target text length classes, target text sentiment classes, target text lexical classes, target text syntactic range classes, and target text syntactic template classes.
[0042] As mentioned earlier, users may not provide all the prompts at once. Therefore, when receiving correction operations from the user for a specified round of prompts, the correction operations also supplement at least some of the translation prompts or cross-language fine-grained control information that are not included in the specified round of prompts. In some embodiments, the user may, for example, provide new prompts for only one type of translation prompt or cross-language fine-grained control information at a time, according to the priority of the control requirements from high to low. This allows the user to more clearly understand the degree to which the generated cross-language target text matches their control requirements.
[0043] Figure 3 shows a schematic diagram of the composition structure of the first learning model according to an embodiment of this application.
[0044] As shown in Figure 3, the first learning model 300 may include a first encoding unit 301, a second encoding unit 302, and a decoding unit 303, as well as a natural language prompt template 304. Based on the received source language text and the prompt information, the trained first learning model 300 generates cross-language target text for display, which may specifically include the following steps.
[0045] First, using the natural language prompt template 304, the prompt information expressed by the user in natural language is integrated into a first continuous text.
[0046] As mentioned above, the embodiments of this application can support user input with richer and broader control dimensions and perspectives than traditional constraint types in the prior art. The prompts provided by the user can be natural language text of any form and content. Subsequently, after preprocessing such as semantic analysis, information extraction and filtering through the natural language prompt template 304, multiple prompts are matched to the corresponding placeholders in the natural language prompt template 304, thereby integrating the natural language prompts provided by the user into a first continuous text P = [p1, p2, ... p...]. t0 ], where p1, p2, ..., p t0 The user prompt information includes translation prompts or cross-language fine-grained control information from t0 angles, and the first continuous text P is in a format that can be recognized and directly applied by subsequent components in the first learning model 300.
[0047] Then, based on the first continuous text P, P is encoded using the first encoding unit 301 to obtain the prompt vector o = [o1, o2, ..., o2]. t0 In this context, o1 is the first component of the cue vector o obtained by encoding p1, o2 is the second component of the cue vector o obtained by encoding p2, and so on. Correspondingly, o t0 It is for p t0 The t0th component of the prompt vector o obtained through encoding, and so on, will not be listed here.
[0048] Furthermore, for example, it is possible to work in parallel based on the source language text X = [x1, x2, ..., x...] input by the user. t1 The second encoding unit 302 encodes X to obtain the source language text vector h = [h1, h2, ..., h...]. t1 Where t1 is the length of the source language text sequence, x1, x2, ..., x t1 Let t1 characters be the source language text X, and h1, h2, ..., h? t1 After encoding X using the second encoding unit 302, it is then compared with x1, x2, ..., x... t1 The components of the corresponding source language text vector h.
[0049] Next, the cue vector o and the source language text vector h are fed into the decoding unit 303 to generate the cross-language target text Y = [y1, y2, ..., y]. t For display purposes, where t is the length of the cross-language target text sequence, and y1, y2, ..., y tt represents the t characters in the cross-language target text sequence. It's important to note that t and t1 are not necessarily the same value, especially when the user prompt contains fine-grained cross-language control information such as target text length, in which case the length of the generated cross-language target text may differ from the source language text.
[0050] In some embodiments, the first encoding unit 301 and the second encoding unit 302 can be implemented using any flexible and adaptable neural network structure. The decoding unit 303 can, for example, use an autoregressive approach to generate cross-language target text Y that meets the control requirements in the prompt vector o based on the source language text vector h and in combination with the prompt vector o.
[0051] In some other embodiments, the first encoding unit 301 and the second encoding unit 302 may not be set as separate components, but may be implemented as a unified encoding unit. In this case, the source language text and prompt information input by the user can be input into the unified encoding unit for unified encoding. This application does not limit this.
[0052] During the training of the first learning model, in order to enable the first learning model to provide matching, high-quality target language text based on user prompts after training, the first learning model can be trained based on both source language text-cross-language target text training sentence pairs and training prompt sets.
[0053] In terms of control information, the prompt information in the training prompt set includes at least translation prompt information and cross-linguistic fine-grained control information. Specifically, the translation prompt information can be constructed, for example, by obtaining word alignment information in the sentence pairs, and the cross-linguistic fine-grained control information can be constructed by extracting linguistic features from the cross-linguistic target text.
[0054] Specifically, the types of translation prompts can include at least target phrases, target short phrases, and target word order. For example, existing statistical translation methods can be used to obtain word alignment information from the source language text to the target language text, and various types of translation prompts can be constructed based on this alignment information.
[0055] For example only, the translation hint information of the target phrase class can specify that a specific phrase in the source language is translated into a specific phrase in the target language in the source language text-cross-language target text training sentence pair. For example, it can be specified that "World Trade Organization" is translated into "世界贸易组织". Another example, the translation hint information of the target phrase class can specify specific phrase restrictions in the target language translation. For example, the translation should start with "在联合国会议上", or the translation should include the specific term "联合国会议", etc. Another example, the translation hint information of the target word order class can specify specific word order restrictions when translating source language phrases. For example, the source language phrase "the United Nations conference" should be translated after the phrase "that was held on November 12th", etc. These are not listed one by one here.
[0056] The types of cross-language fine-grained control information can at least include the target text length class, the target text sentiment class, the target text morphology class, the target text syntactic scope class, and the target text syntactic template class. The linguistic features of the target language text can be obtained through existing language parsing tools, such as Stanza, etc., and various cross-language fine-grained control hint information can be constructed.
[0057] As examples, target text length hints are designed to control the length of the generated cross-language target text. For instance, you could specify a length of 6: "Global warming is a significant issue." Similarly, target text sentiment hints can specify the sentiment of the generated cross-language target text. For example, when the source language is English, the source text "This product is very serviceable." could be rewritten with sentiment hints as either a positive version ("This product is excellent, the user experience is superb") or a negative version ("This product is useful, but has some shortcomings"). Furthermore, target text lexical hints can require the generated cross-language target text to strictly adhere to a given lexical (e.g., part-of-speech) sequence. For instance, given the source text "He ran quickly," if the given lexical hint is adjective-adverb-verb, the generated text might correspond to "He ran nimbly and quickly." For example, target text syntactic range hints can be used to impose specific syntactic structure restrictions on the generated cross-language target text. For instance, it could require that the 5th to 11th words of the generated cross-language target text form a phrasal verb. Similarly, target text syntactic template hints can be used to restrict the generated cross-language target text to follow a specific syntactic template, such as generating the target text according to the template "(ROOT(S(VP(S))))", and so on.
[0058] Figure 4 shows a schematic diagram of a training method for a first learning model according to an embodiment of this application.
[0059] When training the first learning model based on the source language text-cross-language target text training sentence pairs and the training cue set, steps 401-403 as shown in Figure 4 can be followed.
[0060] First, in step 401, it can be determined whether to provide prompt information for the current training statement pair based on the first distribution function.
[0061] In some embodiments, the first distribution function may be, for example, a Bernoulli distribution. By setting the parameters of the Bernoulli distribution, the probability of providing prompts to the training sentence pairs can be controlled as needed. This allows the trained first learning model to not only accurately generate cross-language target text that meets user needs when user prompts are available, but also to directly output accurately translated cross-language target text without sacrificing translation performance when no user prompts are available. Other distribution functions may also be selected for the first distribution function, and this application does not impose any restrictions on this.
[0062] Then, in step 402, if it is determined that prompt information needs to be provided for the current training statement pair, sampling is performed on the training prompt set based on the second distribution function to obtain a sampled prompt information set. The second distribution function can be, for example, any applicable uniform distribution function, thereby ensuring uniform sampling in the training prompt set as much as possible.
[0063] Next, in step 403, the first learning model can be trained by combining the training sentence pairs of each source language text with the cross-language target text with the sampled cue information set corresponding to the training sentence pairs.
[0064] In some embodiments, to ensure as uniform a coverage of prompts across dimensions and angles as possible, when it is determined that prompts should be provided for the current training sentence pair, the type of prompt information to be sampled is first determined from the types of translation prompts and the types of cross-language fine-grained control information based on a second distribution function. Then, in subsets of the training prompt sets corresponding to the types of prompts to be sampled, sampling is performed according to the third distribution function corresponding to each type, and the union of the prompts sampled from each subset is taken as the sampled prompt information set. The third distribution function corresponding to each type can be any applicable uniform distribution function, and they can be the same or different from each other; this application does not impose specific limitations on this.
[0065] In some embodiments, when training the first learning model using the training statement pairs of each source language text-cross-language target text and the corresponding sampled cue information set of the corresponding training statement pairs, the first encoding unit, the second encoding unit, and the decoding unit are jointly optimized, and the training objective is to minimize the negative log-likelihood of the cross-language target text generation.
[0066] Specifically, for example, the decoder can be based on P θ (y|x,P)=P θ (y|h,o) is used to model the conditional probability of the target, that is, let the target text generation state be the first t-1 generated words, h t =y 1:t-1 =[y1,y2,...,y t-1 ], then P θ (y t ) = P θ (y t |y 1:t-1 (x, P), and use methods such as gradient descent, greedy search, or bundle search to select P at step t. θ (y t The largest y t As the current output, y is generated repeatedly. tThis continues until the preset length is reached, or an output termination symbol is encountered. During training and optimization, the goal is to minimize the negative log-likelihood function L of the cross-language target text generation. θ =-∑ t log(P θ (y t Let θ be the target, where θ represents the parameter space including the first encoder, the second encoder, and the decoder, and y' ... t ′ for y t The derivative of P θ Let θ represent a probability function with θ as the parameter space.
[0067] It is worth noting that when training the first learning model according to the method shown in Figure 4, as shown in step 401, the training sentences do not always provide prompts. Therefore, during the training process, the model also learns the ability to accurately generate cross-language target text without relying on user prompts. Thus, according to the embodiments of this application, the processor can receive only the source language text of the cross-language text to be generated by the user, and the processor can generate the cross-language target text for display based solely on the received source language text without relying on prompts from the user.
[0068] According to embodiments of this application, a system for generating cross-language target text based on prompt information is also provided. Figure 5 shows a schematic diagram of some components of the system for generating cross-language target text based on prompt information according to an embodiment of this application.
[0069] As shown in Figure 5, the system 500 includes at least an interface 501, a processor 502, and a display 503. The interface 501 can be configured to acquire the source language text of the cross-language text to be generated, input by the user, and prompt information input by the user. The prompt information supports expression in natural language and includes various types of fine-grained cross-language control information.
[0070] In some embodiments, interface 501 may include a network adapter, cable connector, serial connector, USB connector, parallel connector, high-speed data transmission adapter (such as fiber optic, USB 3.0, Thunderbolt interface, etc.), wireless network adapter (such as WiFi adapter), telecommunications (3G, 4G / LTE, etc.) adapter, etc., and this application does not limit the types of adapters. System 500 can transmit the source language text of the user-inputted cross-language text to be generated and the user-inputted prompt information, etc., to other parts such as processor 502 and display 503 through interface 501. In some embodiments, interface 501 may also receive, for example, a trained first learning model from a learning network training device (not shown), etc., which are not listed here.
[0071] In some embodiments, processor 502 may be configured to perform various operations of the methods for generating cross-language target text based on prompt information as described in various embodiments of this application. In some embodiments, processor 502 may be a processing device including one or more general-purpose processing devices, such as a microprocessor, a central processing unit (CPU), a graphics processing unit (GPU), etc. More specifically, the processor may be a complex instruction set computing (CISC) microprocessor, a reduced instruction set computing (RISC) microprocessor, a very long instruction word (VLIW) microprocessor, a processor running other instruction sets, or a processor running a combination of instruction sets. The processor may also be one or more special-purpose processing devices, such as application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), digital signal processors (DSPs), system-on-a-chip (SoCs), etc.
[0072] In some embodiments, the display 503 may be configured to display at least the generated cross-language target text. In other embodiments, the display 503 may also be configured to display the source language text, and, if the user has provided prompts in previous rounds, to display the prompts from previous rounds and the corresponding cross-language target text, etc., as required, which will not be elaborated here.
[0073] In some embodiments, the display 503 may include a liquid crystal display (LCD), a light-emitting diode display (LED), a plasma display, or any other type of display, and presents a graphical user interface (GUI) on the display for user input and image / data display, etc., which will not be elaborated here.
[0074] According to embodiments of this application, a non-transitory computer-readable storage medium is also provided, on which computer-executable instructions are stored, which, when executed by a processor, implement various operations of the method for generating cross-language target text based on prompt information as described in various embodiments of this application.
[0075] In some embodiments, the aforementioned non-transitory computer-readable media may be, for example, read-only memory (ROM), random access memory (RAM), phase-change random access memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), electrically erasable programmable read-only memory (EEPROM), other types of random access memory (RAM), flash drives or other forms of flash memory, cache, registers, static memory, optical disc read-only memory (CD-ROM), digital versatile optical disc (DVD) or other optical storage, magnetic tape cassette or other magnetic storage devices, or any other possible non-transitory media used to store information or instructions that can be accessed by a computer device.
[0076] The method, system, and medium for generating cross-language target text based on prompt information according to embodiments of this application have better user-friendliness, allowing users to make text control requests in natural language, and are not limited to making all requests at once. They can also progressively iterate or supplement control requests from multiple dimensions and angles by observing and analyzing the generation status of each prompt and the corresponding target text. Furthermore, the first learning model according to embodiments of this application can automatically parse and convert natural language wording and sentence structure into a network-processable format even when there are varying numbers and types of control requests. It robustly identifies user intent and accurately completes text control. The model is applicable to any sequence-to-sequence learning network and can perform multi-task processing such as translation and text rewriting under a unified network architecture, exhibiting strong versatility, universality, and extensibility.
[0077] By comparing with existing translation models, the first learning model according to the embodiments of this application shows better performance in both target text generation quality and control success rate on fine-grained cross-lingual text control tasks. Table 1 shows the comparison results of the first learning model according to the embodiments of this application with existing open-source large models (Llama-2-7B-Chat) and commercial large models (ChatGPT) on target text generation quality and control success rate on fine-grained cross-lingual text control tasks.
[0078] Table 1
[0079] As can be seen from Table 1, the first learning model according to the embodiments of this application outperforms Llama-2-7B-Chat and ChatGPT in terms of model generation quality, control generation quality, and control success rate.
[0080] Furthermore, although exemplary embodiments have been described herein, their scope includes any and all embodiments based on this application that have equivalent elements, modifications, omissions, combinations (e.g., schemes involving intersections of various embodiments), adaptations, or alterations. Elements in the claims will be interpreted broadly based on the language used in the claims and are not limited to the examples described in this specification or during the implementation of this application, which will be interpreted as non-exclusive. Therefore, this specification and examples are intended to be considered illustrative only, and the true scope and spirit are indicated by the full scope of the claims and their equivalents.
[0081] The above description is intended to be illustrative and not restrictive. For example, the above examples (or one or more of them) can be used in combination with each other. Other embodiments can be used by those skilled in the art when reading the above description. Furthermore, in the above detailed description, various features may be grouped together to simplify the application. This should not be construed as an intention that a disclosed feature not claimed is necessary for any claim. Rather, the subject matter of the application may be less than all the features of a particular disclosed embodiment. Thus, the claims are incorporated herein by reference as examples or embodiments, wherein each claim is an independent, separate embodiment, and these embodiments are contemplated as being able to be combined with each other in various combinations or arrangements. The scope of this application should be determined by reference to the claims and the full scope of their equivalents.
[0082] The above embodiments are merely exemplary embodiments of this application and are not intended to limit this application. The scope of protection of this application is defined by the claims. Those skilled in the art can make various modifications or equivalent substitutions to this application within its substance and scope of protection, and such modifications or equivalent substitutions should also be considered to fall within the scope of protection of this application.
Claims
1. A method for generating cross-language target text based on prompt information, characterized in that, include: The processor receives the source language text from the user input to generate the cross-language text; The processor receives prompt information input by the user, wherein the prompt information supports expression in natural language and includes various types of cross-language fine-grained control information; The processor generates cross-language target text for display based on the received source language text and the prompt information, using a trained first learning model.
2. The method according to claim 1, characterized in that, The method also includes, If the user has provided prompts in previous rounds, the prompts from each previous round and their corresponding cross-language target text will be displayed in comparison. Receive correction operations from users for prompts in a specified round, the correction operations being based at least on the deviation between the cross-language target text for the specified round and the prompts for that round; Based on the source language text and the corrected prompt information, the trained first learning model is used again to generate updated cross-language target text for display.
3. The method according to claim 1 or 2, characterized in that, The prompt information also includes translation prompt information, and the method further includes: The system provides guidance to users by offering prompts that enable them to generate cross-language target text, including at least one type of translation prompts and / or at least one type of fine-grained cross-language control information; wherein, The types of translation prompts include target phrases, target short phrases, and target word order. The types of cross-language fine-grained control information include target text length class, target text sentiment class, target text lexical class, target text syntactic range class, and target text syntactic template class.
4. The method according to claim 3, characterized in that, The method also includes, When receiving a correction operation from a user for a specified round of prompts, the correction operation also supplements at least some of the translation prompts or cross-language fine-grained control information that are not included in the specified round of prompts.
5. The method according to claim 1 or 2, characterized in that, The first learning model includes a first encoder, a second encoder, and a decoder. Based on the received source language text and prompt information, it uses the trained first learning model to generate cross-language target text for display. Specifically, it includes: The prompt information is integrated into a first continuous text using a natural language prompt template; Based on the first continuous text, the first encoding unit is used for encoding to obtain a prompt vector; based on the source language text, the second encoding unit is used for encoding to obtain a source language text vector. The prompt vector and the source language text vector are fed into the decoding unit to generate cross-language target text for display.
6. The method according to claim 5, characterized in that, The first learning model is trained based on source language text-cross-language target text training sentence pairs and training prompt set, wherein the prompt information in the training prompt set includes translation prompt information and cross-language fine-grained control information; Furthermore, the translation prompt information is constructed by obtaining word alignment information in the sentence pair; the cross-linguistic fine-grained control information is constructed by extracting linguistic features of the cross-linguistic target text.
7. The method according to claim 6, characterized in that, The first learning model is trained based on source language text-cross-language target text training sentence pairs and a training cue set, specifically including: The first distribution function is used to determine whether to provide prompt information for the current training statement pair; If it is determined that prompt information needs to be provided for the current training statement pair, sampling is performed on the training prompt set based on the second distribution function to obtain the sampled prompt information set for the current training statement pair; The first learning model is trained by combining the training sentence pairs of each source language text with the cross-language target text with the sampled cue information set corresponding to the training sentence pairs.
8. The method according to claim 7, characterized in that, If it is determined that cue information needs to be provided for the current training statement pair, sampling is performed on the training cue set based on the second distribution function to obtain the sampled cue information set for the current training statement pair, specifically including: Given that prompts need to be provided for the current training sentence pair, the type of prompt to be sampled is determined based on the second distribution function, from the types of various translation prompts and the types of cross-lingual fine-grained control information. Within the subsets of the training cue sets corresponding to the types of cue information to be sampled, sampling is performed according to the third distribution function corresponding to each type, and the cue information sampled from each subset is... The union of the information is used as the set of sampled prompt information for the current training statement pair.
9. The method according to claim 7 or 8, characterized in that, The method further includes: when training the first learning model using training statements of each source language text-cross-language target text combined with the sampled cue information set, jointly optimizing the first encoding part, the second encoding part and the decoding part, and taking minimizing the negative log-likelihood of cross-language target text generation as the training objective.
10. The method according to claim 9, characterized in that, The method further includes: The processor receives only the source language text of the cross-language text to be generated from user input; the processor generates the cross-language target text for display based solely on the received source language text and without relying on prompts from the user, using a trained first learning model.
11. A system for generating cross-language target text based on prompt information, characterized in that, include: The interface is configured to: obtain the source language text of the cross-language text to be generated by the user and the prompt information input by the user, wherein the prompt information supports being expressed in natural language and contains multiple types of cross-language fine-grained control information; A processor configured to perform the method for generating cross-language target text based on prompt information as described in any one of claims 1-10; And a display configured to show the generated cross-language target text.
12. A non-transitory computer-readable storage medium, characterized in that, It stores computer-executable instructions that, when executed by a processor, implement the method for generating cross-language target text based on prompt information as described in any one of claims 1-10.
Citation Information
Patent Citations
Machine translation decoding method and device, electronic equipment and storage medium
CN115906877A
Method and system for generating cross-language abstract for long text of source language and medium
CN116187324A
Human-in-loop neural machine translation method and system and readable storage medium
CN117195922A
Method and system for generating cross-language target text based on prompt information and medium
CN118468894A
Predicting the quality of automatic translation of an entire document
US20160124944A1