Knowledge base construction method and apparatus, translation method and apparatus, and device and medium

WO2026179089A1PCT designated stage Publication Date: 2026-09-03MIDEA GRP (SHANGHAI) CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2025/115392
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-02-25
Filing Date
2025-08-18
Publication Date
2026-09-03

Smart Images

  • Figure CN2025115392_03092026_PF_FP_ABST
    Figure CN2025115392_03092026_PF_FP_ABST
Patent Text Reader

Abstract

Provided in the present application are a knowledge base construction method and apparatus, a translation method and apparatus, and a device and a medium. The knowledge base construction method comprises: acquiring a first document in a source language and a second document in a target language; invoking a large language model to extract target text element pairs in a target domain from the first document and the second document, wherein the target text element pairs comprise target text elements in the source language and target text elements in the target language that correspond to each other in terms of content; and on the basis of the target text element pairs, constructing a target knowledge base.
Need to check novelty before this filing date? Find Prior Art

Description

Knowledge base construction methods, translation methods, devices, equipment and media

[0001] This application claims priority to Chinese patent application No. 202510215160.6, filed on February 25, 2025, entitled "Knowledge Base Construction Method, Translation Method, Apparatus, Device and Medium", the contents of which are to be understood as incorporated herein by reference. Technical Field

[0002] This application relates to the field of computer technology, and in particular to a knowledge base construction method, translation method, apparatus, device, and medium. Background Technology

[0003] Machine translation refers to the process of using computer technology to convert text in one language into text in another language. However, in related machine translation technologies, performance is often poor for specific domain translation tasks, making it difficult to achieve accurate translations. Summary of the Invention

[0004] This application provides a knowledge base construction method, translation method, apparatus, device, and medium, which can improve the accuracy of machine translation by utilizing a target knowledge base.

[0005] Firstly, a method for constructing a knowledge base is provided, including:

[0006] Obtain the first document in the source language and the second document in the target language;

[0007] The large language model is invoked to extract target text element pairs of the target domain from the first and second documents. The target text element pairs include target text elements of the source language and target text elements of the target language that are content-comparative.

[0008] Construct a target knowledge base based on target text elements.

[0009] The above approach first obtains a first document in the source language and a second document in the target language. Then, a large language model is used to extract target text element pairs from the first and second documents, representing the target domain. These target text element pairs include target text elements from the source language and target text elements from the target language that are content-complementary. Finally, a target knowledge base is constructed based on these target text element pairs. This allows the target text elements in the knowledge base to provide accurate reference knowledge for the machine translation process, thereby improving translation accuracy.

[0010] In conjunction with the first aspect, in some possible implementations, the large language model is invoked to extract target text element pairs of the target domain from the first and second documents, including:

[0011] Based on the feature information of the first document, the second document, the first text element, and the target domain, construct the first prompt word;

[0012] The large language model is invoked based on the first prompt word so that the large language model can extract the first text element pair of the target domain from the first document and the second document. The first text element pair includes the first text element of the source language and the first text element of the target language that are content-comparative.

[0013] Based on the feature information of the first text element pair and the second text element, as well as the target domain, a second prompt word is constructed.

[0014] The large language model is invoked based on the second prompt word so that the large language model can extract the second text element pair of the target domain from the first text element pair. The second text element pair includes the second text element of the source language and the second text element of the target language that are content-comparative.

[0015] The first text element pair and the second text element pair are identified as the target text element pair.

[0016] Through the above scheme, the first and second text elements are different types of text elements, with the first text element having a higher linguistic structure level than the second. Considering the characteristic of the large language model outputting based on context, and the fact that the first text element has a higher linguistic structure level and is less susceptible to interference from irrelevant content in the first and second documents, the large language model is first invoked to extract the first text element pair from the first and second documents. Then, the large language model is invoked to extract the second text element pair from the first text element pair, while ensuring the accuracy of the second text element pair with a lower linguistic structure level. Finally, the first and second text element pairs are determined as the target text element pairs, improving the accuracy of the target text element pairs.

[0017] Combining the first aspect and the above implementation methods, in some possible implementations, the large language model is invoked to extract target text element pairs of the target domain from the first and second documents, including:

[0018] Based on the feature information of the first document, the second document, and the third text element, as well as the target domain, a third prompt word is constructed.

[0019] The large language model is invoked based on the third cue word so that it can extract third text element pairs of the target domain from the first and second documents. The third text element pairs include third text elements of the source language and third text elements of the target language that are content-comparative.

[0020] The third text element pair is identified as the target text element pair.

[0021] The above approach allows for the extraction of third text element pairs from the first and second documents by invoking a large language model based on the third cue word. These third text element pairs are then identified as target text element pairs. The third text element can be of any type. This effectively improves the extraction efficiency of target text element pairs.

[0022] Combining the first aspect and the above implementation methods, some possible implementation methods involve constructing a target knowledge base based on target text element pairs, including:

[0023] Write the target text element pairs into a pre-created initial structured file to obtain the first structured file;

[0024] The target knowledge base is obtained by importing data from the pre-created initial knowledge base based on the first structured file.

[0025] The above method involves writing target text element pairs into a pre-created initial structured file to obtain a first structured file. Then, data is imported into the pre-created initial knowledge base based on this first structured file to obtain the target knowledge base. Due to the structured storage characteristics of structured files, the accuracy and efficiency of data import are ensured, thereby guaranteeing the quality of the target knowledge base's content.

[0026] Combining the first aspect and the above implementation methods, in some possible implementation methods, data is imported from a pre-created initial knowledge base based on the first structured file to obtain the target knowledge base, including:

[0027] Construct a fourth prompt word based on at least one quality scoring dimension;

[0028] The large language model is invoked based on the fourth prompt word, so that the large language model can score the quality of each target text element pair in the first structured file, and obtain the quality score of each target text element pair in the first structured file.

[0029] Based on the quality scores of each target text element pair in the first structured file, the target text element pairs in the first structured file are filtered to obtain the second structured file.

[0030] The second structured file is imported into the pre-created initial knowledge base to obtain the target knowledge base.

[0031] The above method filters target text element pairs in the first structured file based on at least one quality scoring dimension to obtain a second structured file. Finally, the second structured file is imported into a pre-created initial knowledge base to obtain the target knowledge base. This effectively improves the content quality of the target knowledge base.

[0032] Combining the first aspect and the above implementation methods, in some possible implementation methods, based on the quality scores of each target text element pair in the first structured file, the target text element pairs in the first structured file are filtered to obtain a second structured file, including:

[0033] Determine the baseline value for the quality score of the first structured document;

[0034] Delete the target text elements in the first structured file whose quality scores are lower than the quality score benchmark value to obtain the second structured file.

[0035] By first determining the quality score benchmark value of the first structured file, and then using the quality score benchmark value to guide the screening process of target text element pairs in the first structured file, it is possible to effectively ensure that the second structured file after screening retains high-quality target text element pairs, thereby ensuring the content quality of the target knowledge base subsequently constructed.

[0036] Combining the first aspect and the above implementation methods, in some possible implementation methods, obtaining the first document in the source language and the second document in the target language includes:

[0037] Retrieve source documents based on the target domain;

[0038] Extract the first document in the source language and the second document in the target language from the source documents.

[0039] The above approach allows for the extraction of first and second documents from source documents in the target domain, serving as foundational materials for constructing the target knowledge base. These source documents can be explanatory documents for various products, containing text content in multiple languages. The text content in different languages ​​within the same source document is often strictly correlated, effectively improving the accuracy of translated reference knowledge in the subsequent target knowledge base.

[0040] Secondly, a translation method is provided, characterized by comprising:

[0041] Obtain the original text in the source language;

[0042] Based on the target knowledge base, the target text element pairs corresponding to the original text are determined. The target knowledge base is constructed based on the knowledge base construction method of the first aspect.

[0043] Based on the target text element pairs corresponding to the original text, a large language model is invoked to translate the original text, resulting in the target text in the target language.

[0044] The above approach utilizes a pre-built target knowledge base as a reference to determine the target text element pairs corresponding to the original text. Then, based on these target text element pairs, a large language model is invoked to translate the original text, resulting in the target text in the target language. In this way, the large language model, using the target text element pairs corresponding to the original text as a reference, can accurately perform the translation task of the original text, thereby obtaining high-quality and accurate target text.

[0045] In conjunction with the second aspect, some possible implementations involve determining the target text element pairs corresponding to the original text based on the target knowledge base, including:

[0046] Determine the vector database corresponding to the target knowledge base. The vector database includes the feature vectors of target text element pairs in the target knowledge base.

[0047] Determine the feature vector of the original text;

[0048] The target feature vector is obtained by performing vector retrieval in a vector database based on the feature vector of the original text.

[0049] Based on the target feature vector, the target text element pairs corresponding to the original text are determined.

[0050] By employing the above scheme and using vector databases and vector retrieval, the target text element pairs corresponding to the original text can be determined. These target text element pairs are semantically close to or equivalent to parts of the original text, providing an accurate reference for the subsequent translation process.

[0051] Combining the second aspect and the above implementation methods, in some possible implementation methods, the target text element pairs corresponding to the original text are determined based on the target knowledge base, including:

[0052] Determine the target text element corresponding to the original text;

[0053] Construct the target query statement based on the target text elements corresponding to the original text;

[0054] Based on the target query statement, retrieve the target text element pairs corresponding to the original text from the target knowledge base.

[0055] By using the above method and relevance query statements, the target text element pairs corresponding to the original text can be determined. These target text element pairs are semantically equivalent to parts of the original text, providing an accurate reference for the subsequent translation process.

[0056] Thirdly, a knowledge base construction apparatus is provided, the apparatus comprising:

[0057] The acquisition unit is configured to acquire a first document in the source language and a second document in the target language.

[0058] The calling unit is configured to call the large language model to extract target text element pairs of the target domain from the first document and the second document. The target text element pairs include target text elements of the source language and target text elements of the target language that are content-comparative.

[0059] The building unit is configured to build a target knowledge base based on target text element pairs.

[0060] Fourthly, a translation apparatus is provided, the apparatus comprising:

[0061] The acquisition unit is configured to acquire the raw text in the source language;

[0062] The unit is configured to determine the target text element pairs corresponding to the original text based on the target knowledge base, which is constructed based on the knowledge base construction method of the first aspect.

[0063] The translation unit is configured to translate the original text based on the target text element pairs corresponding to the original text, and then call the large language model to obtain the target text in the target language.

[0064] Fifthly, an electronic device is provided, comprising:

[0065] Memory configured to store executable program code;

[0066] The processor is configured to call and run executable program code from memory, enabling the electronic device to perform methods such as the knowledge base construction method of the first aspect or the translation method of the second aspect.

[0067] Sixthly, a computer-readable storage medium is provided, which stores a computer program that, when executed, implements the knowledge base construction method of the first aspect or the translation method of the second aspect.

[0068] Seventhly, a computer program product storing at least one instruction, which, when executed by a processor, implements the knowledge base construction method of the first aspect or the translation method of the second aspect. Attached Figure Description

[0069] To more clearly illustrate the technical solutions in the embodiments of this application or related technologies, the drawings used in the description of the embodiments or related technologies will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0070] Figure 1 is a flowchart illustrating a knowledge base construction method provided in an embodiment of this application;

[0071] Figure 2 is a schematic diagram illustrating an example of obtaining a first document and a second document according to an embodiment of this application;

[0072] Figure 3 is a flowchart illustrating a method for determining a pair of target text elements according to an embodiment of this application;

[0073] Figure 4 is a schematic diagram illustrating an example of extracting a first text element pair according to an embodiment of this application;

[0074] Figure 5 is a schematic diagram illustrating an example of extracting a second text element pair according to an embodiment of this application;

[0075] Figure 6 is a flowchart illustrating a method for determining a pair of target text elements according to an embodiment of this application;

[0076] Figure 7 is a schematic diagram of a process for constructing a target database based on structured files according to an embodiment of this application;

[0077] Figure 8 is a schematic diagram of a process for constructing a target database based on a quality score structured document according to an embodiment of this application;

[0078] Figure 9 is a flowchart illustrating a translation method provided in an embodiment of this application;

[0079] Figure 10 is a flowchart illustrating a method for determining target text element pairs corresponding to original text based on vector retrieval, according to an embodiment of this application.

[0080] Figure 11 is a flowchart illustrating a process for determining target text element pairs corresponding to original text based on a query, according to an embodiment of this application.

[0081] Figure 12 is a schematic diagram illustrating an example of knowledge base construction and translation provided in an embodiment of this application;

[0082] Figure 13 is a schematic diagram of a knowledge base construction device provided in an embodiment of this application;

[0083] Figure 14 is a schematic diagram of a translation device provided in an embodiment of this application;

[0084] Figure 15 is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation

[0085] To make the features and advantages of this application more apparent and understandable, the technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0086] In the following description, when referring to the accompanying drawings, the same numbers in different drawings denote the same or similar elements unless otherwise indicated. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this application as detailed in the appended claims.

[0087] Hereinafter, the terms "first" and "second" are used for descriptive purposes only and should not be construed as implying or suggesting relative importance or implicitly indicating the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature.

[0088] The following will provide a detailed description of each example. It should be noted that the order of description of the following embodiments is not intended to limit the preferred order of the embodiments.

[0089] Machine translation refers to the process of converting text in one language into text in another language using computer technology. However, machine translation technologies often perform poorly in domain-specific translation tasks, failing to achieve accurate translations. This may be because machine translation lacks sufficiently accurate knowledge as a reference, such as specialized terminology databases, domain-specific language rules, and expression habits. This lack of knowledge causes machine translation systems to struggle to accurately understand the meaning and context of the original text when processing domain-specific texts, thus affecting the accuracy of the translation.

[0090] To address the aforementioned issues, the main solution provided in this application includes: first, obtaining a first document in the source language and a second document in the target language; then, using a large language model to extract target text element pairs from the first and second documents, specifying the target domain. These target text element pairs include target text elements from the source language and target text elements from the target language that are content-comparative. Finally, a target knowledge base is constructed based on these target text element pairs, enabling the target text elements in the knowledge base to provide accurate reference knowledge for the machine translation process, thereby improving translation accuracy.

[0091] Based on the scenario shown in Figure 1, the knowledge base construction method and translation method provided in the embodiments of this application will be described in detail below with reference to Figures 1-12.

[0092] Please refer to Figure 1, which is a flowchart illustrating a knowledge base construction method provided in an embodiment of this application. As shown in Figure 1, the method in this embodiment may include the following steps S101-S103.

[0093] S101, Obtain the first document in the source language and the second document in the target language.

[0094] Specifically, in this embodiment, the source language refers to the language used as the original input language during the translation process, and the target language refers to the language used as the output target during the translation process. For example, when the source language is English, the target language can be Chinese; when the source language is German, the target language can be English; when the source language is Japanese, the target language can be Korean, etc., and so on.

[0095] To extract target text element pairs from the target domain, it is first necessary to obtain a first document in the source language and a second document in the target language. The first document in the source language refers to a document with text content in the source language, while the second document in the target language refers to a document with text content in the target language. It should be noted that the first and second documents are mutually exclusive in content.

[0096] Regarding the process of obtaining the first document in the source language and the second document in the target language, some possible implementations may involve receiving the first document in the source language and the second document in the target language as input; others may involve obtaining the source document of the target domain and then extracting the first document in the source language and the second document in the target language from the source document. For example, when the target domain is the home appliance domain, the source document of the target domain may be the instruction manual of some home appliances in the home appliance domain. The instruction manual of the home appliance includes text content in the source language and text content in the target language, and the text content in the source language and the text content in the target language are mutually complementary. Therefore, the corresponding first document in the source language and second document in the target language can be extracted from the instruction manual of the home appliance.

[0097] S102, invoke the large language model to extract target text element pairs of the target domain from the first document and the second document. The target text element pairs include target text elements of the source language and target text elements of the target language that are content-comparative.

[0098] Specifically, the large language model involved in this embodiment refers to a deep learning model with natural language processing capabilities, such as the GPT series models, the BERT series models, or other models based on the Transformer architecture. Large language models can understand and generate natural language text and have the ability to fine-tune and optimize for specific tasks.

[0099] The target field refers to the specific application field that the target knowledge base is directed at, such as the field of household appliances. In different target fields, the linguistic style, professional terms and expression methods of texts will be different, so it is necessary to construct a corresponding target knowledge base according to the characteristics of the target field.

[0100] It can be understood that text elements can be of multiple types, such as paragraphs, sentences, phrases, words, terms (including phrases and words), etc. These text elements are the basic units constituting documents and texts, and are also the basic objects to be processed in the translation process. The target text elements refer to the text elements that need to be extracted from the first document and the second document, and these text elements are mutually对照 in content, that is, the source-language target text elements and the target-language target text elements are consistent in semantics. For example, the target text element may be a sentence, and in this case, corresponding sentence pairs need to be extracted from the first document and the second document; the target text element may also be a word, and in this case, corresponding word pairs need to be extracted from the documents.

[0101] Target text element pairs include the source-language target text element and the target-language target text element that are mutually对照 in content, which means that in a target text element pair, the source-language target text element and the target-language target text element completely correspond or are similar in content (or semantics) and can form a one-to-one correspondence. For example, in a certain target text element pair, it includes the English word "washing machine" and its corresponding Chinese word "洗衣机".

[0102] S103, constructing a target knowledge base based on target text element pairs.

[0103] Specifically, the target knowledge base involved in this embodiment refers to a database that stores a plurality of text element pairs, and these plurality of text element pairs may involve multiple fields, including the target text element pairs of the target field.

[0104] Regarding the process of constructing a target knowledge base based on target text element pairs, in some possible implementation manners, the target text element pairs can be directly imported into an initial knowledge base in the form of structured files or other forms to obtain the target knowledge base; in some possible implementation manners, the number of target text element pairs is generally multiple, a plurality of target text element pairs can be screened, a part of the target text element pairs are deleted, and then the undeleted target text elements are imported into the initial knowledge base in the form of structured files or other forms to obtain the target knowledge base.

[0105] In this embodiment, a first document in the source language and a second document in the target language are first obtained. Then, a large language model is invoked to extract target text element pairs in the target domain from the first and second documents. These target text element pairs include target text elements from the source language and target text elements from the target language that are content-complementary. Finally, a target knowledge base is constructed based on these target text element pairs. This allows the target text elements in the knowledge base to provide accurate reference knowledge for the machine translation process, thereby improving translation accuracy.

[0106] In one embodiment, step S101 of the embodiment shown in FIG1 can be further refined and may include the following steps.

[0107] Retrieve source documents based on the target domain;

[0108] Extract the first document in the source language and the second document in the target language from the source documents.

[0109] Specifically, in some possible implementations, source documents are stored in a target database, which includes documents from multiple domains, with the target domain being one of those domains. Based on this, a corresponding query can be generated according to the target domain to retrieve the source documents corresponding to that target domain from the target database.

[0110] Furthermore, the process involves extracting a first document in the source language and a second document in the target language from the source document. In some possible implementations, the source document can be converted to Markdown format first. Markdown format supports plain text writing while also allowing the use of simple markup syntax (such as headings, lists, code blocks, etc.) to enhance readability and expressiveness. Markdown also supports various export formats, such as HTML and PDF, facilitating document transfer and management.

[0111] Next, language identification and segmentation are performed on the Markdown source document. Specifically, a language identification model from Natural Language Processing (NLP) can be used to identify the language of the Markdown source document, determining the language of each paragraph or sentence. Then, based on the identified language, the source document is segmented into multiple document fragments in different languages, with each fragment containing text content in the same language. Finally, according to the requirements of the target domain, a first document in the source language and a second document in the target language are extracted from the segmented document fragments.

[0112] In some possible implementations, please refer to Figure 2, which is an example of obtaining a first document and a second document according to an embodiment of this application. As shown in Figure 2, the target database includes various instruction documents in the home appliance field, such as air conditioner instruction documents, refrigerator instruction documents, washing machine instruction documents, etc. Any instruction document can be selected from the target database as a source document, and then a first document in the source language and a second document in the target language are extracted from the source document. It is understood that, assuming the selected source document is an air conditioner instruction document, the target knowledge base subsequently constructed can provide a reference for the translation process of air conditioner-related texts.

[0113] In some possible implementations, relevant document parsing tools can be used to obtain all text fragments of the source document and identify the language of each text fragment. Based on the content and language of each text fragment in the source document, the document's validity is determined using multiple features. If the source document is valid, the relevant document parsing tools are used to obtain each page number and the corresponding text fragment, and then the language of the text fragment corresponding to each page number is identified. Further, the language of each page number in the source document is extracted to form a language sequence. Misidentified languages ​​are corrected based on adjacent languages. Finally, the language boundary page numbers of the language sequence are obtained, and multiple documents in different languages ​​are output based on the source document, including a first document in the source language and a second document in the target language.

[0114] In this embodiment, a first document and a second document can be extracted from source documents in the target domain as foundational materials for subsequently constructing the target knowledge base. The source documents can be explanatory documents for various products and contain text content in multiple languages. The text content in different languages ​​within the same source document is often strictly correlated, which can effectively improve the accuracy of translated reference knowledge in the subsequent target knowledge base.

[0115] Please refer to Figure 3, which is a flowchart illustrating the process of determining target text element pairs according to an embodiment of this application. As shown in Figure 3, the method of this embodiment may include the following steps S201-S205, which can be used as a refinement of step S102 of the embodiment shown in Figure 1.

[0116] S201, Based on the feature information of the first document, the second document, the first text element, and the target domain, construct the first prompt word;

[0117] S202, based on the first prompt word, call the large language model so that the large language model can extract the first text element pair of the target domain from the first document and the second document. The first text element pair includes the first text element of the source language and the first text element of the target language that are mutually contrasting in content.

[0118] S203, construct the second prompt word based on the feature information of the first text element pair, the second text element, and the target domain;

[0119] S204, based on the second prompt word, call the large language model so that the large language model can extract the second text element pair of the target domain from the first text element pair. The second text element pair includes the second text element of the source language and the second text element of the target language that are content-comparative.

[0120] S205, the first text element pair and the second text element pair are determined as the target text element pair.

[0121] Specifically, this embodiment proposes that the first text element and the second text element are different types of text elements, and the first text element is higher than the second text element in terms of language structure. For example, assuming the first text element is a paragraph, then the second text element can be any one of a sentence, phrase, word, or term; assuming the first text element is a sentence, then the second text element can be any one of a phrase, word, or term. It should be noted that terminology is a generalization of phrases and words in certain fields; that is, terminology includes both phrases and words.

[0122] In this context, the language structure hierarchy refers to the hierarchical relationship of text elements in linguistics. Text elements with higher language structure levels are composed of text elements with lower language structure levels. For example, a paragraph is composed of multiple sentences, and a sentence contains multiple phrases and words. In this embodiment, the criterion for determining the hierarchy is: if a type A text element can completely contain a type B text element, and the type B text element is a constituent unit of the type A text element, then the language structure hierarchy of the type A text element is higher than that of the type B text element.

[0123] In some cases, the first text element is a sentence, and the second text element is a term.

[0124] First, based on the feature information of the first document, the second document, and the first text element, as well as the target domain, a first prompt word is constructed. The first prompt word is used to guide the large language model to extract the first text element pair of the target domain from the first document and the second document. The target domain indicates the domain to which the first text element belongs; the feature information of the first text element may include at least one of the following: input / output features of the first text element, extraction standard features, and extraction rule features.

[0125] The input and output features of the first text element indicate the input and output during the extraction process. For example, in the sentence, "You are a professional bilingual extraction assistant. Please help me extract sentences from the ${domain} domain from the documents ${src_lang1} and ${tgt_lang1} respectively, and output them in JSON format; Background: 1. The two documents are one original text and one translation, with identical content. 2. The document content includes titles, paragraph text, tables, formulas, headers, footers, tabs, etc.," in this example, "sentence" refers to the first text element, "${src_lang1}" refers to the first document, ${tgt_lang1} refers to the second document, and ${domain} refers to the target domain.

[0126] The extraction criteria for the first text element are used to indicate the quality standards to be achieved during the extraction process. For example, "Screening criteria: 1. The sentence content is paragraph text and does not contain headings, tables, formulas, headers, footers, tabs, etc. 2. The sentence content meets the requirements of typicality, representativeness, and diversity, and can reflect the style of translation."

[0127] The extraction rule features of the first text element are used to indicate the rules to be followed in the process of extracting the first text element, such as "Extraction rules: 1. Segment the original text and the translation according to the granularity of sentences; 2. Filter out sentences that meet the filtering criteria from the original text; 3. Compare the original text with the translation and find the aligned sentences from the translation", as well as the possible output format.

[0128] Then, based on the first cue word, the large language model is invoked so that, guided by the first cue word, the large language model extracts the first text element pairs of the target domain from the first and second documents. The first text element pairs include the first text elements of the source language and the first text elements of the target language that are content-comparative, meaning that the first text elements of the source language and the first text elements of the target language are semantically consistent.

[0129] Furthermore, based on the feature information of the first text element pair, the second text element, and the target domain, a second prompt word is constructed. The second prompt word is used to guide the large language model to extract the second text element pair of the target domain from the first text element pair. The target domain indicates the domain to which the second text element belongs; the feature information of the second text element may include at least one of the input / output features, extraction standard features, and extraction rule features of the second text element.

[0130] The input and output features of the second text element indicate the input and output during the extraction process. For example, "You are a professional terminology extraction assistant. Please help me extract the terms of the ${domain} domain from the texts ${src_lang2} and ${tgt_lang2} respectively, and output them in JSON format." In this example, "terms" refers to the second text element, ${src_lang2} refers to the first text element in the source language of the first text element pair, ${tgt_lang2} refers to the first text element in the target language of the first text element pair, and ${domain} refers to the target domain.

[0131] The extraction criteria feature of the second text element is used to indicate the quality standard to be achieved during the extraction of the second text element.

[0132] The extraction rule features for the second text element indicate the rules to be followed during the extraction process. For example, "Rule: 1. The term is a professional term in the field, not a frequently occurring general term. 2. If the term cannot be found in two language texts, ignore it. 3. If the term does not exist in the text, the output is []", and the possible output format.

[0133] Then, based on the second cue word, the large language model is invoked so that, guided by the second cue word, the large language model extracts the second text element pair of the target domain from the first text element pair. The second text element pair includes second text elements of the source language and second text elements of the target language that are content-complementary, meaning that the second text elements of the source language and the second text elements of the target language are semantically consistent.

[0134] Finally, the first text element pair and the second text element pair are determined as the target text element pair. That is, the type of the target text element pair can include the type of the first text element pair and the type of the second text element pair.

[0135] In some possible implementations, the pdfminer tool can be used to parse multiple lines of text content in the first and second documents separately. Then, based on the tags of each line of text content in the first and second documents, it can be determined whether it is a heading, ultimately converting it into Markdown format for the first and second documents. Further, multiple batches of first text chunks are extracted from the Markdown format first and second documents, each batch of first text chunks conforming to the input requirements of the large language model (not exceeding the maximum token length). The large language model is then invoked to extract first text element pairs of the target domain from each batch of first text chunks, forming first text element pairs in Markdown format and first text element pairs in JSON object format. Then, multiple batches of second text chunks are extracted from the Markdown format first text element pairs, each batch of second text chunks conforming to the input requirements of the large language model (not exceeding the maximum token length). The large language model is then invoked to extract second text element pairs of the target domain from each batch of second text chunks, forming second text element pairs in JSON object format. Further, the first and second text element pairs in JSON object format can be written to an Excel spreadsheet file.

[0136] To facilitate understanding of the solution in this embodiment, please refer to Figures 4 and 5. Figure 4 is a schematic diagram illustrating an example of extracting a first text element pair provided by an embodiment of this application, and Figure 5 is a schematic diagram illustrating an example of extracting a second text element pair provided by an embodiment of this application.

[0137] As shown in Figure 4, assuming the first text element is a sentence, the first prompt word can be composed of first prompt word component 1 and first prompt word component 2 in Figure 4. Calling the large language model based on the first prompt word yields a JSON object output by the large language model. This JSON object contains pairs of first text elements. Specifically, "Safety Precautions" and "Safety Precautions" constitute one pair of first text elements; "Read this manual carefully before installing or operating your new air conditioning unit. Make sure to save this manual for future reference" and "Read this manual carefully before installing or operating your new air conditioning unit. Make sure to save this manual for future reference" constitute another pair of first text elements.

[0138] As shown in Figure 5, assuming the second text element is a term (including phrases and words), the second prompt word can be composed of second prompt word component 1 and second prompt word component 2 in Figure 5. Calling the large language model based on the second prompt word yields a JSON object output by the large language model, which contains pairs of second text elements. Among them, "air conditioning equipment" and "air conditioning unit" constitute a pair of second text elements.

[0139] In this embodiment, the first text element and the second text element are different types of text elements, and the first text element has a higher linguistic structure level than the second text element. Considering the characteristic of the large language model to output based on context, and the fact that the first text element has a higher linguistic structure level and is less susceptible to interference from irrelevant content in the first and second documents, the large language model is first invoked to extract the first text element pair from the first and second documents. Then, the large language model is invoked to extract the second text element pair from the first text element pair, while ensuring the accuracy of the second text element pair with a lower linguistic structure level.

[0140] For example, when the first text element is a sentence and the second text element is a term, the sentence, as a text element with a higher level of language structure, has rich contextual information and can provide a more comprehensive contextual understanding for the large language model. Therefore, firstly, the large language model is invoked to extract sentence pairs from the first and second documents as the first text element pairs, which can provide an accurate and rich contextual foundation for the subsequent extraction of term pairs. Subsequently, based on the extracted sentence pairs, the large language model is further invoked to extract term pairs as the second text element pairs. Since the term pairs are extracted within the context of the sentence pairs, their accuracy and relevance are improved. The above-mentioned hierarchical extraction strategy not only ensures the accuracy of the target text element pairs but also effectively avoids interference from irrelevant content, thereby improving the quality of the target knowledge base and providing more reliable reference knowledge for the machine translation process.

[0141] Please refer to Figure 6, which is a flowchart illustrating the process of determining target text element pairs according to an embodiment of this application. As shown in Figure 6, the method of this embodiment may include the following steps S301-S303, which can be used as a refinement of step S102 of the embodiment shown in Figure 1.

[0142] S301, construct the third prompt word based on the feature information of the first document, the second document, the third text element, and the target domain;

[0143] S302, invoke the large language model based on the third prompt word, so that the large language model extracts the third text element pairs of the target domain from the first document and the second document. The third text element pairs include the third text elements of the source language and the third text elements of the target language that are content-comparative.

[0144] S303, determine the third text element pair as the target text element pair.

[0145] Specifically, the third text element involved in this embodiment can be any type of text element, such as paragraphs, sentences, phrases, words, terms (including phrases and words), etc. It should be noted that the third text element is not necessarily related to the first and second text elements in the above embodiments at the level of language structure. Furthermore, the third text element may or may not overlap with the first and second text elements in the above embodiments in terms of text element type; this embodiment does not impose any restrictions on this.

[0146] First, based on the feature information of the first document, the second document, and the third text element, as well as the target domain, a third cue word is constructed. This third cue word guides the large language model to extract third text element pairs from the first and second documents, specifying the target domain. The target domain indicates the domain to which the third text element belongs. The feature information of the third text element can include at least one of the following: input / output features, extraction criteria features, and extraction rule features. The input / output features indicate the input and output during the extraction process; the extraction criteria features indicate the quality standards to be achieved; and the extraction rule features indicate the rules to be followed during extraction.

[0147] Then, based on the third cue word, the large language model is invoked so that, guided by the third cue word, the large language model extracts third text element pairs of the target domain from the first and second documents. The third text element pairs include third text elements of the source language and third text elements of the target language that are content-comparative, meaning that the third text elements of the source language and the third text elements of the target language are semantically consistent.

[0148] Finally, the third text element pair is determined as the target text element pair. In other words, the type of the target text element pair can also be any type of text element pair.

[0149] In this embodiment, a large language model can be invoked based on the third cue word to extract the third text element pair of the target domain from the first and second documents. This third text element pair is then identified as the target text element pair. The third text element can be any type of text element. This effectively improves the extraction efficiency of the target text element pair.

[0150] Please refer to Figure 7, which is a flowchart of constructing a target database based on structured files according to an embodiment of this application. As shown in Figure 7, the method of this embodiment may include the following steps S401-S402. Steps S401-S402 can be used as a refinement of step S103 of the embodiment shown in Figure 1.

[0151] S401, Write the target text element pairs into the pre-created initial structured file to obtain the first structured file;

[0152] S402, Import data from the pre-created initial knowledge base based on the first structured file to obtain the target knowledge base.

[0153] Specifically, the structured file involved in this embodiment refers to a file used to store and manage structured data, such as an Excel spreadsheet file with a table structure, an Extensible Markup Language (XML) file with a tree structure, an SQLite database file with a relational database structure, etc., which will not be listed here. The pre-created initial structured file refers to a blank or structured file with a pre-defined basic framework, used to receive and organize the target text element pairs extracted from the first document and the second document.

[0154] The process of writing target text element pairs into a pre-created initial structured file to obtain the first structured file involves arranging and organizing the target text elements from the source language and the target language according to the format requirements of the initial structured file. Then, the organized target text element pairs are stored in the initial structured file to conform to its storage structure and data specifications, ultimately forming the first structured file containing the target text element pairs.

[0155] Taking an initial Excel spreadsheet file (as the initial structured file) with a tabular structure as an example, firstly, the target text elements in the source language and the target text elements in the target text element pairs are organized into two columns of data. Then, these two columns of data are written into the initial Excel file to form the first Excel file containing the target text element pairs (as the first structured file). In some possible implementations, the target text element pairs exist in the form of JSON objects. In this case, the key-value pairs of the JSON object can be mapped to the initial Excel spreadsheet file based on relevant programs or scripts to form the first Excel file containing the target text element pairs.

[0156] Taking the initial XML file with a tree structure (as the initial structured file) as an example, firstly, an XML tag structure matching the target text element pair structure is defined. Then, each target text element pair is encapsulated in the corresponding XML tag to form text data in XML format. Next, the XML formatted text data is written into the initial XML file to obtain the first XML file containing the target text element pairs (as the first structured file).

[0157] Taking the initial SQLite database file (as the initial structured file) of a relational database structure as an example, firstly, a table structure is defined containing two fields: target text element in the source language and target text element in the target language. Then, based on this table structure, each pair of target text elements is inserted into the initial SQLite database file, forming the first SQLite database file (as the first structured file) containing pairs of target text elements.

[0158] Finally, data is imported into the pre-created initial knowledge base based on the first structured file to obtain the target knowledge base. In some possible implementations, data from the first structured file can be directly imported into the initial knowledge base to obtain the target knowledge base. For example, if the initial knowledge base is built on Elasticsearch, the Elasticsearch Application Programming Interface (API) can be used to batch import data from the first structured file into the Elasticsearch index of the initial knowledge base to obtain the target knowledge base. In some possible implementations, the first structured file can be filtered to remove duplicate or low-quality target text element pairs, resulting in a second structured file. Then, data from the second structured file is imported into the initial knowledge base to obtain the target knowledge base.

[0159] In this embodiment, the target text element pairs are written into a pre-created initial structured file to obtain a first structured file. Then, data is imported into the pre-created initial knowledge base based on the first structured file to obtain the target knowledge base. Due to the structured storage characteristics of structured files, the accuracy and efficiency of data import can be ensured, thereby ensuring the content quality of the target knowledge base.

[0160] Please refer to Figure 8, which is a flowchart illustrating the process of constructing a target database based on a quality score structured document according to an embodiment of this application. As shown in Figure 8, the method of this embodiment may include the following steps S501-S504, which can be used as a refinement of step S402 of the embodiment shown in Figure 7.

[0161] S501, construct a fourth prompt word based on at least one quality scoring dimension;

[0162] S502, based on the fourth prompt word, call the large language model so that the large language model can perform quality scoring on each target text element pair in the first structured file, and obtain the quality score of each target text element pair in the first structured file;

[0163] S503, Based on the quality scores of each target text element pair in the first structured file, filter each target text element pair in the first structured file to obtain the second structured file;

[0164] S504, import the second structured file into the pre-created initial knowledge base to obtain the target knowledge base.

[0165] Specifically, the quality scoring dimension involved in this embodiment refers to the dimension used to evaluate the quality of target text element pairs. For example, the quality scoring dimension can be semantic accuracy dimension, translation fluency dimension, etc.

[0166] Regarding the process of constructing a fourth prompt word based on at least one quality score dimension, in some possible implementations, the quality score dimension is pre-recorded in the prompt word template, and the fourth prompt word can be obtained by pairing each target text element in the first structured file.

[0167] It should be noted that there must be at least one quality scoring dimension. If multiple quality scoring dimensions exist, each dimension can be listed in the fourth prompt word, and a corresponding weight can be assigned to each dimension. In this way, the large language model can comprehensively consider the weights of each quality scoring dimension when performing quality scoring, thus obtaining a more accurate and comprehensive scoring result.

[0168] Furthermore, the large language model is invoked based on the fourth prompt word. Guided by the fourth prompt word, the large language model scores the quality of each target text element pair in the first structured file, thus obtaining the quality score of each target text element pair in the first structured file. It is understandable that the quality score can be a specific numerical value, reflecting the overall performance of the target text element pair across various quality scoring dimensions. Alternatively, the quality score can also be an evaluation result in other forms, such as grades (e.g., excellent, good, average, poor) or probability distributions (representing the probability that a target text element pair belongs to different quality grades).

[0169] Furthermore, based on the quality scores of each target text element pair in the first structured file, the target text element pairs in the first structured file are filtered to obtain the second structured file. In some possible implementations, the quality scores of each target text element pair in the first structured file can be sorted in descending order, and the target text element pairs with lower quality scores in the first structured file can be deleted to obtain the second structured file. In some possible implementations, target text element pairs in the first structured file with quality scores lower than the quality score benchmark value can be deleted to obtain the second structured file.

[0170] Finally, the second structured file is imported into the pre-created initial knowledge base to obtain the target knowledge base. For example, if the initial knowledge base is built on Elasticsearch, the application programming interface provided by Elasticsearch can be used to batch import the data from the second structured file into the Elasticsearch index of the initial knowledge base to obtain the target knowledge base.

[0171] In this embodiment, the target text element pairs in the first structured file are filtered based on at least one quality scoring dimension to obtain a second structured file. Finally, the second structured file is imported into a pre-created initial knowledge base to obtain the target knowledge base. This effectively improves the content quality of the target knowledge base.

[0172] In one embodiment, step S503 of the embodiment shown in FIG8 can be further refined and may include the following steps:

[0173] Determine the baseline value for the quality score of the first structured document;

[0174] Delete the target text elements in the first structured file whose quality scores are lower than the quality score benchmark value to obtain the second structured file.

[0175] Specifically, in order to filter target text element pairs in the first structured file, it is first necessary to determine the quality score benchmark value of the first structured file. The quality score benchmark value represents the minimum acceptable standard for the target text element pair in the quality scoring dimension. Only when the quality score of the target text element pair reaches or exceeds the quality score benchmark value is the target text element pair considered to be of high quality and suitable for import into the target knowledge base.

[0176] In some possible implementations, the average quality score of the target text element pairs in the first structured file can be determined, and this average quality score can be used as the baseline quality score for the first structured file. In some possible implementations, the baseline quality score for the first structured file can also be a preset value.

[0177] Furthermore, target text element pairs in the first structured file with quality scores lower than the quality score benchmark are deleted, resulting in the second structured file. This process involves: first, iterating through each target text element pair in the first structured file and obtaining its quality score; then, comparing the quality score of each target text element pair with the quality score benchmark. If a target text element pair has a quality score lower than the benchmark, it is deleted from the first structured file. After this filtering process, the second structured file is obtained, containing the remaining target text element pairs, all of which meet the quality requirements and are suitable for importing into the target knowledge base.

[0178] In this embodiment, a quality score benchmark value for the first structured file is first determined. The quality score benchmark value guides the screening process of target text element pairs in the first structured file, which can effectively ensure that the second structured file after screening retains high-quality target text element pairs, thereby ensuring the content quality of the target knowledge base subsequently constructed.

[0179] Please refer to Figure 9, which is a flowchart illustrating a translation method provided in an embodiment of this application. As shown in Figure 9, the method in this embodiment may include the following steps S601-S603.

[0180] S601, retrieve the original text in the source language;

[0181] S602, Determine the target text element pairs corresponding to the original text based on the target knowledge base;

[0182] S603, based on the target text element pairs corresponding to the original text, calls the large language model to translate the original text and obtain the target text in the target language.

[0183] Specifically, the target knowledge base involved in this embodiment is constructed based on any one of the above-described knowledge base construction methods.

[0184] First, the raw text in the source language needs to be obtained. In some possible implementations, the raw text in the source language can be input by the user. In other possible implementations, the raw text in the source language can also be automatically obtained from other sources, such as web scraping or reading from a database.

[0185] The target knowledge base provides translation reference knowledge for the translation process. In some possible implementations, the target text element pairs corresponding to the original text can be retrieved from the target knowledge base. In other possible implementations, a vector database corresponding to the target knowledge base can be determined, and then target feature vectors can be determined from the vector database through vector retrieval. Based on the target feature vectors, the target text element pairs corresponding to the original text can be determined.

[0186] Furthermore, based on the target text element pairs corresponding to the original text, a large language model can be invoked to translate the original text, obtaining the target text in the target language. Specifically, a fifth cue word can be constructed from the target text element pairs corresponding to the original text. Guided by the fifth cue word, the large language model references the target text element pairs corresponding to the original text and translates the original text to obtain the target text in the target language. It is understandable that because the large language model references the target text element pairs corresponding to the original text during the translation process, it can fully utilize the target domain knowledge in the target knowledge base, making the final target text more accurate and fluent.

[0187] It should be noted that in some possible implementations, it is also possible to choose not to use the target knowledge base, but to directly obtain the original text of the source language, and then call the large language model to translate the original text to obtain the target text of the target language.

[0188] In this embodiment, a pre-built target knowledge base is used as a reference to determine the target text element pairs corresponding to the original text. Then, based on the target text element pairs corresponding to the original text, a large language model is invoked to translate the original text, resulting in target text in the target language. In this way, the large language model, using the target text element pairs corresponding to the original text as a reference, can accurately perform the translation task of the original text, thereby obtaining high-quality and accurate target text.

[0189] Please refer to Figure 10, which is a flowchart illustrating a method for determining the target text element pair corresponding to the original text based on vector retrieval, according to an embodiment of this application. As shown in Figure 10, the method of this embodiment may include the following steps S701-S704, which can be used as a refinement of step S602 in the embodiment shown in Figure 9.

[0190] S701, Determine the vector database corresponding to the target knowledge base. The vector database includes the feature vectors of target text element pairs in the target knowledge base.

[0191] S702, Determine the feature vector of the original text;

[0192] S703: Based on the feature vector of the original text, perform vector retrieval in the vector database to obtain the target feature vector;

[0193] S704, based on the target feature vector, determine the target text element pairs corresponding to the original text.

[0194] Specifically, the content of the vector database involved in this embodiment is mapped from the content of the target knowledge base. For example, a specific algorithm can be used to convert target text element pairs in the target knowledge base into feature vectors and store them in the vector database. Therefore, the vector database includes feature vectors of target text element pairs in the target knowledge base.

[0195] Regarding the process of determining the vector database corresponding to the target knowledge base, some possible implementations include first extracting target text element pairs from the target knowledge base; then, using a pre-trained text embedding model (such as BERT, GPT, etc.) to convert each target text element pair into a feature vector; finally, storing these feature vectors in the vector database and building corresponding indexes to support vector retrieval operations.

[0196] Furthermore, the feature vectors of the original text are determined, and vector retrieval is performed in a vector database based on these feature vectors to obtain the target feature vector. Specifically, the process involves: first, preprocessing the original text, including word segmentation and stop word removal; then, using the same text embedding model as when constructing the vector database, converting the preprocessed original text into feature vectors; next, performing vector retrieval in the vector database, calculating the similarity (e.g., cosine similarity) between the feature vectors of the original text and each feature vector in the database, and finding the feature vector most similar to the feature vectors of the original text, which is the target feature vector.

[0197] Finally, based on the target feature vector, the target text element pairs corresponding to the original text are determined. Specifically, this process involves establishing a mapping relationship between the feature vectors in the vector database and the target text element pairs in the target knowledge base. Based on the target feature vector and this mapping relationship, the target text element pairs corresponding to the original text can be determined.

[0198] It should be noted that this embodiment is applicable to some text elements with a higher level of language structure, such as paragraphs and sentences.

[0199] In this embodiment, a vector database and vector retrieval method are used to determine the target text element pairs corresponding to the original text. The target text element pairs corresponding to the original text are semantically close to or equivalent to some content of the original text, which can provide accurate reference for the subsequent translation process.

[0200] Please refer to Figure 11, which is a flowchart illustrating how to determine the target text element pair corresponding to the original text based on a query, according to an embodiment of this application. As shown in Figure 11, the method of this embodiment may include the following steps S801-S803, which can be used as a refinement of step S602 in the embodiment shown in Figure 9.

[0201] S801, Determine the target text element corresponding to the original text;

[0202] S802, construct the target query statement based on the target text elements corresponding to the original text;

[0203] S803: Based on the target query statement, retrieve the target text element pairs corresponding to the original text from the target knowledge base.

[0204] Specifically, regarding the process of determining the target text element corresponding to the original text, some possible implementations include parsing the original text based on the type of the target text element to extract the corresponding target text element. For example, if the target text element is a sentence, then natural language processing techniques can be used to segment the original text into sentences and extract sentences that meet the requirements as the target text element; if the target text element is a term, then natural language processing techniques can be used to segment the original text into words and extract terms that meet the requirements as the target text element.

[0205] Furthermore, based on the target text elements corresponding to the original text, a target query statement is constructed. This process involves: first, preprocessing the extracted target text elements, such as removing punctuation and standardizing case, to ensure the accuracy and consistency of the query statement. Then, according to the query syntax and rules of the target knowledge base, a query statement containing the target text elements is constructed. For example, if the target knowledge base is built on Elasticsearch, then an Elasticsearch-supported query statement can be constructed for efficient retrieval within the target knowledge base.

[0206] Furthermore, based on the target query statement, the target text element pairs corresponding to the original text are retrieved from the target knowledge base. Specifically, the constructed target query statement is submitted to the target knowledge base for retrieval. The target knowledge base matches and searches for text element pairs in its stored database based on the target text elements in the query statement. If a text element pair matching the target text element is found, it is returned as the target text element pair corresponding to the original text.

[0207] It should be noted that this embodiment is applicable to some text elements with a low level of language structure, such as phrases, words, or terms composed of phrases and words.

[0208] In this embodiment, by using relevant query statements, the target text element pairs corresponding to the original text can be determined. The target text element pairs corresponding to the original text are semantically equivalent to part of the content of the original text, which can provide an accurate reference for the subsequent translation process.

[0209] In one embodiment, based on the embodiments shown in Figures 1-11, please refer to Figure 12, which is a schematic diagram illustrating an example of knowledge base construction and translation provided by an embodiment of this application.

[0210] First, obtain the source documents based on the target domain; then extract the first document in the source language and the second document in the target language from the source documents.

[0211] Based on the feature information of the first document, the second document, the first text element, and the target domain, construct the first prompt word;

[0212] The large language model is invoked based on the first cue word to extract the first text element pair of the target domain from the first and second documents. The first text element pair includes the first text elements of the source language and the first text elements of the target language that are content-comparative. Based on the feature information of the first text element pair, the second text element, and the target domain, a second cue word is constructed. The large language model is then invoked based on the second cue word to extract the second text element pair of the target domain from the first text element pair. The second text element pair includes the second text elements of the source language and the second text elements of the target language that are content-comparative. The first text element pair and the second text element pair are then identified as the target text element pair.

[0213] The target text element pairs are written into a pre-created initial structured file to obtain the first structured file; a fourth prompt word is constructed based on at least one quality scoring dimension; a large language model is invoked based on the fourth prompt word to enable the large language model to perform quality scoring on each target text element pair in the first structured file, resulting in a quality score for each target text element pair in the first structured file; based on the quality scores of each target text element pair in the first structured file, each target text element pair in the first structured file is filtered to obtain the second structured file; the second structured file is imported into a pre-created initial knowledge base to obtain the target knowledge base.

[0214] At this point, the construction of the target knowledge base can be confirmed as complete.

[0215] Furthermore, obtain the original text in the source language.

[0216] The process involves: determining the vector database corresponding to the target knowledge base, which includes the feature vectors of the first text element pairs in the target knowledge base; determining the feature vectors of the original text; performing vector retrieval in the vector database based on the feature vectors of the original text to obtain the target feature vectors; and determining the first text element pairs corresponding to the original text based on the target feature vectors.

[0217] Determine the second text element corresponding to the original text; construct the target query statement based on the second text element corresponding to the original text; and retrieve the pair of second text elements corresponding to the original text from the target knowledge base according to the target query statement.

[0218] Based on the first and second text element pairs corresponding to the original text, the large language model is invoked to translate the original text and obtain the target text in the target language.

[0219] In this embodiment, the first and second text element pairs in the target domain are first determined. A target knowledge base is then constructed based on these pairs, ensuring it contains high-quality reference knowledge within the target domain. During subsequent translation, to meet the translation needs of the original text in the target domain, the corresponding first and second text element pairs can be determined based on the target knowledge base. Using these pairs, a large language model is then invoked to translate the original text, resulting in accurate target text and effectively improving translation accuracy.

[0220] The knowledge base construction apparatus provided in the embodiments of this application will now be described in detail with reference to FIG13. It should be noted that the knowledge base construction apparatus in FIG13 is configured to execute the method of the embodiments shown in FIG1-8 of this application. For ease of explanation, only the parts related to the embodiments of this application are shown; for specific technical details not disclosed, please refer to the embodiments shown in FIG1-8 of this application. Specifically, the knowledge base construction apparatus 10 may include an acquisition unit 11, a calling unit 12, and a construction unit 13, as detailed below:

[0221] Acquisition unit 11 is configured to acquire a first document in the source language and a second document in the target language;

[0222] Calling unit 12 is configured to call the large language model to extract target text element pairs of the target domain from the first document and the second document. The target text element pairs include target text elements of the source language and target text elements of the target language that are content-comparative.

[0223] Unit 13 is configured to build a target knowledge base based on target text element pairs.

[0224] Optionally, in some embodiments, the calling unit 12 may be configured to: construct a first prompt word based on the feature information of the first document, the second document, the first text element, and the target domain; call a large language model based on the first prompt word, so that the large language model extracts a first text element pair of the target domain from the first document and the second document, the first text element pair including a first text element of the source language and a first text element of the target language that are content-comparative; construct a second prompt word based on the feature information of the first text element pair, the second text element, and the target domain; call a large language model based on the second prompt word, so that the large language model extracts a second text element pair of the target domain from the first text element pair, the second text element pair including a second text element of the source language and a second text element of the target language that are content-comparative; and determine the first text element pair and the second text element pair as the target text element pair.

[0225] Optionally, in some embodiments, the calling unit 12 may be configured to: construct a third prompt word based on the feature information of the first document, the second document, the third text element, and the target domain; call a large language model based on the third prompt word so that the large language model extracts the third text element pair of the target domain from the first document and the second document, the third text element pair including the third text element of the source language and the third text element of the target language that are mutually contrasting in content; and determine the third text element pair as the target text element pair.

[0226] Optionally, in some embodiments, the construction unit 13 may be configured to: write target text element pairs into a pre-created initial structured file to obtain a first structured file; and import data into a pre-created initial knowledge base based on the first structured file to obtain a target knowledge base.

[0227] Optionally, in some embodiments, the construction unit 13 may be configured to: construct a fourth prompt word based on at least one quality scoring dimension; invoke a large language model based on the fourth prompt word so that the large language model performs quality scoring on each target text element pair in the first structured file to obtain a quality score for each target text element pair in the first structured file; filter each target text element pair in the first structured file based on the quality score of each target text element pair in the first structured file to obtain a second structured file; and import the second structured file into a pre-created initial knowledge base to obtain a target knowledge base.

[0228] Optionally, in some embodiments, the construction unit 13 may be configured to: determine a quality score benchmark value for a first structured file; delete target text element pairs in the first structured file whose quality scores are lower than the quality score benchmark value, thereby obtaining a second structured file.

[0229] Optionally, in some embodiments, the acquisition unit 11 may be configured to: acquire source documents according to the target domain; and extract a first document in the source language and a second document in the target language from the source documents.

[0230] The effects achievable in this embodiment can be found in the relevant embodiments of the knowledge base construction method described above, and will not be repeated here.

[0231] The translation apparatus provided in the embodiments of this application will now be described in detail with reference to FIG. 14. It should be noted that the translation apparatus in FIG. 14 is configured to execute the method of the embodiments shown in FIG. 9-11 of this application. For ease of explanation, only the parts related to the embodiments of this application are shown; for specific technical details not disclosed, please refer to the embodiments shown in FIG. 9-11 of this application. Specifically, the translation apparatus 20 may include an acquisition unit 21, a determination unit 22, and a translation unit 23, as detailed below:

[0232] Acquisition unit 21 is configured to acquire the original text in the source language;

[0233] Unit 22 is configured to determine the target text element pairs corresponding to the original text based on the target knowledge base, which is constructed using the aforementioned knowledge base construction method.

[0234] Translation unit 23 is configured to translate the original text based on the target text element pairs corresponding to the original text, and obtain the target text in the target language.

[0235] Optionally, in some embodiments, the determining unit 22 may be configured to: determine a vector database corresponding to the target knowledge base, the vector database including feature vectors of target text element pairs in the target knowledge base; determine the feature vector of the original text; perform vector retrieval in the vector database based on the feature vector of the original text to obtain the target feature vector; and determine the target text element pairs corresponding to the original text based on the target feature vector.

[0236] Optionally, in some embodiments, the determining unit 22 may be configured to: determine the target text element corresponding to the original text; construct a target query statement based on the target text element corresponding to the original text; and query the target text element pair corresponding to the original text in the target knowledge base according to the target query statement.

[0237] For the effects achievable in this embodiment, please refer to the relevant embodiments of the translation method described above, which will not be repeated here.

[0238] Accordingly, this application also provides an electronic device 900. Please refer to FIG15, which is a schematic diagram of the structure of an electronic device provided in this application embodiment. The electronic device 900 includes a processor 901 and a memory 902. The processor 901 and the memory 902 are electrically connected.

[0239] The processor 901 is the control center of the electronic device 900. It connects various parts of the electronic device through various interfaces and lines. By running or calling the executable program code stored in the memory 902, and calling the data stored in the memory 902, it performs various functions of the electronic device and processes data, thereby monitoring the electronic device as a whole.

[0240] The memory 902 can be configured to store executable program code and modules. The processor 901 executes various functional applications, knowledge base construction programs, and translation programs by running the executable program code and modules stored in the memory 902. The memory 902 may mainly include a program storage area and a data storage area. The program storage area may store the operating system, executable program code required for at least one function, etc.; the data storage area may store data created based on the use of the electronic device, etc.

[0241] Furthermore, memory 902 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other volatile solid-state storage device. Accordingly, memory 902 may also include a memory controller to provide processor 901 with access to memory 902.

[0242] In this embodiment, the processor 901 in the electronic device 900 loads the instructions corresponding to one or more executable program codes into the memory 902 according to the following steps, and the processor 901 runs the executable program code stored in the memory 902 to realize various functions, as follows:

[0243] Obtain the first document in the source language and the second document in the target language;

[0244] The large language model is invoked to extract target text element pairs of the target domain from the first and second documents. The target text element pairs include target text elements of the source language and target text elements of the target language that are content-comparative.

[0245] Construct a target knowledge base based on target text elements.

[0246] Optionally, when the processor 901 executes the call to the large language model to extract target text element pairs from the first and second documents, it specifically performs the following steps: constructing a first prompt word based on the feature information of the first document, the second document, the first text element, and the target domain; calling the large language model based on the first prompt word so that the large language model extracts the first text element pairs from the first and second documents, the first text element pairs including first text elements in the source language and first text elements in the target language that are content-comparative; constructing a second prompt word based on the feature information of the first text element pairs, the second text element, and the target domain; calling the large language model based on the second prompt word so that the large language model extracts the second text element pairs from the first text element pairs, the second text element pairs including second text elements in the source language and second text elements in the target language that are content-comparative; and determining the first text element pairs and the second text element pairs as target text element pairs.

[0247] Optionally, when the processor 901 executes the call to the large language model to extract target text element pairs from the first and second documents, it specifically performs the following: constructing a third prompt word based on the feature information of the first document, the second document, the third text element, and the target domain; calling the large language model based on the third prompt word so that the large language model extracts the third text element pairs from the first and second documents, the third text element pairs including the third text elements of the source language and the third text elements of the target language that are content-comparative; and determining the third text element pairs as target text element pairs.

[0248] Optionally, when the processor 901 executes the construction of the target knowledge base based on the target text element pairs, it specifically performs the following: writing the target text element pairs into a pre-created initial structured file to obtain a first structured file; and importing data into the pre-created initial knowledge base based on the first structured file to obtain the target knowledge base.

[0249] Optionally, when processor 901 performs data import on a pre-created initial knowledge base based on a first structured file to obtain a target knowledge base, it specifically performs the following: constructing a fourth prompt word based on at least one quality scoring dimension; calling a large language model based on the fourth prompt word so that the large language model performs quality scoring on each target text element pair in the first structured file to obtain a quality score for each target text element pair in the first structured file; filtering each target text element pair in the first structured file based on the quality score of each target text element pair in the first structured file to obtain a second structured file; and importing the second structured file into the pre-created initial knowledge base to obtain the target knowledge base.

[0250] Optionally, when the processor 901 performs a process to filter the target text element pairs in the first structured file based on their quality scores to obtain a second structured file, it specifically performs the following steps: determining the quality score benchmark value of the first structured file; deleting the target text element pairs in the first structured file whose quality scores are lower than the quality score benchmark value, thereby obtaining the second structured file.

[0251] Optionally, when the processor 901 executes the process of obtaining the first document in the source language and the second document in the target language, it specifically performs the following: obtaining the source document according to the target domain; and extracting the first document in the source language and the second document in the target language from the source document.

[0252] Alternatively, processor 901 can also perform:

[0253] Obtain the original text in the source language;

[0254] Based on the target knowledge base, determine the target text element pairs corresponding to the original text;

[0255] Based on the target text element pairs corresponding to the original text, a large language model is invoked to translate the original text, resulting in the target text in the target language.

[0256] Optionally, when the processor 901 executes the process of determining the target text element pair corresponding to the original text based on the target knowledge base, it specifically performs the following steps: determining the vector database corresponding to the target knowledge base, the vector database including the feature vectors of the target text element pairs in the target knowledge base; determining the feature vector of the original text; performing vector retrieval in the vector database based on the feature vector of the original text to obtain the target feature vector; and determining the target text element pair corresponding to the original text based on the target feature vector.

[0257] Optionally, when the processor 901 executes the process of determining the target text element pair corresponding to the original text based on the target knowledge base, it specifically performs the following: determining the target text element corresponding to the original text; constructing a target query statement based on the target text element corresponding to the original text; and retrieving the target text element pair corresponding to the original text from the target knowledge base according to the target query statement.

[0258] For the effects achievable in this embodiment, please refer to the relevant embodiments of the knowledge base construction method or translation method described above, which will not be repeated here.

[0259] This application also provides a computer-readable storage medium storing a computer program. When the computer program is executed, it causes the computer to perform the above-described related method steps to implement any of the knowledge base construction methods or translation methods provided in the above embodiments.

[0260] This application also provides a computer program product that stores at least one instruction, which, when executed by a processor, implements any of the knowledge base construction or translation methods provided in the above embodiments.

[0261] In this embodiment, the computer, computer-readable storage medium, and computer program product are all configured to execute the corresponding methods provided above. Therefore, the effects they can achieve can be referred to the effects in the corresponding methods provided above, and will not be repeated here.

[0262] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A method for constructing a knowledge base, wherein, include: Obtain the first document in the source language and the second document in the target language; The large language model is invoked to extract target text element pairs of the target domain from the first document and the second document. The target text element pairs include target text elements of the source language and target text elements of the target language that are content-comparative. Construct a target knowledge base based on the target text elements.

2. The method according to claim 1, wherein, The step of calling the large language model to extract target text element pairs from the first document and the second document includes: Based on the feature information of the first document, the second document, the first text element, and the target domain, a first prompt word is constructed; Based on the first prompt word, a large language model is invoked to extract a first text element pair of the target domain from the first document and the second document. The first text element pair includes a first text element of the source language and a first text element of the target language that are content-comparative. Based on the feature information of the first text element pair, the second text element, and the target domain, a second prompt word is constructed; The large language model is invoked based on the second prompt word, so that the large language model extracts the second text element pair of the target domain from the first text element pair. The second text element pair includes the second text element of the source language and the second text element of the target language, which are mutually contrasting in content. The first text element pair and the second text element pair are determined as the target text element pair.

3. The method according to claim 1, wherein, The step of calling the large language model to extract target text element pairs from the first document and the second document includes: Based on the feature information of the first document, the second document, the third text element, and the target domain, a third prompt word is constructed; The large language model is invoked based on the third prompt word, so that the large language model extracts the third text element pairs of the target domain from the first document and the second document. The third text element pairs include the third text elements of the source language and the third text elements of the target language that are content-comparative. The third text element pair is determined as the target text element pair.

4. The method according to any one of claims 1 to 3, wherein, The construction of the target knowledge base based on the target text element pairs includes: The target text element pairs are written into a pre-created initial structured file to obtain the first structured file; Based on the first structured file, data is imported into the pre-created initial knowledge base to obtain the target knowledge base.

5. The method according to claim 4, wherein, The step of importing data from a pre-created initial knowledge base based on the first structured file to obtain a target knowledge base includes: Construct a fourth prompt word based on at least one quality scoring dimension; The large language model is invoked based on the fourth prompt word, so that the large language model performs quality scoring on each target text element pair in the first structured file, thereby obtaining the quality score of each target text element pair in the first structured file; Based on the quality scores of each target text element pair in the first structured file, the target text element pairs in the first structured file are filtered to obtain a second structured file; The second structured file is imported into the pre-created initial knowledge base to obtain the target knowledge base.

6. The method according to claim 5, wherein, The step of filtering the target text element pairs in the first structured file based on their quality scores to obtain a second structured file includes: Determine the baseline value for the quality score of the first structured document; The target text element pairs in the first structured file that have a quality score lower than the quality score benchmark value are deleted to obtain the second structured file.

7. The method according to any one of claims 1 to 6, wherein, The acquisition of the first document in the source language and the second document in the target language includes: Retrieve source documents based on the target domain; Extract a first document in the source language and a second document in the target language from the source documents.

8. A translation method, wherein, include: Obtain the original text in the source language; The target text element pairs corresponding to the original text are determined based on the target knowledge base, wherein the target knowledge base is constructed based on the method described in any one of claims 1 to 7; Based on the target text element pairs corresponding to the original text, a large language model is invoked to translate the original text, resulting in the target text in the target language.

9. The method according to claim 8, wherein, The step of determining the target text element pair corresponding to the original text based on the target knowledge base includes: Determine the vector database corresponding to the target knowledge base, wherein the vector database includes feature vectors of target text element pairs in the target knowledge base; Determine the feature vector of the original text; Based on the feature vectors of the original text, vector retrieval is performed in the vector database to obtain the target feature vector; Based on the target feature vector, the target text element pairs corresponding to the original text are determined.

10. The method according to claim 8, wherein, The step of determining the target text element pair corresponding to the original text based on the target knowledge base includes: Determine the target text element corresponding to the original text; Based on the target text elements corresponding to the original text, construct the target query statement; Based on the target query statement, the target text element pair corresponding to the original text is retrieved from the target knowledge base.

11. A knowledge base construction apparatus, wherein, The device includes: The acquisition unit is configured to acquire a first document in the source language and a second document in the target language. The calling unit is configured to call the large language model to extract target text element pairs of the target domain from the first document and the second document. The target text element pairs include target text elements of the source language and target text elements of the target language that are content-comparative. The building unit is configured to build a target knowledge base based on the target text element pairs.

12. A translation device, wherein, The device includes: The acquisition unit is configured to acquire the raw text in the source language; The determining unit is configured to determine the target text element pairs corresponding to the original text based on a target knowledge base, wherein the target knowledge base is constructed based on the method described in any one of claims 1 to 7; The translation unit is configured to translate the original text by calling a large language model based on the target text element pairs corresponding to the original text, thereby obtaining the target text in the target language.

13. An electronic device, wherein, The electronic device includes: Memory configured to store executable program code; A processor configured to call and run the executable program code from the memory, causing the electronic device to perform the method as described in any one of claims 1 to 10.

14. A computer-readable storage medium, wherein, The computer-readable storage medium stores a computer program that, when executed, implements the method as described in any one of claims 1 to 10.

15. A computer program product storing at least one instruction that, when executed by a processor, implements the method of any one of claims 1 to 10.