Knowledge base construction method and device, knowledge base translation method and device, equipment and medium

By building a target knowledge base, using the target text element pairs extracted from the large language model to provide reference for machine translation, it solves the problem of poor translation in specific fields in the existing technology and achieves higher translation accuracy.

CN120069036APending Publication Date: 2025-05-30MIDEA GRP (SHANGHAI) CO LTD +1
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510215160.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-25
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

Existing machine translation techniques perform poorly in handling translation tasks in specific fields and are difficult to achieve accurate translation, mainly due to the lack of sufficiently accurate knowledge as a reference.

Method used

By obtaining documents of the source language and target language, calling large language models to extract target text element pairs of target fields, and building a target knowledge base based on these elements, thereby providing accurate reference knowledge for the machine translation process.

Benefits of technology

It improves the accuracy of machine translation, makes the translation results more accurate and smooth, and is suitable for translation tasks in specific fields.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120069036A_ABST
    Figure CN120069036A_ABST
Patent Text Reader

Abstract

The invention provides a knowledge base construction method and device, a translation method and device, equipment and a medium. The knowledge base construction method comprises the steps of obtaining a first document of a source language and a second document of a target language; calling a large language model to extract a target text element pair of a target field from the first document and the second document, wherein the target text element pair comprises a target text element of a source language and a target text element of a target language which are mutually contrasted in content; and constructing a target knowledge base based on the target text element pair. Based on the scheme, the target text elements in the target knowledge base can provide accurate reference knowledge for the machine translation process, so that the translation accuracy is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and in particular, to a method and device for constructing a knowledge base, a translation method, a device, a device, and a medium. Background Art

[0002] Machine translation refers to the process of using computer technology to convert the text of one language into the text of another language. However, in the related machine translation technology, it often performs poorly in the translation tasks for specific fields and is difficult to achieve accurate translation. Summary of the Invention

[0003] This application provides a method and device for constructing a knowledge base, a translation method, a device, a device, and a medium, which can improve the accuracy of machine translation by using the target knowledge base.

[0004] In a first aspect, a method for constructing a knowledge base is provided, including:

[0005] Obtain a first document in the source language and a second document in the target language;

[0006] Call a large language model to extract target text element pairs in the target field from the first document and the second document, where the target text element pairs include target text elements in the source language and target text elements in the target language that are in content contrast to each other;

[0007] Construct a target knowledge base based on the target text element pairs.

[0008] In a second aspect, a translation method is provided, characterized by including:

[0009] Obtain the original text in the source language;

[0010] Determine the target text element pair corresponding to the original text based on the target knowledge base, where the target knowledge base is constructed based on the knowledge base construction method in the first aspect;

[0011] Based on the target text element pair corresponding to the original text, call a large language model to translate the original text to obtain the target text in the target language.

[0012] In a third aspect, a knowledge base construction device is provided, and the device includes:

[0013] An obtaining unit for obtaining a first document in the source language and a second document in the target language;

[0014] A calling unit for calling a large language model to extract target text element pairs in the target field from the first document and the second document, where the target text element pairs include target text elements in the source language and target text elements in the target language that are in content contrast to each other;

[0015] A building unit for building a target knowledge base based on target text elements.

[0016] In a fourth aspect, a translation device is provided, which includes:

[0017] An acquisition unit for acquiring the original text in the source language;

[0018] A determination unit for determining the target text element pair corresponding to the original text based on the target knowledge base, where the target knowledge base is built based on the knowledge base building method in the first aspect;

[0019] A translation unit for calling a large language model to translate the original text based on the target text element pair corresponding to the original text to obtain the target text in the target language.

[0020] In a fifth aspect, an electronic device is provided, which includes:

[0021] A memory for storing executable program code;

[0022] A processor for calling and running the executable program code from the memory, so that the electronic device executes the knowledge base building method in the first aspect or the translation method in the second aspect.

[0023] In a sixth aspect, a computer-readable storage medium is provided, which stores a computer program. When the computer program is executed, it implements the knowledge base building method in the first aspect or the translation method in the second aspect.

[0024] In a seventh aspect, a computer program product stores at least one instruction. When the at least one instruction is executed by a processor, it implements the knowledge base building method in the first aspect or the translation method in the second aspect.

[0025] The beneficial effects brought by the technical solutions provided in some embodiments of this application at least include: First, acquire the first document in the source language and the second document in the target language, and then call a large language model to extract the target text element pairs in the target field from the first document and the second document. Among them, the target text element pair includes the target text element in the source language and the target text element in the target language that are in content contrast with each other. Finally, build a target knowledge base based on the target text element pair, so that the target text elements in the target knowledge base can provide accurate reference knowledge for the machine translation process, thereby improving the translation accuracy. Description of the Drawings

[0026] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present application. For those skilled in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0027] Figure 1 It is a schematic flowchart of a knowledge base construction method provided by an embodiment of the present application;

[0028] Figure 2 It is an example schematic diagram of obtaining a first document and a second document provided by an embodiment of the present application;

[0029] Figure 3 It is a schematic flowchart of a process for determining a target text element pair provided by an embodiment of the present application;

[0030] Figure 4 It is an example schematic diagram of extracting a first text element pair provided by an embodiment of the present application;

[0031] Figure 5 It is an example schematic diagram of extracting a second text element pair provided by an embodiment of the present application;

[0032] Figure 6 It is a schematic flowchart of a process for determining a target text element pair provided by an embodiment of the present application;

[0033] Figure 7 It is a schematic flowchart of a process for constructing a target database based on a structured file provided by an embodiment of the present application;

[0034] Figure 8 It is a schematic flowchart of a process for constructing a target database based on a structured file with quality scoring provided by an embodiment of the present application;

[0035] Figure 9 It is a schematic flowchart of a translation method provided by an embodiment of the present application;

[0036] Figure 10 It is a schematic flowchart of a process for determining a target text element pair corresponding to an original text based on vector retrieval provided by an embodiment of the present application;

[0037] Figure 11 It is a schematic flowchart of a process for determining a target text element pair corresponding to an original text based on a query provided by an embodiment of the present application;

[0038] Figure 12 It is an example schematic diagram of knowledge base construction and translation provided by an embodiment of the present application;

[0039] Figure 13It is a schematic structural diagram of a knowledge base construction device provided by an embodiment of the present application;

[0040] Figure 14 It is a schematic structural diagram of a translation device provided by an embodiment of the present application;

[0041] Figure 15 It is a schematic structural diagram of an electronic device provided by an embodiment of the present application. Detailed implementation manners

[0042] To make the features and advantages of the present application more obvious and understandable, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative efforts fall within the scope of protection of the present application.

[0043] When the following description refers to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementation manners described in the following exemplary embodiments do not represent all the implementation manners consistent with the present application. On the contrary, they are only examples of devices and methods consistent with some aspects of the present application as detailed in the appended claims.

[0044] Hereinafter, the terms "first" and "second" are only used for descriptive purposes and cannot be construed as implying or suggesting relative importance or implicitly indicating the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of such features.

[0045] The following will be described in detail respectively. It should be noted that the description order of the following embodiments does not limit the preferred order of the embodiments.

[0046] Machine translation refers to the process of using computer technology to convert the text of one language into the text of another language. However, in the related machine translation technology, it often performs poorly in the translation tasks for specific fields and is difficult to achieve accurate translation. The reason may be that machine translation lacks sufficiently accurate knowledge as a reference, such as a professional term library, domain-specific language rules and expression habits, etc. The lack of this knowledge causes the machine translation system to be difficult to accurately understand the meaning and context of the original text when processing the text of a specific field, thereby affecting the accuracy of translation.

[0047] To address the above problems, the main solutions provided by the embodiments of this application include: First, obtain the first document in the source language and the second document in the target language, and then call a large language model to extract target text element pairs in the target domain from the first document and the second document. Among them, the target text element pair includes a target text element in the source language and a target text element in the target language that are content-wise counterparts. Finally, build a target knowledge base based on the target text element pairs, so that the target text elements in the target knowledge base can provide accurate reference knowledge for the machine translation process, thereby improving the translation accuracy.

[0048] Based on Figure 1 the scenario schematic shown below, the knowledge base construction method and translation method provided by the embodiments of this application will be introduced in detail in conjunction with Figure 1 - Figure 12 .

[0049] Please refer to Figure 1 , Figure 1 which is a schematic flowchart of a knowledge base construction method provided by an embodiment of this application. As Figure 1 shown, the method of the embodiment of this application may include the following steps S101 - S103.

[0050] S101, Obtain the first document in the source language and the second document in the target language.

[0051] Specifically, the source language involved in this embodiment refers to the language used as the original input language during the translation process, and the target language refers to the language used as the output target during the translation process. For example, when the source language is English, the target language can be Chinese; when the source language is German, the target language can be English; when the source language is Japanese, the target language can be Korean, etc., which will not be listed one by one here.

[0052] To extract target text element pairs in the target domain, it is first necessary to obtain the first document in the source language and the second document in the target language. Among them, the first document in the source language refers to a document with text content in the source language, and the second document in the target language refers to a document with text content in the target language. It should be noted that the first document and the second document are content-wise counterparts.

[0053] Regarding the process of obtaining the first document in the source language and the second document in the target language, in some possible implementation manners, the first document in the source language and the second document in the target language input can be received; in some possible implementation manners, the source document in the target domain can be obtained, and then the first document in the source language and the second document in the target language can be extracted from the source document. For example, when the target domain is the household appliance domain, the source document in the target domain can be the instruction manuals of some household appliances in the household appliance domain. The instruction manuals of the household appliances include the text content in the source language and the text content in the target language, and the text content in the source language and the text content in the target language are in content contrast with each other. Therefore, the corresponding first document in the source language and the second document in the target language can be extracted from the instruction manuals of the household appliances.

[0054] S102, call a large language model to extract the target text element pairs in the target domain from the first document and the second document. The target text element pairs include the target text elements in the source language and the target text elements in the target language that are in content contrast with each other.

[0055] Specifically, the large language model involved in this embodiment refers to a deep learning model with natural language processing capabilities, such as GPT series models, BERT series models, or other models based on the Transformer architecture, etc. The large language model can understand and generate natural language text and has the ability to be fine-tuned and optimized on specific tasks.

[0056] The target domain refers to the specific application domain targeted by the target knowledge base, such as the household appliance domain. In different target domains, the language style, professional terms, and expression methods of the text will be different. Therefore, it is necessary to construct the corresponding target knowledge base according to the characteristics of the target domain.

[0057] It can be understood that the text elements can be of various types, such as paragraphs, sentences, phrases, words, terms (including phrases and words), etc. These text elements are the basic units that make up the document and text and are also the basic objects to be processed in the translation process. The target text elements refer to the text elements that need to be extracted from the first document and the second document. These text elements are in content contrast with each other, that is, the target text elements in the source language and the target text elements in the target language are semantically consistent. For example, the target text element can be a sentence, and in this case, the corresponding sentence pairs need to be extracted from the first document and the second document; the target text element can also be a word, and in this case, the corresponding word pairs need to be extracted from the document.

[0058] The target text element pair includes the target text element in the source language and the target text element in the target language that are in contrast to each other in content, which means that in the target text element pair, the target text element in the source language and the target text element in the target language are completely corresponding or similar in content (or semantics) and can form a one-to-one correspondence. For example, in a certain target text element pair, it includes the English word "washing machine" and its corresponding Chinese word "washing machine".

[0059] S103. Build a target knowledge base based on the target text element pair.

[0060] Specifically, the target knowledge base involved in this embodiment refers to a database that stores multiple text element pairs. These multiple text element pairs may cover multiple fields, including the target text element pairs in the target field.

[0061] Regarding the process of building a target knowledge base based on the target text element pair, in some possible implementation manners, the target text element pair can be directly imported into the initial knowledge base in a structured file or other form to obtain the target knowledge base; in some possible implementation manners, the number of target text element pairs is generally multiple. Multiple target text element pairs can be screened, and a part of the target text element pairs can be deleted, and then the undeleted target text elements can be imported into the initial knowledge base in a structured file or other form to obtain the target knowledge base.

[0062] In this embodiment, first obtain the first document in the source language and the second document in the target language, and then call the large language model to extract the target text element pairs in the target field from the first document and the second document. Among them, the target text element pair includes the target text element in the source language and the target text element in the target language that are in contrast to each other in content. Finally, build a target knowledge base based on the target text element pair, so that the target text elements in the target knowledge base can provide accurate reference knowledge for the machine translation process, thereby improving the translation accuracy.

[0063] In one embodiment, for Figure 1 The steps of the illustrated embodiment S101 are further refined and may include the following steps.

[0064] Obtain the source document according to the target field;

[0065] Extract the first document in the source language and the second document in the target language from the source document.

[0066] Specifically, in some possible implementation manners, the source document is stored in the target database. The target database includes documents in multiple fields, and the target field is one of the multiple fields. Based on this, a corresponding query statement can be generated according to the target field, and the source document corresponding to the target field can be retrieved from the target database.

[0067] Furthermore, the first document in the source language and the second document in the target language are extracted from the source document. In some possible implementation manners, the source document can be first converted into the Markdown format. The Markdown format features support for writing in plain text format, and at the same time, simple markup syntax (such as headings, lists, code blocks, etc.) can be used to enhance the readability and expressiveness of the document. The Markdown format also supports multiple export formats, such as HTML, PDF, etc., which is convenient for document transmission and management.

[0068] Then, language identification and language segmentation are performed on the source document in the Markdown format. Specifically, a language identification model in natural language processing (NLP) technology can be used to perform language identification on the source document in the Markdown format to determine the language of each paragraph or sentence in the source document. Then, according to the identified languages, the source document is segmented into multiple document fragments in different languages, where each document fragment contains text content in the same language. Finally, according to the requirements of the target field, the first document in the source language and the second document in the target language are extracted from the segmented document fragments.

[0069] In some possible implementation manners, please refer to Figure 2 for an example schematic diagram of obtaining the first document and the second document provided by the embodiments of this application. As Figure 2 shown, the target database includes various instruction documents in the home appliance field, such as air conditioner instruction documents, refrigerator instruction documents, washing machine instruction documents, etc. Any one of the instruction documents can be selected from the target database as the source document, and then the first document in the source language and the second document in the target language are extracted from the source document. It can be understood that assuming the selected source document is an air conditioner instruction document, then the subsequent constructed target knowledge base can provide a reference for the translation process of air conditioner-related texts.

[0070] In some possible implementation manners, relevant document parsing tools can be used to obtain all text fragments of the source document and identify the languages of each text fragment. According to the content and language of each text fragment of the source document, it is determined whether the document is qualified based on multiple features. If the source document is qualified, relevant document parsing tools are used to obtain each page number of the source document and the text fragment corresponding to each page number, and then the languages of the text fragments corresponding to each page number are identified. Further, the languages of each page of the source document are extracted to form a language sequence, the mis-identified languages are corrected according to adjacent languages, and finally the page numbers at the language boundaries of the language sequence are obtained, and multiple documents in different languages are output according to the source document, including the first document in the source language and the second document in the target language.

[0071] In this embodiment, the first document and the second document can be extracted from the source document in the target field as the basic materials for subsequent construction of the target knowledge base. Among them, the source document can be the instruction documents of various products and has text contents in multiple languages. The text contents in different languages in the same source document are often strictly in contrast, which can effectively improve the accuracy of the translation reference knowledge in the subsequent target knowledge base.

[0072] Please refer to Figure 3 , which provides a schematic flowchart of a process for determining target text element pairs according to an embodiment of the present application. As Figure 3 shown, the method of the embodiment of the present application may include the following steps S201-S205, and steps S201-S205 can be used as the refinement steps of step S102 of the embodiment shown in Figure 1 .

[0073] S201, construct a first prompt word based on the first document, the second document, the feature information of the first text element, and the target field;

[0074] S202, call a large language model based on the first prompt word, so that the large language model extracts the first text element pair in the target field from the first document and the second document. The first text element pair includes the first text element in the source language and the first text element in the target language that are in contrast in content;

[0075] S203, construct a second prompt word based on the first text element pair, the feature information of the second text element, and the target field;

[0076] S204, call a large language model based on the second prompt word, so that the large language model extracts the second text element pair in the target field from the first text element pair. The second text element pair includes the second text element in the source language and the second text element in the target language that are in contrast in content;

[0077] S205, determine the first text element pair and the second text element pair as the target text element pair.

[0078] Specifically, in this embodiment, it is proposed that the first text element and the second text element are different types of text elements, and the first text element is higher than the second text element in terms of the language structure level. For example, assuming the first text element is a paragraph, then the second text element can be any one of a sentence, a phrase, a word, or a term; assuming the first text element is a sentence, then the second text element can be any one of a phrase, a word, or a term. It should be noted that a term is a generalization of phrases and words in certain fields, that is, a term includes phrases and words.

[0079] Among them, the language structure level refers to the hierarchical composition relationship of text elements in linguistics. Text elements with a higher language structure level are composed of text elements with a lower language structure level. For example, a paragraph is composed of multiple sentences, and a sentence contains multiple phrases and words. The criterion for determining the level of high and low in this embodiment is: if type A text elements can completely contain type B text elements and type B text elements are the constituent units of type A text elements, then the language structure level of type A text elements is higher than that of type B text elements.

[0080] In some cases, the first text element is a sentence and the second text element is a term.

[0081] First, based on the first document, the second document, the feature information of the first text element, and the target domain, a first prompt is constructed. The first prompt is used to guide the large language model to extract pairs of first text elements in the target domain from the first document and the second document. Among them, the target domain is used to indicate the domain to which the first text element belongs; the feature information of the first text element may include at least one of the input-output feature, the extraction standard feature, and the extraction rule feature of the first text element.

[0082] The input-output feature of the first text element is used to indicate the input and output in the process of extracting the first text element. For example, "You are a professional bilingual extraction assistant. Please help me extract sentences in the ${domain} domain from the ${src_lang1} and ${tgt_lang1} documents respectively and output them in JSON format; Background: 1. The two documents are an original text and a translation respectively, with the same content. 2. The document content includes titles, paragraph texts, tables, formulas, headers, footers, tab characters, etc.". In this example, "sentence" refers to the first text element, "${src_lang1}" refers to the first document, ${tgt_lang1} refers to the second document, and ${domain} refers to the target domain.

[0083] The extraction standard features of the first text element are used to indicate the quality standards to be achieved in the process of extracting the first text element. For example, "Screening criteria: 1. The sentence content is paragraph text and does not contain content such as titles, tables, formulas, headers, footers, and tabs. 2. The sentence content meets the requirements of typicality, representativeness, and diversity and can reflect the translation style."

[0084] The extraction rule features of the first text element are used to indicate the rules to be followed in the process of extracting the first text element. For example, "Extraction rules: 1. Split the source text and the translation into sentence-level granularity; 2. Screen out the sentences that meet the screening criteria from the source text; 3. Compare the source text with the translation and find the aligned sentences in the translation", and the possible output formats that may be set.

[0085] Then, based on the first prompt, the large language model is called to enable the large language model to extract the first text element pairs in the target domain from the first document and the second document under the guidance of the first prompt. The first text element pairs include the first text element in the source language and the first text element in the target language that are in content contrast to each other, that is, the first text element in the source language and the first text element in the target language are semantically consistent.

[0086] Furthermore, based on the first text element pairs, the feature information of the second text element, and the target domain, a second prompt is constructed. The second prompt is used to guide the large language model to extract the second text element pairs in the target domain from the first text element pairs. Among them, the target domain is used to indicate the domain to which the second text element belongs; the feature information of the second text element may include at least one of the input-output features, extraction standard features, and extraction rule features of the second text element.

[0087] The input-output features of the second text element are used to indicate the input and output in the process of extracting the second text element. For example, "You are a professional term extraction assistant. Please help me extract the terms in the ${domain} domain from the ${src_lang2} text and the ${tgt_lang2} text and output them in JSON format." In this example, "terms" refers to the second text element, ${src_lang2} refers to the first text element in the source language of the first text element pair, ${tgt_lang2} refers to the first text element in the target language of the first text element pair, and ${domain} refers to the target domain.

[0088] The extraction standard features of the second text element are used to indicate the quality standards to be achieved in the process of extracting the second text element.

[0089] The extraction rule feature of the second text element is used to indicate the rules to be followed in the process of extracting the second text element. For example, "Rules: 1. The term is a professional term in this field and not a commonly used word that appears frequently. 2. If no corresponding texts in two languages can be found for the term, it is ignored. 3. If the term does not exist in the text, the output is []", and the possible output formats that may be set.

[0090] Then, based on the second prompt, the large language model is called so that the large language model extracts the second text element pair in the target field under the guidance of the second prompt. The second text element pair includes the second text element in the source language and the second text element in the target language that are in contrast to each other in content, that is, the second text element in the source language and the second text element in the target language are semantically consistent.

[0091] Finally, the first text element pair and the second text element pair are determined as the target text element pair. That is to say, the types of the target text element pair can include the types of the first text element pair and the second text element pair.

[0092] In some possible implementation manners, the pdfminer tool can be used to parse the multi-line text contents in the first document and the second document respectively, and then determine whether it is a title according to the labels of the text contents of each line in the first document and the second document, and finally convert them into the first document and the second document in Markdown format. Further, multiple batches of first text chunks are extracted from the first document and the second document in Markdown format, and each batch of first text chunks meets the input requirements of the large language model (not greater than the maximum token length). The large language model is called to extract the first text element pair in the target field from each batch of first text chunks, and form the first text element pair in Markdown format and the first text element pair in JSON object format. Then, multiple batches of second text chunks are extracted from the first text element pair in Markdown format, and each batch of second text chunks meets the input requirements of the large language model (not greater than the maximum token length). The large language model is called to extract the second text element pair in the target field from each batch of second text chunks, and form the second text element pair in JSON object format. Further, the first text element pair and the second text element pair in JSON object format can be written into an Excel table file.

[0093] For ease of understanding the solution of this embodiment, please refer to Figure 4 and Figure 5 , Figure 4 is an example schematic diagram of extracting the first text element pair provided by the embodiment of the present application, Figure 5This is an example schematic diagram of extracting a second text element pair provided in an embodiment of the present application.

[0094] like Figure 4 As shown in FIG. 1 , assuming that the first text element is a sentence, the first prompt word can be Figure 4 The first prompt word component 1 and the first prompt word component 2 in the text are constituted together. Based on the first prompt word, the large language model is called to obtain a JSON object output by the large language model, and the JSON object contains a first text element pair. Among them, "Safety Precautions" and "Safety Precautions" constitute a first text element pair; "Read this manual carefully before installing or operating your new air conditioning unit. Make sure to save this manual for future reference" and "Read this manual carefully before installing or operating your new air conditioning unit. Make sure to save this manual for future reference" constitute a first text element pair.

[0095] like Figure 5 As shown in FIG. 1 , assuming that the second text element is a term (including phrases and words), the second prompt word can be Figure 5 The second prompt word component 1 and the second prompt word component 2 in the text are constituted together, and the large language model is called based on the second prompt word to obtain a JSON object output by the large language model, and the JSON object contains a second text element pair. Among them, "air conditioning equipment" and "air conditioning unit" constitute a second text element pair.

[0096] In this embodiment, the first text element and the second text element are different types of text elements, and the first text element is higher in language structure level than the second text element. Considering the characteristics of the large language model that outputs according to the context, and the language structure level of the first text element is higher and is not easily disturbed by irrelevant content in the first document and the second document, the large language model is first called to extract the first text element pair from the first document and the second document. Then, the large language model is called to extract the second text element pair from the first text element pair, and the accuracy of the second text element pair with a lower language structure level can be ensured at the same time.

[0097] Exemplarily, when the first text element is a sentence and the second text element is a term, the sentence, as a text element with a higher language structure level, has rich context information, which can provide a more comprehensive context understanding for the large language model. Therefore, first, call the large language model to extract sentence pairs from the first document and the second document as the first text element pairs, which can provide an accurate and rich context basis for subsequent extraction of term pairs. Subsequently, based on the extracted sentence pairs, further call the large language model to extract term pairs as the second text element pairs. Since the term pairs are extracted in the context of the sentence pairs, their accuracy and relevance are improved. The above hierarchical extraction strategy not only ensures the accuracy of the target text element pairs but also effectively avoids the interference of irrelevant content, thereby improving the quality of the target knowledge base and providing more reliable reference knowledge for the machine translation process.

[0098] Please refer to Figure 6 , which provides a schematic flowchart of a process for determining target text element pairs according to an embodiment of the present application. As Figure 6 shown, the method of the embodiment of the present application may include the following steps S301 - S303, and steps S301 - S303 may be used as refinement steps for Figure 1 step S102 of the embodiment shown.

[0099] S301, construct a third prompt word based on the feature information of the first document, the second document, the third text element, and the target domain;

[0100] S302, call the large language model based on the third prompt word, so that the large language model extracts third text element pairs of the target domain from the first document and the second document. The third text element pairs include the third text element in the source language and the third text element in the target language that are in contrast to each other in content;

[0101] S303, determine the third text element pairs as the target text element pairs.

[0102] Specifically, the third text element involved in this embodiment may be any type of text element, such as a paragraph, a sentence, a phrase, a word, a term (including a phrase and a word), etc. It should be noted that there is no necessary association between the third text element and the first text element and the second text element in the above embodiment in terms of the language structure level. In addition, the third text element may or may not coincide with the first text element and the second text element in the above embodiment in terms of the text element type, and this embodiment does not limit this.

[0103] First, based on the feature information of the first document, the second document, the third text element, and the target domain, construct a third prompt, which is used to guide the large language model to extract pairs of third text elements in the target domain from the first document and the second document. Here, the target domain is used to indicate the domain to which the third text element belongs; the feature information of the third text element may include at least one of the input-output feature, the extraction standard feature, and the extraction rule feature of the third text element. The input-output feature of the third text element is used to indicate the input and output in the process of extracting the third text element; the extraction standard feature of the third text element is used to indicate the quality standard to be achieved in the process of extracting the third text element; the extraction rule feature of the third text element is used to indicate the rules to be followed in the process of extracting the third text element.

[0104] Then, call the large language model based on the third prompt, so that the large language model extracts pairs of third text elements in the target domain from the first document and the second document under the guidance of the third prompt. The pair of third text elements includes the third text element in the source language and the third text element in the target language that are in contrast to each other in content, that is, the third text element in the source language and the third text element in the target language are semantically consistent.

[0105] Finally, determine the pair of third text elements as the pair of target text elements. That is to say, the type of the pair of target text elements can also be the type of any pair of text elements.

[0106] In this embodiment, the large language model can be called based on the third prompt to extract pairs of third text elements in the target domain from the first document and the second document, and then the pair of third text elements is determined as the pair of target text elements, where the third text element can be any type of text element. In this way, the extraction efficiency of the pair of target text elements can be effectively improved.

[0107] Please refer to Figure 7 , which shows a schematic flowchart of a process for constructing a target database based on a structured file provided by an embodiment of the present application. As Figure 7 shown, the method of the embodiment of the present application may include the following steps S401-S402, and steps S401-S402 may be used as refinement steps of step S103 in the embodiment shown in Figure 1 .

[0108] S401, write the pair of target text elements into a pre-created initial structured file to obtain a first structured file;

[0109] S402, perform data import on a pre-created initial knowledge base based on the first structured file to obtain a target knowledge base.

[0110] Specifically, the structured file involved in this embodiment refers to a file used to store and manage structured data, such as an Excel spreadsheet file with a tabular structure, an Extensible Markup Language (XML) file with a tree structure, a SQLite database file with a relational database structure, etc., which will not be listed one by one here. The initially created structured file refers to a blank or pre-set basic framework structured file for receiving and organizing the target text element pairs extracted from the first document and the second document.

[0111] The process of writing the target text element pairs into the initially created structured file to obtain the first structured file is as follows: Arrange and organize the target text elements in the source language and the target text elements in the target language in the target text element pairs according to the format requirements of the initial structured file. Then, store the organized target text element pairs into the initial structured file to conform to the storage structure and data specifications of the initial structured file, and finally form the first structured file containing the target text element pairs.

[0112] Taking the initially created Excel spreadsheet file (as the initial structured file) with a tabular structure as an example, first organize the target text elements in the source language and the target text elements in the target language in the target text element pairs into two columns of data respectively, and then write these two columns of data into the initial Excel file to form the first Excel file (as the first structured file) containing the target text element pairs. In some possible implementation manners, if the target text element pairs exist in the form of JSON objects, then the key-value pairs of the JSON object can be mapped to the initially created Excel spreadsheet file based on relevant programs or scripts to form the first Excel file containing the target text element pairs.

[0113] Taking the initially created XML file (as the initial structured file) with a tree structure as an example, first define an XML tag structure that matches the structure of the target text element pairs, and then encapsulate each target text element pair in the corresponding XML tag to form text data in XML format. Then write the text data in XML format into the initial XML file to obtain the first XML file (as the first structured file) containing the target text element pairs.

[0114] Taking the initially created SQLite database file (as the initial structured file) with a relational database structure as an example, first define a table structure that includes two fields, namely the target text element in the source language and the target text element in the target language. Then insert each target text element pair into the initially created SQLite database file based on this table structure to form the first SQLite database file (as the first structured file) containing the target text element pairs.

[0115] Finally, data is imported into the pre-created initial knowledge base based on the first structured file to obtain a target knowledge base. In some possible implementation manners, the data in the first structured file can be directly imported into the initial knowledge base to obtain the target knowledge base. For example, if the initial knowledge base is built based on Elasticsearch, then the Application Programming Interface (API) provided by Elasticsearch can be used to batch import the data in the first structured file into the Elasticsearch index of the initial knowledge base to obtain the target knowledge base. In some possible implementation manners, the first structured file can be screened to remove duplicate or low-quality target text element pairs, and then a second structured file is obtained. Then, the data in the second structured file is imported into the initial knowledge base to obtain the target knowledge base.

[0116] In this embodiment, the target text element pairs are written into the pre-created initial structured file to obtain a first structured file, and then data is imported into the pre-created initial knowledge base based on the first structured file to obtain a target knowledge base. Due to the structured storage characteristics of the structured file, the accuracy and efficiency of data import can be ensured, and thus the content quality of the target knowledge base can be ensured.

[0117] Please refer to Figure 8 , which is a schematic flowchart of a process for constructing a target database based on a structured file with quality scores provided by an embodiment of the present application. As Figure 8 shown, the method of the embodiment of the present application may include the following steps S501-S504, and steps S501-S504 may be used as refinement steps of step S402 in the embodiment shown in Figure 7 .

[0118] S501, construct a fourth prompt word based on at least one quality score dimension;

[0119] S502, call a large language model based on the fourth prompt word so that the large language model performs quality scoring on each target text element pair in the first structured file to obtain the quality scores of each target text element pair in the first structured file;

[0120] S503, screen each target text element pair in the first structured file based on the quality scores of each target text element pair in the first structured file to obtain a second structured file;

[0121] S504, import the second structured file into the pre-created initial knowledge base to obtain a target knowledge base.

[0122] Specifically, the quality scoring dimension involved in this embodiment refers to the dimension used to evaluate the quality of the target text element pairs. For example, the quality scoring dimension can be the semantic accuracy dimension, the translation fluency dimension, etc.

[0123] Regarding the process of constructing the fourth prompt based on at least one quality scoring dimension, in some possible implementation manners, the quality scoring dimensions are pre-recorded in the prompt template, and each target text element pair in the first structured file is used to obtain the fourth prompt.

[0124] It should be noted that there is at least one quality scoring dimension. Assuming there are multiple quality scoring dimensions, then each quality scoring dimension can be listed separately in the fourth prompt, and a corresponding weight is assigned to each quality scoring dimension. In this way, when the large language model performs quality scoring, it can comprehensively consider according to the weights of each quality scoring dimension, so as to obtain a more accurate and comprehensive scoring result.

[0125] Furthermore, the large language model is called based on the fourth prompt. Under the guidance of the fourth prompt, the large language model performs quality scoring on each target text element pair in the first structured file, and then obtains the quality scores of each target text element pair in the first structured file. It can be understood that the quality score can be a specific numerical value, which reflects the comprehensive performance of the target text element pair on each quality scoring dimension. In addition, the quality score can also be an evaluation result in other forms, such as a grade (such as excellent, good, general, poor, etc.) or a probability distribution (indicating the probability that the target text element pair belongs to different quality levels).

[0126] Furthermore, based on the quality scores of each target text element pair in the first structured file, each target text element pair in the first structured file is screened to obtain a second structured file. In some possible implementation manners, the quality scores of each target text element pair in the first structured file can be sorted in descending order, and some target text element pairs with lower quality score rankings in the first structured file are deleted to obtain the second structured file. In some possible implementation manners, the target text element pairs with quality scores lower than the quality score benchmark value in the first structured file can be deleted to obtain the second structured file.

[0127] Finally, the second structured file is imported into the pre-created initial knowledge base to obtain the target knowledge base. For example, if the initial knowledge base is built based on Elasticsearch, then the application programming interface provided by Elasticsearch can be used to batch import the data in the second structured file into the Elasticsearch index of the initial knowledge base to obtain the target knowledge base.

[0128] In this embodiment, based on at least one quality scoring dimension, target text element pairs in the first structured file are screened to obtain a second structured file. Finally, the second structured file is imported into a pre-created initial knowledge base to obtain a target knowledge base. In this way, the content quality of the target knowledge base can be effectively improved.

[0129] In one embodiment, for Figure 8 the steps of the illustrated embodiment, step S503 can be further refined and may include the following steps:

[0130] Determine the quality score benchmark value of the first structured file;

[0131] Delete the target text element pairs in the first structured file whose quality scores are lower than the quality score benchmark value to obtain a second structured file.

[0132] Specifically, in order to screen the target text element pairs in the first structured file, it is first necessary to determine the quality score benchmark value of the first structured file. Among them, the quality score benchmark value represents the lowest acceptable standard of the target text element pair in the quality scoring dimension. Only when the quality score of the target text element pair reaches or exceeds the quality score benchmark value, is the target text element pair considered to be of high quality and suitable for being imported into the target knowledge base.

[0133] In some possible implementation manners, the average quality score of the target text element pairs in the first structured file can be determined, and the average quality score of the target text element pairs in the first structured file is determined as the quality score benchmark value of the first structured file. In some possible implementation manners, the quality score benchmark value of the first structured file can also be a preset value.

[0134] Furthermore, delete the target text element pairs in the first structured file whose quality scores are lower than the quality score benchmark value to obtain a second structured file. This process is specifically manifested as follows: First, traverse each target text element pair in the first structured file to obtain its quality score. Then, compare the quality score of each target text element pair with the quality score benchmark value. If the quality score of a certain target text element pair is lower than the quality score benchmark value, delete it from the first structured file. After such a screening process, a second structured file will be obtained. The second structured file contains the remaining target text element pairs, and these target text element pairs all meet the quality requirements and are suitable for being imported into the target knowledge base.

[0135] In this embodiment, the quality score benchmark value of the first structured file is first determined, and the screening process of the target text element pairs in the first structured file is guided by the quality score benchmark value, which can effectively ensure that the second structured file after screening retains high-quality target text element pairs, and further ensure the content quality of the target knowledge base constructed subsequently.

[0136] Please refer to Figure 9 , which is a schematic flowchart of a translation method provided by an embodiment of the present application. As Figure 9 shown, the method of the embodiment of the present application may include the following steps S601 - S603.

[0137] S601, obtain the original text in the source language;

[0138] S602, determine the target text element pairs corresponding to the original text based on the target knowledge base;

[0139] S603, based on the target text element pairs corresponding to the original text, call a large language model to translate the original text to obtain the target text in the target language.

[0140] Specifically, the target knowledge base involved in this embodiment is constructed based on any one of the above-mentioned knowledge base construction methods.

[0141] First, the original text in the source language needs to be obtained. In some possible implementation manners, the original text in the source language can be input by the user. In some possible implementation manners, the original text in the source language can also be automatically obtained from other sources, such as web crawling, database reading, etc.

[0142] The target knowledge base provides translation reference knowledge for the translation process. In some possible implementation manners, the target text element pairs corresponding to the original text can be queried from the target knowledge base. In some possible implementation manners, the vector database corresponding to the target knowledge base can be determined, and then the target feature vector can be determined from the vector database through vector retrieval, and the target text element pairs corresponding to the original text are determined based on the target feature vector.

[0143] Furthermore, based on the target text element pairs corresponding to the original text, a large language model can be called to translate the original text to obtain the target text in the target language. Specifically, a fifth prompt word can be constructed from the target text element pairs corresponding to the original text, and the large language model refers to the target text element pairs corresponding to the original text under the guidance of the fifth prompt word and translates the original text to obtain the target text in the target language. It can be understood that since the large language model refers to the target text element pairs corresponding to the original text during the translation process, the knowledge in the target domain in the target knowledge base can be fully utilized to make the finally obtained target text more accurate and fluent.

[0144] It should be noted that, in some possible implementation manners, it is also possible to choose not to use the target knowledge base, directly obtain the original text in the source language, and call a large language model to translate the original text based on the original text to obtain the target text in the target language.

[0145] In this embodiment, by using the pre-constructed target knowledge base as a reference, the target text element pairs corresponding to the original text are determined based on the target knowledge base, and then the large language model is called to translate the original text based on the target text element pairs corresponding to the original text to obtain the target text in the target language. In this way, the large language model uses the target text element pairs corresponding to the original text as a reference, can accurately perform the translation task of the original text, and then obtain a high-quality and accurate target text.

[0146] Please refer to Figure 10 , which provides a schematic flowchart of a process for determining target text element pairs corresponding to the original text based on vector retrieval according to an embodiment of the present application. As Figure 10 shown, the method of the embodiment of the present application may include the following steps S701-S704, and steps S701-S704 may be used as the refinement steps of step S602 in the embodiment Figure 9 shown.

[0147] S701, determine the vector database corresponding to the target knowledge base, where the vector database includes the feature vectors of the target text element pairs in the target knowledge base;

[0148] S702, determine the feature vector of the original text;

[0149] S703, perform vector retrieval in the vector database based on the feature vector of the original text to obtain the target feature vector;

[0150] S704, determine the target text element pairs corresponding to the original text based on the target feature vector.

[0151] Specifically, the content of the vector database involved in this embodiment is mapped from the content of the target knowledge base. For example, the target text element pairs in the target knowledge base can be converted into feature vectors through a specific algorithm and stored in the vector database. Therefore, the vector database includes the feature vectors of the target text element pairs in the target knowledge base.

[0152] Regarding the process of determining the vector database corresponding to the target knowledge base, in some possible implementation manners, the target text element pairs can be first extracted from the target knowledge base; then, a pre-trained text embedding model (such as BERT, GPT, etc.) is used to convert each target text element pair into a feature vector; finally, these feature vectors are stored in the vector database and corresponding indexes are constructed to support vector retrieval operations.

[0153] Further, determine the feature vector of the original text, and perform vector retrieval in the vector database based on the feature vector of the original text to obtain the target feature vector. The specific process is as follows: First, preprocess the original text, including steps such as word segmentation and stop word removal; then, use the same text embedding model as when constructing the vector database to convert the preprocessed original text into a feature vector; next, perform vector retrieval in the vector database, and find the feature vector that is most similar to the feature vector of the original text by calculating the similarity (such as cosine similarity) between the feature vector of the original text and each feature vector in the vector database, which is the target feature vector.

[0154] Finally, based on the target feature vector, determine the target text element pair corresponding to the original text. The specific process is as follows: There is a mapping relationship between the feature vectors in the vector database and the target text element pairs in the target knowledge base. Based on the target feature vector and this mapping relationship, the target text element pair corresponding to the original text can be determined.

[0155] It should be noted that this embodiment is applicable to some text elements with a relatively high language structure level, such as paragraphs and sentences.

[0156] In this embodiment, by using the vector database and vector retrieval methods, the target text element pair corresponding to the original text can be determined, where the target text element pair corresponding to the original text is semantically close to or equivalent to some content of the original text, and can provide an accurate reference for the subsequent translation process.

[0157] Please refer to Figure 11 , which provides a schematic flow diagram of a process for determining the target text element pair corresponding to the original text based on a query in an embodiment of the present application. As Figure 11 shown, the method of the embodiment of the present application may include the following steps S801 - S803, and steps S801 - S803 may be used as refinement steps for Figure 9 step S602 of the embodiment shown.

[0158] S801, determine the target text element corresponding to the original text;

[0159] S802, construct a target query statement based on the target text element corresponding to the original text;

[0160] S803, query in the target knowledge base according to the target query statement to obtain the target text element pair corresponding to the original text.

[0161] Specifically, regarding the process of determining the target text elements corresponding to the original text, in some possible implementation manners, the original text can be parsed according to the types of the target text elements to extract the target text elements corresponding to the original text. For example, if the target text elements are sentences, the original text can be segmented through natural language processing techniques, and the sentences that meet the requirements can be extracted as the target text elements; if the target text elements are terms, the original text can be tokenized through natural language processing techniques, and the terms that meet the requirements can be extracted as the target text elements.

[0162] Further, based on the target text elements corresponding to the original text, a target query statement is constructed. The specific process is as follows: First, the extracted target text elements are preprocessed, such as removing punctuation marks and unifying case, etc., to ensure the accuracy and consistency of the query statement. Then, according to the query syntax and rules of the target knowledge base, a query statement containing the target text elements is constructed. For example, if the target knowledge base is built based on Elasticsearch, a query statement supported by Elasticsearch can be constructed for efficient retrieval in the target knowledge base.

[0163] Further, according to the target query statement, the target text element pairs corresponding to the original text are retrieved from the target knowledge base. The specific process is as follows: The constructed target query statement is submitted to the target knowledge base for retrieval, and the target knowledge base matches and searches in the stored text element pairs according to the target text elements in the query statement. If a text element pair that matches the target text elements is found, it is returned as the target text element pair corresponding to the original text.

[0164] It should be noted that this embodiment is applicable to some text elements with a relatively low language structure level, such as phrases, words, or terms composed of phrases and words.

[0165] In this embodiment, by using the method of querying with relevant query statements, the target text element pairs corresponding to the original text can be determined, where the target text element pairs corresponding to the original text are semantically equivalent to some parts of the original text and can provide an accurate reference for the subsequent translation process.

[0166] In one embodiment, based on Figure 1 - Figure 11 For the illustrated embodiment, please refer to Figure 12 , Figure 12 which is an example schematic diagram of the knowledge base construction and translation provided by the embodiments of the present application.

[0167] First, the source document is obtained according to the target field; the first document in the source language and the second document in the target language are extracted from the source document.

[0168] Construct a first prompt based on the first document, the second document, the feature information of the first text element, and the target domain.

[0169] Call a large language model based on the first prompt to enable the large language model to extract a first text element pair of the target domain from the first document and the second document. The first text element pair includes a first text element in the source language and a first text element in the target language that are in contrast to each other in content. Construct a second prompt based on the first text element pair, the feature information of the second text element, and the target domain. Call the large language model based on the second prompt to enable the large language model to extract a second text element pair of the target domain from the first text element pair. The second text element pair includes a second text element in the source language and a second text element in the target language that are in contrast to each other in content. Determine the first text element pair and the second text element pair as the target text element pair.

[0170] Write the target text element pair into a pre-created initial structured file to obtain a first structured file. Construct a fourth prompt based on at least one quality scoring dimension. Call the large language model based on the fourth prompt to enable the large language model to perform quality scoring on each target text element pair in the first structured file to obtain the quality scores of each target text element pair in the first structured file. Screen each target text element pair in the first structured file based on the quality scores of each target text element pair in the first structured file to obtain a second structured file. Import the second structured file into a pre-created initial knowledge base to obtain a target knowledge base.

[0171] So far, it can be confirmed that the construction of the target knowledge base is completed.

[0172] Furthermore, obtain the original text in the source language.

[0173] Determine the vector database corresponding to the target knowledge base. The vector database includes the feature vectors of the first text element pairs in the target knowledge base. Determine the feature vector of the original text. Perform vector retrieval in the vector database based on the feature vector of the original text to obtain the target feature vector. Determine the first text element pair corresponding to the original text based on the target feature vector.

[0174] Determine the second text element corresponding to the original text. Construct a target query statement based on the second text element corresponding to the original text. Query the target knowledge base according to the target query statement to obtain the second text element pair corresponding to the original text.

[0175] Call the large language model to translate the original text based on the first text element pair and the second text element pair corresponding to the original text to obtain the target text in the target language.

[0176] In this embodiment, first, the first text element pair and the second text element pair in the target domain are determined, and a target knowledge base is constructed based on the first text element pair and the second text element pair, so that the target knowledge base contains high-quality reference knowledge in the target domain. In the subsequent translation process, for the translation requirements of the original text in the target domain, the first text element pair and the second text element pair corresponding to the original text can be determined based on the target knowledge base, and the large language model can be called based on the first text element pair and the second text element pair corresponding to the original text to translate the original text, and an accurate target text can be obtained, effectively improving the accuracy of translation.

[0177] The following will combine Figure 13 to introduce the knowledge base construction device provided in the embodiments of the present application in detail. It should be noted that Figure 13 the knowledge base construction device in Figure 1 - Figure 8 is used to execute the method of the embodiments shown in the present application Figure 1 - Figure 8 For the sake of convenience of description, only the parts related to the embodiments of the present application are shown. For the specific technical details not disclosed, please refer to the embodiments shown in the present application Figure 1 - Figure 8 Specifically, the knowledge base construction device 10 may include an acquisition unit 11, a call unit 12, and a construction unit 13, which are specifically as follows:

[0178] The acquisition unit 11 is configured to acquire a first document in the source language and a second document in the target language;

[0179] The call unit 12 is configured to call a large language model to extract target text element pairs in the target domain from the first document and the second document, and the target text element pairs include target text elements in the source language and target text elements in the target language that are in contrast to each other in content;

[0180] The construction unit 13 is configured to construct a target knowledge base based on the target text element pairs.

[0181] Optionally, in some embodiments, the calling unit 12 may be configured to: construct a first prompt based on the first document, the second document, the feature information of the first text element, and the target domain; call a large language model based on the first prompt, so that the large language model extracts a first text element pair of the target domain from the first document and the second document, where the first text element pair includes a first text element in the source language and a first text element in the target language that are in content contrast to each other; construct a second prompt based on the first text element pair, the feature information of the second text element, and the target domain; call the large language model based on the second prompt, so that the large language model extracts a second text element pair of the target domain from the first text element pair, where the second text element pair includes a second text element in the source language and a second text element in the target language that are in content contrast to each other; determine the first text element pair and the second text element pair as the target text element pair.

[0182] Optionally, in some embodiments, the calling unit 12 may be configured to: construct a third prompt based on the first document, the second document, the feature information of the third text element, and the target domain; call a large language model based on the third prompt, so that the large language model extracts a third text element pair of the target domain from the first document and the second document, where the third text element pair includes a third text element in the source language and a third text element in the target language that are in content contrast to each other; determine the third text element pair as the target text element pair.

[0183] Optionally, in some embodiments, the constructing unit 13 may be configured to: write the target text element pair into a pre-created initial structured file to obtain a first structured file; perform data import on a pre-created initial knowledge base based on the first structured file to obtain a target knowledge base.

[0184] Optionally, in some embodiments, the constructing unit 13 may be configured to: construct a fourth prompt based on at least one quality scoring dimension; call a large language model based on the fourth prompt, so that the large language model performs quality scoring on each target text element pair in the first structured file to obtain the quality scores of each target text element pair in the first structured file; filter each target text element pair in the first structured file based on the quality scores of each target text element pair in the first structured file to obtain a second structured file; import the second structured file into a pre-created initial knowledge base to obtain a target knowledge base.

[0185] Optionally, in some embodiments, the constructing unit 13 may be configured to: determine a quality score benchmark value of the first structured file; delete the target text element pairs in the first structured file whose quality scores are lower than the quality score benchmark value to obtain a second structured file.

[0186] Optionally, in some embodiments, the obtaining unit 11 may be configured to: obtain a source document according to a target field; extract a first document in the source language and a second document in the target language from the source document.

[0187] For the effects that can be achieved in this embodiment, please refer to the relevant embodiments of the above knowledge base construction method, which will not be elaborated here.

[0188] Next, Figure 14 a detailed introduction to the translation device provided in the embodiments of the present application will be given. It should be noted that, Figure 14 the translation device in Figure 9 - Figure 11 is used to execute the method of the embodiments shown in the present application. For the sake of convenience of description, only the parts related to the embodiments of the present application are shown. For the specific technical details not disclosed, please refer to the embodiments shown in the present application Figure 9 - Figure 11 Specifically, the translation device 20 may include an obtaining unit 21, a determining unit 22, and a translating unit 23, which are specifically as follows:

[0189] The obtaining unit 21 is configured to obtain the original text in the source language;

[0190] The determining unit 22 is configured to determine the target text element pair corresponding to the original text based on the target knowledge base, and the target knowledge base is constructed by the above knowledge base construction method;

[0191] The translating unit 23 is configured to call a large language model to translate the original text based on the target text element pair corresponding to the original text to obtain the target text in the target language.

[0192] Optionally, in some embodiments, the determining unit 22 may be configured to: determine the vector database corresponding to the target knowledge base, where the vector database includes the feature vectors of the target text element pairs in the target knowledge base; determine the feature vector of the original text; perform vector retrieval in the vector database based on the feature vector of the original text to obtain the target feature vector; determine the target text element pair corresponding to the original text based on the target feature vector.

[0193] Optionally, in some embodiments, the determining unit 22 may be configured to: determine the target text element corresponding to the original text; construct a target query statement based on the target text element corresponding to the original text; query the target knowledge base according to the target query statement to obtain the target text element pair corresponding to the original text.

[0194] For the effects that can be achieved in this embodiment, please refer to the relevant embodiments of the above translation method, which will not be elaborated here.

[0195] Correspondingly, an electronic device 900 is further provided in the embodiments of the present application. Please refer to Figure 15 ,Figure 15 It is a schematic structural diagram of an electronic device provided by an embodiment of the present application. The electronic device 900 includes a processor 901 and a memory 902. Among them, the processor 901 is electrically connected to the memory 902.

[0196] The processor 901 is the control center of the electronic device 900, connecting various parts of the entire electronic device through various interfaces and lines. By running or calling the executable program code stored in the memory 902 and calling the data stored in the memory 902, it executes various functions of the electronic device and processes data, thereby monitoring the electronic device as a whole.

[0197] The memory 902 can be used to store executable program code and modules. The processor 901 executes various functional applications and knowledge base construction programs and translation programs by running the executable program code and modules stored in the memory 902. The memory 902 mainly includes a program storage area and a data storage area. Among them, the program storage area can store an operating system, executable program code required for at least one function, etc.; the data storage area can store data created according to the use of the electronic device.

[0198] In addition, the memory 902 can include high-speed random access memory and can also include non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other volatile solid-state storage devices. Correspondingly, the memory 902 can also include a memory controller to provide the processor 901 with access to the memory 902.

[0199] In this embodiment, the processor 901 in the electronic device 900 loads the instructions corresponding to the processes of one or more executable program codes into the memory 902 according to the following steps, and the processor 901 runs the executable program code stored in the memory 902 to implement various functions as follows:

[0200] Obtain a first document in the source language and a second document in the target language;

[0201] Call a large language model to extract target text element pairs in the target field from the first document and the second document. The target text element pairs include target text elements in the source language and target text elements in the target language that are in contrast in content;

[0202] Construct a target knowledge base based on the target text element pairs.

[0203] Optionally, when the processor 901 executes the call to the large language model to extract the target text element pairs in the target domain from the first document and the second document, it specifically executes: constructing a first prompt word based on the first document, the second document, the feature information of the first text element, and the target domain; calling the large language model based on the first prompt word, so that the large language model extracts the first text element pairs in the target domain from the first document and the second document, and the first text element pairs include the first text element in the source language and the first text element in the target language that are contrasted with each other in content; constructing a second prompt word based on the first text element pairs, the feature information of the second text element, and the target domain; calling the large language model based on the second prompt word, so that the large language model extracts the second text element pairs in the target domain from the first text element pairs, and the second text element pairs include the second text element in the source language and the second text element in the target language that are contrasted with each other in content; determining the first text element pairs and the second text element pairs as the target text element pairs.

[0204] Optionally, when the processor 901 executes the call to the large language model to extract the target text element pairs in the target domain from the first document and the second document, it specifically executes: constructing a third prompt word based on the first document, the second document, the feature information of the third text element, and the target domain; calling the large language model based on the third prompt word, so that the large language model extracts the third text element pairs in the target domain from the first document and the second document, and the third text element pairs include the third text element in the source language and the third text element in the target language that are contrasted with each other in content; determining the third text element pairs as the target text element pairs.

[0205] Optionally, when the processor 901 executes constructing the target knowledge base based on the target text element pairs, it specifically executes: writing the target text element pairs into a pre-created initial structured file to obtain a first structured file; importing data into the pre-created initial knowledge base based on the first structured file to obtain the target knowledge base.

[0206] Optionally, when the processor 901 executes importing data into the pre-created initial knowledge base based on the first structured file to obtain the target knowledge base, it specifically executes: constructing a fourth prompt word based on at least one quality scoring dimension; calling the large language model based on the fourth prompt word, so that the large language model performs quality scoring on each target text element pair in the first structured file to obtain the quality scores of each target text element pair in the first structured file; screening each target text element pair in the first structured file based on the quality scores of each target text element pair in the first structured file to obtain a second structured file; importing the second structured file into the pre-created initial knowledge base to obtain the target knowledge base.

[0207] Optionally, when the processor 901 filters each pair of target text elements in the first structured file based on the quality scores of each pair of target text elements in the first structured file to obtain a second structured file, it specifically performs: determining a quality score benchmark value of the first structured file; deleting the pairs of target text elements in the first structured file whose quality scores are lower than the quality score benchmark value to obtain a second structured file.

[0208] Optionally, when the processor 901 executes to obtain a first document in a source language and a second document in a target language, it specifically performs: obtaining a source document according to the target field; extracting the first document in the source language and the second document in the target language from the source document.

[0209] Optionally, the processor 901 may further execute:

[0210] Obtaining the original text in the source language;

[0211] Determining a pair of target text elements corresponding to the original text based on the target knowledge base;

[0212] Based on the pair of target text elements corresponding to the original text, invoking a large language model to translate the original text to obtain the target text in the target language.

[0213] Optionally, when the processor 901 executes to determine a pair of target text elements corresponding to the original text based on the target knowledge base, it specifically performs: determining a vector database corresponding to the target knowledge base, where the vector database includes feature vectors of pairs of target text elements in the target knowledge base; determining a feature vector of the original text; performing vector retrieval in the vector database based on the feature vector of the original text to obtain a target feature vector; determining a pair of target text elements corresponding to the original text based on the target feature vector.

[0214] Optionally, when the processor 901 executes to determine a pair of target text elements corresponding to the original text based on the target knowledge base, it specifically performs: determining target text elements corresponding to the original text; constructing a target query statement based on the target text elements corresponding to the original text; querying in the target knowledge base according to the target query statement to obtain a pair of target text elements corresponding to the original text.

[0215] For the effects that can be achieved in this embodiment, please refer to the relevant embodiments of the above knowledge base construction method or translation method, which will not be elaborated here.

[0216] The embodiment of the present application further provides a computer-readable storage medium, in which a computer program is stored. When the computer program is executed, it causes the computer to execute the above relevant method steps to implement any one of the knowledge base construction methods or translation methods provided in the above embodiments.

[0217] The embodiment of the present application also provides a computer program product. The computer program product stores at least one instruction, and when the at least one instruction is executed by a processor, it implements any one of the knowledge base construction methods or translation methods provided in the above embodiments.

[0218] Among them, the computer, computer-readable storage medium, and computer program product provided in this embodiment are all used to execute the corresponding methods provided above. Therefore, the beneficial effects they can achieve can refer to the beneficial effects in the corresponding methods provided above, and will not be elaborated here.

[0219] The above content is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily think of changes or substitutions, which should all be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. A method for constructing a knowledge base, characterized in that: include: Obtaining a first document in a source language and a second document in a target language; Invoking a large language model to extract a target text element pair in a target domain from the first document and the second document, the target text element pair comprising a target text element in a source language and a target text element in a target language that contrast with each other in content; A target knowledge base is constructed based on the target text element pairs.

2. The method according to claim 1, characterized in that The calling of the large language model to extract target text element pairs in the target domain from the first document and the second document includes: constructing a first prompt word based on the first document, the second document, the feature information of the first text element and the target domain; Invoke a large language model based on the first prompt word, so that the large language model extracts a first text element pair of the target domain from the first document and the second document, wherein the first text element pair includes a first text element in a source language and a first text element in a target language that contrast with each other in content; constructing a second prompt word based on the first text element pair, the feature information of the second text element and the target domain; Calling the large language model based on the second prompt word so that the large language model extracts a second text element pair in the target domain from the first text element pair, wherein the second text element pair includes a second text element in the source language and a second text element in the target language that contrast with each other in content; The first text element pair and the second text element pair are determined as target text element pairs.

3. The method according to claim 1, characterized in that: The calling of the large language model to extract target text element pairs in the target domain from the first document and the second document includes: constructing a third prompt word based on the feature information of the first document, the second document, the third text element, and the target domain; Invoking a large language model based on the third prompt word, so that the large language model extracts a third text element pair in the target domain from the first document and the second document, wherein the third text element pair includes a third text element in a source language and a third text element in a target language that contrast with each other in content; The third text element pair is determined as a target text element pair.

4. The method according to claim 1, characterized in that The constructing a target knowledge base based on the target text element pair comprises: Writing the target text element pair into a pre-created initial structured file to obtain a first structured file; Data is imported into a pre-created initial knowledge base based on the first structured file to obtain a target knowledge base.

5. The method according to claim 4, characterized in that The step of importing data into a pre-created initial knowledge base based on the first structured file to obtain a target knowledge base includes: constructing a fourth prompt word based on at least one quality rating dimension; calling the large language model based on the fourth prompt word, so that the large language model performs a quality score on each of the target text element pairs in the first structured file, and obtains a quality score for each of the target text element pairs in the first structured file; Based on the quality scores of the target text element pairs in the first structured file, the target text element pairs in the first structured file are screened to obtain a second structured file; The second structured file is imported into a pre-created initial knowledge base to obtain a target knowledge base.

6. The method according to claim 5, characterized in that The method of screening the target text element pairs in the first structured file based on the quality scores of the target text element pairs in the first structured file to obtain a second structured file includes: Determining a quality score benchmark value of the first structured document; The target text element pairs whose quality scores in the first structured file are lower than the quality score reference value are deleted to obtain a second structured file.

7. The method according to claim 1, characterized in that The step of obtaining a first document in a source language and a second document in a target language includes: Get source documents based on target domain; A first document in a source language and a second document in a target language are extracted from the source document.

8. A translation method, characterized in that: include: Obtain the original text in the source language; Determine a target text element pair corresponding to the original text based on a target knowledge base, wherein the target knowledge base is constructed based on the method described in any one of claims 1 to 7; Based on the target text element pairs corresponding to the original text, a large language model is called to translate the original text to obtain a target text in a target language.

9. The method according to claim 8, characterized in that The determining, based on the target knowledge base, the target text element pair corresponding to the original text includes: Determine a vector database corresponding to the target knowledge base, wherein the vector database includes feature vectors of target text element pairs in the target knowledge base; Determining a feature vector of the original text; Performing vector search in the vector database based on the feature vector of the original text to obtain a target feature vector; Based on the target feature vector, a target text element pair corresponding to the original text is determined.

10. The method according to claim 8, characterized in that The determining, based on the target knowledge base, the target text element pair corresponding to the original text includes: Determine a target text element corresponding to the original text; Constructing a target query statement based on the target text element corresponding to the original text; According to the target query statement, the target text element pairs corresponding to the original text are queried in the target knowledge base.

11. A knowledge base construction device, characterized in that: The device comprises: An acquisition unit, configured to acquire a first document in a source language and a second document in a target language; A calling unit, configured to call a large language model to extract a target text element pair in a target domain from the first document and the second document, wherein the target text element pair includes a target text element in a source language and a target text element in a target language that are mutually contrasting in content; A construction unit is used to construct a target knowledge base based on the target text element pairs.

12. A translation device, characterized in that: The device comprises: An acquisition unit, used for acquiring the original text in the source language; A determination unit, configured to determine a target text element pair corresponding to the original text based on a target knowledge base, wherein the target knowledge base is constructed based on the method according to any one of claims 1 to 7; The translation unit is used to call the large language model to translate the original text based on the target text element pairs corresponding to the original text to obtain the target text in the target language.

13. An electronic device, characterized in that: The electronic device comprises: A memory for storing executable program codes; A processor, configured to call and run the executable program code from the memory, so that the electronic device executes the method according to any one of claims 1 to 10.

14. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed, the method according to any one of claims 1 to 10 is implemented.

15. A computer program product, wherein the computer program product stores at least one instruction, and when the at least one instruction is executed by a processor, the method according to any one of claims 1 to 10 is implemented.

Citation Information

Cited By

  • Knowledge base construction method and apparatus, translation method and apparatus, and device and medium

    WO2026179089A1