Method and device for translating entity names in text and computer equipment
By constructing associated name groups and utilizing clues and context to determine associations, the problem of identity confusion in cross-language adaptation was solved. This enabled accurate identification and translation of different forms of the same person's name, improving translation accuracy and localization adaptability.
Patent Information
- Application Number
- CN202511525213.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-24
- Publication Date
- 2025-11-21
AI Technical Summary
In the process of cross-language adaptation, the diverse forms of the same person's name can lead to confusion about the identity of the person in the translation. Existing translation methods are unable to accurately identify the same entity, resulting in reduced translation accuracy.
By constructing associated name groups, translating based on the association between entity names, determining associations using clue information and context, and employing entity name extraction models and machine learning models for accurate identification and mapping.
It enables accurate identification of different forms of the same person's name during cross-language adaptation, avoiding confusion of the person's identity in the translation and improving the accuracy and localization adaptability of the translation.
Smart Images

Figure CN120996058A_ABST
Abstract
Description
Technical Field
[0001] This specification relates to the field of information technology processing technology, and in particular to a method, apparatus, computer device, and computer program product for translating entity names in text. Background Technology
[0002] In the cross-language adaptation of long texts such as novels and literary works, the translation of personal names is necessary. Currently, the translation of personal names is direct, such as translating the names literally. However, different language systems have certain differences in language structure, culture, and usage habits, and the translations obtained by directly translating personal names do not conform to the user's language habits. In addition, the same person's name may exist in multiple forms in long texts. For example, the same person may have multiple forms of personal name, such as full name, abbreviation, nickname, title, and alias. Direct translation can easily translate these different forms of personal name into the names of different people, resulting in confusion about the identity of the characters in the translation and reducing the accuracy of the translation.
[0003] In view of this, some embodiments of this specification provide a method, apparatus, computer device, and computer program product for translating entity names in text, aiming to accurately identify different forms of personal names belonging to the same person and avoid confusion of personal identities in the translation. Summary of the Invention
[0004] This specification provides one or more embodiments of a method for translating entity names in text. The method includes: obtaining text to be processed; extracting entity names from the text to be processed; constructing associated name groups by determining whether there is a correlation between entity names, wherein entity names in the same associated name group represent the same entity or different entities with the same characteristics; and mapping at least a portion of the entity names from the source language to the target language based on the associated name groups.
[0005] The method provided according to one or more embodiments of this specification constructs a group of associated names by determining whether there is a correlation between entity names, including: determining clue information based on entity names and text to be processed to characterize whether there is a correlation between two or more entity names; and determining whether there is a correlation between two or more entity names based on the clue information.
[0006] According to one or more embodiments of this specification, the clue information includes: a clue sentence for characterizing whether two or more entity names represent the same entity, and determining whether there is a correlation between two or more entity names based on the clue information includes: determining whether there is a correlation between two or more entity names based on the clue sentence.
[0007] The method provided according to one or more embodiments of this specification, which determines whether there is a relationship between two or more entity names based on a clue sentence, includes: determining the context associated with the clue sentence based on the text to be processed; and determining whether there is a relationship between two or more entity names based on the context associated with the clue sentence.
[0008] According to one or more embodiments of this specification, a method for determining the context associated with a clue sentence based on the text to be processed includes: in response to a clue sentence corresponding to two or more entity names including a clue sentence indicating that there is no association between the two or more entity names, determining the context associated with the clue sentence based on the text to be processed.
[0009] According to one or more embodiments of this specification, the clue information includes: context for characterizing whether two or more entity names represent the same entity; and determining whether there is a relationship between two or more entity names based on the clue information includes: determining whether there is a relationship between two or more entity names based on the context.
[0010] According to one or more embodiments of this specification, after constructing associated name groups by determining whether there is a correlation between entity names, the method further includes: constructing initial name groups by determining whether there is a correlation between entity names; extracting surnames from the initial name groups; and processing the initial name groups based on the surnames to obtain associated name groups.
[0011] The method provided according to one or more embodiments of this specification, which processes an initial name group based on a surname to obtain an associated name group, further includes: determining generic terms based on entity names in the text to be processed; and filtering generic terms from the initial name group to obtain the associated name group.
[0012] According to one or more embodiments of this specification, the method for mapping entity names in an associated name group from a source language to a target language includes: mapping surnames from a source language to a target language to obtain a surname mapping table; and mapping entity names in an associated name group from a source language to a target language based on the surname mapping table and the associated name group to obtain an entity name mapping table.
[0013] According to one or more embodiments of this specification, the method for extracting entity names from text to be processed includes: extracting entity names from text to be processed using an entity name extraction model, wherein the entity name extraction model is trained based on a training sample set, the training sample set includes a corpus labeled with term categories, the corpus is obtained based on a classification template, and the classification template includes terms and context containing the terms.
[0014] One or more embodiments of this specification also provide an entity name translation device in text, the device comprising: an acquisition module for acquiring text to be processed; an extraction module for extracting entity names from the text to be processed; a construction module for constructing associated name groups by determining whether there is a correlation between entity names, wherein entity names in the same associated name group represent the same entity or different entities having the same characteristics; and a mapping module for mapping at least a portion of the entity names from the source language to the target language based on the associated name groups.
[0015] One or more embodiments of this specification also provide a computer device, the computer device including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, it is able to implement the entity name translation method in the text described in some embodiments of this specification.
[0016] One or more embodiments of this specification also provide a computer program product, including a computer program that, when at least a portion of the computer program is executed by a processor, enables the implementation of the entity name translation method in text described in some embodiments of this specification. Attached Figure Description
[0017] This specification will be further described by way of exemplary embodiments, which will be described in detail with reference to the accompanying drawings. The same numbers in the drawings denote the same structures or steps.
[0018] Figure 1 This is an exemplary flowchart illustrating a method for translating entity names in text according to some embodiments of this specification.
[0019] Figure 2 This is a schematic diagram of textual evidence according to some embodiments of this specification.
[0020] Figure 3 This is a schematic diagram of an association name grouping according to some embodiments of this specification.
[0021] Figure 4 This is an exemplary flowchart of a method for obtaining associated name groupings according to some embodiments of this specification.
[0022] Figure 5 This is an embodiment of the entity name mapping representation shown in this specification.
[0023] Figure 6 This is an exemplary block diagram of an entity name translation device in text according to some embodiments of this specification. Detailed Implementation
[0024] To more clearly illustrate the technical solutions of the embodiments in this specification, the embodiments will be described in detail below with reference to the accompanying drawings. Obviously, the content described below are some examples or embodiments of this specification. For those skilled in the art, without creative effort, the technical solutions or means disclosed in this specification can be applied to other scenarios based on this technical content.
[0025] It should be understood that the terms "system," "device," "unit," and / or "module" used in this specification are a method of distinguishing different components, elements, parts, sections, or assemblies at different levels. However, if other terms can achieve the same purpose, they may be replaced by other expressions.
[0026] Unless otherwise specified, the technical terms used to describe components, elements, etc. in this specification are not singular but may include plural. Generally speaking, terms such as "comprising" or "including" only indicate that explicitly identified steps, elements, or components are included, and these steps, elements, and components do not constitute an exclusive list, as the described method or apparatus may also include other steps or components.
[0027] This specification uses flowcharts to illustrate the operational steps performed by the apparatus or system of related embodiments. However, unless otherwise specified, the order in which these steps are described should not be construed as a limitation on the order of execution. Those skilled in the art can adjust the order of these steps based on the knowledge and information conveyed by the embodiments in this specification. Such adjustments include, but are not limited to, reversing the order of steps, merging multiple steps, and splitting a step.
[0028] Cross-language text adaptation refers to the process of converting and localizing source language texts according to the language habits and cultural background of the target language, without altering the core content of the source language text. This allows the text to naturally and accurately convey its intended meaning within the target language's cultural context. In the cross-language adaptation of long texts such as novels and literary works, the translation of entity names is necessary. An entity can refer to an independently existing, identifiable "thing" or "object." Entities can include concrete entities, such as people, places, organizations, institutions, items, and products. They can also include abstract entities, such as events, works, skills, and techniques. An entity name is an identifier used to identify, refer to, or name a specific entity. Examples include names of people, places, institutions, events, works, and skills.
[0029] In some related embodiments, the forms in which personal names appear in novels and literary works are diverse, such as full names, abbreviations, nicknames, appellations, aliases, etc., rather than simple names in the traditional sense, making it difficult to extract personal names (such as full names) and their variants (such as abbreviations, nicknames, appellations, aliases, etc.). In some related embodiments, the same character in a long text may appear in many variants. Some current translation methods mainly rely on string matching, statistical features, dependency analysis, etc., and it is difficult to accurately identify that these variants belong to the same entity. In some related embodiments, during cross - language localization processing, different appellation forms of the same character are often localized separately, resulting in confusion in the identities of characters in the translated text and a lack of systematic mapping and adaptation to the surname system of the target language, making the adaptation of surname culture poor. In some related embodiments, the translation of entity names can be direct translation, for example, translating the entity name literally from the grammatical or lexical level. However, there are certain differences in language structure, culture, or language usage habits between different language systems, and the translated text obtained by directly translating entity names does not conform to the user's language habits and cannot be localized. For example, when translating the Chinese name "Li Hua" into a Japanese name, direct translation will result in "李華", while a more localized translation is to translate "Li Hua" into "藤原 Shin".
[0030] Therefore, some embodiments of this specification propose a method for translating entity names in a text. The method includes obtaining the text to be processed, extracting the entity names in the text to be processed, constructing associated name groups by determining whether there is an association between entity names, and mapping at least some of the entity names in the entity names from the source language to the target language based on the associated name groups. Since the entity names in the same associated name group represent the same entity or different entities with the same characteristics, language mapping can be performed based on the relationship between the entity names in the associated name group during language mapping, so as to accurately identify that these different forms of personal names belong to the same person and avoid confusion in the identities of characters in the translated text.
[0031] Figure 1 It is an exemplary flowchart of a method for translating entity names in a text shown in some embodiments of this specification. Figure 1The illustrated process 100 can be executed by a computing device, for example, by a text entity name translation device 600 deployed on the computing device. In some embodiments, the computing device can be a server for storing, analyzing, and processing text. The server can include a local server or a cloud server; depending on different service needs, a local server corresponding to a specific region can be deployed in one or more regions. In some embodiments, the server can be a single computer or a computing cluster composed of multiple computers, thereby providing more powerful computing power and more efficient response to user service requests. In some embodiments, the computing device can be a terminal device, including but not limited to desktop computers, smartphones, laptops, VR (Virtual Reality) devices, tablets, smart TVs, in-vehicle terminals, etc. In some embodiments, part of process 100 can be executed by a computing device acting as a user end, while another part can be executed by a computing device acting as a server end. Figure 1 As shown, process 100 may include the following steps.
[0032] Step 110: Obtain the text to be processed. In some embodiments, step 110 may be implemented by the acquisition module 610.
[0033] In some embodiments, source text can be obtained and structured to obtain structured text content, which is then used as the text to be processed.
[0034] For example, a data access device can acquire source text sent by a terminal device. The data access device can convert source text in different formats into standard format text, enabling flexible data input methods. For instance, source text in different formats such as Excel, CSV, and TXT files can be imported locally into the data access device, which then performs format conversion. Another example is that the data access device receives text data from the network via an API (Application Programming Interface) and performs format conversion on the incoming text data. Furthermore, a text preprocessing engine automatically segments and extracts structured content from standard format text. For example, it extracts the text content corresponding to each part of the standard format text according to the structure of chapters, headings, and paragraphs, forming structured text content, which can then be used as the text to be processed.
[0035] In some embodiments, basic entity names and their corresponding contexts can be initially extracted from standard-formatted text. For example, a text preprocessing engine can extract the name "Wang Baoxiang" and the corresponding context "Wang Baoxiang looked back..." from standard-formatted text. The text preprocessing engine can extract names from standard-formatted text using deep learning models (such as BiLSTM-CRF (Bidirectional Long Short-Term Memory Network-Conditional Random Field Model)) or based on dictionary matching.
[0036] In some embodiments, when extracting entity names from standard format text, overlapping entity names can be filtered. For example, if the standard format text includes "Wang Baoxiang looked back...", the entity names extracted from this text include "Wang Baoxiang" and "Baoxiang". Since the entity names "Wang Baoxiang" and "Baoxiang" are overlapping entities, the overlapping entity name "Baoxiang" can be filtered out, retaining the entity name "Wang Baoxiang", and extracting the context "Wang Baoxiang looked back..." corresponding to the entity name "Wang Baoxiang". In some embodiments, the structured text content, the extracted entity names, and the context corresponding to the entity names can be used as the text to be processed.
[0037] Step 120: Extract entity names from the text to be processed. In some embodiments, step 120 can be implemented by extraction module 620.
[0038] In some embodiments, entity names can be extracted from the text to be processed using an entity name extraction model. Examples of such models include the Qwen model, the LLama model, and the Baichuan model.
[0039] In some embodiments, the entity name extraction model may include the Qwen-72B-Instruct model. The text to be processed can be input into the Qwen-72B-Instruct model, and the Qwen-72B-Instruct model outputs a list of entity names (e.g., full name, abbreviation, nickname, title, alias, etc.) containing the full name of the entity.
[0040] In some embodiments, entity names and their corresponding categories can be extracted from the text to be processed using an entity name extraction model. Categories can include male character names, female character names, titles, full names, aliases, etc. For example, the entity name "Miss Su" is categorized as a female character name, and the entity name "Achen" is categorized as a title. A full name can be the character's full name, such as "Su Nanqing," and an alias can be another name for the character besides the full name, such as "Qingqing."
[0041] In some embodiments, the computing device can also label names based on the categories corresponding to entity names, labeling entity names classified as full names as parent names and entity names classified as aliases as child names. For example, if the names identified from a certain chapter of the text to be processed include Su Nanqing, Miss Su, Qingqing, Anti, and Black Cat, where Su Nanqing is classified as a full name and the other names are classified as aliases, then Su Nanqing can be labeled as a parent name, and Miss Su, Qingqing, Anti, and Black Cat can be labeled as child names.
[0042] In some embodiments, the entity name extraction model can be obtained by training a first preliminary machine learning model based on a first training sample set. In some embodiments, the first preliminary machine learning model can be an untrained machine learning model. The first training sample set may include a corpus labeled with term categories, and terms may include different forms of entity names, such as full names, abbreviations, nicknames, titles, aliases, etc. Term categories may include male character names, female character names, appellations, full names, aliases, etc.
[0043] In some embodiments, the corpus can be obtained based on a classification template. The classification template can be the input format of a language model used to filter terms in a term candidate set. For example, terms can be extracted from different text fragments or novels using a generative language model (e.g., ChatGPT-3.5), and the results can be merged and duplicate terms removed to obtain a preliminary term candidate set. Then, common terms in the preliminary term candidate set are filtered using the TF-IDF algorithm, and high-frequency prefixes and suffixes in the preliminary term candidate set are filtered using a word segmenter to obtain an intermediate term candidate set. Then, a classification template is constructed and used as the input template format of an encoding language model to further filter the terms in the intermediate term candidate set, obtaining the corpus. The classification template can include terms and context containing the terms, for example, term + separator + context before the term + term + context after the term. The encoding language model can be, for example, BERT (Bidirectional Encoder Representations from Transformers). The input template format for an encoding-based language model can be set as term + separator + context before the term + term + context after the term. Terms from the intermediate term candidate set and their corresponding contexts can be input into the encoding-based language model according to this template format. The encoding-based language model analyzes the input terms and their corresponding contexts, determines whether they are terms that need to be filtered, removes the terms that need to be filtered from the intermediate term candidate set to obtain a term set, and then categorizes the remaining terms. For example, the category label can indicate whether the remaining term is a term, whether it is a male or female name, whether it is a title, whether it is a full name or an alias, etc., to obtain the final term set. Furthermore, based on the final term set, terms in the original text (e.g., text fragments, novels) can be annotated to obtain a large amount of term annotation data, which can then be used as the first training sample set.
[0044] In some embodiments, the first preliminary machine learning model may be a trained machine learning model (i.e., a previously trained entity name extraction model, which may be referred to as the preliminary entity name extraction model). The training data used to obtain the preliminary entity name extraction model is at least partially different from the samples in the first training sample set. The entity name extraction model can be obtained by fine-tuning the preliminary entity name extraction model using the first training sample set. For example, the Qwen model can be selected as the preliminary entity name extraction model, and multiple candidate entity name extraction models can be obtained by performing comprehensive full-parameter fine-tuning on multiple scales of the Qwen model (e.g., 7B, 14B, 32B, and 72B). The final entity name extraction model is determined by quality evaluation of the multiple candidate entity name extraction models. The quality evaluation may include precision, recall, inference speed, etc.
[0045] The fine-tuning process of the initial entity name extraction model may include: inputting term annotation data from the training sample set into the entity name extraction model for iterative training, and systematically optimizing key hyperparameters during training, such as learning rate, batch size, and learning rate scheduler. During training, gradients are calculated through backpropagation of the loss function to continuously optimize the model parameters of the initial entity name extraction model. After multiple iterations of training and model parameter optimization, an entity name extraction model or a candidate entity name extraction model is obtained.
[0046] Furthermore, experimental results show that the 72B-scale Qwen model (i.e., the Qwen-72B-Instruct model) performs best in terms of accuracy. Considering the inference latency and computational cost of large-scale models in actual deployment, AWQ quantization technology can also be applied to the Qwen-72B-Instruct model. The quantized Qwen-72B-finetuned-AWQ model maintains high accuracy while significantly improving inference speed, achieving the best balance between model performance and running efficiency.
[0047] In some embodiments, entity names in the text to be processed can also be extracted based on a preset entity name library. For example, the text to be processed can be matched with entity names in the preset entity name library to obtain the successfully matched entity names.
[0048] In other embodiments, entity names in the text to be processed can also be extracted based on regular expressions. For example, a surname database and a given name database can be established, and any combination of the surname and given name databases can be matched against the text to be processed to obtain the successfully matched entity names.
[0049] Step 130 involves constructing associated name groups by determining whether there is a correlation between entity names. In some embodiments, step 130 can be implemented by the construction module 630.
[0050] In some embodiments, it can be first determined whether there is a relationship between entity names, and then a group of related entity names can be constructed. Entity names in the same group of related names can represent the same entity or different entities with the same characteristics. When translating based on the group of related names, the translation can be based on the relationship between people in the group of related names, so as to accurately identify that these different forms of names belong to the same person, avoid confusion of the identity of people in the translation, improve the accuracy of translation, and achieve localized translation.
[0051] The presence of a relationship between entity names indicates whether the entity names refer to the same entity or whether there are different entities with common characteristics. When entity names are related, it means that the entity names refer to the same entity or that there are common characteristics between the entity names, such as having the same last name and / or having the same second character in the first name.
[0052] In some embodiments, initial name groups can be constructed by determining whether there is a relationship between entity names. Furthermore, the initial name groups can be designated as associated name groups. When there is a relationship between two or more entity names, it means that the two or more entity names represent the same entity and are grouped into the same initial name group. When there is no relationship between two or more entity names, it means that the two or more entity names represent different entities, and therefore, the two or more entity names are not in the same initial name group.
[0053] In some embodiments, surnames can be extracted from the initial name groups, and the initial name groups can be processed based on the surnames to obtain associated name groups. For example, multiple initial name groups can be clustered based on surnames, grouping initial name groups with the same surname into the same associated name group. For a detailed explanation of this part, please refer to the description in process 400 below.
[0054] In some embodiments, two or more entity names can refer to a parent entity name and one or more child entity names. The parent entity name can be the fully qualified name of the entity, such as "Su Nanqing". The child entity name can be an alias of the entity, such as "Miss Su", "Anti", etc. In other embodiments, the two or more entity names can also be any entity names, without specific limitations.
[0055] Let's first explain how to determine whether there is a relationship between entity names.
[0056] In some embodiments, clue information representing the existence of a relationship between two or more entity names can be determined based on entity names and the text to be processed, and the relationship between the two or more entity names can be determined based on the clue information. In some embodiments, the clue information may include a clue sentence representing whether two or more entity names refer to the same entity. For example, "'Miss Su, you are pregnant.' The doctor's words were like a thunderbolt, making the drowsy Su Nanqing suddenly open her eyes wide: '...What?'" can indicate that "Miss Su" and "Su Nanqing" are the same person. As another example, "Su Nanqing glanced at Mrs. Su and said coldly, 'Please leave!'" can indicate that "Su Nanqing" and "Mrs. Su" are not the same person.
[0057] In some embodiments, the existence of a relationship between two or more entity names can be determined based on clue sentences. For example, a first inference model can extract clue sentences from the text to be processed to characterize whether a relationship exists between two or more entity names, and determine whether a relationship exists between two or more entity names based on the clue sentences. Exemplarily, the first inference model can be a Reasoning COT model based on a Generative Large Language Model (LLM). A set of entity names whose relationship needs to be identified, extracted from the text to be processed, and the text to be processed can be input into the first inference model. The first inference model can extract clue sentences from the text to be processed based on this input set of entity names, determine whether a relationship exists between two or more entity names in the set of entity names based on the clue sentences, and output a set of textual evidence. The textual evidence may include a set of entity names (including two or more entity names), the judgment result of whether a relationship exists between each entity name in the set of entity names, and clue sentences characterizing whether a relationship exists between each entity name in the set of entity names. Clue sentences may include supporting sentences or opposing sentences. Supporting sentences are sentences that support the idea that two or more entity names represent the same entity; opposing sentences are sentences that oppose the idea that two or more entity names represent the same entity.
[0058] Figure 2 This is a schematic diagram of textual evidence according to some embodiments of this specification. For example, entity names "Su Nanqing," "Miss Su,"..."Anti" extracted from a section of the text to be processed, along with the text to be processed or the corresponding section text, can be input into the first inference model. The first inference model outputs as follows: Figure 2The example given is a set of textual evidence. For instance, the textual evidence includes the entity names "Su Nanqing" and "Miss Su," and the judgment result is "Support," indicating that the entity names "Su Nanqing" and "Miss Su" are related. The corresponding supporting sentence is: "'Miss Su, you are pregnant.' The doctor's words were like a thunderbolt, making the drowsy Su Nanqing suddenly open her eyes wide: '...What?'"
[0059] In some embodiments, a first inference model can be obtained by supervised fine-tuning and reinforcement learning training of a preliminary second machine learning model. The preliminary second machine learning model may include an untrained machine learning model or a previously trained first inference model (also referred to as a preliminary first inference model). For example, the preliminary first inference model can be supervisedly fine-tuned based on a second training sample set, which may include samples of two or more entity names, text fragment samples containing the two or more entity names, and annotation information for the text fragments. The annotation information may include annotation data for sentences in the text fragments that indicate the relationship between the two or more entity names (e.g., whether they represent the same entity), and may also include annotation data for whether the two or more entity names in the text fragments are related (e.g., support or rejection). Exemplarily, data from the second training sample set can be input into the preliminary first inference model for iterative training. During training, gradients are calculated through backpropagation using a loss function, continuously optimizing the model parameters of the preliminary first inference model to obtain a fine-tuned first inference model. The loss function can include the KL divergence loss function. The KL divergence loss function is used to balance the model's learning of new tasks with the retention of original knowledge, ensuring that the fine-tuned first inference model does not deviate too much from the original first inference model (i.e., the initial first inference model), thus preserving its general capabilities. After multiple iterations of training and model parameter optimization, the fine-tuned first inference model is obtained.
[0060] In some embodiments, reinforcement learning algorithms can be used to further improve the accuracy of the content generated by the first inference model, based on fine-tuning training. For example, the reinforcement learning algorithm could be the GRPO algorithm. During reinforcement learning training, text fragments from the first ten chapters of several popular online novels can be selected as reinforcement learning samples. These samples can include samples of two or more entity names and text fragment samples containing those two or more entity names. During reinforcement learning, the reinforcement learning samples can be input into the fine-tuned first inference model, and the output of the fine-tuned first inference model can be scored based on a rule-based reward template. This provides accurate feedback signals to the first inference model, guiding the generation of high-quality, formatted clue sentences. Further, the advantage value is calculated based on the score, and the policy parameters are updated. The value network is updated based on the loss function to obtain the first inference model. For example, the format of the reward template can be as follows: <item> <fullname> Full name 1< / fullname> <nickname> Alias 1< / nickname> <type> support< / type> <clues> Clue sentence a< / clues> <clues> Clue sentence b< / clues> < / item> .
[0061] In some embodiments, scoring the output of the fine-tuned first inference model can be achieved by parsing the output and determining whether the format of the output is correct based on the reward template. For example, if the format of the output matches the format of the reward template, it is judged as correct and awarded 0.1 points; if the format of the output does not match the format of the reward template, it is judged as incorrect and awarded 0 points. Alternatively, the item elements in the output can be traversed. Item elements can include the full name, alias, type (support or reject), and clue sentence. These elements can be verified to conform to specific rules. For example, verifying whether the "full name" and "alias" appear together in the clue sentence adds 0.4 points. It can also be verified whether the clue sentence appears completely in the original text; adding another 0.5 points if it does.
[0062] In some embodiments, after obtaining the clue sentence, the context associated with the clue sentence can be determined based on the text to be processed, and the relationship between two or more entity names can be determined based on the context associated with the clue sentence. For example, a query can be performed in the text to be processed based on the clue sentence to obtain the context within a specified window containing the clue sentence. The specified window can be a range in the text to be processed, for example, from N sentences before the clue sentence to M sentences after the clue sentence, or from N lines before the line containing the clue sentence to M lines after the line containing the clue sentence.
[0063] In some embodiments, when there are multiple clue sentences corresponding to the same entity name group, the context corresponding to the clue sentences can be concatenated. For example, the context corresponding to multiple clue sentences can be linearly concatenated according to their order in the original text to enhance the completeness of the context information. This allows the subsequent second inference model to perform inference based on the complete context information when determining whether there is a correlation between two or more entity names based on the context associated with the clue sentences, thereby improving the accuracy of the inference results.
[0064] In some embodiments, when there are multiple clue sentences corresponding to the same entity name group, and the contexts corresponding to the clue sentences overlap, the overlapping contexts can also be filtered. For example, a certain entity name group includes two names, and there are two corresponding clue sentences, located in the third and fifth lines of the text to be processed, respectively. Assuming the context associated with the clue sentences is from the two lines before the line number of the clue sentence to the two lines after the line number of the clue sentence, then the context associated with the third-line clue sentence is the content of the first, second, third, fourth, and fifth lines, and the context associated with the fifth-line clue sentence is the content of the third, fourth, fifth, sixth, and seventh lines. The contexts of the third, fourth, and fifth lines overlap and need to be filtered. The filtered contexts are then concatenated to obtain the context content of the first to seventh lines.
[0065] In some embodiments, a second inference model can be used to determine whether a relationship exists between two or more entity names based on the context associated with the clue sentence. In some embodiments, the second inference model and the first inference model can be the same model. In some embodiments, the second inference model and the first inference model can be different inference models.
[0066] For example, the second reasoning model can be a Reasoning COT model based on a Generative Large Language Model (LLM). The context associated with the clue sentence and the entity names in the clue sentence whose relevance needs to be determined can be input into the second reasoning model, which then outputs the judgment result. For instance, when the second reasoning model outputs "yes," it indicates that the entity names in the clue sentence are related; when the second reasoning model outputs "no," it indicates that the entity names in the clue sentence are not related.
[0067] In some embodiments, a second inference model can be obtained by supervised fine-tuning and reinforcement learning training of a preliminary third machine learning model. The preliminary third machine learning model may include an untrained machine learning model or a previously trained second inference model (also referred to as a preliminary second inference model). In some embodiments, the preliminary third machine learning model may be a first inference model.
[0068] The preliminary second inference model can be trained under supervised conditions based on a third training sample set. This third training sample set can include samples of two or more entity names, context samples containing those two or more entity names, and annotation information for the context samples. The annotation information can include labels indicating whether the context sample supports a correlation between the two or more entity names. For example, data from the third training sample set can be input into the preliminary second inference model for iterative training. During training, gradients are calculated through backpropagation of the loss function to continuously optimize the model parameters. After multiple iterations of training and model parameter optimization, a fine-tuned second inference model is obtained.
[0069] In some embodiments, reinforcement learning algorithms can be used to further improve the accuracy of the content generated by the second inference model, based on fine-tuning training. For example, the reinforcement learning algorithm could be the GRPO algorithm. During reinforcement learning training, text fragments from the first ten chapters of several popular online novels can be selected as reinforcement learning samples. These samples can include samples of two or more entity names and text fragment samples containing those two or more entity names. During reinforcement learning, these samples can be input into the fine-tuned third inference model, and a large language model (LLM) can be used to reward the output of the fine-tuned second inference model. For example, a correct judgment earns 0.5 points, and an incorrect judgment deducts 0.5 points. Furthermore, an advantage value is calculated based on the rating of the output content, and the policy parameters are updated. The value network is updated based on the loss function to obtain the second inference model.
[0070] After determining whether there is a relationship between two or more entity names based on the clue sentence, a second judgment can be made based on the context associated with the clue sentence to further determine whether there is a relationship between two or more entity names. This decomposes the entity name recognition problem of different forms into two independent but related stages: "clue sentence extraction" and "relationship disambiguation". The data flow between the stages simplifies the complex task and further improves the determination of the relationship between entity names.
[0071] In some embodiments, in response to the fact that the clue sentences corresponding to two or more entity names include clue sentences indicating that there is no association between the two or more entity names, the context associated with the clue sentences can be determined based on the text to be processed, and whether there is an association between the two or more entity names can be determined based on the context associated with the clue sentences. For example, if the text evidence output by the first inference model includes the entity names "Su Nanqing" and "Miss Su", and the corresponding clue sentences include clue sentences opposing the association between the entity names "Su Nanqing" and "Miss Su", then the context associated with all the clue sentences corresponding to the entity names "Su Nanqing" and "Miss Su" can be extracted from the text to be processed, and the association between the entity names "Su Nanqing" and "Miss Su" can be further determined based on the context. For example, the context and the entity names "Su Nanqing" and "Miss Su" can be input into the second inference model, and the second inference model can output a judgment result, which can include whether there is an association between the entity names "Su Nanqing" and "Miss Su" or whether there is no association.
[0072] In other embodiments, in response to the fact that the clue sentences corresponding to two or more entity names simultaneously include both clue sentences indicating that there is no correlation between the two or more entity names and clue sentences indicating that there is a correlation between the two or more entity names, the context associated with the clue sentences can be determined based on the text to be processed, and whether there is a correlation between the two or more entity names can be determined based on the context associated with the clue sentences. For example, if the text evidence output by the first inference model includes the entity names "Su Nanqing" and "Miss Su", and the corresponding clue sentences simultaneously include both clue sentences opposing the correlation between the entity names "Su Nanqing" and "Miss Su" and clue sentences supporting the correlation between the entity names "Su Nanqing" and "Miss Su", then the context associated with all the clue sentences corresponding to the entity names "Su Nanqing" and "Miss Su" can be extracted from the text to be processed, and the correlation between the entity names "Su Nanqing" and "Miss Su" can be further determined based on the context. For example, the context and the entity names "Su Nanqing" and "Miss Su" can be input into the second inference model. The second inference model can output a judgment result, which may include whether there is a relationship or no relationship between the entity names "Su Nanqing" and "Miss Su".
[0073] In some embodiments, the clue information can also be context used to characterize whether two or more entity names represent the same entity, such as "Su Nanqing walked into the hospital... 'Miss Su, you are pregnant.' The doctor's words were like a thunderbolt, making the drowsy Su Nanqing suddenly open her eyes wide: '...What?'...Su Nanqing finally decided to go somewhere first." In some embodiments, the existence of a relationship between two or more entity names can be determined based on the context. For example, a third inference model can be used to extract context from the text to be processed to characterize whether a relationship exists between two or more entity names, and the relationship between two or more entity names can be determined based on the context.
[0074] In some embodiments, the third inference model, the second inference model, and the first inference model may be the same inference model.
[0075] In some embodiments, the third inference model, the second inference model, and the first inference model can be different inference models.
[0076] In some embodiments, the third inference model and the second inference model can be the same inference model. In some embodiments, the third inference model and the second inference model can be different inference models.
[0077] For example, the third reasoning model can be a Reasoning COT model based on a Generative Large Language Model (LLM). A set of entity names extracted from the text to be processed, along with the text itself, can be input into the third reasoning model. Based on this set of entity names, the third reasoning model can directly extract context from the text, determine whether there is a relationship between two or more entity names within the set, and output a set of textual evidence. This textual evidence may include the entity name set (containing two or more entity names), the judgment result regarding the relationship between each entity name in the entity name set, and the context used to characterize the relationship between the entity names in the entity name set.
[0078] In some embodiments, a third inference model can be obtained by supervising fine-tuning and reinforcement learning training of a first or second inference model. In some embodiments, the first or second inference model can be supervised based on a fourth training sample set. This fourth training sample set may include samples of two or more entity names, text fragment samples containing the two or more entity names, and annotation information for the text fragments. The annotation information may include annotation data of the context representing the relationship between the two or more entity names (e.g., whether they represent the same entity) in the text fragment, and may also include annotation data of the relationship type representing the two or more entity names (e.g., support or rejection) in the text fragment. For explanations regarding supervising fine-tuning training of the first or second inference model based on the fourth training sample set and reinforcement learning of the inference model, refer to the preceding descriptions of supervising fine-tuning training of the first inference model based on the second training sample set and reinforcement learning of the first inference model, the difference being that the text lengths of the clue sentences and the context are different.
[0079] In some embodiments, associated name groups can be constructed for entity names that are confirmed to be related. Entity names in the same associated name group (here, the associated name group can refer to the initial name group) represent the same entity. For example, binary relation pairs can be constructed based on the related entity names, and associated name groups can be constructed based on each binary relation pair. For instance, if the entity names "Su Nanqing" and "Miss Su" are determined to be related, and the entity names "Su Nanqing" and "Anti" are determined to be related, then binary relation pairs (Su Nanqing, Miss Su) and (Su Nanqing, Anti) can be constructed respectively, and an associated name group (Su Nanqing, Miss Su, Anti) can be constructed based on these binary relation pairs. All entity names in this associated name group represent the same person.
[0080] Figure 3 This is a schematic diagram illustrating an associated name grouping according to some embodiments of this specification. For example, associated name groups (where associated name groups can refer to initial name groups) can be constructed based on binary relation pairs (Su Nanqing, Miss Su), (Su Nanqing, Qingqing)...(Su Nanqing, QQ). Figure 3 As shown. Furthermore, the associated name grouping can also include the number of supporting and opposing statements for each binary relation pair, as well as the final judgment result after secondary judgment. For example, for the entity names "Su Nanqing" and "Miss Su," there are 298 supporting clues and 5 opposing clues. After secondary judgment based on the context associated with the clues, the judgment result is "yes," indicating that the entity names "Su Nanqing" and "Miss Su" are related and represent the same person.
[0081] In one or more embodiments of this specification, when the entity names in the associated name group represent the same entity, the associated name group may refer to the initial name group; when the entity names in the associated name group represent different entities with the same characteristics, the associated name group may refer to the associated name group obtained after processing the initial name group based on the surname.
[0082] Step 140: Based on the associated name grouping, at least a portion of the entity names in the entity names are mapped from the source language to the target language. In some embodiments, step 140 may be implemented by the mapping module 640.
[0083] In some embodiments, a Large Language Model (LLM) can be used to translate associated name groups (e.g., initial name groups). For example, a Large Language Model could be ChatGPT-3.5. When using an LLM for translation, LLM prompts can be constructed, for example, indicating that all entity names in the associated name group (e.g., the initial name group) represent the same person, requiring the LLM to translate all entity names in the associated name group. The LLM maps the entity names in the associated name group from the source language to the target language based on the prompts, obtaining an entity name mapping table.
[0084] In some embodiments, surnames can be mapped from the source language to the target language to obtain a surname mapping table, and entity names in the associated name groups can be mapped from the source language to the target language based on the surname mapping table and associated name groups to obtain an entity name mapping table.
[0085] In some embodiments, a Large Language Model (LLM) can be invoked to map surnames in a surname set from the source language to the target language in a localized manner, and a surname mapping table can be constructed. For example, for the surname "Zhang", with Chinese as the source language and Japanese as the target language, an LLM prompt can be constructed, requiring the model to translate it into a common surname in Japanese and follow Japanese cultural conventions. Based on the prompt, the Large Language Model translates the surname "Zhang" into "Sato".
[0086] In some embodiments, an entity name mapping table can be obtained by invoking a Large Language Model (LLM) to map entity names in each associated name group from the source language to the target language based on a surname mapping table and associated name groupings. For example, an LLM prompt can be constructed, requiring the Large Language Model to perform localization translation based on the surname mapping table and associated name groupings, mapping entity names from the source language to the target language, and constructing the entity name mapping table. Figure 5 This is an embodiment of the entity name mapping representation shown in this specification. For example... Figure 5 As shown, the entity name mapping table contains the entity names before translation (e.g., ...). Figure 5The original terminology and the translated entity names (such as...) Figure 5 The translation of terms in the text can also include the number of chapters in which entity names appear (e.g., ...). Figure 5 The number of chapters in which the Chinese term appears and the text in which the entity name is located (e.g., Figure 5 (The original text of the Chinese terminology is located in the main body). Additionally, the entity name mapping table can also include classifications of entity names (such as...). Figure 5 (The terminology classification in the text) The classification of the entity name can be determined when extracting the entity name in step 120.
[0087] In some embodiments, the entity name mapping table can also be validated. For example, it can be validated to ensure that the surnames and given names of people in the entity name mapping table are correctly translated and that there are no omissions in the translation.
[0088] Figure 4 This is an exemplary flowchart of a method for obtaining associated name groupings according to some embodiments of this specification. Figure 4 The illustrated process 400 can be executed by a computing device, for example, by a server-side computing device. In some embodiments, process 400 can be implemented by an entity name translation device 600 deployed on a computing device. Figure 4 As shown, in some embodiments, process 400 may include the following steps.
[0089] Step 410 involves constructing initial name groups by determining whether there are any relationships between entity names. In some embodiments, step 410 can be implemented by construction module 630.
[0090] In some embodiments, an initial name group can be constructed in accordance with the method for constructing associated name groups in steps 110-130, wherein the entity names in the same initial name group represent the same entity.
[0091] Step 420: Extract the last name from the initial name group. In some embodiments, step 420 can be implemented by the construction module 630.
[0092] In some embodiments, the surnames of entity names in the initial name group can be extracted using a fourth inference model to form a surname set. For example, the fourth inference model can be a Reasoning COT model. The initial name group can be input into the fourth inference model, which extracts the surnames of the entity names in the initial name group and constructs the surname set. For example, the surname set could be {"Zhang", "Li", "Ouyang"}. In other embodiments, the surnames of all entity names in the text to be processed can also be extracted to form a surname set.
[0093] In some embodiments, the fourth inference model can be obtained by performing supervised fine-tuning training on a preliminary fourth machine learning model. The preliminary fourth machine learning model can include an untrained machine learning model or a previously trained fourth inference model (also referred to as a preliminary fourth inference model). For example, the preliminary fourth inference model can be subjected to supervised fine-tuning training based on a fifth training sample set, which can include names and annotation information for surnames. The annotation information can include the surnames in the names. Exemplarily, the data in the fifth training sample set can be input into the preliminary fourth inference model for iterative training. During the training process, the gradient is calculated through backpropagation of the loss function, and the model parameters of the preliminary fourth inference model are continuously optimized to obtain the fourth inference model. In some embodiments, the surnames of the entity names in the initial name groups can also be extracted through a preset surname dictionary. For example, the entity names in the initial name groups can be matched in the preset surname dictionary, and when the match is successful, the corresponding surnames are extracted.
[0094] Step 430: Process the initial name groups based on the surnames to obtain associated name groups. In some embodiments, step 430 can be implemented by a building block 630.
[0095] In some embodiments, a relationship graph can be constructed and connectivity analysis can be performed based on each initial name group and the surnames to obtain associated name groups. The entity names in the same associated name group represent different entities with the same characteristics. The same characteristics can be having the same surname and / or the second character in the name being the same. For example, "Su Nanfeng" and "Su Nanbei". Exemplarily, a surname node can be introduced in the relationship graph construction, and the entity names associated with the surname are used as nodes, and edges are established with the surname node through the relationship of "belonging to the same surname". Further, based on the above, taking each entity name as a node, edges are connected according to the association between the entity names confirmed in the initial name groups. For example, edges are established between two entity names representing the same entity, thereby forming a comprehensive graph structure including surname connections and associated entity name connections. By performing connectivity analysis on this comprehensive graph structure, each maximally connected subgraph is obtained, and the maximally connected subgraphs are used as the associated name groups. The associated name groups contain both the relationships of entity names belonging to the same surname and the relationships of entity names representing the same entity. Subsequently, when performing translation based on the associated name groups, local translation can be performed based on the relationships contained in the associated name groups.
[0096] In some embodiments, generic terms can be further determined based on entity names in the text to be processed, and filtered from the initial name grouping to obtain associated name groups. Entity names in the same associated name group represent different entities with the same characteristics. These same characteristics can include having the same surname. Generic terms can refer to general appellations that do not have a clear individual reference. In some embodiments, generic terms in the form of "surname + appellation" or "first name + appellation" can be extracted from the text to be processed, such as Young Master Huo, Old Lady Huo, Mr. Gu, General Manager Ma, Senior Brother Qi, Miss Black Cat, etc. For example, generic terms in the text to be processed can be extracted using a fifth reasoning model, and a list of generic terms can be established. The fifth reasoning model can be the Reasoning COT model.
[0097] In some embodiments, a fifth inference model can be obtained by supervised fine-tuning training of a preliminary sixth machine learning model. The preliminary sixth machine learning model may include an untrained machine learning model or a previously trained fifth inference model (also referred to as a preliminary fifth inference model). For example, the preliminary fifth inference model can be supervisedly fine-tuned based on a sixth training sample set, which may include text content containing generic terms and annotation information for these generic terms. The annotation information may include generic terms in the text content. Exemplarily, data from the sixth training sample set can be input into the preliminary fifth inference model for iterative training. During training, gradients are calculated through backpropagation of the loss function, continuously optimizing the model parameters of the preliminary fifth inference model to obtain the fifth inference model. In some embodiments, the fourth and fifth inference models may be the same inference model or different inference models.
[0098] In some embodiments, generic terms in the initial name group can also be extracted through a preset generic term library. For example, the entity names in the initial name group can be matched in the preset generic term library, and when a match is successful, the corresponding generic term can be extracted.
[0099] In some embodiments, surname nodes can be introduced into the relationship graph construction, and the entity names associated with those surnames can be used as nodes, establishing edge connections with the surname nodes based on the "belonging to the same surname" relationship. Then, based on the above, each entity name is used as a node, and edge connections are established according to the associations between entity names confirmed in the initial name grouping. Further, the edges in the relationship graph are filtered according to the entity names in the generic reference list, thus forming a comprehensive graph structure that includes surname connections and associated entity name connections, while filtering out generic references. Connectivity analysis is performed on this comprehensive graph structure to obtain various maximally connected subgraphs, which are then used as associated name groups. By transforming the relationships between entity names into a graph structure, automatic clustering of different forms of entity names is achieved through multi-dimensional relationship fusion and connected component analysis, thereby improving translation accuracy and realizing localized translation.
[0100] This specification also provides a device for translating entity names in text. Figure 6 This is an exemplary block diagram of an entity name translation device in text according to some embodiments of this specification. Figure 6 As shown, in some embodiments, the entity name translation device 600 in the text may include an acquisition module 610, an extraction module 620, a construction module 630, and a mapping module 640.
[0101] The acquisition module 610 is used to acquire the text to be processed.
[0102] Extraction module 620 is used to extract entity names from the text to be processed.
[0103] The construction module 630 is used to construct associated name groups by determining whether there is a correlation between the entity names, wherein the entity names in the same associated name group represent the same entity or different entities with the same characteristics.
[0104] The mapping module 640 is used to map at least a portion of the entity names in the entity names from the source language to the target language based on the associated name grouping.
[0105] In some optional embodiments, the construction module 630 can also be used to determine, based on entity names and text to be processed, clue information to characterize whether there is a relationship between two or more entity names; and to determine, based on the clue information, whether there is a relationship between two or more entity names.
[0106] In some optional embodiments, the clue information includes: a clue sentence for characterizing whether two or more entity names represent the same entity. The construction module 630 can also be used to determine whether there is a correlation between two or more entity names based on the clue sentence.
[0107] In some alternative embodiments, the construction module 630 may also be used to determine the context associated with the clue sentence based on the text to be processed; and to determine whether there is a relationship between two or more entity names based on the context associated with the clue sentence.
[0108] In some optional embodiments, the construction module 630 may also be used to determine the context associated with the clue sentences based on the text to be processed in response to the clue sentences corresponding to two or more entity names including clue sentences that indicate that there is no association between the two or more entity names.
[0109] In some optional embodiments, the clue information includes: context for characterizing whether two or more entity names represent the same entity; the construction module 630 can also be used to determine whether there is a relationship between two or more entity names based on the context.
[0110] In some optional embodiments, the construction module 630 can also be used to construct an initial name group by determining whether there is a correlation between entity names; extract the surnames from the initial name group; and process the initial name group based on the surnames to obtain a related name group.
[0111] In some alternative embodiments, the construction module 630 can also be used to determine generic terms based on entity names in the text to be processed; and to filter generic terms from the initial name group to obtain associated name groups.
[0112] In some optional embodiments, the mapping module 640 can also be used to map surnames from a source language to a target language to obtain a surname mapping table; and based on the surname mapping table and the associated name grouping, to map entity names in the associated name group from the source language to the target language to obtain an entity name mapping table.
[0113] In some optional embodiments, the extraction module 620 can also be used to extract entity names from the text to be processed by an entity name extraction model, the entity name extraction model being trained based on a training sample set, the training sample set including a corpus labeled with term categories, the corpus being obtained based on a classification template, the classification template including terms and the context containing the terms.
[0114] For more information on each module, please refer to [link / reference]. Figures 1-5 The relevant explanations will not be repeated here. It should be understood that... Figure 6The apparatus and modules shown can be implemented in various ways. For example, in some embodiments, the apparatus and modules can be implemented in hardware, software, or a combination of both. The hardware portion can be implemented using dedicated logic; the software portion can be stored in memory and executed by an appropriate instruction execution device, such as a microprocessor or dedicated-design hardware. Those skilled in the art will understand that the methods and apparatus described above can be implemented using computer-executable instructions and / or included in the control code of a processor, such as in the memory of a disk, CD, or DVD-ROM. The apparatus and modules described in this specification can be implemented not only by hardware circuitry such as very large-scale integrated circuits or gate arrays, semiconductors such as logic chips or transistors, or programmable hardware devices such as field-programmable gate arrays or programmable logic devices, but also by software, for example, executed by various types of processors, or by a combination of the aforementioned hardware circuitry and software (e.g., firmware).
[0115] It should be noted that the above description of the device and its modules is for convenience only and should not be construed as limiting this specification to the embodiments described. It is understood that those skilled in the art, after understanding the principle of the device, may arbitrarily combine the various modules without departing from this principle to form sub-devices connected to other modules. Alternatively, some modules may be split to obtain more modules or multiple units under a single module. Such modifications are all within the scope of this specification.
[0116] Some embodiments of this specification also provide a computer device, which includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it can implement this specification. Figures 1-5 The method shown.
[0117] Some embodiments of this specification also provide a computer program product, including computer instructions that, when at least a portion of the computer instructions are executed by a processor, can implement this specification. Figures 1-5 The method is illustrated. In some embodiments, the computer program product may relate only to computer instructions, which may be carried on a storage medium or processing device. In other embodiments, the computer program product may also be a storage medium or processing device containing the aforementioned computer instructions. The processing device may include one or more processors, and the storage medium.
[0118] In some embodiments, the processor may be a combination of one or more of the following processors: central processing unit (CPU), application-specific integrated circuit (ASIC), application-specific instruction set processor (ASIP), graphics processing unit (GPU), physical processing unit (PPU), digital signal processor (DSP), field-programmable gate array (FPGA), programmable logic device (PLD), programmable logic controller (PLC), reduced instruction set computer (RISC), and microprocessor.
[0119] In some embodiments, the storage medium may include one or more combinations of the following: mass storage, removable storage, volatile read-write memory, and read-only memory (ROM). Exemplary mass storage may include disks, optical disks, solid-state drives, etc. Exemplary removable storage may include flash drives, floppy disks, optical disks, memory cards, compressed hard disks, magnetic tapes, etc. Exemplary volatile read-write memory may include random access memory (RAM). Exemplary RAM may include dynamic random access memory (DRAM), dual data rate synchronous dynamic random access memory (DDRSDRAM), static random access memory (SRAM), silicon controlled retrieval memory (T-RAM), and zero-capacitance memory (Z-RAM), etc. Exemplary read-only memory may include masked read-only memory (MROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), compressed hard disk read-only memory (CD-ROM), and digital multifunction hard disk read-only memory, etc.
[0120] The beneficial effects that the embodiments of this specification may bring include, but are not limited to: (1) by obtaining the text to be processed, extracting the entity names in the text to be processed, constructing associated name groups by determining whether there is a correlation between the entity names, and mapping at least some of the entity names in the entity names from the source language to the target language based on the associated name groups. Since the entity names in the same associated name group represent the same entity or different entities with the same characteristics, language mapping can be performed based on the relationship between the entity names in the associated name group, thereby accurately identifying that these different forms of personal names belong to the same person and avoiding confusion of the identity of the person in the translation; (2) after determining whether there is a correlation between two or more entity names based on the clue sentence, a second judgment can be made based on the context associated with the clue sentence to further determine the correlation between two or more entity names. The problem of identifying entity names in different forms is decomposed into two independent but related stages: "clue sentence extraction" and "relationship disambiguation". The complex task is simplified by the data flow between the stages, and the judgment of the relationship between entity names is further improved. (3) By transforming the relationship between entity names into a graph structure, the automatic clustering of entity names in different forms is achieved through multi-dimensional relationship fusion and connected component analysis, thereby improving the accuracy of translation and realizing localized translation. (4) By performing surname mapping, constructing associated name groups and verifying the entity name mapping table, a three-layer localization architecture is formed, which solves the consistency problem of batch translation of large-scale entity names. The task scale is decomposed by grouping, which solves the problem of too many names in ultra-long texts and realizes the localized translation of personal names. It should be noted that different embodiments may produce different beneficial effects. In different embodiments, the beneficial effects may be any one or a combination of the above, or any other possible beneficial effects.
[0121] The basic concepts have been described above. It is obvious that the detailed disclosure above is merely illustrative and does not constitute a limitation of this specification. Although not explicitly stated herein, various modifications, improvements, and corrections may be made to this specification by those skilled in the art. Such modifications, improvements, and corrections are taught in this specification and therefore remain within the spirit and scope of the exemplary embodiments described herein.
Claims
1. A method for translating entity names in text, characterized in that, The method includes: Get the text to be processed; Extract entity names from the text to be processed; By determining whether there is a correlation between the entity names, associated name groups are constructed. Entity names in the same associated name group represent the same entity or different entities with the same characteristics; and Based on the associated name grouping, at least a portion of the entity names in the entity names are mapped from the source language to the target language.
2. The method according to claim 1, characterized in that, The step of constructing associated name groups by determining whether there is a correlation between the entity names includes: Based on the entity names and the text to be processed, determine clues to characterize whether there is a relationship between two or more entity names; and Based on the clue information, determine whether there is a correlation between the two or more entity names.
3. The method according to claim 2, characterized in that, The clue information includes: Clue sentences used to indicate whether two or more entity names represent the same entity. Determining whether there is a correlation between the two or more entity names based on the aforementioned clue information includes: Based on the clue sentence, determine whether there is a correlation between the two or more entity names.
4. The method according to claim 3, characterized in that, Determining whether there is a correlation between the two or more entity names based on the clue sentence includes: Determine the context associated with the clue sentence based on the text to be processed; and The existence of a relationship between the two or more entity names is determined based on the context associated with the clue sentence.
5. The method according to claim 4, characterized in that, The step of determining the context associated with the clue sentence based on the text to be processed includes: In response to the clue sentences corresponding to the two or more entity names including the clue sentences indicating that there is no correlation between the two or more entity names, the context associated with the clue sentences is determined based on the text to be processed.
6. The method according to claim 2, characterized in that, The clue information includes: Context used to indicate whether two or more entity names represent the same entity; and Determining whether there is a correlation between the two or more entity names based on the aforementioned clue information includes: Determine whether there is a relationship between the two or more entity names based on the context.
7. The method according to claim 1, characterized in that, The step of constructing associated name groups by determining whether there is a correlation between the entity names includes: Initial name groups are constructed by determining whether there are relationships between entity names; Extract the surname from the initial name group; The initial name group is processed based on the surname to obtain the associated name group.
8. The method according to claim 7, characterized in that, The process of processing the initial name group based on the surname to obtain the associated name group further includes: Determine generic terms based on entity names in the text to be processed; and The associated name group is obtained by filtering the generic terms from the initial name group.
9. The method according to claim 7 or 8, characterized in that, The step of mapping entity names in the associated name group from the source language to the target language includes: Map the surname from the source language to the target language to obtain a surname mapping table; Based on the surname mapping table and the associated name grouping, the entity names in the associated name grouping are mapped from the source language to the target language to obtain the entity name mapping table.
10. The method according to claim 1, characterized in that, The extraction of entity names from the text to be processed includes: The entity names in the text to be processed are extracted by an entity name extraction model, which is trained on a training sample set, including a corpus labeled with term categories. The corpus is obtained based on a classification template, which includes terms and the context containing the terms.
11. A device for translating entity names in text, characterized in that, The device includes: The acquisition module is used to acquire the text to be processed; The extraction module is used to extract entity names from the text to be processed; The construction module is used to construct associated name groups by determining whether there is a correlation between the entity names, wherein the entity names in the same associated name group represent the same entity or different entities with the same characteristics; The mapping module is used to map at least a portion of the entity names in the entity names from the source language to the target language based on the associated name grouping.
12. A computer device, characterized in that, The computer device includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, it is able to implement the method as described in any one of claims 1 to 10.
13. A computer program product, characterized in that, It includes a computer program that, when at least a portion of the computer program is executed by a processor, enables the implementation of the method as described in any one of claims 1 to 10.
Citation Information
Patent Citations
Academic institution name entity alignment method based on text features
CN112016328A
Knowledge graph generation method and device, terminal and storage medium
CN112836057A
Text data-oriented alias entity rapid identification method and system
CN116720520A
Translation semantic correction method and system based on large language model
CN120258014A
Semantic compatibility checking for automatic correction and discovery of named entities
US20090204596A1