Translation semantic correction method and system based on large language model

Through the translation semantic bias correction method based on the large language model, the knowledge graph and user feedback optimized translation is used to solve the semantic mistranslation problem of machine translation in complex contexts, and high accuracy and adaptive translation effects are achieved.

CN120258014APending Publication Date: 2025-07-04HANGZHOU SHUNSHUN ZHIXING TECH CO LTD

Patent Information

Application Number
CN202510732964.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-04
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

Existing machine translation methods are difficult to accurately understand semantics in complex contexts, especially when dealing with proper nouns, culturally specific expressions and polysemes, and lack the integration of dynamic knowledge, making it difficult to adapt to the translation needs of professional fields.

Method used

The translation semantic deviation correction method based on large language models is used to identify named entities, construct knowledge graphs, compare translation results and graphs, conduct quality evaluation, dynamically optimize translation results, and iteratively optimize them in combination with user feedback.

Benefits of technology

It improves the accuracy and coherence of translation, adapts to complex contexts, enhances the adaptability and intelligence level of the translation system, reduces mistranslation and bias, and improves the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120258014A_ABST
    Figure CN120258014A_ABST
Patent Text Reader

Abstract

The invention provides a translation semantic correction method and system based on a large language model. Belongs to the technical field of natural language processing. The method comprises the following steps: collecting data; performing named entity recognition by using a large language model, extracting specific entities from a source language text, and analyzing a relationship between the entities through LLM to generate an entity pair relationship; identifying and extracting attributes of the entities, and associating the attributes with the corresponding entities; integrating the extracted entities, relationships and attributes according to a graph structure, and constructing a preliminary knowledge graph; inputting the source language text into the model for translation; terms in the preliminary translation result are compared with the knowledge graph, and potential semantic deviation is recognized. And performing quality evaluation on the translation result. By integrating the identified and extracted relationship and attribute between the entities into the knowledge graph, the semantic relationship between the entities can be understood more deeply, so that the subsequent translation is more in line with the actual context.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention provides a method and system for correcting translation semantics based on a large language model, belonging to the technical field of natural language processing. Background Art

[0002] Traditional machine translation methods mainly rely on statistical models or neural machine translation (NMT) frameworks to achieve the mapping from the source language to the target language through training with large-scale parallel corpora. However, such methods have limited semantic understanding capabilities in complex contexts. Especially when dealing with proper nouns, culture-specific expressions, and polysemous words, semantic deviations are likely to occur due to the lack of context association or insufficient domain knowledge. For example, entities such as personal names, place names, and organization names may cause mistranslations or even logical contradictions during cross-language conversion if domain knowledge is not combined. In addition, the existing technology lacks support for integrating dynamic knowledge and is difficult to adapt to the translation requirements of updated professional terms or specific fields (such as academic papers, news). Summary of the Invention

[0003] The present invention provides a method and system for correcting translation semantics based on a large language model to solve the problems mentioned in the above background art: A method for correcting translation semantics based on a large language model proposed by the present invention, the method includes: S1. Collect data; S2. Use a large language model for named entity recognition, extract specific entities from the source language text, analyze the relationships between entities through the large language model to generate entity pair relationships; identify and extract the attributes of the entities, and associate these attributes with the corresponding entities; integrate the extracted entities, relationships, and attributes according to the graph structure to construct a preliminary knowledge graph; S3. Input the source language text into the model for translation; S4. Compare the terms in the preliminary translation result with the knowledge graph to identify potential semantic deviations; S5. Evaluate the quality of the translation result.

[0004] The translation semantics correction system based on a large language model proposed by the present invention includes a memory, a processor, and a computer program stored on the memory and executable on the processor. The processor executes the program to implement the method for correcting translation semantics based on a large language model as described in any one of the above.

[0005] Advantages of the present invention: The technical solution proposed by the present invention does not need to rely on predefined rules or simple pattern matching, can more accurately identify entities and their relationships in complex contexts, and can more deeply understand the semantics and context of the text, avoiding translation deviations caused by inaccurate entity recognition or incorrect relationship understanding, thereby significantly improving the accuracy of translation. By integrating the relationships and attributes between the identified and extracted entities into the knowledge graph, the semantic connections between entities can be more deeply understood, making the subsequent translation more in line with the actual context; by analyzing the relationships and attributes between entities, the model can better understand the context of the source language text and reduce translation errors caused by unclear context; thus improving the coherence and professionalism of translation. Brief Description of the Drawings

[0006] Figure 1 It is a flowchart of the method described in the present invention. Detailed Embodiments

[0007] The following describes the preferred embodiments of the present invention with reference to the accompanying drawings. It should be understood that the preferred embodiments described herein are only used to illustrate and explain the present invention, and are not used to limit the present invention.

[0008] An embodiment of the present invention, as Figure 1 shown, a translation semantic deviation correction method based on a large language model, the method includes: S1. Collect bilingual text data of the source language text and the target language text through multiple channels, where the multiple channels include books, news, academic papers, and websites; and preprocess the collected data; S2. Use a large language model (LLM) for named entity recognition (NER) to extract specific entities from the source language text, where the specific entities include personal names, locations, and organizations; analyze the relationships between entities through the LLM to generate entity pair relationships (EPR); identify and extract the attributes of entities, such as the age of a person, the population of a place, etc., and associate these attributes with the corresponding entities; integrate the extracted entities, relationships, and attributes according to the graph structure to construct a preliminary knowledge graph; S3. Use the large language model as a translation engine to input the source language text into the model for translation; and perform preliminary optimization on the translation result, where the preliminary optimization includes grammar correction and vocabulary replacement; S4. Compare the terms in the preliminary translation result with the knowledge graph to identify potential semantic deviations; dynamically optimize the preliminary translation result according to the accurate translation and context information in the knowledge graph; If there is a translation in the knowledge graph that exactly matches the term, directly replace it; If there is a translation in the knowledge graph that partially matches the term, then reason based on the context information and select the most appropriate translation; If there is no translation in the knowledge graph that matches the term, then use the reasoning ability of the large language model for translation and add it to the knowledge graph for subsequent use; Conduct an overall inspection of the optimized translation results to ensure semantic coherence and natural expression.

[0009] S5. Use a translation quality assessment model based on deep learning to evaluate the quality of the optimized translation results; according to the evaluation results, continuously iterate and optimize the translation algorithm; collect user feedback to understand users' satisfaction with the translation quality and improvement suggestions; integrate user feedback into the optimization process of the translation algorithm to continuously improve the translation quality; and output the final translation results.

[0010] The working principle and effect of the above technical solution are as follows: By integrating the translation engine of the large language model and the knowledge graph, semantic deviations in translation can be effectively corrected to ensure that the translation results are more accurate and natural; at the same time, the accuracy of translation is improved and the naturalness of translation is enhanced; the knowledge graph provides background information and entity relationship support for translation, and can avoid common mistranslations and deviations when dealing with complex texts, reducing mistranslations and deviations, thus saving time and improving the overall efficiency; the system can continuously obtain feedback from new translation cases and improve the knowledge graph, realizing the self-learning and evolution of translation, improving the adaptive ability, accuracy and intelligence level of the translation system, and improving the quality, flexibility and user experience of translation. With the accumulation of the knowledge graph, the translation quality will gradually improve, enhancing the long-term application value of the system; the system can make appropriate translation decisions according to the specific context through the context analysis of terms, avoiding simple word-for-word translation, and improving the flexibility and accuracy of translation; through the user feedback mechanism, the system can adapt to the needs of different users, continuously adjust and optimize the translation algorithm, making the translation results more and more in line with the actual usage needs, and improving user satisfaction; this method can adapt to translation tasks in various different fields and languages, especially in the translation of named entities and professional terms, and can provide more accurate translation results.

[0011] In one embodiment of the present invention, the S1 includes: S11. Collect bilingual text data of the source language text and the target language text from multiple channels (such as books, news, academic papers, professional websites, and social media) to ensure the diversity and richness of the data sources; and conduct key data collection for specific fields (such as science and technology, law, medicine, etc.) to improve the professionalism and accuracy of the translation results; ensure that the collected data is timely, reflecting the current language usage habits and the development dynamics of the field; S12. Preprocess the real-time collected data, and the preprocessing includes: Data cleaning: Remove noises such as advertisements, junk information, and irrelevant links in the text, and retain pure bilingual text data.

[0012] Text standardization: Unify the text format, such as punctuation marks, capitalization, spaces, etc., to ensure the consistency of the data.

[0013] Word segmentation and part-of-speech tagging: Use natural language processing technology to segment the text and tag the part of speech, providing a basis for subsequent steps.

[0014] Syntactic analysis and semantic understanding: Perform syntactic analysis and semantic understanding on the text, extract sentence structures, grammatical relationships, and semantic information, providing support for subsequent translation and semantic correction.

[0015] The working principle and effect of the above technical solution are as follows: By collecting bilingual texts from different fields and multiple channels, the translation system can comprehensively cover various language expressions, improving the system's adaptability to various language phenomena and fields; The process of data cleaning and standardization removes messy information, making the input data cleaner, providing a high-quality data source for model training, and indirectly improving the accuracy of model translation. The standardization process ensures that texts in different formats can be processed in the same model, reducing errors caused by format differences; Through word segmentation, part-of-speech tagging, and syntactic analysis, the system can better understand the text structure and semantics, enhancing the intelligence of the translation process; Especially in the case of long sentences or complex structures, syntactic analysis and semantic understanding can ensure the correct recognition of each sentence component during translation, thus avoiding simple word-for-word translation and improving the translation quality; The data collection for specific fields (such as technology, law, medicine, etc.) effectively improves the professionalism and accuracy of the translation system in these fields. The system can identify and correctly translate professional terms and industry-specific language expressions, reducing translation errors caused by misunderstanding domain vocabulary or terms; The guarantee of timeliness enables the translation system to handle current language habits and the latest domain dynamics, ensuring that the translation results conform to the current language usage trends and avoiding outdated expressions.

[0016] In one embodiment of the present invention, the S2 includes: S21. Use a large language model to perform named entity recognition, extract key entities, and the key entities include personal names, locations, and organizations; and verify the identified key entities to remove duplicate entities, ensuring the uniqueness and accuracy of the entities; S22. Classify the entities according to the entity type (such as personal names, locations, organizations) and add tags to them; S23. Analyze the relationships between entities through LLM, extract entity-pair relationships (EPR), such as "someone was born in a certain place"; verify the extracted relationships and organize the relationships to form a structured relationship network; S24. Identify the attributes of entities based on LLM and rule matching techniques, such as age, population, and area; verify the identified attributes and associate the attributes with the corresponding entities to form attribute-entity pairs; S25. Integrate the extracted entities, relationships, and attributes according to the graph structure to construct a preliminary knowledge graph; optimize the preliminarily constructed knowledge graph, including removing redundant information, supplementing missing relationships, and adjusting entity hierarchies; S26. Store the optimized knowledge graph in a database and establish an efficient indexing mechanism.

[0017] The working principle and effects of the above technical solution are as follows: Through the efficient recognition of the LLM, named entities can be accurately extracted from the text, laying a good foundation for the subsequent construction of the knowledge graph, improving the accuracy and efficiency of named entity recognition, and providing a reliable basis for the construction of the knowledge graph; The deduplication and verification processes ensure the uniqueness of each entity in the graph, avoiding information redundancy caused by entity duplication; Using the LLM for named entity recognition can better handle polysemous words, homonyms, and ambiguous entities in the context compared to traditional methods, improving the accuracy and comprehensiveness of recognition; The addition of entity classification and labels helps with the subsequent recognition, query, and management of entities, enhancing the efficiency and accuracy of queries. For example, all entities of the "location" type can be quickly filtered out in the database, supporting flexible queries and retrievals. By adding labels to entities, the data structure can be made clearer, providing standardized information for the construction of the knowledge graph and facilitating subsequent processing and optimization; Using the deep semantic analysis of the LLM, the system can identify complex entity relationships in the text, which is more accurate and efficient than traditional relationship extraction methods; The extracted relationships provide rich information for the construction of the knowledge graph, enabling a better description of the interconnections between entities; It improves the accuracy and efficiency of entity relationship extraction, accelerates the data processing process, and also enhances the applicability of the knowledge graph in multi-domain applications; Verifying the relationships ensures that only accurate and reasonable relationships are in the graph, avoiding incorrect or unreasonable relationships from being incorporated into the knowledge graph and guaranteeing the credibility and practicality of the graph; Combining the LLM with rule matching technology can efficiently extract various attributes of relevant entities from the text, and the verification process ensures the accuracy of the extracted attributes, preventing incorrect or incomplete attribute information from entering the graph; Associating the attributes with the entities can provide richer descriptions of the entities, and these attributes can be queried and analyzed in the knowledge graph, providing knowledge support in more dimensions, enhancing the accuracy of entity descriptions, and also increasing the diversity of queries, analyses, inferences, and applications. Through continuous optimization, the finally formed knowledge graph has high accuracy, integrity, and scalability, and can provide support for various application scenarios (such as intelligent question answering, recommendation systems, etc.); Redundant information is removed during the optimization process, making the graph more concise and refined; Missing relationships are supplemented to ensure the comprehensiveness of the graph's information; The entity hierarchy is adjusted to make the graph structure more reasonable, facilitating subsequent analysis and expansion. An efficient storage structure and indexing mechanism are adopted to ensure that the knowledge graph can be quickly accessed and retrieved in a large-scale data environment.

[0018] In one embodiment of the present invention, S23 includes: S231. Based on the deep understanding ability of the large language model (LLM), perform context analysis on the text containing entities; Through the context information, preliminarily judge the possible associations or relationships between entities; S232. Based on the learning ability of the LLM, mine common entity relationship patterns from historical data or large-scale corpora, including but not limited to "someone was born in a certain place", "an organization is located in a certain place", "someone holds a position in an organization", etc.; through pattern matching, locate the possible entity relationships in the text. S233. According to the diversity of entity types and relationship patterns, formulate targeted relationship extraction strategies. For example, for person name and location entities, adopt an event extraction strategy based on the timeline; for person name and organization entities, adopt an extraction strategy based on job descriptions; utilize the semantic parsing ability of the LLM to deeply parse the text and accurately extract the relationships between entities; evaluate the confidence of the extracted relationships, and evaluate the reliability and accuracy of the relationships by calculating indicators such as the occurrence frequency of the relationship in the text and the context support degree. S234. Use texts or data from different sources to cross-validate the extracted relationships. For example, for the relationship "someone was born in a certain place", it can be verified by comparing data such as the person's birth certificate and household registration information; remove duplicates for duplicate or redundant relationships, and merge relationships with the same or similar meanings. S235. According to the type and importance of the relationships, construct a relationship hierarchy. For example, regard "father-son relationship", "husband-wife relationship", etc. as the basic relationship hierarchy, and regard "superior-subordinate relationship", "master-apprentice relationship", etc. as the extended relationship hierarchy. Through hierarchy construction, the relationship network between entities can be better understood and displayed; display the extracted and verified relationships in a graphical way to form a structured relationship network; through visualization tools, display the information, including intuitively displaying the relationships between entities, the hierarchy and importance of the relationships, etc.; further optimize the relationship network, including removing noisy relationships, supplementing missing relationships, adjusting the relationship hierarchy, etc. Through optimization, the accuracy and integrity of the relationship network can be further improved.

[0019] S236. Store the optimized relationship network in a database, and establish an indexing mechanism for the relationship network database. The indexing mechanism can include but not be limited to indexing based on entity names, indexing based on relationship types, etc.

[0020] The working principle and effects of the above technical solution are as follows: Through the deep understanding ability based on the large language model (LLM), it can effectively and accurately identify the relationships between different entities from the text, improving the accuracy of text information extraction and enhancing the multi-language processing ability; By using the semantic parsing ability of the LLM, it can extract common entity relationship patterns from a large amount of historical data or corpus; Through pattern matching, not only can the accuracy of relationship extraction be improved, but also the most appropriate extraction strategy can be selected according to the different types of entities, improving the adaptability and flexibility of the system; Through the confidence evaluation mechanism, the reliability of the extracted relationships can be evaluated according to indicators such as the occurrence frequency and context support degree in the text. At the same time, through cross-validation and redundancy removal, the accuracy and consistency of the relationships are ensured. By merging duplicate or similar relationships, the integration and accuracy of the data are further improved; Laying out the entity relationships in layers helps users to more clearly understand the complex network structure between entities; The relationship network not only supports the basic relationship level, but also covers more complex extended relationship levels, helping users to deeply explore the internal connections between entities from different dimensions, improving the clarity of user understanding and the visualization effect of the relationship network. The introduction of graphical display and visualization tools makes the presentation of information more intuitive and understandable, further improving the user experience; The optimized relationship network is stored in the database, and the indexing mechanism ensures the efficiency of retrieval. Different types of indexes (such as indexes based on entity names or relationship types) guarantee fast searching and data access, support the efficient management of large-scale data, accelerate the analysis and exploration of complex relationship networks, and optimize the utilization efficiency of storage resources; By further optimizing and supplementing the relationship network, the relationship network between entities can be continuously improved to make it more accurate and complete.

[0021] In one embodiment of the present invention, S231 includes: Preprocess the text containing entities, where the preprocessing includes word segmentation, part-of-speech tagging, and named entity recognition; Using the deep understanding ability of the LLM, embed the entities and their context information into a high-dimensional vector space; The vectors can capture the semantic features and context relationships of the entities in the text, providing a basis for subsequent relationship judgment. For each entity in the text, calculate its contextual similarity with other entities through cosine similarity; Entity pairs with high similarity may have associations or relationships; Based on historical data or large-scale corpus, mine common entity association rules; The association rules are used to reflect common association patterns between entities, such as "person name - location" (place of birth), "person name - organization" (affiliated institution), etc.

[0022] Match the entity pairs in the text with the mined association rules to preliminarily judge the possible associations or relationships between entities; at the same time, utilize the prediction ability of the LLM to predict the relationships of entity pairs that have not been matched with rules but have similar contexts; Consider the dynamics of the context, that is, the relationships between entities may change with the change of the text content; utilize the time series analysis ability of the LLM to capture the time clues in the context and make a preliminary judgment on the dynamic changes of entity relationships; Combine the context information in multiple texts or data sources. By constructing a cross-context entity relationship graph, use the path and connection information in the graph to reason about the cross-context relationships between entities; Use semantic role labeling technology to identify the semantic roles of entities and words in the text, such as agents, patients, tools, etc.; based on factors such as context similarity, association rule matching degree, time clues, etc., evaluate the confidence of the preliminarily judged entity relationships; relationships with high confidence are more likely to be verified and retained in subsequent steps; Sort the preliminarily judged entity relationships according to the confidence and importance of the relationships, and set the priorities of different relationships so as to process and verify them in sequence in subsequent steps.

[0023] The working principle and effects of the above technical solution are as follows: By preprocessing and deeply understanding the entities in the text, this application can more accurately identify the entities and their semantic features in the context. Utilizing the deep understanding ability and high-dimensional vector space embedding of the LLM, it can precisely capture the potential relationships between entities, conduct efficient analysis, enhance the context sensitivity, improve the cross-domain adaptability, and deeply understand the semantic connections between entities. Traditional methods mainly rely on rules, dictionaries, fixed word vectors, or statistical models, unable to dynamically capture context information, and usually lacking in-depth semantic understanding and reasoning capabilities, resulting in insufficient understanding of the semantic relationships between entities; Based on cosine similarity calculation and historical data mining, it can identify the potential associations between entities, especially for entity pairs that do not match the rules. Through the prediction ability of the LLM, relationship inference is carried out, improving the system's processing ability for complex texts and unknown entity relationships; Considering that entity relationships may change with the text content, time series analysis is used to track and judge the dynamic changes of entity relationships; It can reflect the changes in the relationships between entities at different times and in different contexts, making relationship extraction more in line with the actual context, enhancing the flexibility of the system. Traditional steps mainly focus on static entity recognition and matching, as well as similarity calculation based on simple rules or algorithms, lacking dynamic recognition and matching, which will lead to insufficient and inaccurate mining; By constructing a cross-context entity relationship graph, information from different texts or data sources can be combined to conduct cross-context relationship reasoning. The introduction of the graph enables the system to process more complex multi-source information and infer potential connections between entities in multiple dimensions, improving the depth of data integration and analysis, enhancing the data integration ability, and improving the efficiency of entity recognition and relationship extraction; Through semantic role labeling technology, in-depth parsing of entities and vocabulary in the text is carried out to further clarify the specific relationships between entities (such as agent, patient, tool, etc.), thereby improving the accuracy of relationship recognition, making the relationships in the text not only simple entity pairs but also rich in semantic information, enhancing the depth of text understanding, improving data quality and credibility; By conducting confidence evaluation on the initially judged entity relationships, it can ensure that relationships with high confidence are preferentially retained, improving the accuracy of relationship judgment. Based on a multi-factor confidence evaluation mechanism, the interference of incorrect relationships is effectively reduced, making the final result more reliable; Sorting and setting priorities for entity relationships can ensure that the most important relationships are focused on in subsequent processing, optimizing the processing flow and system efficiency, avoiding unnecessary calculations and repetitive work, and improving the overall efficiency. The priority sorting of traditional methods usually relies on simple statistics or rule bases, lacking flexibility and depth, and only sorting based on the results of preliminary judgments, resulting in the inability to conduct in-depth comprehensive sorting through multiple factors.

[0024] In one embodiment of the present invention, the S24 includes: Preprocess the text or data source containing entities, including text cleaning, word segmentation, part-of-speech tagging, etc.; construct an attribute dictionary based on domain knowledge and common attributes; the dictionary contains information such as the name, type, and possible value range of the attributes, providing a basis for subsequent attribute recognition; Utilize the semantic understanding and generation capabilities of a large language model (LLM) to extract attributes from the preprocessed text; among them, the LLM can identify implicit attribute information in the text, such as the "population" attribute in the sentence "The population of this city is approximately XX million"; combine rule matching technology, and use the predefined attribute dictionary and rule base to precisely match the attributes in the text; the rule base contains information such as attribute names and value patterns; Perform fusion processing on multiple attribute descriptions that may exist for the same entity. If there are conflicts between attributes, make a decision based on factors such as context information and the reliability of the attribute source; Verify the identified attributes, including value range, type matching, etc.; correct or remove attribute values that do not meet expectations; associate the verified attributes with the corresponding entities to form attribute-entity pairs; For key entities, check whether they have all necessary attributes; if some important attributes are missing, try to supplement this information from other data sources or texts; Standardize the identified attributes, such as unifying units, formats, etc. Construct an attribute hierarchy based on the importance and relevance of the attributes; for example, distinguish between basic attributes (such as name, age) and extended attributes (such as occupation, hobbies) to better display and query in the knowledge graph; and remove redundant attribute-entity pairs, such as duplicate attribute descriptions or insignificant attribute information.

[0025] The working principle and effects of the above technical solution are as follows: By combining domain knowledge, common attributes, a predefined attribute dictionary, and rule matching technology, it is possible to accurately identify the attributes of relevant entities in the text, reducing the ambiguity and errors in attribute recognition and improving the accuracy of recognition; By performing preprocessing steps such as text cleaning, word segmentation, and part-of-speech tagging on the text, it is possible to ensure that the subsequent attribute extraction process is more efficient, while reducing the interference of noise information and improving the processing efficiency; For the fusion processing of multiple attribute descriptions of the same entity, it is possible to ensure effective adjudication through context information and the reliability of attribute sources in the event of attribute conflicts, avoiding the problems of attribute conflicts or omissions, thereby providing more accurate attribute information, improving the quality of data and the robustness of the system, and also providing stronger support for various intelligent systems (such as decision support, personalized recommendation, cross-domain data integration, etc.); By verifying the value range and type matching of attribute values, it is possible to effectively identify and correct attribute information that does not meet expectations, ensuring that each attribute in the knowledge graph meets the actual requirements, enhancing the accuracy and credibility of the graph; For the missing attributes of key entities, by attempting to supplement information from other data sources or texts, it is ensured that the attributes of entities in the knowledge graph are as complete as possible, providing more comprehensive data support for subsequent applications; Unifying the units and formats of attributes and constructing an attribute hierarchy based on the importance and relevance of attributes helps to better manage and display attribute information, improving the readability, query efficiency, and usage experience of the knowledge graph; By removing redundant attribute-entity pairs, irrelevant information is reduced, the knowledge graph is refined, the structure of the graph is optimized, making queries clearer and more intuitive, and the data quality is improved; This solution can be customized for specific domains by constructing a dictionary of domain knowledge and common attributes, enabling the technical solution to not only adapt to the data extraction requirements of different domains but also have strong scalability and be able to play a role in a variety of application scenarios; By integrating technologies such as attribute recognition, standardization, verification, and redundancy processing, the quality of the generated knowledge graph is ultimately improved, making it more accurate and reliable in subsequent analysis, query, and reasoning processes, further enhancing the potential of data mining and artificial intelligence applications.

[0026] In one embodiment of the present invention, S25 includes: Integrate the entity, relationship, and attribute data extracted in steps S21 to S24, and perform data cleaning to remove invalid, duplicate, or redundant information; According to the construction goal and application scenario of the knowledge graph, design the graph structure, including the representation methods of entity nodes, relationship edges, and attribute information, as well as the hierarchical structure and storage method of the graph, etc.; Add the cleaned entity data to the graph as nodes of the graph; During the addition process, the uniqueness and accuracy of the entities need to be considered to ensure that each entity corresponds to only one node in the graph; Connect the entity nodes through relationship edges according to the extracted relationship data; during the connection process, it is necessary to verify the accuracy and rationality of the relationships to ensure that the relationship edges can correctly reflect the associations between entities; associate the extracted attribute information with the corresponding entity nodes as additional information of the nodes; during the addition process, it is necessary to ensure the accuracy and integrity of the attribute information and the corresponding relationship between the attributes and the entities. Identify and remove redundant information in the graph, including duplicate entity nodes, relationship edges, and attribute information, etc. By removing redundant information, the graph structure can be simplified, and the clarity and readability of the graph can be improved; according to the requirements of the integrity and accuracy of the graph, supplement the missing relationships in the graph by analyzing the relationship patterns between existing entity nodes or by combining external data sources for relationship reasoning and mining. Adjust the entity hierarchy in the graph according to the type and importance of the entities; for example, entities with greater global influence (such as well-known organizations, historical figures, etc.) can be placed at a higher level, and weight is assigned to the relationship edges in the graph by analyzing factors such as the occurrence frequency of the relationship edges and context information.

[0027] The working principle and effects of the above technical solution are as follows: By integrating and cleaning entity, relationship, and attribute data to remove invalid, duplicate, or redundant information, the quality and accuracy of the knowledge graph can be significantly improved. Data cleaning can ensure that there is no redundant data in the graph, guaranteeing the high quality and credibility of the information. By reasonably designing the structure of the graph, including the representation of nodes, relationship edges, and attribute information, as well as the hierarchical structure and storage method, the knowledge graph can be made more efficient, intuitive in storage, query, and reasoning, facilitating subsequent expansion and maintenance, reducing redundant information, improving storage efficiency, and lowering subsequent maintenance costs. By ensuring that each entity corresponds to only one node, data redundancy can be avoided, the accuracy of the graph can be ensured, the problem of entity duplication can be reduced, and subsequent knowledge reasoning and querying can be made more concise and fast. By verifying the accuracy and rationality of the relationship edges, ensuring that the relationships between entities can accurately reflect the associations in the real world, the knowledge graph can provide more effective data support for analysis and reasoning, improving the scalability of the knowledge graph. Associating the extracted attribute information with the entity nodes and ensuring its accuracy and integrity can ensure that each entity has a complete description, providing richer information support for subsequent reasoning and intelligent analysis. By removing redundant entity, relationship, and attribute information, the structure of the graph becomes more concise and clear, significantly improving the readability and query efficiency of the graph, enabling the knowledge graph to serve practical applications more efficiently, and enhancing the intelligent service ability of the knowledge graph. By integrating external data sources, the deficiencies of the data sources can be made up for, enhancing the adaptability of the graph to diverse scenarios. By adjusting the hierarchy according to the type and importance of entities, the key entities in the graph can be made more prominent, which is helpful for efficient querying and reasoning. Entities with greater global influence, such as well-known organizations and historical figures, can be preferentially displayed, improving the usability and intuitiveness of the graph. By assigning weights according to factors such as the frequency of occurrence of relationship edges and context information, the relationship structure in the graph can be further optimized, enhancing the reasoning efficiency of the graph, and ensuring that more important relationships are processed first, which is beneficial for subsequent intelligent analysis and applications.

[0028] In one embodiment of the present invention, the S3 includes: S31. Evaluate mainstream large language models, and select a model with strong translation ability and good domain adaptability as the translation engine; according to the requirements of a specific domain, perform customization processing on the translation engine, such as adding a domain dictionary, adjusting translation strategies, etc., to improve the professionalism and accuracy of translation; S32. Input the source language text into the translation engine for translation to generate a preliminary translation result; perform a grammar check on the preliminary translation result to correct grammar errors and ensure that the sentence is smooth and the structure is complete; S33. Replace and optimize the words in the translation result according to the expression habits and domain characteristics of the target language to improve the accuracy and fluency of the translation; S34. Adjust and optimize the paragraph and text structure of the translation result to ensure the overall semantic coherence and clear logic.

[0029] The working principle and effects of the above technical solution are as follows: By selecting a large language model with strong translation ability and good domain adaptability and performing customized processing, the accuracy and professionalism of the translation can be significantly improved. The customized processing enables the translation to better adapt to the terms and language habits of a specific field, thereby ensuring that the translation result meets the requirements of the target field; By performing grammar checking and correction on the preliminary translation result, grammar errors can be effectively eliminated, making the translation result more fluent and natural. This not only ensures the language quality of the translation but also improves the reading experience of the readers; By replacing and optimizing the words, the translation result is made more in line with the expression habits and domain characteristics of the target language. For professional fields, adjusting the terms and expression methods can make the translation more accurate and professional, thereby improving the trust and readability of the translation; By adjusting and optimizing the paragraph and text structure, the overall semantic coherence of the translation result is ensured and the logic is clear. This helps to avoid ambiguity or confusion caused by structural problems, thereby improving the translation quality while ensuring the effectiveness of information transmission; By carefully adjusting and optimizing each link in the translation process, it is ensured that the translation is not only accurate but also fluent and easy to understand. For translations involving complex structures or specific industry terms, fluency is particularly important, which can improve the readability and acceptance of the translation; The customized processing of the translation engine (such as adding a domain dictionary, adjusting the translation strategy) can enhance the adaptability of the translation engine to different types of texts. The terms and language structures in specific fields are better processed, avoiding translation mistakes or misunderstandings that may occur in a general translation engine; By continuously adjusting and optimizing the translation result, ensuring high-quality output of the translation, it can better serve the actual application scenarios. For example, in fields such as business, law, or medicine, high-quality translation can ensure accurate information transmission and avoid misunderstandings, thereby supporting more effective decision-making and communication.

[0030] In one embodiment of the present invention, the S4 includes: S41. Construct a professional term library according to the domain characteristics and translation requirements, including common terms, professional terms, abbreviations, etc.; Compare the terms in the preliminary translation result with the knowledge graph and the term library to identify potential semantic deviations; S42. Classify the identified semantic deviations, such as inaccurate term translation, context mismatch, semantic confusion, etc.; For terms that exactly match the knowledge graph or the term library, directly replace them to ensure the accuracy of term translation; S43. For partially matched terms, conduct reasoning and analysis by combining context information, and select the most appropriate translation for replacement; for terms not existing in the knowledge graph and the term base, utilize the reasoning ability of the large language model for translation, and add the new translation to the knowledge graph and the term base for subsequent use; S44. Conduct an overall semantic check on the optimized translation results to ensure semantic coherence and natural expression, while avoiding introducing new semantic deviations.

[0031] The working principle and effects of the above technical solution are as follows: By constructing a professional term base and combining a knowledge graph for term matching, the accuracy of term translation can be ensured. This not only avoids term translation errors but also improves the precision of translation in professional fields, reducing the risk of misunderstanding or misuse; By identifying and classifying semantic deviations in the preliminary translation results, potential translation problems such as inaccurate term translation, context mismatch, or semantic confusion can be discovered in a timely manner, and then targeted corrections can be made. This meticulous processing helps improve the translation quality; For partially matched terms, by conducting reasoning and analysis by combining context information, the translation can be flexibly adjusted to ensure that the term is more in line with the actual context. This context-driven translation optimization can make the translation results more natural, fluent, and understandable; For newly emerged terms, the translation can be carried out through the reasoning ability of the large language model and incorporated into the knowledge graph and the term base. This dynamic update mechanism ensures that the term base and the knowledge graph can continue to expand and improve to adapt to the ever-changing field requirements; During the translation process, through multi-level semantic checks and optimizations, the overall semantic coherence of the translation results can be guaranteed, avoiding logical loopholes or ambiguities in the translation and ensuring the accurate transmission of information. At the same time, the natural expression of the translation is also enhanced, making it more in line with the expression habits of the target language; Through automated term comparison, semantic reasoning, and context analysis, not only can the translation quality be ensured, but also the translation efficiency can be greatly improved. The precise control of terms and the reasonable reasoning of the context during the translation process enable high-quality translation to be completed in a relatively short time, thereby improving work efficiency; Since the solution involves the construction of a professional term base and a knowledge graph, it can adapt to the translation needs of a variety of different fields. With the continuous improvement of the term base and the knowledge graph, the system can flexibly handle various translation scenarios to ensure the accurate translation of various documents and materials.

[0032] In one embodiment of the present invention, the S41 includes: Deeply analyze the professional knowledge, industry standards, and common terms in the target field, including professional terms, common expressions, industry abbreviations, etc.; Through channels such as consulting professional literature, industry reports, and standard specifications, widely collect relevant terms to ensure the comprehensiveness and accuracy of the term base.

[0033] Based on the collected terms, construct a professional term library and standardize it, including unifying the spelling, definitions, translation rules, etc. of the terms to ensure the consistency and accuracy of the terms in different contexts. At the same time, establish an indexing mechanism for the term library to facilitate subsequent quick searching and comparison.

[0034] Integrate the term library with the existing knowledge graph, and utilize the semantic relationships and logical reasoning capabilities of the knowledge graph to enhance the functionality and practicality of the term library; through the knowledge graph, the associations and context relationships between terms can be further understood, providing strong support for subsequent translation and semantic checking; Compare the preliminary translation results with the term library and the knowledge graph to identify potential semantic deviations, including problems such as inaccurate term translation, context mismatch, semantic confusion, etc.; through comparison, promptly discover and correct errors and deviations in the translation to improve the accuracy of the translation; Classify the identified semantic deviations, such as inaccurate term translation, context mismatch, semantic confusion, etc., and prioritize them according to their impact on the translation quality; for high-priority semantic deviations, correct them first; Introduce the intelligent learning ability of the large language model to continuously learn and optimize the term library and the knowledge graph; by continuously learning and analyzing new translation data and domain knowledge, the large language model can automatically discover new terms and translation rules and incorporate them into the term library and the knowledge graph; at the same time, the large language model can also automatically adjust and optimize the term translation and matching strategies according to user feedback and translation quality evaluation results to achieve self-optimization and continuous improvement of the term library.

[0035] The working principle and effects of the above technical solution are as follows: By deeply analyzing the professional knowledge, industry standards, and common terms in the target field, a complete term library is constructed. Combining with the knowledge graph can greatly improve the accuracy and consistency of term translation. The standardized processing of the term library ensures that term translation is not affected by personal understanding or context changes, thereby avoiding translation errors or confusion. By combining the term library with the knowledge graph, the relationships and contexts between terms can be understood, thus improving the semantic coherence of the translated content. The semantic associations in the knowledge graph help to maintain the consistency of terms and concepts in translation, making the translation results not only accurate but also natural and fluent. By comparing with the term library and the knowledge graph, the system can automatically identify and correct potential semantic biases in translation, such as inaccurate term translation, context mismatch, semantic confusion, etc. The automated comparison greatly improves the accuracy and efficiency of translation and reduces the burden of manual inspection. By classifying and prioritizing the identified semantic biases, it can ensure that the errors that have a greater impact on translation quality are corrected in a timely manner. This intelligent classification mechanism improves the efficiency of translation quality control and ensures that the most important translation problems are solved first. Introducing the intelligent learning ability of the large language model enables the term library and the knowledge graph to be continuously updated and improved. With the accumulation of new translation data and domain knowledge, the system can automatically discover new terms, translation rules, and incorporate them into the existing knowledge system. This continuous self-learning and optimization function ensures that the term library and translation strategies keep up with the times and adapt to the changing translation needs. Through the automated comparison and optimization process, the translation speed can be greatly improved. The standardization of the term library and the support of the knowledge graph ensure the consistency of translation quality, avoid repetitive labor and time waste, and thus improve the overall translation efficiency. The construction of the term library and the knowledge graph is cross-domain, so the system can adapt to the translation needs of multiple different fields. As the system continuously learns new domain knowledge, it can provide accurate translation support for various industries. Through the feedback mechanism of the large language model, the translation strategy can be optimized according to the specific needs and feedback of users, enabling users to obtain personalized and high-quality translation results. This user-centered adaptive optimization enhances the overall translation experience.

[0036] In one embodiment of the present invention, S5 includes: S51. Construct a translation quality evaluation model based on deep learning technology, including multiple dimensions such as accuracy evaluation, fluency evaluation, and semantic coherence evaluation; randomly extract samples from the optimized translation results and perform manual annotation as the training data and test data of the evaluation model. S52. Use the annotated data to train and test the evaluation model, continuously adjust the model parameters to improve the accuracy of evaluation; use the trained evaluation model to evaluate the quality of the optimized translation results and generate a detailed quality evaluation report. S53. Analyze the quality assessment report to identify deficiencies and potential improvement points in the translation algorithm; based on the evaluation results and analysis, iteratively optimize the translation algorithm, including adjusting the model structure, improving translation strategies, increasing domain adaptability, etc.; S54. Continuously introduce new bilingual text data and domain knowledge to continuously update and optimize the translation model and evaluation model; establish a multi-channel user feedback mechanism, including online surveys, user reviews, professional reviews, etc., to collect user feedback in a timely manner; organize and analyze the user feedback data to extract user satisfaction with translation quality and improvement suggestions; S55. Integrate user feedback into the optimization process of the translation algorithm and evaluation model to continuously improve translation quality and user satisfaction; before outputting the final translation result, conduct a final quality check on the translation result to ensure the accuracy and fluency of the translation result; S56. Adjust the format and layout of the translation result according to user needs to meet specific user requirements; output and deliver the final translation result in the format and through the channels specified by the user.

[0037] The working principle and effects of the above technical solution are as follows: By constructing a translation quality evaluation model based on deep learning, the quality of translation results can be comprehensively evaluated from multiple dimensions (accuracy, fluency, semantic coherence, etc.). The method of randomly sampling and manually annotating provides the model with real and representative data, thus improving the accuracy of the evaluation results; Using the trained evaluation model to evaluate the quality of translation results can automatically generate detailed quality reports and help the development team quickly identify problems existing in the translation. This method reduces the time and effort of manual inspection and provides objective and systematic quality feedback; By analyzing the quality evaluation report, deficiencies and potential improvement points in the translation algorithm can be identified, and then the translation model can be iteratively optimized. This includes measures such as adjusting the model structure, improving translation strategies, and enhancing domain adaptability, so as to continuously improve the performance of the translation system; By continuously introducing new bilingual text data and domain knowledge, the translation model and evaluation model can be continuously updated and optimized. This dynamic improvement mechanism enables the system to adapt to changing translation needs, improving translation quality and accuracy; By establishing a multi-channel user feedback mechanism (including online surveys, user reviews, and professional reviews, etc.), feedback from users on translation quality can be collected in a timely manner, and adjustments and optimizations can be made according to users' suggestions. This mechanism ensures the important position of user needs and satisfaction in the translation algorithm, thereby enhancing the user experience; Adjusting the format and typesetting of the translation results according to user needs to meet specific requirements. Through this process, it can be ensured that the translation results are not only accurate in content but also meet the user's format and typesetting standards, providing personalized translation services; Through the final quality inspection link, the accuracy and fluency of the final translation results can be ensured, thereby improving the credibility and professionalism of the translation output. This step helps ensure high standards and consistency of the translation results and avoids potential quality problems; By combining translation quality evaluation, continuous improvement, and personalized adjustment, the translation quality can be continuously improved, thereby increasing user satisfaction. The integration of user feedback helps the development team understand actual needs and improvement directions, ensuring that the system meets the ever-changing market demands.

[0038] An embodiment of the present invention, a translation semantic correction system based on a large language model, includes a memory, a processor, and a computer program stored on the memory and executable on the processor. The processor executes the program to implement the translation semantic correction method based on a large language model as described in any one of the above.

[0039] Obviously, those skilled in the art can make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalent technologies, the present invention also intends to include these changes and modifications.

Claims

1. A translation semantic correction method based on large language models, characterized in that, The method includes: S1. Collect data; S2. Use a large language model for named entity recognition, extract specific entities from the source language text, analyze the relationships between entities through the large language model, and generate entity pair relationships; identify and extract the attributes of entities, and associate these attributes with the corresponding entities; integrate the extracted entities, relationships, and attributes according to the graph structure to construct a preliminary knowledge graph; S3. Input the source language text into the model for translation; S4. Compare the terms in the preliminary translation result with the knowledge graph to identify potential semantic deviations; S5. Conduct a quality assessment of the translation result.

2. The method for correcting translation semantics based on a large language model according to claim 1, wherein The above S1 includes: S11. Collect bilingual text data of the source language text and the target language text from multiple channels; S12. Preprocess the data collected in real time.

3. The method for correcting translation semantics based on a large language model according to claim 1, wherein, The above S2 includes: S21. Use a large language model for named entity recognition, extract key entities, and verify the identified key entities; S22. Classify the entities according to entity types and add labels to them; S23. Analyze the relationships between entities through the LLM, extract entity pair relationships; verify the extracted relationships, and organize the relationships to form a structured relationship network; S24. Identify the attributes of entities based on the LLM and rule matching technology, verify the identified attributes, and at the same time associate the attributes with the corresponding entities to form attribute-entity pairs; S25. Integrate the extracted entities, relationships, and attributes according to the graph structure to construct a preliminary knowledge graph; optimize the preliminarily constructed knowledge graph; S26. Store the optimized knowledge graph in the database and establish an efficient indexing mechanism.

4. The method for correcting translation semantics based on a large language model according to claim 3, wherein The above S23 includes: S231. Based on the in-depth understanding ability of the LLM, conduct a context analysis of the text containing entities; through the context information, preliminarily judge the associations or relationships existing between entities; S232. Based on the learning ability of the LLM, mine common entity relationship patterns from historical data or large-scale corpora; through pattern matching, locate the possible entity relationships in the text; S233. According to the diversity of entity types and relationship patterns, formulate relationship extraction strategies; use the semantic parsing ability of the LLM to deeply parse the text and extract the relationships between entities; conduct a confidence assessment of the extracted relationships; S234. Use texts or data from different sources to cross-validate the extracted relationships, remove duplicates for duplicate or redundant relationships, and merge relationships with the same or similar meanings; S235. According to the type and importance of the relationships, construct a relationship hierarchy; display the extracted and verified relationships in a graphical way to form a structured relationship network; display the information through a visualization tool; S236. Store the optimized relationship network in the database and establish an indexing mechanism for the relationship network database.

5. The method for correcting translation semantics based on a large language model according to claim 4, wherein The above S231 includes: Preprocess the text containing entities, and use the in-depth understanding ability of the LLM to embed the entities and their context information into a high-dimensional vector space; For each entity in the text, calculate its contextual similarity with other entities through cosine similarity; based on historical data or large-scale corpora, mine common entity association rules; Match the entity pairs in the text with the mined association rules to preliminarily judge the possible associations or relationships between entities; at the same time, utilize the prediction ability of the LLM to predict the relationships of entity pairs that are contextually similar but not matched to the rules; Utilize the time series analysis ability of the LLM to capture the time clues in the context and make a preliminary judgment on the dynamic changes of entity relationships; Combine the contextual information in multiple texts or data sources, construct a cross-context entity relationship graph, and infer the cross-context relationships between entities using the path and connection information in the graph; Utilize semantic role labeling technology to identify the semantic roles of entities and vocabulary in the text, and evaluate the confidence of the preliminarily judged entity relationships; Rank the preliminarily judged entity relationships according to the confidence and importance of the relationships, and set the priorities for different relationships.

6. The method for correcting translation semantics based on a large language model according to claim 1, wherein The S3 includes: S31. Select a model as the translation engine; S32. Generate a preliminary translation result; S33. Replace and optimize the vocabulary in the translation result; S34. Adjust and optimize the paragraph and discourse structures of the translation result.

7. The method for correcting translation semantics based on a large language model according to claim 1, characterized in that The S4 includes: S41. Construct a professional term library and identify potential semantic biases; S42. Classify the identified semantic biases; S43. Perform targeted translation; S44. Conduct an overall semantic check on the optimized translation result.

8. The method for correcting translation semantics based on a large language model according to claim 7, characterized in that, The S41 includes: Deeply analyze the professional knowledge, industry standards, and common terms in the target field; Based on the collected terms, construct a professional term library; Integrate the term library with the existing knowledge graph; Compare the preliminary translation result with the term library and the knowledge graph to identify potential semantic biases; Classify the identified semantic biases and rank them according to their impact on translation quality; Continuously learn and optimize the term library and the knowledge graph.

9. The method for correcting translation semantics based on a large language model according to claim 1, wherein The S5 includes: S51. Construct a translation quality evaluation model based on deep learning technology; S52. Train and test the evaluation model; S53. Iteratively optimize the translation algorithm; S54. Organize and analyze the user feedback data, and extract the user's satisfaction with translation quality and improvement suggestions; S55. Integrate the user feedback into the optimization process of the translation algorithm and the evaluation model; S56. Adjust the format and typeset the translation result.

10. A translation semantic correction system based on a large language model, characterized in that, It includes a memory, a processor, and a computer program stored on the memory and executable on the processor. The processor executes the program to implement the translation semantic rectification method based on a large language model as described in any one of claims 1-9.

Citation Information

Patent Citations

  • Machine translation-based construction method for Chinese semantic knowledge base

    CN105677913A

  • Deep reading method and system based on knowledge graph

    CN118193748A

  • Big data knowledge graph system for outdoor melons and vegetables

    CN118861315A

  • Intelligent machine translation method and device based on pre-trained large language model

    CN119227699A

  • Knowledge graph construction method and system based on large model technology

    CN119494390A

Cited By

  • Translation optimization method and system of knowledge graph assisted semantic enhancement translation model

    CN120725033A

  • Composite knowledge chain construction method for scientific research logic representation

    CN120930751A

  • A composite knowledge chain construction method for scientific research logic representation

    CN120930751B

  • Method and device for translating entity names in text and computer equipment

    CN120996058A

  • Automatic subtitle translation method and device based on large model and storage medium

    CN121189338A