Legal translation method and system based on neural network

Through the neural network-based legal translation method, the problems of insufficient accuracy and economic infeasibility of existing legal translation software are solved, and high accuracy and feasibility of large-scale translation are achieved.

CN120218087AInactive Publication Date: 2025-06-27CHENGDU LUYOUYOU TRANSLATION SERVICE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510279942.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-11
Publication Date
2025-06-27
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing legal translation software is insufficiently accurate when processing legal texts, and still requires manual assistance, and large-scale translation tasks are not economically feasible.

Method used

A neural network-based legal translation method is adopted, and the legal documents to be translated are received, preprocessed and vocabulary mapping are performed, and the vocabulary meaning correlation is dynamically matched using the language model, and integrated verification is carried out through the functional verification model to output the translated literature.

Benefits of technology

It improves the accuracy of legal translation, reduces the dependence on manual translation, and increases the feasibility of large-scale legal document translation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120218087A_ABST
    Figure CN120218087A_ABST
Patent Text Reader

Abstract

The invention relates to a legal translation method and system based on a neural network, and is applied to the technical field of legal translations, and the method comprises the steps: receiving a to-be-translated legal document, preprocessing the to-be-translated legal document, and obtaining a to-be-translated document; mapping the to-be-translated literature with a preset vocabulary database, and determining an unambiguous clear vocabulary and a multi-meaning complex-meaning vocabulary; performing association degree dynamic matching on the meaning of the clear vocabulary and the meaning of the remeaning vocabulary according to a preset language model, and determining an optimal vocabulary meaning corresponding to the remeaning vocabulary; and according to a preset functional verification model, carrying out integration verification on the clear vocabulary and the complex-meaning vocabulary, and outputting a translated document.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of legal translation, and particularly to a legal translation method and system based on a neural network. Background Art

[0002] With the acceleration of the globalization process, commercial exchanges between different countries and regions have become increasingly frequent. However, in the fields of international business, legal services, intellectual property protection, etc., language barriers remain a major challenge. Especially when it comes to cross-border legal affairs, accurate translation of legal texts becomes crucial.

[0003] Traditional translation of legal documents usually relies on human translation, which is not only time-consuming and laborious but also costly. In addition, due to the professionalism of legal terms and the differences between different legal systems, even experienced translators may hardly avoid making errors.

[0004] In recent years, natural language processing (NLP) technology has made remarkable progress. In particular, the application of machine learning and deep learning algorithms enables computers to better understand and generate human language. However, most current translation software or platforms still have many deficiencies in processing legal texts, such as being unable to accurately capture the meaning of legal terms and not being able to adapt to the language habits of specific legal domains.

[0005] Regarding the above related technologies, it is considered that the current legal translation software has insufficient translation accuracy and still requires human-assisted translation, which is not economically feasible for large-scale translation tasks. Summary of the Invention

[0006] To address the problem that the current legal translation software has insufficient translation accuracy and still requires human-assisted translation, which is not economically feasible for large-scale translation tasks, this application provides a legal translation method and system based on a neural network.

[0007] In a first aspect, a legal translation method based on a neural network provided by this application adopts the following technical solutions: including:

[0008] Receiving a legal document to be translated, preprocessing the legal document to be translated to obtain a document to be translated;

[0009] Mapping the document to be translated with a preset vocabulary database to determine unambiguous clear words and polysemous words with multiple meanings;

[0010] Dynamically matching the meanings of the clear words and the polysemous words according to a preset language model to determine the best word meaning corresponding to the polysemous words;

[0011] Integrate and verify the explicit vocabulary and the polysemous vocabulary according to a preset functional verification model, and output a translated document.

[0012] Preferably, mapping the document to be translated with a preset vocabulary database to determine explicit vocabulary without ambiguity and polysemous vocabulary with multiple meanings, includes:

[0013] Establish multiple sliding windows with different lengths in the document to be translated according to a preset fixed phrase hash table;

[0014] Screen the fixed phrases in the document to be translated according to the sliding window to determine the fixed phrases;

[0015] Remove duplicate vocabulary in the vocabulary to be translated and combine with the fixed phrases to generate a vocabulary list to be translated;

[0016] Map the vocabulary database with the vocabulary list to be translated to determine and translate the explicit vocabulary and the polysemous vocabulary, and the explicit vocabulary includes the fixed phrases.

[0017] Preferably, dynamically matching the relevance of the meanings of the explicit vocabulary and the polysemous vocabulary according to a preset language model to determine the best vocabulary meaning corresponding to the polysemous vocabulary, includes:

[0018] Establish a context association window for the polysemous vocabulary in the document to be translated;

[0019] Vectorize the vocabulary in the context association window according to the language model to obtain vocabulary embedding vectors;

[0020] Take a weighted average value according to the vocabulary embedding vectors to obtain a context vector;

[0021] Construct candidate meaning vectors for each meaning in the polysemous vocabulary according to the language model;

[0022] Calculate the relevance of all the candidate meaning vectors and the context vector using a preset cosine similarity formula;

[0023] Select the vocabulary meaning corresponding to the candidate meaning vector with the highest relevance, that is, determine the best vocabulary meaning corresponding to the polysemous vocabulary.

[0024] Preferably, after vectorizing the vocabulary in the context association window according to the language model to obtain vocabulary embedding vectors, it further includes:

[0025] If there are multiple polysemous vocabulary in the context association window, perform dependency syntactic analysis according to a preset dependency parser to obtain vocabulary dependency relationships;

[0026] Determine the feature vector of a word according to the lexical dependency relationship and the lexical embedding vector;

[0027] According to the feature vectors corresponding to different meanings of the polysemous word, perform vector fusion based on the attention weighting algorithm to obtain a feature fusion vector;

[0028] Calculate the correlation degree of the candidate meaning vector according to the candidate meaning vector of the target polysemous word, the feature fusion vectors of other polysemous words, and the feature vector.

[0029] Preferably, the formula of the attention weighting algorithm is as follows:

[0030]

[0031]

[0032] Wherein, Si is the original attention score of the i-th related word, Sj is the original attention score of the j-th related word, ai is the normalized attention score, exp is the exponential function, Vi is the feature vector of the i-th word meaning, and FusedVector represents the feature fusion vector.

[0033] Preferably, after calculating the correlation degree of the candidate meaning vector according to the candidate meaning vector of the target polysemous word, the feature fusion vectors of other polysemous words, and the feature vector, it further includes:

[0034] Select the word meaning with the highest correlation degree to determine the meaning of the target polysemous word;

[0035] Determine the feature vector of the target polysemous word according to the meaning of the target polysemous word;

[0036] Calculate the correlation degrees of the candidate meaning vectors corresponding to other polysemous words according to the candidate meaning vectors, the feature fusion vectors, and the feature vectors of other polysemous words, until the meanings of all polysemous words in the context association window are determined.

[0037] Preferably, the integrating and validating the definite words and the polysemous words according to a preset functional validation model and outputting a translated document includes:

[0038] Arrange and integrate the definite words, the polysemous words, and the fixed phrases according to the word order of the document to be translated and the dependency statement analysis to obtain an initial translation file;

[0039] Calculate the perplexity of each sentence according to a preset perplexity formula, and rearrange the sentences with a perplexity greater than the threshold value;

[0040] Adjust the grammar of each sentence according to a preset grammar checking tool, and output an adjusted translation file;

[0041] Search for similar sentences in the historical database with a synonymous word ratio greater than a preset ratio in the adjusted translation file;

[0042] Detect the matching degree of the similar sentences according to a preset matching degree algorithm, and rearrange the translation sentences with a matching degree less than a preset value;

[0043] Until all the translation sentences are verified, output the translation document according to the format of the legal document to be translated;

[0044] Preferably, the perplexity formula is specifically as follows:

[0045]

[0046] Wherein, w1, w2,..., wn represent the vocabulary sequence in the sentence, p(wi∣w<i) represents the conditional probability of the vocabulary wi given the previous vocabulary, and P represents the perplexity;

[0047] The matching degree algorithm formula is specifically as follows:

[0048]

[0049] m(s, l i ) = cos(s, l i );

[0050] Wherein, i represents the number of the similar sentences, m represents the matching score of the sentence s and the similar sentence li, M(s) represents the total matching score of the sentence s and all the similar sentences, s represents the vocabulary embedding vector of the whole sentence, and li represents the vocabulary embedding vector of each vocabulary.

[0051] In a second aspect, a legal translation device based on a neural network according to the present application adopts the following technical solutions, including:

[0052] A file receiving module, configured to receive a legal document to be translated, preprocess the legal document to be translated, and obtain a document to be translated;

[0053] A vocabulary mapping module, configured to map the document to be translated with a preset vocabulary database to determine unambiguous clear vocabulary and polysemous vocabulary with multiple meanings;

[0054] A lexical association module, configured to dynamically match the association degrees between the meanings of the explicit vocabulary and the polysemous vocabulary according to a preset language model, and determine the best lexical meaning corresponding to the polysemous vocabulary;

[0055] A function verification module, configured to integrally verify the explicit vocabulary and the polysemous vocabulary according to a preset functional verification model, and output a translated document.

[0056] In a third aspect, the present application further provides a control device, including:

[0057] It includes a memory and a processor, and a computer program capable of being loaded and executed by the processor, such as the above-mentioned neural network-based legal translation method, is stored on the memory.

[0058] In a fourth aspect, the present application further provides a computer-readable storage medium, storing a computer program capable of being loaded and executed by the processor, such as the above-mentioned neural network-based legal translation method.

[0059] In summary, in the present application, the user selects a legal field and inputs a legal document to be translated on the initial interface. The system performs preprocessing steps such as text recognition and sorting on the legal document to be translated. Then, the system selects a corresponding vocabulary database according to the legal field, and establishes a mapping relationship between the vocabulary database and the vocabulary on the document to be translated. At the same time, the system screens the fixed vocabulary in the document to be translated according to the fixed phrase hash table, reducing the subsequent vocabulary translation volume. Then, the system calculates the embedding vector of each vocabulary according to the BERT language model, and calculates the feature fusion vector of the polysemous vocabulary. The system calculates the association degrees between the polysemous vocabulary and other vocabularies according to the vocabulary embedding vector, feature fusion vector, feature vector, etc., so as to determine the best lexical meaning with the highest association degree in the polysemous vocabulary. After the system translates the vocabulary, it sorts and integrates the vocabulary according to the document to be translated to form an initial translated document. Finally, the system performs function verification on the initial translated document in the functional space, so as to achieve the effect of accurately translating legal documents, thereby reducing the usage rate of manual labor in translating legal documents and increasing the feasibility of translating large-scale legal documents. Description of the Drawings

[0060] Figure 1 is a schematic flowchart of a neural network-based legal translation method.

[0061] Figure 2 is a structural block diagram of a neural network-based legal translation device.

[0062] Description of the reference numerals: 210, document receiving module; 220, vocabulary mapping module; 230, lexical association module; 240, function verification module. Detailed Embodiments

[0063] The following will further elaborate on Figure 1 - Figure 2 this application in detail.

[0064] The core of the legal translation software lies in its independently constructed legal database. This database will integrate legal text data in Chinese and English from multiple sources to ensure the comprehensiveness and accuracy of the translation basis. The steps for establishing the database include:

[0065] Data collection: The construction of the database will obtain legal text data from multiple channels. These data sources include the Chinese-English text database of national laws, authoritative institutions in the field of international law, case files of major law firms, and publicly released legal documents. This multi-level data collection ensures the extensiveness and authority of the database.

[0066] Data sorting and tagging: After the data collection is completed, the system will classify and sort the data through natural language processing (NLP) technology, and tag each legal provision, case, or clause. For example, the provisions in different legal fields (such as contract law, intellectual property law, international trade law) will be classified separately to facilitate quick access to relevant content during translation.

[0067] Update and maintenance mechanism: Legal texts are dynamically changing, and the database must have the ability to be updated in real time. The system will establish an automated data update mechanism to regularly obtain the latest legal texts from authoritative sources and integrate them. The database also needs to conduct regular quality checks to ensure the accuracy and consistency of the data.

[0068] Based on this, the embodiment of this application discloses a legal translation method based on a neural network. The execution entity is the control system. The system selects the corresponding vocabulary database according to the legal field, and establishes a mapping relationship between the vocabulary database and the vocabulary in the document to be translated. According to the BERT language model, the embedding vector of each vocabulary is calculated, and the feature fusion vector of polysemous vocabulary is calculated. The system calculates the correlation degree between polysemous vocabulary and other vocabulary according to the vocabulary embedding vector, feature fusion vector, feature vector, etc., so as to determine the best meaning of the polysemous vocabulary and achieve the effect of translating legal documents.

[0069] Referring to Figure 1 , the embodiment of this application at least includes steps S10 to S40.

[0070] S10, receive the legal document to be translated, preprocess the legal document to be translated, and obtain the document to be translated.

[0071] S20, map the document to be translated with the preset vocabulary database to determine the clear vocabulary without ambiguity and the polysemous vocabulary with multiple meanings.

[0072] S30. Dynamically match the relevance between the meanings of explicit words and polysemous words according to a preset language model, and determine the best word meaning corresponding to the polysemous word.

[0073] S40. Integrate and verify the explicit words and polysemous words according to a preset functional verification model, and output the translated document.

[0074] Among them, the language model used in this application is the BERT language model. The BERT language model after a large amount of training can accurately calculate the relationship degree of different words in legal regulations, increasing the accuracy of translation.

[0075] Specifically, the user selects the legal field and inputs the legal document to be translated on the initial interface. The system performs preprocessing steps such as text recognition and sorting on the legal document to be translated. Then the system selects the corresponding vocabulary database according to the legal field, and establishes a mapping relationship between the vocabulary in the vocabulary database and the words in the document to be translated. At the same time, the system screens the fixed words in the document to be translated according to the fixed phrase hash table, reducing the subsequent vocabulary translation volume. Then the system calculates the embedding vector of each word according to the BERT language model, and calculates the feature fusion vector of the polysemous word. The system calculates the relevance between the polysemous word and other words according to the word embedding vector, feature fusion vector, feature vector, etc., so as to determine the word meaning with the best relevance in the polysemous word. After the system translates the words, it sorts and integrates the words according to the document to be translated to form an initial translated document. Finally, the system performs functional verification on the initial translated document in the functional space, so as to achieve the effect of accurately translating legal documents, thereby reducing the usage rate of manual labor in translating legal documents and increasing the feasibility of translating large-scale legal documents.

[0076] In some embodiments, step S20 specifically includes the following steps: establish a context association window for the review words in the document to be translated; perform vector representation on the words in the context association window according to the language model to obtain word embedding vectors; perform weighted average value taking according to the word embedding vectors to obtain context vectors; construct candidate meaning vectors for each meaning in the polysemous word according to the language model; use a preset cosine similarity formula to calculate the relevance between all candidate meaning vectors and the context vectors; select the word meaning corresponding to the candidate meaning vector with the highest relevance, that is, determine the best word meaning corresponding to the polysemous word.

[0077] Specifically, the mapping neural network is the technical core of this legal translation software. It can translate complex legal languages into the target language while preserving the legal meaning and the accuracy of the context. The establishment process includes the following steps: Semantic understanding and learning: The first step of the mapping neural network is the in-depth learning of legal languages. By training the deep learning model, the system can understand the semantic differences of legal terms between different languages and establish corresponding mapping relationships. For example, the English legal term "Consideration" corresponds to "consideration" in Chinese law. The system will learn and master this corresponding relationship through a large number of samples. Neural network architecture design: The neural network of this system will adopt a multi-layer neural network architecture, and each layer focuses on specific translation tasks. The initial layer is responsible for the translation of basic vocabulary, the middle layer processes complex sentence structures, and the high layer focuses on the precise conversion of legal terms. Each layer of the network is continuously optimized through a feedback mechanism to ensure that the final translated result conforms to legal logic and is smooth and easy to understand. Dynamic adjustment of mapping relationships: The translation of legal languages requires extremely high accuracy. Therefore, the mapping neural network will have the ability of dynamic adjustment. During the translation process, the system will combine the latest data in the legal database and the feedback of users to optimize the mapping relationship in real time. This means that when translating for the first time, the system may give several alternative translations and gradually adjust the mapping relationship according to the user's selection and feedback to improve the accuracy of translation.

[0078] Actually, when the user selects the quick translation mode, in order to save translation time, the system automatically calculates the context vector based on the explicit vocabulary in the context association window, and then selects the vocabulary meaning with the highest degree of association according to the degree of association between the context vector and the candidate meaning vector, so as to achieve the effect of quickly translating polysemous words, facilitating the rapid translation of a large number of legal documents while ensuring accuracy.

[0079] In some embodiments, when considering the issue of more accurate translation, the corresponding processing steps are as follows: If there are multiple polysemous words in the context association window, dependency syntactic analysis is performed according to a preset dependency parser to obtain the lexical dependency relationship; According to the lexical dependency relationship and the lexical embedding vector, the feature vector of the word is determined; According to the feature vectors corresponding to different meanings of the polysemous word, vector fusion is performed based on the attention weighting algorithm to obtain the feature fusion vector; According to the candidate meaning vector of the target polysemous word, the feature fusion vectors of other polysemous words, and the feature vector, the degree of association of the candidate meaning vector is calculated.

[0080] In practice, when the user selects the accurate translation mode, the system performs dependency syntactic analysis on the words in the context association window. Based on the word dependency relationships and word embedding vectors, the feature vectors of the words are determined. Then, based on the multiple feature vectors of the polysemous words, attention-weighted fusion is performed to obtain the feature vectors. Finally, the system calculates the association degree based on the candidate meaning vectors of the target polysemous words, other feature fusion vectors, and the feature vectors, thereby taking into account the influence of the polysemous words in the context window when calculating the association degree, and further facilitating the dynamic determination of the meanings of polysemous words in order to increase the accuracy of polysemous word translation.

[0081] Furthermore, the attention-weighted algorithm formula is as follows:

[0082]

[0083] Among them, Si is the original attention score of the i-th relevant word, Sj is the original attention score of the j-th relevant word, ai is the normalized attention score, exp is the exponential function, Vi is the feature vector of the meaning of the i-th word, and FusedVector represents the feature fusion vector.

[0084] By adopting the attention-weighted algorithm, the weights of the meanings of different words in the sentence can be changed, avoiding the situation where the weights of all words are the same in the average value, and greatly increasing the accuracy of the calculation of the feature fusion vector.

[0085] In some embodiments, considering the problem of dynamically translating polysemous words, the corresponding processing steps are as follows: select the word meaning with the highest association degree to determine the meaning of the target polysemous word; determine the feature vector of the target polysemous word according to the meaning of the target polysemous word; calculate the association degrees of the candidate meaning vectors corresponding to other polysemous words in turn according to the candidate meaning vectors, feature fusion vectors, and feature vectors of other polysemous words until the meanings of all polysemous words in the context association window are determined.

[0086] Specifically, after the system translates the meaning of a polysemous word, this polysemous word becomes a clear word, so it is necessary to re-establish the word embedding vector and the feature vector in order to increase the accuracy of the translation of the next target polysemous word.

[0087] In some embodiments, step S40 specifically includes the following steps: arranging and integrating explicit vocabulary, polysemous vocabulary, and fixed phrases in accordance with the vocabulary order and dependency statement analysis of the document to be translated to obtain an initial translation file; calculating the perplexity of each statement according to a preset perplexity formula, and rearranging the statements with a perplexity greater than the threshold; adjusting the grammar of each statement according to a preset grammar checking tool to output an adjusted translation file; searching for similar statements with synonymous vocabulary greater than a preset ratio in the historical database in the adjusted translation file; detecting the matching degree of the similar statements according to a preset matching degree algorithm, and rearranging the translation statements with a matching degree less than the preset value; until all translation statements are verified, and outputting the translated document in the format of the legal document to be translated.

[0088] Specifically, the system sorts the translated vocabulary according to the dependency statement analysis to facilitate the fluency of the statement. Then the system calculates the perplexity of the statement and rearranges the statements with a larger perplexity. Then the system adjusts the grammar of each statement according to the grammar checking tool to increase the grammatical rigor of the statement. At the same time, the system detects the matching degree of the similar statements according to the matching degree algorithm. If the total matching degree is too small, it indicates that there is a sorting problem with the translation statement, and it is rearranged, so as to conduct a functional verification of the translation statement to increase the fluency, grammatical accuracy, and statement function accuracy of the statement translation.

[0089] Actually, to ensure the accuracy and reliability of the translation, the software system will introduce two key components: the "translation space" and the "functional verification space". Establishment of the translation space: The translation space is a multi-dimensional language conversion model. The system will simulate the translation process of legal texts in Chinese-English and English-Chinese translations within this space. By inputting the mapping neural network into the translation space, the system can conduct simulation tests to ensure that the translation of each legal text is not only linguistically correct but also logically rigorous in terms of law. Construction of the functional verification space: The functional verification space is the link for the final confirmation of the translation result. The system will verify whether the translation is completely consistent with the original text in terms of legal meaning by comparing relevant articles and cases in the legal database. The functional verification space will also test the application scenarios of the translation, such as whether it can be effectively used in actual legal contracts or legal documents. Through this mechanism, the system can further identify potential translation errors and ensure that each translation result can meet the requirements of legal practice.

[0090] Furthermore, the perplexity formula is specifically as follows:

[0091]

[0092] Among them, w1, w2, …, wn represent the sequence of words in a sentence, p(wi∣w<i) represents the conditional probability of word wi given the previous words, and P represents perplexity;

[0093] The matching degree algorithm formula is specifically as follows:

[0094]

[0095] m(s, l i ) = cos(s, l i );

[0096] Among them, i represents the number of similar sentences, m represents the matching score between sentence s and similar sentence li, M(s) represents the total matching score between sentence s and all similar sentences, s represents the word embedding vector of the entire sentence, and li represents the word embedding vector of each word.

[0097] The implementation principle of a legal translation method based on neural network in an embodiment of this application is as follows: The user selects a legal field and inputs a legal document to be translated through the initial interface. The system performs preprocessing steps such as text recognition and sorting on the legal document to be translated. Then, the system selects a corresponding vocabulary database according to the legal field and establishes a mapping relationship between the vocabulary in the vocabulary database and the words in the document to be translated. At the same time, the system screens the fixed words in the document to be translated according to the fixed phrase hash table to reduce the subsequent vocabulary translation volume. Then, the system calculates the embedding vector of each word according to the BERT language model and calculates the feature fusion vector of polysemous words. The system calculates the correlation degree between polysemous words and other words based on the word embedding vector, feature fusion vector, feature vector, etc., so as to determine the best meaning of the polysemous word. After the system translates the words, it sorts and integrates the words according to the document to be translated to form an initial translation document. Finally, the system performs functional verification on the initial translation document in the functional space, so as to achieve the effect of accurately translating legal documents, thereby reducing the usage rate of manual labor in translating legal documents and increasing the feasibility of translating large-scale legal documents.

[0098] Figure 1 It is a schematic flowchart of a legal translation method based on neural network in an embodiment. It should be understood that although Figure 1 the steps in the flowchart are shown in sequence according to the arrows, these steps are not necessarily executed in the order indicated by the arrows; unless otherwise clearly stated in this article, the execution of these steps has no strict order limit, and these steps can be executed in other orders; and Figure 1At least a part of the steps may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily executed and completed at the same moment, but can be executed at different moments. The execution order of these sub-steps or stages is not necessarily sequential, but can be executed alternately or in turn with at least a part of other steps or sub-steps or stages of other steps.

[0099] Based on the same technical concept, referring to Figure 2 , this embodiment of the present application further provides a legal translation device based on a neural network, adopting the following technical solutions. The device includes:

[0100] A document receiving module 210, configured to receive a legal document to be translated, preprocess the legal document to be translated, and obtain a literature to be translated;

[0101] A vocabulary mapping module 220, configured to map the literature to be translated with a preset vocabulary database to determine unambiguous clear vocabulary and polysemous vocabulary with multiple meanings;

[0102] A vocabulary association module 230, configured to dynamically match the meanings of the clear vocabulary and the polysemous vocabulary according to a preset language model to determine the best vocabulary meaning corresponding to the polysemous vocabulary;

[0103] A function verification module 240, configured to integrally verify the clear vocabulary and the polysemous vocabulary according to a preset functional verification model, and output a translated literature.

[0104] In some embodiments, the vocabulary mapping module 220 is specifically configured to establish multiple sliding windows of different lengths in the literature to be translated according to a preset fixed phrase hash table;

[0105] Screen the fixed phrases in the literature to be translated according to the sliding windows to determine the fixed phrases;

[0106] Remove duplicate vocabulary in the vocabulary to be translated and combine with the fixed phrases to generate a vocabulary list to be translated;

[0107] Map according to the vocabulary database and the vocabulary list to be translated to determine and translate the clear vocabulary and the polysemous vocabulary. The clear vocabulary includes fixed phrases.

[0108] In some embodiments, the vocabulary association module 230 is specifically configured to establish a context association window in the literature to be translated according to the polysemous vocabulary;

[0109] Perform vector representation on the vocabulary in the context association window according to the language model to obtain a vocabulary embedding vector;

[0110] Perform weighted average value taking according to the vocabulary embedding vector to obtain a context vector;

[0111] Construct candidate meaning vectors for each meaning in the polysemous vocabulary according to the language model;

[0112] Use a preset cosine similarity formula to calculate the correlation between all candidate meaning vectors and the context vector;

[0113] Select the lexical meaning corresponding to the candidate meaning vector with the highest correlation, that is, determine the best lexical meaning corresponding to the polysemous vocabulary.

[0114] In some embodiments, the lexical association module 230 is further configured to perform dependency syntactic analysis according to a preset dependency parser to obtain lexical dependency relationships if there are multiple polysemous words in the context association window;

[0115] Determine the feature vector of the word according to the lexical dependency relationship and the lexical embedding vector;

[0116] Perform vector fusion according to the attention weighting algorithm based on the feature vectors corresponding to different meanings of the polysemous vocabulary to obtain a feature fusion vector;

[0117] Calculate the correlation of the candidate meaning vector according to the candidate meaning vector of the target polysemous word, the feature fusion vectors of other polysemous words, and the feature vectors.

[0118] In some embodiments, the attention weighting algorithm formula is as follows:

[0119]

[0120] Where Si is the original attention score of the i-th related word, Sj is the original attention score of the j-th related word, ai is the normalized attention score, exp is the exponential function, Vi is the feature vector of the i-th word meaning, and FusedVector represents the feature fusion vector.

[0121] In some embodiments, the lexical association module 230 is further configured to select the lexical meaning with the highest correlation to determine the meaning of the target polysemous word;

[0122] Determine the feature vector of the target polysemous word according to the meaning of the target polysemous word;

[0123] Calculate the correlation of the candidate meaning vectors corresponding to other polysemous words in turn according to the candidate meaning vectors, feature fusion vectors, and feature vectors of other polysemous words until the meanings of all polysemous words in the context association window are determined.

[0124] In some embodiments, the function verification module 240 is specifically configured to arrange and integrate the explicit words, polysemous words, and fixed phrases according to the lexical order of the document to be translated and the dependency sentence analysis to obtain an initial translation file;

[0125] Calculate the perplexity of each statement according to a preset perplexity formula, and rearrange the statements with perplexity greater than the threshold value;

[0126] Adjust the grammar of each statement according to a preset grammar checking tool, and output an adjusted translation file;

[0127] Search for similar statements with synonymous words in the adjusted translation file greater than a preset ratio in the historical database;

[0128] Detect the matching degree of similar statements according to a preset matching degree algorithm, and rearrange the translation statements with a matching degree less than the preset value;

[0129] Until all the translation statements are verified, output a translation document according to the format of the legal document to be translated.

[0130] In some embodiments, the perplexity formula is specifically as follows:

[0131]

[0132] Wherein, w1, w2, …, wn represent the vocabulary sequence in the sentence, p(wi∣w<i) represents the conditional probability of vocabulary wi given the previous vocabulary, and P represents the perplexity;

[0133] The matching degree algorithm formula is specifically as follows:

[0134]

[0135] m(s, l i ) = cos(s, l i );

[0136] Wherein, i represents the number of similar statements, m represents the matching score of sentence s and similar statement li, M(s) represents the total matching score of sentence s and all similar statements, s represents the vocabulary embedding vector of the whole sentence, and li represents the vocabulary embedding vector of each vocabulary.

[0137] The embodiments of the present application also disclose a control device.

[0138] Specifically, the control device includes a memory and a processor, and a computer program capable of being loaded and executed by the processor for the above-mentioned neural network-based legal translation method is stored on the memory.

[0139] The embodiments of the present application also disclose a computer-readable storage medium.

[0140] Specifically, the computer-readable storage medium stores a computer program that can be loaded and executed by a processor, such as the above-mentioned neural network-based legal translation method. The computer-readable storage medium includes, for example, various media that can store program codes, such as USB flash drives, external hard drives, read-only memories (ROM), random access memories (RAM), magnetic disks, or optical discs.

[0141] The above are all preferred embodiments of this application. The protection scope of this application is not limited accordingly. Therefore, all equivalent changes made according to the structure, shape, and principle of this application shall be covered within the protection scope of this application.

Claims

1. A legal translation method based on neural network, characterized in that: include: Receiving legal documents to be translated, preprocessing the legal documents to be translated, and obtaining documents to be translated; Mapping the document to be translated with a preset vocabulary database to determine unambiguous and clear words and polysemous words with multiple meanings; Dynamically matching the meanings of the explicit words and the polysemous words according to a preset language model to determine the best meaning of the word corresponding to the polysemous word; The explicit vocabulary and the polysemous vocabulary are integrated and verified according to a preset functional verification model, and the translation document is output.

2. A legal translation method based on neural network according to claim 1, characterized in that: The mapping of the document to be translated with a preset vocabulary database to determine unambiguous and clear words and multi-meaning words includes: Establishing a plurality of sliding windows of different lengths in the document to be translated according to a preset fixed phrase hash table; Screening fixed phrases in the document to be translated according to the sliding window to determine fixed phrases; Removing repeated words from the words to be translated and combining them with fixed phrases to generate a word list to be translated; The explicit vocabulary and the polysemous vocabulary are determined and translated according to mapping between the vocabulary database and the vocabulary list to be translated, wherein the explicit vocabulary includes the fixed phrase.

3. A legal translation method based on neural network according to claim 2, characterized in that: The dynamically matching the meanings of the explicit words and the polysemous words according to the preset language model to determine the best word meaning corresponding to the polysemous words includes: Establishing a context association window in the document to be translated according to the review vocabulary; Performing vector representation on the vocabulary of the context window according to the language model to obtain a vocabulary embedding vector; Taking a weighted average value according to the vocabulary embedding vector to obtain a context vector; constructing a candidate meaning vector for each meaning in the polysemous vocabulary according to the language model; Calculate the correlation between all the candidate meaning vectors and the context vector using a preset cosine similarity formula; The lexical meaning corresponding to the candidate meaning vector corresponding to the highest degree of association is selected, that is, the optimal lexical meaning corresponding to the multi-meaning word is determined.

4. A legal translation method based on neural network according to claim 3, characterized in that: After performing vector representation on the vocabulary of the context window according to the language model to obtain the vocabulary embedding vector, the method further includes: If there are multiple polysemous words in the context association window, performing dependency syntactic analysis according to a preset dependency parser to obtain lexical dependency relations; Determining a feature vector of a vocabulary according to the vocabulary dependency and the vocabulary embedding vector; According to the feature vectors corresponding to different meanings of the polysemous words, vector fusion is performed according to an attention weighted algorithm to obtain a feature fusion vector; The relevance of the candidate meaning vector is calculated based on the candidate meaning vector of the target polysemous word, the feature fusion vectors of other polysemous words, and the feature vector.

5. A neural network-based legal translation method according to claim 4, characterized in that: The attention weighted algorithm formula is as follows: Among them, Si is the original attention score of the i-th related word, Sj is the original attention score of the j-th related word, ai is the normalized attention score, exp is the exponential function, Vi is the feature vector of the i-th word meaning, and FusedVector represents the feature fusion vector.

6. A neural network-based legal translation method according to claim 4, characterized in that: After calculating the correlation degree of the candidate meaning vector of the target polysemous word, the feature fusion vectors of the other polysemous words, and the feature vector, the following steps are further included: Select the word meaning with the highest correlation degree to determine the meaning of the target polysemous word; Determine the feature vector of the target polysemous word according to the meaning of the target polysemous word; According to the candidate meaning vectors, the feature fusion vectors, and the feature vectors of the other polysemous words, calculate the correlation degrees of the candidate meaning vectors corresponding to the other polysemous words in turn until the meanings of all the polysemous words in the context association window are determined.

7. A neural network-based legal translation method according to claim 4, characterized in that: The integrating and verifying the explicit words and the polysemous words according to a preset functional verification model and outputting a translation document includes: Arrange and integrate the explicit words, the polysemous words, and the fixed phrases according to the word order of the document to be translated and the dependency statement analysis to obtain an initial translation file; Calculate the perplexity of each sentence according to a preset perplexity formula, and rearrange the sentences with perplexity greater than the threshold; Adjust the grammar of each sentence according to a preset grammar checking tool and output an adjusted translation file; Search in the historical database for similar sentences with a synonym ratio greater than a preset ratio in the adjusted translation file; Detect the matching degree of the similar sentences according to a preset matching degree algorithm, and rearrange the translation sentences with a matching degree less than the preset value; Until all the translation sentences are verified, output the translation document according to the format of the legal document to be translated.

8. A neural network-based legal translation method according to claim 7, characterized in that: The specific form of the perplexity formula is as follows: Where w1, w2, …, wn represent the vocabulary sequence in the sentence, p(wi∣w<i) represents the conditional probability of the vocabulary wi given the previous vocabulary, and P represents the perplexity; The specific form of the matching degree algorithm formula is as follows: m(s,l i )=cos(s,l i ); Where i represents the number of the similar sentences, m represents the matching score of the sentence s and the similar sentence li, M(s) represents the total matching score of the sentence s and all the similar sentences, s represents the vocabulary embedding vector of the whole sentence, and li represents the vocabulary embedding vector of each vocabulary.

9. A legal translation device based on a neural network, characterized in that: The device includes: A file receiving module, configured to receive a legal document to be translated, preprocess the legal document to be translated, and obtain a document to be translated; A vocabulary mapping module, configured to map the document to be translated with a preset vocabulary database to determine explicit words without ambiguity and polysemous words with multiple meanings; A vocabulary association module, configured to perform dynamic matching of the correlation degrees of the meanings of the explicit words and the polysemous words according to a preset language model to determine the best word meaning corresponding to the polysemous word; A functional verification module, configured to perform integrated verification on the explicit words and the polysemous words according to a preset functional verification model and output a translation document.

10. A control device, characterized in that: The device includes: It includes a memory and a processor, and a computer program capable of being loaded and executed by the processor as described in any one of claims 1 to 8 is stored on the memory.