Business communication intelligent translation system based on natural language processing

By designing an intelligent business communication translation system based on natural language processing, the problem of translation ambiguity in business communication is solved, fast and accurate business text translation is achieved, and the efficiency and accuracy of business communication is improved.

CN120106101APending Publication Date: 2025-06-06RIZHAO VOCATIONAL & TECHNICAL UNIVERSITY
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510311259.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-17
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

The existing intelligent translation technology has translation ambiguity in business communication, especially the inaccurate translation caused by polysynonyms and contextual dependence and cultural differences in industry terms, which affects business decision-making and cooperative relationships.

Method used

Design a business communication intelligent translation system based on natural language processing, including text acquisition module, term division module, industry determination module, context analysis module, syntax analysis module and text translation module. Through the collaborative work of these modules, the system can accurately identify industry terms, analyze context, build syntax trees, and translate them in combination with dynamically adjusted context information.

Benefits of technology

It achieves rapid and accurate translation of business texts, overcomes language barriers and translation ambiguity problems, greatly improves the efficiency and reliability of business text translation, ensures accurate understanding and communication of information, and thus improves the efficiency and accuracy of business communication.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120106101A_ABST
    Figure CN120106101A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of intelligent translation, and discloses a business communication intelligent translation system based on natural language processing. Comprising a text acquisition module used for acquiring a business text; the term division module is used for dividing the business text to obtain industry terms; the industry determination module determines the industry to which the business text belongs based on the industry terms; the context analysis module is used for performing text analysis on the business text to obtain context information; the syntactic analysis module is used for performing syntactic analysis on the business text according to the context information and constructing a syntactic tree; the text translation module is used for translating the business text by fusing the belonging industry, context information and the syntactic tree; according to the method, the business text can be quickly and accurately translated, and the problem of translation ambiguity in a traditional translation mode is solved; therefore, the efficiency and the reliability of business text translation are greatly improved, accurate understanding and transmission of information are ensured, and the efficiency and the accuracy of business communication are further improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of intelligent translation technology, and more specifically, to a business communication intelligent translation system based on natural language processing. Background Art

[0002] In the era of globalization, the internationalization trend of business communication is becoming increasingly prominent. When companies communicate with international customers, partners and suppliers, language barriers often become an important factor hindering effective communication. Especially in the business environment, rapid response and accurate understanding are crucial. Companies need to be able to quickly translate various information such as meeting minutes, business emails, contract documents, etc. Traditional translation methods are not only time-consuming and labor-intensive, but also lead to inaccurate translation due to human factors, affecting business decisions and the establishment of cooperative relationships. With the rapid development of science and technology, the advancement of natural language processing (NLP) and artificial intelligence (AI) technologies has promoted the rise of intelligent translation technology, making real-time and accurate cross-language communication possible.

[0003] For example, the patent with announcement number CN205910694U discloses a Chinese-Russian wood professional vocabulary translation system; it includes a personal service end, a server and a customer service platform; the server includes a sound sensor, an optical sensor array, an amplifier, a frame converter, a PLC circuit board, a storage circuit, and an analog signal converter. The optical sensor and the sound sensor are connected to the analog signal converter, the input end of the frame converter is connected to the optical sensor and the sound sensor, and the output end of the frame converter is connected to the PLC circuit board through the amplifier, and the PLC circuit board is connected to the storage circuit; an analog signal converter is set on the PLC circuit board, and the analog signal converter is set with a network port and / or a telephone interface to connect to the network or telephone line; this invention can accurately and quickly translate according to Chinese and Russian, and optimize the professional terms of the wood industry, ensuring that the recognition rate is 99% or above, which serves as a bridge for communication between China and Russia and solves the communication problems between China and Russia.

[0004] However, although the above technology can be applied to business communication translation, there are still translation ambiguity problems in practical applications, such as polysemy and context dependence, cultural differences in industry terms, etc. The above technology does not explain the specific technical solution for controlling the semantic ambiguity rate, which will lead to inaccurate translation results, thereby affecting the communication effect of both parties, causing misunderstandings or wrong decisions, and reducing the efficiency and trust of business communication.

[0005] In view of this, the present invention proposes a business communication intelligent translation system based on natural language processing to solve the above problems. Summary of the invention

[0006] In order to overcome the above-mentioned defects of the prior art and to achieve the above-mentioned purpose, the present invention provides the following technical solution: a business communication intelligent translation system based on natural language processing, comprising:

[0007] A text acquisition module is used to acquire business texts;

[0008] The term segmentation module is used to segment business texts and obtain industry terms;

[0009] Industry determination module, which determines the industry to which the business text belongs based on industry terminology;

[0010] The context analysis module is used to analyze business texts and obtain context information;

[0011] The syntax analysis module performs syntax analysis on business texts based on context information and constructs a syntax tree;

[0012] The text translation module integrates the industry, context information and syntax tree to translate business texts.

[0013] Furthermore, the step of acquiring industry terms includes:

[0014] Step S101: cleaning the business text to obtain a standardized text;

[0015] Step S102: performing word segmentation processing on the standardized text to obtain independent words;

[0016] Step S103: filtering out industry terms from independent words;

[0017] In step S101, text cleaning includes deleting punctuation marks and special characters;

[0018] In step S102, the method for obtaining independent words includes:

[0019] A conditional probability model is constructed, the standardized text is input into the conditional probability model, and a word label sequence is output, in which the word labels in the word label sequence correspond one-to-one to the characters in the standardized text; the word labels include B, M, E and S, where B indicates that the corresponding character is the beginning part of a word, M indicates that the corresponding character is the middle part of a word, E indicates that the corresponding character is the end part of a word, and S indicates that the corresponding character is a single-character word; according to the word label sequence, a word set is constructed, in which each character in the word set constitutes an independent word.

[0020] Furthermore, the step of constructing a word set includes:

[0021] Step S201: obtaining the word tag that is at the front of the word tag sequence and marking it as the current tag;

[0022] Step S202: if the current label is B, then a new word set is constructed and marked as the latest set, and the corresponding characters are added to the latest set; if the current label is M or E, then the corresponding characters are added to the latest set; if the current label is S, then a new word set is constructed, but not marked as the latest set, and the corresponding characters are added to the new word set;

[0023] Step S203: deleting the current tag from the word tag sequence;

[0024] Step S204: looping steps S201 to S203 until there is no word tag in the word tag sequence, and then the loop ends.

[0025] Furthermore, in step S103, the method of filtering out industry terms from independent words includes:

[0026] A term library is preset, and the term library includes a set of industry terminology corresponding to each industry; a pre-trained word vector model is used to convert industry terms and independent words in the term library into corresponding word vectors respectively; the word vectors corresponding to the industry terms are marked as first vectors, and the word vectors corresponding to the independent words are marked as second vectors; the cosine similarity between each second vector and each first vector is calculated, and marked as similarity values; a similarity threshold is preset, and each similarity value is compared with the similarity threshold respectively; the second vector corresponding to the similarity value with a value greater than the similarity threshold is marked as a matching vector, and the second vector corresponding to the value less than or equal to the similarity threshold is not marked; from all independent words, the independent words corresponding to the matching vector are screened out and used as industry terms.

[0027] Furthermore, the method for determining the industry to which the business text belongs includes:

[0028] All industry terms in the term base are marked as classification terms, and the number of documents studied when obtaining classification terms is counted and marked as the number of documents; the number of times each classification term appears in each document is counted and marked as the frequency of occurrence; the number of words in each document is counted and marked as the total number of words; based on the frequency of occurrence and the total number of words, the relative frequency of each classification term in each document is calculated; based on the number of documents, the inverse document frequency of each classification term is calculated; based on the relative frequency and the inverse document frequency, the term weight of each classification term for each document is calculated;

[0029] Add up the term weights of the same classification terms in the corresponding literature of each industry in turn to obtain the industry weight of each classification term for each industry; mark the selected industry terms as screened terms, and add up the industry weights of all screened terms corresponding to the same industry in turn to obtain the total weight value of the business text for each industry; sort all the total weight values from largest to smallest, and take the industry corresponding to the largest total weight value as the industry to which the business text belongs.

[0030] Further, the expression of the relative frequency is: In the formula, xp(a, b) is the relative frequency of classification term a in literature b, cp(a, b) is the occurrence frequency of classification term a in literature b, cz(b) is the total number of words in literature b, a ∈ [1, A], b ∈ [1, C], A is the total number of classification terms in the thesaurus, and C is the number of literatures;

[0031] The expression of the inverse document frequency is: In the formula, nw(a) is the inverse document frequency of classification term a, and wx(a) is the number of literatures containing classification term a;

[0032] The expression of the term weight is: sq(a, b) = xp(a, b) × nw(a); in the formula, sq(a, b) is the term weight of classification term a for literature b.

[0033] Further, the method for obtaining context information includes:

[0034] Divide the business text into sentences according to punctuation marks to obtain independent sentences; count the number of independent words in each independent sentence and mark it as the word count; count the number of industry terms in each independent sentence and mark it as the term count; calculate the complexity fz corresponding to each independent sentence according to the word count and the term count;

[0035] Preset a complexity threshold and a normal window size W 0 , the complexity threshold includes a first threshold TH 1 and a second threshold TH 2 , 0 < TH 1 < TH 2 ; compare the complexity of each independent sentence with the complexity threshold respectively; if TH 1 ≤ fz ≤ TH 2 , then set the initial window size of the corresponding independent sentence to W 1 ; if TH 1 > fz, then set the initial window size of the corresponding independent sentence to W 2 ; if TH 2 < fz, then set the initial window size of the corresponding independent sentence to W 3 ;

[0036] According to the initial window size of each independent sentence, the context sentence corresponding to each independent sentence is obtained; wherein the context sentence includes the previous sentence and the following sentence, and W 1 , W 2 , W 3 All marked as W 4 , for an independent sentence with index i, i∈[1,I], I is the number of independent sentences in the business text; the previous sentence includes The independent sentence with index i-1 to the independent sentence with index i-1, and the latter sentence includes the independent sentence with index i+1 to the independent sentence with index i-1. independent sentences; count the number of independent sentences in the context sentences corresponding to each independent sentence and mark them as the number of sentences; add the complexity of all independent sentences in the context sentences corresponding to each independent sentence in turn, and then divide it by the corresponding number of sentences to obtain the context complexity corresponding to each independent sentence; calculate the final window size of each independent sentence according to the context complexity of each independent sentence; according to the final window size of each independent sentence, re-obtain the context sentences corresponding to each independent sentence and use them as context information.

[0037] Furthermore, the complexity calculation method includes:

[0038]

[0039] In the formula, fz is the complexity, cs is the number of words, ss is the number of terms, α, β, γ, δ are weight coefficients, and p and q are exponential coefficients;

[0040] W 1 =W 0 , W 2 =W 0 -θ 1 ×(TH 1 -fz), W 3 =W 0 +θ 2 ×(fz-TH 2 );where θ 1 ,θ 2 is the adjustment factor;

[0041] The calculation method of the final window size includes:

[0042]

[0043] Where W 5 is the final window size, sz is the context complexity, k 1 , k 2 is the adjustment factor.

[0044] Furthermore, the method for constructing a syntax tree includes:

[0045] Perform word segmentation on each independent sentence, obtain independent words in each independent sentence, and mark them as sentence items; take the sentence items corresponding to each independent sentence as a set of item sets, and the item sets correspond to the independent sentences one by one; input each item set into the trained part-of-speech tagging model, and predict the corresponding set label, the part-of-speech tagging model is a deep neural network model; the set label is the digital label corresponding to the part-of-speech set, different part-of-speech sets have different digital labels, and the part-of-speech set includes the part-of-speech corresponding to each sentence item in the item set; according to the predicted set label, obtain the part-of-speech set corresponding to each item set;

[0046] Using a pre-trained word vector model, each sentence item is converted into a corresponding word vector and marked as a third vector; the third vector corresponding to each independent sentence is used as a set of vector sets, and the vector set corresponds to the independent sentence one by one; different numerical labels are set for different parts of speech and marked as part-of-speech tags; the parts of speech in each part-of-speech set are replaced with corresponding part-of-speech tags and marked as a replacement set; the independent sentences in the context information corresponding to each independent sentence are marked as context sentences, and the vector set and replacement set corresponding to the context sentence corresponding to each independent sentence are used as a context set; the replacement set, vector set and context set corresponding to the same independent sentence are merged as a merged set, and the merged set corresponds to the independent sentence one by one; each merged set is input into the trained Bi LSTM model respectively, and the corresponding dependency set is output; the dependency set includes the dependent words and dependency type of each sentence item in the independent sentence; according to the dependency set, the syntactic tree corresponding to each independent sentence is constructed.

[0047] Furthermore, the method for translating a business text includes:

[0048] Different digital labels are set for different industries and marked as industry labels; different digital labels are set for different languages ​​and marked as language labels; the target language is obtained, and the corresponding language label is obtained according to the target language, and marked as the target label; different digital labels are set for different dependency types and marked as type labels; in the dependency set corresponding to each independent sentence, the dependency type is replaced with the corresponding relationship label, the dependency word is replaced with the corresponding third vector, and the dependency set after the replacement is marked as a relationship replacement set; the context set, relationship replacement set, industry label and target label corresponding to each independent sentence in the business text are taken as a set of sets to be translated, and the sets to be translated correspond to the independent sentences in the business text one by one; each set to be translated is input into the trained Transformer model respectively, and the corresponding translation set is output, and the translation set includes the translation result of each sentence item in the independent sentence in the target language.

[0049] The technical effects and advantages of the business communication intelligent translation system based on natural language processing of the present invention are as follows:

[0050] By accurately identifying industry terms in business texts, it is possible to effectively identify the industry in the business text, thereby better understanding the professional content and background information of the business text; by using the method of adaptively adjusting the size of the context window, it is possible to accurately capture the context information of each independent sentence; by performing syntactic analysis on the business text and constructing a syntax tree, and translating in combination with dynamically adjusted context information, it is possible to quickly and accurately translate the business text, overcoming language barriers and translation ambiguity problems in traditional translation methods; thereby greatly improving the efficiency and reliability of business text translation, ensuring accurate understanding and communication of information, and further improving the efficiency and accuracy of business communication, providing intelligent and efficient support for enterprises' cross-language international business communication, and promoting the smooth progress of international business cooperation. BRIEF DESCRIPTION OF THE DRAWINGS

[0051] Figure 1 A schematic diagram of an intelligent translation system for business communication based on natural language processing according to Embodiment 1 of the present invention;

[0052] Figure 2 This is a flow chart of a method for constructing a word set according to Embodiment 1 of the present invention. DETAILED DESCRIPTION

[0053] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0054] Example 1

[0055] See also Figure 1 As shown, the intelligent translation system for business communication based on natural language processing described in this embodiment includes a text acquisition module, a terminology division module, an industry determination module, a context analysis module, a syntax analysis module and a text translation module; each module is connected by wired and / or wireless means to realize data transmission between modules.

[0056] The text acquisition module is used to acquire business texts.

[0057] Business texts include meeting minutes (including discussion points, decisions, action items, etc.), business emails (including cooperation proposals, customer feedback, etc.), contract documents (including agreements, terms, etc.), etc. Business texts are obtained through user uploads and support multiple file formats (such as DOC, TXT, PDF, etc.); at the same time, the system allows batch upload of multiple business texts to improve translation efficiency.

[0058] The term segmentation module is used to segment business texts and obtain industry terms.

[0059] Steps to capture industry terminology include:

[0060] Step S101: cleaning the business text to obtain a standardized text;

[0061] Step S102: performing word segmentation processing on the standardized text to obtain independent words;

[0062] Step S103: Filter out industry terms from independent words.

[0063] In the above step S101, text cleaning includes deleting punctuation marks (such as periods, commas, exclamation marks, etc.) and special characters (such as &, #, $, etc.), and text cleaning is achieved through regular expressions.

[0064] In the above step S102, the method of obtaining independent words includes:

[0065] A conditional probability model (CRF) is constructed, and the standardized text is input into the conditional probability model to output a word label sequence. The word labels in the word label sequence correspond one-to-one to the characters in the standardized text. The conditional probability model is a prior art, and the specific construction process is not described in detail here; the word labels include B, M, E and S, where B indicates that the corresponding character is the beginning part of a word, M indicates that the corresponding character is the middle part of a word, E indicates that the corresponding character is the end part of a word, and S indicates that the corresponding character is a single-character word; according to the word label sequence, a word set is constructed, in which each character in the word set constitutes an independent word.

[0066] like Figure 2 As shown, the steps of constructing a word set include:

[0067] Step S201: obtaining the word tag that is at the front of the word tag sequence and marking it as the current tag;

[0068] Step S202: if the current label is B, then a new word set is constructed and marked as the latest set, and the corresponding characters are added to the latest set; if the current label is M or E, then the corresponding characters are added to the latest set; if the current label is S, then a new word set is constructed but not marked as the latest set, and the corresponding characters are added to the new word set; it should be understood that only the latest constructed word set is marked as the latest set, that is, when a new word set is constructed, the word set originally marked as the latest set is unmarked as the latest set;

[0069] Step S203: deleting the current tag from the word tag sequence;

[0070] Step S204: looping steps S201 to S203 until there is no word tag in the word tag sequence, and then the loop ends.

[0071] For example, the standardized text is "I love Beijing Tiananmen", and after inputting into the conditional probability model, the obtained word tag sequence is SSBEBME; therefore, the word set is "I", "love", "Beijing" and "Tiananmen".

[0072] In the above step S103, the method of filtering out industry terms from independent words includes:

[0073] A term library is preset, and the term library includes a set of industry terms corresponding to each industry; the industry terms in the term library are obtained by technical personnel in this field through literature research; a pre-trained word vector model (such as GloVe, Word2Vec, etc.) is used to convert the industry terms and independent words in the term library into corresponding word vectors respectively; the word vectors corresponding to the industry terms are marked as first vectors, and the word vectors corresponding to the independent words are marked as second vectors; the cosine similarity between each second vector and each first vector is calculated, and marked as similarity values. The calculation process of cosine similarity is a prior art and will not be described in detail here; a similarity threshold is preset, and the similarity threshold is pre-set by technical personnel in this field according to actual conditions; each similarity value is compared with the similarity threshold, and the second vector corresponding to the similarity value with a value greater than the similarity threshold is marked as a matching vector, and the second vector corresponding to the value less than or equal to the similarity threshold is not marked; from all independent words, the independent words corresponding to the matching vector are screened out and used as industry terms.

[0074] The industry determination module determines the industry to which the business text belongs based on industry terminology.

[0075] Methods for determining the industry to which a business text belongs include:

[0076] All industry terms in the term base are marked as classification terms. The number of documents studied when obtaining classification terms is counted and marked as the number of documents. The number of times each classification term appears in each document is counted and marked as the frequency of occurrence. The number of words in each document is counted and marked as the total number of words. Based on the frequency of occurrence and the total number of words, the relative frequency of each classification term in each document is calculated. The expression of relative frequency is: Where xp(a,b) is the relative frequency of classification term a in document b, cp(a,b) is the frequency of occurrence of classification term a in document b, cz(b) is the total number of words in document b, a∈[1,A], b∈[1,C], A is the total number of classification terms in the term base, and C is the number of documents. Based on the number of documents, the inverse document frequency of each classification term is calculated. The expression of inverse document frequency is: Where nw(a) is the inverse document frequency of classification term a, wx(a) is the number of documents containing classification term a, which is obtained by the technical personnel in the field when presetting the term base; the term weight of each classification term for each document is calculated according to the relative frequency and the inverse document frequency; the expression of the term weight is: sq(a,b)=xp(a,b)×nw(a); where sq(a,b) is the term weight of classification term a for document b;

[0077] The term weights of the same classification terms in the corresponding documents of each industry are added up in sequence to obtain the industry weight of each classification term for each industry; the screened industry terms are marked as screening terms, and the industry weights of all the screening terms corresponding to the same industry are added up in sequence to obtain the total weight of the business text for each industry; all the total weight values ​​are sorted from large to small, and the industry corresponding to the front-most total weight value is taken as the industry to which the business text belongs.

[0078] The context analysis module is used to perform text analysis on business texts and obtain context information.

[0079] Methods for obtaining contextual information include:

[0080] According to punctuation marks, the business text is divided into sentences to obtain independent sentences; the number of independent words in each independent sentence is counted and marked as the number of words; the number of industry terms in each independent sentence is counted and marked as the number of terms; according to the number of words and the number of terms, the complexity fz corresponding to each independent sentence is calculated; the calculation method of the complexity fz includes:

[0081]

[0082] In the formula, cs is the number of words, ss is the number of terms, α, β, γ, and δ are all weight coefficients, and p and q are both exponential coefficients. The specific values of the weight coefficients and exponential coefficients in the formula can be set according to the actual situation. The weight coefficients and exponential coefficients reflect the influence degrees of the number of words and the number of terms on the complexity of an independent sentence. Those skilled in the art can preset the corresponding weight coefficients and exponential coefficients according to the actual influence degrees of the number of words and the number of terms on the complexity of an independent sentence, so as to accurately evaluate the translation complexity of each independent sentence. It should be understood that both the number of words and the number of terms are influence parameters of complexity. Among them, the more the number of words, the longer the corresponding independent sentence, that is, it carries more information, so the complexity during translation is greater, and vice versa. The more the number of terms, the stronger the professionalism of the corresponding independent sentence, and targeted translation needs to be carried out according to the knowledge of a specific field, so the complexity during translation is greater, and vice versa.

[0083] Preset the complexity threshold and the normal window size W 0 , the complexity threshold includes the first threshold TH 1 and the second threshold TH 2 , 0 < TH 1 < TH 2 , the complexity threshold and the normal window size are both preset by those skilled in the art according to the actual situation; compare the complexity of each independent sentence with the complexity threshold respectively; if TH 1 ≤ fz ≤ TH 2 , then set the initial window size of the corresponding independent sentence to W 1 ; if TH 1 > fz, then set the initial window size of the corresponding independent sentence to W 2 ; if TH 2 < fz, then set the initial window size of the corresponding independent sentence to W 3 ;

[0084] Among them, W 1 = W 0 , W 2 = W 0 - θ 1 × (TH 1 - fz), W 3 = W 0 + θ 2 × (fz - TH 2 ); in the formula, θ 1 , θ 2 are adjustment factors, and the adjustment factors are preset by those skilled in the art according to the actual situation.

[0085] According to the initial window size of each independent sentence, the context sentence corresponding to each independent sentence is obtained; wherein the context sentence includes the previous sentence and the following sentence, and W 1 , W 2 , W 3 All marked as W 4 , for an independent sentence with index i (i.e., the i-th sentence), i∈[1,I], where I is the number of independent sentences in the business text; the previous sentence includes the one with index The independent sentence with index i-1 to the independent sentence with index i-1, and the latter sentence includes the independent sentence with index i+1 to the independent sentence with index i-1. independent sentences; count the number of independent sentences in the context sentences corresponding to each independent sentence and mark them as the number of sentences; add the complexity of all independent sentences in the context sentences corresponding to each independent sentence in turn, and then divide it by the corresponding number of sentences to obtain the context complexity corresponding to each independent sentence; calculate the final window size of each independent sentence according to the context complexity of each independent sentence; according to the final window size of each independent sentence, re-obtain the context sentences corresponding to each independent sentence and use them as context information.

[0086] The calculation method of the final window size includes:

[0087]

[0088] Where W 5 is the final window size, sz is the context complexity, k 1 , k 2 It is an adjustment factor, which is preset by technicians in this field according to actual conditions.

[0089] It should be noted that W 1 , W 2 , W 3 , W 5 If the value is not an integer during the calculation process, it will be rounded up.

[0090] The syntax analysis module performs syntax analysis on business texts based on context information and constructs a syntax tree;

[0091] Methods for constructing syntax trees include:

[0092] Perform word segmentation on each independent sentence, obtain the independent words in each independent sentence, and mark them as sentence items; it should be noted that the method of word segmentation on independent sentences is consistent with the method of word segmentation on standardized text; the sentence items corresponding to each independent sentence are taken as a set of item sets, and the item sets correspond to the independent sentences one by one; each item set is input into the trained part-of-speech tagging model, and the corresponding set label is predicted; the set label is the digital label corresponding to the part-of-speech set, and different part-of-speech sets have different digital labels. The part-of-speech set includes the part-of-speech corresponding to each sentence item in the item set, such as noun, verb, adjective, etc.; according to the predicted set label, obtain the part-of-speech set corresponding to each item set;

[0093] The training process of the part-of-speech tagging model includes:

[0094] Collect r groups of word sets in advance, set corresponding set labels for the r groups of word sets, where r is an integer greater than 1, and convert the word sets and the corresponding set labels into a corresponding set of feature vectors; the set labels corresponding to the word sets are collected by those skilled in the art in the art in the process of marking the parts of speech of sentence words in the past, and the parts of speech are marked for each sentence word in each group of word sets in turn, and the parts of speech corresponding to the sentence words of each word set are taken as a group of part-of-speech sets, and different digital labels are set for different part-of-speech sets, and marked as set labels, and corresponding set labels are set for the r groups of word sets in turn;

[0095] Each set of feature vectors is used as the input of the part-of-speech tagging model. The part-of-speech tagging model outputs a set of predicted set labels corresponding to each set of terms, and uses the actual set labels corresponding to each set of terms as the prediction target. The actual set labels are the pre-set set labels corresponding to the term sets. The training goal is to minimize the sum of the prediction errors of all term sets. The calculation formula for the prediction error is η d =(μ d -ε d ) 2 , where η d is the prediction error, d is the group number of the feature vector corresponding to the term set, μ d is the predicted set label corresponding to the d-th set of terms, ε d is the actual set label corresponding to the dth group of word items; the part-of-speech tagging model is trained until the sum of the prediction errors reaches convergence and the training is stopped.

[0096] The above-mentioned part-of-speech tagging model is specifically a deep neural network model; it includes an input layer, a hidden layer and an output layer; each hidden layer includes multiple neurons, each neuron is connected to the neurons in the next layer, and the connection contains weights, which determine the importance and influence of data transmission in the neural network; an activation function is applied to each neuron between the hidden layer and the output layer, and the activation function introduces nonlinearity, allowing the network to learn more complex patterns and features.

[0097] Using a pre-trained word vector model, each sentence item is converted into a corresponding word vector and marked as a third vector; the third vector corresponding to each independent sentence is used as a set of vector sets, and the vector set corresponds to the independent sentence one by one; different parts of speech are set with different numerical labels and marked as part-of-speech tags; the parts of speech in each part-of-speech set are replaced with corresponding part-of-speech tags and marked as a replacement set; the independent sentences in the context information corresponding to each independent sentence are marked as context sentences, and the vector set and replacement set corresponding to the context sentence of each independent sentence are used as a context set; the replacement set, vector set and context set corresponding to the same independent sentence are merged as a merged set, and the merged set corresponds to the independent sentence one by one; The union sets are respectively input into the trained BiLSTM model, and the corresponding dependency set is output. It should be noted that the BiLSTM model is a prior art and will not be described in detail here. The dependency set includes the dependent words and dependency types of each sentence item in the independent sentence. The dependent word is another sentence item that a sentence item in the independent sentence depends on, indicating the relationship between words in the grammatical structure. The dependency type is the specific relationship between the sentence item and its dependent word, such as the subject-predicate relationship (describing the relationship between the subject and the predicate), the verb-object relationship (describing the relationship between the verb and its direct object), the preposition-object relationship (describing the relationship between the preposition and its object), etc. According to the dependency set, a syntactic tree corresponding to each independent sentence is constructed.

[0098] The text translation module integrates the industry, context information and syntax tree to translate business texts.

[0099] Methods for translating business texts include:

[0100] Different digital labels are set for different industries and marked as industry labels; different digital labels are set for different languages ​​and marked as language labels; the target language is obtained, the corresponding language label is obtained according to the target language, and marked as the target label, and the target language is set by the user; different digital labels are set for different dependency types and marked as type labels; in the dependency set corresponding to each independent sentence, the dependency type is replaced with the corresponding relationship label, the dependency word is replaced with the corresponding third vector, and the dependency set after the replacement is marked as a relationship replacement set; the context set, relationship replacement set, industry label and target label corresponding to each independent sentence in the business text are used as a set of sets to be translated, and the sets to be translated correspond to the independent sentences in the business text one by one; each set to be translated is input into a trained Transformer model (such as T5 model, BART model, etc.), and the corresponding translation set is output, and the translation set includes the translation results of each sentence item in the independent sentence in the target language; it should be noted that the Transformer model is a prior art, and the specific training process is not described in detail here.

[0101] This embodiment can effectively identify the industry in the business text by accurately identifying the industry terms in the business text, so as to better understand the professional content and background information of the business text; adopt the method of adaptively adjusting the context window size to accurately capture the context information of each independent sentence; by performing syntactic analysis on the business text and constructing a syntax tree, and translating in combination with the dynamically adjusted context information, the business text can be translated quickly and accurately, overcoming the language barriers and translation ambiguity problems in traditional translation methods; thereby greatly improving the efficiency and reliability of business text translation, ensuring the accurate understanding and communication of information, and further improving the efficiency and accuracy of business communication, providing intelligent and efficient support for enterprises' cross-language international business communication, and promoting the smooth progress of international business cooperation.

[0102] Example 2

[0103] The present application also provides an electronic device. The electronic device may include one or more processors and one or more memories. The memories store computer-readable codes, and when the computer-readable codes are executed by the one or more processors, the business communication intelligent translation system based on natural language processing as described above may be executed.

[0104] The method or system according to the implementation mode of the present application can also be implemented with the help of the architecture of the electronic device shown in the present application. The electronic device may include a bus, one or more CPUs, ROM, RAM, a communication port connected to a network, input / output, a hard disk, etc. A storage device in the electronic device, such as a ROM or a hard disk, can store a business communication intelligent translation system based on natural language processing provided by the present application. Furthermore, the electronic device may also include a user interface. Of course, the architecture shown in the present application is only exemplary. When implementing different devices, one or more components in the electronic device shown in the present application may be omitted according to actual needs.

[0105] Example 3

[0106] One embodiment of the present application discloses a computer-readable storage medium. Computer-readable instructions are stored on the computer-readable storage medium. When the computer-readable instructions are executed by a processor, a business communication intelligent translation system based on natural language processing according to an embodiment of the present application described with reference to the above figures can be executed. The storage medium includes, but is not limited to, for example, volatile memory and / or non-volatile memory. Volatile memory may include, for example, random access memory (RAM) and cache memory (cache), etc. Non-volatile memory may include, for example, read-only memory (ROM), hard disk, flash memory, etc.

[0107] In addition, according to the embodiments of the present application, the process described above with reference to the flowchart can be implemented as a computer software program. For example, the present application provides a non-transitory machine-readable storage medium, which stores machine-readable instructions, and the machine-readable instructions can be executed by a processor to execute instructions corresponding to the method steps provided by the present application, for example: a business communication intelligent translation system based on natural language processing. When the computer program is executed by a central processing unit (CPU), the above functions defined in the method of the present application are executed.

[0108] The above is only a specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed by the present invention, which should be included in the protection scope of the present invention. Therefore, the protection scope of the present invention should be based on the protection scope of the claims.

[0109] Finally: The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the protection scope of the present invention.

Claims

1. A business communication intelligent translation system based on natural language processing, characterized in that: include: A text acquisition module is used to acquire business texts; The term segmentation module is used to segment business texts and obtain industry terms; Industry determination module, which determines the industry to which the business text belongs based on industry terminology; The context analysis module is used to analyze business texts and obtain context information; The syntax analysis module performs syntax analysis on business texts based on context information and constructs a syntax tree; The text translation module integrates the industry, context information and syntax tree to translate business texts.

2. According to claim 1, a business communication intelligent translation system based on natural language processing is characterized in that: The steps of obtaining industry terminology include: Step S101: cleaning the business text to obtain a standardized text; Step S102: performing word segmentation processing on the standardized text to obtain independent words; Step S103: filtering out industry terms from independent words; In step S101, text cleaning includes deleting punctuation marks and special characters; In step S102, the method for obtaining independent words includes: A conditional probability model is constructed, the standardized text is input into the conditional probability model, and a word label sequence is output, in which the word labels in the word label sequence correspond one-to-one to the characters in the standardized text; the word labels include B, M, E and S, where B indicates that the corresponding character is the beginning part of a word, M indicates that the corresponding character is the middle part of a word, E indicates that the corresponding character is the end part of a word, and S indicates that the corresponding character is a single-character word; according to the word label sequence, a word set is constructed, in which each character in the word set constitutes an independent word.

3. According to claim 2, a business communication intelligent translation system based on natural language processing is characterized in that: The step of constructing a word set comprises: Step S201: Obtain the word tag that is at the front of the word tag sequence and mark it as the current tag; Step S202: if the current label is B, then a new word set is constructed and marked as the latest set, and the corresponding characters are added to the latest set; if the current label is M or E, then the corresponding characters are added to the latest set; if the current label is S, then a new word set is constructed, but not marked as the latest set, and the corresponding characters are added to the new word set; Step S203: deleting the current tag from the word tag sequence; Step S204: looping steps S201 to S203 until there is no word tag in the word tag sequence, and then the loop ends.

4. The business communication intelligent translation system based on natural language processing according to claim 3 is characterized in that: In step S103, the method of filtering out industry terms from independent words includes: A term library is preset, and the term library includes a set of industry terminology corresponding to each industry; a pre-trained word vector model is used to convert industry terms and independent words in the term library into corresponding word vectors respectively; the word vectors corresponding to the industry terms are marked as first vectors, and the word vectors corresponding to the independent words are marked as second vectors; the cosine similarity between each second vector and each first vector is calculated, and marked as similarity values; a similarity threshold is preset, and each similarity value is compared with the similarity threshold respectively; the second vector corresponding to the similarity value with a value greater than the similarity threshold is marked as a matching vector, and the second vector corresponding to the value less than or equal to the similarity threshold is not marked; from all independent words, the independent words corresponding to the matching vector are screened out and used as industry terms.

5. The business communication intelligent translation system based on natural language processing according to claim 4 is characterized in that: The method for determining the industry to which the business text belongs includes: All industry terms in the term base are marked as classification terms. The number of research documents when obtaining classification terms is counted and marked as the number of documents. The number of times each classification term appears in each document is counted and marked as the frequency of occurrence. The number of words in each document is counted and marked as the total number of words. According to the frequency of occurrence and the total number of words, the relative frequency of each classification term in each document is calculated. According to the number of documents, the inverse document frequency of each classification term is calculated. According to the relative frequency and the inverse document frequency, the term weight of each classification term for each document is calculated. The term weights of the same classification terms in the documents corresponding to each industry are added in sequence to obtain the industry weight of each classification term for each industry. The selected industry terms are marked as screened terms, and the industry weights corresponding to the same industry of all screened terms are added in sequence to obtain the total weight value of the business text for each industry. All the total weight values are sorted from largest to smallest, and the industry corresponding to the largest total weight value is used as the industry to which the business text belongs.

6. The business communication intelligent translation system based on natural language processing according to claim 5, characterized in that: The relative frequency expression is: Where xp(a,b) is the relative frequency of classification term a in document b, cp(a,b) is the frequency of occurrence of classification term a in document b, cz(b) is the total number of words in document b, a∈[1,A], b∈[1,C], A is the total number of classification terms in the term base, and C is the number of documents; The expression of the inverse document frequency is: Where nw(a) is the inverse document frequency of classification term a, and wx(a) is the number of documents containing classification term a; The expression of the term weight is: sq(a,b) = xp(a,b) × nw(a); where sq(a,b) is the term weight of classification term a for document b.

7. The intelligent translation system for business communication based on natural language processing according to claim 6, characterized in that: The method for obtaining context information includes: According to punctuation marks, the business text is divided into sentences to obtain independent sentences. The number of independent words in each independent sentence is counted and marked as the number of words. The number of industry terms in each independent sentence is counted and marked as the number of terms. According to the number of words and the number of terms, the complexity fz of each independent sentence is calculated. A preset complexity threshold and a normal window size W0 are set. The complexity threshold includes a first threshold TH1 and a second threshold TH2, where 0 < TH1 < TH2. The complexity of each independent sentence is compared with the complexity threshold respectively. If TH1 ≤ fz ≤ TH2, the initial window size of the corresponding independent sentence is set to W1. If TH1 > fz, the initial window size of the corresponding independent sentence is set to W2. If TH2 < fz, the initial window size of the corresponding independent sentence is set to W3. According to the initial window size of each independent sentence, the context sentence corresponding to each independent sentence is obtained; wherein the context sentence includes the previous sentence and the next sentence, W1, W2, and W3 are all marked as W4, for the independent sentence with index i, i∈[1,I], I is the number of independent sentences in the business text; the previous sentence includes the index The independent sentence from index i to the independent sentence with index i-1, and the latter sentence includes the independent sentence with index i+1 to the independent sentence with index independent sentences; count the number of independent sentences in the context sentences corresponding to each independent sentence and mark them as the number of sentences; add the complexity of all independent sentences in the context sentences corresponding to each independent sentence in turn, and then divide it by the corresponding number of sentences to obtain the context complexity corresponding to each independent sentence; calculate the final window size of each independent sentence according to the context complexity of each independent sentence; according to the final window size of each independent sentence, re-obtain the context sentences corresponding to each independent sentence and use them as context information.

8. The business communication intelligent translation system based on natural language processing according to claim 7 is characterized in that: The calculation method of the complexity includes: where fz is the complexity, cs is the number of words, ss is the number of terms, α, β, γ, δ are all weight coefficients, and p, q are all exponential coefficients. W1 = W0, W2 = W0 - θ1 × (TH1 - fz), W3 = W0 + θ2 × (fz - TH2); where θ1, θ2 are adjustment factors. The calculation method of the final window size includes: where W5 is the final window size, sz is the context complexity, and k1, k2 are adjustment factors.

9. The business communication intelligent translation system based on natural language processing according to claim 8, characterized in that: The method for constructing a syntactic tree includes: Perform word segmentation on each independent sentence, obtain independent words in each independent sentence, and mark them as sentence items; take the sentence items corresponding to each independent sentence as a set of item sets, and the item sets correspond to the independent sentences one by one; input each item set into the trained part-of-speech tagging model, and predict the corresponding set label, the part-of-speech tagging model is a deep neural network model; the set label is the digital label corresponding to the part-of-speech set, different part-of-speech sets have different digital labels, and the part-of-speech set includes the part-of-speech corresponding to each sentence item in the item set; according to the predicted set label, obtain the part-of-speech set corresponding to each item set; Using a pre-trained word vector model, each sentence item is converted into a corresponding word vector and marked as a third vector; the third vector corresponding to each independent sentence is used as a set of vector sets, and the vector set corresponds to the independent sentence one by one; different numerical labels are set for different parts of speech and marked as part-of-speech tags; the parts of speech in each part-of-speech set are replaced with corresponding part-of-speech tags and marked as a replacement set; the independent sentences in the context information corresponding to each independent sentence are marked as context sentences, and the vector set and replacement set corresponding to the context sentence corresponding to each independent sentence are used as a context set; the replacement set, vector set and context set corresponding to the same independent sentence are merged as a merged set, and the merged set corresponds to the independent sentence one by one; each merged set is input into the trained BiLSTM model respectively, and the corresponding dependency set is output; the dependency set includes the dependent words and dependency type of each sentence item in the independent sentence; according to the dependency set, the syntactic tree corresponding to each independent sentence is constructed.

10. The intelligent translation system for business communication based on natural language processing according to claim 9, characterized in that: The method for translating a business text comprises: Different digital labels are set for different industries and marked as industry labels; different digital labels are set for different languages ​​and marked as language labels; the target language is obtained, and the corresponding language label is obtained according to the target language, and marked as the target label; different digital labels are set for different dependency types and marked as type labels; in the dependency set corresponding to each independent sentence, the dependency type is replaced with the corresponding relationship label, the dependency word is replaced with the corresponding third vector, and the dependency set after the replacement is marked as a relationship replacement set; the context set, relationship replacement set, industry label and target label corresponding to each independent sentence in the business text are taken as a set of sets to be translated, and the sets to be translated correspond to the independent sentences in the business text one by one; each set to be translated is input into the trained Transformer model respectively, and the corresponding translation set is output, and the translation set includes the translation result of each sentence item in the independent sentence in the target language.

Citation Information

Patent Citations

  • Sino -Russian timber specialized vocabulary translation system

    CN205910694U