Business text real-time translation method and system based on artificial intelligence
By adopting multi-agent collaborative architecture and dynamic termbase interface module in business text translation technology, the problems of imbalance between scenario generalization and specialization, insufficient static feature adaptation, efficiency and real-time bottlenecks and lack of user feedback in the existing technology are solved, and efficient, accurate and real-time translation effects are achieved in specific business scenarios.
Patent Information
- Application Number
- CN202510374657.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-27
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2045-03-27
AI Technical Summary
The existing business text translation technology has problems such as imbalance in scenario generalization and specialization, insufficient static feature adaptation, efficiency and real-time bottlenecks, and lack of user feedback mechanisms in actual business scenarios.
Using the real-time translation method of business text based on artificial intelligence, multiple network models are designed by collecting data in different business scenarios, and using loss functions that integrate overall loss, semantic loss, stylistic loss and cultural loss for training, several agents are obtained. User text data is collected in real time, and word frequency statistics and style classification are used using the TF-IDF algorithm and BERT model, match the scene and select the appropriate agent for translation, and dynamically replace the agent according to user feedback.
Adaptive translation in specific business scenarios is realized, the accuracy and real-time of translation are improved, the user experience is enhanced, and the translation strategy can be dynamically adjusted to meet the needs of different scenarios.
Smart Images

Figure CN120218090A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of text translation, and in particular to a real-time business text translation method and system based on artificial intelligence. Background Art
[0002] Neural machine translation is abbreviated as NMT. Existing business text translation technologies are mainly based on the NMT framework, with the Transformer model as the core, and realize end-to-end text conversion through training on a large-scale bilingual corpus. The general NMT model captures context dependencies through the self-attention mechanism and performs excellently in open-domain translation tasks. In specific implementation, after the input text is extracted with semantic features by the encoder, the decoder generates the target language sequence, and the model parameters are optimized through the loss function. Some improvement methods use domain adaptation technology to inject a small amount of business domain data based on the general model to improve the translation accuracy of professional terms.
[0003] However, the above technologies have significant defects in actual business scenarios. Imbalance between scenario generalization and specialization: A single model is difficult to balance the different needs of different scenarios, resulting in term mistranslation and style mismatch; Insufficient adaptation of static features: Traditional methods rely on fixed term libraries and weight assignments and cannot dynamically adjust translation strategies. Bottlenecks in efficiency and real-time performance: Deep Transformer models have high computational complexity, and the average response time exceeds 3 seconds, making it difficult to meet the requirements of high-real-time scenarios such as business negotiations; Lack of user feedback mechanism: When translation errors occur, the model cannot be replaced and regenerated through user feedback. Summary of the Invention
[0004] In order to overcome the deficiencies of the prior art, the purpose of the present invention is to provide a real-time business text translation method and system based on artificial intelligence to achieve adaptive translation in specific business scenarios.
[0005] To achieve the above purpose, the present invention provides the following solutions:
[0006] A real-time business text translation method based on artificial intelligence, comprising:
[0007] Collect business text data translated in different business scenarios, design a network model for each of the business scenarios, and train the network model with the business text data of the same business scenario as the network model based on a loss function that integrates overall loss, semantic loss, style loss, and cultural loss to obtain a number of agents;
[0008] Real-time collect communication text during the communication between the user and the customer or business documents uploaded by the user to obtain text data to be translated;
[0009] Use the TF-IDF algorithm and the BERT model to perform word frequency statistics, context information extraction, and style classification on the text data to be translated, obtain word frequency data, previous context keywords, and text format types, and perform application scenario matching on the text data to be translated according to the word frequency data, the previous context keywords, and the text format types to obtain scenario classification probability data;
[0010] Match according to the scenario classification probability data and the agent to determine the translation main model;
[0011] Use the translation main model to translate the text data to be translated to obtain a real-time translation result;
[0012] Collect the feedback information of the user. If the feedback information is a translation error, then replace the agent according to the scenario classification probability data, determine the replaced agent as the translation main model, and return to the step "Use the translation main model to translate the text data to be translated to obtain a real-time translation result".
[0013] Preferably, the construction process of the agent includes:
[0014] Construct an improved Transformer network; the improved Transformer network includes: an input processing layer, an encoder layer, a decoder layer, and a linear projection layer connected in sequence; the input processing layer includes: a word embedding module, a position encoding module, and a dynamic term library interface module connected in sequence; the encoder layer includes: a multi-head self-attention module that fuses a cultural perception attention head and a local attention window, an intertextuality enhancement module, a feed-forward network module, an encoding residual connection and layer normalization module; the decoder layer includes: an input processing module that fuses a template buffer, a masked multi-head self-attention module, a standard cross-sequence attention module, a sequence annotation head module, and a decoding residual connection and layer normalization module;
[0015] Respectively use the word embedding module, the position encoding module, and the dynamic term library interface module to perform word vector mapping, hybrid position encoding, and professional term retrieval and word embedding on the pre-collected standard translation text data to obtain an input vector;
[0016] Use the cultural perception attention head to load the pre-injected external knowledge graph, use the local attention window to limit the attention range, and perform weighted vector splicing and linear transformation fusion on the input vector by using the multi-head self-attention module according to the external knowledge graph and the attention range to obtain a fusion result;
[0017] Use the intertextuality enhancement module to perform cross-sentence keyword extraction on the fusion result to obtain a context-enhanced vector;
[0018] The context enhancement vector is non-linearly activated using the feed-forward network module, and the output of the feed-forward network module is integrated using the encoding residual connection and layer normalization module to obtain an encoded output;
[0019] The fusion template buffer in the input processing module is used to load the high-frequency term template;
[0020] The masked multi-head self-attention module is used to mask subsequent sequence positions of the encoded output to obtain a context representation;
[0021] Based on the high-frequency term template, the standard cross-sequence attention module is used to extract alignment relationships from the context representation to obtain a cross-modal representation;
[0022] The sequence annotation head module is used to predict fixed-format labels for the cross-modal representation, and the output of the sequence annotation head module is integrated using the decoding residual connection and layer normalization module to obtain a decoded output;
[0023] The linear projection layer is used to map the decoded output to the target language space to obtain the model translation output.
[0024] Preferably, communication texts during the communication between the user and the customer or business documents uploaded by the user are collected in real time to obtain text data to be translated, including:
[0025] Guide the user to extend the information acquisition plugin to the communication platform;
[0026] When the user selects single-sentence translation, the target sentence selected by the user on the communication platform is extracted to obtain the text data to be translated;
[0027] When the user selects real-time translation, all the speech texts in the user's chat window on the communication platform are extracted based on the user name label classification to obtain the text data to be translated;
[0028] When the user communicates through voice or video on the communication platform, the voice information in the communication window created by the user is monitored using cloud speech recognition algorithms, speaker separation techniques, and noise suppression algorithms, and the voice information is converted into text form to obtain the text data to be translated;
[0029] When the business document is a non-text document, OCR technology is used to extract the text information of the business document, and the BERT-Corrector model is used to correct spelling and fill in missing parts of the extracted text information to obtain the text data to be translated.
[0030] Preferably, the calculation formula of the scene classification probability data is as follows:
[0031] S q = p1·sim TF (F0, F 0q ) + p2·sim KW (K0, K 0q ) + p3·sim FMT (T0, T 0q );
[0032] Among them, S q is the matching degree of the text data to be translated to the q-th type of business scenario; p1, p2, and p3 are the first weight, the second weight, and the third weight respectively; sim TF (·), sim KW (·), sim FMT (·) represent cosine similarity calculation, Jaccard similarity calculation, and Boolean matching respectively; F0, K0, and T0 represent the word frequency data, the above-mentioned keyword, and the text format type respectively; F 0q , K 0q , T 0q represent the benchmark word frequency, the keyword set, and the format type set of the q-th type of business scenario respectively.
[0033] Preferably, the business scenarios include: advertising slogans, business letters, and business contracts.
[0034] Preferably, the expression of the dynamic term library interface module is:
[0035] D = LayerNorm(E + σ(W d ·[E;t])⊙t);
[0036] Among them, D is the output of the dynamic term library interface module; LayerNorm(·) represents layer normalization calculation; E is the word embedding vector; σ(·) is the Sigmoid function; W d is the gating weight matrix; t is the term embedding retrieved from the dynamic term library; ⊙ represents element-wise multiplication;
[0037] The expression of the cultural perception attention head is:
[0038]
[0039] Among them, Attention c (Q, K, V) represents the output of the cultural perception attention head; Q, K, and V are the query matrix, the key matrix, and the value matrix respectively; softmax(·) is the normalization matrix; d k is the attention dimension; Kkq is the knowledge graph key matrix; T represents the transpose operation;
[0040] The expression of the input processing module is:
[0041] T = Concat(H, T embed ) · W t ;
[0042] where T is the output of the input processing module; Concat(·) represents the concatenation operation; H is the input vector; T embed is the high-frequency term template embedding; W t is the projection matrix.
[0043] Preferably, the text format types include: letters, contracts, reports, technical documents, manuscripts, proposals, statements, and unformatted text.
[0044] Preferably, it further includes:
[0045] Calculating the loss function according to the model translation output and the standard translation text data to obtain the total loss value; the expression of the loss function is:
[0046] L total = αL gobol + βL semanitc + γL style + δL cutre ;
[0047] where L gobol = CrossEntropy(Y pred , Y rtue );
[0048]
[0049] L culture = ∑ c∈C GAT(c, G) · Accuracy(c);
[0050] [α, β, γ, δ] = Softmax(f ω (s));
[0051] L total is the total loss value; α, β, γ, δ are the global translation weight, the technical term translation weight, the style weight, and the cultural adaptation weight respectively; L gobol , L semanitc , L style , L cutre , L pred , Yrtue They are the predicted probability distribution and the standard translation distribution respectively; CrossEntropy(·) represents the calculation of cross entropy; T represents the dynamic term base; N is the number of terms in the current batch; w i is the i-th target language word; I(·) is the indicator function, which outputs 1 when the internal condition is satisfied and 0 otherwise; D KL (·) represents the calculation of relative entropy; p θ (w i ) and p ref (w i ) represent the predicted probability distribution of w i and the standard distribution of w i in the dynamic term base respectively; φ s (·) represents the encoder layer; X src and Y pred represent the source language text and the target language predicted text respectively; C represents the set of culture-loaded words; c is a single element in the set of culture-loaded words; G represents the cultural knowledge graph; GAT(c, G) is the cultural sensitivity weight output by the attention network; Accuracy(c) is the translation accuracy of cultural words; f ω (s) represents the weight assignment network;
[0052] Based on the total loss value, the overall loss, the semantic loss, the stylistic loss, and the cultural loss, calculate the loss gradient and perform gradient clipping to obtain gradient data. According to the gradient data, use the gradient backpropagation strategy to iteratively optimize the model parameters of the improved Transformer network, and dynamically adjust the learning rate of the gradient backpropagation strategy using the Adam optimizer during the iteration process to obtain the trained agent.
[0053] Preferably, the non-text files include: pdf files, JPEG pictures, and PNG pictures; the cloud speech recognition algorithm is the ASR model; the speaker separation technology includes: gated convolutional network and deep clustering analysis algorithm; the noise suppression algorithm is the DCCRN model.
[0054] Preferably, a real-time business text translation system based on artificial intelligence includes:
[0055] A model construction subsystem, configured to collect the business text data translated under different business scenarios, design a network model for each business scenario, and train the network model based on the loss function using the business text data in the same business scenario as the network model to obtain a plurality of agents;
[0056] A data acquisition subsystem, configured to collect in real time the communication text during the communication between the user and the customer or the business document uploaded by the user, so as to obtain the text data to be translated;
[0057] A scenario classification subsystem, configured to use the TF-IDF algorithm and the BERT model to perform word frequency statistics, context information extraction and style classification on the text data to be translated, so as to obtain the word frequency data, the above-mentioned keyword and the text format type, and perform application scenario matching on the text data to be translated according to the word frequency data, the above-mentioned keyword and the text format type, so as to obtain the scenario classification probability data;
[0058] A model matching subsystem, configured to match according to the scenario classification probability data and the intelligent agent to determine the translation main model;
[0059] A text translation subsystem, configured to use the translation main model to translate the text data to be translated, so as to obtain a real-time translation result;
[0060] A user feedback subsystem, configured to collect the feedback information of the user. If the feedback information is a translation error, the intelligent agent is replaced according to the scenario classification probability data, and the replaced intelligent agent is determined as the translation main model, and the process returns to the step of "using the translation main model to translate the text data to be translated to obtain a real-time translation result".
[0061] The present invention discloses the following technical effects:
[0062] The present invention provides a real-time business text translation method and system based on artificial intelligence. Through a multi-intelligent agent collaborative architecture, the problem that most existing translation systems use a generalization model, resulting in poor business text translation effects, is solved, and the translation task of a specific scenario is realized; through the TF-IDF algorithm and the BERT model, the scheduling problem of different intelligent agents is solved, and the dynamic scheduling and accurate matching of intelligent agents are realized; by integrating the loss functions of overall loss, semantic loss, style loss and cultural loss, the problem that conventional translation models have poor effects in terms of semantics, style and culture is solved, and the improvement of the translation accuracy and equivalence of the model is realized; by collecting user feedback and replacing intelligent agents, the problem that users are not satisfied with the translation results of the matched intelligent agents is solved, and the improvement of the user experience and the switching of translation styles are realized. Description of the Drawings
[0063] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0064] Figure 1 Schematic diagram of the real-time translation process of business texts based on artificial intelligence provided by the embodiments of the present invention;
[0065] Figure 2 Schematic diagram of the construction process of the intelligent agent provided by the embodiments of the present invention;
[0066] Figure 3 Schematic diagram of the process for obtaining text data to be translated provided by the embodiments of the present invention;
[0067] Figure 4 Schematic diagram of the model training process provided by the embodiments of the present invention;
[0068] Figure 5 Schematic diagram of the real-time translation system of business texts based on artificial intelligence provided by the embodiments of the present invention. Detailed implementation manners
[0069] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.
[0070] The object of the present invention is to provide a method and system for real-time translation of business texts based on artificial intelligence to achieve adaptive translation in specific business scenarios.
[0071] To make the above objects, features, and advantages of the present invention more obvious and understandable, the following will further describe the present invention in detail with reference to the drawings and specific implementation manners.
[0072] Figure 1 Schematic diagram of the real-time translation process of business texts based on artificial intelligence provided by the embodiments of the present invention, as Figure 1 shown, the present invention provides a method for real-time translation of business texts based on artificial intelligence, including:
[0073] Step 100: Collect the translated business text data in different business scenarios, design a network model for each business scenario, and use the business text data that is the same as the business scenario of the network model to train the network model based on a loss function that integrates the overall loss, semantic loss, stylistic loss, and cultural loss, to obtain a number of intelligent agents;
[0074] Step 200: Real-time collect the communication text during the communication between the user and the customer or the business documents uploaded by the user to obtain the text data to be translated;
[0075] Step 300: Use the TF-IDF algorithm and the BERT model to perform word frequency statistics, context information extraction, and stylistic classification on the text data to be translated, obtain the word frequency data, previous context keywords, and text format types, and perform application scenario matching on the text data to be translated according to the word frequency data, previous context keywords, and text format types to obtain the scenario classification probability data;
[0076] Step 400: Match according to the scenario classification probability data and the intelligent agents to determine the translation main model;
[0077] Step 500: Use the translation main model to translate the text data to be translated to obtain the real-time translation result;
[0078] Step 600: Collect the feedback information of the user. If the feedback information is a translation error, then replace the intelligent agent according to the scenario classification probability data, determine the replaced intelligent agent as the translation main model, and return to the step "Use the translation main model to translate the text data to be translated to obtain the real-time translation result".
[0079] Reference Figure 2 , the construction process of the intelligent agent includes:
[0080] Step 501: Construct an improved Transformer network; the improved Transformer network includes: an input processing layer, an encoder layer, a decoder layer, and a linear projection layer connected in sequence; the input processing layer includes: a word embedding module, a position encoding module, and a dynamic term library interface module connected in sequence; the encoder layer includes: a multi-head self-attention module that integrates cultural perception attention heads and local attention windows, an intertextuality enhancement module, a feed-forward network module, an encoding residual connection and layer normalization module connected in sequence; the decoder layer includes: an input processing module that integrates a template buffer, a masked multi-head self-attention module, a standard cross-sequence attention module, a sequence annotation head module, and a decoding residual connection and layer normalization module connected in sequence;
[0081] Step 502: Respectively use the word embedding module, position encoding module, and dynamic term library interface module to perform word vector mapping, hybrid position encoding, and professional term retrieval and word embedding on the pre-collected standard translation text data to obtain an input vector;
[0082] Step 503: Use the culture-aware attention head to load the pre-injected external knowledge graph, use the local attention window to limit the attention range, and use the multi-head self-attention module to perform weighted vector splicing and linear transformation fusion on the input vector according to the external knowledge graph and the attention range to obtain a fusion result;
[0083] Step 504: Use the intertextuality enhancement module to extract cross-sentence keywords from the fusion result to obtain a context-enhanced vector;
[0084] Step 505: Use the feed-forward network module to perform non-linear activation processing on the context-enhanced vector, and use the encoding residual connection and layer normalization module to integrate the output of the feed-forward network module to obtain an encoding output;
[0085] Step 506: Use the fusion template buffer in the input processing module to load high-frequency term templates;
[0086] Step 507: Use the masked multi-head self-attention module to mask the subsequent sequence positions of the encoding output to obtain a context representation;
[0087] Step 508: Use the standard cross-sequence attention module to extract alignment relationships from the context representation according to the high-frequency term template to obtain a cross-modal representation;
[0088] Step 509: Use the sequence annotation head module to predict fixed-format tags for the cross-modal representation, and use the decoding residual connection and layer normalization module to integrate the output of the sequence annotation head module to obtain a decoding output;
[0089] Step 510: Use the linear projection layer to map the decoding output to the target language space to obtain the model translation output.
[0090] Reference Figure 3 , in real-time collect the communication text during the communication process between the user and the customer or the business documents uploaded by the user to obtain the text data to be translated, including:
[0091] Step 201: Guide the user to extend the information acquisition plugin to the communication platform;
[0092] Step 202: When the user selects single-sentence translation, extract the target sentence selected by the user on the communication platform to obtain the text data to be translated;
[0093] Step 203: When the user selects real-time translation, all the speech texts in the user's chat window on the communication platform are extracted based on the user name tag classification to obtain the text data to be translated;
[0094] Step 204: When the user uses the voice communication method or video communication method of the communication platform, the voice information in the communication window created by the user is monitored by using the cloud voice recognition algorithm, speaker separation technology and noise suppression algorithm, and the voice information is converted into text form to obtain the text data to be translated;
[0095] Step 205: When the business document is a non-text document, the text information of the business document is extracted by using OCR technology, and the BERT-Corrector model is used to correct the spelling and fill in the missing parts of the extracted text information to obtain the text data to be translated.
[0096] Specifically, the calculation formula of the scenario classification probability data is:
[0097] S q =p1·sim TF F0,F 0q +p2·sim KW K0,K 0q +p3·sim FMT T0,T 0q ;
[0098] Among them, S q is the matching degree of the text data to be translated for the qth type of business scenario; p1, p2, and p3 are the first weight, the second weight, and the third weight respectively; sim TF (·), sim KW (·), sim FMT (·) respectively represent cosine similarity calculation, Jaccard similarity calculation, and Boolean matching; F0, K0, and T0 respectively represent word frequency data, previous keywords, and text format types; F 0q , K 0q , T 0q respectively represent the benchmark word frequency, keyword set, and format type set of the qth type of business scenario.
[0099] Optionally, the business scenarios include: advertising slogans, business letters, and business contracts.
[0100] Specifically, the expression of the dynamic term library interface module is:
[0101] D=LayerNorm(E+σ(W d ·E;t)⊙t);
[0102] Among them, D is the output of the dynamic term library interface module; LayerNorm(·) represents layer normalization calculation; E is the word embedding vector; σ(·) is the Sigmoid function; W d is the gating weight matrix; t is the term embedding retrieved by the dynamic term library; ⊙ represents element-wise multiplication;
[0103] The expression of the culture-aware attention head is:
[0104]
[0105] Among them, Attention c (Q, K, V) represents the output of the culture-aware attention head; Q, K, and V are the query matrix, key matrix, and value matrix respectively; softmax(·) is the normalization matrix; d k is the attention dimension; K kq is the knowledge graph key matrix; T represents the transpose operation;
[0106] The expression of the input processing module is:
[0107] T = Concat(H, T embed ·W t ;
[0108] Among them, T is the output of the input processing module; Concat(·) represents the concatenation operation; H is the input vector; T embed is the high-frequency term template embedding; W t is the projection matrix.
[0109] Optionally, the text format types include: letters, contracts, reports, technical documents, manuscripts, proposals, statements, and unformatted text.
[0110] Reference Figure 4 , and also includes:
[0111] Step 511: Calculate the loss function based on the model translation output and the standard translation text data to obtain the total loss value; the expression of the loss function is:
[0112] L total = αL gobol + βL semanitc + γL style + δL cutre ;
[0113] Among them, L gobol = CrossEntropy(Y pred , Y rtue );
[0114]
[0115] L culture = ∑ c∈C GAT(c, G)·Accuracy(c);
[0116] [α, β, γ, δ] = Softmax(f ω (s));
[0117] L total is the total loss value; α, β, γ, δ are the global translation weight, the technical term translation weight, the style weight, and the cultural adaptation weight respectively; L gobol , L semanitc , L style , L cutre are the overall loss, the semantic loss, the style loss, and the cultural loss respectively; Y pred , Y rtue are the predicted probability distribution and the standard translation distribution respectively; CrossEntropy(·) represents the cross-entropy calculation; T represents the dynamic term base; N is the number of terms in the current batch; w i is the i-th target language word; I(·) is the indicator function that outputs 1 when the internal condition is met and 0 otherwise; D KL (·) represents the relative entropy calculation; p θ (w i ), p ref (w i ) represent the predicted probability distribution of w i and the standard distribution of w i in the dynamic term base respectively; φ s (·) represents the encoder layer; X src , Y pred represent the source language text and the target language predicted text respectively; C represents the set of culture-loaded words; c is a single element in the set of culture-loaded words; G represents the cultural knowledge graph; GAT(c, G) is the cultural sensitivity weight output by the attention network; Accuracy(c) is the translation accuracy of the culture word; f ω (s) represents the weight assignment network;
[0118] Step 512: Calculate the loss gradient and perform gradient clipping based on the total loss value, the overall loss, the semantic loss, the style loss, and the cultural loss to obtain gradient data. Then, use the gradient backpropagation strategy to iteratively optimize the model parameters of the Transformer improved network according to the gradient data, and dynamically adjust the learning rate of the gradient backpropagation strategy using the Adam optimizer during the iteration process to obtain the trained agent.
[0119] Preferably, the non-text files include: pdf files, JPEG pictures, and PNG pictures; the cloud voice recognition algorithm is the ASR model; the speaker separation technology includes: gated convolutional network and deep clustering analysis algorithm; the noise suppression algorithm is the DCCRN model.
[0120] Reference Figure 5 , an artificial intelligence-based real-time business text translation system, including:
[0121] A model construction subsystem, which is used to collect business text data completed in translation under different business scenarios, design a network model for each business scenario, and train the network model based on the loss function using the business text data with the same business scenario as the network model to obtain several agents;
[0122] A data acquisition subsystem, which is used to collect communication text during the communication process between the user and the customer in real time or business files uploaded by the user to obtain text data to be translated;
[0123] A scenario classification subsystem, which is used to use the TF-IDF algorithm and the BERT model to perform word frequency statistics, context information extraction, and style classification on the text data to be translated, obtain word frequency data, previous keywords, and text format types, and perform application scenario matching on the text data to be translated according to the word frequency data, previous keywords, and text format types to obtain scenario classification probability data;
[0124] A model matching subsystem, which is used to match according to the scenario classification probability data and the agents to determine the translation main model;
[0125] A text translation subsystem, which is used to translate the text data to be translated using the translation main model to obtain a real-time translation result;
[0126] A user feedback subsystem, which is used to collect the feedback information of the user. If the feedback information is a translation error, then replace the agent according to the scenario classification probability data, determine the replaced agent as the translation main model, and return to the step of "using the translation main model to translate the text data to be translated to obtain a real-time translation result".
[0127] Specifically, communication platform plugin extension and text listening:
[0128] 1) Plugin integration: Guide the user to install an information acquisition plugin and embed it into the interface layer of the target communication platform to listen to the text input event of the chat window.
[0129] 2) Single sentence translation: After the user selects the target sentence, the plugin extracts the sentence as the text to be translated.
[0130] 3) Real-time translation: Continuously capture all the speech texts within the specified chat window according to the user name tag to form a real-time text stream.
[0131] Furthermore, speech-to-text conversion for voice / video communication:
[0132] 1) Noise suppression: Denoise the original audio to separate the pure speech from the background noise.
[0133] 2) Speaker separation: Use a gated convolutional network to extract voiceprint features and combine with a deep clustering algorithm to distinguish the speech segments of different speakers.
[0134] 3) Speech recognition: Convert the denoised speech stream into text through a cloud ASR model (end-to-end speech recognition model), supporting multi-language recognition and real-time streaming processing.
[0135] 4) Text formatting: Segment the converted text by timestamp and speaker tag to generate structured text data to be translated.
[0136] Specifically, the processing of non-text business documents:
[0137] 1) OCR text extraction: Use optical character recognition (OCR) technology to parse the text content in the document; perform adaptive layout analysis for complex formats such as tables and handwritten texts, and retain the original layout logic.
[0138] 2) Text error correction and completion: Use the BERT-Corrector model to correct the spelling of the text extracted by OCR and repair recognition errors; predict missing characters based on context semantics, such as completing incomplete words caused by blurred images.
[0139] Preferably, the ASR model is based on an end-to-end deep learning architecture, combining an acoustic model and a language model. The acoustic model extracts the temporal features of speech signals through a convolutional neural network (CNN) or Transformer, capturing acoustic patterns at the phoneme level; the language model is based on a recurrent neural network or self-attention mechanism to perform semantic error correction and context completion on the converted text. The model realizes low-latency inference through cloud distributed computing resources, supports multi-language recognition, and can optimize the recognition accuracy of professional terms for business scenarios. The gated convolutional network dynamically adjusts the activation area of the convolutional kernel through a gating mechanism (such as a GLU unit) to extract local features in the speech signal and distinguish the timbre and intonation features of different speakers. Deep clustering analysis: Maps the spectral features of speech signals to a high-dimensional space, and separates the feature distributions of different speakers through a clustering algorithm (such as K-means) to generate independent speech segments. DCCRN (Deep Complex Convolutional Recurrent Network) is based on a complex-domain neural network, processing both the amplitude and phase information of speech signals at the same time. Extracts time-frequency domain features through a complex convolutional layer, combines a recurrent neural network (such as LSTM) to model temporal dependencies, and dynamically separates noise from clean speech. The model adopts a mask learning strategy to predict the noise mask and reconstruct the denoised speech signal, which is especially suitable for the conference noise or environmental noise commonly found in business scenarios.
[0140] Furthermore, the dynamic term library interface module dynamically fuses general word embeddings and professional term embeddings through a gating mechanism, and combines layer normalization to stabilize the training process. This module flexibly calls the professional term library according to the context to solve the problem of mistranslation of professional vocabulary in business texts and improve the accuracy and consistency of term translation.
[0141] Specifically, the cultural perception attention head introduces an external knowledge graph in the attention mechanism through double-key matrix weighting, enhancing the sensitivity to culture-loaded words (such as etiquette terms and regional expressions).
[0142] Optionally, optimize the accuracy and efficiency of scenario classification: The combined calculation cost of TF-IDF and BERT is high. To improve efficiency and reduce costs, the DistilBERT or ALBERT model can be used; in addition, a business term graph can be constructed with the help of a graph neural network to improve the classification accuracy through node relationships. The specific implementation idea is: Replace the BERT in the scenario classification module with DistilBERT and fuse GNN features; collect business corpora for domain adaptive pre-training, construct a term co-occurrence graph, and extract key node embeddings.
[0143] Preferably, set up a real-time update mechanism for the dynamic term library, extract new terms from user feedback, automatically update the term library or regularly crawl the industry term database, and retain historical versions to support error rollback.
[0144] Specifically, when dealing with mixed-scenario texts (business communications without fixed scenarios), the outputs of multiple agents are weighted and fused according to scenario probabilities, and the optimal result is selected using confidence through a voting mechanism.
[0145] Preferably, this embodiment improves on the traditional Transformer model, which may lead to a reduction in the response speed of the final model (the latency can be directly reduced by increasing computing power), which is not conducive to the real-time translation function. To improve the response speed of the model, knowledge distillation can be performed on the maturely trained model to train a lightweight student model to imitate the original agent; 8-bit quantization of the model parameters and pruning of redundant layers can be carried out to obtain a low-parameter model.
[0146] The beneficial effects of the present invention are as follows:
[0147] Through the multi-agent collaborative architecture, the present invention realizes the translation tasks in specific scenarios and improves the applicability of translation; through the TF-IDF algorithm and the BERT model, the dynamic scheduling and accurate matching of agents are realized, improving the reliability of the translation model; by integrating the loss functions of overall loss, semantic loss, stylistic loss, and cultural loss, the sensitivity of the translation model in terms of semantics, style, and culture is improved, enhancing the accuracy and equivalence of model translation; through collecting user feedback and agent replacement, the improvement of user experience and the switching of translation styles are realized.
[0148] Each embodiment in this specification is described in a progressive manner. The key point of each embodiment is to illustrate the differences from other embodiments. The same or similar parts among the embodiments can be referred to each other. For the system disclosed in the embodiments, since it corresponds to the method disclosed in the embodiments, the description is relatively simple, and the relevant parts can be referred to the description of the method part.
[0149] Specific examples are used in this article to elaborate on the principles and implementation manners of the present invention. The descriptions of the above embodiments are only used to help understand the method and its core idea of the present invention; at the same time, for those of ordinary skill in the art, based on the idea of the present invention, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation to the present invention.
Claims
1. A method for real-time translation of business text based on artificial intelligence, characterized in that: include: Collecting business text data translated in different business scenarios, designing a network model for each business scenario, and training the network model using the business text data that is the same as the business scenario of the network model based on a loss function that integrates overall loss, semantic loss, stylistic loss, and cultural loss to obtain a plurality of intelligent agents; Collect communication texts between users and customers or business documents uploaded by users in real time to obtain text data to be translated; Using the TF-IDF algorithm and the BERT model to perform word frequency statistics, context information extraction, and style classification on the text data to be translated, word frequency data, previous keywords, and text format types are obtained; and application scenario matching is performed on the text data to be translated according to the word frequency data, the previous keywords, and the text format types to obtain scenario classification probability data; According to the scene classification probability data and the intelligent agent, a translation subject model is determined; Translating the text data to be translated using the translation subject model to obtain a real-time translation result; Collect user feedback information. If the feedback information is a translation error, replace the agent according to the scene classification probability data, determine the replaced agent as the translation subject model, and return to the step of "using the translation subject model to translate the text data to be translated to obtain a real-time translation result".
2. The method for real-time translation of business text based on artificial intelligence according to claim 1, characterized in that: The process of constructing the intelligent agent includes: Constructing a Transformer improved network; the Transformer improved network includes: an input processing layer, an encoder layer, a decoder layer and a linear projection layer connected in sequence; the input processing layer includes: a word embedding module, a position encoding module, and a dynamic terminology library interface module connected in sequence; the encoder layer includes: a multi-head self-attention module that integrates cultural perception attention heads and local attention windows, an intertextuality enhancement module, a feedforward network module, an encoding residual connection and a layer normalization module that are connected in sequence; the decoder layer includes: an input processing module that integrates a template buffer, a masked multi-head self-attention module, a standard cross-sequence attention module, a sequence annotation head module, and a decoding residual connection and a layer normalization module that are connected in sequence; Using the word embedding module, the position encoding module, and the dynamic terminology library interface module to perform word vector mapping, hybrid position encoding, professional terminology retrieval, and word embedding on the pre-collected standard translation text data, respectively, to obtain an input vector; Using the culture-aware attention head to load the pre-injected external knowledge graph, using the local attention window to limit the attention range, and using a multi-head self-attention module to perform weighted vector concatenation and linear transformation fusion on the input vector according to the external knowledge graph and the attention range to obtain a fusion result; Using the intertextuality enhancement module to extract cross-sentence keywords from the fusion result to obtain a context enhancement vector; Using the feedforward network module to perform nonlinear activation processing on the context enhancement vector, and using the coding residual connection and layer normalization module to integrate the output of the feedforward network module to obtain a coded output; Using the fusion template buffer in the input processing module to load the high-frequency term template; The masked multi-head self-attention module is used to perform subsequent sequence position masking on the encoded output to obtain a context representation; Extracting alignment relationships of the context representations using the standard cross-sequence attention module according to the high-frequency term template to obtain a cross-modal representation; Using the sequence labeling header module to predict fixed format labels for the cross-modal representation, and using the decoding residual connection and layer normalization module to integrate the output of the sequence labeling header module to obtain a decoded output; The decoded output is mapped to the target language space using the linear projection layer to obtain a model translation output.
3. The method for real-time translation of business text based on artificial intelligence according to claim 1, characterized in that: Collect the communication texts between users and customers or the business documents uploaded by users in real time to obtain the text data to be translated, including: Guide users to extend information acquisition plug-ins to communication platforms; When the user selects single sentence translation, extracting the target sentence selected by the user on the communication platform to obtain the text data to be translated; When the user selects real-time translation, all speech texts in the user chat window on the communication platform are extracted based on the user name tag classification to obtain the text data to be translated; When the user communicates via voice or video on the communication platform, the cloud-based voice recognition algorithm, speaker separation technology, and noise suppression algorithm are used to monitor the voice information in the communication window created by the user, and the voice information is converted into text to obtain the text data to be translated; When the business document is a non-text document, the text information of the business document is extracted using OCR technology, and the BERT-Corrector model is used to perform spelling correction and missing filling on the extracted text information to obtain the text data to be translated.
4. The method for real-time translation of business text based on artificial intelligence according to claim 1, characterized in that: The calculation formula of the scene classification probability data is: S q =p1·sim TF (F0,F 0q )+p2·sim KW (K0,K 0q )+p3·sim FMT (T0,T 0q ); Among them, S q is the matching degree of the text data to be translated to the business scenario of the qth category; p1, p2, p3 are the first weight, the second weight, and the third weight respectively; sim TF (·),sim KW (·),sim FMT (·) respectively represent cosine similarity calculation, Jaccard similarity calculation and Boolean matching; F0, K0, T0 respectively represent the word frequency data, the above keyword and the text format type; F 0q , K 0q 、T 0q They respectively represent the benchmark word frequency, keyword set and format type set of the qth business scenario.
5. The method for real-time translation of business text based on artificial intelligence according to claim 1, characterized in that: The business scenarios include: advertising slogans, business letters and business contracts.
6. The method for real-time translation of business text based on artificial intelligence according to claim 1, characterized in that: The expression of the dynamic term base interface module is: D=LayerNorm(E+σ(W d ·[E;t)]⊙t); Where D is the output of the dynamic term library interface module; LayerNorm(·) represents layer normalization calculation; E is the word embedding vector; σ(·) is the Sigmoid function; W d is the gated weight matrix; t is the term embedding for dynamic term base retrieval; ⊙ represents element-by-element multiplication; The expression of the culture-aware attention head is: Among them, Attention c (Q, K, V) represents the output of the culture-aware attention head; Q, K, V are the query matrix, key matrix, and value matrix respectively; softmax(·) is the normalization matrix; d k is the attention dimension; K kq is the knowledge graph key matrix; T represents the transposition operation; The expression of the input processing module is: T=Concat(H,T embed )·W t ; Wherein, T is the output of the input processing module; Concat(·) represents the concatenation operation; H is the input vector; T embed is the high-frequency term template embedding; W t is the projection matrix.
7. The method for real-time translation of business text based on artificial intelligence according to claim 1, characterized in that: The text format types include: letters, contracts, reports, technical documents, manuscripts, proposals, statements and unformatted texts.
8. The method for real-time translation of business text based on artificial intelligence according to claim 2, characterized in that: Also includes: The loss function is calculated according to the model translation output and the standard translation text data to obtain a total loss value; the expression of the loss function is: L total =αL gobol +βL semanitc +γL style +δL cutre ; Among them, L gobol =CrossEntropy(Y pred ,Y rtue ); L culture =∑ c∈C GAT(c,G)·Accuracy(c); [a,b,c,d]=Softmax(f ω (s)); L total is the total loss value; α, β, γ, and δ are global translation weight, professional term translation weight, style weight, and cultural adaptation weight, respectively; L gobol , L semanitc , L style , L cutre are the overall loss, the semantic loss, the stylistic loss, and the cultural loss respectively; Y pred , Y rtue are the predicted probability distribution and the standard translation distribution respectively; CrossEntropy(·) represents the cross entropy calculation; T represents the dynamic term base; N represents the number of terms in the current batch; w i is the i-th target language word; I(·) is an indicative function, which outputs 1 when the internal condition is met, otherwise it outputs 0; D KL (·) indicates relative entropy calculation; p θ (w i ), p ref (w i ) represent w i The predicted probability distribution and w i The standard distribution in the dynamic term base; φ s (·) represents the encoder layer; X src , Y pred Represent the source language text and the target language predicted text respectively; C represents the set of culturally loaded words; c is a single element in the set of culturally loaded words; G represents the cultural knowledge graph; GAT(c,G) is the cultural sensitivity weight output by the attention network; Accuracy(c) is the accuracy of cultural word translation; f ω (s) represents the weight distribution network; Loss gradient calculation and gradient clipping are performed according to the total loss value, the overall loss, the semantic loss, the stylistic loss, and the cultural loss to obtain gradient data, and the model parameters of the Transformer improved network are iteratively optimized using a gradient back propagation strategy according to the gradient data. During the iterative process, the learning rate of the gradient back propagation strategy is dynamically adjusted using an Adam optimizer to obtain the trained intelligent agent.
9. The method for real-time translation of business text based on artificial intelligence according to claim 3, characterized in that: The non-text files include: PDF files, JPEG images and PNG images; the cloud-based speech recognition algorithm is an ASR model; the speaker separation technology includes: a gated convolutional network and a deep clustering analysis algorithm; the noise suppression algorithm is a DCCRN model.
10. A business text real-time translation system based on artificial intelligence, applied to the business text real-time translation method based on artificial intelligence according to claim 1, the system comprising: A model building subsystem, used for collecting the business text data translated in different business scenarios, designing a network model for each business scenario, and training the network model based on the loss function using the business text data that is the same as the business scenario of the network model to obtain a plurality of intelligent agents; A data acquisition subsystem is used to collect the communication texts between users and customers or the business documents uploaded by users in real time to obtain the text data to be translated; A scene classification subsystem, for performing word frequency statistics, context information extraction, and genre classification on the text data to be translated using the TF-IDF algorithm and the BERT model to obtain the word frequency data, the preceding keywords, and the text format type, and performing application scene matching on the text data to be translated according to the word frequency data, the preceding keywords, and the text format type to obtain the scene classification probability data; A model matching subsystem, used for matching the scene classification probability data with the intelligent agent to determine the translation subject model; A text translation subsystem, used to translate the text data to be translated using the translation subject model to obtain a real-time translation result; The user feedback subsystem is used to collect the feedback information of the user. If the feedback information is a translation error, the intelligent agent is replaced according to the scene classification probability data, the replaced intelligent agent is determined as the translation subject model, and the step of "using the translation subject model to translate the text data to be translated to obtain a real-time translation result" is returned.
Citation Information
Patent Citations
Speech translation method, system and equipment fusing text semantic features
CN112800782A
Concept verification method and device based on pre-training language model, equipment and medium
CN118133797A
Language model and agent combined translation method
CN118643845A
Instruction-driven end-side agent control method and system
CN119272801A
Multi-agent question answering system based on large language model
CN119294520A
Cited By
Cross-end translation method and device and electronic equipment
CN121173792A
Multi-agent collaborative cross-language system translation method and system based on loAs
CN121189343A
A multi-agent collaborative multi-modal document translation method and system
CN122528921A
A multi-agent collaborative multi-modal document translation method and system
CN122528921B