Translation method and system based on semantic energy conservation network, medium, equipment and program product
By introducing semantic energy conservation networks into machine translation, using Hamiltonian deep neural network and VI Verlet integrators for lexical information evolution, the problem of insufficient semantic transmission in special contexts is solved, and more accurate translation results are achieved.
Patent Information
- Application Number
- CN202510149161.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-11
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2045-02-11
AI Technical Summary
Existing machine translation technology based on Transformer model is difficult to effectively transmit semantic information in source language text in certain special contexts, resulting in semantic loss.
The translation method based on semantic energy conservation network is adopted to evolve vocabulary information through Hamiltonian deep neural network (HDNN) and VI Verlet integrators, and the evolved relative position and activity amount are input into the Transformer-based translation network to achieve the final generation of translation results.
It effectively avoids semantic deficiencies in the translation process and improves the translation accuracy of the source language to the target language, especially in long sentences or complex grammatical structures.
Smart Images

Figure CN120181099A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of machine translation, and particularly relates to a translation method, system, medium, device and program product based on a semantic energy conservation network. Background Art
[0002] In modern life, cross - language communication is becoming increasingly frequent. For this reason, machine translation technology is also gradually popularized and developed.
[0003] Although the Transformer model has achieved remarkable results in machine translation, in some special contexts, the Transformer translation network can only achieve correct literal translation from the source language to the target language, but it is difficult to effectively convey some semantic information in the source - language text to the target language, resulting in semantic loss.
[0004] Therefore, it is necessary to provide an improved technical solution for the above - mentioned deficiencies of the existing technology. Summary of the Invention
[0005] The purpose of the present application is to provide a translation method, system, medium, device and program product based on a semantic energy conservation network to solve or alleviate the problems existing in the above - mentioned existing technology.
[0006] To achieve the above purpose, the present application provides the following technical solutions:
[0007] In a first aspect, the present application provides a translation method based on a semantic energy conservation network, including:
[0008] Obtain the text to be translated, and determine the relative position of each word in the sentence to be translated and its corresponding feature vector;
[0009] Map the relative position of each word in the sentence to be translated to the generalized coordinates in Hamiltonian dynamics, map the feature vector of each word in the sentence to be translated to the active quantity corresponding to the relative position, and use an energy conversion network to convert the sentence to be translated into an energy - form representation;
[0010] According to this energy - form representation, calculate the Hamiltonian of each word in the initial state and the total Hamiltonian of the sentence to be translated as the initial energy of each word and the initial total energy of the sentence to be translated;
[0011] Use the Xavier method for weight initialization, and input the initialized result into the Hamiltonian deep neural network HDNN to extract the feature of word information with the goal of semantic energy conservation, and obtain the relative position and active quantity evolving at each time point;
[0012] Input the processing result of HDNN into the VI Verlet integrator, and perform further evolution according to the time step to obtain the relative position and activity after further evolution;
[0013] Input the relative position and activity after further evolution into the Transformer-based translation network to obtain the translation result;
[0014] Use the energy evaluation network to recalculate the Hamiltonian of each word and the total Hamiltonian of the sentence to be translated, and calculate the energy difference to obtain the final result.
[0015] Combined with the first aspect, in some possible implementation manners, the specific form of the energy representation is:
[0016] E 0i = H(p i , q i ),
[0017]
[0018] In the formula, E 0i is the initial energy of the i-th word, obtained by H(p i , q i ), H is the Hamiltonian, p i is the activity of the i-th word, q i is the relative position of the i-th word, i = 1, 2,..., n, and E tot0 is the initial total energy of the sentence to be translated.
[0019] Combined with the first aspect, in some possible implementation manners, the expression form of the Hamiltonian H is:
[0020]
[0021] In the formula, U and V are parameter matrices, q is the vector composed of the relative positions of all words, and p is the vector composed of the activities of all words.
[0022] Combined with the first aspect, in some possible implementation manners, HDNN is a time-varying Hamiltonian system; the expression of the dynamic equation of the time-varying Hamiltonian system is as follows:
[0023]
[0024] In the formula, is the state vector, t is the time, J(t) is a time-varying skew-symmetric matrix, and y (0) is the initial state vector, including the initial energy of each word and the initial total energy of the sentence to be translated.
[0025] In combination with the first aspect, in some possible embodiments, in the Hamiltonian deep neural network (HDNN), the extended Hamiltonian is discretized using the S-IE discretization method to obtain the relative positions and activity levels of each word at each time point. The expression is as follows:
[0026]
[0027] q i+1 = q i - hLM q,i σ(M p,i q i + b p,i ),
[0028] In the formula, i represents the i-th time step, p i , p i+1 are the relative positions of adjacent time steps, q i , q i+1 are the activity levels of adjacent time steps, h is the time step size, L and M are parameter matrices, L T , M T are the transpose matrices of L and M respectively, M p,i is the parameter matrix specifically represented at the i-th time step related to the variable p i , and M q,i is the parameter matrix specifically represented at the i-th time step related to the variable q i , b q,i represents the bias term related to the variable q i and having a specific value at the i-th time step, b p,i , represents the bias term related to the variable p i and having a specific value at the i-th time step.
[0029] In the second aspect, the present embodiment provides a translation system based on a semantic energy conservation network, including:
[0030] An acquisition module, configured to acquire the text to be translated and determine the relative position of each word in the sentence to be translated and its corresponding feature vector;
[0031] A mapping module, configured to map the relative position of each word in the sentence to be translated to the generalized coordinates in Hamiltonian dynamics, map the feature vector of each word in the sentence to be translated to the activity level corresponding to the relative position, and use the energy conversion network to convert the sentence to be translated into an energy form representation;
[0032] A calculation module, configured to calculate the Hamiltonian of each word in the initial state and the total Hamiltonian of the sentence to be translated according to the energy form representation, as the initial energy of each word and the initial total energy of the sentence to be translated;
[0033] The first evolution module is used to perform weight initialization using the Xavier method and input the initialized result into the Hamiltonian deep neural network (HDNN) to extract the features of lexical information with the goal of semantic energy conservation, obtaining the relative position and activity amount evolved at each time point.
[0034] The second evolution module is used to input the processing result of the HDNN into the VI Verlet integrator for further evolution according to the time step, obtaining the relative position and activity amount after further evolution.
[0035] The translation module is used to input the relative position and activity amount after further evolution into the Transformer-based translation network to obtain the translation result.
[0036] The evaluation module is used to recalculate the Hamiltonian of each lexical item and the total Hamiltonian of the sentence to be translated using the energy evaluation network, and calculate the energy difference to obtain the final result.
[0037] In a third aspect, this embodiment provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the translation method based on the semantic energy conservation network provided in any of the above embodiments are implemented.
[0038] In a fourth aspect, this embodiment provides an electronic device, including: a memory, a processor, and a program stored in the memory and executable on the processor. When the processor executes the program, the steps of the translation method based on the semantic energy conservation network provided in any of the above embodiments are implemented.
[0039] In a fifth aspect, this embodiment provides a computer program product, including computer-executable instructions for causing a computer to execute the steps of the translation method based on the semantic energy conservation network provided in any of the above embodiments.
[0040] The technical solution of the embodiments of this application has the following beneficial effects:
[0041] In an embodiment of the first aspect, the vocabulary of the input source language sentence is represented in the form of energy. From the perspective of energy conservation, the Hamiltonian deep neural network (HDNN) and the VI Verlet integrator are used as energy evolution networks to process the evolution of vocabulary positions and activity amounts, and the evolved relative positions and activity amounts are input into the Transformer-based translation network to obtain the translation result. This method combines the idea of Hamiltonian dynamics and the advantages of deep learning, evolves through semantic energy conservation, ensures the stability and consistency of information, avoids semantic loss or information distortion during translation, and uses HDNN for vocabulary information evolution and feature extraction, enabling the model to better understand the context relationship of the input text, especially in long sentences or complex grammatical structures, which helps to improve the accuracy of Transformer during translation.
[0042] By using the Xavier initialization method in HDNN, it can ensure that the model has a reasonable parameter distribution at the beginning of training, which helps to accelerate the training process and improve convergence.
[0043] Based on the evolution of HDNN, the VI Verlet integrator is used for further evolution. Utilizing the reversibility characteristic of the VI Verlet integrator, it is convenient to trace the changes in the vocabulary representation amounts (relative position q and activity amount p) according to the time step, thereby further improving the translation accuracy.
[0044] The technical effects that can be achieved by any one of the second to fifth aspects described above can refer to the description of the beneficial effects in the first aspect above, and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS
[0045] Figure 1 It is a schematic flowchart of a translation method based on a semantic energy conservation network according to some embodiments of the present application.
[0046] Figure 2 It is a schematic structural diagram of a semantic energy conservation network according to some embodiments of the present application.
[0047] Figure 3 It is a schematic flowchart of a translation method based on a semantic energy conservation network according to another embodiment of the present application.
[0048] Figure 4 It is a schematic block diagram of a translation system based on a semantic energy conservation network according to another embodiment of the present application.
[0049] Figure 5 It is a schematic diagram of an electronic device according to another embodiment of the present application.
[0050] Description of the reference numerals:
[0051] 1 - Energy evolution module, 11 - Energy conversion network, 12 - Hamiltonian deep neural network HDNN, 13 - VI Verlet integrator, 2 - Transformer translation network, 21 - Multi - head attention mechanism, 22 - Residual connection and normalization processing, 23 - Feed - forward neural network, 26 - Empirical data augmentation, 3 - Energy evaluation network. Detailed implementation manners
[0052] As described in the background art, in some special contexts, the existing machine - learning translation technology based on Transformer technology is prone to semantic loss. For example, for the Chinese text: "Three common cobblers can equal Zhuge Liang.", the translated English text is: "Three common cobblers can equal Zhuge Liang.", which is literally correct, but it is difficult to reflect the relationship between the fool and the wise, resulting in semantic loss.
[0053] In view of this, the present disclosure provides a translation method, system, medium, device and program product based on a semantic energy conservation network, introducing the Hamiltonian mechanics system, using the Hamiltonian deep neural network HDNN to evolve lexical information with the goal of semantic energy conservation, then inputting the result evolved by the HDNN into the VI Verlet integrator for further evolution, and using the Transformer translation network for translation to alleviate the semantic loss problem of the Transformer translation network and improve the translation accuracy from the source language to the target language.
[0054] The terms "first", "second", "third" and "fourth" etc. in the description and claims of this application and the accompanying drawings are used to distinguish different objects, rather than to describe a specific order. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non - exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units is not limited to the listed steps or units, but optionally further includes steps or units not listed, or optionally further includes other steps or units inherent to these processes, methods, products or devices.
[0055] Referring to "embodiments" herein means that the specific features, structures or characteristics described in connection with the embodiments can be included in at least one embodiment of this application. The phrase appears in various places in the specification does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art explicitly and implicitly understand that the embodiments described herein can be combined with other embodiments.
[0056] The embodiments of the present application will be described below with reference to the accompanying drawings.
[0057] This embodiment provides a translation method based on a semantic energy conservation network. As Figure 1 shown, the method includes the following steps:
[0058] Step S101: Obtain the text to be translated, and determine the relative position of each word in the sentence to be translated and its corresponding feature vector.
[0059] In this embodiment, obtaining the text to be translated means obtaining the source language text to be translated. The text to be translated can be obtained through the input box provided by the user interface or other user interfaces provided by the system, or by uploading the document to be translated (such as a PDF or Word document) or by invoking the document through an API. The specific obtaining method is not limited in this embodiment.
[0060] In this embodiment, word embedding technology and position encoding technology can be used to determine the relative position of each word in the sentence to be translated and its corresponding feature vector. For example, Word2Vec, GloVe, or context-aware BERT, GPT can be used to map each word in the sentence to be translated to a high-dimensional feature vector to capture the semantic information of the word, and at the same time, position encoding is used to identify the relative position of each word in the sentence.
[0061] For ease of description, let q=(q1…q n ) represent the relative positions of each word, and let p=(p1…p n ) represent the feature vectors of each word.
[0062] Step S102: Map the relative position of each word in the sentence to be translated to the generalized coordinates in Hamiltonian dynamics, map the feature vector of each word in the sentence to be translated to the active quantity corresponding to the relative position, and use the energy conversion network to convert the sentence to be translated into an energy form representation.
[0063] In this embodiment, the role of the energy conversion network is to map the word positions and feature vectors of the source language to the concepts in Hamiltonian mechanics and represent them in the form of energy.
[0064] In Hamiltonian dynamics, p represents the generalized momentum and q represents the generalized coordinates.
[0065] For example, assume that sentence A to be translated contains N words, and the relative positions of each word are q1…q n , and the corresponding active quantities (i.e., feature vectors in NLP) are p1…p n, the energy conversion network takes the relative positions and activity levels of each word as inputs, calculates the Hamiltonian H of each word to obtain the energy of the word, and sums up the energies of all words to get the total energy of the sentence.
[0066] Step S103: According to this energy form representation, calculate the Hamiltonian of each word in the initial state and the total Hamiltonian of the sentence to be translated, as the initial energy of each word and the initial total energy of the sentence to be translated.
[0067] Exemplarily, the initial energy of each word is represented by E 0i and the total energy of the sentence is represented by E tot0 . The calculation formulas are as follows:
[0068] E 0i = H(p i , q i ),
[0069]
[0070] wherein, E 0i is the initial energy of the i-th word, obtained through the Hamiltonian H of the system, p i is the activity level of the i-th word, q i is the relative position of the i-th word, i = 1, 2,..., n, and R tot0 is the initial total energy of the sentence to be translated.
[0071] The Hamiltonian H of the system, also known as the Hamiltonian function, here, a specific form of H is given, corresponding to the physical model of a mass ball with energy conservation. The expression of H is:
[0072]
[0073] wherein, U and V are parameter matrices, q is a vector composed of the relative positions of all words, and p is a vector composed of the activity levels of all words.
[0074] Step S104: Use the Xavier method for weight initialization and input the initialized result into the Hamiltonian deep neural network HDNN to extract the feature of word information with the goal of semantic energy conservation, and obtain the relative position and activity level evolving at each time point.
[0075] The Xavier method accelerates the training process and improves the performance of the model by reasonably setting the initial values of the weights of the neural network layers so that the variances of the outputs and gradients of each layer are kept as stable as possible. In this embodiment, by introducing the parameter matrices U and V into the Hamiltonian function and adopting the Xavier method to initialize the parameter matrices, it is convenient for feature optimization and semantic understanding.
[0076] Furthermore, when initializing the parameter matrix using the Xavier method, the parameter matrix can be initialized as a weight matrix following a uniform distribution, or it can be initialized as a weight matrix following a normal distribution. This embodiment does not limit this.
[0077] Input the result of Xavier initialization into the Hamiltonian deep neural network HDNN to further process the lexical information, making the energy close to being conserved.
[0078] It should be noted that the Hamiltonian deep network HDNN is a network built based on Hamiltonian dynamics. Its core idea is to apply Hamiltonian mechanics to the training of neural networks and describe the evolution of the system by introducing the Hamiltonian H.
[0079] Among them, HDNN is a time-varying Hamiltonian system, and the expression of the dynamic equation of the time-varying Hamiltonian system is as follows:
[0080]
[0081] In the formula, is the state vector, t is the time, J(t) is the time-varying skew-symmetric matrix, and y (0) is the initial state vector, including the initial energy of each word and the initial total energy of the sentence to be translated.
[0082] To ensure the marginal stability of the system, the S-IE discretization method is used for discretization processing: T is the end point of the motion time interval corresponding to the dynamic model. The time sampling period h = T / n, where n represents the number of layers of the HDNN network. The sampling period h is also called the step size of the time step. Then the total duration T is divided into n discrete time steps. Each time step corresponds to a layer of the neural network and also corresponds to an update of the system state. By introducing the parameter matrices L, M, and the bias b, the discretization processing gives:
[0083]
[0084] In the formula: σ represents a common activation function, such as tanh(·), ReLU(·), etc.
[0085] For the convenience of solution, assume:
[0086]
[0087] The relative position and activity level of each word at each time point are obtained, and the expression is as follows:
[0088]
[0089] q i+1 = q i - hLMq,i σ(M p,i q i +b p,i ),
[0090] In the formula, i represents the i-th time step, p i 、p i+1 are the relative positions of adjacent time steps, q i 、q i+1 are the activity amounts of adjacent time steps, h is the time step size, L and M are parameter matrices, L T 、M T are the transposed matrices of L and M respectively, M p,i is the parameter matrix specifically represented at the i-th time step related to the variable p i 、M q,i is the parameter matrix specifically represented at the i-th time step related to the variable q i ,b q,i represents the bias term related to the variable q i and having a specific value at the i-th time step, b p,i ,represents the bias term related to the variable p i and having a specific value at the i-th time step.
[0091] In this embodiment, through the application of the Hamiltonian deep neural network HDNN, the enhancement and extraction of lexical semantic information are strengthened.
[0092] Step S105: Input the processing result of the HDNN into the VI Verlet integrator, and perform further evolution according to the time step size to obtain the further evolved relative position and activity amount.
[0093] In this embodiment, the VI Verlet integrator realizes further evolution according to the time step size by combining the Hamiltonian equation and the symplectic property based on the input lexical position information q and activity amount p:
[0094]
[0095] q i+1 =2q i -q i-1 +p i ′t 2 ,
[0096] In the formula, p i ′ is the first-order derivative of p i with respect to time t.
[0097] Finally, the further evolved relative positions and activity amounts are obtained. At this point, the energy evolution ends. Therefore, the relative positions and activity amounts output by the VI Verlet integrator are also referred to as the final results of the energy evolution.
[0098] Step S106: Input the further evolved relative positions and activity amounts into the Transformer-based translation network to obtain the translation result.
[0099] In this embodiment, the further evolved relative positions and activity amounts are the relative positions and activity amounts after being processed by the VI Verlet integrator, that is, the final results of the energy evolution.
[0100] The Transformer-based translation network, also known as the Transformer translation network, is a deep learning model based on the self-attention mechanism. It processes the input and output sequences through the self-attention mechanism and parallel design, and is commonly used in natural language processing (NLP) tasks, especially in machine translation, so as to capture long-distance dependencies more efficiently during the training process.
[0101] Among them, the Transformer translation network includes a multi-head attention mechanism (MHA), which is used to receive the evolved relative position q and activity amount p, and capture the long-distance dependencies in the input sentence through the attention mechanism. The calculation formula of the multi-head attention mechanism is as follows:
[0102]
[0103] In the formula, Q1, K1, and V1 are the query, key, and value matrices respectively, and d k is the dimension of the key and query.
[0104] Among them, Q1 corresponds to the relative position q in the final result of the energy evolution, K1 corresponds to the activity amount p in the final result of the energy evolution, and V1 = (v1,..., v i ) is the output result of the attention mechanism.
[0105] Step S107: Use the energy evaluation network (NN) to recalculate the Hamiltonian of each vocabulary and the total Hamiltonian of the sentence to be translated, and calculate the energy difference to obtain the final result.
[0106] In this embodiment, the energy evaluation network recalculates the Hamiltonian of each vocabulary and the total Hamiltonian of the sentence to be translated according to V1 output by the multi-head attention mechanism and the relative position q output by the VI Verlet integrator according to the following formula:
[0107]
[0108] wherein, v i and q i are respectively the output of the attention mechanism and the relative position of the i-th word in V1 and q, and E 1i is the Hamiltonian of the i-th word after recalculation, and E tot1 is the total Hamiltonian of the sentence to be translated after recalculation.
[0109] That is to say, the ability evaluation network re-evaluates the energy of each word and the total energy of the sentence by comparing the final result of the energy evolution with the output result of the Transformer translation network.
[0110] Calculate the energy difference between E tot1 and E tot0 , that is, evaluate the energy loss, and obtain the final result, that is, evaluate the translation effect from the energy perspective.
[0111] The translation method based on the semantic energy conservation network provided in this embodiment specifically implements its steps through the semantic energy conservation network. The structure of the semantic energy conservation network is as Figure 2 shown, including the following three parts: an energy evolution module 1, a Transformer translation network 2, and an energy evaluation network 3 (NN). Among them, the energy evolution module 1 is used to strengthen the expression and extraction of word information; the Transformer translation network 2 is used to receive the output of the energy evolution module and realize the translation from the source language to the target language; the energy evaluation network 3 (NN) is used to evaluate the translation loss.
[0112] Specifically, the energy evolution module 1 includes: an energy conversion network (ETN) 11, a Hamiltonian deep neural network 12 (HDNN), and a VI Verlet integrator 13 (VI). The energy conversion network 11 is used to convert the text input in the source language into an energy form representation; the Hamiltonian deep neural network 12 (HDNN) and the VI Verlet integrator 13 are used to evolve the word information in the energy representation form to strengthen the feature expression ability and feature extraction ability of the word information.
[0113] The Transformer translation network 2 includes modules such as the multi-head attention mechanism 21 (MHA), residual connection and normalization processing 22 (Add&Norm), feed-forward neural network 23 (FFN), and empirical data augmentation 26 (EDA). The multi-head attention mechanism 21 receives the lexical information (including relative position and activity amount) output by the energy evolution module, processes it using the multi-head attention mechanism, and sequentially passes through the residual connection and normalization processing 22 (Add&Norm) and the feed-forward neural network 23 (FFN), and then uses the residual connection and normalization processing 22 (Add&Norm) again to obtain the intermediate result 1. At the same time, the empirical data of the target language is input into the Transformer translation network 2, processed by the multi-head attention mechanism 21 (MHA) to obtain the intermediate result 2. The intermediate result 1 and the intermediate result 2 are combined and input into the empirical data augmentation 26 (EDA) module. The output result of the EDA module is processed by the feed-forward neural network 23 (FFN) to obtain the output result of the Transformer translation network 2.
[0114] Finally, the energy evaluation network 3 (NN) calculates the loss between the output result of the Transformer translation network 2 and the output result of the energy evolution module 1 to obtain the final output for evaluating the translation effect.
[0115] As an example, as Figure 3 shown, the method provided in this embodiment can be executed according to the following steps:
[0116] Step 1: Input the source language and obtain the energy of each sentence;
[0117] Among them, the energy of each sentence is realized by mapping the relative position of the vocabulary and the feature vector to the energy form through the energy conversion network.
[0118] Step 2: Train in the model to minimize the energy loss;
[0119] Here, the model refers to the semantic energy conservation network model, including the energy evolution module 1, the Transformer translation network 2, and the energy evaluation network 3. The detailed results of each part are referred to the foregoing and will not be elaborated here.
[0120] Step 3: Re-evaluate the energy in the neural network;
[0121] Here, the neural network refers to the semantic energy conservation network. Re-evaluating the energy means re-calculating the energy of each vocabulary based on the output result of the energy evolution module 1 and the translation result of the Transformer translation network 2, and then calculating the total energy of the sentence.
[0122] Step 4: Output the translation result, and the semantic deviation is manifested as energy loss.
[0123] In summary, the translation method based on the semantic energy conservation network provided in this embodiment combines the physics theory with artificial intelligence language translation by applying the Hamiltonian mechanics system to the field of language translation. Specifically, the semantic energy conservation network is used to perform the translation task, including the energy evolution module 1, the Transformer translation network 2, and the energy evaluation network 3, to achieve the task of translating the source language. The Hamiltonian mechanics system is introduced into the energy evolution module 1, and the Hamiltonian function is usually used to represent the total energy of the system. For a dynamic system, if no external force does work, the energy of the system is constant. According to this theory, in this embodiment, it is proposed that in the process of neural network translation, ideally, the energy is conserved in the semantic space between the source language and the target language. Therefore, the semantic energy conservation network is trained with the goal of minimizing energy loss, so that the translation from the source language to the target language is literally correct, and at the same time, the semantics of the source language can be better transmitted to the target language, alleviating the problem of semantic loss in the translation process as much as possible.
[0124] It should be noted that the "Hamiltonian function" involved in this application is a different concept from the "Hamiltonian operator" and "Hamiltonian matrix element" which combine traditional Hamiltonian with artificial intelligence, borrowing different physics ideas. The "Hamiltonian operator" and "Hamiltonian matrix element" are concepts in quantum mechanics and do not belong to the category of Hamiltonian dynamics. In quantum mechanics, the Hamiltonian operator is used to represent the total energy of the system and describes the time evolution of the quantum state through the Schrödinger equation. The Hamiltonian matrix element is the matrix element that describes the transition probability or amplitude between system states in quantum mechanics, representing the transition between the i-th and j-th quantum states in quantum mechanics, and they are the matrix representation of the Hamiltonian operator in the selected basis. While the Hamiltonian function in this application is a concept in classical mechanics and belongs to a part of Hamiltonian dynamics, and Hamiltonian dynamics uses the Hamiltonian function to describe the state and evolution of the system. Although the names of all three contain the word "Hamiltonian" and are all related to the energy of the system, they belong to two different theoretical systems of classical mechanics and quantum mechanics respectively.
[0125] It also needs to be noted that in some existing technologies, the Hamiltonian evolution layer is mainly used to process the generalized potential and generalized momentum, mainly relying on the Hamiltonian neural network and the Euler integrator. In this application, the vocabulary of the input source language sentence is represented in the form of energy. From the perspective of energy conservation, the energy evolution network is used to process the vocabulary position and activity information. By using the Hamiltonian deep neural network combined with the VI Verlet integrator, it can better utilize the properties of the symplectic structure to ensure the accuracy and stability of numerical simulation, and better ensure the constancy of energy in the semantic space, which is different from the understanding and combination angle of the idea of Hamiltonian classical dynamics in the existing technology.
[0126] Based on the same inventive concept, this embodiment also provides a translation system based on a semantic energy conservation network, as Figure 4 shown. The system includes: an acquisition module 301, a mapping module 302, a calculation module 303, a first evolution module 304, a second evolution module 305, a translation module 306, and an evaluation module 307.
[0127] Wherein:
[0128] The acquisition module 301 is configured to acquire the text to be translated and determine the relative position of each word in the sentence to be translated and its corresponding feature vector;
[0129] The mapping module 302 is configured to map the relative position of each word in the sentence to be translated into a generalized coordinate in Hamiltonian dynamics, map the feature vector of each word in the sentence to be translated into an active quantity corresponding to the relative position, and use an energy conversion network to convert the sentence to be translated into an energy form representation;
[0130] The calculation module 303 is configured to calculate the Hamiltonian of each word in the initial state and the total Hamiltonian of the sentence to be translated according to the energy form representation, as the initial energy of each word and the initial total energy of the sentence to be translated;
[0131] The first evolution module 304 is configured to perform weight initialization using the Xavier method and input the initialized result into the Hamiltonian deep neural network HDNN to extract the feature of the word information with the goal of semantic energy conservation, and obtain the relative position and active quantity evolved at each time point;
[0132] The second evolution module 305 is configured to input the processing result of the HDNN into the VI Verlet integrator for further evolution according to the time step, and obtain the further evolved relative position and active quantity;
[0133] The translation module 306 is configured to input the further evolved relative position and active quantity into a translation network based on Transformer to obtain a translation result;
[0134] The evaluation module 307 is configured to recalculate the Hamiltonian of each word and the total Hamiltonian of the sentence to be translated using an energy evaluation network, and calculate the energy difference to obtain a final result.
[0135] The translation system based on the semantic energy conservation network provided in this embodiment can implement the steps and processes of the translation method based on the semantic energy conservation network provided in any of the above embodiments and achieve the same technical effects, which will not be elaborated here one by one.
[0136] The translation method based on the semantic energy conservation network provided by any of the above embodiments can be applied to Figure 5 the electronic device 500 shown in the figure. The electronic device 500 may be, but is not limited to, mobile terminals such as mobile phones, tablet computers, handheld computers, personal digital assistants (PDAs), smart home devices such as smart TVs and smart cameras, wearable devices such as smart bracelets, smart watches, and smart glasses, or other computer devices such as desktop computers, laptop computers, ultra-mobile personal computers (UMPCs), netbooks, and smart screens.
[0137] As Figure 5 shown in the figure, the electronic device 500 may include one or more of the following components: a processor 501, a memory 503, a communication interface 502, and a communication bus 504. Among them, the memory 503 may be connected to the processor 501 through the bus 504. The bus may transmit data between the processor 501 and the memory 503. The bus may be divided into an address bus, a data bus, a control bus, etc.
[0138] The processor 501 may include one or more processing cores. The processor 501 can connect various parts within the entire electronic device 500 using various interfaces and circuits. By running or executing instructions, programs, code sets, or instruction sets stored in the memory 503, and by invoking the data stored in the memory 503, it can perform various functions of the electronic device 500 and process data. Exemplarily, the processor 501 may include an application processor (AP), a modem processor, a CPU, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a video codec, a digital signal processor (DSP), a field-programmable gate array (FPGA), a programmable logic array (PLA), and / or a neural-network processing unit (NPU), etc. Among them, the CPU mainly processes the operating system, user interface, and application programs, etc.; the GPU is responsible for rendering and drawing the content to be displayed; the NPU is used to implement artificial intelligence (AI) functions; the modem is used to process wireless communication. Different processing units can be independent devices or integrated in one or more processors. For example, the multiple processing units shown above are all integrated in one SoC, or the AP is a separate semiconductor chip and other processing units are integrated in one SoC. This application does not make any limitations in this regard.
[0139] The memory 503 may include a random access memory (RAM), may also include a read-only memory (ROM), and may further include a non-transitory computer-readable storage medium. The memory 503 can be used to store instructions, programs, code, code sets, or instruction sets. The memory 503 may include a program storage area and a data storage area. Among them, the program storage area can store instructions for implementing the operating system, instructions for at least one function, such as the three-dimensional rough fracture non-linear flow field simulation method, the three-dimensional rough fracture long-term corrosion mechanism analysis method, etc.; the data storage area can store data created according to the use of the electronic device 500, such as the input data for flow field simulation, etc.
[0140] In addition, those skilled in the art can understand that the structure of the electronic device 500 shown in the above drawings does not limit the electronic device 500. The electronic device may include more or fewer components than shown in the drawings, or combine certain components, or have different component arrangements. For example, the electronic device 500 further includes components such as a microphone, a speaker, a radio frequency circuit, a sensor, an audio circuit, a power supply, and a Bluetooth module, which will not be elaborated here.
[0141] The embodiment of the present application also provides a computer program product including computer-executable instructions. In one embodiment, the computer-executable instructions are used to cause a computer to execute the functions in the above method embodiment.
[0142] The computer-executable instructions can be stored in a computer-readable storage medium. The embodiment of the present application also provides a computer-readable storage medium storing executable instructions. In one embodiment, the computer-executable instructions are used to cause a computer to execute the functions in the above method embodiment.
[0143] The computer-readable storage medium provided by the embodiment of the present application may be a random access memory (RAM), a flash memory, a read-only memory (ROM), a programmable ROM (PROM), an erasable PROM (EPROM), an electrically erasable PROM (EEPROM), a register, a hard disk, a removable hard disk, a CD-ROM, or any other form of computer-readable storage medium well-known in the art.
[0144] The computer-executable instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer program or instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center in a wired or wireless manner. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center integrating one or more available media. The available medium can be a magnetic medium, such as a floppy disk, a hard disk, or a magnetic tape; it can also be an optical medium, such as a digital video disc (DVD); or it can be a semiconductor medium, such as a solid-state drive.
[0145] The above are only the preferred embodiments of the present application and are not intended to limit the present application. For those skilled in the art, various changes and modifications can be made to the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included within the protection scope of the present application.
Claims
1. A translation method based on a semantic energy conservation network, characterized in that: include: Obtain the text to be translated, and determine the relative position of each word in the sentence to be translated and its corresponding feature vector; The relative position of each word in the sentence to be translated is mapped to the generalized coordinates in Hamiltonian dynamics, the feature vector of each word in the sentence to be translated is mapped to the activity corresponding to the relative position, and the sentence to be translated is converted into energy form using an energy conversion network; According to the energy form, the Hamiltonian of each word in the initial state and the total Hamiltonian of the sentence to be translated are calculated as the initial energy of each word and the initial total energy of the sentence to be translated; The Xavier method is used to initialize the weights, and the initialization results are input into the Hamiltonian deep neural network HDNN. The feature extraction of vocabulary information is carried out with the goal of semantic energy conservation, and the relative position and activity of the evolution at each time point are obtained. The processing result of HDNN is input into VI Verlet integrator, and further evolved according to the time step to obtain the relative position and activity after further evolution; The further evolved relative position and activity are input into the Transformer-based translation network to obtain the translation result; The energy evaluation network is used to recalculate the Hamiltonian of each word and the total Hamiltonian of the sentence to be translated, and the energy difference is calculated to obtain the final result.
2. The method according to claim 1, characterized in that The energy form is specifically expressed as: E 0i =H(p i ,q i ), In the formula, E 0i is the initial energy of the i-th word, through H(p i ,q i ) is obtained, H is the Hamiltonian, p i is the activity of the i-th word, q i is the relative position of the i-th word, i = 1, 2, ..., n, E tot0 is the initial total energy of the sentence to be translated.
3. The method according to claim 2, characterized in that The Hamiltonian H is expressed as: Where U and V are parameter matrices, q is a vector consisting of the relative positions of all words, and p is a vector consisting of the activity of all words.
4. The method according to claim 1, characterized in that: HDNN is a time-varying Hamiltonian system; the dynamic equation of the time-varying Hamiltonian system is expressed as follows: In the formula, is the state vector, t is the time, M(t) is the time-varying skew-symmetric matrix, y (0) is the initial state vector, including the initial energy of each word and the initial total energy of the sentence to be translated.
5. The method according to claim 4, characterized in that In the Hamiltonian deep neural network HDNN, the S-IE discretization method is used to discretize the extended Hamiltonian to obtain the relative position and activity of each word at each time point. The expression is as follows: q i+1 =q i -hLM q,i σ(M p,i q i +b p,i ), Where i represents the i-th time step, p i 、p i+1 is the relative position of adjacent time steps, q i ,q i+1 is the activity of adjacent time steps, h is the time step, L and M are parameter matrices, L T 、M T are the transposed matrices of L and M respectively, M p,i is related to the variable p i The parameter matrix M that is specifically represented at the i-th time step q,i is related to the variable q i The parameter matrix with specific representation at the i-th time step, b q,i Represents the variable q i is a bias term that is relevant and has a specific value at the i-th time step, b p,i Represents the variable p i is related to , and is a bias term with a specific value at the i-th time step.
6. A translation system based on a semantic energy conservation network, characterized in that: include: An acquisition module is used to acquire the text to be translated and determine the relative position of each word in the sentence to be translated and its corresponding feature vector; A mapping module is used to map the relative position of each word in the sentence to be translated into a generalized coordinate in Hamiltonian dynamics, map the feature vector of each word in the sentence to be translated into an activity corresponding to the relative position, and use an energy conversion network to convert the sentence to be translated into an energy form representation; A calculation module, used for calculating the Hamiltonian of each word in the initial state and the total Hamiltonian of the sentence to be translated according to the energy form representation, as the initial energy of each word and the initial total energy of the sentence to be translated; The first evolution module is used to initialize weights using the Xavier method and input the initialization results into the Hamiltonian deep neural network HDNN to extract features of vocabulary information with the goal of semantic energy conservation, and obtain the relative position and activity of the evolution at each time point; The second evolution module is used to input the processing result of HDNN into the VI Verlet integrator, further evolve according to the time step, and obtain the relative position and activity after further evolution; The translation module is used to input the further evolved relative positions and activity amounts into the Transformer-based translation network to obtain the translation results; The evaluation module is used to recalculate the Hamiltonian of each word and the total Hamiltonian of the sentence to be translated using the energy evaluation network, and calculate the energy difference to obtain the final result.
7. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the translation method based on the semantic energy conservation network as described in any one of claims 1 to 5 are implemented.
8. An electronic device, characterized in that: include: A memory, a processor, and a program stored in the memory and executable on the processor, wherein when the processor executes the program, the steps of the translation method based on the semantic energy conservation network as described in any one of claims 1 to 5 are implemented.
9. A computer program product, characterized in that The method comprises computer executable instructions for causing a computer to execute the steps of the translation method based on the semantic energy conservation network as described in any one of claims 1 to 5.
Citation Information
Patent Citations
Non-autoregressive neural machine translation decoding method and device, equipment and storage medium
CN114611505A
Arbitrary continuous time perception model for simulating mixed synaptic transfer and training method
CN116861974A
Method and device for calculating energy of target system and medium
CN117114119A
KDDIN-based aero-engine residual life prediction method
CN117371331A
Gradient-based quantum-assisted Hamiltonian learning
CN118176512A