Language translation data processing method and system based on large model
Through the big model-based language translation data processing method, the pre-trained language model extracts word semantic features and constructs a dependency tree structure, the problem of insufficient implicit semantic association recognition in long text translation is solved, and a higher quality cross-segment semantic guidance translation is achieved.
Patent Information
- Application Number
- CN202510264085.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-06
- Publication Date
- 2025-07-04
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing machine translation techniques are difficult to identify implicit semantic associations when processing long texts, resulting in insufficient logical coherence and semantic accuracy, especially in long sentences, composite sentences or chapter-level translation scenarios.
Through the big model-based language translation data processing method, the pre-trained language model extracts word semantic features, encodes segmentation, and constructs a dependency tree structure, and combines the dependencies between explicit and implicit semantics to generate paragraph labels for translation processing.
The boundary division of text fragments is optimized, the semantic consistency and logical coherence of the translation process are enhanced, and the translation quality of long and complex texts is improved.
Smart Images

Figure CN120258010A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of data processing. More specifically, this application relates to a method and system for processing language translation data based on a large model. Background Art
[0002] Data processing is a core link in ensuring translation quality and efficiency during machine translation. Its main function is to clean, structure, and extract features from the original text data to improve the training effect of machine learning or deep learning models. Specifically, data processing includes removing noise (such as special characters and meaningless symbols), word segmentation (dividing text into semantic units), part-of-speech tagging (tagging word categories), syntactic analysis (constructing sentence structure relationships), alignment processing (establishing the corresponding relationship between the source language and the target language), etc. These steps not only improve the readability of the data but also reduce the computational complexity of the translation model and enhance the ability to understand the context.
[0003] In existing machine translation technologies, text translation usually relies on sequence-to-sequence models or Transformer models based on self-attention mechanisms. These models are mainly based on sentence-level or paragraph-level language modeling. However, the main problem with existing translation data processing methods when dealing with long texts is the weak ability to recognize implicit semantic associations. Especially in long sentence, complex sentence, or discourse-level translation scenarios, it is difficult to maintain logical coherence and semantic accuracy. Existing technologies usually rely on end-to-end training of neural networks but fail to fully utilize the dependency syntactic relationships and semantic features of the text, lacking effective semantic guidance for the division of text segments, resulting in the omission of implicit context associations during translation and affecting the overall translation quality. Therefore, how to achieve cross-segment semantic-guided translation of long and complex texts has become a difficult problem faced by the industry. Summary of the Invention
[0004] This application provides a method and system for processing language translation data based on a large model, which can achieve cross-segment semantic-guided translation of long and complex texts.
[0005] In a first aspect, this application provides a method for processing language translation data based on a large model, including the following steps: Receiving the language translation data to be processed; Performing semantic feature extraction on the words in the language translation data based on a pre-trained large language model to obtain the semantic features of each word, and then encoding and segmenting the language translation data according to all the semantic features to obtain multiple text segments, and determining the segmentation loss of each text segment; Perform dependency syntactic analysis on each text segment to obtain the dependency relationships between the words in the text segment, and then determine the dependency tree structure of each text segment. Determine the explicit semantic dependency relationships between every two text segments based on the semantic similarity between every two text segments and the dependency tree structures of the corresponding text segments; Perform implicit association on the segment semantics between each text segment according to the full text information of the language translation data to obtain the implicit semantic association degrees between each text segment, and then perform dependency analysis on the implicit semantic relationships between every two text segments based on all the implicit semantic association degrees and the segmentation losses of each text segment to obtain the implicit semantic dependency relationships between every two text segments; Construct the paragraph tags of the language translation data through the explicit semantic and implicit semantic dependency relationships between every two text segments, and perform translation processing on the data to be translated based on the paragraph tags.
[0006] Preferably, semantic feature extraction is performed on the words in the language translation data based on a pre-trained large language model, and the semantic features of each word are obtained, specifically including: Perform word segmentation processing on the language translation data to obtain the word vectors of each word; Extract the semantic features of each word from all the word vectors through the hidden layer of the large language model.
[0007] Preferably, encoding segmentation is performed on the language translation data according to all the semantic features to obtain multiple text segments, and the segmentation losses of each text segment are determined, specifically including: Determine the encoding value of each word according to the semantic feature similarity between adjacent words; Perform segmentation on the language translation data according to all the encoding values to obtain multiple text segments; Predict the segmentation losses of each text segment during the segmentation process through a preset loss model.
[0008] Preferably, perform dependency syntactic analysis on each text segment to obtain the dependency relationships between the words in the text segment, and then determine the dependency tree structure of each text segment, specifically including: For each text segment, determine the co-occurrence probability between every two words in the text segment; Perform syntactic analysis on the text segment to identify the syntactic relationships between the words in the text segment; Determine the dependency relationships between the words in the text segment through the co-occurrence probability between every two words in the text segment and the syntactic relationships; Identify the core word of the text segment and use the core word as the root node of the dependency tree; Construct the dependency tree structure of the text fragment based on the dependency relationships between the words in the text fragment and the root node, and then determine the dependency tree structure of each text fragment.
[0009] Preferably, determining the explicit semantic dependency relationship between every two text fragments according to the semantic similarity between every two text fragments and the dependency tree structure of the corresponding text fragment specifically includes: For every two text fragments, determine the Pearson correlation coefficient of the syntactic structure between the two text fragments according to the dependency tree structures of the two text fragments; Extract the feature vectors of the explicit semantics of the two text fragments through a pre-trained large language model; Determine the semantic similarity between the two text fragments according to the Euclidean distance between the two feature vectors; Determine the explicit semantic dependency relationship between the two text fragments through the Pearson correlation coefficient and the semantic similarity, and then determine the explicit semantic dependency relationship between every two text fragments.
[0010] Preferably, perform implicit association on the fragment semantics between each text fragment according to the full text information of the language translation data to obtain the implicit semantic association degree between each text fragment specifically includes: For every two text fragments, determine the association relationship of the words between the two text fragments in the word order structure according to the full text information of the language translation data; Extract the feature vectors of the implicit semantics of the two text fragments through a pre-trained large language model; Determine the implicit semantic association degree between the two text fragments according to the association relationship of the words between the two text fragments in the word order structure and the feature vectors of the implicit semantics of the two text fragments, and then determine the implicit semantic association degree between every two text fragments.
[0011] Preferably, perform dependency analysis on the implicit semantic relationship between every two text fragments based on all the implicit semantic association degrees and the segmentation losses of each text fragment to obtain the implicit semantic dependency relationship between every two text fragments specifically includes: For every two text fragments, obtain the segmentation losses of the two text fragments, and compensate the implicit semantics of the corresponding text fragments through the two segmentation losses respectively to obtain the semantic compensation values of the two text fragments; Determine the implicit semantic association compensation coefficient between the two text fragments according to the two semantic compensation values; Determine the implicit semantic dependency relationship between the two text fragments through the implicit semantic association degree between the two text fragments and the association compensation coefficient, and then determine the implicit semantic dependency relationship between every two text fragments.
[0012] Second aspect, the present application provides a large model-based language translation data processing system, including: A receiving module, configured to receive the language translation data to be processed; A processing module, configured to extract semantic features of words in the language translation data based on a pre-trained language large model, obtain semantic features of each word, and then encode and segment the language translation data according to all the semantic features to obtain multiple text segments, and determine the segmentation loss of each text segment; The processing module is further configured to perform dependency syntactic analysis on each text segment to obtain the dependency relationship between each word in the text segment, and then determine the dependency tree structure of each text segment, and determine the explicit semantic dependency relationship between each two text segments according to the semantic similarity between each two text segments and the dependency tree structure of the corresponding text segment; The processing module is further configured to implicitly associate the segment semantics between each text segment according to the full text information of the language translation data, obtain the implicit semantic association degree between each text segment, and then perform dependency analysis on the implicit semantic relationship between each two text segments based on all the implicit semantic association degrees and the segmentation loss of each text segment to obtain the implicit semantic dependency relationship between each two text segments; An execution module, configured to construct a paragraph label of the language translation data through the explicit semantic and implicit semantic dependency relationships between each two text segments, and perform translation processing on the data to be translated based on the paragraph label.
[0013] Third aspect, the present application provides a computer device, the computer device includes a memory and a processor, the memory stores code, and the processor is configured to obtain the code and execute the above-mentioned large model-based language translation data processing method.
[0014] Fourth aspect, the present application provides a computer-readable storage medium, the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the above-mentioned large model-based language translation data processing method is implemented.
[0015] The technical solution provided by the disclosed embodiments of the present application has the following beneficial effects: In the embodiments of the present application, first, language translation data to be processed is received; semantic features of words in the language translation data are extracted based on a pre-trained large language model to obtain semantic features of each word, and then the language translation data is encoded and segmented according to all the semantic features to obtain multiple text segments, and the segmentation loss of each text segment is determined; dependency syntactic analysis is performed on each text segment to obtain the dependency relationships between the words in the text segment, and then the dependency tree structure of each text segment is determined. The explicit semantic dependency relationship between every two text segments is determined according to the semantic similarity between every two text segments and the dependency tree structure of the corresponding text segments; the fragment semantics between the text segments is implicitly associated according to the full text information of the language translation data to obtain the implicit semantic association degree between the text segments, and then the implicit semantic relationship between every two text segments is analyzed based on all the implicit semantic association degrees and the segmentation losses of the text segments to obtain the implicit semantic dependency relationship between every two text segments; the paragraph label of the language translation data is constructed through the explicit semantic and implicit semantic dependency relationships between every two text segments, and the translation processing of the data to be translated is performed based on the paragraph label.
[0016] It can be seen that in the present application, the paragraph label of the language translation data is constructed through the explicit semantic and implicit semantic dependency relationships between every two text segments, and the translation processing of the data to be translated is performed based on the paragraph label. First, the semantic features of words are extracted through a pre-trained large language model, and the text is encoded and segmented based on the semantic features, so that the segment division not only depends on the syntactic structure, but also combines deep semantic information, thereby optimizing the boundary division of the text segments and reducing the possibility of semantic fragmentation between segments. Secondly, dependency syntactic analysis is performed within each text segment to construct a dependency tree structure to clarify the grammatical relationships between words, so that more accurate semantic transfer can be performed based on syntactic dependencies during the translation process, avoiding mistranslation caused by syntactic ambiguity. In addition, by calculating the semantic similarity between text segments and combining the dependency tree structure, the explicit semantic dependency relationship is clarified, so that the translation process not only focuses on local lexical matching, but also can optimize the translation strategy based on the overall syntactic structure. Then, the implicit semantic association degree between the text segments is analyzed according to the full text information of the language translation data, and then the implicit semantic dependency relationship is constructed, so that machine translation can capture the context information across sentences and text segments, thereby enhancing the coherence and semantic consistency of the text. Finally, by combining the explicit semantic dependency relationship and the implicit semantic dependency relationship, a paragraph-level label is constructed, so that the translation system can recognize the semantic structure at the text segment level in addition to the word level, and achieve a translation result that is more in line with the context logic. In summary, this solution can realize cross-segment semantic-guided translation of long and complex texts, thereby improving the translation quality of long texts and complex sentence patterns. Brief Description of the Drawings
[0017] Figure 1 is an exemplary flowchart of a large model-based language translation data processing method shown in some embodiments of the present application; Figure 2 is a schematic flowchart of determining the dependency relationship of explicit semantics shown in some embodiments of the present application; Figure 3 is a schematic flowchart of implementing language translation data processing shown in some embodiments of the present application; Figure 4 is a schematic structural diagram of a large model-based language translation data processing system shown in some embodiments of the present application; Figure 5 is a schematic structural diagram of a computer device implementing a large model-based language translation data processing method shown in some embodiments of the present application. Detailed implementation manners
[0018] To better understand the technical solutions of the present application, the technical solutions of the present application will be described in detail below in conjunction with the accompanying drawings of the specification and specific implementation manners.
[0019] Refer to Figure 1 , this figure is an exemplary flowchart of a large model-based language translation data processing method shown in some embodiments of the present application. The large model-based language translation data processing method 100 mainly includes the following steps: In step 101, receive the language translation data to be processed.
[0020] It should be noted that the language translation data in the present application is text-type data. In other embodiments, it may also be other types of data, which are not specifically limited here.
[0021] Specifically, receiving the language translation data to be processed can be implemented in the following manner, that is: the language translation data to be processed can be received from the input interface of the translation software, and the format of the language translation data is text format.
[0022] In step 102, based on the pre-trained language large model, extract semantic features of the words in the language translation data to obtain the semantic features of each word. Then, encode and segment the language translation data according to all the semantic features to obtain multiple text segments, and determine the segmentation loss of each text segment.
[0023] In some embodiments, extracting semantic features of the words in the language translation data based on the pre-trained language large model to obtain the semantic features of each word can be implemented by the following steps: Perform word segmentation processing on the language translation data to obtain the word vectors of each word; Extract the semantic features of each word from all word vectors through the hidden layer of the language model.
[0024] It should be noted that the language model in this application can be an autoencoder model or an autoregressive model. The training process of the language model usually includes three stages: preprocessing, pre-training, and fine-tuning. Each stage is closely linked to improve the language understanding ability of the model. First, in the preprocessing stage, it is necessary to construct a large-scale corpus, use a tokenization algorithm (such as WordPiece, SentencePiece) to segment the text, convert it into a vocabulary index sequence, and convert the text into a continuous vector through One-hot encoding or word embedding (Embedding). At the same time, construct a training data format suitable for different tasks. For example, Bidirectional Encoder Representations from Transformers (BERT) uses the Masked Language Model (MLM) and Next Sentence Prediction (NSP) tasks to construct the input. Second, in the pre-training stage, the model uses self-supervised learning to perform representation learning on the text through a deep Transformer architecture (composed of multiple layers of self-attention mechanisms and feedforward neural networks (FFN)). For example, in the MLM task of BERT, the model will randomly mask some words and train the model to predict the masked content, enabling it to learn to infer words based on the context. In addition, during the training process, gradient descent is used to optimize the cross-entropy loss, and an adaptive optimization algorithm (such as AdamW) is used to update the model parameters to improve the model convergence speed and generalization ability. Finally, in the fine-tuning stage, the pre-trained language model is applied to specific downstream tasks, such as machine translation, text summarization, question answering systems, etc. Through transfer learning, it is further trained on the target task dataset to adapt to specific scenarios. For example, in the machine translation task, an encoder-decoder architecture can be used for fine-tuning to optimize the translation quality.
[0025] In specific implementation, the tokenization process is performed on the language translation data to obtain the word vectors of each word, which can be achieved in the following way: The WordPiece tokenization algorithm can be used to perform sub-word level segmentation on the input language translation data to obtain multiple words, and then each word is mapped to an ID to convert it into an index sequence that can be input into the language model. Then, this index sequence is input into the pre-trained language model, and through the embedding layer of the language model, the index sequence is converted into a vector, and this vector is used as the word vector. This word vector not only contains the basic information of the word but also its relationship in the context. The semantic features of each word are extracted from all the word vectors through the hidden layer of the language large model, which can be achieved in the following way: The language large model uses multiple layers of Transformer encoders to calculate the relationship between each word and other words in the sentence through the self-attention mechanism, capture the context-dependent information of the word, and then use the last hidden state of the language large model (such as the output of the last Transformer layer of BERT) as the semantic feature of the corresponding word.
[0026] It should be noted that the semantic feature in this application is a feature that reflects the meaning of a word in a specified context and its relationship with other language units.
[0027] In some embodiments, encoding and segmenting the language translation data based on all the semantic features to obtain multiple text segments, and determining the segmentation loss of each text segment can be achieved through the following steps: Determine the encoding value of each word according to the similarity of the semantic features between adjacent words; Segment the language translation data based on all the encoding values to obtain multiple text segments; Predict the segmentation loss of each text segment during the segmentation process through a preset loss model.
[0028] It should be noted that the encoding value in this application is a numerical value used to quantitatively represent the similarity degree of the semantic information of adjacent words, which reflects the semantic features of the word in a specified context; the segmentation loss in this application is an index that measures the degree of semantic information loss caused by the segmentation strategy during the text segmentation of the language translation data. It is used to evaluate whether the text segments after segmentation can completely retain the original semantic structure, ensure the semantic coherence between adjacent segments, and optimize the segmentation strategy to reduce the loss of important semantics.
[0029] It should also be noted that the loss model in this application refers to a machine learning model used to evaluate the degree of information loss of text segments during semantic processing. Its main function is to measure the accuracy and integrity of input data in the encoding segmentation task, ensuring that the translation model can effectively optimize the fidelity of semantic expression; during the evaluation process, the cross-entropy loss in the prior art can be used to measure the semantic information loss caused by text segmentation.
[0030] When specifically implemented, the encoding value of each word can be determined according to the similarity of semantic features between adjacent words in the following way, that is: select a word as the selected word, take the word adjacent to the left of the selected word as the left adjacent word, and take the word adjacent to the right of the selected word as the right adjacent word. The semantic vectors of the selected word, the left adjacent word, and the right adjacent word can be extracted using a pre-trained large language model. Then, the cosine similarity is used to calculate the similarity between the semantic vectors of the selected word and the left adjacent word, and this similarity is used as the left similarity. The cosine similarity is used to calculate the similarity between the semantic vectors of the selected word and the right adjacent word, and this similarity is used as the right similarity. Then, the average value between the left similarity and the right similarity is used as the encoding value of the selected word, and then the encoding values of the remaining words are obtained; segmenting the language translation data based on all the encoding values to obtain multiple text segments can be implemented in the following way, that is: by setting an encoding threshold, when the encoding value of a word is lower than this encoding threshold, it is determined that the word is the starting point of a new text segment, and then the language translation data is divided into multiple coherent segments, and the divided segments are used as text segments; predicting the segmentation loss of each text segment during the segmentation process using a preset loss model can be implemented in the following way, that is: the cross-entropy loss function in the preset loss model can be used to calculate the loss value of each text segment during the segmentation process, that is, the cross-entropy loss value of each text segment can be used as the segmentation loss of the corresponding text segment.
[0031] It should be noted that the training process of the loss model in this application mainly includes steps such as data preprocessing, model construction, forward propagation, loss calculation, gradient backpropagation, and parameter update, ensuring that the model can learn reasonable text segment division and semantic dependency relationships. The specific training process is as follows: Data preprocessing includes text cleaning, feature extraction, and feature annotation; Model construction includes parameter settings for the input layer, encoding layer, and prediction layer; Forward propagation can calculate the segmentation probability and dependency relationships of text segments; Loss calculation can use the cross-entropy loss function to calculate the loss of each text segment; Gradient backpropagation and optimization can update parameters to improve the model performance; Model evaluation can use metrics to measure the model quality.
[0032] In step 103, dependency syntactic analysis is performed on each text segment to obtain the dependency relationships between the words in the text segment, and then the dependency tree structure of each text segment is determined. The explicit semantic dependency relationship between every two text segments is determined according to the semantic similarity between every two text segments and the dependency tree structure of the corresponding text segments.
[0033] In some embodiments, the determination of the dependency tree structure of each text segment by performing dependency syntactic analysis on each text segment to obtain the dependency relationships between the words in the text segment can be achieved by the following steps: For each text segment, determine the co-occurrence probability between every two words in the text segment; Perform syntactic analysis on the text segment to identify the syntactic relationships between the words in the text segment; Determine the dependency relationships between the words in the text segment through the co-occurrence probability between every two words in the text segment and the syntactic relationships; Identify the core word of the text segment and use the core word as the root node of the dependency tree; Construct the dependency tree structure of the text segment through the dependency relationships between the words in the text segment and the root node, and then determine the dependency tree structure of each text segment.
[0034] It should be noted that the co-occurrence probability in this application is an index for measuring the probability that two words appear simultaneously in a given context; the syntactic relationship in this application is an index for measuring the grammatical structure connection between the words in the text segment; the dependency relationship in this application is an index for measuring the mutual dependency relationship between the words in a sentence; the dependency tree structure in this application is a representation method that connects the words in a sentence into a tree-like structure by dependency relationships.
[0035] In specific implementation, the co-occurrence probability between every two words in a text fragment can be determined in the following manner, that is: the co-occurrence times between every two words in the text fragment can be queried through a large-scale corpus in the prior art, and the ratio between the co-occurrence times and the average occurrence times of every two words appearing alone is used as the co-occurrence probability between every two words in the text fragment; syntactic analysis of the text fragment to identify the syntactic relationship between each word in the text fragment can be implemented in the following manner, that is: a dependency syntactic analysis model (such as a dependency parser based on BERT+BiLSTM or a dependency analysis method based on Transition-Based Parsing) can be used to parse the structural relationship of the text fragment, output the syntactic role of each word (such as subject, predicate, object), and then mark the syntactic role of each word to obtain the syntactic role table of the words in the text fragment, and thus the syntactic relationship between each word in the text fragment can be reflected through the syntactic role table; determining the dependency relationship between each word in the text fragment based on the co-occurrence probability between every two words in the text fragment and the syntactic relationship can be implemented in the following manner, that is: for every two words, the link relationship composed of the co-occurrence probability between the two words and the syntactic role relationship between the words (the syntactic role relationship can be obtained by querying the syntactic role table) can be used as the dependency relationship between the two words, and thus the dependency relationship between every two words can be obtained. For example: the co-occurrence probability between "bank" and "deposit" is "A", and "deposit" is the object of "bank", then a dependency relationship (A(bank→deposit)) is formed; identifying the core word of the text fragment and using the core word as the root node of the dependency tree can be implemented in the following manner, that is: a TF-IDF model in the prior art can be used to identify the core word of the text fragment, such as the verb or subject in the subject-predicate structure, and the verb or subject is used as the core word; constructing the dependency tree structure of the text fragment based on the dependency relationship between each word in the text fragment and the root node can be implemented in the following manner, that is: a minimum spanning tree algorithm in the prior art can be used to orderly connect each word to the root node according to the dependency relationship between each word in the text fragment to obtain a directed tree structure, and this directed tree structure is used as the dependency tree structure of the text fragment.
[0036] In some embodiments, referring to Figure 2 as shown, this figure is a schematic flowchart of determining the dependency relationship of explicit semantics in some embodiments of the present application. In this embodiment, determining the dependency relationship of explicit semantics between every two text fragments based on the semantic similarity between every two text fragments and the dependency tree structure of the corresponding text fragment can be implemented through the following steps: In step 1031, for every two text fragments, the Pearson correlation coefficient of the syntactic structure between the two text fragments is determined according to the dependency tree structures of the two text fragments; In step 1032, feature vectors of the explicit semantics of two text segments are extracted by a pre-trained large language model; In step 1033, the semantic similarity between the two text segments is determined according to the Euclidean distance between the two feature vectors; In step 1034, the dependency relationship of the explicit semantics between the two text segments is determined by the Pearson correlation coefficient and the semantic similarity, and then the dependency relationship of the explicit semantics between each two text segments is determined.
[0037] It should be noted that explicit semantics refers to the directly expressed and clearly visible semantic information, which is usually directly conveyed through forms such as words and sentence structures in language. For example, the meanings expressed by "cat" and "chair" and the action "sit" between them in the sentence "The cat is sitting on the chair" are explicit semantics, which are directly presented through clear language units and syntactic structures.
[0038] It should also be noted that the semantic similarity in this application is an indicator to measure the similarity degree between two text segments at the explicit semantic level; the dependency relationship of the explicit semantics in this application is the mutual dependency relationship directly shown between text segments at the explicit semantic level.
[0039] In specific implementation, the Pearson correlation coefficient of the syntactic structure between the two text segments can be determined according to the dependency tree structures of the two text segments in the following way, that is: the Pearson correlation coefficient of the dependency tree structures between the two text segments can be used as the Pearson correlation coefficient of the syntactic structure between the two text segments; the semantic similarity between the two text segments can be determined according to the Euclidean distance between the two feature vectors in the following way, that is: the Euclidean distance between the two feature vectors can be used as the semantic similarity between the two text segments, and a smaller Euclidean distance means that the two text segments are closer in explicit semantics; the dependency relationship of the explicit semantics between the two text segments can be determined by the Pearson correlation coefficient and the semantic similarity in the following way, that is: the product of the Pearson correlation coefficient and the semantic similarity can be used as the dependency value of the explicit semantics between the two text segments, that is, the dependency relationship of the explicit semantics between the two text segments can be described by this dependency value.
[0040] In step 104, the implicit association of the segment semantics between each text segment is performed according to the full text information of the language translation data to obtain the association degree of the implicit semantics between each text segment, and then the dependency analysis of the implicit semantic relationship between each two text segments is performed based on all the association degrees of the implicit semantics and the segmentation loss of each text segment to obtain the dependency relationship of the implicit semantics between each two text segments.
[0041] In some embodiments, the implicit association of the fragment semantics between each text fragment is obtained according to the full text information of the language translation data. The following steps can be used to achieve the association degree of the implicit semantics between each text fragment: For every two text fragments, determine the association relationship of the words between the two text fragments in terms of word order structure according to the full text information of the language translation data; Extract the feature vectors of the implicit semantics of the two text fragments through a pre-trained large language model; Determine the association degree of the implicit semantics between the two text fragments according to the association relationship of the words between the two text fragments in terms of word order structure and the feature vectors of the implicit semantics of the two text fragments, and then determine the association degree of the implicit semantics between every two text fragments.
[0042] It should be noted that implicit semantics refers to semantic information that is not directly expressed or clearly stated, but can be indirectly inferred through context, reasoning, or background knowledge. Implicit semantics often depends on the context or context of the sentence and is not directly reflected on the language surface. For example, the sentence "She went out with an umbrella" does not explicitly mention the information of "raining", but it can be inferred from common sense and context that "raining" is the implicit semantics.
[0043] It should also be noted that the association degree of the implicit semantics in this application describes the correlation strength shown by two text fragments through context information and potential semantic connections without direct syntactic matching.
[0044] In specific implementation, to determine the correlation relationship of the word order structure between two text segments according to the full text information of the language translation data, the following method can be adopted, that is: the word order statistical methods in the prior art (such as n-gram model, sliding window) or syntactic analysis (such as dependency syntactic analysis) can be used to extract the word order features of the words in the two text segments, and then the co-occurrence matrix in the prior art is used to calculate the relative position relationship of the word order features between the two text segments, and this relative position relationship is used as the correlation relationship of the words in the two text segments in the word order structure; to extract the feature vectors of the implicit semantics of the two text segments through a pre-trained large language model can be achieved in the following way, that is: the pre-trained large language model can be used to extract the implicit semantic feature vectors of each text segment. This large language model is trained through a large-scale corpus and can generate context-related semantic representations of the text segments. These vectors can capture the implicit information in the text segments, such as potential emotions, intentions or implicit relationships; to determine the correlation degree of the implicit semantics between the two text segments according to the correlation relationship of the words in the two text segments in the word order structure and the feature vectors of the implicit semantics of the two text segments can be achieved in the following way, that is: the cosine similarity of the feature vectors between the implicit semantics of the two text segments can be calculated using cosine similarity, and the product of this cosine similarity and the correlation relationship of the words in the two text segments in the word order structure is used as the correlation degree of the implicit semantics between the two text segments.
[0045] In some embodiments, based on the correlation degrees of all implicit semantics and the segmentation losses of each text segment, a dependency analysis is performed on the implicit semantic relationships between each pair of text segments, and the dependency relationships of the implicit semantics between each pair of text segments can be obtained by the following steps: For each pair of text segments, obtain the segmentation losses of the two text segments, and compensate the implicit semantics of the corresponding text segments through the two segmentation losses respectively to obtain the semantic compensation values of the two text segments; Determine the correlation compensation coefficient of the implicit semantics between the two text segments according to the two semantic compensation values; Determine the dependency relationship of the implicit semantics between the two text segments through the correlation degree of the implicit semantics between the two text segments and the correlation compensation coefficient, and further determine the dependency relationship of the implicit semantics between each pair of text segments.
[0046] It should be noted that the semantic compensation value in this application is an index that measures the loss of semantic information caused by text segmentation and compensates for the semantics of text segments through model adjustment; the associated compensation coefficient in this application is an index used to adjust the implicit semantic association strength between text segments to compensate for semantic deviation caused by segmentation loss or context fragmentation; the dependency relationship of implicit semantics in this application is used to represent the semantic dependency structure formed between two text segments through context information, semantic reasoning, or potential theme connections without direct syntactic or lexical connections.
[0047] In specific implementation, the implicit semantics of the corresponding text segments are compensated by two segmentation losses respectively. The semantic compensation values of the two text segments can be obtained in the following way, that is: the word vector interpolation method in the prior art can be used to map the segmentation losses into the semantic space of the corresponding text segments respectively, and the mapping values are used as the semantic compensation values of the text segments, so that the semantic compensation values of the two text segments can be obtained; the associated compensation coefficient of the implicit semantics between the two text segments can be determined according to the two semantic compensation values in the following way, that is: the cosine similarity in the prior art can be used to calculate the similarity between the two semantic compensation values, and the similarity is used as the associated compensation coefficient of the implicit semantics between the two text segments; the dependency relationship of the implicit semantics between the two text segments can be determined by the association degree of the implicit semantics between the two text segments and the associated compensation coefficient in the following way, that is: the weighted sum value between the association degree of the implicit semantics between the two text segments and the associated compensation coefficient can be used as the dependency relationship of the implicit semantics between the two text segments, where the weighted coefficients of the association degree and the associated compensation coefficient are adjustable parameters and can be adjusted according to historical experimental data, and their value ranges are between 0 and 1.
[0048] In step 105, the paragraph labels of the language translation data are constructed based on the dependency relationships of the explicit semantics and the implicit semantics between every two text segments, and the translation processing of the data to be translated is performed based on the paragraph labels.
[0049] In some embodiments, the construction of the paragraph labels of the language translation data based on the dependency relationships of the explicit semantics and the implicit semantics between every two text segments can be implemented by the following steps: Construct the paragraph labels of the explicit semantics during the translation process of the language translation data through the dependency relationships of the explicit semantics between every two text segments; Construct the paragraph labels of the implicit semantics during the translation process of the language translation data through the dependency relationships of the implicit semantics between every two text segments; Fuse the paragraph labels of the explicit semantics and the paragraph labels of the implicit semantics to obtain the paragraph labels of the language translation data.
[0050] It should be noted that the paragraph tags in this application are annotation information used to identify the semantic relationships and structural levels between text segments. The paragraph tags can be used in tasks such as machine translation, text summarization, and automatic question answering to help the translation model understand the logical structure of the text, improve semantic coherence, and context consistency.
[0051] It should also be noted that, as shown in Figure 3 the figure is a schematic flowchart of implementing language translation data processing in some embodiments of this application. The flowchart shows how to generate paragraph tags to assist in the processing of language translation data through the analysis of explicit and implicit semantic dependency relationships. Specifically, when implemented, the paragraph tags of explicit semantics in the translation process of the language translation data can be constructed by the explicit semantic dependency relationship between every two text segments in the following way, that is: the explicit semantic dependency relationship between adjacent text segments can be mapped to the explicit semantic dependency value between adjacent text segments, and this dependency value can be used as the explicit semantic tag value between adjacent text segments. Then, the list formed by all tag values in the order of the text segments in the language translation data is used as the paragraph tags of explicit semantics in the translation process of the language translation data; the paragraph tags of implicit semantics in the translation process of the language translation data can be constructed by the implicit semantic dependency relationship between every two text segments in the following way, that is: the implicit semantic dependency relationship between adjacent text segments can be mapped to the implicit semantic dependency value between adjacent text segments, and this dependency value can be used as the implicit semantic tag value between adjacent text segments. Then, the list formed by all tag values in the order of the text segments in the language translation data is used as the paragraph tags of implicit semantics in the translation process of the language translation data; the paragraph tags of the language translation data can be obtained by fusing the paragraph tags of explicit semantics and the paragraph tags of implicit semantics in the following way, that is: the list tags composed of the paragraph tags of explicit semantics and the paragraph tags of implicit semantics can be used as the paragraph tags of the language translation data, and the specific form of the list tags is: list tags = [paragraph tags of explicit semantics: paragraph tags of implicit semantics].
[0052] It should be noted that the translation processing of the data to be translated based on the paragraph tags in this application refers to translating the language translation data to be translated according to the paragraph tags. Specifically, in implementation: First, according to the paragraph tags of the text to be translated, analyze the relationships (explicit or implicit semantics) between text fragments, and determine the translation strategy based on these relationships. For example, for text fragments labeled as "causal relationship", the translation model should prioritize maintaining the coherence of the causal logic, while for paragraphs with a "supplementary explanation" relationship, appropriate expansion or explanation can be adopted during translation; Then, according to the explicit and implicit semantic information in the paragraph tags, combined with the context and theme in the text, integrate the semantics during the translation process. This process can use a pre-trained large language model to extract context information, and at the same time, combine the topic model and semantic similarity calculation to perform semantic compensation on each paragraph to ensure the semantic consistency of each paragraph during the translation process.
[0053] On the other hand, in some embodiments, this application provides a large model-based language translation data processing system. Refer to Figure 4 , this figure is a schematic structural diagram of a large model-based language translation data processing system shown according to some embodiments of this application. The large model-based language translation data processing system 400 includes: a receiving module 401, a processing module 402, and an execution module 403, which are described as follows: Receiving module 401. In this application, the receiving module 401 is mainly used to receive the language translation data to be processed; Processing module 402. In this application, the processing module 402 is used to extract semantic features of words in the language translation data based on a pre-trained large language model, obtain the semantic features of each word, and then encode and segment the language translation data based on all the semantic features to obtain multiple text fragments, and determine the segmentation loss of each text fragment; In this application, the processing module 402 is further used to perform dependency syntactic analysis on each text fragment to obtain the dependency relationships between the words in the text fragment, and then determine the dependency tree structure of each text fragment, and determine the explicit semantic dependency relationship between each two text fragments according to the semantic similarity between each two text fragments and the dependency tree structure of the corresponding text fragment; In this application, the processing module 402 is further used to implicitly associate the fragment semantics between each text fragment according to the full text information of the language translation data to obtain the implicit semantic association degree between each text fragment, and then perform dependency analysis on the implicit semantic relationship between each two text fragments based on all the implicit semantic association degrees and the segmentation loss of each text fragment to obtain the implicit semantic dependency relationship between each two text fragments; The execution module 403. In this application, the execution module 403 is mainly used to construct the paragraph tags of the language translation data through the explicit semantic and implicit semantic dependencies between every two text segments, and perform the translation processing of the data to be translated based on the paragraph tags.
[0054] In addition, this application also provides a computer device, which includes a memory and a processor. The memory stores code, and the processor is configured to obtain the code and execute the above-mentioned large model-based language translation data processing method.
[0055] In some embodiments, refer to Figure 5 , this figure is a schematic structural diagram of a computer device for implementing the large model-based language translation data processing method according to some embodiments of this application. The large model-based language translation data processing method in the above embodiments can be implemented by Figure 5 the computer device shown. The computer device 500 includes at least one processor 501, a communication bus 502, a memory 503, and at least one communication interface 504.
[0056] The processor 501 can be a general-purpose central processing unit (CPU) or an application-specific integrated circuit (ASIC).
[0057] The communication bus 502 can be used to transmit information between the above components.
[0058] The memory 503 can be a read-only memory (ROM) or other types of static storage devices that can store static information and instructions, a random access memory (RAM) or other types of dynamic storage devices that can store information and instructions, or an electrically erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM) or other optical disc storage, optical disc storage (including compact discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), magnetic disks or other magnetic storage devices, or any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto. The memory 503 can exist independently and be connected to the processor 501 through the communication bus 502. The memory 503 can also be integrated with the processor 501.
[0059] Among them, the memory 503 is used to store the program code for executing the solution of this application, and is controlled by the processor 501 for execution. The processor 501 is used to execute the program code stored in the memory 503. The program code may include one or more software modules. The above-described method for processing language translation data based on a large model in the embodiment can be implemented by one or more software modules in the processor 501 and the program code in the memory 503.
[0060] The communication interface 504, using any device such as a transceiver, is used to communicate with other devices or communication networks, such as Ethernet, radio access network (RAN), wireless local area networks (WLAN), etc.
[0061] In a specific implementation, as an embodiment, the computer device may include multiple processors, and each of these processors may be a single-core (single-CPU) processor or a multi-core (multi-CPU) processor. The processor here may refer to one or more devices, circuits, and / or processing cores for processing data (such as computer program instructions).
[0062] The above computer device may be a general-purpose computer device or a special-purpose computer device. In a specific implementation, the computer device may be a desktop computer, a laptop computer, a network server, a personal digital assistant (PDA), a mobile phone, a tablet computer, a wireless terminal device, a communication device, or an embedded device. The embodiments of this application do not limit the type of the computer device.
[0063] In addition, this application also provides a computer-readable storage medium, which stores a computer program, and when the computer program is executed by a processor, it implements the above-described method for processing language translation data based on a large model.
[0064] Although the preferred embodiments of this application have been described, those skilled in the art can make additional changes and modifications to these embodiments once they know the basic creative concepts. Therefore, the appended claims are intended to be interpreted to include the preferred embodiments and all changes and modifications falling within the scope of this application.
[0065] Obviously, those skilled in the art can make various changes and modifications to this application without departing from the spirit and scope of this application. Thus, if these modifications and variations of this application fall within the scope of the claims of this application and their equivalent technologies, this application is also intended to include these changes and modifications.
Claims
1. A method for processing language translation data based on a large model, characterized in that, Including the following steps: Receiving language translation data to be processed; Based on a pre-trained large language model, extracting semantic features of words in the language translation data to obtain semantic features of each word, and then encoding and segmenting the language translation data based on all semantic features to obtain multiple text segments, and determining the segmentation loss of each text segment; Performing dependency syntactic analysis on each text segment to obtain the dependency relationships between words in the text segment, and then determining the dependency tree structure of each text segment, and determining the explicit semantic dependency relationship between every two text segments according to the semantic similarity between every two text segments and the dependency tree structure of the corresponding text segments; Implicitly associating the segment semantics between each text segment according to the full text information of the language translation data to obtain the implicit semantic association degree between each text segment, and then performing dependency analysis on the implicit semantic relationship between every two text segments based on all implicit semantic association degrees and the segmentation loss of each text segment to obtain the implicit semantic dependency relationship between every two text segments; Constructing a paragraph label of the language translation data through the explicit semantic and implicit semantic dependency relationships between every two text segments, and performing translation processing on the data to be translated based on the paragraph label.
2. The method according to claim 1, characterized in that, Based on a pre-trained large language model, extracting semantic features of words in the language translation data to obtain semantic features of each word specifically includes: Performing word segmentation on the language translation data to obtain word vectors of each word; Extracting semantic features of each word from all word vectors through the hidden layer of the large language model.
3. The method according to claim 1, characterized in that, Encoding and segmenting the language translation data based on all semantic features to obtain multiple text segments, and determining the segmentation loss of each text segment specifically includes: Determining the encoding value of each word according to the semantic similarity between adjacent words; Segmenting the language translation data based on all encoding values to obtain multiple text segments; Predicting the segmentation loss of each text segment during the segmentation process through a preset loss model.
4. The method according to claim 1, characterized in that, Performing dependency syntactic analysis on each text segment to obtain the dependency relationships between words in the text segment, and then determining the dependency tree structure of each text segment specifically includes: For each text segment, determining the co-occurrence probability between every two words in the text segment; Performing syntactic analysis on the text segment to identify the syntactic relationships between words in the text segment; Determining the dependency relationships between words in the text segment through the co-occurrence probability between every two words in the text segment and the syntactic relationships; Identifying the core word of the text segment and using the core word as the root node of the dependency tree; Constructing the dependency tree structure of the text segment through the dependency relationships between words in the text segment and the root node, and then determining the dependency tree structure of each text segment.
5. The method according to claim 1, characterized in that, Determining the explicit semantic dependency relationship between every two text segments according to the semantic similarity between every two text segments and the dependency tree structure of the corresponding text segments specifically includes: For every two text segments, determine the Pearson correlation coefficient of the syntactic structure between the two text segments according to the dependency tree structures of the two text segments; Extract the feature vectors of the explicit semantics of the two text segments through a pre-trained large language model; Determine the semantic similarity between the two text segments according to the Euclidean distance between the two feature vectors; Determine the dependency relationship of the explicit semantics between the two text segments through the Pearson correlation coefficient and the semantic similarity, and further determine the dependency relationship of the explicit semantics between every two text segments.
6. The method according to claim 1, wherein Perform implicit association on the segment semantics between each text segment according to the full-text information of the language translation data, and obtain the association degree of the implicit semantics between each text segment, specifically including: For every two text segments, determine the association relationship of the words between the two text segments in the word order structure according to the full-text information of the language translation data; Extract the feature vectors of the implicit semantics of the two text segments through a pre-trained large language model; Determine the association degree of the implicit semantics between the two text segments according to the association relationship of the words between the two text segments in the word order structure and the feature vectors of the implicit semantics of the two text segments, and further determine the association degree of the implicit semantics between every two text segments.
7. The method according to claim 1, wherein Based on the association degrees of all implicit semantics and the segmentation losses of each text segment, perform dependency analysis on the implicit semantic relationships between every two text segments, and obtain the dependency relationships of the implicit semantics between every two text segments, specifically including: For every two text segments, obtain the segmentation losses of the two text segments, and compensate the implicit semantics of the corresponding text segments through the two segmentation losses respectively to obtain the semantic compensation values of the two text segments; Determine the association compensation coefficient of the implicit semantics between the two text segments according to the two semantic compensation values; Determine the dependency relationship of the implicit semantics between the two text segments through the association degree of the implicit semantics between the two text segments and the association compensation coefficient, and further determine the dependency relationship of the implicit semantics between every two text segments.
8. A language translation data processing system based on a large model, characterized in that, Including: A receiving module for receiving the language translation data to be processed; A processing module for extracting semantic features of the words in the language translation data based on a pre-trained large language model to obtain the semantic features of each word, and then encoding and segmenting the language translation data according to all the semantic features to obtain multiple text segments, and determining the segmentation losses of each text segment; The processing module is further configured to perform dependency syntactic analysis on each text segment to obtain the dependency relationships between the words in the text segment, and then determine the dependency tree structure of each text segment, and determine the dependency relationship of the explicit semantics between every two text segments according to the semantic similarity between every two text segments and the dependency tree structure of the corresponding text segment; The processing module is further configured to implicitly associate the segment semantics between each text segment according to the full text information of the language translation data, obtain the association degree of the implicit semantics between each text segment, and then perform a dependency analysis on the implicit semantic relationship between every two text segments based on the association degrees of all implicit semantics and the segmentation loss of each text segment, so as to obtain the dependency relationship of the implicit semantics between every two text segments; The execution module is configured to construct the paragraph label of the language translation data through the dependency relationship of the explicit semantics and the implicit semantics between every two text segments, and perform the translation processing of the data to be translated based on the paragraph label.
9. A computer device, the computer device includes a memory and a processor, the memory stores code, characterized in that, The processor is configured to obtain the code and execute the large model-based language translation data processing method according to any one of claims 1 to 7.
10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the large model-based language translation data processing method according to any one of claims 1 to 7.
Citation Information
Cited By
Method and system for reversely generating design book through codes driven by large language model
CN120723297A
Teaching resource recommendation method and system based on natural language processing
CN121724811A