Cross-language text fusion intelligent alignment method and system
By combining deep learning models with XML tags, the problems of semantic ambiguity and format differences in mixed Chinese and English texts are solved, high-precision cross-language text alignment and multimodal fusion are achieved, and the accuracy and efficiency of text alignment are improved.
Patent Information
- Application Number
- CN202510908768.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-02
- Publication Date
- 2025-10-10
AI Technical Summary
Existing technologies have difficulty effectively solving the problems of semantic ambiguity, format differences, and multimodal information fusion when processing mixed Chinese and English texts, especially lacking targeted processing capabilities in cross-language text alignment.
It adopts a multilingual text feature extraction model based on deep learning, combined with XML tags and attention mechanism, strengthens cross-language semantic associations through a multi-head attention mechanism, builds a hierarchical alignment model, realizes character-level and paragraph-level format coordination, and monitors and corrects errors in real time.
It improves the accuracy and efficiency of cross-language text alignment, solves the problem of semantic ambiguity, achieves high-precision multimodal fusion, and ensures the standardization and consistency of text format.
Smart Images

Figure CN120764481A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of battery technology, and in particular to a cross-language text fusion intelligent alignment method and system. Background Art
[0002] In the era of the deep integration of globalization and digitalization, multilingual information exchange scenarios are becoming increasingly complex, and the use of mixed Chinese and English texts in academic papers, business documents, and translation materials is becoming increasingly common. As a key component in the efficient processing of multilingual information, cross-language text fusion and alignment technology aims to address the semantic, grammatical, and formatting issues faced by texts in different languages, ensuring consistency and readability when mixed Chinese and English texts are used.
[0003] CN119150225A discloses a method, device, and electronic device for aligning and fusing multimodal entity features. The method comprises: inputting a knowledge graph to be aligned into a multimodal embedding module for feature encoding to obtain a multimodal embedding, wherein the knowledge graph to be aligned includes multimodal data, and the multimodal embedding includes at least two of structural embedding, text embedding, and visual embedding; inputting the multimodal embedding into a cross-modal attention fusion module for feature fusion and extraction to obtain a cross-modal attention fusion embedding; and inputting the multimodal embedding into a multimodal adaptive feature fusion module for feature fusion to obtain an adaptive feature fusion embedding; and aligning and fusing entities in the knowledge graph to be aligned based on a multimodal early fusion embedding, the cross-modal attention fusion embedding, and the adaptive feature fusion embedding, wherein the multimodal early fusion embedding is obtained by cascading the multimodal embeddings in the embedding layer.
[0004] CN116257643B discloses a cross-language entity alignment method, device, equipment and readable storage medium. The cross-language entity alignment method includes the following steps: obtaining a cross-language knowledge graph to be fused, and obtaining a first alignment seed corresponding to the cross-language knowledge graph; translating the text in the cross-language knowledge graph into a unified language text, and performing preliminary alignment on the entity vectors corresponding to the unified language text to obtain a preliminary alignment result; determining the similarity between the entity vectors, and using the unified language text corresponding to the similarity greater than or equal to the first preset similarity as the second alignment seed; adjusting the entity vectors in the preliminary alignment result in batches according to text similarity and / or semantic similarity according to the first alignment seed and the second alignment seed; the step of adjusting the entity vectors in the preliminary alignment result in batches according to text similarity and / or semantic similarity according to the first alignment seed and the second alignment seed includes: using the corresponding vectors of the first alignment seed and the second alignment seed as label vectors; according to a preset loss function and the label vector, in an iterative calculation manner, According to text similarity and / or semantic similarity, the entity vectors in the preliminary alignment result are adjusted in batches until the loss value corresponding to the preset loss function reaches a preset threshold, wherein, in the process of adjusting the entity vectors, the preset loss function is optimized by a preset gradient descent method; the preset loss function includes a preset text loss function and a preset semantic loss function, and the step of adjusting the entity vectors in the preliminary alignment result in batches according to text similarity and / or semantic similarity in an iterative calculation manner based on the preset loss function and the label vector includes: determining similar entity vectors in the preliminary alignment result whose similarity is less than the first preset similarity and greater than or equal to the second preset similarity; adjusting the similar entity vectors according to text similarity and semantic similarity according to the label vector in an alternating iterative calculation manner using the preset text loss function and the preset semantic loss function; aligning the entity vector with the highest similarity among the adjusted entity vectors to obtain a target alignment result.
[0005] CN115760023B discloses a digital collaborative design environment intelligent alignment system and method, the digital collaborative design environment intelligent alignment system includes: a network service end unit, which is arranged in a server and is connected to a health monitoring unit and a design environment synchronization start unit through a network; is used to manage various standard design environments in a synchronization folder in the server, and generates a document version stamp file of all design environments at the time of this release under a cloud design environment set unit each time a service is started or restarted, and the document version stamp file is used for difference comparison calculation in the design environment synchronization start unit; a health monitoring unit, which is arranged in a server and is connected to the network service end unit through a network; is used to monitor the program in the server, and automatically repair the program if an interruption or error occurs in the program; a cloud design environment set unit, which is connected to the network service end unit through a network, and is used to obtain and save the document version stamp file of all design environments released by the network service end unit; the design environment synchronization start unit, which is arranged in a client, includes a storage module, and is used to receive data from the network service end unit in real time, and receive the received data. The collected data is saved in the storage module for comparative calculation of the design environment differences between the server and the client; according to the integrated design environment required by the design software unit, a synchronization update request is sent to the network server unit in real time to obtain the design environment data in the cloud design environment set unit for synchronously updating the design environment in the design software unit; the design software unit is arranged in the client, including a design environment module, an environment update module, and a file push module, which is connected to the network of the design environment synchronization start unit for generating and updating specific design data, requesting and obtaining the design environment synchronization start unit for updating data of the design environment module; the environment update module modifies and updates the data in the design environment module; the file push module modifies and updates the files and folders of the design software unit according to the data of the design environment synchronization start unit; the cloud design environment set unit includes multiple design environment document version stamp files arranged according to generation time; the design software unit includes multiple modeling and design software for jointly participating in collaborative design based on the Internet.
[0006] Traditional text alignment methods rely primarily on rule engines or statistical models. Rule-based methods use preset Chinese and English typesetting rules, such as character spacing and paragraph indentation standards, to achieve alignment. However, when faced with complex language structures and diverse document formats, these rules lack universality and scalability, making them difficult to adapt to dynamically changing text scenarios. Summary of the Invention
[0007] Long-term practice has shown that statistical methods, which utilize large-scale parallel corpora to learn language alignment patterns, have achieved some success in sentence- and paragraph-level alignment, but lack the ability to specifically address issues such as semantic ambiguity and formatting differences in mixed Chinese and English texts. Furthermore, existing technologies often focus on text alignment in a single language or translation scenario, with limited capabilities for character- and paragraph-level format coordination in mixed Chinese and English typesetting, as well as cross-modal integration, such as the fusion of text with images, tables, and other information.
[0008] In view of this, the present invention aims to propose a cross-language text fusion intelligent alignment method, comprising:
[0009] Step S1, obtaining text data including at least two encoding types, preprocessing the text data, and identifying XML tags in the text data;
[0010] Step S2: constructing a multilingual text feature extraction model based on deep learning, wherein the multilingual text feature extraction model includes a multilingual pre-trained language model, and encoding the extracted text information using the multilingual pre-trained language model to obtain semantic feature vectors of each language in the text information;
[0011] Step S3: using XML to annotate the semantic feature vector, and distinguishing the XML tags from the language in the text information to generate a text sequence with tag identification;
[0012] Step S4: constructing an alignment model that integrates language semantics and format information. The alignment model takes the semantic feature vector and the text sequence with label identification as input, calculates the semantic association between different languages through an attention mechanism, and combines the format information represented by the XML tags to generate a text alignment method;
[0013] Step S5: Use a text alignment method to align the cross-language text in the target document, and monitor the integration of text information and XML tags in real time; if no garbled errors or XML confusion are detected, the aligned target document is output.
[0014] Preferably, in step S5, if a garbled text error or XML confusion is detected, an error prompt is triggered, and the erroneous text or the confused XML tag is corrected according to the error prompt.
[0015] Preferably, in step S1, during the pre-processing of the text data, noise characters and special symbols in the text data are removed, and word segmentation or sentence segmentation is performed on the text data.
[0016] Preferably, in step S1, regular expression matching or a deep learning-based sequence labeling model is used to identify XML tags.
[0017] Preferably, a sequence labeling model based on deep learning assigns a label to each character in the text, and labels the text at the character level so that each character corresponds to a label;
[0018] Map characters to numeric indices and build a character vocabulary; map labels to numeric indices and build a label vocabulary; encode the character vocabulary and label vocabulary in batches and input them into the trained Transformer model; obtain the global semantic and contextual information of the characters through the self-attention mechanism and output feature vectors;
[0019] The Transformer model outputs the predicted label index for each character, converts the index into the actual label according to the label vocabulary, and parses the XML tags.
[0020] Preferably, in step S3, a first label is used to record the language type corresponding to the semantic feature vector, a second label is used to mark the semantic category, and a third label is used to record the association between the semantic feature vector and the original text information.
[0021] Preferably, in step S2, a trained mBERT-based text classification model is used.
[0022] Preferably, in step S4, the alignment model for integrating language semantics and format information includes:
[0023] The semantic processing layer receives the semantic feature vectors input by the multilingual text feature extraction model, automatically assigns weights through the attention mechanism, and calculates semantic association information;
[0024] The format parsing layer is used to parse the XML tags in the text sequence with tag identifiers and extract the format information;
[0025] The alignment strategy generation layer is used to integrate semantic association information and format information to obtain the text alignment method.
[0026] The present invention also provides a system based on the above-mentioned cross-language text fusion intelligent alignment method, the system comprising:
[0027] an acquisition unit, configured to acquire text data including at least two encoding types, preprocess the text data, and identify XML tags in the text data;
[0028] A modeling unit is configured to construct a multilingual text feature extraction model based on deep learning, wherein the multilingual text feature extraction model includes a multilingual pre-trained language model, and uses the multilingual pre-trained language model to encode the extracted text information to obtain semantic feature vectors of each language in the text information;
[0029] A tagging unit is used to tag the semantic feature vector using XML, distinguish the XML tags from the language characters in the text information, and generate a text sequence with tag identification;
[0030] A strategy unit is used to build an alignment model that integrates language semantics and format information. The alignment model takes the semantic feature vector and the text sequence with label identification as input, calculates the semantic association between different languages through the attention mechanism, and combines the format information represented by the XML tags to generate a text alignment method;
[0031] The alignment unit is used to align the cross-language text in the target document using a text alignment method, and monitor the integration of text information and XML tags in real time; if no garbled errors or XML confusion are detected, the aligned target document is output.
[0032] The present invention provides a machine-readable storage medium having instructions stored thereon, the instructions being used to enable a machine to execute the above-mentioned cross-language text fusion intelligent alignment method of the present application.
[0033] The cross-language text fusion intelligent alignment method disclosed in the present invention, through steps S1-S5, first obtains text data containing at least two encoding types, and identifies XML tags therein after preprocessing; then builds a multilingual text feature extraction model based on deep learning, encodes the text using a multilingual pre-trained language model, and obtains a semantic feature vector; then annotates the semantic feature vector with XML, distinguishes XML tags from language characters, and generates a labeled text sequence; then builds an alignment model that integrates semantics and format information, takes the semantic feature vector and the labeled text sequence as input, and generates an alignment method by means of an attention mechanism and XML format information; finally, uses the method to align cross-language texts, monitors the fusion status in real time, and outputs the aligned document when there are no garbled errors or XML confusion. The present invention also discloses a system for executing the above method. In practical applications, the method and system accurately parse text semantics and format information by preprocessing and identifying XML tags of texts of multiple encoding types, combining deep semantic feature extraction and XML annotation of multi-language pre-training models, and constructing a hierarchical alignment model with the Transformer architecture as the core. It strengthens cross-language semantic associations through a multi-head attention mechanism, and realizes character-level and paragraph-level format coordination based on XML tag weight distribution and conditional constraints. In multi-language mixed typesetting scenarios, the accuracy of format alignment is significantly improved, and the problem of semantic ambiguity is effectively solved. In terms of cross-modal processing, the model's ability to parse XML structured information can be extended to the layout relationship processing of text, images, and tables, so that the comprehensive alignment efficiency in multi-modal fusion scenarios is significantly improved. A real-time monitoring and error correction mechanism is further adopted to ensure the accuracy of the fusion of text and XML tags, avoid garbled code and tag confusion problems, and realize high-precision cross-language text alignment from semantics to format, from single modality to multi-modality.
[0034] Other features and advantages of the present invention will be described in detail in the following detailed description. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] The accompanying drawings, which constitute a part of the present invention, are used to provide a further understanding of the present invention. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute improper limitations on the present invention.
[0036] In the attached figure:
[0037] Fig. 1 A schematic flow chart of a cross-language text fusion intelligent alignment method according to an embodiment of the present invention;
[0038] Fig. 2 The figure is a business process diagram of a cross-language text fusion intelligent alignment method according to an embodiment of the present invention. DETAILED DESCRIPTION
[0039] The specific embodiments of the present invention will be described in detail with reference to the accompanying drawings. It should be noted that the embodiments described herein are only used to explain and illustrate the technical solutions of the present invention and do not constitute any limitation on the scope of protection of the present invention.
[0040] To facilitate those skilled in the art to deeply understand the technical solution of the present invention, the following will provide a comprehensive and clear description of the technical solution in conjunction with the drawings in the embodiments. It should be noted that the described embodiments are only partial examples of the technical solution of the present invention and are not the entire content. Based on the embodiments disclosed in the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work shall fall within the scope of protection of the present invention.
[0041] In addition, in the specification, claims and drawings of the present invention, terms such as "first", "second", "third", etc. are mainly used to distinguish similar technical features, rather than to limit a specific order or execution sequence. In appropriate scenarios, the data referred to by these terms can be interchangeable to meet the different application scenario requirements of the embodiments of the present invention. At the same time, "including", "having" and their derivative words are intended to cover non-exclusive combinations of technical elements. For example, a method, system, product or device that includes multiple steps or units is not limited to the steps or units that are explicitly listed, but also includes other steps or units that are not explicitly listed but are inherent parts of the technical solution.
[0042] Traditional text alignment technology systems mainly rely on two major technical paths: rule engines and statistical models. Among them, rule-based alignment methods build a standardized alignment framework by pre-setting Chinese and English typesetting specifications such as character spacing and paragraph indentation. However, faced with complex and changing language structures, such as nested clauses, dialect variants, and diverse document formats, such as special typesetting and mixed encoding, its rule system is difficult to cover all scenarios, exposing the dual bottlenecks of insufficient universality and limited scalability, and unable to effectively respond to dynamically changing text processing needs. Statistical-based alignment solutions rely on large-scale parallel corpora and learn language alignment patterns in a data-driven manner. However, when processing mixed Chinese and English texts, there is a lack of effective strategies for resolving semantic ambiguity and adapting to format differences. It is difficult to accurately capture the subtle semantic differences in cross-language expressions, and it is also unable to take into account the differences in typesetting habits of different languages. In addition, existing technologies are mostly limited to text alignment research in single language scenarios or traditional translation fields. In multilingual mixed typesetting scenarios, there are still significant technical shortcomings in precise character-level alignment control, paragraph-level format collaborative optimization, and deep fusion processing of text with multimodal information such as images and tables. The present invention provides a cross-language text fusion intelligent alignment method, such as Figs. 1-2 As shown, the cross-language text fusion intelligent alignment method includes:
[0043] Step S1, obtaining text data including at least two encoding types, preprocessing the text data, and identifying XML tags in the text data;
[0044] Step S2: constructing a multilingual text feature extraction model based on deep learning, wherein the multilingual text feature extraction model includes a multilingual pre-trained language model, and encoding the extracted text information using the multilingual pre-trained language model to obtain semantic feature vectors of each language in the text information;
[0045] Step S3: using XML to annotate the semantic feature vector, and distinguishing the XML tags from the language in the text information to generate a text sequence with tag identification;
[0046] Step S4: constructing an alignment model that integrates language semantics and format information. The alignment model takes the semantic feature vector and the text sequence with label identification as input, calculates the semantic association between different languages through an attention mechanism, and combines the format information represented by the XML tags to generate a text alignment method;
[0047] Step S5: Use a text alignment method to align the cross-language text in the target document, and monitor the integration of text information and XML tags in real time; if no garbled errors or XML confusion are detected, the aligned target document is output.
[0048] The cross-language text fusion intelligent alignment method accurately analyzes text semantics and format information through pre-processing and XML tag recognition of multi-coding type texts, combined with deep semantic feature extraction and XML annotation of multi-language pre-training models, and constructs a hierarchical alignment model with Transformer architecture as the core. It strengthens cross-language semantic associations through a multi-head attention mechanism, and realizes character-level and paragraph-level format coordination based on XML tag weight distribution and conditional constraints. In multi-language mixed typesetting scenarios, the accuracy of format alignment is significantly improved, and the semantic ambiguity problem is effectively solved. In terms of cross-modal processing, the model's ability to parse XML structured information can be extended to the layout relationship processing of text, images, and tables, which significantly improves the comprehensive alignment efficiency in multi-modal fusion scenarios. A real-time monitoring and error correction mechanism is further adopted to ensure the accuracy of text and XML tag fusion, avoid garbled code and tag confusion problems, and realize high-precision cross-language text alignment from semantics to format, from single modality to multi-modality.
[0049] In order to quickly obtain abnormal situations in the generated text and provide real-time feedback information, targeted corrections can be made based on error prompts. In a more preferred embodiment of the present invention, in step S5, if a garbled error or XML confusion is detected, an error prompt is triggered, and the erroneous text or confused XML tags are corrected according to the error prompt. The introduced error detection and correction mechanism can quickly identify abnormal situations such as garbled errors and XML tag confusion by monitoring the fusion state of text information and XML tags in real time, and trigger accurate error prompts in time to clearly mark the error location and type. The active error feedback mechanism not only optimizes the XML annotation process, but also avoids alignment failures or document format confusion problems caused by text anomalies, ensures the integrity of text semantics and format standardization, and improves the quality of cross-language documents.
[0050] In order to effectively remove noise characters and special symbols, eliminate meaningless or interfering content in the text, avoid the negative impact of garbled characters, format control characters, etc. on subsequent processing, make the text data purer and more standardized, and facilitate subsequent operations such as XML tag recognition and semantic feature extraction. In a more preferred embodiment of the present invention, in step S1, the text data is pre-processed to remove noise characters and special symbols in the text data, and the text data is segmented or sentence-processed. Noise characters and special symbols are accurately identified and deleted according to specific rules. For common garbled characters, format control characters, line feed characters\n, tab characters\t, carriage return characters\r, as well as special symbols such as mathematical symbols, punctuation variants, and technical symbols, corresponding regular expression patterns are used for matching. Segmentation or sentence processing can divide continuous text into units that conform to the logic of language expression, making it easier for subsequent multilingual text feature extraction models based on deep learning to more accurately capture text semantic information. For example, a dictionary-based word segmentation method is to build a dictionary containing a large number of words, scan the text, match the character sequence in the text with the words in the dictionary, and if the match is successful, it is divided into one word. Word segmentation methods based on statistical models use large-scale corpora to train language models, learn the probability of word occurrence and contextual relationships, and determine word segmentation boundaries by calculating the probability of character sequences. For example, hidden Markov models (HMMs) and conditional random fields (CRFs) are commonly used for word segmentation. Common punctuation marks in Chinese and English, such as the period (.), question mark (?), and exclamation point (!), often mark the end of a sentence. By identifying these punctuation marks, text can be segmented into different sentences.
[0051] Regular expression matching relies on concise and efficient pattern matching rules to quickly locate XML tag structures. For text data with standardized formats and fixed tag types, it can achieve accurate extraction of tags with extremely low computing resource consumption, greatly improving processing speed, and is suitable for scenarios with high processing efficiency requirements. In order to ensure the reliability of XML tag recognition, and then guarantee subsequent text semantic feature extraction, format information parsing and other operations, in the more preferred case of the present invention, in step S1, regular expression matching or a sequence annotation model based on deep learning is used to identify XML tags. For texts with irregular or missing punctuation, a language model based on deep learning, such as the mBERT model (Multi-lingualBERT), can be used. It is a pre-trained language model that can handle multiple languages to predict sentence boundaries. By training the model on large-scale annotated data, it learns the semantic and grammatical features of sentences, thereby accurately judging the end position of sentences. In application, the pre-trained model is called with the help of the transformers library to perform sentence prediction. The sequence labeling model based on deep learning has powerful semantic understanding and pattern learning capabilities. Through training on a large amount of labeled data, it can effectively identify XML tags with variable formats and complex nesting levels in complex texts. Even in the face of noise interference or non-standard label writing, it can accurately distinguish tags from ordinary texts with its generalization ability, especially when processing low-quality and diverse text data. The XML start tag begins with < and ends with >, with the tag name and possible attributes in the middle. Use the regular expression / <([^>]+)> / , where / is the delimiter of the regular expression, < and > are used to match the start and end boundaries of the tag, and ([^>]+) means matching any non-> characters between the two >, that is, the tag name and attribute part. Brackets are used to capture the matched content for subsequent extraction. For example, for <div class=″container">, which accurately matches and extracts the div class="container". To identify complete XML elements, use the regular expression / <([^>]+)>[^<]*<\八1> / s, where \1 is a backreference to ensure the start and end tags have the same name, [^<]* matches the text between the tags, and the s flag allows .* to match any character, including line breaks, to handle XML elements that span multiple lines.
[0052] In the process of identifying XML tags using a sequence tagging model based on deep learning, each XML text is annotated at the character level, and tag categories are defined, such as "B-StartTag" (starting character of the start tag), "I-StartTag" (character inside the start tag), "B-EndTag" (starting character of the end tag), "I-EndTag" (character inside the end tag), and "O" (non-tag character). For example, for , the character "<" is labeled as "B-StartTag", "d", "i", "v" are labeled as "I-StartTag", and ">" is labeled as "I-StartTag". The model is selected as a bidirectional long short-term memory network combined with a conditional random field model BiLSTM-CRF and a model based on a Transformer such as BERT-CRF. In BiLSTM-CRF, BiLSTM learns the context information of the text from the front and back directions, and CRF can optimize the output of BiLSTM by utilizing the dependency relationship between labels to obtain a more reasonable label sequence. The model based on the Transformer can better capture the global semantic information of the text through a powerful self-attention mechanism. In a more preferred case, the hyperparameters are set, the learning rate is set to 0.001-0.01, the batch size is set to 16, 32, etc. The loss function and the optimizer are defined. For the BiLSTM-CRF model, the negative log-likelihood loss function of CRF is used to calculate the difference between the predicted label sequence and the true label sequence, and the optimizer is usually Adam. In the training loop, the training data is input into the model in batches. Calculate the loss, update the model parameters through back propagation, and regularly evaluate the performance of the model on the validation set, and adjust the hyperparameters according to the evaluation results.
[0053] In order to convert text information into structured data suitable for model processing through character-level fine labeling and digital index coding. The process of constructing a character and label vocabulary enables the model to learn a variety of text content and label systems, and after training with a large amount of labeled data, it can adapt to XML texts of different sources and different encoding formats. In a more preferred case of the present application, the sequence labeling model based on deep learning assigns a label to each character in the text, and labels the text at the character level, so that each character corresponds to a label;
[0054] Map characters to numerical indexes and build a character vocabulary; map tags to numerical indexes and build a tag vocabulary; encode the character vocabulary and tag vocabulary in batches and input them into the trained Transformer model; obtain the global semantics and contextual information of the characters through the self-attention mechanism and output the feature vector. In the process of building the character vocabulary, assign a unique numerical index to each character, starting from O and increasing in sequence to form a mapping relationship between characters and numerical indexes. For example, the character < corresponds to index O, a corresponds to index 1, and so on. Summarize all label categories that appear in the labeled data. Similarly, assign a unique numerical index to each tag and build a mapping from label to numerical index. For example, "B-StartTag" corresponds to index O, "I-StartTag" corresponds to index 1, and so on, which are stored as tag_vocab = {'B-StartTag': O, 'I-StartTag': 1,...}. Character encoding is to map each character in the labeled data to the corresponding numerical index according to the character vocabulary and convert the text into a numerical sequence. For example, the text The encoding is [O, 1, 2, 3]. If we assume that the index is O, the index of d is 1, the index of i is 2, and the index of v is 3. Label encoding is to convert the label corresponding to each character into a digital index according to the label vocabulary to obtain a label digital sequence. For example, the character sequence If the tags are "B-StartTag", "I-StartTag", "I-StartTag", "I-StartTag", then they are encoded as [0, 1, 1, 1]. The encoded character sequences and tag sequences are grouped according to the set batch size, preferably 32 or 64. For sequences with insufficient length, fill them with a padding value (such as 0) to ensure that the length of the sequences in the same batch of the input model is consistent.
[0055] The Transformer model outputs a predicted label index for each character. This index is converted to an actual label according to a label vocabulary, and the XML tags are parsed. Pre-trained models such as BERT and RoBERTa are used as the foundational architecture. Fully connected layers and classifiers are added on top to output a probability distribution of labels for each character. The model is quickly built using the HuggingFace Transformers library. During the training loop, batches of data are fed into the model, the loss is calculated, and the model parameters are updated using the backpropagation algorithm. The XML text to be recognized is encoded using the same character encoding as the training data and fed into the trained Transformer model in batches. The model outputs a probability distribution of labels for each character, and the label index with the highest probability is taken as the prediction. The predicted label indexes are converted to actual label names according to the label vocabulary, resulting in a sequence of predicted labels for each character. The predicted label sequence is then iterated over, extracting the complete XML tag structure based on the label type and name, and parsing it into a complete XML element.
[0056] The self-attention mechanism gives the model the ability to dynamically focus on key semantic information and can flexibly adjust the attention paid to each character according to the text content. When identifying XML tags, it can not only make judgments based on the tag grammatical structure, but also make comprehensive decisions based on the text semantic information. The recognition effect is particularly outstanding for tags with close semantic associations, such as custom tags in specific semantic scenarios, significantly improving the model's understanding and parsing capabilities of XML tags in complex semantic environments.
[0057] To clearly distinguish semantic feature vectors from different languages and avoid information confusion during multilingual processing, accurate language attribute identifiers are provided for subsequent cross-language text analysis, machine translation, and other tasks. This allows the model to perform differentiated processing based on different language characteristics, thereby improving the compatibility and accuracy of multilingual processing. In a more preferred embodiment of the present invention, in step S3, a first label is used to record the language type corresponding to the semantic feature vector, a second label is used to annotate the semantic category, and a third label is used to record the association between the semantic feature vector and the original text information. Using the first label to record the language type clearly distinguishes semantic feature vectors from different languages and avoids information confusion during multilingual processing. The second label annotates the semantic category, systematically categorizing semantic features. Whether they are entities, events, or abstract concepts, they can be quickly located and retrieved, facilitating the efficient extraction of specific semantic content in applications such as information retrieval and knowledge graph construction. The third label records the association with the original text, establishing a mapping between semantic features and text instances. This allows semantic analysis results to be traced back to specific contexts, facilitating verification of semantic extraction accuracy, and enabling more reasonable semantic matching and content generation based on the structure and content of the original text. The first label is used to record the language type corresponding to the semantic feature vector. Establish a standardized language code table. For example, refer to the ISO639-1 standard where "zh" represents Chinese and "en" represents English. According to the needs of semantic analysis, build a semantic category system to determine the content of the second label. The second label uses a hierarchical classification structure to mark semantic categories, which are first divided into entity categories, such as people, places, institutions, etc. Distinguish between major categories such as entity classes, event classes, and attribute classes, and then further subdivide them. The third label records the semantic feature vector, which is a combined label containing information such as text position index and context fragments. For example, use the format of "[starting character index: ending character index]-original text fragment".
[0058] In order to label the semantic feature vectors and provide support for subsequent text alignment, semantic analysis, etc. In a more preferred embodiment of the present invention, in step S2, a trained mBERT-based text classification model is used.
[0059] In order to better integrate the alignment model of language semantics and format information, in a more preferred embodiment of the present invention, in step S4, the alignment model integrating language semantics and format information includes a semantic processing layer, a format parsing layer, and an alignment strategy generation layer.
[0060] To leverage the attention mechanism to dynamically assign weights to the semantic feature vectors output by the multilingual text feature extraction model, and accurately capture the semantic associations between cross-language texts, the semantic processing layer receives the semantic feature vectors from the multilingual text feature extraction model and automatically assigns weights through the attention mechanism to calculate semantic association information.
[0061] To deeply parse XML tags in tagged text sequences, quickly extract formatting information such as character spacing and paragraph indentation, and effectively resolve formatting confusion caused by differences in typesetting rules between different languages, the format parsing layer parses XML tags in tagged text sequences and extracts formatting information.
[0062] In order to organically integrate semantic association information with format information, it not only ensures the coherence of the text semantic logic, but also achieves precise format coordination at the character level and paragraph level. The alignment strategy generation layer integrates semantic association information and format information to obtain a text alignment method. More preferably, the domain expert knowledge and language typesetting rules are combined to build a priori knowledge rule base. In the alignment strategy generation process, these rules are incorporated into the algorithm as constraints to guide the model to generate text alignment methods that conform to language habits and professional standards. For example, in the alignment of scientific and technological documents, paragraph indentation, title format, etc. are strictly constrained according to the typesetting rules of academic papers. More preferably, after each text alignment method is generated, the alignment strategy generation layer adjusts the internal parameters and fusion strategy based on the reward feedback obtained, and continuously optimizes the alignment scheme. After multiple rounds of iterative learning, the model can automatically explore the optimal way to integrate semantic and format information under different conditions.
[0063] The present invention also provides a system based on the above-mentioned cross-language text fusion intelligent alignment method, the system comprising:
[0064] an acquisition unit, configured to acquire text data including at least two encoding types, preprocess the text data, and identify XML tags in the text data;
[0065] A modeling unit is configured to construct a multilingual text feature extraction model based on deep learning, wherein the multilingual text feature extraction model includes a multilingual pre-trained language model, and uses the multilingual pre-trained language model to encode the extracted text information to obtain semantic feature vectors of each language in the text information;
[0066] A tagging unit is used to tag the semantic feature vector using XML, distinguish the XML tags from the language characters in the text information, and generate a text sequence with tag identification;
[0067] A strategy unit is used to build an alignment model that integrates language semantics and format information. The alignment model takes the semantic feature vector and the text sequence with label identification as input, calculates the semantic association between different languages through the attention mechanism, and combines the format information represented by the XML tags to generate a text alignment method;
[0068] The alignment unit is used to align the cross-language text in the target document using a text alignment method, and monitor the integration of text information and XML tags in real time; if no garbled errors or XML confusion are detected, the aligned target document is output.
[0069] The system efficiently removes noise data through refined pre-processing of texts of multiple encoding types, and uses regular expression matching or deep learning sequence annotation models to accurately identify XML tags. The powerful semantic representation capabilities of the multi-language pre-trained model allow for deep semantic feature extraction of the text, and through XML structured annotation, it achieves dual deconstruction of text semantics and format information, effectively eliminating semantic ambiguity. Based on the weight distribution and conditional constraint strategy of XML tags, precise format coordination at the character and paragraph levels is achieved, and the accuracy of format alignment is significantly improved in multi-language mixed typesetting scenarios. It fully utilizes the parsing advantages of XML structured information, conducts real-time monitoring and intelligent error correction, dynamically verifies the fusion status of text and XML tags, and promptly discovers and avoids problems such as garbled characters and tag confusion, achieving high-precision cross-language text alignment from semantic understanding to format presentation.
[0070] The present invention provides an electronic device, at least one processor; and
[0071] a memory communicatively connected to the at least one processor; wherein,
[0072] The memory stores instructions that can be executed by the at least one processor. The instructions are executed by the at least one processor to enable the at least one processor to perform the above-mentioned cross-language text fusion intelligent alignment method.
[0073] The present invention provides a machine-readable storage medium having instructions stored thereon, the instructions being used to enable a machine to execute the cross-language text fusion intelligent alignment method as described above in the present application.
[0074] It should be noted that for the aforementioned method embodiments, for simplicity of description, they are all expressed as a series of action combinations. However, those skilled in the art should be aware that the present invention is not limited by the order of the actions described, because according to the present invention, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in this specification are all preferred embodiments, and the actions and modules involved are not necessarily required by the present invention.
[0075] In the above embodiments, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0076] In addition, the functional units in the various embodiments of the present invention may be integrated into a single processing unit, each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.
[0077] The foregoing description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Those skilled in the art will readily appreciate that various modifications and variations of the present invention are possible. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present invention are intended to be within the scope of protection of the present invention.
Claims
1. A cross-language text fusion intelligent alignment method, characterized by: The cross-language text fusion intelligent alignment method includes: Step S1, obtaining text data including at least two encoding types, preprocessing the text data, and identifying XML tags in the text data; Step S2: constructing a multilingual text feature extraction model based on deep learning, wherein the multilingual text feature extraction model includes a multilingual pre-trained language model, and encoding the extracted text information using the multilingual pre-trained language model to obtain semantic feature vectors of each language in the text information; Step S3: using XML to annotate the semantic feature vector, and distinguishing the XML tags from the language in the text information to generate a text sequence with tag identification; Step S4: constructing an alignment model that integrates language semantics and format information. The alignment model takes the semantic feature vector and the text sequence with label identification as input, calculates the semantic association between different languages and texts through an attention mechanism, and combines the format information represented by the XML tags to generate a text alignment method. Step S5: Use a text alignment method to align the cross-language text in the target document, and monitor the fusion of text information and XML tags in real time; if no garbled code errors or XML confusion are detected, output the aligned target document.
2. The cross-language text fusion intelligent alignment method according to claim 1, characterized in that: In step S5, if garbled text errors or XML confusion are detected, an error prompt is triggered, and the erroneous text or confused XML tags are corrected according to the error prompt.
3. The cross-language text fusion intelligent alignment method according to claim 1, characterized in that: In step S1, during the preprocessing of the text data, noise characters and special symbols in the text data are removed, and the text data is segmented into words or sentences.
4. The cross-language text fusion intelligent alignment method according to claim 1, characterized in that: In step S1, regular expression matching or a deep learning-based sequence labeling model is used to identify XML tags.
5. The cross-language text fusion intelligent alignment method according to claim 4 is characterized in that: The deep learning-based sequence labeling model assigns a label to each character in the text and labels the text at the character level so that each character corresponds to a label; Map characters to numeric indices and build a character vocabulary; map labels to numeric indices and build a label vocabulary; encode the character vocabulary and label vocabulary in batches and input them into the trained Transformer model; obtain the global semantic and contextual information of the characters through the self-attention mechanism and output feature vectors; The Transformer model outputs the predicted label index for each character, converts the index into the actual label according to the label vocabulary, and parses the XML tags.
6. The cross-language text fusion intelligent alignment method according to claim 1, characterized in that: In step S3, a first label is used to record the language type corresponding to the semantic feature vector, a second label is used to mark the semantic category, and a third label is used to record the association between the semantic feature vector and the original text information.
7. The cross-language text fusion intelligent alignment method according to any one of claims 1 to 6, characterized in that: In step S2, the trained mBERT-based text classification model is used.
8. The cross-language text fusion intelligent alignment method according to any one of claims 1 to 6, characterized in that: In step S4, the alignment model for integrating language semantics and format information includes: The semantic processing layer receives the semantic feature vectors input by the multilingual text feature extraction model, automatically assigns weights through the attention mechanism, and calculates semantic association information; The format parsing layer is used to parse the XML tags in the text sequence with tag identifiers and extract the format information; The alignment strategy generation layer is used to integrate semantic association information and format information to obtain the text alignment method.
9. A system based on the cross-language text fusion intelligent alignment method according to any one of claims 1 to 8, characterized in that: The system comprises, an acquisition unit, configured to acquire text data including at least two encoding types, preprocess the text data, and identify XML tags in the text data; A modeling unit is configured to construct a multilingual text feature extraction model based on deep learning, wherein the multilingual text feature extraction model includes a multilingual pre-trained language model, and uses the multilingual pre-trained language model to encode the extracted text information to obtain semantic feature vectors of each language in the text information; A tagging unit is used to tag the semantic feature vector using XML, distinguish the XML tags from the language characters in the text information, and generate a text sequence with tag identification; A strategy unit is used to construct an alignment model that integrates language semantics and format information. The alignment model takes the semantic feature vector and the text sequence with label identification as input, calculates the semantic association between different languages through the attention mechanism, and combines the format information represented by the XML tag to generate a text alignment method; The alignment unit is used to align the cross-language text in the target document using a text alignment method, and monitor the integration of text information and XML tags in real time; if no garbled errors or XML confusion are detected, the aligned target document is output.
10. A machine-readable storage medium, characterized in that The machine-readable storage medium stores instructions for enabling a machine to execute the cross-language text fusion intelligent alignment method according to any one of claims 1 to 8 of the present invention.
Citation Information
Patent Citations
Multi-modal entity feature alignment fusion method and device and electronic equipment
CN119150225A
Cited By
Multi-language user matching recommendation method and system based on multi-dimensional label fusion
CN121327249A
A multi-language user matching recommendation method and system based on multi-dimensional label fusion
CN121327249B
Language psychological disorder patient rehabilitation state evaluation system based on multi-agent cooperation
CN121331389A
Joint representation fusion method, system, medium, and product for alignment of dual-persona theory
CN122471373A