Website cluster globalization intelligent language translation method
Patent Information
- Application Number
- CN202610899646.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-22
- Publication Date
- 2026-09-29
AI Technical Summary
站点部署阶段无法自动适配不同目标站点的数据导入配置,依赖人工逐一配置字段映射关系,操作繁琐且易出现人为失误,制约站群全球化部署的整体进度,因此如何提升网站站群全球化翻译的质量与部署效率,成为了亟待解决的技术问题
1.本发明通过对网站站群进行文档对象模型树遍历,提取文本节点及其深度值、兄弟节点顺序索引与挂载路径,完整保留页面文本的原生结构化特征,基于预设行业归属标识对页面内容文本集进行行业标签化处理,并依据预置行业语境适配规则库匹配对应行业的专属术语映射表与标准句式模板完成行业语境适配转换,实现翻译内容与原始页面布局的精准匹配,统一行业术语表述,使翻译内容贴合目标行业的表达习惯。
Smart Images

Figure CN122840073A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of website cluster translation technology, and in particular to a global intelligent language translation method for website clusters. Background Technology
[0002] Website cluster globalization language translation technology supports the batch construction of multilingual websites for enterprises and the reach of users in global markets. Current technology lacks structured and industry-specific fine-tuning in the page text extraction and translation processing stages. It simply scrapes the plain text content of the page, losing information about the text's layout and structure within the page. This results in misaligned translated content and elements, negatively impacting the user's browsing experience and the effectiveness of content delivery. Furthermore, it fails to differentiate between industry-specific expressions, using a uniform translation standard, leading to inconsistent industry terminology and sentence structures, and insufficient content professionalism.
[0003] Existing technologies have significant shortcomings in translation task scheduling and site deployment. The sequential or parallel allocation of translation tasks is unreasonable, resulting in low system resource utilization and lengthy processing times for large-scale site clusters. After translation generation, systematic page layout validation and element position correction are not performed, easily leading to issues such as element overlap and overflow. During the site deployment phase, automatic adaptation to data import configurations for different target sites is not possible, relying on manual configuration of field mapping relationships one by one. This process is cumbersome and prone to human error, hindering the overall progress of global site cluster deployment. Therefore, improving the quality and deployment efficiency of global translation for website clusters has become an urgent technical problem to be solved. Summary of the Invention
[0004] This invention provides a global intelligent language translation method for website clusters to solve the problems mentioned in the background art.
[0005] To achieve the above objectives, this invention provides a global intelligent language translation method for website clusters, comprising: S01. Obtain the text set of page content, text block boundary information, text block arrangement order, and text block embedding position of the website group to be processed; S02. Based on the preset industry affiliation identifier, the page content text set is processed by industry tagging to obtain the text data to be translated of the page content text set; S03. Based on a pre-set industry context adaptation rule library, perform industry context adaptation conversion on the text data to be translated to obtain industry-corrected text data of the text data to be translated. S04. Based on the text block boundary information, the industry-corrected text data is segmented to obtain industry-corrected text blocks of the industry-corrected text data. The industry-corrected text blocks are allocated to multiple translation execution threads to synchronously perform language conversion operations, thereby obtaining the translated text blocks of the text data to be translated. S05. Based on the arrangement order of the text blocks and the embedding position of the text blocks, perform structural fusion on the translated text blocks to obtain the translated page of the translated text blocks; S06. Compare the page element composition of the translated page and the website group to be processed. Based on the generated structural deviation list, adjust the element positions of the translated page and perform an adaptation test with the data import configuration of the target site to correct the data field mapping relationship of the target site and obtain the site deployment data package.
[0006] In a preferred embodiment, obtaining the page content text set, text block boundary information, text block arrangement order, and text block embedding position of the website group to be processed includes: The document object model tree is traversed for the websites in the website group to be processed, text nodes are extracted, and the depth value and sibling node order index of each text node in the document object model tree are recorded. The text content of all the text nodes is aggregated to form the page content text set of the website to be processed; Based on adjacent text nodes with the same parent node, the text nodes are grouped into one or more text blocks. According to the depth value of the first text node in the text block and the order index of the sibling nodes, the text block arrangement order of the relative positions of the text blocks in the page is generated. Record the correspondence between the text block and its parent node in the document object model tree, and use it as the text block embedding position of the website to be processed; The position offsets of the start and end text nodes of each text block in the document object model tree are used to determine the text block boundary information of the website to be processed.
[0007] In a preferred embodiment, the step of performing industry-specific tagging processing on the page content text set based on a preset industry affiliation identifier to obtain the text data to be translated from the page content text set includes: Obtain the preset industry affiliation identifier, which has a hierarchical industry coding structure; The text set of page content is segmented into words, and text feature words are extracted. The text feature words are then matched with the industry feature word library in the industry affiliation identifier to obtain the industry relevance weight of the text feature words. Based on the industry relevance weights, the page content text set is tagged and labeled to obtain the labeled text content set of the page content text set; The labeled text content set is integrated and transformed with the industry tags to generate the text data to be translated with a data pair structure.
[0008] In a preferred embodiment, the process of performing industry-specific context adaptation transformation on the text data to be translated based on a pre-set industry context adaptation rule base to obtain industry-corrected text data includes: The industry affiliation identifier of the text data to be translated is parsed, and the industry affiliation identifier is retrieved from the pre-set industry context adaptation rule base to obtain the industry sentence template and industry terminology mapping table of the industry affiliation identifier. The sentence sequence of the text data to be translated is traversed, and each sentence is structurally matched with the industry sentence template to obtain the segment of the text data to be translated that needs to be corrected. Based on the industry terminology mapping table, the terminology of the segment to be corrected is replaced, and the sentence structure of the replaced segment is adjusted to obtain the corrected sentence of the segment to be corrected. All the corrected statements are merged according to their original order in the text data to be translated to obtain the industry-corrected text data of the text data to be translated.
[0009] In a preferred embodiment, the step of segmenting the industry-corrected text data based on the text block boundary information to obtain industry-corrected text blocks of the industry-corrected text data includes: Parse the text block boundary information to obtain the starting character offset and ending character offset of the text block boundary information, and form an offset pair set; The industry-corrected text data is serialized, all text characters in the industry-corrected text data are extracted, and arranged in the original order in the industry-corrected text data to obtain the character sequence to be segmented in the industry-corrected text data. Traverse the character sequence to be segmented, and mark the corresponding segmentation start position and segmentation end position in the character sequence to be segmented according to the offset of each group of start character offset and end character offset in the set. On the character sequence to be segmented, starting from each segmentation start position and ending at its paired segmentation end position, a continuous character subsequence is segmented, and the character subsequence is used as the industry correction text block of the industry correction text data.
[0010] In a preferred embodiment, the step of allocating the industry-corrected text block to multiple translation execution threads and synchronously performing language conversion operations to obtain the translated text block of the text data to be translated includes: Obtain the number of currently available translation execution threads, and divide the industry-corrected text block into allocation groups corresponding to the number of translation execution threads based on the number of translation execution threads; The allocation group is sent to the corresponding translation execution thread, and the translation task carrying the industry-corrected text block is passed to the translation execution thread; When the translation execution thread receives the translation task, it performs a conversion operation from the source language to the target language on the received industry-corrected text block, generating a translated text block corresponding to each industry-corrected text block.
[0011] In a preferred embodiment, the step of structurally fusing the translated text blocks based on their arrangement order and embedding position to obtain the translated page of the translated text blocks includes: The text block embedding position is parsed to obtain the mounting path of each text block in the original page document object model tree, and the correspondence between the text block embedding position and the translated text block is established. The order of the translated text blocks on the page is determined based on the arrangement order of the text blocks; Construct the skeleton of the document object model for the translated page, and fill the content of the translated text block into the document object model tree node pointed to by the mounting path in the order described above; The translated page is generated by serializing all nodes of the document object model skeleton that have been filled with content.
[0012] In a preferred embodiment, the step of comparing the page element composition of the translated page and the website group to be processed, and adjusting the element positions of the translated page according to the generated structural deviation list, includes: Extract the mounting path and layout position coordinates of each page element in the translated page to form a position mapping relationship of the translated page elements; The original pages corresponding to the translated pages in the website group to be processed are analyzed, and the mounting paths and layout coordinates of each page element in the original pages are extracted to form the original page element position mapping relationship. Page elements with the same mounting path between the translated page element position mapping and the original page element position mapping are identified as comparison element groups. In the comparison element group, the position deviation vector between the layout position coordinates of the page element in the translated page and the layout position coordinates of the page element in the original page is used as the structural deviation amount of the comparison element group. Based on the pre-built page layout constraint relationship diagram and the structural deviation, the comparison element group is collaboratively corrected to obtain the correction offset vector of the comparison element group. The mounting paths and corresponding correction adjustment vectors of all element pairs whose initial position deviation vectors exceed the preset offset threshold are compiled into the structural deviation list. Traverse the list of structural deviations, locate the deviation elements in the translation page according to the mounting path, and use the correction adjustment vector to reverse the layout position coordinates of the deviation elements to obtain the corrected translation page.
[0013] In a preferred embodiment, the collaborative correction includes: ; In the formula, Indicates the first Corrected offset vector for group comparison of element groups Indicates the first Structural deviation of the group of elements. Indicated in the page layout constraint diagram, it is related to the first The set of other aligned element groups that have constraints on the aligned element groups. Indicates the preset first Group comparison of element group with the first The constraint strength coefficient between the group comparison elements is... This is the preset deviation balance factor.
[0014] In a preferred embodiment, the adaptation check with the data import configuration of the target site is performed to correct the data field mapping relationship of the target site, resulting in a site deployment data package, including: Obtain the data import configuration of the target site, parse the target page structure template defined in the data import configuration, and extract the field identifiers and field hierarchy paths that need to be filled in the target page structure template; Obtain the adjusted translation page, parse the page structure of the translation page, and extract the field identifiers and field hierarchy paths of each content block in the translation page; The field identifiers and field hierarchy paths required to be filled in the target page structure template are matched and compared with the field identifiers and field hierarchy paths of each content block in the translated page. Mapping conflict items with mismatched field identifiers or inconsistent field hierarchy paths are identified, and a mapping conflict list is generated. For the mapping conflict list, based on the preset field mapping rule library, the target site field identifier corresponding to the translated page field identifier in the mapping conflict item is found, and a corrected field mapping relationship is established to replace the original field mapping relationship; The data of each content block in the translated page is filled into the corresponding field positions of the target page structure template according to the corrected field mapping relationship, and then packaged into the site deployment data package.
[0015] Compared with the prior art, the present invention has the following beneficial effects: 1. This invention performs document object model tree traversal on website groups to extract text nodes and their depth values, sibling node order indexes, and mounting paths, fully preserving the original structural features of the page text. Based on preset industry affiliation identifiers, the page content text set is tagged with industry tags, and the industry context adaptation conversion is completed by matching the corresponding industry's exclusive terminology mapping table and standard sentence templates according to a preset industry context adaptation rule library. This achieves accurate matching between the translated content and the original page layout, unifies industry terminology, and makes the translated content conform to the expression habits of the target industry.
[0016] 2. This invention employs a multi-threaded parallel translation scheduling method. It dynamically allocates industry-corrected text blocks based on the number of available translation execution threads in the system to synchronously perform source-to-target language conversion operations. It constructs a page element composition comparison and structural deviation correction process, collaboratively adjusts element positions based on the page layout constraint diagram, and automatically completes adaptation checks with the target site's data import configuration to correct field mapping relationships. This generates a deployment data package that can be directly imported into the target site, improving system resource utilization and translation processing speed, maintaining the consistency of the overall page layout, simplifying the site deployment process, and reducing the risk of human error. Attached Figure Description
[0017] Figure 1 This is a flowchart illustrating a method for intelligent language translation of a website cluster for globalization, provided in an embodiment of the present invention. The realization of the objective, functional features and advantages of the present invention will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation
[0018] It should be understood that the specific embodiments described herein are merely illustrative of the invention and are not intended to limit the invention.
[0019] This application provides a method for intelligent language translation of a globalized website cluster. The executing entity of this method includes, but is not limited to, at least one of the following electronic devices that can be configured to execute the method provided in this application: a server, a terminal, etc. In other words, the method for intelligent language translation of a globalized website cluster can be executed by software or hardware installed on a terminal device or a server device. The server includes, but is not limited to, a single server, a server cluster, a cloud server, or a cluster of cloud servers. The server can be an independent server or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms.
[0020] Reference Figure 1 The diagram shown is a flowchart illustrating a method for intelligent language translation of a website group for globalization, according to an embodiment of the present invention. In this embodiment, the method includes: S01. Obtain the text set of page content, text block boundary information, text block arrangement order, and text block embedding position of the website group to be processed; In this embodiment of the invention, obtaining the page content text set, text block boundary information, text block arrangement order, and text block embedding position of the website group to be processed includes: The document object model tree is traversed for the websites in the website group to be processed, text nodes are extracted, and the depth value and sibling node order index of each text node in the document object model tree are recorded. The text content of all the text nodes is aggregated to form the page content text set of the website to be processed; Based on adjacent text nodes with the same parent node, the text nodes are grouped into one or more text blocks. According to the depth value of the first text node in the text block and the order index of the sibling nodes, the text block arrangement order of the relative positions of the text blocks in the page is generated. Record the correspondence between the text block and its parent node in the document object model tree, and use it as the text block embedding position of the website to be processed; The position offsets of the start and end text nodes of each text block in the document object model tree are used to determine the text block boundary information of the website to be processed.
[0021] For one of the websites in the website group to be processed, a Document Object Model (DOM) tree traversal is performed to extract the text nodes containing all visible text content. The traversal starts from the root node of the DOM tree and proceeds downwards along the node branches. When a node is reached, it is determined whether it contains direct text content. If it does, the node is marked as a text node, and its depth value in the DOM tree is recorded. The depth value represents the number of node levels traversed from the root node to the current text node. Simultaneously, the sibling node order index of this text node under its parent node is recorded, representing the sequence number of this text node among all child nodes of the same parent node. The text content contained within all extracted text nodes of the website to be processed is aggregated, and the strings within each text node are concatenated according to their natural order in the document to form the page content text set of the website to be processed.
[0022] Based on adjacent text nodes with the same parent node, text nodes belonging to the same parent node and consecutive in sibling node order index are grouped into a single group, with each group constituting a text block. All the divided text blocks are traversed. For each text block, the depth value and sibling node order index of the first text node within the block are obtained. The depth value of the first text node determines the page hierarchy of the text block, and the sibling node order index determines the order of appearance of the text blocks within the same hierarchy. The relative position of the text blocks on the page is generated by combining the hierarchy depths and sibling node order indices of all text blocks. This text block arrangement order indicates the order in which each text block is arranged from front to back in the page layout. The correspondence between each text block and its parent node in the Document Object Model (DOM) tree is recorded. The mounting path is a complete description of all intermediate nodes traversed from the root node of the DOM tree to the text block's parent node. This correspondence is used as the text block embedding position in the website to be processed, clearly defining the exact mounting node of each text block within the page document structure. The starting and ending text nodes within each text block are determined. The starting text node is the text node with the smallest sibling node sequential index in the text block, and the ending text node is the text node with the largest sibling node sequential index in the text block. The positional offsets of the starting and ending text nodes in the Document Object Model (DOM) tree are obtained. The positional offset of the starting text node is used as the starting boundary of the text block, and the positional offset of the ending text node is used as the ending boundary of the text block. The combination of these two boundary values constitutes the text block boundary information of the website to be processed. The text block boundary information defines the start and end range of each text block in the original page content character sequence.
[0023] The beneficial effects are that by explicitly traversing the document object model tree to extract text nodes and recording their depth values and sibling node order indices, the aggregation process of the page content text set fully preserves the original page's text content and hierarchical relationships, avoiding omissions and disordered text information. Text blocks are grouped based on adjacent text nodes with the same parent node, and the text block arrangement order is generated according to the depth value of the first text node and the sibling node order index. This ensures that the division of text blocks strictly follows the natural organization of the page document structure. The text block embedding position is clearly defined by recording the mounting path, specifying the precise location of each text block in the document tree. The text block boundary information is precisely defined by the position offsets of the starting and ending text nodes, accurately limiting the start and end range of each text block. These operations ensure, from the source, the complete acquisition and accurate transmission of the text structure information and page layout relationships of the website group to be processed.
[0024] S02. Based on the preset industry affiliation identifier, the page content text set is processed by industry tagging to obtain the text data to be translated of the page content text set; In this embodiment of the invention, the process of performing industry-labeling processing on the page content text set based on a preset industry affiliation identifier to obtain the text data to be translated from the page content text set includes: Obtain the preset industry affiliation identifier, which has a hierarchical industry coding structure; The text set of page content is segmented into words, and text feature words are extracted. The text feature words are then matched with the industry feature word library in the industry affiliation identifier to obtain the industry relevance weight of the text feature words. Based on the industry relevance weights, the page content text set is tagged and labeled to obtain the labeled text content set of the page content text set; The labeled text content set is integrated and transformed with the industry tags to generate the text data to be translated with a data pair structure.
[0025] The system acquires a pre-defined industry classification identifier. This identifier employs a hierarchical industry coding structure, divided into three levels from coarse to fine: major industry category code, intermediate industry category code, and minor industry category code. Each level corresponds to a fixed-length numeric or alphanumeric code segment. These three levels of code segments are concatenated in a predetermined order to form the complete industry classification identifier. The system then performs word segmentation on the page content text set. This segmentation uses a dictionary-based forward maximum matching method, matching continuous text strings within the page content text set against a pre-defined general word segmentation dictionary. Strings that successfully match the dictionary are extracted in descending order of length as text feature words. All extracted text feature words form a text feature word set.
[0026] Each text feature word in the text feature word set is associated and matched with the industry feature word library corresponding to the industry affiliation identifier. The industry feature word library is a vocabulary set pre-constructed for each industry's complete industry affiliation identifier. The industry feature word library stores frequently occurring technical terms and business-specific vocabulary within that industry. The specific method of association and matching is to perform a complete string match comparison on each text feature word in the industry feature word library. When a text feature word is completely identical to a term in the industry feature word library, it is considered a successful match. The successfully matched text feature word and the industry affiliation identifier level to which the matched industry feature word library belongs are recorded. The text feature word is assigned an industry relevance weight according to the industry affiliation identifier level at which the match occurs. Text feature words that match successfully at the major category level receive the first weight value, text feature words that match successfully at the intermediate category level receive the second weight value, and text feature words that match successfully at the minor category level receive the third weight value. The first weight value is less than the second weight value, and the second weight value is less than the third weight value. If a text feature word matches successfully at multiple levels, the highest weight value is retained. Finally, the industry relevance weight corresponding to each text feature word is obtained.
[0027] The page content text set is tagged and classified according to industry relevance weights. The page content text set is divided into several content segments to be tagged by sentences or natural paragraphs. For each content segment to be tagged, the industry relevance weight of all text feature words contained in the content segment is calculated. The industry affiliation identifier corresponding to the text feature word with the highest industry relevance weight is used as the industry label of the content segment to be tagged. If a content segment to be tagged contains multiple text feature words with the same highest industry relevance weight but they correspond to different industry affiliation identifiers, the industry affiliation identifier corresponding to the text feature word that appears most frequently among these text feature words is selected as the industry label. The content segments to be tagged after being labeled with industry labels are collected in their original order to form the labeled text content set of the page content text set. Each labeled content segment in the labeled text content set carries its corresponding industry label. The process involves integrating and transforming the annotated text content set with industry tags. Specifically, for each annotated content fragment in the annotated text content set, the text content of the annotated content fragment is used as the text part of the data pair structure, and the industry tag carried by the annotated content fragment is used as the tag part of the data pair structure. The text part and the tag part are combined to form a complete data pair. All data pairs are arranged according to the original order of the annotated content fragments in the page content text set to generate text data to be translated with a data pair structure.
[0028] The beneficial effects are as follows: By adopting a hierarchical industry coding structure for industry classification, the determination of industry classification is accurately locked layer by layer from broad category to subcategory, avoiding the ambiguity of classification caused by a single-level classification. A dictionary-based forward maximum matching method is used for word segmentation, and text feature words are compared and matched against an industry feature word library based on complete string consistency. Different industry relevance weights are assigned according to the level of the match, ensuring that the classification labeling accurately reflects the relevance of the page content to the industry. Integrating the labeled text content set and industry labels into a data pair structure establishes a clear one-to-one binding relationship between each content fragment in the text data to be translated and the industry classification label, providing an accurate and unambiguous basis for industry classification for subsequent contextual adaptation and conversion.
[0029] S03. Based on a pre-set industry context adaptation rule library, perform industry context adaptation conversion on the text data to be translated to obtain industry-corrected text data of the text data to be translated. In this embodiment of the invention, the step of performing industry-specific context adaptation transformation on the text data to be translated based on a pre-set industry context adaptation rule base to obtain industry-corrected text data of the text data to be translated includes: The industry affiliation identifier of the text data to be translated is parsed, and the industry affiliation identifier is retrieved from the pre-set industry context adaptation rule base to obtain the industry sentence template and industry terminology mapping table of the industry affiliation identifier. The sentence sequence of the text data to be translated is traversed, and each sentence is structurally matched with the industry sentence template to obtain the segment of the text data to be translated that needs to be corrected. Based on the industry terminology mapping table, the terminology of the segment to be corrected is replaced, and the sentence structure of the replaced segment is adjusted to obtain the corrected sentence of the segment to be corrected. All the corrected statements are merged according to their original order in the text data to be translated to obtain the industry-corrected text data of the text data to be translated.
[0030] The industry affiliation identifier carried by each data pair in the text data to be translated is parsed. The three levels of code segments of the industry affiliation identifier, namely the major industry code, the intermediate industry code, and the minor industry code, are extracted in sequence and concatenated to form a complete industry affiliation identifier retrieval key. The search is performed using a pre-defined industry context adaptation rule base based on the complete industry affiliation identifier search key. This rule base is a rule storage structure organized according to the hierarchical coding structure of industry affiliation identifiers. Each complete industry affiliation identifier search key corresponds to a rule entry. Each rule entry contains two parts: a set of industry sentence templates and a mapping table of industry terms. The industry sentence template set includes four commonly used sentence structures within the industry: declarative, interrogative, imperative, and heading sentences. Each sentence structure template specifies a fixed order and composition of the subject, predicate, object, attributive, and adverbial phrases. The industry terminology mapping table is a two-column lookup table. The first column stores non-standard, commonly used expressions within the industry, and the second column stores the corresponding standard industry terms. After retrieving a rule entry that matches the industry affiliation identifier search key, the industry sentence template set and industry terminology mapping table for that rule entry are retrieved.
[0031] The text data to be translated is traversed in a sentence sequence. The data pairs in the text data are arranged in the order of their appearance in the original page content text set. The text parts of each data pair are extracted one by one in this order, and each text part is divided into independent sentence units using periods, question marks, and exclamation marks as delimiters. For each segmented sentence unit, the sentence unit is matched structurally with the declarative sentence template, interrogative sentence template, imperative sentence template, and title sentence template in the industry sentence template set. The specific method of structural matching is to split the sentence unit according to the syntactic components of subject, predicate, object, attributive, and adverbial. It is then determined whether the arrangement of the syntactic components after splitting is consistent with the arrangement of components specified by the current industry sentence template. If the arrangement of the syntactic components of the split sentence unit can completely correspond to the arrangement order specified by a certain industry sentence template, the match is considered successful. The original text segment containing the sentence unit is marked as a segment to be corrected, and the type identifier of the industry sentence template that successfully matched the segment to be corrected is recorded. If a sentence unit fails to match the structure of any of the four industry sentence templates, the original text segment containing that sentence unit will be retained as a segment that does not require correction and will not be marked as a segment to be corrected.
[0032] For each marked segment to be corrected, terminology replacement is performed on the segment according to the industry terminology mapping table. The specific operation of terminology replacement is to traverse every word in the segment to be corrected, and search and compare each word in the first column of the industry terminology mapping table. When a word in the segment to be corrected is exactly the same as a non-standard idiomatic expression in the first column of the industry terminology mapping table, the non-standard idiomatic expression is replaced with the corresponding industry standard term in the second column of the industry terminology mapping table. Words in the segment to be corrected that do not match in the first column of the industry terminology mapping table remain unchanged. After completing the terminology replacement, the replaced fragment is obtained. Based on the industry sentence template type identifier recorded in the fragment to be corrected, the syntactic component arrangement order specified by the industry sentence template corresponding to that type identifier is obtained. The sentence structure of each syntactic component in the replaced fragment is adjusted according to the arrangement order specified by the industry sentence template. The adjustment method is to rearrange the subject, predicate, object, attributive, and adverbial components in the replaced fragment in order. For missing syntactic components in the replaced fragment, corresponding placeholders or conjunctions are inserted according to the conventional filling method of the component in that position in the industry sentence template. For redundant syntactic components in the replaced fragment, the redundant syntactic components are deleted, and finally, the corrected sentence is formed. All corrected sentences and all fragments that do not need to be corrected are merged according to the original order of the data pairs in the text data to be translated. During the merging, the correspondence between the internal text part of each data pair and its industry label remains unchanged. After the merging is completed, the industry corrected text data of the text data to be translated is formed.
[0033] The beneficial effects are as follows: By parsing the industry affiliation identifiers in the text data to be translated and concatenating them into hierarchical coded segments as search keys, a search is performed in the industry context adaptation rule base. The obtained industry sentence template set covers four sentence structures: declarative, interrogative, imperative, and heading. The industry terminology mapping table provides a complete correspondence between non-standard idiomatic expressions and industry standard terms. After breaking down each sentence in the text data to be translated into syntactic components such as subject, verb, object, modifier, and adverbial, and matching their arrangement order with the industry sentence templates, it is possible to accurately identify the segments that do not conform to industry standard expressions and need to be corrected. Based on the industry terminology mapping table, the segments to be corrected are replaced word by word, and the syntactic components of the replaced segments are readjusted according to the arrangement order specified by the industry sentence templates. This ensures that the corrected sentences conform to industry standard expression habits in terms of both vocabulary usage and sentence structure. The corrected industry-specific text data shows a significant improvement in semantic accuracy and industry standardization.
[0034] S04. Based on the text block boundary information, the industry-corrected text data is segmented to obtain industry-corrected text blocks of the industry-corrected text data. The industry-corrected text blocks are allocated to multiple translation execution threads to synchronously perform language conversion operations, thereby obtaining the translated text blocks of the text data to be translated. In this embodiment of the invention, the step of segmenting the industry-corrected text data based on the text block boundary information to obtain industry-corrected text blocks of the industry-corrected text data includes: Parse the text block boundary information to obtain the starting character offset and ending character offset of the text block boundary information, and form an offset pair set; The industry-corrected text data is serialized, all text characters in the industry-corrected text data are extracted, and arranged in the original order in the industry-corrected text data to obtain the character sequence to be segmented in the industry-corrected text data. Traverse the character sequence to be segmented, and mark the corresponding segmentation start position and segmentation end position in the character sequence to be segmented according to the offset of each group of start character offset and end character offset in the set. On the character sequence to be segmented, starting from each segmentation start position and ending at its paired segmentation end position, a continuous character subsequence is segmented, and the character subsequence is used as the industry correction text block of the industry correction text data.
[0035] The process of allocating the industry-corrected text block to multiple translation execution threads and synchronously performing language conversion operations to obtain the translated text block of the text data to be translated includes: Obtain the number of currently available translation execution threads, and divide the industry-corrected text block into allocation groups corresponding to the number of translation execution threads based on the number of translation execution threads; The allocation group is sent to the corresponding translation execution thread, and the translation task carrying the industry-corrected text block is passed to the translation execution thread; When the translation execution thread receives the translation task, it performs a conversion operation from the source language to the target language on the received industry-corrected text block, generating a translated text block corresponding to each industry-corrected text block.
[0036] The text block boundary information is parsed. The text block boundary information contains multiple sets of one-to-one corresponding start character offsets and end character offsets. Each start character offset indicates the starting character position number of the text block in the original page content text set, and each end character offset indicates the ending character position number of the text block in the original page content text set. All pairs of start character offsets and end character offsets are extracted and arranged in ascending order of start character offsets to form an offset pair set. Each pair of offsets in the offset pair set corresponds to the start and end range of a text block. The industry-corrected text data is serialized by iterating through all the text content and extracting each character, punctuation mark, and number. Invisible format control characters such as newlines and spaces in the original text are ignored. Each extracted visible character is then arranged in its original order in the industry-corrected text data to form a one-dimensional character sequence. This one-dimensional character sequence is the character sequence to be segmented in the industry-corrected text data. The position of each character in the character sequence to be segmented is uniquely identified by an incrementing integer position number starting from the beginning of the sequence.
[0037] Traverse the character sequence to be segmented, read the start character offset and end character offset of each offset pair in the offset pair set, take the value indicated by the start character offset as the segmentation start position number, and take the value indicated by the end character offset as the segmentation end position number. Count backwards from the beginning of the character sequence to be segmented. When the count value equals the segmentation start position number, the character position pointed to is marked as the segmentation start position. Continue counting backwards. When the count value equals the segmentation end position number, the character position pointed to is marked as the segmentation end position. Mark the corresponding segmentation start position and segmentation end position for each offset pair in the offset pair set in the character sequence to be segmented. All segmentation start and end position marking operations are completed group by group according to the order of the offset pairs in the offset pair set. On the character sequence to be segmented, starting from the first segmentation start position, all characters from this segmentation start position to its paired segmentation end position are extracted, including the characters at the segmentation start position and the characters at the segmentation end position, forming a continuous character subsequence. Then, the next set of offset pairs is processed sequentially, starting from the next segmentation start position and ending at its paired segmentation end position, extracting all characters to form another continuous character subsequence, until all offset pairs in the offset pair set have been processed. Each segmented continuous character subsequence is a sector-corrected text block in the sector-corrected text data, and the set of all sector-corrected text blocks constitutes the sector-corrected text block set of the sector-corrected text data.
[0038] The system retrieves the number of currently available translation execution threads. Each translation execution thread is an independent parallel processing unit running within the translation execution environment. It queries the total number of currently idle translation execution threads and uses this total number as the number of currently available threads. Based on the number of translation execution threads, the industry-specific text block set is divided. The division method involves sequentially allocating all industry-specific text blocks in the set according to their generation order. The first industry-specific text block is assigned to the first allocation group, the second to the second, and so on. Once the last allocation group is reached, the allocation process resumes from the first allocation group, continuing until all industry-specific text blocks have been allocated. This results in allocation groups equal to the number of translation execution threads. Each allocation group contains one or more industry-specific text blocks, arranged in the order of their allocation. Each allocation group is sent to a corresponding translation execution thread. The translation execution thread is then given a translation task containing all the industry-corrected text blocks within the allocation group. In addition to the text content of the industry-corrected text blocks, the translation task also includes a target language identifier and a source language identifier. The source language identifier is uniformly the language used in the original page content, and the target language identifier is the target language to be translated and output. When each translation execution thread receives a translation task, it performs a conversion operation from the source language to the target language for each received industry-corrected text block. The specific process of the conversion operation is that the translation execution thread internally calls a pre-set bilingual translation resource. The bilingual translation resource stores the mapping relationship between source language vocabulary and phrases and target language vocabulary and phrases, as well as the conversion rules between source language sentence structure and target language sentence structure. The translation execution thread first matches and searches for each word and phrase in the industry-corrected text block on the source language side of the bilingual translation resource to obtain the corresponding target language vocabulary and phrase. Then, it adjusts the word order and grammatical form of the sentences after word replacement according to the sentence structure conversion rules in the bilingual translation resource, and finally generates the corresponding translation text block. The text content in the translation text block is in the target language expression form. The character length of the translation text block may be different from the character length of the corresponding industry-corrected text block, but the semantic content carried by the translation text block corresponds to the corresponding industry-corrected text block. All translation execution threads perform the above translation operations in parallel. After all translation execution threads have completed their respective translation tasks, the generated translated text blocks are reassembled according to the original order of the industry-corrected text blocks in the industry-corrected text block set to obtain the translated text block set of the text data to be translated.
[0039] The beneficial effects are as follows: By parsing the boundary information of text blocks, the starting and ending character offsets are extracted to form an offset pair set. Character serialization is then performed on the industry-corrected text data to obtain a sequence of characters to be segmented, uniquely identified by an increasing position number. This establishes a precise positional mapping between the boundary information of the text blocks and the character sequence. Traversing the offset pair set, the starting and ending positions of the segmentation are marked on the character sequence to be segmented, and continuous character subsequences are extracted as industry-corrected text blocks, ensuring that the segmented industry-corrected text blocks completely correspond to the original text blocks in terms of content scope. The number of currently available translation execution threads is obtained, and the industry-corrected text blocks are evenly distributed to each translation execution thread in a circular allocation manner to perform language conversion operations in parallel. Internally, each translation execution thread calls bilingual translation resources to complete vocabulary matching and replacement and sentence structure adjustment. This parallel processing method significantly shortens the overall time of the site-group-level translation task, while maintaining consistency in translation quality among the translated text blocks output by each thread under the same industry context parameters.
[0040] S05. Based on the arrangement order of the text blocks and the embedding position of the text blocks, perform structural fusion on the translated text blocks to obtain the translated page of the translated text blocks; In this embodiment of the invention, the step of structurally fusing the translated text blocks based on the text block arrangement order and the text block embedding position to obtain the translated page of the translated text blocks includes: The text block embedding position is parsed to obtain the mounting path of each text block in the original page document object model tree, and the correspondence between the text block embedding position and the translated text block is established. The order of the translated text blocks on the page is determined based on the arrangement order of the text blocks; Construct the skeleton of the document object model for the translated page, and fill the content of the translated text block into the document object model tree node pointed to by the mounting path in the order described above; The translated page is generated by serializing all nodes of the document object model skeleton that have been filled with content.
[0041] The text block embedding locations are parsed. Each text block embedding location records the mapping relationship between its mounting path and its parent node in the original page's Document Object Model (DOM) tree. The mounting path starts from the root node of the DOM tree and proceeds downwards through a sequence of intermediate node names, with each level of node names connected by a predefined path separator to form a complete path string. The mounting paths stored in the text block embedding locations are read line by line to obtain the complete mounting path of each text block in the original page's DOM tree. Simultaneously, based on the identifier mapping relationship preserved during the text block generation process, each translated text block is mapped to the industry-corrected text block upon which it was generated. Then, the industry-corrected text block is mapped to the original text block. This mapping chain establishes the correspondence between the text block embedding locations and the translated text blocks. In other words, each translated text block corresponds to a specific mounting path, which indicates the target node location where the translated text block content should be placed.
[0042] The order of the translated text blocks on the page is determined by the arrangement of the text blocks. This arrangement records the relative position of each text block based on its hierarchy and sibling node order in the original page. The text block identifier sequence in the arrangement indicates the top-to-bottom and front-to-back order of the text blocks in the original page. The text block identifier sequences are then extracted sequentially according to their original order, and each identifier is replaced with its corresponding translated text block. This results in a sequence of translated text blocks' order on the page, which determines their final display order on the translated page. The document object model (Document Object Model) skeleton of the translated page is constructed. The Document Object Model skeleton is a tree-like node structure starting from the root node. The initial state of the Document Object Model skeleton is empty. First, the root node of the translated page is created. The root node is the top-level container node of the Document Object Model tree. The node type of the root node is consistent with the root node type of the original page. Then, the required intermediate nodes and leaf nodes are created sequentially on the corresponding mounting paths according to the order of the translated text blocks under the root node. If an intermediate node indicated in a mounting path has not yet been created in the Document Object Model skeleton, the missing intermediate node is created layer by layer from the root node according to the hierarchical order of the mounting path until all nodes at all levels of the mounting path exist in the Document Object Model skeleton. When creating nodes, the node type and attribute information of the corresponding nodes in the original page's Document Object Model tree are retained.
[0043] Following the order of the translated text blocks, the content of each translated text block is sequentially filled into the document object model tree node pointed to by the corresponding mounting path. The filling method involves inserting the text content of the translated text block as a text child node of that node. If the node already has a text child node, the content of the translated text block is appended to the existing text child node. The text child node stores all the character content of the translated text block, and the character encoding of the text child node is consistent with the character encoding of the translated text block. During the filling process, the text formatting tags inside the translated text block are not modified. After all the translated text blocks have been processed sequentially, all target nodes in the document object model skeleton of the translated page will contain the corresponding translated text block content. The serialization process involves serializing all nodes in the document object model (DOM) skeleton of the translated page. Starting from the root node of the DOM skeleton, the serialization recursively traverses the root node and all its descendant nodes. For each node, the node type label, node attributes, and text content are converted into a string representation according to a predefined format. The node type label is enclosed in angle brackets, and node attributes are appended to the node type label in pairs of name and value. The text content is appended directly after the closing position of the node label. Nested child nodes are recursively processed in the same format and placed before the closing label of their parent nodes. All the serialized strings are then concatenated in the traversal order to form a complete page file string, which is the final translated page.
[0044] The beneficial effect is that by parsing the text block embedding position, the complete mounting path of each text block in the original page's document object model tree is obtained, and a correspondence between the text block embedding position and the translated text block is established. This ensures that each translated text block can be accurately mapped to the predetermined mounting node in the document tree. When constructing the skeleton of the translated page's document object model, intermediate nodes and leaf nodes are created layer by layer on the corresponding mounting path according to the order of the translated text blocks. Then, the content of the translated text blocks is filled into the target node in the form of text child nodes. The filling process is strictly executed according to the text block arrangement order, so that the final translated page maintains the consistency of the node hierarchy and content display order with the original page in terms of document structure. This avoids structural misalignment or order disorder caused by changes in text length after translation. The generated translated page can be directly rendered and displayed in the target environment without manual secondary adjustments.
[0045] S06. Compare the page element composition of the translated page and the website group to be processed. Based on the generated structural deviation list, adjust the element positions of the translated page and perform an adaptation test with the data import configuration of the target site to correct the data field mapping relationship of the target site and obtain the site deployment data package.
[0046] In this embodiment of the invention, the step of comparing the page element composition of the translated page and the website group to be processed, and adjusting the element positions of the translated page according to the generated structural deviation list, includes: Extract the mounting path and layout position coordinates of each page element in the translated page to form a position mapping relationship of the translated page elements; The original pages corresponding to the translated pages in the website group to be processed are analyzed, and the mounting paths and layout coordinates of each page element in the original pages are extracted to form the original page element position mapping relationship. Page elements with the same mounting path between the translated page element position mapping and the original page element position mapping are identified as comparison element groups. In the comparison element group, the position deviation vector between the layout position coordinates of the page element in the translated page and the layout position coordinates of the page element in the original page is used as the structural deviation amount of the comparison element group. Based on the pre-built page layout constraint relationship diagram and the structural deviation, the comparison element group is collaboratively corrected to obtain the correction offset vector of the comparison element group. The mounting paths and corresponding correction adjustment vectors of all element pairs whose initial position deviation vectors exceed the preset offset threshold are compiled into the structural deviation list. Traverse the list of structural deviations, locate the deviation elements in the translation page according to the mounting path, and use the correction adjustment vector to reverse the layout position coordinates of the deviation elements to obtain the corrected translation page.
[0047] The collaborative correction includes: ; In the formula, Indicates the first Corrected offset vector for group comparison of element groups Indicates the first Structural deviation of the group of elements. Indicated in the page layout constraint diagram, it is related to the first The set of other aligned element groups that have constraints on the aligned element groups. Indicates the preset first Group comparison of element group with the first The constraint strength coefficient between the group comparison elements is... This is the preset deviation balance factor.
[0048] The data import configuration with the target site is subjected to an adaptation test to correct the data field mapping relationship of the target site, resulting in a site deployment data package, including: Obtain the data import configuration of the target site, parse the target page structure template defined in the data import configuration, and extract the field identifiers and field hierarchy paths that need to be filled in the target page structure template; Obtain the adjusted translation page, parse the page structure of the translation page, and extract the field identifiers and field hierarchy paths of each content block in the translation page; The field identifiers and field hierarchy paths required to be filled in the target page structure template are matched and compared with the field identifiers and field hierarchy paths of each content block in the translated page. Mapping conflict items with mismatched field identifiers or inconsistent field hierarchy paths are identified, and a mapping conflict list is generated. For the mapping conflict list, based on the preset field mapping rule library, the target site field identifier corresponding to the translated page field identifier in the mapping conflict item is found, and a corrected field mapping relationship is established to replace the original field mapping relationship; The data of each content block in the translated page is filled into the corresponding field positions of the target page structure template according to the corrected field mapping relationship, and then packaged into the site deployment data package.
[0049] Extract the mounting paths and layout coordinates of each page element in the translated page to form a mapping relationship between the translated page elements. The page element of the translated page refers to the node corresponding to each visible content block in the document object model tree of the translated document. The extraction process starts from the root node of the document object model tree of the translated document and traverses each node in a depth-first manner. For each node traversed, obtain the complete mounting path of the node in the document object model tree of the translated document. The mounting path is obtained by recording the name of each intermediate node layer by layer from the root node to the current node and connecting them with a path separator. The layout coordinates are obtained by reading the x-coordinate and y-coordinate of the top left corner of the rectangular area occupied by the current node in the rendered layout of the translated document, as well as the width and height of the rectangular area. These four values are combined to form the layout coordinates of the page element. The mounting path of each page element is used as the key and the layout coordinates are used as the value to establish the mapping relationship between the translated page elements. The process involves parsing the original pages corresponding to the translated pages within the website group to be processed. Based on the file name or page identifier of the translated page, the corresponding original page file is located within the website group. The mounting paths and layout coordinates of each page element in the original page are extracted using the same method as the extraction of the translated page element position mapping relationship. Starting from the root node of the original page's document object model tree, each node is traversed using a depth-first approach. The complete mounting path of each node is obtained, and the four values of the top-left corner x-coordinate, top-left corner y-coordinate, width, and height of the node's rectangular area are read from the rendered layout of the original page as the layout coordinates. The mounting path of each original page element is used as the key, and the layout coordinates as the value, to establish the original page element position mapping relationship.
[0050] The position mapping relationship of translated page elements is compared with that of the original page elements. For each mounting path appearing in the position mapping relationship of translated page elements, it is checked whether there is an identical mounting path in the position mapping relationship of the original page elements. If so, a pair of translated page elements and original page elements with the same mounting path are determined as a comparison element group. The layout position coordinates of the translated page elements in the comparison element group consist of four values: horizontal coordinate, vertical coordinate, width, and height. The layout position coordinates of the original page elements also consist of four values: horizontal coordinate, vertical coordinate, width, and height. The difference between the horizontal coordinate of the translated page element and the horizontal coordinate of the original page element is calculated as the horizontal deviation. The difference between the vertical coordinate of the translated page element and the vertical coordinate of the original page element is calculated as the vertical deviation. The difference between the width of the translated page element and the width of the original page element is calculated as the width deviation. The difference between the height of the translated page element and the height of the original page element is calculated as the height deviation. The horizontal deviation and the vertical deviation are combined to form a position deviation vector, which corresponds to the structural deviation of the comparison element group. The pre-built page layout constraint graph is a graph structure that represents the layout constraint relationships between page elements. The vertices in the graph structure are the pairs of elements, and the edges in the graph structure represent the layout constraint relationships between two pairs of elements. The types of layout constraint relationships include left and right alignment constraints, top and bottom alignment constraints, equal spacing constraints, and relative position constraints. Each edge is marked with a preset constraint strength coefficient. The magnitude of the constraint strength coefficient indicates the degree of mutual influence between the two pairs of elements when adjusting the layout.
[0051] Based on the pre-built page layout constraint relationship diagram and the structural deviation of each comparison element group, a collaborative correction is performed on all comparison element groups. The calculation process for the collaborative correction is as follows: ; In the formula, Indicates the first Corrected offset vector for group comparison of element groups Indicates the first Structural deviation of the group of elements. This indicates the relationship between the page layout constraint diagram and the first... The set of other aligned element groups that have constraints on the aligned element groups. Indicates the preset first Group comparison of element group with the first The constraint strength coefficient between the group comparison elements is... This is a preset deviation balance factor. For any given... Group comparison of elements to obtain the first group Structural deviation of the group of elements Then, find all relationships with the first element in the page layout constraint diagram. Other pairs of elements that have layout constraints with the same pair of elements constitute the set of pairs of elements. For the set of comparison elements Each pair of comparison elements Obtain the comparison element group Structural deviation And its relationship with the first Group comparison of constraint strength coefficients between element groups Each pair of comparison elements Structural deviation With the corresponding constraint strength coefficient Multiply the results, sum them up, and then divide the sum by the total constraint strength coefficients. The sum of the quotients is the first quotient. The weighted average of the deviations of the adjacent constraints on the group comparison elements is used to adjust the deviation balance factor. With the The structural deviation of the group compared to the element group itself Multiply, and simultaneously Subtract the deviation balance factor The obtained difference is multiplied by the weighted average of the deviations of adjacent constraints, and then the two products are added together. The result of the sum is the first value. Correction offset vector of group comparison element group Deviation balance factor It is a pre-set fixed value, ranging from zero to one, the deviation balance factor. A larger value indicates a greater tendency to retain the first value after collaborative correction. The effect of group ratio on the deviation of the element group itself, deviation balance factor The smaller the value, the more likely the collaborative correction will be to balance the deviation of the element group by referring to the surrounding constraints.
[0052] The correction offset vector obtained by collaborative correction includes two components: the correction horizontal offset and the correction vertical offset. The correction horizontal offset in the correction offset vector is added to the horizontal coordinate of the original translated page element in the corresponding comparison element group to obtain the adjusted horizontal coordinate. The correction vertical offset is added to the vertical coordinate of the original translated page element in the corresponding comparison element group to obtain the adjusted vertical coordinate. The width and height values remain unchanged from the original width and height of the translated page element. The adjusted horizontal coordinate, the adjusted vertical coordinate, and the original width and height are combined to form the adjusted layout position coordinates. Obtain the preset offset threshold, which is a pre-set upper limit for lateral and longitudinal offset. Compare the absolute value of the lateral offset in the position deviation vector of each comparison element group with the upper limit for lateral offset, and compare the absolute value of the longitudinal offset with the upper limit for longitudinal offset. If the absolute value of the lateral offset is greater than the upper limit for lateral offset or the absolute value of the longitudinal offset is greater than the upper limit for longitudinal offset, then gather the mounting path corresponding to the comparison element group and the correction adjustment vector corresponding to the comparison element group. The correction adjustment vector consists of two components: the correction lateral offset and the correction longitudinal offset. The correction lateral offset is the negative value of the correction lateral offset in the correction offset vector, and the correction longitudinal offset is the negative value of the correction longitudinal offset in the correction offset vector. Gather all mounting paths that meet the conditions and their corresponding correction adjustment vectors into a structural deviation list. Iterate through each entry in the structural deviation list, locate the corresponding deviation element in the translated page according to the mounting path in the entry, use the correction adjustment vector of the entry to adjust the layout position coordinates of the deviation element in reverse, directly add the correction horizontal offset in the correction adjustment vector to the current horizontal coordinate of the deviation element, and directly add the correction vertical offset to the current vertical coordinate of the deviation element to complete the layout position coordinate correction of the deviation element. After all deviation elements are corrected, the corrected translated page is obtained.
[0053] Retrieve the target site's data import configuration. The target site's data import configuration is a structured configuration file that describes how the target site receives and stores externally imported page content. Parse the target page structure template defined in the data import configuration. The target page structure template consists of a set of field definition entries. Each field definition entry contains a field identifier and a field hierarchy path. The field identifier is a unique name string of the field in the target site, and the field hierarchy path is a description of the field's hierarchical position in the target site's page structure. Extract all field identifiers and field hierarchy paths that are required to be filled in the target page structure template. The process involves obtaining the adjusted translated page, parsing its page structure (represented by content block nodes in the document object model tree), and traversing each content block node starting from the root node. For each content block node, the process retrieves its field identifier and field hierarchy path. The field identifier is obtained from the corresponding node identifier or preset semantic annotation in the original page, and the field hierarchy path is extracted from the mounting path of the content block node in the document object model tree of the translated page. This process extracts all field identifiers and field hierarchy paths for each content block in the translated page.
[0054] The field identifiers and field hierarchy paths required to be filled in the target page structure template are matched and compared with the field identifiers and field hierarchy paths of each content block in the translated page. The specific matching method is to search in the content blocks of the translated page for each field required to be filled in the target page structure template to see if there is a corresponding content block with the same field identifier and field hierarchy path. If no corresponding content block with the same field identifier is found, it is recorded as a mapping conflict item. The type of mapping conflict item is field identifier mismatch. If a corresponding content block with the same field identifier is found but the field hierarchy path is different, it is also recorded as a mapping conflict item. The type of mapping conflict item is inconsistent field hierarchy path. All identified mapping conflict items are summarized to generate a mapping conflict list. For each mapping conflict item in the mapping conflict list, processing is performed according to a preset field mapping rule library. The preset field mapping rule library stores the correspondence between field identifiers that the translated page may use and field identifiers required by the target site. For mapping conflicts of the type of field identifier mismatch, the target site field identifier corresponding to the field identifier of the translated page in the field mapping rule library is searched for in the field mapping rule library, and a corrected field mapping relationship from the field identifier of the translated page to the field identifier of the target site is established. For mapping conflicts of the type of inconsistent field hierarchy path, the corresponding target site field hierarchy path is searched for in the field mapping rule library, and a corrected field mapping relationship from the field hierarchy path of the translated page to the field hierarchy path of the target site is established. The original field mapping relationship is replaced with the corrected field mapping relationship. The data of each content block in the translated page is extracted block by block. Based on the field identifier of each content block, the corresponding target site field identifier and field hierarchy path are found in the corrected field mapping relationship. The text content and formatting information of the content block are filled into the corresponding field positions in the target page structure template specified by the target site field identifier and field hierarchy path. After all content blocks are filled, the filled target page structure template and all the data filled in it are packaged into a whole data package. During the packaging, the original file format and directory structure of the target page structure template are maintained. After the packaging is completed, the site deployment data package is obtained.
[0055] The beneficial effects are as follows: By extracting the mounting paths and layout coordinates of page elements from both the translated and original pages and establishing position mapping relationships, page elements with the same mounting paths are identified as comparison element groups, and position deviation vectors are calculated, accurately quantifying the positional offset of page elements before and after translation. Collaborative correction is performed based on the page layout constraint relationship diagram. When calculating the correction offset vector, both the structural deviation of the comparison element group itself and the deviation of adjacent constraint comparison element groups weighted by constraint strength coefficients are considered, ensuring that layout adjustments take into account the mutual constraints between elements. The correction adjustment vector is used to reverse the position of the deviation elements, so that the corrected translated page maintains layout coordination while restoring the visual effect of the original page. By importing the target site data configuration and comparing it with the corrected translated page for field identification and field hierarchy path adaptation, correction relationships are found in the field mapping rule base for mapping conflicts, and new field mappings are established. This ensures that the translated page content is accurately filled into the corresponding field positions of the target page structure template. Finally, the encapsulated site deployment data package can be directly imported into the target site for deployment.
[0056] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above, and that the present invention can be implemented in other specific forms without departing from the spirit or essential characteristics of the present invention.
[0057] This application embodiment can acquire and process relevant data based on artificial intelligence technology. Artificial intelligence is the theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to obtain optimal results.
[0058] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention.
Claims
1. A global intelligent language translation method for website clusters, characterized in that, The method includes: S01. Obtain the text set of page content, text block boundary information, text block arrangement order, and text block embedding position of the website group to be processed; S02. Based on the preset industry affiliation identifier, the page content text set is processed by industry tagging to obtain the text data to be translated of the page content text set; S03. Based on a pre-set industry context adaptation rule library, perform industry context adaptation conversion on the text data to be translated to obtain industry-corrected text data of the text data to be translated. S04. Based on the text block boundary information, the industry-corrected text data is segmented to obtain industry-corrected text blocks of the industry-corrected text data. The industry-corrected text blocks are allocated to multiple translation execution threads to synchronously perform language conversion operations, thereby obtaining the translated text blocks of the text data to be translated. S05. Based on the arrangement order of the text blocks and the embedding position of the text blocks, perform structural fusion on the translated text blocks to obtain the translated page of the translated text blocks; S06. Compare the page element composition of the translated page and the website group to be processed. Based on the generated structural deviation list, adjust the element positions of the translated page and perform an adaptation test with the data import configuration of the target site to correct the data field mapping relationship of the target site and obtain the site deployment data package.
2. The intelligent language translation method for globalizing website clusters as described in claim 1, characterized in that, The process of obtaining the page content text set, text block boundary information, text block arrangement order, and text block embedding position of the website group to be processed includes: The document object model tree is traversed for the websites in the website group to be processed, text nodes are extracted, and the depth value and sibling node order index of each text node in the document object model tree are recorded. The text content of all the text nodes is aggregated to form the page content text set of the website to be processed; Based on adjacent text nodes with the same parent node, the text nodes are grouped into one or more text blocks. According to the depth value of the first text node in the text block and the order index of the sibling nodes, the text block arrangement order of the relative positions of the text blocks in the page is generated. Record the correspondence between the text block and its parent node in the document object model tree, and use it as the text block embedding position of the website to be processed; The position offsets of the start and end text nodes of each text block in the document object model tree are used to determine the text block boundary information of the website to be processed.
3. The intelligent language translation method for globalizing website clusters as described in claim 1, characterized in that, The process of tagging the page content text set with industry-specific identifiers to obtain the text data to be translated from the page content text set includes: Obtain the preset industry affiliation identifier, which has a hierarchical industry coding structure; The text set of page content is segmented into words, and text feature words are extracted. The text feature words are then matched with the industry feature word library in the industry affiliation identifier to obtain the industry relevance weight of the text feature words. Based on the industry relevance weights, the page content text set is tagged and labeled to obtain the labeled text content set of the page content text set; The labeled text content set is integrated and transformed with the industry tags to generate the text data to be translated with a data pair structure.
4. The intelligent language translation method for globalizing website clusters as described in claim 1, characterized in that, The method, based on a pre-built industry context adaptation rule base, performs industry context adaptation transformation on the text data to be translated, obtaining industry-corrected text data of the text data to be translated, including: The industry affiliation identifier of the text data to be translated is parsed, and the industry affiliation identifier is retrieved from the pre-set industry context adaptation rule base to obtain the industry sentence template and industry terminology mapping table of the industry affiliation identifier. The sentence sequence of the text data to be translated is traversed, and each sentence is structurally matched with the industry sentence template to obtain the segment of the text data to be translated that needs to be corrected. Based on the industry terminology mapping table, the terminology of the segment to be corrected is replaced, and the sentence structure of the replaced segment is adjusted to obtain the corrected sentence of the segment to be corrected. All the corrected statements are merged according to their original order in the text data to be translated to obtain the industry-corrected text data of the text data to be translated.
5. The intelligent language translation method for globalizing website clusters as described in claim 1, characterized in that, The step of segmenting the industry-corrected text data based on the text block boundary information to obtain industry-corrected text blocks of the industry-corrected text data includes: Parse the text block boundary information to obtain the starting character offset and ending character offset of the text block boundary information, and form an offset pair set; The industry-corrected text data is serialized, all text characters in the industry-corrected text data are extracted, and arranged in the original order in the industry-corrected text data to obtain the character sequence to be segmented in the industry-corrected text data. Traverse the character sequence to be segmented, and mark the corresponding segmentation start position and segmentation end position in the character sequence to be segmented according to the offset of each group of start character offset and end character offset in the set. On the character sequence to be segmented, starting from each segmentation start position and ending at its paired segmentation end position, a continuous character subsequence is segmented, and the character subsequence is used as the industry correction text block of the industry correction text data.
6. The intelligent language translation method for globalizing website clusters as described in claim 1, characterized in that, The process of allocating the industry-corrected text block to multiple translation execution threads and synchronously performing language conversion operations to obtain the translated text block of the text data to be translated includes: Obtain the number of currently available translation execution threads, and divide the industry-corrected text block into allocation groups corresponding to the number of translation execution threads based on the number of translation execution threads; The allocation group is sent to the corresponding translation execution thread, and the translation task carrying the industry-corrected text block is passed to the translation execution thread; When the translation execution thread receives the translation task, it performs a conversion operation from the source language to the target language on the received industry-corrected text block, generating a translated text block corresponding to each industry-corrected text block.
7. The intelligent language translation method for globalizing website clusters as described in claim 1, characterized in that, The step of structurally fusing the translated text blocks based on their arrangement order and embedding position to obtain the translated page of the translated text blocks includes: The text block embedding position is parsed to obtain the mounting path of each text block in the original page document object model tree, and the correspondence between the text block embedding position and the translated text block is established. The order of the translated text blocks on the page is determined based on the arrangement order of the text blocks; Construct the skeleton of the document object model for the translated page, and fill the content of the translated text block into the document object model tree node pointed to by the mounting path in the order described above; The translated page is generated by serializing all nodes of the document object model skeleton that have been filled with content.
8. The intelligent language translation method for globalizing website clusters as described in claim 1, characterized in that, The step of comparing the page element composition of the translated page and the website group to be processed, and adjusting the element positions of the translated page according to the generated structural deviation list, includes: Extract the mounting path and layout position coordinates of each page element in the translated page to form a position mapping relationship of the translated page elements; The original pages corresponding to the translated pages in the website group to be processed are analyzed, and the mounting paths and layout coordinates of each page element in the original pages are extracted to form the original page element position mapping relationship. Page elements with the same mounting path between the translated page element position mapping and the original page element position mapping are identified as comparison element groups. In the comparison element group, the position deviation vector between the layout position coordinates of the page element in the translated page and the layout position coordinates of the page element in the original page is used as the structural deviation amount of the comparison element group. Based on the pre-built page layout constraint relationship diagram and the structural deviation, the comparison element group is collaboratively corrected to obtain the correction offset vector of the comparison element group. The mounting paths and corresponding correction adjustment vectors of all element pairs whose initial position deviation vectors exceed the preset offset threshold are compiled into the structural deviation list. Traverse the list of structural deviations, locate the deviation elements in the translation page according to the mounting path, and use the correction adjustment vector to reverse the layout position coordinates of the deviation elements to obtain the corrected translation page.
9. The intelligent language translation method for globalizing website clusters as described in claim 8, characterized in that, The collaborative correction includes: ; In the formula, Indicates the first Corrected offset vector for group comparison of element groups Indicates the first Structural deviation of the group of elements. Indicated in the page layout constraint diagram, it is related to the first The set of other aligned element groups that have constraints on the aligned element groups. Indicates the preset first Group comparison of element group with the first The constraint strength coefficient between the group comparison elements is... This is the preset deviation balance factor.
10. The intelligent language translation method for globalizing website clusters as described in claim 1, characterized in that, The data import configuration with the target site is subjected to an adaptation test to correct the data field mapping relationship of the target site, resulting in a site deployment data package, including: Obtain the data import configuration of the target site, parse the target page structure template defined in the data import configuration, and extract the field identifiers and field hierarchy paths that need to be filled in the target page structure template; Obtain the adjusted translation page, parse the page structure of the translation page, and extract the field identifiers and field hierarchy paths of each content block in the translation page; The field identifiers and field hierarchy paths required to be filled in the target page structure template are matched and compared with the field identifiers and field hierarchy paths of each content block in the translated page. Mapping conflict items with mismatched field identifiers or inconsistent field hierarchy paths are identified, and a mapping conflict list is generated. For the mapping conflict list, based on the preset field mapping rule library, the target site field identifier corresponding to the translated page field identifier in the mapping conflict item is found, and a corrected field mapping relationship is established to replace the original field mapping relationship; The data of each content block in the translated page is filled into the corresponding field positions of the target page structure template according to the corrected field mapping relationship, and then packaged into the site deployment data package.