A Multimodal Translation Method in the Context of Cross-border E-commerce
By identifying and analyzing links in e-commerce platform pages, combining natural language processing and multi-language page mapping tables, automatic link mapping and effectiveness monitoring are realized across languages, solving the problem of link processing in cross-border e-commerce scenarios, and improving user experience and translation system performance.
Patent Information
- Application Number
- CN202510033414.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-09
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-01-09
AI Technical Summary
In the cross-border e-commerce scenario, multimodal translation methods are difficult to effectively handle various links in the e-commerce platform page, resulting in a large number of dead or wrong links appearing in the page after translation, affecting the user experience and reducing the performance of the translation system.
By obtaining the content of the e-commerce platform page, identify and extract the absolute path, relative path and dynamic parameter information of product links, classified links and activity links. Natural language processing technology is used to analyze the link context, judge the target page type, and obtain the corresponding page address of the target language version from the pre-established multilingual page mapping table. Formulate processing rules for different types of links, automatically convert links in batches, generate links in the target language, and conduct effectiveness monitoring and automated testing.
It realizes cross-language automatic mapping and effective maintenance of e-commerce platform links, improves the user experience of multilingual version websites, and improves the performance of the translation system.
Smart Images

Figure CN119415787B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of information technology, and in particular to a multimodal translation method in a cross-border e-commerce scenario. Background Art
[0002] In cross-border e-commerce scenarios, multimodal translation approaches face the technical challenge of handling the various links found on e-commerce platform pages. These links, including product links, category links, and event links, are often deeply embedded in various parts of the page and tightly coupled with other elements. When performing multimodal translation on a page, not only must the text content be translated, but these links must also remain valid after translation, correctly redirecting to the corresponding pages in the target language. Link processing involves multiple technical considerations. First, the correspondence between pages in different language versions is complex: a single source language page may correspond to multiple target language pages, and vice versa. Second, links come in a variety of formats, including absolute or relative paths, and may contain dynamic parameters. Furthermore, links can change over time; for example, a link to a product may become invalid due to product delisting. These factors complicate link processing during multimodal translation. Improper link handling can result in a large number of broken or incorrect links on the translated page, severely impacting the user experience. Furthermore, the efficiency of link processing directly impacts the performance of the entire translation system. How to ensure the correctness of links while efficiently replacing them, achieve accurate link processing, and provide users with a smooth cross-language browsing experience is a technical problem that needs to be solved urgently. Summary of the Invention
[0003] The present invention provides a multimodal translation method in a cross-border e-commerce scenario, which mainly includes:
[0004] Obtain the page content of the e-commerce platform, identify product links, category links and activity links from the page content, and extract the absolute path, relative path and dynamic parameter information of the link; for the identified links, use natural language processing technology to analyze the link context and determine whether the target page type pointed to by the link is a product details page, a product list page or an activity theme page; based on the determined target page type, obtain the corresponding page address of the link in the target language version from the pre-established multilingual page mapping table to generate link mapping relationship data; for the generated link mapping relationship data, formulate processing rules for different types of links, use the processing rules to automatically convert the links in batches to obtain target language version links; monitor the validity of the obtained target language version links, and if an invalid link is detected, analyze the user behavior log and product information update record, and map the invalid link to the latest valid link; use automated testing tools to simulate user click behavior, perform access testing on the mapped links, check whether they jump to the target page correctly, and determine whether they are the correct target language version pages based on the similarity of the page content.
[0005] The technical solution provided by the embodiment of the present invention may have the following beneficial effects:
[0006] The present invention discloses a multimodal translation method for a cross-border e-commerce scenario. The method first identifies products, categories, and active links from the content of an e-commerce platform page, and extracts the path and parameter information of the link. Then, natural language processing technology is used to analyze the link context and determine the type of target page to which the link points. Based on the determination result, the corresponding page address of the link in the target language version is obtained from a pre-established multilingual page mapping table, and link mapping relationship data is generated. Then, processing rules for different types of links are formulated, and the links are automatically converted in batches to obtain target language version links. The validity of the converted links is monitored. If an invalid link is detected, the user behavior log and product information update record are analyzed, and the invalid link is mapped to the latest valid link. Finally, an automated testing tool is used to simulate user click behavior, and an access test is performed on the mapped link to check whether the jump is correct and determine the similarity of the page content. The present invention realizes cross-language automatic mapping and validity maintenance of e-commerce platform links, thereby improving the user experience of multilingual version websites. BRIEF DESCRIPTION OF THE DRAWINGS
[0007] Figure 1 This is a flowchart of a multimodal translation method in a cross-border e-commerce scenario of the present invention.
[0008] Figure 2 Schematic diagram of a multimodal translation method in a cross-border e-commerce scenario according to the present invention.
[0009] Figure 3This is another schematic diagram of a multimodal translation method in a cross-border e-commerce scenario according to the present invention. DETAILED DESCRIPTION
[0010] The technical solutions of the present invention will be clearly and completely described below in conjunction with the embodiments. Obviously, the embodiments described are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0011] like Figure 1-3 In this embodiment, a multimodal translation method in a cross-border e-commerce scenario may specifically include:
[0012] Step S101: Obtain the content of the e-commerce platform page, identify product links, category links and activity links from the page content, and extract the absolute path, relative path and dynamic parameter information of the link.
[0013] Obtain the page content of the e-commerce platform and parse the page content into a DOM tree structure; identify the corresponding link nodes from the DOM tree based on pre-defined feature rules for product links, category links, and active links; for each identified link node, determine whether it is an absolute path or a relative path; if the link node begins with "http: / / " or "https: / / ", it is determined to be an absolute path; otherwise, it is determined to be a relative path; for link nodes determined to be relative paths, convert the relative path to an absolute path based on the URL of the current page; extract the dynamic parameter part from the URL of the link node through a pre-defined regular expression pattern to obtain the parameter name and parameter value; store the extracted absolute path, relative path, and dynamic parameter information of the link node in a structured manner to build a link information database; the table structure of the link information database includes link ID, link URL, link type, and dynamic parameter fields; Based on the manually annotated training data set, the features of the link nodes are extracted, and one of the naive Bayes or decision tree algorithms is trained to obtain a link classification model; feature extraction is performed on the newly collected link nodes, and the extracted features are input into the link classification model to obtain the classification results of the link nodes; the classification results are stored in the link information database to improve the metadata information of the link nodes; and link analysis and application are performed based on the link information database.
[0014] Exemplarily, taking a well-known cross-border e-commerce platform as an example, the crawler program can simulate a browser to access the home page, obtain the HTML source code, and then use a parsing library such as Beautiful Soup to convert it into a DOM tree. This can conveniently extract various elements on the page. Identifying link nodes according to predefined rules is the key. For example, product links usually contain keywords such as "product" or "item", category links may have the word "category", and activity links often contain "activity" or "promotion". The target links can be quickly located through these features. It is important to determine whether the link is an absolute path. Taking "https: / / www.example.com / product / 12345" as an example, it starts with "https: / / " and is an absolute path. While " / category / electronics" is a relative path and needs to be concatenated into a complete link according to the current page URL. This can ensure that a complete and valid URL is used in subsequent processing. Extracting dynamic parameters from the link can obtain more information. For example, in "https: / / www.example.com / search?keyword=手机&page=2", the two parameters "keyword" and "page" and their values can be extracted. These parameters often contain important business meanings, such as search keywords, page numbers, etc. Storing the extracted link information in a database is convenient for subsequent use. A table named "links" can be designed, which contains fields such as id, url, type, params, etc. This can efficiently store and query a large amount of link data and provide a basis for subsequent analysis. Using machine learning algorithms to classify links is an effective method to improve accuracy. A batch of training data can be manually labeled. For example, " / product / 12345" is labeled as a product link, " / category / electronics" is labeled as a category link, etc. Then, features of the links are extracted, such as URL length, keyword occurrence times, etc., and a machine learning algorithm such as Naive Bayes or Decision Tree is trained to obtain a link classification model. This can automatically classify new unknown links. Analyzing based on the link database can obtain a lot of valuable information. For example, by analyzing the quantity and distribution of product links, the product structure of the platform can be understood; by tracking the changes of activity links, the promotion strategy of the platform can be grasped; by analyzing the link sequence accessed by users, the browsing path and interest preferences of users can be mined. These analysis results can provide more auxiliary information for multimodal translation in the cross-border e-commerce web page scenario, so as to obtain a more accurate and refined translation effect.
[0015] In step S102 , natural language processing technology is used to analyze the link context of the identified link to determine whether the target page to which the link points is a product details page, a product list page, or an activity topic page.
[0016] Get the target web page link to be analyzed, and get the HTML content of the target web page through the requests library; pre-process the HTML content, use regular expressions to remove JavaScript and CSS noise, use the jieba library for Chinese word segmentation, and use the THULAC tool for part-of-speech tagging; use the TextRank algorithm to extract keywords and key phrases from the pre-processed page content; obtain the feature word library of product details page, product list page and event special page pre-built using TF-IDF; calculate the cosine similarity between the target page content and the feature words of the three types of pages, and determine the page type with the highest similarity as the type of the target page; if the target page is determined to be a product details page, use Regular expressions are used to match HTML tags and their contents corresponding to the key attributes of product name, price, and specification parameters to extract structured product information fields. If the target page is judged to be a product list page, XPath is used to parse the HTML node tree to obtain the link, thumbnail URL, name, and price information of each product in the product list, and the steps are recursively called for each product link to obtain product information. If the target page is judged to be an event special page, regular expressions are used to match the HTML content corresponding to the key attributes of event theme, time, and participation method to extract structured event information fields. The parsed structured product information or event information is stored in the MongoDB database with the URL as the primary key, and the page type is recorded at the same time.
[0017] For example, taking a well-known cross-border e-commerce platform as an example, you can use Python's requests library to send HTTP requests to obtain the HTML content of the target page. For example, for a product details page, you may get raw content containing a large number of JavaScript, CSS, and HTML tags. In order to extract useful information, you need to pre-process the HTML content. Using regular expressions can effectively remove noise such as JavaScript and CSS. For example, you can use the pattern " <script>.*?< / script> "and" <style>。*?< / style>" to match and delete these contents. The preprocessed text is clearer and easier to analyze later. Chinese word segmentation is a key step in understanding the content of a page. The jieba library is an excellent Chinese word segmentation tool that can accurately segment continuous Chinese text into meaningful words. For example, "Huawei's latest smartphone" may be segmented into "Huawei / latest / smartphone". This segmentation result lays the foundation for subsequent text analysis. Part-of-speech tagging further enriches the linguistic information of the text. The THULAC tool can tag parts of speech for each segmentation result, such as nouns, verbs, adjectives, etc. This is crucial for understanding the role of words in sentences. For example, in "Red Limited Edition Mobile Phone", "red" and "limited edition" will be marked as adjectives, while "mobile phone" will be marked as a noun. The TextRank algorithm is a keyword extraction method based on a graph model. method. It regards the words in the text as nodes in the graph, and the co-occurrence relationship between words as edges. Through iterative calculation, the importance score of each word can be obtained. For example, in a page describing a smartphone, words such as "processor", "camera", and "battery life" may be extracted as keywords. Page type identification is an important part of e-commerce data analysis. Through the pre-built feature word library, the similarity between the target page and different types of pages can be calculated. For example, a product details page may contain feature words such as "specifications", "parameters", and "add to shopping cart", while a product list page may contain words such as "sort", "filter", and "pagination". By calculating cosine similarity, the page type can be accurately determined. For product details pages, structured product information needs to be extracted. Regular expressions can be used to accurately match product attributes in HTML tags. For example, you can use the pattern "<h1.*?> (.*?)" to extract the product name,<spanclass=“price”> (.*?)" to extract the price. This structured information provides the basis for subsequent product analysis and comparison. Parsing the product list page requires more complex HTML node traversal. XPath can be used to locate elements in HTML documents. For example, the XPath expression " / / div[@class='product-item']" can be used to select all product items, and then extract the link, image, name, price and other information of each product separately. Event theme pages usually contain rich marketing information. Regular expressions can be used to extract key attributes of the event, such as "Event time: (\d{4}-\d{2}-\d{2})" can match the event date. This information is crucial for analyzing marketing strategies and effects. Finally, the extracted structured data is stored in the MongoDB database. MongoDB's document storage feature is very suitable for storing complex product information. For example, a product document may contain fields such as "name", "price", and "specifications".Using the URL as the primary key ensures data uniqueness and traceability. This systematic data collection and analysis process provides e-commerce platforms with rich data support. For example, analyzing keywords on product detail pages can optimize search engines; comparing prices and specifications of different products can enable intelligent recommendations; and tracking visits and conversion rates on campaign pages can evaluate marketing effectiveness. Based on this data, multimodal translation results can be refined to improve translation accuracy, making the translated web content more aligned with user browsing habits and providing users with a smooth cross-language browsing experience.
[0018] Step S103 : According to the determined target page type, the corresponding page address of the link in the target language version is obtained from a pre-established multi-language page mapping table to generate link mapping relationship data.
[0019] Obtain the URL address of the link to be processed, and use regular expression matching to determine the target page type of the link; for the determined target page type, obtain a pre-established multilingual page mapping table, which records the URL addresses of corresponding pages in different language versions; obtain the corresponding page URL address that matches the target page type and target language version from the multilingual page mapping table; if the corresponding page URL address is found in the multilingual page mapping table, associate the link URL address with the obtained corresponding page URL address to generate link mapping relationship data; otherwise, record the unmatched link URL address for subsequent manual processing. Use the jieba word segmentation tool to segment the anchor text of the link, calculate the importance score of each word based on the TF-IDF algorithm, and select the words with the highest scores as keywords; store the extracted keywords as attribute fields of the link mapping relationship data together with the URL address. The K-means clustering algorithm is used to group link mapping data based on the similarity between the link URL and the anchor text. The edit distance between URL addresses and the jaccard similarity between anchor text keywords are calculated, and data with high similarity are grouped into the same cluster, resulting in a collection of link mapping data for different topics. The generated link mapping data is stored in a pre-established MySQL relational database. The database table structure includes the link URL, target page type, target language version, corresponding page URL, and anchor text keyword fields. The application queries the database to implement link jumps for multilingual web pages. Simultaneously, the stored multilingual link mapping relationships are used to perform machine translation on the web page content, generating web pages in different languages.
[0020] For example, the target page type is first determined by matching the URL address with a regular expression. For example, for a product detail page, the pattern " / product / \d+" might be used, while for a product list page, " / category / \w+" might be used. This method allows for quick and accurate identification of different page types. After determining the page type, the system queries a pre-established multilingual page mapping table. This mapping table records the URL addresses of corresponding pages in different language versions. For example, the Chinese version of the "Mobile Phones" category page might be " / category / shouji," while the English version would be " / category / mobile-phones." Using this mapping, the system can accurately establish link mappings between different language versions. If a matching URL is found in the mapping table, the system generates link mapping data. Links for which no matching URL is found in the mapping table are recorded and manually matched later. This may occur on newly launched pages or pages for special marketing campaigns. By regularly reviewing these unmatched links, the mapping table can be continuously refined, improving the system's coverage. This process involves not only URL matching but also anchor text processing. Using the Jieba word segmentation tool to segment anchor text effectively processes Chinese text. For example, "latest smartphone" might be segmented into "latest smartphone / smartphone / mobilephone." Term importance is then calculated using the TF-IDF algorithm. In this example, "smartphone" and "mobile phone" appear more frequently and are likely to receive higher scores because they better represent the core content of the link. The extracted keywords are stored along with the URL, forming rich link mapping data. This data is not only used for multilingual redirection but also for content analysis and search optimization. For example, keyword analysis can reveal the product features that users are most interested in, allowing for optimized product descriptions and marketing strategies. To better organize and manage this data, the system uses the K-means clustering algorithm to group link mapping data. This process considers the edit distance of URL addresses and the Jaccard similarity of anchor text keywords. For example, all links related to "smartphone" might be clustered together, while links related to "laptop" might form another cluster. This clustering method helps uncover underlying topic structures, which is valuable for content organization and recommendation systems. Finally, this processed and clustered link mapping data is stored in a MySQL database. The database table structure includes fields such as the link URL, target page type, target language version, corresponding page URL, and anchor text keywords. This structured storage not only facilitates query and management but also provides a foundation for subsequent data analysis and application. Through this system, e-commerce platforms can achieve seamless multilingual web page redirection.For example, when a user switches from the Chinese version to the English version, the system automatically maps the current page to the corresponding English version, ensuring a consistent user experience. Furthermore, this system provides a foundation for machine translation. By storing multilingual link mappings, the system can more accurately translate webpage content and generate high-quality multilingual versions. This not only improves translation efficiency but also ensures the accuracy of professional terminology and brand names, which is crucial for maintaining brand image and providing accurate information.
[0021] Step S104 : formulating processing rules for different types of links based on the generated link mapping relationship data, and automatically converting the links in batches using the processing rules to obtain target language version links.
[0022] Obtain link data in the source and target languages, and establish a link mapping relationship model using a decision tree algorithm; use the link mapping relationship model to predict the full link data to obtain preliminary link mapping relationship data; formulate corresponding processing rules for different link types in the link mapping relationship data, including rules for link structure conversion and rules for translating the page content corresponding to the link; use a distributed computing framework to batch process source language links, perform structural conversion on each link according to the processing rules, and call a machine translation API to translate the page content into the target language to obtain the target language version of the link data; perform quality inspection on the target language version of the link data, automatically check the links using predefined quality assessment rules, and manually spot check to identify problematic links; for problematic links discovered during quality inspection, use a heuristic algorithm to optimize the link conversion rules and machine translation parameters, and re-perform link conversion processing through multiple rounds of iteration until the preset quality standards are met; apply the optimized link conversion configuration to the full link data to generate the final target language version of the link, and push the link data to the multilingual website system through the Web service API.
[0023] For example, first, link data in the source and target languages is obtained, including URLs, page titles, and keywords. For example, on an e-commerce website, the URL for a product page in the source language (Chinese) might be "https: / / example.com / zh / products / smartphone," while the corresponding URL in the target language (English) might be "https: / / example.com / en / products / smartphone." Next, the structural features and content attributes of the links are analyzed. Structural features include URL pattern and path depth. In the example above, the URL pattern is " / language code / products / product name," with a path depth of 3. Content attributes include page title, description, and keywords. For example, the title of a Chinese page might be "smartphone," while the title of an English page is "Smartphone." Using the structural features and content attributes of links as input features and the link correspondences as labels, a decision tree algorithm is used to train a model capable of predicting link mappings. This model learns how to map source language URLs to target language URLs. The training data can be a manually annotated set of link correspondences. Once the model is trained, it can be used to predict link mappings for the entire link data set. Processing rules for different types of links require specific processing. For example, for product pages, the product name in the URL may need to be translated from Chinese to English. For category pages, the entire path structure may need to be adjusted. Furthermore, the translation of page content, including titles, body copy, and keywords, must also be considered. Distributed computing frameworks such as MapReduce can efficiently process large amounts of link data. During processing, the link structure is transformed according to pre-defined rules, and machine translation APIs such as the Google Translate API are used to translate the page content. This generates link data in the target language. To ensure the quality of generated links, quality checks are required. This includes automated checks for link validity and content integrity, as well as manual spot checks of selected links. For example, checks can be performed to verify that the generated English URL is accessible, the page content is complete, and the translation is accurate. Links identified as problematic can be optimized using heuristic algorithms. The conversion rules and translation parameters of the machine translation system to be optimized are obtained as initial input to the heuristic algorithm. The business content is analyzed to extract key attributes, and the relationships between attributes are determined to identify highly correlated attribute combinations. Based on the attribute combinations, an attribute association model is constructed using one of the machine learning algorithms, such as decision trees, support vector machines, and neural networks. The attribute association model is then combined with the conversion rules and translation parameters, and updated conversion rules and translation parameters are obtained through iterative optimization.Based on the updated conversion rules and translation parameters, the translation results are optimized by integrating contextual logic and thought chains to improve translation coherence and accuracy. The optimized translation results are evaluated by calculating metrics such as the BLEU score and semantic similarity to determine whether the translation quality meets the expected target. If the translation quality does not meet the expected target, further iterative optimization is performed. If the translation quality does meet the expected target, the optimized translation model is output, completing the heuristic algorithm optimization process. For example, if the translation of a certain product name is frequently incorrect, the translation engine parameters can be adjusted or a specialized terminology dictionary can be used. Through multiple rounds of iterative optimization, the quality of link conversion can be continuously improved. Finally, the optimized link data is pushed to the multilingual website system, and the link relationship data is imported into the search engine database. This not only connects web pages in different languages, but also ensures that users can find the corresponding language version of the page when searching. This entire process not only improves the user experience of the multilingual website but also helps enhance the website's internationalization and search engine optimization effectiveness.
[0024] Step S105 : monitoring the validity of the obtained target language version link. If an invalid link is detected, analyzing the user behavior log and product information update record, and mapping the invalid link to the latest valid link.
[0025] Get all product detail page links of the target language version to form an initial link list; for each link in the initial link list, determine whether it can be accessed normally by sending an HTTP request; if the status code returned by the HTTP request is 200, mark the link as a valid link; otherwise, mark it as a broken link and record the relevant information of the broken link; for each broken link, find out from the user behavior log of the business system whether there is a latest access record for the product ID and page ID corresponding to the broken link; if there is a latest access record, extract the link in the record as the potential latest valid link for the broken link; if no latest access record is found in the user behavior log, find the latest update record of the product corresponding to the broken link in the product information update record, and extract the updated product detail page link as the potential latest valid link for the broken link; for For invalid links that still cannot find the latest valid links, use the item-based collaborative filtering algorithm based on the category and keyword features of the products to find the N products with the highest feature similarity from the link library, and obtain the detail page links of these N products as potential replacement links for the invalid links; filter out invalid links for all potential latest valid links and potential replacement links; establish a mapping relationship between the filtered valid links and the corresponding invalid links to form a mapping table of invalid links to valid links; traverse all links on the page, and replace the invalid links with corresponding valid links in the mapping table; update the page after the link is replaced, and store the mapping table persistently in the database; for invalid links that cannot find corresponding valid links, remove them from the page, and log the information of these invalid links that cannot be processed, and review them manually on a regular basis.
[0026] For example, first, obtain all product detail page links in the target language. For example, a cross-border e-commerce website might have tens of thousands of links to product detail pages on its English version, such as "https: / / example.com / en / products / summer-dress-2023." Next, verify the validity of these links. By sending an HTTP request and checking the returned status code, you can determine whether the link is accessible. For example, if the link returns a 200 status code, it indicates the link is valid; if it returns a 404, it indicates the link is invalid, possibly due to product removal or page migration. For invalid links, you need to find alternatives. First, you can check user behavior logs. Suppose a user recently browsed to a summer dress with ID 12345, but the corresponding old link is no longer valid. By analyzing the logs, you might find that the user actually visited a new link, "https: / / example.com / en / products / summer-dress-2023-updated." This new link is likely the latest valid link for the product. If you can't find relevant information in the user behavior logs, you can look at the product information update records. For example, you may discover that product ID 12345 was recently updated, with the new detail page link becoming "https: / / example.com / en / products / summer-dress-2023-new-collection." This information can help find the latest valid link. If a replacement link is still unavailable, you can use an item-based collaborative filtering algorithm. Suppose the broken link is for a red dress. Based on characteristics such as color, style, and season, you can find the most similar product in the link library. For example, you may find a similar pink dress with the link "https: / / example.com / en / products / pink-summer-dress-2023." This link can serve as a potential replacement. After obtaining potential replacement links, you need to verify their validity again. This step helps filter out links that may also be broken, ensuring that only accessible pages are provided to users. Next, you need to establish a mapping between broken links and valid links. For example, you might get a mapping like this: {"https: / / example.com / en / products / summer-dress-2023":"https: / / example.com / en / products / summer-dress-2023-new-collection"}. This mapping will guide how to update the links on the web page.When updating page links, it's necessary to traverse all links on the page and replace them based on the mapping relationship. This process includes not only links to product detail pages, but may also include links in navigation bars, related recommendations, and other locations. The updated page will provide users with a more accurate browsing experience. For broken links that cannot be replaced, they need to be removed from the page. This prevents users from clicking on invalid links and improves the user experience. At the same time, the information about these unprocessed links should be recorded for subsequent manual review. This link management process not only improves website usability but also helps maintain the website's SEO effectiveness. By promptly updating and managing links, the number of dead links encountered by search engine crawlers can be reduced, improving the website's overall quality score. In addition, this process can help promptly detect changes in product inventory or classification, providing valuable information for inventory management and product strategy.
[0027] Step S106: Use an automated testing tool to simulate user click behavior, perform access testing on the mapped link, check whether it correctly jumps to the target page, and determine whether it is the correct target language version page based on the page content similarity.
[0028] Obtain all page links from the source language website and, based on predefined language mappings, generate corresponding target language links for each source language page link. For each source-target language link pair in the link mapping table, use automated testing tools to simulate user click behavior, triggering a redirect from the source language page to the target language page, and obtain the actual target language page content. Extract key content features from the response page, and use language detection tools to determine the page language type and confirm whether it is in the correct target language. Compute similarity between the text content of the response page and the text content of the intended target language page. Using the TF-IDF algorithm, convert the text of both pages into a term frequency-inverse document frequency matrix and calculate the cosine similarity between the two matrices, resulting in a similarity score between 0 and 1. A predefined similarity threshold is set and the similarity score for each link pair is evaluated. If the similarity exceeds the threshold, the link redirection is considered correct and the corresponding link mapping does not need to be adjusted. If the similarity is less than the threshold, the link redirection is considered problematic and the link pair is added to the list for optimization. For each source language-target language link pair in the list to be optimized, the latest content of the source language page is retrieved, and the source language page is translated into the target language through the machine translation interface to obtain a new target language page.
[0029] For example, first obtain all the page links of the source language website, which can be achieved through a website crawler. Take a cross-border e-commerce website as an example. Suppose that the Chinese version of the website has 100,000 product pages, and each page has a unique URL, such as "https: / / example.com / zh / products / summer-dress-2023". Next, generate the target language link according to the predefined language mapping relationship. For example, the mapping rule from Chinese to English may be to replace " / zh / " in the URL with " / en / ". In this way, the above Chinese link will be mapped to "https: / / example.com / en / products / summer-dress-2023". In this way, an initial link mapping table containing 100,000 pairs of Chinese and English links can be obtained. In order to verify the accuracy of these link mappings, it is necessary to simulate the actual browsing behavior of users. Using automated testing tools such as Selenium, you can write scripts to simulate users clicking the language switch button on the Chinese page, and then check whether the English page after the jump is as expected. This process can not only verify the validity of the link, but also detect whether the language switching function is working properly. After obtaining the target language page, it is necessary to extract the key content features of the page. This includes <title> The page title in the tag,< / title> <meta> Keywords in tags, <h1> arrive< / h1> <h6>The title content, and Paragraph text within the <title> tag. These elements typically contain core information about the page and can be used to determine the page's language and content. Language detection is a crucial step in ensuring the correct translation of a page. Using tools such as Google Language Detector, you can analyze the extracted text and determine its language. For example, if a page expected to be in English is detected as Chinese, this may indicate a link mapping error or an incorrect translation. To more accurately assess translation quality, it is necessary to calculate the text similarity between the source and target language pages. The TF-IDF (Term Frequency-Inverse Document Frequency) algorithm is an effective method. It converts page text into numerical vectors and then calculates the cosine similarity between these vectors to produce a score between 0 and 1. For example, if the similarity between a Chinese page and its English counterpart is 0.9, this indicates a high-quality translation; however, a similarity of only 0.3 may indicate serious translation issues. Setting a similarity threshold (such as 0.8) can help quickly identify links that require optimization. Link pairs with similarities below the threshold require further processing. This may involve re-acquiring the latest content of the source language page and generating a new target language page using a machine translation service such as the Google Translate API. By comparing the similarity between the new machine-translated page and the original target language page, a decision can be made as to whether the link mapping needs to be updated. If the new page's similarity is significantly higher than the original page, for example, increasing from 0.3 to 0.85, consideration can be given to replacing the original, incorrect link with the new machine-translated page. This optimization process is repeated until all link pairs meet the preset similarity threshold. Ultimately, an optimized link mapping table is generated, which serves as the basic configuration data for the website's language version switch. A detailed optimization report is also generated, including information such as the original similarity for each link pair, the optimized similarity, and overall similarity improvement statistics. This approach not only improves the quality of multilingual websites but also helps webmasters quickly identify and resolve translation issues. For example, if pages in a particular product category are found to have generally low similarity, this may indicate that the translation of specialized terminology in that category needs improvement. By analyzing the optimization report, webmasters can make targeted improvements to translation quality and enhance the user experience.
[0030] It should be noted that the above examples are merely specific embodiments of the present invention. Obviously, the present invention is not limited to the above examples and is subject to numerous variations. All variations that can be directly derived or conceived by a person skilled in the art from the disclosure of the present invention should be considered within the scope of protection of the present invention. < / h6>
Claims
1. A multimodal translation method in a cross-border e-commerce scenario, characterized in that: The method comprises: Obtain the page content of the e-commerce platform, identify product links, category links and activity links from the page content, and extract the absolute path, relative path and dynamic parameter information of the link; For the identified links, natural language processing technology is used to analyze the link context to determine whether the target page the link points to is a product details page, a product list page, or an event feature page; According to the determined target page type, the corresponding page address of the link in the target language version is obtained from a pre-established multi-language page mapping table to generate link mapping relationship data; it also includes obtaining the URL address of the link to be processed, and using regular expression matching to determine the target page type of the link; for the determined target page type, a pre-established multi-language page mapping table is obtained, and the mapping table records the URL addresses of the corresponding pages in different language versions; from the multi-language page mapping table, the corresponding page URL address that matches the target page type and the target language version is obtained; if the corresponding page URL address is found in the multi-language page mapping table, the URL address of the link is associated with the obtained corresponding page URL address to generate link mapping relationship data; otherwise, the unmatched link URL address is recorded for subsequent manual processing; the anchor text of the link is segmented using the jieba word segmentation tool, and each word is calculated based on the TF-IDF algorithm The importance scores of the words are calculated, and the words with the highest scores are selected as keywords; the extracted keywords are used as attribute fields of the link mapping relationship data and stored together with the URL address; the link mapping relationship data are grouped using the K-means clustering algorithm according to the similarity between the URL address and the anchor text of the link; the edit distance between the URL addresses and the jaccard similarity between the anchor text keywords are calculated, and the data with high similarity are divided into the same cluster to obtain link mapping relationship data sets of different topics; the generated link mapping relationship data is stored in a pre-established MySQL relational database, and the database table structure includes link URL, target page type, target language version, corresponding page URL and anchor text keyword field; the application realizes link jump of multi-language web pages by querying the database; at the same time, the stored multi-language link mapping relationship is used to machine translate the web page content to generate web pages in different language versions; Formulate processing rules for different types of links for the generated link mapping relationship data, and use the processing rules to automatically convert the links in batches to obtain target language version links; the processing rules include structure conversion rules for links and translation rules for page content corresponding to the links; Monitor the validity of the target language version links. If a broken link is detected, analyze the user behavior log and product information update records to map the broken link to the latest valid link. Use automated testing tools to simulate user click behavior, perform access tests on mapped links, check whether they jump to the target page correctly, and determine whether it is the correct target language version page based on the similarity of page content.
2. The method according to claim 1, characterized in that The method of obtaining the content of the e-commerce platform page, identifying product links, category links and activity links from the page content, and extracting the absolute path, relative path and dynamic parameter information of the link includes: Get the page content of the e-commerce platform and parse the page content into a DOM tree structure; According to the pre-defined characteristic rules of product links, category links and activity links, the corresponding link nodes are identified from the DOM tree; For each identified link node, determine whether it is an absolute path or a relative path; If the link node starts with "http: / / " or "https: / / ", it is considered an absolute path; Otherwise, it is determined to be a relative path; For link nodes that are determined to be relative paths, convert the relative paths into absolute paths based on the URL of the current page; From the URL of the link node, match and extract the dynamic parameter part through the pre-defined regular expression pattern to obtain the parameter name and parameter value; The absolute path, relative path and dynamic parameter information of the extracted link nodes are stored in a structured manner to build a link information database; The table structure of the link information database includes link ID, link URL, link type and dynamic parameter fields; Based on the manually annotated training data set, the features of the link nodes are extracted, and one of the naive Bayes or decision tree algorithms is trained to obtain a link classification model; Extract features from the newly collected link nodes, input the extracted features into the link classification model, and obtain the classification results of the link nodes; The classification results are stored in the link information database to improve the metadata information of the link nodes; Perform link analysis and application based on link information database.
3. The method according to claim 1, characterized in that: The method of analyzing the link context using natural language processing technology for the identified link to determine whether the target page type pointed to by the link is a product details page, a product list page, or an activity topic page includes: Get the target web page link to be analyzed, and obtain the HTML content of the target web page through the requests library; Preprocess HTML content, use regular expressions to remove JavaScript and CSS noise, use the jieba library for Chinese word segmentation, and use the THULAC tool for part-of-speech tagging; For the preprocessed page content, the TextRank algorithm is used to extract keywords and key phrases; Get the feature word library of product detail page, product list page and activity topic page pre-built using TF-IDF; Calculate the cosine similarity between the target page content and the three types of page feature words, and determine the page type with the highest similarity as the type of the target page; If the target page is determined to be a product details page, regular expressions are used to match the HTML tags and their contents corresponding to the key attributes of product name, price, and specification parameters to extract structured product information fields; If the target page is determined to be a product list page, XPath is used to parse the HTML node tree to obtain the link, thumbnail URL, name and price information of each product in the product list, and the steps are recursively called for each product link to obtain product information; If the target page is determined to be an event-themed page, regular expressions are used to match the HTML content corresponding to the key attributes of event theme, time, and participation method to extract structured event information fields; The parsed structured product information or activity information is stored in the MongoDB database with the URL as the primary key, and the page type is recorded at the same time.
4. The method according to claim 1, characterized in that The method of formulating processing rules for different types of links for the generated link mapping relationship data, and using the processing rules to automatically convert the links in batches to obtain target language version links includes: Obtain link data of source language and target language, and use decision tree algorithm to establish link mapping relationship model; Use the link mapping relationship model to predict the full amount of link data to obtain preliminary link mapping relationship data; Use a distributed computing framework to batch process source language links, perform structural transformation on each link according to processing rules, and call the machine translation API to translate the page content into the target language to obtain the link data of the target language version; Perform quality checks on the target language version of the link data, automatically check the links using predefined quality assessment rules, and manually spot check to identify problematic links; For problematic links found in quality inspection, we use heuristic algorithms to optimize link conversion rules and machine translation parameters, and re-convert links for multiple rounds until the preset quality standards are met. Apply the optimized link conversion configuration to the full link data to generate the final target language version link, and push the link data to the multilingual website system through the Web service API.
5. The method according to claim 1, characterized in that The validity of the obtained target language version link is monitored. If an invalid link is detected, the user behavior log and product information update record are analyzed to map the invalid link to the latest valid link, including: Get all product detail page links of the target language version to form an initial link list; For each link in the initial link list, determine whether it can be accessed normally by sending an HTTP request; If the status code returned by the HTTP request is 200, the link is marked as a valid link; Otherwise, mark it as a broken link and record relevant information of the broken link; For each invalid link, check the user behavior log of the business system to see whether there is a latest access record of the product ID and page ID corresponding to the invalid link; If there is a latest access record, extract the link in the record as the potential latest valid link of the invalid link; If the latest access record is not found in the user behavior log, the latest update record of the product corresponding to the invalid link is searched in the product information update record, and the updated product details page link is extracted as the potential latest valid link of the invalid link; For invalid links that still cannot find the latest valid links, we use the item-based collaborative filtering algorithm to find the N products with the highest feature similarity from the link library according to the category and keyword features of the products they belong to, and obtain the detail page links of these N products as potential replacement links for the invalid links; For all potential latest valid links and potential alternative links, filter out invalid links; Establishing a mapping relationship between the filtered valid links and the corresponding invalid links to form a mapping table from invalid links to valid links; Traverse all the links on the page, and replace the invalid links with the corresponding valid links for those that have corresponding valid links in the mapping table; Update the page after link replacement and store the mapping table persistently in the database; For invalid links that cannot find corresponding valid links, they will be removed from the page. At the same time, the information of these invalid links that cannot be processed will be recorded in the log and reviewed manually on a regular basis.
6. The method according to claim 1, characterized in that The automated testing tool is used to simulate user click behavior, perform access testing on the mapped link, check whether the target page is correctly jumped to, and determine whether it is the correct target language version page based on the similarity of the page content, including: Obtain all page links of the source language website, and generate the target language page link corresponding to each source language page link according to the predefined language mapping relationship; For each source language-target language link pair in the link mapping table, an automated testing tool is used to simulate user click behavior, trigger a jump from the source language page to the target language page, and obtain the actual response target language page content; Extract key content features of the response page, use language detection tools to determine the page language type, and confirm whether it is the correct target language; Calculate the similarity between the text content of the response page and the text content of the expected target language page; Using the TF-IDF algorithm, the texts of the two pages are converted into term frequency-inverse document frequency matrices, and the cosine similarity of the two matrices is calculated to obtain a similarity score between 0 and 1; Set a predefined similarity threshold and judge the similarity score of each link pair; If the similarity is greater than the threshold, the link jump is determined to be correct and the corresponding link mapping does not need to be adjusted; If the similarity is less than the threshold, it is determined that there is a problem with the link jump, and the link pair is recorded in the list to be optimized; For each source language-target language link pair in the list to be optimized, the latest content of the source language page is retrieved, and the source language page is translated into the target language through a machine translation interface to obtain a new target language page.
Citation Information
Patent Citations
Webpage favorite management method, device and system
CN105069011A
Dynamic page conversion method and device
CN106156080A
Webpage language switching method and device and terminal device
CN110362370A