Non-intrusive translation method and related device, electronic equipment and storage medium
Through the client, the virtual tree structure is generated and translated and restored, the problem of the impact of translated web pages on the security and accuracy of the original web page in the existing technology is solved, and a non-invasive and high-accuracy translation is achieved.
Patent Information
- Application Number
- CN202411185642.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-08-27
- Publication Date
- 2025-05-23
AI Technical Summary
The prior art causes the security and accuracy of the original web page to be reduced when generating and translating web pages.
The client responds to the translation instructions to parse the web page to be translated, builds a virtual tree structure, and preprocesses each element to be translated, generates the data to be translated and the data to be restored, and is translated and restored independently of the original web page.
It realizes non-invasive translation and restoration of the original web page, protects the privacy and security of the original web page, and improves the accuracy of restoration of multimodal elements in the virtual web page.
Smart Images

Figure CN120031049A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of natural language processing, and in particular to a non-intrusive translation method and related devices, electronic devices and storage media. Background Art
[0002] With the rapid development of Internet technology, the global level of informatization continues to improve. The Internet has become the core channel for information dissemination. Multilingual support plays a vital role in it and is indispensable for improving user experience and promoting cross-cultural communication.
[0003] Due to the diversity of languages, users of different languages often face obstacles in obtaining and understanding information during cross-language and cross-cultural communication. In the prior art, web pages are usually translated using general machine translation technology on the server side or specific vertical scenario translation plug-in tools based on the browser side.
[0004] However, the above methods need to modify the source code of the web page or rely on specific browser plug-ins, which causes a certain intrusive impact on the original web page and cannot protect the security and accuracy of the original web page. Summary of the invention
[0005] The main technical problem solved by the present application is to provide a non-intrusive translation method and related devices, electronic devices and storage media, which can solve the problem that the security and accuracy of the original web page are reduced due to the generation of translated web pages by the prior art.
[0006] To solve the above technical problems, the first technical solution adopted in the present application is to provide a non-intrusive translation method, including: the client parses the web page to be translated in response to the translation instruction, and constructs a virtual tree structure based on the obtained multiple elements to be translated; preprocesses each element to be translated to generate multiple data to be translated; wherein each data to be translated includes a unique identifier and a text to be translated; obtains multiple data to be restored corresponding to the multiple data to be translated; wherein each data to be restored includes a unique identifier and a translation text; restores each data to be restored into restored data; based on the unique identifier corresponding to each restored data, renders the multiple restored data into the virtual tree structure to generate a virtual translated web page.
[0007] The step of parsing the web page to be translated by the client in response to the translation instruction and constructing a virtual tree structure based on the obtained multiple elements to be translated comprises: the client dynamically identifies the web page to be translated in response to the translation instruction to obtain a web page element set; wherein the web page element set comprises at least one type of web page element and a label corresponding to the web page element; the type comprises at least one of a text type, an image type and a font package type; traversing the web page element set by the label to construct a virtual tree structure based on the identified multiple elements to be translated; preprocessing each element to be translated to generate multiple data to be translated comprises: classifying the elements to be translated according to the label, Based on the classification result, each element to be translated is preprocessed accordingly to generate multiple data to be translated including unique identifiers and translation texts; the step of obtaining multiple data to be restored corresponding to the multiple data to be translated includes: sending multiple translation requests carrying the data to be translated to the first-level cache of the client, or sending multiple translation requests to the first-level cache and sending at least part of the translation requests to the server through the first-level cache to receive the multiple data to be restored including unique identifiers and translation texts returned by the first-level cache or the server; the step of restoring each data to be restored into restored data includes: restoring each data to be restored into restored data of a corresponding type based on the classification result.
[0008] Among them, the steps of classifying the elements to be translated according to the labels, and performing corresponding preprocessing on each element to be translated based on the classification results to generate a plurality of data to be translated including a unique identifier and a translation text include: determining the type corresponding to each element to be translated according to the label; in response to the element to be translated being a text type, extracting the text information to be translated in the element to be translated to generate the text to be translated; in response to the element to be translated being a picture type, sending the element to be translated to a server so that the server performs picture recognition on the element to be translated, and generates the text to be translated based on the text information returned by the server; in response to the element to be translated being a font package type, performing character recognition and calculation on the element to be translated, and converting the element to be translated into a picture based on the recognition result and the calculation result; sending the picture to the server so that the server performs picture recognition on the picture, and generates the text to be translated based on the text information returned by the server; generating a corresponding unique identifier based on each text to be translated, and generating the data to be translated based on the unique identifier, the type corresponding to the label, the type corresponding to the element to be translated, and the text to be translated; and constructing a queue of data to be translated based on a plurality of data to be translated.
[0009] The step of sending a plurality of translation requests carrying data to be translated to a first-level cache of a client, or sending a plurality of translation requests to a first-level cache and sending at least part of the translation requests to a server through the first-level cache to receive a plurality of data to be restored including a unique identifier and a translation text returned by the first-level cache or the server, comprises: sending a plurality of translation requests to a first-level cache of a client in sequence; querying whether there is data to be restored corresponding to the data to be translated in the first-level cache based on each translation request; in response to the existence of data to be restored corresponding to the data to be translated in the first-level cache, receiving the data to be restored returned by the first-level cache; in response to the absence of data to be restored corresponding to the data to be translated in the first-level cache, sending the translation request to a second-level cache of a server through the first-level cache, and receiving the data to be restored returned by the second-level cache; building a data queue to be restored based on the plurality of data to be restored; wherein the data to be restored includes a unique identifier, a type corresponding to a tag, a type corresponding to an element to be translated, a target language identifier, a translation text, and location information.
[0010] Among them, the step of restoring each to-be-restored data into the corresponding type of restored data based on the classification result includes: inputting the to-be-restored data in the to-be-restored data queue into the multimodal restoration software package in sequence according to the first-in-first-out principle; wherein the multimodal restoration software package encapsulates a variety of restoration algorithms for processing multimodal data, and the multiple restoration algorithms include a text restoration algorithm, an image restoration algorithm, and a font package restoration algorithm; based on the type of the to-be-restored data, calling the corresponding restoration algorithm to restore the to-be-restored data to output the restored data.
[0011] Among them, the step of calling the corresponding restoration algorithm based on the type of the data to be restored to restore the data to output the restored data includes: in response to the data to be restored being of text type, calling the text restoration algorithm to process the translated text in the data to be restored to avoid overflow of the translated text; in response to the data to be restored being of image type, calling the image restoration algorithm to scale the image generated based on the translated text in the data to be restored; in response to the data to be restored being of font package type, calling the font package restoration algorithm to restore the image generated based on the translated text in the data to be restored based on the calculation result corresponding to the font package type.
[0012] Among them, after the step of determining the type corresponding to each element to be translated according to the label, it includes: in response to the element to be translated being a picture type or a font package type, and the server has not extracted text information from the picture corresponding to the element to be translated, each element to be translated from which text information has not been extracted is merged into the same Sprite image, and the coordinates of each element to be translated in the Sprite image are recorded in a virtual tree structure; based on the unique identifier corresponding to each restored data, multiple restored data are rendered into the virtual tree structure to generate a virtual translation web page, including: splitting each element to be translated in the Sprite image, and rendering each element to be translated in the Sprite image into the virtual translation web page based on the corresponding coordinates.
[0013] To solve the above technical problems, the second technical solution adopted in the present application is to provide a non-intrusive translation method, including: a server receives multiple translation requests carrying data to be translated sent by a client; wherein the data to be translated is generated by the client based on multiple elements to be translated identified by the web page to be translated, and each data to be translated includes a unique identifier and a text to be translated; based on the multiple translation requests, multiple data to be restored corresponding to the multiple data to be translated are obtained; wherein each data to be restored includes a unique identifier and a translation text; the multiple data to be restored are sent to the client, so that the client restores each data to be restored into restored data, and according to the unique identifier corresponding to each restored data, the multiple restored data are rendered into a virtual tree structure constructed by the client based on the multiple elements to be translated, so as to generate a virtual translated web page.
[0014] Among them, the step of obtaining multiple data to be restored corresponding to multiple data to be translated based on multiple translation requests includes: based on each translation request, querying whether there is data to be restored corresponding to the data to be translated in the secondary cache of the server; in response to the existence of data to be restored corresponding to the data to be translated in the secondary cache, sending the data to be restored to the primary cache of the client; in response to the absence of data to be restored corresponding to the data to be translated in the secondary cache, sending a translation request to the cloud platform, so that the cloud platform translates the data to be translated based on the translation request; receiving and storing the data to be restored returned by the cloud platform through the secondary cache, and sending the data to be restored to the primary cache of the client.
[0015] Among them, before the server receives a plurality of data to be translated and corresponding translation requests sent by the client, it includes: establishing a correction theme library in the server; wherein the correction theme library stores a plurality of professional fields and professional terms and translation knowledge corresponding to the professional fields, and each professional field carries at least one classification label; after the server receives a plurality of translation requests carrying data to be translated sent by the client, it includes: based on each text to be translated, storing the corresponding translation request in the most matching professional field; after the step of receiving and storing the data to be restored returned by the cloud platform through the secondary cache, and sending the data to be restored to the primary cache of the client, it includes: storing the data to be restored in the most matching professional field; performing manual correction processing on the translation data stored in the correction theme library, and synchronously storing the translation data after manual correction processing in the secondary cache of the server.
[0016] To solve the above technical problems, the third technical solution adopted in the present application is to provide a client, including: a construction module, which is used to parse the web page to be translated in response to the translation instruction, and to construct a virtual tree structure based on the obtained multiple elements to be translated; a preprocessing module, which is used to preprocess each element to be translated to generate multiple data to be translated; wherein each data to be translated includes a unique identifier and a text to be translated; a first acquisition module, which is used to obtain multiple data to be restored corresponding to the multiple data to be translated; wherein each data to be restored includes a unique identifier and a translation text; a restoration module, which is used to restore each data to be restored into restored data; a generation module, which is used to render the multiple restored data into the virtual tree structure based on the unique identifier corresponding to each restored data, so as to generate a virtual translated web page.
[0017] To solve the above technical problems, the fourth technical solution adopted in the present application is to provide a server, comprising: a receiving module, used to receive multiple translation requests carrying data to be translated sent by a client; wherein the data to be translated is generated by the client based on multiple elements to be translated identified by the web page to be translated, and each data to be translated includes a unique identifier and a text to be translated; a second acquisition module, used to obtain multiple data to be restored corresponding to the multiple data to be translated based on the multiple translation requests; wherein each data to be restored includes a unique identifier and a translation text; a sending module, used to send the multiple data to be restored to the client, so that the client restores each data to be restored into restored data, and according to the unique identifier corresponding to each restored data, renders the multiple restored data into a virtual tree structure constructed by the client based on the multiple elements to be translated, so as to generate a virtual translation web page.
[0018] In order to solve the above technical problems, the fifth technical solution adopted in the present application is to provide an electronic device, including: a memory, used to store program data, and when the stored program data is executed, the steps in the non-intrusive translation method as described in any one of the above items are implemented; a processor, used to execute the program instructions stored in the memory to implement the steps in the non-intrusive translation method as described in any one of the above items.
[0019] In order to solve the above technical problems, the sixth technical solution adopted in the present application is to provide a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps in the non-intrusive translation method as described in any one of the above items are implemented.
[0020] The beneficial effects of the present application are as follows: Different from the prior art, the present application provides a non-intrusive translation method and related devices, electronic devices and storage media, and parses the web page to be translated by the client in response to the translation instruction, and constructs a virtual tree structure based on the obtained multiple elements to be translated, and can accommodate multiple elements to be translated through a virtual area independent of the web page to be translated, thereby reducing the impact on the original web page. By preprocessing each element to be translated, generating multiple data to be translated including a unique identifier and a text to be translated, and obtaining multiple data to be restored including a unique identifier and a translation text corresponding to the multiple data to be translated, the elements to be translated can be matched one by one with the data to be translated and the data to be restored through the unique identifier, so as to ensure the accurate storage and reference of the relevant data in the virtual number structure, thereby improving the consistency and accuracy of the data. Each data to be restored is restored to restored data, and multiple restored data are rendered into a virtual tree structure based on the unique identifier corresponding to each restored data, so as to improve the restoration accuracy of each element to be translated in the virtual translation web page. The present application can realize the non-intrusive translation and restoration of the original web page through the client script, thereby protecting the privacy and security of the original web page and improving the accuracy of the restoration of multimodal elements in the virtual web page. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0022] Figure 1 It is a principle block diagram of an implementation method of the non-intrusive translation system of the present application;
[0023] Figure 2 yes Figure 1 A signal flow diagram of an implementation method of a non-intrusive translation system;
[0024] Figure 3 It is a flowchart of the first implementation method of the non-intrusive translation method of the present application;
[0025] Figure 4 It is a flowchart of the second implementation method of the non-intrusive translation method of the present application;
[0026] Figure 5 It is a flowchart of the third implementation method of the non-intrusive translation method of the present application;
[0027] Figure 6 is a flowchart of the fourth implementation method of the non-intrusive translation method of the present application;
[0028] Figure 7 This is a workflow diagram of an application scenario of the non-intrusive translation method of the present application;
[0029] Figure 8 It is a schematic diagram of a virtual translation web page generated by a PC based on a non-intrusive translation method;
[0030] Fig. 9 It is a schematic diagram of a virtual translation webpage generated by the APP based on the non-intrusive translation method;
[0031] Fig.10 This is a schematic diagram of the structure of an implementation method of the client of the present application;
[0032] Fig.11 It is a structural diagram of an implementation method of the server of the present application;
[0033] Fig.12 It is a structural schematic diagram of an embodiment of the electronic device of the present application;
[0034] Fig.13 It is a structural schematic diagram of an implementation method of a computer-readable storage medium of the present application. DETAILED DESCRIPTION
[0035] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of this application.
[0036] The terms used in the embodiments of the present application are only for the purpose of describing specific embodiments, and are not intended to limit the present application. The singular forms of "a", "said", and "the" used in the embodiments of the present application and the appended claims are also intended to include plural forms, unless otherwise clearly indicated above, and "multiple" generally includes at least two, but does not exclude the inclusion of at least one.
[0037] It should be understood that the term "and / or" used in this article is only a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone. In addition, the character " / " in this article generally indicates that the associated objects before and after are in an "or" relationship.
[0038] It should be understood that the terms "include", "comprises" or any other variations used herein are intended to cover non-exclusive inclusion, so that a process, method, article or device that includes a series of elements includes not only those elements, but also includes other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of more restrictions, the elements defined by the sentence "includes..." do not exclude the presence of other identical elements in the process, method, article or device that includes the elements.
[0039] This application first provides a non-intrusive translation system.
[0040] Specifically, see Figure 1 , Figure 1 It is a principle block diagram of an implementation of the non-intrusive translation system of the present application. In this implementation, the non-intrusive translation system 100 includes a client 101, a server 102 and a cloud platform 103 connected to each other.
[0041] In this embodiment, the client 101 refers to a client that has standard web page display capabilities and supports hypertext markup language, for example, any one or more of a personal computer (PC), a tablet computer (PAD), a mobile phone, a smart display, and an application (APP).
[0042] In this implementation, the server 102 supports online storage and database services, data processing and analysis, secure access control, etc. This application mainly relates to the database services and data processing functions of the cloud platform.
[0043] In some implementations, the server 102 includes a server engine and a server cache. The server engine is used to identify and generate images, and the server cache is used to obtain translation results.
[0044] In this embodiment, the cloud platform 103 is a translation platform based on cloud computing, which can provide multi-language translation and conversion capabilities. The cloud platform 103 is loosely coupled with the server 102, and provides a translation interface for the server 102. The server 102 calls the translation service of the cloud platform 103 through the translation interface. Among them, the cloud platform 103 can be a cloud platform provided by any manufacturer, and this application does not limit this.
[0045] In some implementations, multiple small language translation models are integrated on the cloud platform 103. In some specific implementations, the cloud platform 103 is integrated with translation interfaces of small AI (Artificial Intelligence) models in the field of multi-language translation, which can meet the translation needs of multiple languages (such as Chinese, English, Japanese, Russian, Arabic, etc.).
[0046] In some implementations, a large language model (LLM) is integrated on the cloud platform 103 to handle translation requirements of high standards or specific fields to ensure the accuracy of the translation.
[0047] See also Figure 2 , Figure 2 yes Figure 1 Signal flow diagram of an implementation method of a non-intrusive translation system. In this implementation method, the client 101 sends a picture link to the server 102, so that the server engine of the server 102 identifies the picture content based on the picture link, and sends the extracted text information to the client 101, so that the client 101 generates data to be translated based on the text information. The client 101 sends a translation request carrying the data to be translated to the server 102, and the server 102 queries the server cache based on the translation request. In response to finding the corresponding data to be restored in the server cache, the server 102 sends the data to be restored to the client 101. In response to not finding the corresponding data to be restored in the server cache, the server 102 sends the translation request to the cloud platform 103, so that the cloud platform 103 calls the corresponding model based on the translation request to translate the data to be translated, receives and stores the data to be restored returned by the cloud platform 103, and then sends the data to be restored to the client 101.
[0048] See also Figure 3 , Figure 3 1 is a flowchart of the first implementation of the non-intrusive translation method of the present application. In this implementation, the execution subject of the method is a client, and the method includes:
[0049] S11: The client parses the web page to be translated in response to the translation instruction, and constructs a virtual tree structure based on the obtained multiple elements to be translated.
[0050] In this implementation, the web page to be translated is an HTML (Hyper Text Markup Language) web page, and the HTML web page includes a plurality of web page elements and corresponding HTML tags.
[0051] In this implementation, the plurality of web page elements include web page elements that need to be translated and web page elements that do not need to be translated, and the elements to be translated are web page elements that need to be translated.
[0052] In this embodiment, the virtual tree structure is a DOM (Document Object Model) document object model constructed in a virtual web page independent of the web page to be translated. Adding or reducing elements to be translated in the virtual tree structure will not affect the web page to be translated.
[0053] It can be understood that by constructing a virtual tree structure, multiple elements to be translated can be accommodated in a virtual area independent of the web page to be translated, thereby reducing the impact on the original web page.
[0054] S12: Preprocess each element to be translated to generate a plurality of data to be translated; wherein each data to be translated includes a unique identifier and a text to be translated.
[0055] In this implementation, each element to be translated is preprocessed to extract the text to be translated in the element to be translated, and a unique identifier is generated based on the text to be translated.
[0056] In some implementations, the unique identifier is a hash value.
[0057] It is understandable that by converting the text to be translated into a hash value, the elements to be translated can be matched one-to-one with the data to be translated through the hash value, so as to ensure the accurate storage and reference of the relevant data in the virtual number structure. Furthermore, by retrieving the corresponding data to be translated through the hash value, the overall retrieval efficiency of the virtual tree structure can be improved.
[0058] S13: Acquire a plurality of data to be restored corresponding to the plurality of data to be translated; wherein each data to be restored includes a unique identifier and a translation text.
[0059] In this embodiment, the data to be restored corresponding to the data to be translated is first searched based on the local cache of the client, and in response to the existence of the corresponding data to be restored in the local cache of the client, the data to be restored is directly returned from the local cache.
[0060] Further, in response to the client not having the corresponding data to be restored in the local cache, a translation request carrying the data to be translated is sent to the server, so that the server searches the server cache based on the translation request. In response to finding the corresponding data to be restored in the server cache, the data to be restored is sent to the client through the server.
[0061] Furthermore, in response to not finding the corresponding data to be restored in the server cache, the server sends a translation request to the cloud platform, so that the cloud platform calls the corresponding model to translate the data to be translated based on the translation request, receives and stores the data to be restored returned by the cloud platform through the server, and then sends the data to be restored to the client through the server.
[0062] It can be understood that through the above-mentioned multi-stage chain processing, unnecessary real-time translation requests can be reduced, thereby improving the response speed and overall efficiency of translation.
[0063] It can be understood that by making the data to be restored include a unique identifier, the data to be translated and the data to be restored can be matched one-to-one through the unique identifier, so as to further improve the consistency and accuracy in the data processing process.
[0064] S14: Restoring each to-be-restored data into restored data.
[0065] In this implementation manner, each piece of data to be restored is adjusted accordingly to be processed into restored data.
[0066] S15: Based on the unique identifier corresponding to each restored data, render the multiple restored data into a virtual tree structure to generate a virtual translated web page.
[0067] It can be understood that rendering multiple restored data into a virtual tree structure based on a unique identifier corresponding to each restored data can improve the restoration accuracy of each element to be translated in the virtual translation web page.
[0068] In this implementation manner, the virtual translated web page is independent of the web page to be translated.
[0069] In some implementations, the virtually translated web page may be overlaid on the web page to be translated to present the translated web page to the user.
[0070] It can be understood that since the present application does not need to modify the source code of the web page, nor does it need to rely on specific browser plug-ins, it can reduce the degree of intrusive integration coupling, thereby protecting the privacy and security of the original web page, and ensuring the accuracy of restoring multiple elements to be translated in the original web page.
[0071] Different from the prior art, this embodiment parses the web page to be translated by the client responding to the translation instruction, and constructs a virtual tree structure based on the obtained multiple elements to be translated, and can accommodate multiple elements to be translated by a virtual area independent of the web page to be translated, thereby reducing the impact on the original web page. By preprocessing each element to be translated, generating multiple data to be translated including a unique identifier and a text to be translated, and obtaining multiple data to be restored including a unique identifier and a translation text corresponding to multiple data to be translated, the elements to be translated can be matched one by one with the data to be translated and the data to be restored through the unique identifier, so as to ensure the accurate storage and reference of the relevant data in the virtual number structure, thereby improving the consistency and accuracy of the data. Each data to be restored is restored to restored data, and multiple restored data are rendered into a virtual tree structure based on the unique identifier corresponding to each restored data, which can improve the restoration accuracy of each element to be translated in the virtual translation web page. This application can realize the non-intrusive translation and restoration of the original web page through the client script, thereby protecting the privacy and security of the original web page, and improving the accuracy of the restoration of multimodal elements in the virtual web page.
[0072] See also Figure 4 , Figure 4 1 is a flow chart of the second implementation of the non-intrusive translation method of the present application. In this implementation, the execution subject of the method is a client, and the method includes:
[0073] S21: The client dynamically identifies the web page to be translated in response to the translation instruction to obtain a set of web page elements; wherein the set of web page elements includes at least one type of web page element and a label corresponding to the web page element; the type includes at least one of a text type, an image type and a font package type.
[0074] In this embodiment, after receiving the translation instruction input by the user, the client dynamically identifies the web page to be translated based on the translation instruction to obtain the web page attributes of the web page to be translated and the multiple web page elements and corresponding tags included in the web page to be translated.
[0075] Dynamic identification refers to identifying web page elements in a web page line by line or page by page when the web page to be translated changes dynamically (for example, the page is scrolled down), so that the web page elements in the web page element set change dynamically.
[0076] The web page to be translated is an HTML (Hyper Text Markup Language) web page, and the HTML web page includes a plurality of web page elements and corresponding HTML tags.
[0077] Among them, web page attributes include the character encoding of the web page, the language identifier of the web page, the height of the visible area of the web page, the width of the visible area of the web page, the height of the full text of the web page, the width of the full text of the web page, and the vertical and horizontal distances of the current page relative to the upper left corner of the display area.
[0078] Among them, HTML tags include <title> Label,< / title> <meta> Label, Label, Label, <href>Label, <input> Label, <textarea> Label,< / textarea> <select>Tags and <option> At least one type of tag.
[0079] Specifically, a tag is a title that defines a document and is displayed on the title bar or tab page of a browser. It should only contain text. If it contains a tag, any tags it contains will be ignored. A tag is an auxiliary tag that is located at the head of a document and does not contain any content. A tag is an inline container that is used to mark a part of a text or a part of a document. Tags are used to embed images in HTML pages. Technically speaking, the image is not actually inserted into the web page, but the image is linked to the web page, that is, the tag creates a container for referencing the image. A tag is a URL (Uniform Resource Locator) that specifies the target of a hyperlink.< / option> < / select> <textarea> The tag is used to define a multi-line text input control. Text of any length can be entered in the text input field.< / textarea> <select>A label represents a control that provides a menu of options. <option>The tag defines the options in the selection list.
[0080] In this embodiment, when the client first loads the web page to be translated, it automatically identifies the language identifier of the web page based on the web page attributes of the web page, uses the language corresponding to the language identifier as the source language, and stores the language identifier in the local cache (Local Storage) of the client. When the web page is subsequently translated, the source language is always used as the reference language to be translated to avoid the problem of reduced translation quality caused by continuous translation between multiple languages.
[0081] In some embodiments, the translation instruction carries a target language identifier specified by the user, and subsequent translations translate the elements to be translated in the target language corresponding to the target language identifier. In other embodiments, the translation instruction does not carry the target language identifier specified by the user, and the system language of the user device where the client is located is used as the target language. In some other embodiments, if the system language cannot be identified, the remaining languages recently used and stored in the local cache are used as the target language, and this application does not limit this.
[0082] Wherein, the source language and the target language can be the languages corresponding to any country, such as Chinese, English, Japanese, Russian, Arabic, etc., which is not limited in this application.
[0083] S22: Traverse the web page element set through the label to build a virtual tree structure based on the identified multiple elements to be translated.
[0084] In this embodiment, whileNodes (loop statement) is used to traverse each label in the web page element set to identify multiple elements to be translated and multiple web page elements that do not need to be translated.
[0085] Wherein, the web page elements that do not need to be translated include web page elements marked as web page annotation content, web page elements whose language identifier in the label information is consistent with the target language identifier, and web page elements marked with an identifier that does not need to be translated.
[0086] It can be understood that by traversing the web page element set through the label to identify multiple elements to be translated, multiple web page elements that do not need to be translated can be removed to reduce subsequent translation tasks and restoration tasks, thereby further improving the response speed and overall efficiency of translation.
[0087] In this embodiment, the JS (JavaScript) interface MutationObserver is used to filter the elements to be translated and add them to the virtual tree structure.
[0088] MutationObserver is a JavaScript API (Application Programming Interface) used to monitor changes in the DOM tree. It provides an asynchronous way to monitor operations such as adding, deleting, and changing attributes of DOM elements, as well as modifications to text nodes. Through MutationObserver, changes in the DOM can be captured in real time and corresponding responses can be made.
[0089] It can be understood that since the web page element set is dynamically changing, the multiple elements to be translated identified based on this and the constructed virtual tree structure are also dynamically changing, which is convenient for real-time processing of newly added data and is conducive to improving the accuracy and timeliness of web page translation.
[0090] S23: Classify the elements to be translated according to the labels, and perform corresponding preprocessing on each element to be translated based on the classification results to generate multiple data to be translated including unique identifiers and translation texts.
[0091] In this embodiment, the type corresponding to each element to be translated is first determined according to the label.
[0092] In some embodiments, if the label corresponding to the element to be translated is a label,.< / option> < / select> <textarea> In some other implementations, if the label corresponding to the element to be translated is< / textarea> label, the classification result is that the element to be translated is of the image type. In some other implementations, if the label corresponding to the element to be translated is tag, the classification result is that the element to be translated is of font package type.
[0093] In this implementation, after the type of each element to be translated is determined, corresponding preprocessing is performed based on its corresponding type.
[0094] In some implementations, in response to the element to be translated being of text type, the text information to be translated in the element to be translated is extracted to generate the text to be translated.
[0095] Specifically, the client performs cleaning operations such as special symbol filtering, word segmentation, and sentence segmentation on the text information to be translated to generate the text to be translated.
[0096] In some other implementations, in response to the element to be translated being of image type, the element to be translated is sent to a server, so that the server performs image recognition on the element to be translated and generates text to be translated based on text information returned by the server.
[0097] Specifically, the client sends the image link corresponding to the element to be translated to the server engine, so that the server engine recognizes the corresponding image content in the image link based on OCR (Optical Character Recognition) technology to extract the text information to be translated, and returns the text information to be translated to the client.
[0098] In some other embodiments, in response to the element to be translated being a font package type, character recognition and calculation are performed on the element to be translated, and the element to be translated is converted into an image based on the recognition result and the calculation result. The image is sent to a server so that the server performs image recognition on the image and generates a text to be translated based on the text information returned by the server.
[0099] Specifically, the client performs character recognition on the elements to be translated through Unicode (character set), and then calculates the rectangular area occupied by the HTML tag based on the coordinates x and y of the HTML tag corresponding to the element to be translated, and then converts the elements to be translated of the font package type into images based on the character recognition results and the rectangular area through canvas drawing, and sends the images to the server engine.
[0100] It can be understood that this embodiment implements differentiated preprocessing for elements to be translated of different modalities (types), which can fully consider the differences between different types of web page elements to effectively extract the text information included in the elements to be translated, avoid semantic deviations in subsequent translation, and thus improve the accuracy of translation.
[0101] Furthermore, a corresponding unique identifier (hash value) is generated based on each text to be translated, and data to be translated is generated according to the unique identifier, the type corresponding to the tag, the type corresponding to the element to be translated, and the text to be translated.
[0102] If the element to be translated is of image type or font package type, the data to be translated also includes an image position identifier and a container position size, and the image position identifier and container size information are represented by coordinate x, coordinate y, width (width) and height (height).
[0103] Specifically, the definition structure of the data to be translated is as follows:
[0104] [hash_id, flag_type, word_type, content, path, position_size].
[0105] Among them, hash_id refers to the unique identifier (hash value); flag_type refers to the tag type; word_type refers to the type of the element to be translated; content refers to the content corresponding to the text to be translated; path refers to the image location identifier; position_size refers to the container position size.
[0106] The following is an example structure of the data to be translated:
[0107]
[0108]
[0109] Among them, json (JavaScript Object Notation) refers to a lightweight data exchange format, and each array item represents a data to be translated, where the first element is a unique identifier (hash_id), the second element is the tag type, the third element is the type corresponding to the element to be translated, the fourth element is the content corresponding to the text to be translated, the fifth element (if any) indicates the image position, and the sixth element (if any) describes the position and size of the container.
[0110] In this implementation, a queue of data to be translated is constructed based on a plurality of data to be translated.
[0111] Furthermore, in this embodiment, in response to the element to be translated being a picture type or a font package type, and the server has not extracted text information from the picture corresponding to the element to be translated, each element to be translated from which text information has not been extracted is merged into the same sprite image, and the coordinates of each element to be translated in the sprite image are recorded in a virtual tree structure.
[0112] Among them, due to potential problems such as cross-domain and canvas pollution, the pictures are converted into base64 (based on 64 printable characters to represent binary data) encoding before being merged.
[0113] S24: Send multiple translation requests carrying data to be translated to the first-level cache of the client, or send multiple translation requests to the first-level cache and send at least part of the translation requests to the server through the first-level cache to receive multiple data to be restored including unique identifiers and translation texts returned by the first-level cache or the server.
[0114] In this embodiment, multiple translation requests are first sent to the first-level cache of the client in sequence, and based on each translation request, it is queried in the first-level cache whether there is data to be restored corresponding to the data to be translated. In response to the existence of data to be restored corresponding to the data to be translated in the first-level cache, the data to be restored returned by the first-level cache is received.
[0115] The first-level cache is the local cache mentioned above, and the first-level cache stores the historical translation data returned from the second-level cache.
[0116] Specifically, if the same client has previously translated the web page to be translated based on the same target language, the corresponding translation data (including translation requests, data to be translated, and data to be restored) is stored in the client's first-level cache. By querying based on the translation request, the corresponding data to be restored can be quickly hit without continuing to send translation requests to the server, thereby reducing unnecessary real-time translation requests and subsequently improving the response speed and overall efficiency of the translation.
[0117] In this embodiment, in response to the absence of data to be restored corresponding to the data to be translated in the first-level cache, a translation request is sent to the second-level cache of the server through the first-level cache, and the data to be restored returned by the second-level cache is received.
[0118] The second-level cache is the server cache mentioned above, and the second-level cache stores historical translation requests sent by the first-level cache and translation data received from the cloud platform.
[0119] In some implementations, in response to the existence of data to be restored corresponding to the data to be translated in the secondary cache, the data to be restored is sent to the primary cache of the client.
[0120] In some other embodiments, in response to the absence of data to be restored corresponding to the data to be translated in the secondary cache, a translation request is sent to the cloud platform, so that the cloud platform translates the data to be translated based on the translation request. The data to be restored returned by the cloud platform is received and stored through the secondary cache, and the data to be restored is sent to the primary cache of the client.
[0121] In this implementation, the data to be restored includes a unique identifier, a type corresponding to a tag, a type corresponding to an element to be translated, a target language identifier, a translation text, and location information.
[0122] Specifically, the definition structure of the data to be restored is as follows:
[0123] [hash_id, flag_type, word_type, languages, content, path, position_size].
[0124] Among them, hash_id refers to the unique identifier (hash value); flag_type refers to the tag type; word_type refers to the type of the element to be translated; languages refers to the target language; content refers to the content corresponding to the translated text; path refers to the image location identifier; position_size refers to the container position size.
[0125] The following is an example structure of the data to be translated:
[0126]
[0127]
[0128] Among them, json refers to a lightweight data exchange format, and each array item represents a data to be restored, where the first element is a unique identifier (hash_id), the second element is the tag type, the third element is the type corresponding to the element to be translated, the fourth element is the target language, the fifth element is the content corresponding to the translated text, the sixth element (if exists) indicates the image position, and the seventh element (if exists) describes the position and size of the container.
[0129] In this implementation, if the element to be translated is of a picture type or a font package type, the server engine regenerates a picture that conforms to the original picture format based on the translated text obtained in the secondary cache, and generates data to be restored based on the picture.
[0130] Furthermore, in this implementation, a queue of data to be restored is constructed based on a plurality of data to be restored.
[0131] It can be understood that, through the unique identifier corresponding to each to-be-translated element, each to-be-translated data in the to-be-translated data queue can be matched one-to-one with the to-be-restored data in the to-be-restored data queue.
[0132] S25: Restoring each to-be-restored data into restored data of a corresponding type based on the classification result.
[0133] In this embodiment, the data to be restored in the data queue to be restored are first input into the multimodal restoration software package in sequence according to the first-in-first-out principle. The multimodal restoration software package encapsulates a variety of restoration algorithms for processing multimodal data, including a text restoration algorithm, an image restoration algorithm, and a font package restoration algorithm.
[0134] The multimodal restoration software package encapsulates multiple restoration processing chains, each of which corresponds to a restoration algorithm for processing corresponding types of elements to be translated.
[0135] In this embodiment, the multimodal restoration software package only encapsulates multiple restoration processing chains for processing web page elements. In other embodiments, the restoration processing chains for more modes such as video and audio can be expanded and encapsulated, and this application does not limit this.
[0136] Furthermore, based on the type of the data to be restored, a corresponding restoration algorithm is called to restore the data to be restored, so as to output restored data.
[0137] In some implementations, in response to the data to be restored being of text type, a text restoration algorithm is called to process the translation text in the data to be restored to avoid overflow of the translation text.
[0138] In some specific implementations, if the translation text is longer than the text to be translated and the style is distorted, the text restoration algorithm processes the text to be translated according to different adjustment priorities. The adjustment priorities are as follows: abbreviation of professional terms > adjusting text size > displaying partial text information.
[0139] The above adjustment priority is explained by taking an example. For example, the English translation of central processing unit is Central Processing Unit, and CPU is preferred to be abbreviated. If the translation text is without abbreviation, the text size can be reduced. If the translation text is longer and the omission of part of the text does not affect the overall expression, only part of the text information can be displayed.
[0140] In some other implementations, in response to the data to be restored being of a picture type, a picture restoration algorithm is called to scale the picture generated based on the translated text in the data to be restored.
[0141] In some specific implementations, if the translated text is longer than the text to be translated and its style is distorted, the size of the new image generated by the server engine based on the translated text is larger than the size of the original image. At this time, the new image is scaled accordingly through an image restoration algorithm to avoid the new image obscuring other web page elements.
[0142] In some other implementations, in response to the data to be restored being of a font package type, a font package restoration algorithm is called to restore images generated based on the translated text in the data to be restored based on a calculation result corresponding to the font package type.
[0143] In some specific implementations, if the translated text is longer than the text to be translated and the style is deformed, the size of the new image generated by the server engine based on the translated text is larger than the size of the original image. At this time, the font package restoration algorithm is used to call the rectangular area calculated based on the HTML tag when the font package was previously converted, and the new image is restored to the size corresponding to the rectangular area through canvas drawing to avoid the new image obstructing other web page elements.
[0144] It can be understood that this embodiment performs different types of restoration for different types of data to be restored, which can reduce the restoration differences between different types of web page elements to avoid problems such as inconsistent fonts and disordered layout, thereby improving the restoration effect of subsequent layout styles.
[0145] S26: Based on the unique identifier corresponding to each restored data, render the multiple restored data into a virtual tree structure to generate a virtual translated web page.
[0146] In this implementation, in order to avoid monitoring all element changes and being built into a virtual tree structure, the DOM to be replaced is cached into a "temporary process" object during translation and rendering. For these DOMs in rendering, even if MutationObserver monitors them, it is necessary to filter out these DOMs and not perform translation and restoration.
[0147] In this embodiment, in addition to rendering and restoring multiple restored data with translated texts, each element to be translated in the Sprite image is also split, and each element to be translated in the Sprite image is rendered into a virtual translation web page based on the corresponding coordinates to further improve the accuracy of style restoration.
[0148] Different from the prior art, this embodiment parses the web page to be translated by the client in response to the translation instruction, and constructs a virtual tree structure based on the obtained multiple elements to be translated, which can accommodate multiple elements to be translated through a virtual area independent of the web page to be translated, thereby reducing the impact on the original web page. Further, differentiated preprocessing is implemented for different types of elements to be translated, and multiple data to be translated including unique identifiers and texts to be translated are generated, which can fully consider the differences between different types of web page elements, so as to effectively extract the text information included in the elements to be translated, and avoid semantic deviations in subsequent translation, thereby improving the accuracy of translation. Further, multiple data to be restored including unique identifiers and translation texts corresponding to multiple data to be translated are obtained, and the elements to be translated, the data to be translated, and the data to be restored can be matched one by one through the unique identifier to ensure the accurate storage and reference of the relevant data in the virtual number structure, thereby improving the consistency and accuracy of the data. Further, different types of restoration are performed for different types of data to be restored, which can reduce the restoration differences between different types of web page elements, so as to avoid problems such as inconsistent fonts and disordered typesetting, thereby subsequently improving the restoration effect of the layout style. Furthermore, each to-be-restored data is restored into restored data, and multiple restored data are rendered into a virtual tree structure based on a unique identifier corresponding to each restored data, which can improve the restoration accuracy of each to-be-translated element in the virtual translated web page. The present application can achieve non-intrusive translation and restoration of the original web page through client scripts, thereby protecting the privacy and security of the original web page and improving the accuracy of restoration of multimodal elements in the virtual web page.
[0149] See also Figure 5 , Figure 5 1 is a flow chart of the third implementation method of the non-intrusive translation method of the present application. In this implementation method, the execution subject of the method is a server, and the method includes:
[0150] S31: The server receives multiple translation requests carrying data to be translated sent by the client; wherein the data to be translated is generated by the client based on multiple elements to be translated identified by the web page to be translated, and each data to be translated includes a unique identifier and a text to be translated.
[0151] The specific process of identifying the elements to be translated and the specific process of generating the data to be translated can be found in the descriptions of S11 to S12 and S21 to S23, which will not be described in detail here.
[0152] S32: Based on the multiple translation requests, a plurality of data to be restored corresponding to the multiple data to be translated is obtained; wherein each data to be restored includes a unique identifier and a translation text.
[0153] In this implementation manner, the server obtains a plurality of data to be restored corresponding to a plurality of data to be translated based on the secondary cache and the cloud platform.
[0154] S33: Sending multiple data to be restored to the client, so that the client restores each data to be restored into restored data, and rendering the multiple restored data into a virtual tree structure constructed by the client based on multiple elements to be translated according to a unique identifier corresponding to each restored data, so as to generate a virtual translated web page.
[0155] The specific restoration process of the data to be restored and the specific rendering process of the restored data can be found in the descriptions of S14 to S15 and S25 to S26, which will not be described in detail here.
[0156] Different from the prior art, this implementation method caches the translation requests sent by the client through the server, and sends the translated data to be restored to the client, which can store a large amount of translation data and facilitate the timely return of the second hit data to be restored, thereby improving the efficiency of translation implementation.
[0157] See also Figure 6 , Figure 6 1 is a flowchart of the fourth implementation method of the non-intrusive translation method of the present application. In this implementation method, the execution subject of the method is a server, and the method includes:
[0158] S41: Establishing a correction subject library in the server; wherein the correction subject library stores a plurality of professional fields and corresponding professional terms and translation knowledge in the professional fields, and each professional field carries at least one classification label.
[0159] In this implementation, the correction theme library adopts a three-level theme design, and the code for each level is a 3-digit standard compilation, which supports expansion at the same level.
[0160] In some embodiments, the correction subject library includes 5 primary categories, 22 secondary categories, and 94 tertiary categories.
[0161] In some specific ways, the first-level codes are C001, C002, C003, C004, and C005, and the corresponding classification labels (or first-level names) are living, education, travel, work, business, and medical health. The second-level codes corresponding to the first-level code C001 are C001001, C001002, C001003, etc., and the corresponding classification labels (or second-level names) are life shopping, transportation, hotel reservation, etc. The third-level codes corresponding to the second-level code C001001 are C001001001, C001001002, C001001003, etc., and the corresponding classification labels (or third-level names) are online shopping platforms, quality shopping malls, surrounding supermarkets, etc.
[0162] S42: The server receives multiple translation requests carrying data to be translated sent by the client; wherein the data to be translated is generated by the client based on multiple elements to be translated identified by the web page to be translated, and each data to be translated includes a unique identifier and a text to be translated.
[0163] S43: Based on each text to be translated, the corresponding translation request is stored in the most matching professional field.
[0164] In this embodiment, a query is performed in the correction subject library based on the text to be translated to determine the professional terms that have the highest match with the text to be translated and the classification labels of the professional fields corresponding to the translation knowledge, and the translation requests are classified based on the classification labels with the highest match, and the translation requests are stored in the most matching professional fields.
[0165] In this embodiment, corresponding long translations and abbreviated translations (abbreviations) are also stored for professional terms corresponding to each professional field, so as to more accurately classify translation requests.
[0166] S44: Based on each translation request, query whether there is data to be restored corresponding to the data to be translated in the secondary cache of the server.
[0167] S45: In response to the existence of data to be restored corresponding to the data to be translated in the secondary cache, the data to be restored is sent to the primary cache of the client.
[0168] In some specific implementations, if client A has previously translated the web page to be translated based on the same target language, the corresponding translation data is stored in the secondary cache of the server. Client B (two different clients from client A) then requests to translate the web page to be translated based on the same target language. After receiving the translation request, the secondary cache of the server can query the stored historical translation requests to quickly hit the corresponding data to be restored without continuing to send translation requests to the cloud platform, thereby reducing unnecessary real-time translation requests, and further improving the response speed and overall efficiency of the translation.
[0169] S46: In response to the absence of the to-be-translated data corresponding to the to-be-translated data in the secondary cache, sending a translation request to the cloud platform, so that the cloud platform translates the to-be-translated data based on the translation request.
[0170] In this embodiment, after receiving a translation request, the cloud platform can call a translation interface to translate the translation data using a small translation model corresponding to the target language, or call a large language model to translate the translation data using the large language model.
[0171] If you want to call a large language model for translation, you need to build a large language model prompt word project in the server in advance.
[0172] The following is an example structure for building a large language model prompt word project:
[0173] #Limited Roles
[0174] You are now a multilingual translation expert assistant
[0175] #Assigning tasks
[0176] Please translate according to the current input data items. The data item identifiers are described as follows:
[0177] #Description content
[0178] Translation source language: {src_type}
[0179] Translation target language: {tgt_type}
[0180] Translation source content: {src_content}
[0181] Translation target content: {tgt_content}
[0182] #Construct one-shot
[0183] Please study and understand the following carefully:
[0184] You are a multilingual translation expert assistant. Now you need to translate this {src_type} content {src_content} into {tgt_type} content. The translated content is the data item {tgt_content}
[0185] #Enter the specific content to be translated
[0186] The input data items are as follows:
[0187] Translation source language: {Chinese}
[0188] Translation target language: {English}
[0189] Translation source content: {Technology is top-notch, products are local}
[0190] Translation target content: {tgt_content}
[0191] #Confirmation and task execution
[0192] Do you understand the above content? If you do, please start translating:
[0193] #Large Model Language LLM Translation Result Example
[0194] {Technology soars to the sky; products stand firm on the ground.}
[0195] S47: Receive and store the data to be restored returned by the cloud platform through the secondary cache, and send the data to be restored to the primary cache of the client.
[0196] S48: The data to be restored is stored in the most matching professional field.
[0197] In this implementation, the data to be translated and the data to be restored corresponding to each translation request are stored in the most matching professional field to facilitate subsequent secondary hits, thereby improving the overall translation speed.
[0198] S49: Manually correct the translation data stored in the correction subject library, and synchronously store the manually corrected translation data in the secondary cache of the server.
[0199] In this implementation, manual correction processing may be performed on the translation data stored in the correction subject library based on the user's correction instruction or periodic self-checking.
[0200] It can be understood that by manually correcting the translation data stored in the correction subject library, the translation result can be made more accurate.
[0201] It can be understood that the manually corrected translation data is synchronously stored in the secondary cache of the server, so that after receiving the same translation request in the future, a more accurate translation result can be quickly returned to the client, thereby improving the accuracy and timeliness of the translation, and then realizing the efficient operation of multi-level translation.
[0202] See also Figure 7 , Figure 7 It is a workflow diagram of an application scenario of the non-intrusive translation method of the present application. In this embodiment, the client responds to the translation instruction to parse the web page to be translated, and constructs a virtual tree structure based on the obtained multiple elements to be translated. Then, the types of elements to be translated are classified according to the tags, and each element to be translated is pre-processed accordingly based on the classification results. In response to the element to be translated being a text type, the text information to be translated in the element to be translated is extracted to generate the text to be translated. In response to the element to be translated being a picture type, the element to be translated is sent to the server so that the server performs picture recognition on the element to be translated, and generates the text to be translated based on the text information returned by the server. In response to the element to be translated being a font package type, the element to be translated is subjected to character recognition and calculation, and the element to be translated is converted into a picture based on the recognition result and the calculation result, and then the picture is sent to the server so that the server performs picture recognition on the picture, and generates the text to be translated based on the text information returned by the server. Generate multiple data to be translated including unique identifiers and translation texts, and construct a data queue to be translated based on multiple data to be translated. Multiple translation requests carrying data to be translated are sent to the first-level cache in sequence, and at least part of the translation requests are sent to the second-level cache of the server through the first-level cache to receive multiple data to be restored including unique identifiers and translation texts returned by the first-level cache or the second-level cache. Then, a queue of data to be restored is constructed based on the multiple data to be restored, and each data to be restored in the queue of data to be restored is restored into corresponding types of restored data in sequence based on the classification result. Finally, based on the unique identifier corresponding to each restored data, the multiple restored data are rendered into a virtual tree structure to generate a virtual translation web page.
[0203] Specifically, see Figure 8 and Fig. 9 , Figure 8 This is a schematic diagram of a virtual translation web page generated by a PC based on a non-intrusive translation method. Fig. 9 It is a schematic diagram of a virtual translation web page generated by the APP based on the non-intrusive translation method.
[0204] Different from the prior art, this embodiment parses the web page to be translated by the client in response to the translation instruction, and constructs a virtual tree structure based on the obtained multiple elements to be translated, and can accommodate multiple elements to be translated through a virtual area independent of the web page to be translated, thereby reducing the impact on the original web page. By preprocessing each element to be translated, generating multiple data to be translated including a unique identifier and a text to be translated, and obtaining multiple data to be restored including a unique identifier and a translation text corresponding to the multiple data to be translated, the elements to be translated can be matched one by one with the data to be translated and the data to be restored through the unique identifier, so as to ensure the accurate storage and reference of the relevant data in the virtual number structure, thereby improving the consistency and accuracy of the data. Each data to be restored is restored into restored data, and multiple restored data are rendered into the virtual tree structure based on the unique identifier corresponding to each restored data, which can improve the restoration accuracy of each element to be translated in the virtual translation web page.
[0205] Correspondingly, the present application provides related devices of the non-intrusive translation method.
[0206] See also Fig.10 , Fig.10 It is a schematic diagram of the structure of an implementation of the client of the present application. In this implementation, the client 50 includes a construction module 51, a pre-processing module 52, a first acquisition module 53, a restoration module 54 and a generation module 55.
[0207] The construction module 51 is used to parse the web page to be translated in response to the translation instruction, and to construct a virtual tree structure based on the obtained multiple elements to be translated.
[0208] The preprocessing module 52 is used to preprocess each element to be translated to generate a plurality of data to be translated, wherein each data to be translated includes a unique identifier and a text to be translated.
[0209] The first acquisition module 53 is used to acquire a plurality of data to be restored corresponding to the plurality of data to be translated, wherein each data to be restored includes a unique identifier and a translation text.
[0210] The restoration module 54 is used to restore each to-be-restored data into restored data.
[0211] The generating module 55 is used to render the plurality of restored data into a virtual tree structure based on the unique identifier corresponding to each restored data, so as to generate a virtual translated web page.
[0212] For the specific process, please refer to the relevant text descriptions in S11 to S15 and S21 to S26, which will not be repeated here.
[0213] See also Fig.11 , Fig.11 It is a schematic structural diagram of an embodiment of the server of the present application. In this embodiment, the server 60 includes a receiving module 61, a second obtaining module 62, and a sending module 63.
[0214] The receiving module 61 is configured to receive multiple translation requests carrying data to be translated sent by the client. Among them, the data to be translated is generated by the client based on multiple elements to be translated recognized from the web page to be translated, and each data to be translated includes a unique identifier and the text to be translated.
[0215] The second obtaining module 62 is configured to obtain multiple data to be restored corresponding to the multiple data to be translated based on the multiple translation requests. Among them, each data to be restored includes a unique identifier and the translated text.
[0216] The sending module 63 is configured to send the multiple data to be restored to the client, so that the client restores each data to be restored into restored data, and based on the unique identifier corresponding to each restored data, renders the multiple restored data into the virtual tree structure constructed by the client based on the multiple elements to be translated, so as to generate a virtual translated web page.
[0217] Among them, for the specific process, please refer to the relevant text descriptions in S31 - S33 and S41 - S49, which will not be elaborated here.
[0218] Different from the prior art, in the present application, the client 50 responds to the translation instruction to parse the web page to be translated, and constructs a virtual tree structure based on the obtained multiple elements to be translated, and can accommodate multiple elements to be translated through a virtual area independent of the web page to be translated, thereby reducing the impact on the original web page. By preprocessing each element to be translated, generating multiple data to be translated including a unique identifier and the text to be translated, and obtaining multiple data to be restored including a unique identifier and the translated text corresponding to the multiple data to be translated based on the client 50 and the server 60, the element to be translated can be in one-to-one correspondence with the data to be translated and the data to be restored through the unique identifier, so as to ensure the accurate storage and reference of relevant data in the virtual number structure, thereby improving the consistency and accuracy of the data. Restoring each data to be restored into restored data, and rendering the multiple restored data into the virtual tree structure based on the unique identifier corresponding to each restored data can improve the restoration accuracy of each element to be translated in the virtual translated web page. The present application can achieve non-invasive translation and restoration of the original web page through the script of the client 50, thereby protecting the privacy and security of the original web page and improving the restoration accuracy of multi-modal elements in the virtual web page.
[0219] Correspondingly, the present application provides an electronic device.
[0220] Please refer to Fig.12 , Fig.12 Schematic diagram of the structure of an electronic device of the present application. Fig.12 As shown, in this embodiment, the electronic device 70 includes a memory 71 and a processor 72 .
[0221] In this embodiment, the memory 71 is used to store program data, and when the program data is executed, the steps in the non-intrusive translation method described above are implemented. The processor 72 is used to execute the program instructions stored in the memory 71 to implement the steps in the non-intrusive translation method described above.
[0222] Specifically, the processor 72 is used to control itself and the memory 71 to implement the steps in the non-intrusive translation method as described above. The processor 72 can also be called a CPU (Central Processing Unit). The processor 72 may be an integrated circuit chip with signal processing capabilities. The processor 72 can also be a general-purpose processor, a digital signal processor (Digital Signal Processor, DSP), an application-specific integrated circuit (Application Specific Integrated Circuit, ASIC), a field-programmable gate array (Field-Programmable Gate Array, FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc. In addition, the processor 72 can be implemented by multiple integrated circuit chips.
[0223] Different from the prior art, the present embodiment parses the web page to be translated by the processor 72 in response to the translation instruction, and constructs a virtual tree structure based on the obtained multiple elements to be translated, and can accommodate multiple elements to be translated by a virtual area independent of the web page to be translated, thereby reducing the impact on the original web page. By pre-processing each element to be translated, multiple data to be translated including a unique identifier and a text to be translated are generated, and multiple data to be restored including a unique identifier and a translation text corresponding to multiple data to be translated are obtained, and the elements to be translated can be corresponded to the data to be translated and the data to be restored one by one through the unique identifier, so as to ensure the accurate storage and reference of the relevant data in the virtual number structure, thereby improving the consistency and accuracy of the data. Each data to be restored is restored to restored data, and multiple restored data are rendered into the virtual tree structure based on the unique identifier corresponding to each restored data, so as to improve the restoration accuracy of each element to be translated in the virtual translation web page. The present application can realize the non-intrusive translation and restoration of the original web page through the client script, thereby protecting the privacy and security of the original web page, and improving the accuracy of the restoration of multimodal elements in the virtual web page.
[0224] Correspondingly, the present application provides a computer-readable storage medium.
[0225] See also Fig.13 , Fig.13 It is a structural schematic diagram of an implementation method of a computer-readable storage medium of the present application.
[0226] The computer-readable storage medium 80 includes a computer program 801 stored on the computer-readable storage medium 80, and the computer program 801 implements the steps in the non-intrusive translation method as described above when executed by the above-mentioned processor. Specifically, if the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium 80. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a computer-readable storage medium 80, including several instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor (processor) to perform all or part of the steps of each implementation method of the present application. The aforementioned computer-readable storage medium 80 includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk and other media that can store program codes.
[0227] In the several embodiments provided in the present application, it should be understood that the disclosed methods and devices can be implemented in other ways. For example, the device implementation described above is only schematic. For example, the division of modules or units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, and the indirect coupling or communication connection of devices or units can be electrical, mechanical or other forms.
[0228] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the present embodiment.
[0229] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.
[0230] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions to enable a computer device (which can be a personal computer, server, or network device, etc.) or a processor (processor) to perform all or part of the steps of each implementation method of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk and other media that can store program code.
[0231] The above description is only an implementation method of the present application, and does not limit the patent scope of the present application. Any equivalent structure or equivalent process transformation made using the contents of the present application specification and drawings, or directly or indirectly used in other related technical fields, are also included in the patent protection scope of the present application. < / href>
Claims
1. A non-intrusive translation method, characterized in that: include: The client parses the web page to be translated in response to the translation instruction, and constructs a virtual tree structure based on the obtained multiple elements to be translated; Preprocessing each of the elements to be translated to generate a plurality of data to be translated; wherein each data to be translated includes a unique identifier and a text to be translated; Acquire a plurality of data to be restored corresponding to the plurality of data to be translated; wherein each of the data to be restored includes the unique identifier and the translation text; Restoring each of the to-be-restored data into restored data; Based on the unique identifier corresponding to each of the restored data, a plurality of the restored data are rendered into the virtual tree structure to generate a virtual translated web page.
2. The non-intrusive translation method according to claim 1, characterized in that: The client parses the web page to be translated in response to the translation instruction, and constructs a virtual tree structure based on the obtained multiple elements to be translated, including: The client dynamically identifies the web page to be translated in response to the translation instruction to obtain a set of web page elements; wherein the set of web page elements includes at least one type of web page element and a label corresponding to the web page element; the type includes at least one of a text type, an image type and a font package type; Traversing the web page element set by using the tags to construct the virtual tree structure based on the identified multiple elements to be translated; The step of preprocessing each of the elements to be translated to generate a plurality of data to be translated comprises: Classifying the elements to be translated according to the labels, and performing corresponding preprocessing on each element to be translated based on the classification result, so as to generate a plurality of data to be translated including the unique identifier and the translation text; The step of obtaining a plurality of data to be restored corresponding to the plurality of data to be translated comprises: Sending multiple translation requests carrying the data to be translated to the first-level cache of the client, or sending multiple translation requests to the first-level cache and sending at least part of the translation requests to the server through the first-level cache, so as to receive multiple data to be restored including the unique identifier and the translation text returned by the first-level cache or the server; The step of restoring each of the to-be-restored data into restored data comprises: Based on the classification result, each of the to-be-restored data is restored into restored data of a corresponding type.
3. The non-intrusive translation method according to claim 2, characterized in that: The step of classifying the elements to be translated according to the labels and performing corresponding preprocessing on each element to be translated based on the classification result to generate a plurality of data to be translated including the unique identifier and the translation text comprises: Determining the type corresponding to each of the elements to be translated according to the label; In response to the element to be translated being of the text type, extracting text information to be translated from the element to be translated to generate the text to be translated; In response to the element to be translated being of the image type, sending the element to be translated to the server, so that the server performs image recognition on the element to be translated, and generates the text to be translated based on the text information returned by the server; In response to the element to be translated being of the font package type, character recognition and calculation are performed on the element to be translated, and the element to be translated is converted into an image based on the recognition result and the calculation result; the image is sent to the server so that the server performs image recognition on the image, and the text to be translated is generated based on the text information returned by the server; Generate the corresponding unique identifier based on each of the to-be-translated texts, and generate the to-be-translated data according to the unique identifier, the type corresponding to the tag, the type corresponding to the to-be-translated element, and the to-be-translated text; A queue of data to be translated is constructed based on the plurality of data to be translated.
4. The non-intrusive translation method according to claim 3, characterized in that: The step of sending a plurality of translation requests carrying the data to be translated to the first-level cache of the client, or sending a plurality of translation requests to the first-level cache and sending at least part of the translation requests to the server through the first-level cache to receive a plurality of data to be restored including the unique identifier and the translation text returned by the first-level cache or the server, comprises: Sending the plurality of translation requests sequentially to the first-level cache of the client; Based on each of the translation requests, querying in the first-level cache whether there is the data to be restored corresponding to the data to be translated; In response to the existence of the to-be-restored data corresponding to the to-be-translated data in the first-level cache, receiving the to-be-restored data returned by the first-level cache; In response to the absence of the to-be-restored data corresponding to the to-be-translated data in the first-level cache, sending the translation request to the second-level cache of the server through the first-level cache, and receiving the to-be-restored data returned by the second-level cache; A queue of data to be restored is constructed based on the plurality of data to be restored; wherein the data to be restored includes the unique identifier, the type corresponding to the tag, the type corresponding to the element to be translated, the target language identifier, the translated text and location information.
5. The non-intrusive translation method according to claim 4, characterized in that: The step of restoring each of the to-be-restored data into restored data of a corresponding type based on the classification result comprises: The data to be restored in the data queue to be restored are sequentially input into the multimodal restoration software package according to the first-in-first-out principle; wherein the multimodal restoration software package encapsulates a plurality of restoration algorithms for processing multimodal data, and the plurality of restoration algorithms include a text restoration algorithm, an image restoration algorithm, and a font package restoration algorithm; The corresponding restoration algorithm is called based on the type of the data to be restored to restore the data to be restored, so as to output the restored data.
6. The non-intrusive translation method according to claim 5, characterized in that: The step of invoking the corresponding restoration algorithm based on the type of the data to be restored to restore the data to be restored so as to output the restored data includes: In response to the data to be restored being of the text type, calling the text restoration algorithm to process the translation text in the data to be restored to avoid overflow of the translation text; In response to the data to be restored being of the image type, calling the image restoration algorithm to scale the image generated based on the translated text in the data to be restored; In response to the data to be restored being of the font package type, the font package restoration algorithm is called to restore the image generated based on the translated text in the data to be restored based on the calculation result corresponding to the font package type.
7. The non-intrusive translation method according to claim 6, characterized in that: After the step of determining the type corresponding to each element to be translated according to the tag, the method further comprises: In response to the element to be translated being of the image type or the font package type, and the server not extracting the text information from the image corresponding to the element to be translated, merging each element to be translated from which the text information has not been extracted into the same sprite image, and recording the coordinates of each element to be translated in the sprite image in the virtual tree structure; The step of rendering the plurality of restored data into the virtual tree structure based on the unique identifier corresponding to each restored data to generate a virtual translated web page comprises: Each of the elements to be translated in the Sprite image is split, and each of the elements to be translated in the Sprite image is rendered into the virtual translation webpage based on the corresponding coordinates.
8. A non-intrusive translation method, characterized in that: include: The server receives multiple translation requests carrying data to be translated sent by the client; wherein the data to be translated is generated by the client based on multiple elements to be translated identified by the web page to be translated, and each of the data to be translated includes a unique identifier and a text to be translated; Based on the multiple translation requests, a plurality of data to be restored corresponding to the multiple data to be translated is obtained; wherein each of the data to be restored includes the unique identifier and the translation text; The plurality of data to be restored are sent to the client, so that the client restores each of the data to be restored into restored data, and according to the unique identifier corresponding to each of the restored data, the plurality of restored data are rendered into a virtual tree structure constructed by the client based on the plurality of elements to be translated, so as to generate a virtual translated web page.
9. The non-intrusive translation method according to claim 8, characterized in that: The step of acquiring a plurality of data to be restored corresponding to the plurality of data to be translated based on the plurality of translation requests comprises: Based on each of the translation requests, querying in the secondary cache of the server whether there is the data to be restored corresponding to the data to be translated; In response to the existence of the data to be restored corresponding to the data to be translated in the secondary cache, sending the data to be restored to the primary cache of the client; In response to the absence of the to-be-translated data corresponding to the to-be-translated data in the secondary cache, sending the translation request to the cloud platform, so that the cloud platform translates the to-be-translated data based on the translation request; The data to be restored returned by the cloud platform is received and stored through the secondary cache, and the data to be restored is sent to the primary cache of the client.
10. The non-intrusive translation method according to claim 9, characterized in that: Before the server receives a plurality of to-be-translated data and corresponding translation requests sent by the client, the method includes: Establishing a correction subject library in the server; wherein the correction subject library stores a plurality of professional fields and professional terms and translation knowledge corresponding to the professional fields, and each of the professional fields carries at least one classification label; After the server receives a plurality of translation requests carrying data to be translated sent by the client, the server includes: Based on each of the texts to be translated, storing the corresponding translation request in the most matching professional field; After the step of receiving and storing the data to be restored returned by the cloud platform through the secondary cache and sending the data to be restored to the primary cache of the client, the method further comprises: Storing the data to be restored in the professional field that best matches the data; Manual correction processing is performed on the translation data stored in the correction subject library, and the translation data after the manual correction processing is synchronously stored in the secondary cache of the server.
11. A client, characterized in that: include: A construction module, used for parsing the web page to be translated in response to the translation instruction, and constructing a virtual tree structure based on the obtained multiple elements to be translated; A preprocessing module, used for preprocessing each of the elements to be translated to generate a plurality of data to be translated; wherein each data to be translated includes a unique identifier and a text to be translated; A first acquisition module is used to acquire a plurality of data to be restored corresponding to the plurality of data to be translated; wherein each of the data to be restored includes the unique identifier and a translation text; A restoration module, used for restoring each of the to-be-restored data into restored data; A generating module is used for rendering the plurality of restored data into the virtual tree structure based on the unique identifier corresponding to each restored data, so as to generate a virtual translated web page.
12. A server, characterized in that: include: A receiving module, configured to receive a plurality of translation requests carrying data to be translated sent by a client; wherein the data to be translated is generated by the client based on a plurality of elements to be translated identified by the web page to be translated, and each of the data to be translated includes a unique identifier and a text to be translated; A second acquisition module is used to acquire a plurality of data to be restored corresponding to the plurality of data to be translated based on the plurality of translation requests; wherein each of the data to be restored includes the unique identifier and the translation text; The sending module is used to send the multiple data to be restored to the client, so that the client restores each of the data to be restored into restored data, and renders the multiple restored data into a virtual tree structure constructed by the client based on the multiple elements to be translated according to the unique identifier corresponding to each of the restored data, so as to generate a virtual translated web page.
13. An electronic device, characterized in that: include: A memory for storing program data, wherein when the stored program data is executed, the steps of the non-intrusive translation method according to any one of claims 1 to 7 or the non-intrusive translation method according to any one of claims 8 to 10 are implemented; A processor is used to execute the program instructions stored in the memory to implement the steps in the non-intrusive translation method according to any one of claims 1 to 7 or the non-intrusive translation method according to any one of claims 8 to 10.
14. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the non-intrusive translation method according to any one of claims 1 to 7 or the non-intrusive translation method according to any one of claims 8 to 10 are implemented.