Text translation method and device, electronic equipment and storage medium
By introducing cultural knowledge bases and policy databases for verification into the neural machine translation algorithm, the problems of cultural semantic bias and content violations in cross-regional translation are solved, thereby improving the accuracy and flexibility of text translation results.
Patent Information
- Application Number
- CN202511572541.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-30
- Publication Date
- 2026-01-30
Smart Images

Figure CN121435993A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of computer, and in particular, to a text translation method and device, electronic equipment and storage medium. BACKGROUND
[0002] With the development of artificial intelligence technology, machine translation (MT) algorithm has become a key text processing tool to break the language barrier. From the early rule-based machine translation (RBMT) algorithm to the current mainstream neural machine translation (NMT) algorithm, although the fluency of the text translation process has been significantly improved, there are still the following limitations:
[0003] The existing NMT algorithm can only complete cross-language conversion of the text, and lacks adaptive adjustment of the text content to cross-regional cultural habits, policy compliance, etc., resulting in the final translation result may have cultural sensitivity deficiency, text content does not comply with local policy regulations, etc. SUMMARY
[0004] The present application provides a text translation method, device, electronic equipment and storage medium, which can solve the problems of cultural semantic deviation and content violation risk in cross-regional translation text, eliminate the ambiguity in the translated text, and improve the accuracy and flexibility of the text translation result.
[0005] According to an aspect of the present application, a text translation method is provided, the method comprising:
[0006] adopting a neural machine translation algorithm to translate the original text to obtain an initial translation text;
[0007] inputting the initial translation text into a pre-constructed cultural knowledge base, and outputting a content adjustment strategy corresponding to the initial translation text through the cultural knowledge base;
[0008] processing the initial translation text according to the content adjustment strategy to obtain an intermediate translation text, and then calling a pre-constructed policy database to verify the intermediate translation text;
[0009] generating a target translation text corresponding to the original text according to the verification result of the intermediate translation text.
[0010] Optionally, before adopting the neural machine translation algorithm to translate the original text to obtain the initial translation text, it further comprises:
[0011] collecting professional cultural data corresponding to different regions; wherein the professional cultural data comprises at least one of the following:
[0012] expert knowledge data, cultural academic data, media public data, and user feedback data for translation samples;
[0013] According to the association relationship between each cultural item in the professional cultural data, the cultural knowledge base is constructed.
[0014] Optionally, the initial translation text is input into the pre-constructed cultural knowledge base, and a content adjustment strategy corresponding to the initial translation text is output through the cultural knowledge base, including:
[0015] The region code, text content and text type corresponding to the initial translation text are input into the pre-constructed cultural knowledge base;
[0016] Through the cultural knowledge base, the cultural suitability score corresponding to the initial translation text is determined according to the region code, text content and text type, and the content adjustment strategy is output according to the cultural suitability score.
[0017] Optionally, before calling the pre-constructed policy database to verify the intermediate translation text, it further includes:
[0018] Through the large language model, the text paragraph and / or text diction in the intermediate translation text are modified to make the expression mode corresponding to the intermediate translation text meet the preset requirements.
[0019] Optionally, calling the pre-constructed policy database to verify the intermediate translation text includes:
[0020] According to the policy database and the preset multiple verification dimensions, the large language model determines the corresponding violation probability of the intermediate translation text under different verification dimensions;
[0021] According to the verification result of the intermediate translation text, the target translation text corresponding to the original text is generated, including:
[0022] According to the violation probability of the intermediate translation text under different verification dimensions, the non-compliant text content in the intermediate translation text is filtered to obtain the target translation text.
[0023] Optionally, before inputting the initial translation text into the pre-constructed cultural knowledge base, it further includes:
[0024] Extracting key metadata in the initial translation text and preprocessing the initial translation text; the preprocessing at least includes text cleaning processing and text standardization processing;
[0025] The key metadata includes text publishing time, author information and text keywords.
[0026] According to another aspect of the present application, there is provided a text translation apparatus, the apparatus comprising:
[0027] an initial translation module configured to translate the original text using a neural machine translation algorithm to obtain an initial translated text;
[0028] a culture adaptation module configured to input the initial translated text into a pre-constructed culture knowledge base, and output a content adjustment strategy corresponding to the initial translated text through the culture knowledge base;
[0029] a compliance verification module configured to process the initial translated text according to the content adjustment strategy to obtain an intermediate translated text, and then call a pre-constructed policy database to verify the intermediate translated text;
[0030] a target text generation module configured to generate a target translated text corresponding to the original text according to a verification result of the intermediate translated text.
[0031] According to another aspect of the present application, there is provided an electronic device, the electronic device comprising:
[0032] at least one processor; and
[0033] a memory communicatively connected to the at least one processor; wherein
[0034] the memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor to enable the at least one processor to perform the text translation method according to any one of the embodiments of the present application.
[0035] According to another aspect of the present application, there is provided a computer readable storage medium storing computer instructions for enabling a processor to implement the text translation method according to any one of the embodiments of the present application when executed by the processor.
[0036] According to another aspect of the present application, there is provided a computer program product comprising a computer program for implementing the text translation method according to any one of the embodiments of the present application when executed by a processor.
[0037] The technical scheme provided by the embodiment of the present application comprises the following steps: an original text is translated into an initial translation text by using a neural machine translation algorithm; the initial translation text is input into a pre-constructed cultural knowledge base, and a content adjustment strategy corresponding to the initial translation text is output by the cultural knowledge base; the initial translation text is processed into an intermediate translation text according to the content adjustment strategy; the intermediate translation text is verified by calling a pre-constructed policy database; and a target translation text corresponding to the original text is generated according to the verification result of the intermediate translation text. The technical scheme provides a text translation method that is suitable for different regional cultures and conforms to regional policy rules, and can solve the problems of cultural semantic deviation and content violation risk in cross-regional translation texts, eliminate ambiguity in the translation text, and improve the accuracy and flexibility of the text translation result.
[0038] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present application, nor is it intended to limit the scope of the present application. Other features of the present application will become apparent from the following description. BRIEF DESCRIPTION OF DRAWINGS
[0039] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings needed in the embodiment description will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.
[0040] Figure 1 is a flowchart of a text translation method according to an embodiment of the present application;
[0041] Figure 2 is a flowchart of another text translation method according to an embodiment of the present application;
[0042] Figure 3 is a structural schematic diagram of a text translation device according to an embodiment of the present application;
[0043] Figure 4 is a structural schematic diagram of an electronic device implementing the text translation method according to an embodiment of the present application. DETAILED DESCRIPTION
[0044] In order to enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings of the embodiments of the present application. Obviously, the described embodiments are only some embodiments of the present application, not all. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor should be within the scope of protection of the present application.
[0045] It should be noted that the terms "first", "second", and the like in the description and in the claims of the present application and in the above-described drawings are intended to distinguish similar objects and not necessarily describe a particular sequential or chronological order. It should be understood that the data thus used can be interchanged, where appropriate, so that the embodiments of the application described herein can be practiced in other than the illustrated or described order. Furthermore, the terms "comprise" and "have" and any variations thereof are intended to cover non-exclusive inclusion, for example, processes, methods, systems, products, or devices that include a list of steps or units not necessarily limited to those clearly listed, but can include other steps or units not clearly listed or inherent to such processes, methods, products, or devices.
[0046] Figure 1 A flowchart of a text translation method provided for an embodiment of the present application, the embodiment can be applicable to the case of translating texts across cultures, the method can be executed by a text translation device, which can be realized in the form of hardware and / or software, and configured in an electronic device. As shown in the figure, the method comprises: Figure 1
[0047] Step 110: translating the original text by using a neural machine translation algorithm to obtain an initial translation text.
[0048] In this step, the original text can be a news release or a product manual, etc., and the embodiment does not limit this.
[0049] Specifically, after obtaining the original text, an encoder in the neural machine translation algorithm can be used to map the language sequence in the original text into a fixed-length semantic vector, and then a decoder can be used to decode the semantic vector into a target language sequence, thereby completing the language conversion of the original text and obtaining the initial translation text.
[0050] Step 120: inputting the initial translation text into a pre-constructed cultural knowledge base, and outputting a content adjustment strategy corresponding to the initial translation text through the cultural knowledge base.
[0051] In the embodiment, the cultural knowledge base is a structured, extensible, and iterative knowledge base, which is used to provide the initial translation text with deep cultural knowledge beyond words, so that the content of the translation text is not only correct, but also conforms to the cultural habits of the target region.
[0052] In this step, specifically, the target region corresponding to the initial translation text can be identified, the content in the initial translation text is analyzed according to the cultural habits of the target region through the cultural knowledge base, and a content adjustment strategy corresponding to the initial translation text is output according to the analysis result. Optionally, the content adjustment strategy can include the mapping relationship between the original content and the target content.
[0053] In one embodiment of the present embodiment, before the original text is translated using the neural machine translation algorithm to obtain the initial translation text, it further includes: collecting professional cultural data corresponding to different regions, and constructing the cultural knowledge base according to the association relationship between each cultural item in the professional cultural data. The professional cultural data includes at least one of the following: expert knowledge data, cultural academic data, media public data, and user feedback data for translation samples.
[0054] In a specific embodiment, expert knowledge data, such as the knowledge data of anthropologists and linguists about different regional cultures, can be obtained through structured interviews and questionnaires. After obtaining the expert knowledge data, the expert knowledge data can be verified through a preset verification process, such as verifying whether the data source meets the preset requirements.
[0055] In addition, cultural academic data corresponding to different regions can also be obtained, including cross-cultural research monographs, academic papers, academic reports, and cultural promotion materials. At the same time, big data mining and crawling algorithms can be used to obtain cultural data publicly disclosed by different media platforms, including news texts reported by news websites (such as news texts about the same event in different regions), social media publicly disclosed regional cultural data, local product introduction data in different regions, subtitles of films and television shows in different regions, and publicly disclosed data of tourism websites in different regions. The present embodiment does not limit this.
[0056] Specifically, when obtaining media public data, a natural language processing (NLP) algorithm can be used to extract key cultural pattern data and cultural sensitive data from unstructured text publicly disclosed by media platforms. Optionally, the NLP algorithm can include sentiment analysis algorithms and topic modeling algorithms.
[0057] The advantage of such a setting is that by obtaining professional cultural data of different regions, the authority and consistency of the subsequent text translation result in the professional field can be ensured, so that the text translation method can be flexibly applied to different professional fields, such as medical diagnosis, technical maintenance, and international conferences.
[0058] In this embodiment, the translation text feedback function can also be provided to users in different regions through a preset visual interface. Specifically, the translation samples corresponding to different regions can be displayed to users in different regions. After browsing the translation samples, the users can return evaluation data through a "translation feedback" button, such as "translation content is impolite" and "translation content is inconsistent with local cultural habits". After obtaining the evaluation data of the users for the translation samples, the evaluation data can be input as training data into the cultural knowledge base to optimize the cultural knowledge in the cultural knowledge base.
[0059] In a specific embodiment, after the professional cultural data is collected, a graph database or a non-relational database (Not Only SQL, NOSQL) can be used to construct the cultural knowledge base according to the association relationships between the cultural entries in the professional cultural data.
[0060] Optionally, in the cultural knowledge base, a preset format specification (for example, JSON Schema) can be used to store each cultural entry, thereby ensuring the uniformity of the cultural entry format and facilitating the calling of the text translation process.
[0061] Step 130: processing the initial translation text according to the content adjustment strategy to obtain an intermediate translation text, and then calling a pre-constructed policy database to verify the intermediate translation text.
[0062] In this embodiment, the initial translation text can be processed according to the content adjustment strategy to obtain an intermediate translation text, for example, adjusting the expression manner (style, sentence pattern, etc.) of the cultural content in the initial translation text to ensure that the content of the intermediate translation text is appropriate and humanized.
[0063] Taking a news release as an example, the direct translation title of the news release can be adjusted to a more eye-catching title that conforms to the local news habits of the target region.
[0064] In this step, the legal and regulatory database of the target region in the policy database can be called. According to the legal and regulatory database and the context information in the intermediate translation text, the compliance of the intermediate translation text is verified by a large language model, and a compliance report and a violation list are generated.
[0065] Step 140: generating a target translation text corresponding to the original text according to the verification result of the intermediate translation text.
[0066] In this step, specifically, the non-compliant text content in the intermediate translation text can be filtered according to the above-mentioned violation list, thereby obtaining the final translation text (i.e., the target translation text) corresponding to the original text.
[0067] The technical scheme provided by the embodiment of the present application comprises the following steps: an original text is translated into an initial translation text by using a neural machine translation algorithm; the initial translation text is input into a pre-constructed cultural knowledge base; a content adjustment strategy corresponding to the initial translation text is output by the cultural knowledge base; the initial translation text is processed into an intermediate translation text according to the content adjustment strategy; the intermediate translation text is verified by calling a pre-constructed policy database; and a target translation text corresponding to the original text is generated according to the verification result of the intermediate translation text. The technical scheme provides a text translation method that is suitable for different regional cultures and conforms to regional policy rules, and can solve the problems of cultural semantic deviation and content violation risk in cross-regional translation texts, eliminate ambiguity in the translation text, and improve the accuracy and flexibility of the text translation result.
[0068] On the basis of the above-mentioned embodiment, the embodiment further provides a method for maintaining and iteratively updating the cultural knowledge base, which specifically comprises the following steps: new cultural knowledge data is automatically identified by continuously monitoring global major news and social networks by using a crawling algorithm; a successful translation text case is used as a reward signal to train a content adjustment strategy output model, so that the model outputs a content adjustment strategy that is more in line with different regional cultural habits; and cultural entries in the cultural knowledge base are updated in real time according to the update result of the local knowledge base in different regions.
[0069] Optionally, in the embodiment, after the cultural knowledge base is updated each time, the version number of the cultural knowledge base, the update content, the target region involved in the update content, and the update reason can be recorded.
[0070] Figure 2 A flowchart of another text translation method provided by the embodiment of the present application is shown in FIG. 2, which comprises the following steps: Figure 2
[0071] In step 210, an original text is translated into an initial translation text by using a neural machine translation algorithm, and key metadata in the initial translation text is extracted and the initial translation text is preprocessed.
[0072] In the embodiment, the key metadata comprises text publishing time, author information, and text keywords, and the preprocessing at least comprises text cleaning processing and text standardization processing.
[0073] In step 220, a region code corresponding to the initial translation text, text content, and text type are input into a pre-constructed cultural knowledge base.
[0074] In this step, the region code, text content and text type corresponding to the initial translation text can be input to the cultural knowledge base through a preset interface (Representational State Transfer API, RESTful API).
[0075] Step 230, determining, by the cultural knowledge base, a cultural suitability score corresponding to the initial translation text according to the region code, text content and text type, and outputting a content adjustment strategy according to the cultural suitability score.
[0076] In this step, specifically, the cultural suitability score corresponding to the initial translation text can be determined by the following formula :
[0077]
[0078] wherein, represents the initial translation text, represents a cultural entry in the cultural knowledge base, represents a feature vector generated according to the region code, text content and text type, represents a feature vector of the cultural entry.
[0079] In this embodiment, optionally, if the cultural suitability score corresponding to the initial translation text is low, the content in the initial translation text can be analyzed according to the cultural entry corresponding to the target region, and a content adjustment strategy can be output according to the analysis result.
[0080] Step 240, processing the initial translation text to obtain an intermediate translation text according to the content adjustment strategy, and modifying the text paragraphs and / or text expressions in the intermediate translation text by a large language model, so that the expression mode corresponding to the intermediate translation text meets the preset requirements.
[0081] In this step, the intermediate translation text can be obtained by processing the initial translation text according to the content adjustment strategy through a natural language generation (NLG) model, and then the unqualified content in the intermediate translation text can be automatically modified, for example, the paragraphs and expressions in the intermediate translation text can be modified, so as to ensure the consistency and fluency of the expression mode of the intermediate translation text.
[0082] Step 250, determining, by a large language model, a violation probability corresponding to the intermediate translation text under different verification dimensions according to a pre-constructed policy database and a plurality of preset verification dimensions.
[0083] In this step, the violation probability of the intermediate translation text under different verification dimensions can be determined by the following formula :
[0084]
[0085] wherein, is a preset evaluation function, is a weight value, is a bias value, is a feature vector of the intermediate translation text under different verification dimensions.
[0086] In one specific embodiment, the verification dimensions can include data privacy (such as phone number, email address, ID number, IP address) dimension, financial and health industry sensitive data dimension, policy sensitive data dimension, absolute language dimension, and data security dimension, etc.
[0087] Step 260, filtering the non-compliant text content in the intermediate translation text according to the violation probability of the intermediate translation text under different verification dimensions, to obtain a target translation text.
[0088] In this embodiment, taking a news release as an example, the original news release can be obtained first, the neural machine translation algorithm is used to translate the original news release, then the cultural content in the news translation is optimized and adjusted based on the cultural knowledge base, the news translation is checked for compliance through the policy database of the target region, the non-compliant text content in the news translation is filtered, and finally the standardized news release adapted for dissemination in the target region is output.
[0089] The technical scheme provided by the embodiment of the application comprises the following steps: an original text is translated by using a neural machine translation algorithm to obtain an initial translation text, key metadata in the initial translation text is extracted, the initial translation text is preprocessed, region codes, text content and text types corresponding to the initial translation text are input into a cultural knowledge base, a cultural suitability score corresponding to the initial translation text is determined by using the cultural knowledge base, a content adjustment strategy is output according to the cultural suitability score, an intermediate translation text is obtained by processing the initial translation text according to the content adjustment strategy, a text paragraph and / or a text expression in the intermediate translation text are modified by using a large language model, so that an expression mode corresponding to the intermediate translation text meets a preset requirement, a policy database and a plurality of preset check dimensions are used to determine a rule violation probability corresponding to the intermediate translation text under different check dimensions by using the large language model, and a target translation text is obtained by filtering non-compliant text content in the intermediate translation text according to the rule violation probability corresponding to the intermediate translation text under different check dimensions. The technical scheme can solve the problems of cultural semantic deviation and content rule violation risk in cross-region translation texts, eliminate ambiguity in the translation text, and improve the accuracy and flexibility of the text translation result.
[0090] Figure 3 A structural schematic diagram of a text translation device provided by the embodiment of the application is shown in FIG. 1. The device is applied to an electronic device, such as a computer. Figure 3 As shown in the figure, the device comprises an initial translation module 310, a cultural adaptation module 320, a compliance verification module 330 and a target text generation module 340.
[0091] The initial translation module 310 is configured to translate an original text by using a neural machine translation algorithm to obtain an initial translation text.
[0092] The cultural adaptation module 320 is configured to input the initial translation text into a pre-constructed cultural knowledge base, and output a content adjustment strategy corresponding to the initial translation text by using the cultural knowledge base.
[0093] The compliance verification module 330 is configured to process the initial translation text according to the content adjustment strategy to obtain an intermediate translation text, and then call a pre-constructed policy database to verify the intermediate translation text.
[0094] The target text generation module 340 is configured to generate a target translation text corresponding to the original text according to a verification result of the intermediate translation text.
[0095] The technical scheme provided by the embodiment of the application comprises the following steps: an original text is translated into an initial translation text by using a neural machine translation algorithm; the initial translation text is input into a pre-constructed cultural knowledge base; a content adjustment strategy corresponding to the initial translation text is output by the cultural knowledge base; an intermediate translation text is obtained by processing the initial translation text according to the content adjustment strategy; the intermediate translation text is verified by calling a pre-constructed policy database; and a target translation text corresponding to the original text is generated according to the verification result of the intermediate translation text. The technical scheme provides a text translation mode that is adapted to different regional cultures and meets regional policy rules, and can solve the problems of cultural semantic deviation and content violation risk in cross-regional translation texts, eliminate ambiguity in the translation text, and improve the accuracy and flexibility of the text translation result.
[0096] On the basis of the above-mentioned embodiment, the device further comprises:
[0097] The knowledge base construction module is configured to collect professional cultural data corresponding to different regions, and construct the cultural knowledge base according to the association relationship between each cultural item in the professional cultural data. The professional cultural data comprises at least one of the following: expert knowledge data, cultural academic data, media public data, and feedback data of a user for a translation sample.
[0098] The text modification module is configured to modify a text paragraph and / or a text diction in the intermediate translation text by using a large language model before verifying the intermediate translation text by calling the pre-constructed policy database, so that the expression manner corresponding to the intermediate translation text meets a preset requirement.
[0099] The initial translation module 310 comprises:
[0100] The preprocessing unit is configured to extract key metadata in the initial translation text and pre-process the initial translation text. The preprocessing at least comprises text cleaning processing and text standardization processing. The key metadata comprises a text publishing time, author information, and a text keyword.
[0101] The cultural adaptation module 320 comprises:
[0102] The text input unit is configured to input a region code, text content, and a text type corresponding to the initial translation text into a pre-constructed cultural knowledge base.
[0103] The strategy output unit is configured to determine a cultural suitability score corresponding to the initial translation text according to the region code, the text content, and the text type by using the cultural knowledge base, and output a content adjustment strategy according to the cultural suitability score.
[0104] The compliance verification module 330 comprises:
[0105] The probability assessment unit is used to determine the probability of violation of the intermediate translated text under different verification dimensions by using a large language model based on the policy database and multiple preset verification dimensions.
[0106] The target text generation module 340 includes:
[0107] The text filtering unit is used to filter non-compliant text content in the intermediate translated text according to the violation probability corresponding to different verification dimensions, so as to obtain the target translated text.
[0108] The above-described apparatus can execute the methods provided in all the foregoing embodiments of the present invention, and has the corresponding functional modules and beneficial effects for executing the above methods. Technical details not described in detail in the embodiments of the present invention can be found in the methods provided in all the foregoing embodiments of the present invention.
[0109] Figure 4 A schematic diagram of an electronic device 10, which can be used to implement embodiments of the present invention, is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices (e.g., helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.
[0110] like Figure 4 As shown, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12 or a random access memory (RAM) 13, communicatively connected to the at least one processor 11. The memory stores computer programs executable by the at least one processor. The processor 11 can perform various appropriate actions and processes based on the computer program stored in the ROM 12 or loaded from storage unit 18 into the RAM 13. The RAM 13 can also store various programs and data required for the operation of the electronic device 10. The processor 11, ROM 12, and RAM 13 are interconnected via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.
[0111] A plurality of components in the electronic device 10 are connected to the I / O interface 15, including: an input unit 16, such as a keyboard, a mouse, etc.; an output unit 17, such as various types of displays, speakers, etc.; a storage unit 18, such as a magnetic disk, an optical disk, etc.; and a communication unit 19, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 19 allows the electronic device 10 to exchange information / data with other devices through a computer network, such as the Internet, and / or various telecommunication networks.
[0112] The processor 11 can be various general and / or special purpose processing components with processing and computing capabilities. Some examples of the processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various specialized artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The processor 11 performs various methods and processes described above, such as the text translation method.
[0113] In some embodiments, the text translation method can be implemented as a computer program tangibly embodied in a computer readable storage medium, such as the storage unit 18. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 10 via the ROM 12 and / or the communication unit 19. When the computer program is loaded onto the RAM 13 and executed by the processor 11, one or more steps of the text translation method described above can be performed. Alternatively, in other embodiments, the processor 11 can be configured to perform the text translation method by any other appropriate means, such as by means of firmware.
[0114] Various implementations of the systems and techniques described above can be realized in digital electronic circuitry, integrated circuitry, a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), a system on a chip (SOC), a programmable logic device (CPLD), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include implementation in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be special or general purpose, coupled to receive data and instructions from, and to transmit data and instructions to, a storage system, at least one input device, and at least one output device.
[0115] Computer programs for implementing the methods of the present application can be written in any combination of one or more programming languages. These computer programs can be provided to a processor of a general purpose computer, special purpose computer, or other programmable data processing apparatus, such that the computer program
[0116] In the context of the present application, a computer-readable storage medium can be a tangible medium that can contain or store a computer program for use by or in connection with an instruction execution system, apparatus, or device. A computer-readable storage medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. Alternatively, a computer-readable storage medium can be a machine-readable signal medium. More specific examples of a machine-readable storage medium will include one or more lines of a program of instructions in a transitory signal, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0117] To provide for interaction with a user, the systems and techniques described here can be implemented on an electronic device having a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the electronic device. Other kinds of devices can be used to provide for interaction with a user as well; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form, including acoustic, speech, or tactile input.
[0118] The systems and techniques described herein can be implemented in a computing system that includes a back end component, e.g., as a data server, or that includes a middleware component, e.g., an application server, or that includes a front end component, e.g., a user computer having a graphical user interface or a Web browser through which a user can interact with an implementation of the systems and techniques described herein, or any combination of such back end, middleware, or front end components. The components of the system can be interconnected by any form or medium of digital data communication, e.g., a communication network. Examples of communication networks include a local area network (LAN), a wide area network (WAN), a blockchain network, and the Internet.
[0119] The computing system can include clients and servers. A client and server are generally remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other. A server can be a cloud server, also known as a cloud computing server or cloud host, which is a host product in the cloud computing service system, to solve the defects of large management difficulty and weak business scalability in traditional physical host and VPS service.
[0120] It should be understood that the various forms of flow shown above can be re-ordered, added to, or deleted from without departing from the scope of the present disclosure. For example, the steps recited in the present disclosure can be executed in parallel, executed in series, or executed in different orders, as long as the desired results of the technical solutions of the present disclosure can be achieved, and the present disclosure is not limited herein.
[0121] The above detailed description does not constitute a limitation on the protection scope of the present application. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent replacements, and improvements made within the spirit and principles of the present application shall be included in the protection scope of the present application.
Claims
1. A method of translating text, characterized by, The method comprises: adopting a neural machine translation algorithm to translate the original text to obtain an initial translation text; inputting the initial translation text into a pre-constructed cultural knowledge base, and outputting a content adjustment strategy corresponding to the initial translation text through the cultural knowledge base; processing the initial translation text according to the content adjustment strategy to obtain an intermediate translation text, and then calling a pre-constructed policy database to verify the intermediate translation text; generating a target translation text corresponding to the original text according to the verification result of the intermediate translation text.
2. The method of claim 1, wherein, Before adopting the neural machine translation algorithm to translate the original text to obtain the initial translation text, the method further comprises: collecting professional cultural data corresponding to different regions; wherein the professional cultural data comprises at least one of the following: expert knowledge data, cultural academic data, media public data, and user feedback data for translation samples; constructing the cultural knowledge base according to the association relationship between each cultural item in the professional cultural data.
3. The method of claim 1, wherein, Inputting the initial translation text into a pre-constructed cultural knowledge base, and outputting a content adjustment strategy corresponding to the initial translation text through the cultural knowledge base, comprises: inputting the region code, text content and text type corresponding to the initial translation text into the pre-constructed cultural knowledge base; determining the cultural suitability score corresponding to the initial translation text according to the region code, text content and text type through the cultural knowledge base, and outputting the content adjustment strategy according to the cultural suitability score.
4. The method of claim 1, wherein, Before calling the pre-constructed policy database to verify the intermediate translation text, the method further comprises: modifying the text paragraphs and / or text expressions in the intermediate translation text through a large language model, so that the expression mode corresponding to the intermediate translation text meets the preset requirements.
5. The method of claim 1, wherein, Calling the pre-constructed policy database to verify the intermediate translation text comprises: determining the violation probability corresponding to the intermediate translation text under different verification dimensions according to the policy database and the preset multiple verification dimensions through a large language model; According to the verification result of the intermediate translation text, generating a target translation text corresponding to the original text, comprises: filtering the non-compliant text content in the intermediate translation text according to the violation probability corresponding to the intermediate translation text under different verification dimensions to obtain the target translation text.
6. The method of claim 1, wherein, Before inputting the initial translation text into the pre-constructed cultural knowledge base, the method further comprises: extracting key metadata in the initial translation text and preprocessing the initial translation text; the preprocessing at least comprises: text cleaning processing, text standardization processing; The key metadata comprises text publishing time, author information and text keywords.
7. A text translation apparatus characterized by comprising: The device comprises: an initial translation module configured to adopt a neural machine translation algorithm to translate an original text to obtain an initial translation text; a cultural adaptation module configured to input the initial translation text into a pre-constructed cultural knowledge base, and output a content adjustment strategy corresponding to the initial translation text through the cultural knowledge base; A compliance verification module is configured to process the initial translation text according to the content adjustment policy to obtain intermediate translation text, and then call a pre-constructed policy database to verify the intermediate translation text; A target text generation module is configured to generate target translation text corresponding to the original text according to a verification result of the intermediate translation text.
8. An electronic device, comprising: The electronic device comprises: at least one processor; and a memory connected to the at least one processor in communication; wherein The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor to enable the at least one processor to execute the text translation method of any one of claims 1-6.
9. A computer-readable storage medium, characterized in that, The computer readable storage medium stores computer instructions for enabling the processor to implement the text translation method of any one of claims 1-6 when executed.
10. A computer program product, characterised in that, The computer program product comprises a computer program, and the computer program implements the text translation method according to any one of claims 1-6 when executed by the processor. The computer program product comprises a computer program, and the computer program implements the text translation method according to any one of claims 1-6 when executed by the processor.