Intelligent translation method and device

Through intelligent translation methods, the problem of difficulty in transmitting cultural elements in multimedia content is solved, and more efficient localization and user experience improvement of multimedia content is achieved.

CN119940370APending Publication Date: 2025-05-06SHANGHAI BILIBILI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510121817.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-24
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

It is difficult for multimedia content to effectively convey cultural elements during the translation process, resulting in difficulty for users in different countries and regions to understand and affect the viewing experience.

Method used

Through intelligent translation methods, the cultural elements in the multimedia content are automatically identified and replaced with target cultural elements that are easier to understand and accept. Combined with text translation and cultural element replacement, the translated target multimedia content is generated.

Benefits of technology

It improves the localization effect of multimedia content, enables users in different countries and regions to better understand and integrate content, and improves the user's viewing experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119940370A_ABST
    Figure CN119940370A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides an intelligent translation method, which comprises the following steps: acquiring multimedia content to be translated, the multimedia content to be translated comprising an image to be processed; obtaining text content and cultural elements according to the to-be-processed image; according to the target translation language, translating the text content to obtain target text content; obtaining target culture elements for replacing the culture elements according to the target translation language; and obtaining translated target multimedia content based on the to-be-processed image, the target text content and the target culture element. By automatically identifying the culture elements in the multimedia content and replacing the culture elements with the target culture elements easier to understand and accept, users in different countries and regions can better understand and integrate the content when watching the multimedia content. By automatically replacing the culture elements in the multimedia content, the localization effect of the multimedia content is efficiently improved, and therefore the watching experience of a user is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application relate to the field of translation technology, and in particular, to an intelligent translation method, apparatus, computer device, computer-readable storage medium, and computer program product. Background Art

[0002] With the continuous development of globalization, the spread of multimedia content (such as comics, animation, etc.) in different countries and regions is increasing. Multimedia content usually contains cultural elements of a specific country or region. However, the translation process of multimedia content is easily affected by the translator, making it difficult for cultural elements in multimedia content to be understood by users from other countries or regions, which in turn affects the user's viewing experience.

[0003] It should be noted that the above content is not necessarily prior art, nor is it intended to limit the scope of patent protection of this application. Summary of the invention

[0004] The embodiments of the present application provide an intelligent translation method, apparatus, computer device, computer-readable storage medium, and computer program product to solve or alleviate one or more of the technical problems raised above.

[0005] One aspect of an embodiment of the present application provides an intelligent translation method, the method comprising: Acquire multimedia content to be translated, wherein the multimedia content to be translated includes an image to be processed; Acquire text content and cultural elements according to the image to be processed; According to the target translation language, the text content is translated to obtain the target text content; According to the target translation language, obtaining a target cultural element for replacing the cultural element; and Based on the image to be processed, the target text content and the target cultural element, the translated target multimedia content is acquired.

[0006] Optionally, obtaining text content and cultural elements according to the image to be processed includes: Inputting the image to be processed into a pre-trained cultural element recognition model; and One or more cultural elements are determined from the image to be processed by using the cultural element recognition model.

[0007] Optionally, obtaining a target cultural element for replacing the cultural element according to the target translation language includes: According to the target translation language, determining the target cultural information corresponding to the target translation language; Based on the target cultural information and the cultural elements, the target cultural elements are acquired from a preset cultural knowledge base.

[0008] Optionally, obtaining translated target multimedia content based on the image to be processed, the target text content and the target cultural element includes: Obtaining size information of the text content in the image to be processed; According to the character length of the target text content and the size information of the text content, the attributes of the target text content are adjusted; wherein the attributes of the target text content include font style and / or font size.

[0009] Optionally, the text content includes a plurality of sentences; and translating the text content according to a target translation language to obtain a target text content includes: Based on the target translation language and the text translation model, the plurality of sentences are translated to obtain a plurality of target sentences; Obtaining context information of each of the target sentences; Optimizing and adjusting the target sentence according to the context information of the target sentence; The target text content includes: a plurality of target sentences corresponding to the plurality of sentences and after being optimized.

[0010] Optionally, the method further comprises: Acquiring feedback data for the target multimedia content; Based on the feedback data, the cultural element recognition model and / or the text translation model is adjusted.

[0011] Another aspect of the embodiments of the present application provides an intelligent translation device, the device comprising: an acquisition module, configured to acquire multimedia content to be translated, wherein the multimedia content to be translated includes an image to be processed; and Acquire text content and cultural elements according to the image to be processed; A translation module, used for translating the text content according to a target translation language to obtain a target text content; The acquisition module is further used to acquire a target cultural element for replacing the cultural element according to the target translation language; and Based on the image to be processed, the target text content and the target cultural element, the translated target multimedia content is acquired.

[0012] Another aspect of an embodiment of the present application provides a computer device, including: at least one processor; and a memory communicatively coupled to the at least one processor; Wherein: the memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method as described above.

[0013] Another aspect of an embodiment of the present application provides a computer-readable storage medium, wherein the computer-readable storage medium stores computer instructions, and when the computer instructions are executed by a processor, the method described above is implemented.

[0014] Another aspect of an embodiment of the present application provides a computer program product, including a computer program, which implements the method described above when executed by a processor.

[0015] The above technical solution adopted in the embodiment of the present application may include the following advantages: by automatically identifying cultural elements in multimedia content and replacing them with target cultural elements that are easier to understand and accept, users from different countries and regions can better understand and integrate into the content when watching multimedia content. By automatically replacing cultural elements in multimedia content, the localization effect of multimedia content is efficiently improved, thereby improving the user's viewing experience. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] The accompanying drawings exemplarily illustrate the embodiments and constitute a part of the specification, and together with the text description of the specification, are used to explain the exemplary implementation of the embodiments. The embodiments shown are for illustrative purposes only and do not limit the scope of the claims. In all drawings, the same reference numerals refer to similar but not necessarily identical elements.

[0017] Figure 1 The flowchart of the intelligent translation method according to the first embodiment of the present application is schematically shown; Figure 2 Schematically shows Figure 1 Flow chart of sub-steps of step S102; Figure 3 Schematically shows Figure 1 Flow chart of sub-steps of step S106; Figure 4 Schematically shows Figure 1 Flow chart of sub-steps of step S108; Figure 5 Schematically shows Figure 1 Flow chart of sub-steps of step S104; Figure 6 Another flow chart of the intelligent translation method according to the first embodiment of the present application is schematically shown; Figure 7A diagram schematically shows an application example of the intelligent translation method according to the first embodiment of the present application; Figure 8 A diagram schematically shows an application example of the intelligent translation method according to the first embodiment of the present application; Fig. 9 A diagram schematically shows an application example of the intelligent translation method according to the first embodiment of the present application; Fig.10 A diagram schematically shows an application example of the intelligent translation method according to the first embodiment of the present application; Fig.11 A block diagram of an intelligent translation device according to the second embodiment of the present application is schematically shown; and Fig.12 The hardware architecture diagram of the computer device according to the third embodiment of the present application is schematically shown. DETAILED DESCRIPTION

[0018] In order to make the purpose, technical solutions and advantages of the present application more clearly understood, the present application is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not intended to limit the present application. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in the field without making creative work are within the scope of protection of the present application.

[0019] It should be noted that the descriptions involving "first", "second", etc. in the embodiments of the present application are only for descriptive purposes and cannot be understood as indicating or implying their relative importance or implicitly indicating the number of technical features indicated. Therefore, the features defined as "first" and "second" may explicitly or implicitly include at least one of the features. In addition, the technical solutions between the various embodiments can be combined with each other, but they must be based on the ability of ordinary technicians in the field to implement them. When the combination of technical solutions is contradictory or cannot be implemented, it should be deemed that such combination of technical solutions does not exist and is not within the scope of protection required by this application.

[0020] In the description of the present application, it should be understood that the numerical labels before the steps do not indicate the order in which the steps are executed, but are only used to facilitate the description of the present application and to distinguish each step, and therefore should not be understood as a limitation on the present application.

[0021] First, the following terms are explained: Localization: The process of adapting a product, content or service to a specific region, language and cultural background. This includes adjusting the content, design and functionality to suit the cultural habits, legal regulations and market demands of the target region.

[0022] Optical Character Recognition: It is a technology that converts text in an image into editable text through computer technology. It is usually used to extract text information from scanned documents, photos or handwritten text.

[0023] OpenCV: is an open source computer vision and machine learning software library that provides a large number of image and video processing functions. It can be used for image recognition, feature extraction, target tracking, image restoration, machine learning and other tasks.

[0024] Secondly, in order to facilitate those skilled in the art to understand the technical solutions provided in the embodiments of the present application, the relevant technologies are described below: The traditional comic translation and localization process requires a lot of manual work, including text translation, adjustment of cultural elements, and modification of images. It requires not only high language skills, but also a deep understanding of the culture and audience of the target translation language. During the translation process, translators need to translate the text sentence by sentence and modify some images in the comics according to language and cultural differences in order to better connect with the target culture. However, the translation and localization process of comics is not only time-consuming and labor-intensive, but also easily affected by the translator's personal understanding and cultural background, resulting in low translation efficiency and poor localization effect, thus affecting the experience of the target audience.

[0025] To this end, the embodiment of the present application provides an intelligent translation technology solution. In this technical solution, by automatically identifying cultural elements in multimedia content and replacing them with target cultural elements that are easier to understand and accept, users from different countries and regions can better understand and integrate into the content when watching multimedia content. By automatically replacing cultural elements in multimedia content, the localization effect of multimedia content is efficiently improved, thereby improving the user's viewing experience. See below for details.

[0026] The technical solutions of the present application are described below through multiple embodiments. It should be noted that these embodiments can be implemented in a variety of different forms and should not be construed as being limited to the embodiments described here.

[0027] Embodiment 1 Figure 1 The flowchart of the intelligent translation method according to the first embodiment of the present application is schematically shown.

[0028] like Figure 1 As shown, the intelligent translation method may include steps S100 to S108, wherein: Step S100, obtaining multimedia content to be translated, wherein the multimedia content to be translated includes an image to be processed; Step S102, acquiring text content and cultural elements according to the image to be processed; Step S104, translating the text content according to the target translation language to obtain the target text content; Step S106, obtaining a target cultural element for replacing the cultural element according to the target translation language; and Step S108: acquiring translated target multimedia content based on the image to be processed, the target text content and the target cultural element.

[0029] The intelligent translation method provided in this embodiment automatically identifies cultural elements in multimedia content and replaces them with target cultural elements that are easier to understand and accept, so that users from different countries and regions can better understand and integrate into the content when watching multimedia content. By automatically replacing cultural elements in multimedia content, the localization effect of multimedia content is efficiently improved, thereby improving the user's viewing experience.

[0030] The following combination Figure 1 , each step in steps S100~S108 and other optional steps are explained in detail.

[0031] Step S100 , obtaining multimedia content to be translated, wherein the multimedia content to be translated includes an image to be processed.

[0032] In some embodiments, the multimedia content to be translated may be a content form that integrates multiple media elements (such as text, images, sounds, videos, etc.), such as comics, cartoons, TV series, or games, etc. Correspondingly, the image to be processed may refer to the comic image of each page in the comic or the image of each frame in the cartoon, TV series, or game.

[0033] Step S102 , according to the image to be processed, obtaining text content and cultural elements.

[0034] In some embodiments, the text content may be a dialogue box in a comic, a subtitle or character dialogue in a cartoon or film and television drama, etc. Cultural elements may be cultural elements with a specific cultural background in the image to be processed. Cultural elements may be symbols, objects, scenes, character costumes, buildings, etc. in the image, such as big red lanterns, Spring Festival couplets, and glutinous rice balls, as well as Christmas trees, Santa Claus, hamburgers, and Notre Dame de Paris. The text in the image to be processed can be identified and converted into an editable text format through optical character recognition (OCR) technology. At the same time, combined with the image recognition model, the cultural elements in the image to be processed are automatically identified and marked.

[0035] Step S104 , according to the target translation language, the text content is translated to obtain the target text content.

[0036] In some embodiments, the target translation language may be Chinese, English, French, Japanese, or Spanish, etc. The text content may be text content in a comic, for example, a dialogue or narration text between characters in the comic.

[0037] By translating the text content to be translated from the source language into the target translation language, the target text content that conforms to the expression habits of the target translation language can be obtained, so that the multimedia content can be understood by audiences using different languages.

[0038] Step S106 , according to the target translation language, obtaining a target cultural element for replacing the cultural element.

[0039] According to the target translation language (i.e. the cultural background of the audience), target cultural elements suitable for replacing the original cultural elements are obtained, so that the cultural elements originally associated with the source language can be replaced with target cultural elements that are easier for the audience to understand and accept, thereby ensuring the adaptability and attractiveness of multimedia content in different cultural backgrounds and efficiently realizing the localization of multimedia content.

[0040] In some embodiments, a cultural element may refer to an image with unique meaning associated with a certain cultural background in multimedia content. For example, in a Japanese comic, a cultural element associated with Japanese culture such as a "Japanese-style dining table" may appear. If the Japanese comic is to be translated for foreign audiences, the cultural element may be replaced with a target cultural element similar to the "Japanese-style dining table" in the cultural background associated with the target translation language, such as a flowing banquet.

[0041] Step S108 , based on the image to be processed, the target text content and the target cultural elements, obtaining the translated target multimedia content.

[0042] By integrating the original content of the image to be processed, the translated target text content and the adjusted target cultural elements, the target multimedia content after translation and localization is generated, so that the target audience can better understand the target multimedia content. For example, the target text content replaces the original text content, and the target cultural elements replace the original cultural elements. There are many ways to replace, such as directly covering with a mask map, or first clearing the original text content / cultural elements, and then covering the embedding with the target text content / target cultural elements.

[0043] For example, in a Chinese comic, there is a scene depicting a host inviting friends to his home for tea, with traditional "tea sets" on the table, such as exquisite purple clay teapots, teacups, and tea leaves. "Tea sets" are typical Chinese cultural elements and symbolize China's tea culture. When the target translation language of this comic is English (i.e. the audience is Western audiences), since traditional Chinese tea sets are not common in Western culture, directly retaining "purple clay teapots" or "Chinese teacups" may make the audience feel unfamiliar, and thus cannot be well integrated into the comic plot. Therefore, the "purple clay teapots" and "teacups" in the comic can be replaced with target cultural elements that are more in line with Western culture, such as "coffee cups" and "coffee pots".

[0044] In an optional embodiment, if Figure 2 As shown, step S102 includes S200~S204: Step S200: input the image to be processed into a pre-trained cultural element recognition model.

[0045] In some embodiments, images representing different cultural backgrounds can be collected as pre-training data for the cultural element recognition model, for example, images representing Chinese culture such as red lanterns, tea sets, and dragon-shaped carvings. Further, using the pre-training data, a convolutional neural network model (such as ResNet, EfficientNet, etc.) can be used for training, combined with a region extraction network (such as Faster R-CNN) to identify and locate cultural elements in the image.

[0046] By inputting the image to be processed into a pre-trained model for analysis, the cultural elements contained in the image to be processed can be identified. The cultural element recognition model can be trained using deep learning methods, so that the cultural element recognition model can recognize cultural elements associated with the cultural background in the image to be processed, such as unique buildings, clothing or objects. The cultural element recognition model can realize automatic recognition of cultural elements and provide accurate and efficient data support for the localization of multimedia content.

[0047] Step S202: determining one or more cultural elements from the image to be processed by using the cultural element recognition model.

[0048] Exemplarily, a comic image can be input into a cultural element recognition model to identify cultural elements in the comic. Specifically, the comic image can be preprocessed using tools such as OpenCV, for example, scaling it to the cultural element recognition model input size of 224×224 pixels, and expanding the dimension to adapt to the input format of the cultural element recognition model. Through the cultural element recognition model, the type, confidence, and location information of the cultural element can be output. For example, the comic image contains a scene with a hanging red lantern drawn in the scene. The cultural element recognition model can output the category of the cultural element as a red lantern, with a confidence of 92%, and the location information: [x_min=50, y_min=100, x_max=150, y_max=200], thereby determining the cultural elements in the image to be processed.

[0049] In this embodiment, by inputting the image to be processed into a pre-trained cultural element recognition model, the cultural elements in the image are automatically extracted, thereby improving the efficiency and accuracy of cultural element recognition and avoiding omissions or deviations that may be caused by manual localization. Based on the pre-trained cultural element recognition model, cultural elements related to the cultural background in the image to be processed can be quickly located, thereby ensuring the adaptability of multimedia content in different cultural environments and improving the user's viewing experience.

[0050] In an optional embodiment, if Figure 3 As shown, step S106 also includes: Step S300: determining target cultural information corresponding to the target translation language according to the target translation language.

[0051] Step S302: based on the target cultural information and the cultural elements, the target cultural elements are acquired from a preset cultural knowledge base.

[0052] Different target translation languages ​​correspond to different cultural backgrounds. Audiences under different cultural backgrounds have different acceptance and understanding of multimedia content. Therefore, it is necessary to analyze the cultural background of the target translation language to determine how to select target cultural elements that are adapted to the corresponding cultural background to adjust the original cultural elements, so as to better meet the understanding needs of the target audience. For example, when the target translation language is Chinese, the corresponding target cultural information is Chinese culture, English corresponds to Western culture, Japanese corresponds to Japanese culture, Korean corresponds to Korean culture, etc.

[0053] In some embodiments, the preset cultural knowledge base may include cultural elements under various cultural backgrounds, for example, food elements (such as sushi, hamburgers, or dumplings), clothing elements (such as cheongsam, kimono, suits, or hanbok), or architectural elements (such as the Eiffel Tower, the Great Wall, and the Statue of Liberty).

[0054] For example, in a Japanese comic, there is a scene showing "sushi" on a "Japanese-style dining table". First, the type of target cultural element can be determined based on the type of cultural element and the target cultural information. For example, "sushi" corresponds to Chinese food culture. Furthermore, the target cultural elements related to "sushi" in Chinese food culture can be searched in a preset cultural knowledge base. Based on the target cultural information and cultural elements, the cultural knowledge base can output, for example, "dim sum" or "dumplings" as alternatives (i.e., target cultural elements), and finally replace the cultural elements in the original comic with target cultural elements that are more in line with Chinese culture.

[0055] In this embodiment, the corresponding target cultural information is determined by the target translation language, thereby ensuring that the cultural background of the target audience can be fully considered when replacing cultural elements. By searching for target cultural elements that match the target cultural information in the preset cultural knowledge base, the target multimedia content is closer to the cognition and acceptance of the target audience at the visual and cultural levels, thereby improving the localization effect of the target multimedia content.

[0056] In an optional embodiment, if Figure 4 As shown, step S108 includes: Step S400: obtaining size information of the text content in the image to be processed.

[0057] Step S402, adjusting the attributes of the target text content according to the character length of the target text content and the size information of the text content; wherein the attributes of the target text content include font style and / or font size.

[0058] In the process of translating multimedia content, the text content needs to be replaced with the target text content corresponding to the target translation language. By obtaining the size information of the original text content, it can be ensured that the replaced target text content is adapted to the position of the original text content, so as not to affect the overall image layout and visual effects.

[0059] For example, when replacing a cultural element, if the cultural element contains text content (for example, a logo, signboard or decorative text in the cultural element), not only the cultural element needs to be replaced, but also the translated target text content needs to be added to the target cultural element. Ensure that the font, size, line spacing, etc. of the replaced text are consistent with the text content on the original cultural element, so as to ensure the coordination of the image layout. By obtaining the size information of the text content in the image to be processed, it can be ensured that the layout of the translated target text content matches the original text content, avoiding the situation where the text is too large or too small and is not suitable for visual display. For example, in an American comic, a certain picture contains the cultural element of a road sign, which has a green background and white fonts and says "Main Street". In the process of translating into Chinese, the road sign itself needs to be replaced with a style that is more in line with the Chinese cultural background (such as a blue background with white fonts and arrows or pinyin information), and the original English "Main Street" (text content) is translated into Chinese "Main Street" (target text content). Since the length of the target text content is different from the text content, it can be adapted and adjusted according to the length of the translated text and the size of the cultural elements, ultimately ensuring that the replaced localized road signs are visually consistent with Chinese culture. In addition, the font style of the target text content can also be appropriately adjusted to better suit the reading habits of Chinese audiences. For example, Chinese can use Kaiti or a font style that is more suitable for road signs.

[0060] In this embodiment, the size information of the text content in the image to be processed is obtained, and the target text content is dynamically adjusted according to the character length of the target text content, so as to achieve coordinated adaptation between the target text content and the image to be processed. By adjusting the font style and / or font size, the target text content is ensured to be consistent with the original image in visual effect, avoiding the problem of typesetting disorder caused by text length differences, enhancing the readability and aesthetics of multimedia content, and further improving the user's reading experience.

[0061] In an optional embodiment, if Figure 5 As shown, step S104 includes S500~S504: Step S500: translating the plurality of sentences based on the target translation language and the text translation model to obtain a plurality of target sentences.

[0062] In some embodiments, the text translation model can be a neural machine translation (NMT) model implemented based on PyTorch. The NMT model achieves high-precision translation from the source language to the target translation language through an encoder-decoder architecture, in which the encoder is used to extract semantic information from the source sentence, and the decoder generates sentences in the target language based on the output of the encoder. During the training process, the NMT model can combine specific vocabulary in a specific field (such as onomatopoeia common in the comics field) and a vocabulary of expressions to further optimize the model's processing capabilities for unique expressions in a specific field. In addition, in order to improve the fluency of the translation results and adapt to the cultural background of the target language, post-processing steps can also be combined. For example, the language model is used to optimize the grammar and style of the translation results to ensure that the output sentences conform to the grammatical rules and cultural habits of the target language, so that the translation results are more natural and fluent.

[0063] Step S502: Obtain context information of each target sentence.

[0064] The context information of the target sentence can include the content of the sentences before and after the target sentence, the tone of the character, the cultural background and environmental information in the comics, or the emotional tone of the text. Based on the context information of the target sentence, the text translation model can improve its ability to grasp the meaning and tone of the target sentence and avoid semantic conflicts between the sentences before and after.

[0065] Step S504, optimizing and adjusting the target sentence according to the context information of the target sentence; wherein the target text content includes: a plurality of target sentences corresponding to the plurality of the sentences and after being optimized and adjusted.

[0066] Based on the context information of the target sentence, the text translation model can optimize and adjust each target sentence so that the target sentence conforms to the grammar, style and cultural characteristics of the target language, thereby improving the accuracy of the translation results. For example, in the translation of a Chinese comic, character A uses a humorous slang in a certain scene. According to the context information, the text translation model can obtain the emotional expression tendency of the context and translate the slang into a slang with similar meaning in English, rather than directly translating it into the literal meaning. It ensures that the target sentence not only retains the original sense of humor, but also conforms to the context and cultural characteristics of the target language. By optimizing and adjusting each target sentence through context information, the generated target text content is more natural and fluent, and closer to the expression habits of the target translation language.

[0067] In this embodiment, by obtaining the context information of each target sentence, the target sentence is optimized and adjusted to better adapt the target sentence to the expression habits and cultural characteristics of the target language, thereby avoiding the semantic conflict or inconsistent style between the upper and lower sentences, thereby further improving the accuracy of the target text content.

[0068] In an optional embodiment, if Figure 6 As shown, the method also includes: Step S600: Acquire feedback data for the target multimedia content.

[0069] In some embodiments, feedback data can be collected through user interaction, questionnaires, scoring systems, etc., and used to evaluate the translation quality, cultural adaptability, and overall user satisfaction of the target multimedia content. For example, users can provide feedback through the application, such as scoring the accuracy of a translation, proposing a suggestion for replacing a cultural element, or marking the content in the translation that does not conform to the target culture. Feedback data can serve as a key data source for optimizing translation and / or localization models to further improve translation and localization results.

[0070] Step S602: adjusting the cultural element recognition model and / or the text translation model based on the feedback data.

[0071] By way of example, if the feedback data indicates that the replacement of cultural elements does not conform to the habits or acceptance of the target culture, the cultural element recognition model can adjust the recognition accuracy of the cultural elements based on the feedback, or update the cultural background information in the cultural knowledge base. Similarly, if the feedback received by the text translation model indicates that the translation of certain sentences fails to accurately convey the meaning of the original text or does not conform to the expression habits of the target language, the translation model can further optimize its translation strategy based on the feedback data, including adjusting the translation style, vocabulary selection or grammatical structure. By continuously iterating and optimizing the model, the quality of the translation results is gradually improved, and the adaptability of the cultural element recognition model and / or the text translation model to diverse cultural backgrounds is enhanced.

[0072] In this embodiment, by obtaining feedback data for the target multimedia content, the cultural element recognition model and / or the text translation model are adjusted, thereby optimizing the recognition accuracy or translation accuracy of the model according to the feedback, thereby ensuring that the translation result is not only more consistent with the grammar and cultural background of the target language, but also further improving the understanding and acceptance of the target audience.

[0073] In order to make this application easier to understand, the following Figures 7 to 10 An exemplary application is provided.

[0074] S11, input the text content in the comic into the text translation module, and automatically obtain the translation result of the text content based on the neural machine translation engine.

[0075] S12, optimizing and adjusting the translation result based on the context-aware translation adjustment mechanism.

[0076] S13, inputting the comic into an image processing module, obtaining cultural elements in the comic, and adjusting the cultural elements to obtain an adjusted comic image.

[0077] S14, combining the translation result and the comic image after adjusting the cultural elements to obtain a translated and localized comic.

[0078] S15, uploading the translated and localized comics for users to browse.

[0079] S16, obtaining feedback information from the user after browsing, optimizing the text translation module and the image processing module based on the feedback information, and outputting the comic with adjusted translation results and localized results as the final output.

[0080] In this exemplary application, by automatically identifying cultural elements in multimedia content and replacing them with target cultural elements that are easier to understand and accept, users from different countries and regions can better understand and integrate into the content when watching multimedia content. By automatically replacing cultural elements in multimedia content, the localization effect of multimedia content (such as comics) is effectively improved, thereby improving the user's viewing experience.

[0081] Embodiment 2 Fig.11 The block diagram of the intelligent translation device according to the second embodiment of the present application is schematically shown. The device can be divided into one or more program modules, one or more program modules are stored in a storage medium, and are executed by one or more processors to complete the embodiment of the present application. The program module referred to in the embodiment of the present application refers to a series of computer program instruction segments that can complete specific functions. The following description will specifically introduce the functions of each program module in this embodiment. Fig.11 As shown, the device 1600 may include: an acquisition module 1610 and a translation module 1620, wherein: An acquisition module 1610 is used to acquire multimedia content to be translated, wherein the multimedia content to be translated includes an image to be processed; and Acquire text content and cultural elements according to the image to be processed; A translation module 1620, configured to translate the text content according to a target translation language to obtain a target text content; The acquisition module 1610 is further used to acquire a target cultural element for replacing the cultural element according to the target translation language; and Based on the image to be processed, the target text content and the target cultural element, the translated target multimedia content is acquired.

[0082] In an optional embodiment, obtaining text content and cultural elements according to the image to be processed includes: Inputting the image to be processed into a pre-trained cultural element recognition model; and One or more cultural elements are determined from the image to be processed by using the cultural element recognition model.

[0083] In an optional embodiment, obtaining a target cultural element for replacing the cultural element according to the target translation language includes: According to the target translation language, determining the target cultural information corresponding to the target translation language; Based on the target cultural information and the cultural elements, the target cultural elements are acquired from a preset cultural knowledge base.

[0084] In an optional embodiment, obtaining the translated target multimedia content based on the image to be processed, the target text content and the target cultural element includes: Obtaining size information of the text content in the image to be processed; According to the character length of the target text content and the size information of the text content, the attributes of the target text content are adjusted; wherein the attributes of the target text content include font style and / or font size.

[0085] In an optional embodiment, the text content includes a plurality of sentences; and translating the text content according to a target translation language to obtain the target text content includes: Based on the target translation language and the text translation model, the plurality of sentences are translated to obtain a plurality of target sentences; Obtaining context information of each of the target sentences; Optimizing and adjusting the target sentence according to the context information of the target sentence; The target text content includes: a plurality of target sentences corresponding to the plurality of sentences and after being optimized.

[0086] In an optional embodiment, the device 1600 is further used for: Acquiring feedback data for the target multimedia content; Based on the feedback data, the cultural element recognition model and / or the text translation model is adjusted.

[0087] Embodiment 3 Fig.12The schematic diagram of the hardware architecture of a computer device 10000 suitable for implementing the intelligent translation method according to the third embodiment of the present application is shown schematically. In some embodiments, the computer device 10000 may be a terminal device such as a smart phone, a wearable device, a tablet computer, a personal computer, a vehicle terminal, a game console, a virtual device, a workbench, a digital assistant, a set-top box, a robot, etc. In other embodiments, the computer device 10000 may be a rack server, a blade server, a tower server, or a cabinet server (including an independent server, or a server cluster composed of multiple servers), etc. Fig.12 As shown, the computer device 10000 includes but is not limited to: a memory 10010, a processor 10020, and a network interface 10030 that can communicate with each other through a system bus. Among them: The memory 10010 includes at least one type of computer-readable storage medium, and the readable storage medium includes flash memory, hard disk, multimedia card, card-type memory (such as SD or DX memory), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic memory, magnetic disk, optical disk, etc. In some embodiments, the memory 10010 can be an internal storage module of the computer device 10000, such as a hard disk or memory of the computer device 10000. In other embodiments, the memory 10010 can also be an external storage device of the computer device 10000, such as a plug-in hard disk equipped on the computer device 10000, a smart memory card (Smart Media Card, SMC), a secure digital (Secure Digital, SD) card, a flash card (Flash Card), etc. Of course, the memory 10010 can also include both the internal storage module of the computer device 10000 and its external storage device. In this embodiment, the memory 10010 is generally used to store the operating system and various application software installed in the computer device 10000, such as program codes of the intelligent translation method, etc. In addition, the memory 10010 can also be used to temporarily store various data that have been output or will be output.

[0088] In some embodiments, the processor 10020 may be a central processing unit (CPU), a controller, a microcontroller, a microprocessor, or other chips. The processor 10020 is generally used to control the overall operation of the computer device 10000, such as performing control and processing related to data interaction or communication with the computer device 10000. In this embodiment, the processor 10020 is used to run the program code stored in the memory 10010 or process data.

[0089] The network interface 10030 may include a wireless network interface or a wired network interface, and the network interface 10030 is generally used to establish a communication link between the computer device 10000 and other computer devices. For example, the network interface 10030 is used to connect the computer device 10000 to an external terminal through a network, and to establish a data transmission channel and a communication link between the computer device 10000 and the external terminal. The network may be a wireless or wired network such as an intranet, the Internet, the Global System of Mobile communication (GSM), Wideband Code Division Multiple Access (WCDMA), 4G network, 5G network, Bluetooth, Wi-Fi, etc.

[0090] It should be pointed out that Fig.12 Only a computer device having components 10010 - 10030 is shown, but it should be understood that implementation of all of the components shown is not a requirement, and more or fewer components may alternatively be implemented.

[0091] In this embodiment, the intelligent translation method stored in the memory 10010 may also be divided into one or more program modules and executed by one or more processors (such as the processor 10020) to complete the embodiment of the present application.

[0092] Embodiment 4 An embodiment of the present application further provides a computer-readable storage medium having a computer program stored thereon, wherein when the computer program is executed by a processor, the steps of the intelligent translation method in the embodiment are implemented.

[0093] In this embodiment, the computer-readable storage medium includes flash memory, hard disk, multimedia card, card-type memory (for example, SD or DX memory, etc.), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic memory, disk, optical disk, etc. In some embodiments, the computer-readable storage medium can be an internal storage unit of a computer device, such as a hard disk or memory of the computer device. In other embodiments, the computer-readable storage medium can also be an external storage device of a computer device, such as a plug-in hard disk equipped on the computer device, a smart memory card (Smart Media Card, SMC), a secure digital (Secure Digital, SD) card, a flash card (Flash Card), etc. Of course, the computer-readable storage medium can also include both the internal storage unit of the computer device and its external storage device. In this embodiment, the computer-readable storage medium is generally used to store an operating system and various application software installed on the computer device, such as the program code of the intelligent translation method in the embodiment, etc. In addition, the computer-readable storage medium can also be used to temporarily store various types of data that have been output or are to be output.

[0094] Embodiment 5 An embodiment of the present application also provides a computer program product, including a computer program, which implements the method in the above embodiment when executed by a processor.

[0095] Obviously, those skilled in the art should understand that the modules or steps of the above-mentioned embodiments of the present application can be implemented by general-purpose computer devices, they can be concentrated on a single computer device, or distributed on a network composed of multiple computer devices, optionally, they can be implemented by executable program codes of computer devices, so that they can be stored in a storage device and executed by the computer device, and in some cases, the steps shown or described can be executed in a different order from that herein, or they can be made into individual integrated circuit modules, or multiple modules or steps therein can be made into a single integrated circuit module for implementation. In this way, the embodiments of the present application are not limited to any specific combination of hardware and software.

[0096] It should be noted that the above are only preferred embodiments of the present application, and the patent protection scope of the present application is not limited thereto. Any equivalent structure or equivalent process transformation made using the contents of the specification and drawings of the present application, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present application.

Claims

1. An intelligent translation method, characterized in that: The method comprises: Acquire multimedia content to be translated, wherein the multimedia content to be translated includes an image to be processed; Acquire text content and cultural elements according to the image to be processed; According to the target translation language, the text content is translated to obtain the target text content; According to the target translation language, obtaining a target cultural element for replacing the cultural element; and Based on the image to be processed, the target text content and the target cultural element, the translated target multimedia content is acquired.

2. The method according to claim 1, characterized in that According to the image to be processed, text content and cultural elements are obtained, including: Inputting the image to be processed into a pre-trained cultural element recognition model; and One or more cultural elements are determined from the image to be processed by using the cultural element recognition model.

3. The method according to claim 1, characterized in that: According to the target translation language, obtaining a target cultural element for replacing the cultural element includes: According to the target translation language, determining the target cultural information corresponding to the target translation language; Based on the target cultural information and the cultural elements, the target cultural elements are acquired from a preset cultural knowledge base.

4. The method according to claim 3, characterized in that Based on the image to be processed, the target text content and the target cultural element, the translated target multimedia content is obtained, including: Obtaining size information of the text content in the image to be processed; According to the character length of the target text content and the size information of the text content, the attributes of the target text content are adjusted; wherein the attributes of the target text content include font style and / or font size.

5. The method according to any one of claims 2 to 4, characterized in that: The text content includes a plurality of sentences; according to the target translation language, the text content is translated to obtain the target text content, including: Based on the target translation language and the text translation model, the plurality of sentences are translated to obtain a plurality of target sentences; Obtaining context information of each of the target sentences; Optimizing and adjusting the target sentence according to the context information of the target sentence; The target text content includes: a plurality of target sentences corresponding to the plurality of sentences and after being optimized.

6. The method according to claim 5, characterized in that The method further comprises: Acquiring feedback data for the target multimedia content; Based on the feedback data, the cultural element recognition model and / or the text translation model is adjusted.

7. An intelligent translation device, characterized in that: The device comprises: an acquisition module, configured to acquire multimedia content to be translated, wherein the multimedia content to be translated includes an image to be processed; and Acquire text content and cultural elements according to the image to be processed; A translation module, used for translating the text content according to a target translation language to obtain a target text content; The acquisition module is further used to acquire a target cultural element for replacing the cultural element according to the target translation language; and Based on the image to be processed, the target text content and the target cultural element, the translated target multimedia content is acquired.

8. A computer device, characterized in that: include: at least one processor; and a memory communicatively connected to the at least one processor; wherein: The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 6.

9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer instructions, and when the computer instructions are executed by a processor, the method according to any one of claims 1 to 6 is implemented.

10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to claims 1 to 6 are implemented.