Translation method, translation device, storage medium and electronic device

CN115860014BActive Publication Date: 2026-08-21IFLYTEK CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211472758.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-17
Publication Date
2026-08-21
Estimated Expiration
2042-11-17

AI Technical Summary

Technical Problem

然而,通常情况下,拍摄得到的图像中可能同时包含物体和文本,并且,文本在物体上,这种情况下,具有拍照翻译功能的应用程序很难从中选择出用户实际需要翻译的翻译目标,从而给用户带来了极大的不便

Benefits of technology

[0015] The translation method provided in this application first acquires an image to be translated, wherein the image includes at least one translatable object, and the translatable object contains translatable text. Further, the actual translation target desired by the user is determined from the translatable object and the translatable text, and the actual translation target is translated to obtain the translation result. Through the solution in this application embodiment, the actual translation target desired by the user can be translated without requiring the user to accurately acquire the image to be translated, thereby obtaining the desired translation result. This method is not only accurate but also convenient and fast.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115860014B_ABST
    Figure CN115860014B_ABST
Patent Text Reader

Abstract

The application provides a translation method, a translation device, a storage medium and an electronic device, and relates to the technical field of image processing. The translation method comprises the following steps: acquiring a to-be-translated image, the to-be-translated image comprising at least one translatable object, and the translatable object comprising translatable text; determining an actual translation target from the translatable object and the translatable text; and translating the actual translation target to obtain a translation result. Through the scheme in the application, the actual translation target that the user wants to translate is translated without the user accurately collecting the to-be-translated image, so that the translation result that the user wants to obtain is obtained, which is not only accurate in translation, but also convenient and fast.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image processing technology, specifically to a translation method, translation device, storage medium, and electronic device. Background Technology

[0002] With the popularization of hardware devices such as translation machines and learning machines, applications (APPs) with photo translation functions have become well-known. Correspondingly, photo translation functions have greatly facilitated people's daily learning and work communication.

[0003] In practical applications of photo translation functions, it is usually necessary to use a camera device to capture images of the area containing the translation target. However, the captured image may often contain both objects and text, with the text on top of the object. In such cases, it is difficult for the photo translation application to select the actual target that the user needs to translate, causing significant inconvenience to the user. Summary of the Invention

[0004] To address the aforementioned technical problems, this application is proposed. Embodiments of this application provide a translation method, a translation apparatus, a storage medium, and an electronic device.

[0005] In a first aspect, one embodiment of this application provides a translation method, the method comprising: acquiring an image to be translated, the image to be translated including at least one translatable object, and the translatable object containing translatable text; determining an actual translation target from the translatable object and the translatable text; translating the actual translation target to obtain a translation result.

[0006] In conjunction with the first aspect, in some implementations of the first aspect, determining the actual translation target from the translatable object and the translatable text includes: determining the outline-complete object and the outline-complete text from the translatable object and the translatable text based on the image information of the image to be translated; if the number of outline-complete objects and outline-complete texts is 1, then determining the actual translation target from the outline-complete objects and outline-complete texts based on their respective position information in the image to be translated.

[0007] In conjunction with the first aspect, in some implementations of the first aspect, the location information includes proportion or distance from the center position of the image to be translated. Based on the location information of the outline complete object and the outline complete text in the image to be translated, the actual translation target is determined from the outline complete object and the outline complete text, including: determining the outline complete object and the outline complete text whose proportion is greater than a preset proportion threshold as the actual translation target; or, determining the outline complete object and the outline complete text whose distance from the center position of the image to be translated is less than a preset distance threshold as the actual translation target.

[0008] In conjunction with the first aspect, in some implementations of the first aspect, the location information includes proportion, distance to the center position of the image to be translated, and distance between the outline-complete object and the outline-complete text. Based on the location information of the outline-complete object and the outline-complete text in the image to be translated, the actual translation target is determined from the outline-complete object and the outline-complete text, including: calculating the translation score of the outline-complete object based on at least one of the proportion of the outline-complete object in the image to be translated, the distance to the center position of the image to be translated, and the distance between the outline-complete object and the outline-complete text; calculating the translation score of the outline-complete text based on at least one of the proportion of the outline-complete text in the image to be translated, the distance to the center position of the image to be translated, and the distance between the outline-complete object and the outline-complete text; and determining the actual translation target from the outline-complete object and the outline-complete text based on the translation scores of the outline-complete object and the outline-complete text.

[0009] In conjunction with the first aspect, in certain implementations of the first aspect, determining the actual translation target from the translatable object and the translatable text includes: determining the outline-complete object and the outline-complete text from the translatable object and the translatable text based on image information of the image to be translated; and determining the actual translation target from the outline-complete object and the outline-complete text based on user information.

[0010] In conjunction with the first aspect, in some implementations of the first aspect, the user information includes translation vocabulary information. Based on the user information, the actual translation target is determined from the outline complete object and the outline complete text, including: determining the object translation word corresponding to the outline complete object and the text translation word corresponding to the outline complete text; determining whether the object translation word and the text translation word are within the translation vocabulary information; and determining the outline complete object or outline complete text within the translation vocabulary information as the actual translation target.

[0011] In conjunction with the first aspect, in certain implementations of the first aspect, determining the actual translation target from the translatable object and the translatable text includes: sending a translation selection prompt to the user so that the user selects the actual translation target from the translatable object and the translatable text; and determining the actual translation target from the translatable object and the translatable text in response to the user's translation selection operation.

[0012] Secondly, one embodiment of this application provides a translation apparatus, which includes: an acquisition module for acquiring an image to be translated, the image to be translated including at least one translatable object, and the translatable object containing translatable text; a determination module for determining an actual translation target from the translatable object and the translatable text; and a translation module for translating the actual translation target to obtain a translation result.

[0013] Thirdly, one embodiment of this application provides a computer-readable storage medium storing a computer program for performing the translation method described in the first aspect.

[0014] Fourthly, one embodiment of this application provides an electronic device, the electronic device comprising: a processor; a memory for storing processor-executable instructions; the processor being configured to perform the translation method described in the first aspect.

[0015] The translation method provided in this application first acquires an image to be translated, wherein the image includes at least one translatable object, and the translatable object contains translatable text. Further, the actual translation target desired by the user is determined from the translatable object and the translatable text, and the actual translation target is translated to obtain the translation result. Through the solution in this application embodiment, the actual translation target desired by the user can be translated without requiring the user to accurately acquire the image to be translated, thereby obtaining the desired translation result. This method is not only accurate but also convenient and fast. Attached Figure Description

[0016] The above and other objects, features, and advantages of this application will become more apparent from the more detailed description of the embodiments of this application in conjunction with the accompanying drawings. The drawings are provided to further illustrate the embodiments of this application and form part of the specification. They are used together with the embodiments of this application to explain this application and do not constitute a limitation thereof. In the drawings, the same reference numerals generally represent the same components or steps.

[0017] Figure 1 The diagram shown is a schematic representation of an implementation environment to which the embodiments of this application apply.

[0018] Figure 2 The diagram shown is a flowchart illustrating a translation method provided in an exemplary embodiment of this application.

[0019] Figure 3 The diagram shown is a flowchart illustrating the process of determining the actual translation target provided in another exemplary embodiment of this application.

[0020] Figure 4 The diagram shown is a flowchart illustrating the process of determining the actual translation target according to another exemplary embodiment of this application.

[0021] Figure 5 The diagram shown is a flowchart illustrating the process of determining the actual translation target provided in another exemplary embodiment of this application.

[0022] Figure 6 The diagram shown is a flowchart illustrating the process of determining the actual translation target provided in another exemplary embodiment of this application.

[0023] Figure 7 The diagram shown is a flowchart illustrating the process of determining the actual translation target provided in another exemplary embodiment of this application.

[0024] Figure 8 The diagram shown is a structural schematic of a translation apparatus provided in an exemplary embodiment of this application.

[0025] Figure 9 The diagram shown is a structural schematic of an electronic device provided in an embodiment of this application. Detailed Implementation

[0026] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0027] Application Overview

[0028] With the widespread use of hardware devices such as translation machines and learning machines, applications (APPs) with photo translation functions have become well-known, greatly facilitating people's daily learning and work communication. In related photo translation methods, users typically take a picture of the target area they want to translate and upload the result to the translation application to obtain the translation. However, if the picture contains both a translatable object and translatable text, and the translatable text is on top of the translatable object, the translation application cannot confirm the actual target the user wants to translate. This leads to users needing to repeatedly take pictures of the target area to ensure that the picture contains only the translatable target they want to translate.

[0029] For example, if a user takes a picture of a shirt with the words "I love China" on it, the translation app cannot choose to translate either "shirt" or "I love China." Similarly, if a user takes a picture of a pavilion with couplets on either side, the app cannot choose to translate either "pavilion" or the text on the couplets. Likewise, if a user takes a picture of a book, the app cannot choose to translate either "book" or the text on the book.

[0030] Based on the above, this application provides a translation method. First, an image to be translated is acquired, which includes at least one translatable object and translatable text. The actual translation target is determined from the translatable object and the translatable text. The actual translation target is then translated to obtain the translation result. Through the solution in this application, the actual translation target required by the user can be ultimately determined from the translatable object and the translatable text, avoiding the need for repeated photography. Furthermore, it intelligently obtains the actual translation target the user wants to translate, and then translates the actual translation target to obtain the desired translation result. This method is not only easy to implement but also accurate, convenient, and fast.

[0031] Exemplary application scenarios

[0032] The translation method provided in this application can be executed by an electronic device, which can be a terminal, such as a smartphone, tablet computer, or desktop computer. Alternatively, the electronic device can also be a server, such as a standalone physical server, a server cluster consisting of multiple servers, or a cloud server capable of cloud computing and cloud translation.

[0033] This application provides a schematic diagram of the implementation environment for a translation method. In this embodiment, the electronic device is a server. Specifically, as shown... Figure 1 As shown, this implementation environment includes terminal 11 and server 12, and there is a communication connection between terminal 11 and server 12.

[0034] Terminal 11 can be a smartphone, tablet, desktop computer, etc. Terminal 11 can acquire images to be translated taken by the user using its own camera system. The images to be translated contain both translatable objects and translatable text, with the translatable text appearing on the translatable objects. A translation application is deployed on server 12. Server 12 can be a physical machine or a virtual machine, and there can be one or more servers. This embodiment does not limit the type or number of servers.

[0035] Furthermore, the implementation environment of the translation method in this application can also include only a terminal, on which a translation application is deployed. In this case, the terminal uses its own camera system to acquire the image to be translated taken by the user, then uses its own translation application to determine the actual translation target that the user wants to translate, and translates the actual translation target to obtain the translation result.

[0036] Exemplary methods

[0037] Figure 2 The diagram shown is a flowchart illustrating a translation method provided in an exemplary embodiment of this application. Figure 2 As shown in the embodiments of this application, the translation method includes the following steps.

[0038] Step S210: Obtain the image to be translated. The image to be translated includes at least one translatable object, and the translatable object contains translatable text.

[0039] Specifically, a translatable object refers to a translatable object contained in an image to be translated, such as a cup, a computer, a T-shirt, or shoes.

[0040] Step S220: Determine the actual translation target from the translatable object and the translatable text.

[0041] A translatable object may contain one or more translatable texts; this application does not limit the specific number of translatable texts. For example, an image to be translated contains one translatable object, and the translatable object contains three translatable texts. The actual translation target needs to be determined from the one translatable object and the three translatable texts.

[0042] Step S230: Translate the actual translation target to obtain the translation result.

[0043] For example, the actual translation target can be translated into English, Japanese, Korean, or other languages. This application does not limit the language of the translation result.

[0044] In this embodiment, an image to be translated is first acquired. The image includes at least one translatable object, and the translatable object contains translatable text. The actual translation target is determined from the translatable object and the translatable text. The actual translation target is then translated to obtain the translation result. The solution in this application allows for the final determination of the user's desired actual translation target from the translatable object and the translatable text, avoiding the need for repeated photography. Furthermore, it intelligently obtains the user's desired translation target and translates it to obtain the desired translation result. This approach is not only easy to implement but also accurate, convenient, and fast.

[0045] Figure 3 The diagram shown is a flowchart illustrating the process of determining the actual translation target according to another exemplary embodiment of this application. Figure 2 Extending from the illustrated embodiment Figure 3 The illustrated embodiment will be described in detail below. Figure 3 The illustrated embodiments and Figure 2 The differences between the embodiments shown are not repeated here, and the similarities are not repeated here.

[0046] Step S310: Based on the image information of the image to be translated, determine the outline-complete object and outline-complete text from the translatable object and translatable text.

[0047] Specifically, the criteria for determining whether an object or text has a complete outline are based on the outline integrity conditions set by the translation application. For example, if the size of a translatable object accounts for 80% or more of its empirical data, the outline of the translatable object is considered complete. Similarly, if the size of translatable text accounts for 80% or more of its empirical data, the outline of the translatable text is considered complete. Alternatively, the integrity of the outlines of translatable objects and translatable text can also be determined using an integrity detection model.

[0048] Step S320: If the number of contour-complete objects and contour-complete texts is 1, then the actual translation target is determined from the contour-complete objects and contour-complete texts based on their respective position information in the image to be translated.

[0049] Location information refers to the distribution of outlined objects and outlined text within the image to be translated. For example, the location information of an outlined object and the location information of outlined text can be determined by calculating the number and position of each pixel in the image to be translated; similarly, the location information of an outlined object can be determined by calculating the number and position of the smallest bounding rectangle of the outlined object in the image to be translated, and the location information of outlined text can be determined by calculating the number and position of the smallest bounding rectangle of the outlined text in the image to be translated.

[0050] In this embodiment, the first step is to identify the fully outlined objects and text contained in the image to be translated. Further, when there is only one fully outlined object and one fully outlined text, the actual translation target is determined based on the positional information of each object and text within the image. Positional information can, to some extent, represent the user's preference for the desired actual translation target. Therefore, the solution in this embodiment can more accurately determine the actual translation target, meeting the user's translation needs and making the determination of the actual translation target more intelligent, scientific, and rational.

[0051] Figure 4 The diagram shown is a schematic representation of a process for determining the actual translation target, provided in another exemplary embodiment of this application. Figure 3 Extending from the illustrated embodiment Figure 4 The illustrated embodiment will be described in detail below. Figure 4 The illustrated embodiments and Figure 3 The differences between the embodiments shown are not repeated here, and the similarities are not repeated here.

[0052] like Figure 4 As shown in this embodiment, the location information includes proportion or distance from the center position of the image to be translated. Based on the location information of the outline-complete object and the outline-complete text in the image to be translated, the actual translation target is determined from the outline-complete object and the outline-complete text, including the following steps.

[0053] Step S410: The objects with complete outlines and the text with complete outlines that have a proportion greater than a preset proportion threshold are identified as the actual translation targets.

[0054] For example, the proportion of a fully outlined object in the image to be translated is determined based on the number of pixels occupied by the fully outlined object in the image to be translated, and the total number of pixels contained in the image to be translated. Similarly, the proportion of fully outlined text in the image to be translated is determined based on the number of pixels occupied by the fully outlined text in the image to be translated, and the total number of pixels contained in the image to be translated.

[0055] For example, if the number of pixels occupied by the complete outline object in the image to be translated is 40, and the image to be translated contains 100 pixels, the calculated percentage of the complete outline text is 40%.

[0056] For example, the preset percentage threshold is 80%. If the percentage of a complete outline object is greater than 80%, the complete outline object is determined as the actual translation target; if the percentage of a complete outline text is greater than 80%, the complete outline text is determined as the actual translation target; if the percentages of both the complete outline object and the complete outline text are greater than 80%, both the complete outline object and the complete outline text are determined as the actual translation targets.

[0057] Step S420: The objects with complete outlines and the text with complete outlines that are less than the center position of the image to be translated are determined as the actual translation targets.

[0058] For example, the distance between the center position of the outlined object and the center position of the image to be translated is determined based on the pixel position of the outlined object in the image to be translated and the pixel position of the center point of the image to be translated. Similarly, the distance between the center position of the outlined text and the center position of the image to be translated is determined based on the pixel position of the outlined text in the image to be translated and the pixel position of the center point of the image to be translated.

[0059] For example, if the image to be translated is 10*10 pixels, meaning it contains 100 pixels, the position of each pixel in the image can be represented by (a, b), where a ∈ [1, 10] and b ∈ [1, 10]. Further, pixel (5, 5) can be determined as the center point of the image to be translated. Based on the pixel positions occupied by the complete outline object in the image to be translated, the pixel position of the center point of the complete outline object is calculated, and the distance between the center point of the complete outline object and the center position of the image to be translated is further calculated.

[0060] For example, the preset distance threshold can be specifically set according to the size of the image to be translated, as well as the size of the outlined object and the outlined text. Furthermore, the methods for determining distance and proportion mentioned in the embodiments of this application are merely examples; those skilled in the art can choose appropriate methods for calculating distance and proportion based on specific circumstances.

[0061] In this embodiment, the actual translation target is determined from the complete outline object and the complete outline text by determining the proportion of the complete outline object and the distance between them and the center position of the image to be translated, making the translation process more intelligent, streamlined, and convenient.

[0062] Figure 5 The diagram shown is a schematic representation of a process for determining the actual translation target provided in another exemplary embodiment of this application. Figure 3 Extending from the illustrated embodiment Figure 5 The illustrated embodiment will be described in detail below. Figure 5 The illustrated embodiments and Figure 3 The differences between the embodiments shown are not repeated here, and the similarities are not repeated here.

[0063] like Figure 5 As shown in this embodiment, the location information includes proportion, distance to the center of the image to be translated, and distance between the outlined object and the outlined text. Based on the location information of the outlined object and the outlined text in the image to be translated, the actual translation target is determined from the outlined object and the outlined text, including the following steps.

[0064] Step S510: Calculate the translation score of the outline-complete object based on at least one of the proportion of the outline-complete object in the image to be translated, the distance between the outline-complete object and the center position of the image to be translated, and the distance between the outline-complete text.

[0065] For example, an empirical threshold range [T1, T2] for contour-complete objects is determined, and the score data for contour-complete objects is further determined based on the empirical threshold range of contour-complete objects, the proportion of contour-complete objects in the image to be translated, the distance between the contour-complete objects and the center position of the image to be translated, and the distance between the contour-complete text.

[0066] The score data for a complete outline object can be determined using a scoring formula built into the translation application. For example, the score for a complete outline object... Where a1 represents the distance to the center of the image to be translated, a2 represents the proportion of the object with a complete outline in the image to be translated, and c represents the distance to the text with a complete outline.

[0067] Alternatively, a pre-trained neural network model can be used to determine the score data of an object with a complete outline. For example, the image to be translated can be directly input into the neural network model to obtain the score data of the object with a complete outline.

[0068] Step S520: Calculate the translation score of the outline-complete text based on at least one of the proportion of the outline-complete text in the image to be translated, the distance between the outline-complete text and the center position of the image to be translated, and the distance between the outline-complete objects.

[0069] Similarly, the empirical threshold range [K1, K2] of the contour-complete text is determined. Based on the empirical threshold range of the contour-complete text, the proportion of the contour-complete text in the image to be translated, the distance between the contour-complete text and the center position of the image to be translated, and the distance between the contour-complete object and the contour-complete object, the score data of the contour-complete object is further determined.

[0070] As described in step S510, the score data for the outlined complete text can be determined using a scoring formula built into the translation application. For example, the score for the outlined complete text... Where b1 represents the distance to the center of the image to be translated, b2 represents the proportion of the outlined text in the image to be translated, and c represents the distance to the outlined object.

[0071] Similarly, the image to be translated can be directly input into the neural network model mentioned in step S510 to obtain the score data of the complete outline text.

[0072] Step S530: Based on the translation scores of the outline-complete object and the outline-complete text, determine the actual translation target from the outline-complete object and the outline-complete text.

[0073] Specifically, the difference between the score data of the complete outline object and the score data of the complete outline text is determined. If the difference is greater than or equal to a preset difference threshold, the one with the higher score in the complete outline object and the complete outline text is determined as the actual translation target. Otherwise, the actual translation target can be determined from the complete outline object and the complete outline text based on the user's historical information.

[0074] For example, if the score of the complete outline object is 72 and the score of the complete outline text is 84, then the difference between the score of the complete outline object and the score of the complete outline text is equal to 6. The preset difference threshold is 10, so the complete outline text with a score of 84 can be identified as the actual translation target.

[0075] A user's historical information includes at least one of the following: translation preference information, translation vocabulary information, and context information.

[0076] Translation preference information refers to whether a user prefers translating objects or text. This preference information can be determined based on the user's historical interactions with the translation app. For example, if a user prefers using the translation app for foreign language learning through image recognition, or if the user has previously chosen object recognition for language learning, their translation preference will naturally lean towards translating objects.

[0077] The term "new vocabulary information" refers to the user's vocabulary dictionary stored on the electronic device that downloaded the translation application. Of course, the electronic device also stores the user's existing vocabulary. The more a word appears in the vocabulary dictionary as a target for translation, the greater its weight in the final translation.

[0078] Scene information refers to the user's current scene as depicted in the image to be translated, such as a travel scene or a conference scene. Different scenes have different translation weights for objects with complete outlines and for text with complete outlines.

[0079] For example, the user's historical information can be input into the trained network model to obtain the user's translation preference weights, and then calculated with the scores of the outline-complete object and the outline-complete text obtained in the aforementioned embodiments. Based on the calculation results, the actual translation target is finally determined.

[0080] For example, the score of a complete outline object is 75, and the score of a complete outline text is 80. The difference between the two scores is insufficient to determine the actual translation target. In this case, the translation weight calculated by the network model is 0.8 for the complete outline object and 0.6 for the complete outline text. Therefore, the final score of the complete outline object is 75 * 0.8 = 60, and the final score of the complete outline text is 48. The final score difference between the two is 12. The preset difference threshold is 10. At this point, the complete outline object can be determined as the final actual translation target.

[0081] In this embodiment, the actual translation target is determined based on the scores of the outlined object and the outlined text. Specifically, the closer the outlined object is to the center point of the image to be translated, the farther it is from the outlined text, and the closer its proportion is to the empirical threshold range for the object, the higher its score. Similarly, the closer the outlined text is to the center point of the image to be translated, the farther it is from the outlined object, and the closer its proportion is to the empirical threshold range for the text, the higher its score. The solution in this application allows for accurate determination of the actual translation target.

[0082] Figure 6 The diagram shown is a schematic representation of a process for determining the actual translation target provided in another exemplary embodiment of this application. Figure 2 Extending from the illustrated embodiment Figure 6 The illustrated embodiment will be described in detail below. Figure 6 The illustrated embodiments and Figure 2 The differences between the embodiments shown are not repeated here, and the similarities are not repeated here.

[0083] like Figure 6 As shown in the embodiments of this application, determining the actual translation target from the translatable object and the translatable text includes the following steps.

[0084] Step S610: Based on the image information of the image to be translated, determine the outline-complete object and outline-complete text from the translatable object and translatable text.

[0085] As shown above, it is possible to utilize Figure 3 The method in the illustrated embodiment determines outline-complete objects and outline-complete text from translatable objects and translatable text.

[0086] Step S620: Based on user information, determine the actual translation target from the outline-complete object and the outline-complete text.

[0087] Specifically, user information includes translation vocabulary information. Further, it determines the object translation words corresponding to the fully outlined object and the text translation words corresponding to the fully outlined text, and whether the object translation words and text translation words are within the translation vocabulary information. The fully outlined object or fully outlined text within the translation vocabulary information is then identified as the actual translation target.

[0088] In this embodiment, user information is introduced to determine the actual translation target, making the actual translation target more in line with the user's needs.

[0089] Figure 7 The diagram shown is a schematic representation of a process for determining the actual translation target provided in another exemplary embodiment of this application. Figure 2 Extending from the illustrated embodiment Figure 7 The illustrated embodiment will be described in detail below. Figure 7 The illustrated embodiments and Figure 2 The differences between the embodiments shown are not repeated here, and the similarities are not repeated here.

[0090] like Figure 7 As shown in the embodiments of this application, determining the actual translation target from the translatable object and the translatable text includes the following steps.

[0091] Step S710: Send a translation selection prompt to the user so that the user can select the actual translation target from the translatable objects and translatable text.

[0092] Specifically, the translation selection prompts can be text-based or voice-based.

[0093] For example, the text prompt could be "Do you want to translate ×× object / or ×× text?"

[0094] Step S720: In response to the user's translation selection operation, determine the translatable object and the actual translation target in the translatable text.

[0095] Specifically, users can select translations via voice or touchscreen.

[0096] In this embodiment of the application, the user can make a selection by translating the selection prompt information, and the user can also get the translation result they need. Moreover, this solution is simple and easy to implement.

[0097] In an exemplary embodiment of this application, determining the actual translation target from the translatable object and the translatable text further includes: using a translation target detection model to determine the actual translation target from the translatable object and the translatable text.

[0098] The translation detection model can output scores for both the translatable object and the translatable text, or it can directly output the final determined translation target.

[0099] For example, the image to be translated can be directly input into the translation target detection model, or the image to be translated and user information can be input into the translation target detection model simultaneously to obtain the scores corresponding to the translatable object and the translatable text, respectively. The translation application determines the actual translation target based on the score results and translates the actual translation target. Alternatively, the translation detection model can directly output the actual translation target corresponding to the image to be translated, and the translation application translates the actual translation target to obtain the translation result.

[0100] Furthermore, the user's translation selection information can be used as the label for the network model to be trained. By collecting the user's photo data and translation selection information as training data, the network model can be trained to obtain a translation object detection model.

[0101] In this embodiment, the translation target detection model can accurately and easily obtain the translation result corresponding to the actual translation target that the user wants to translate. Furthermore, the translation target detection model has a wide range of applications and can detect any image to be translated.

[0102] In an exemplary embodiment of this application, if the image to be translated contains multiple outlined objects and multiple outlined texts, or multiple outlined objects and one outlined text, or one outlined object and multiple outlined texts, then the outlined object with the largest proportion in the image to be translated is determined; the outlined text with the largest proportion in the image to be translated is determined; based on the outlined object with the largest proportion and the outlined text with the largest proportion, the target type that the user wants to translate is determined; and the outlined objects and outlined texts contained in the image to be translated that are of the same target type that the user wants to translate are determined as the actual translation target that the user wants to translate.

[0103] Following the method described in the previous example, the positional distribution data corresponding to each complete outline object is determined. Based on the positional distribution data of each complete outline object, the proportion of each complete outline object in the image to be translated is determined, and the complete outline object with the largest proportion is further identified. For example, the proportion of the complete outline object with the largest proportion is 45%.

[0104] Following the method described in the previous example, the positional distribution data corresponding to each complete outline text is determined. Based on the positional distribution data corresponding to each complete outline text, the proportion of each complete outline text in the image to be translated is determined, and the complete outline text with the largest proportion is further identified. For example, the proportion of the complete outline text with the largest proportion is 32%.

[0105] Specifically, the target type for translation is determined by identifying the largest percentage of fully outlined objects or text. Continuing with the previous example, if the largest percentage of fully outlined objects is 45% and the largest percentage of fully outlined text is 32%, then the fully outlined object representing 45% of the data is identified as the target type for translation. In other words, the user wants to translate an object.

[0106] For example, if the largest proportion of the data is a complete outline object, then the complete outline object with the largest proportion is determined as the target type to be translated, and all complete outline objects contained in the image to be translated are determined as the actual translation targets that the user wants to translate.

[0107] The solution in this application embodiment enables the intelligent selection of the actual translation target required by the user when the image to be translated contains multiple objects with complete outlines or multiple texts with complete outlines, thereby making the translation result more accurate.

[0108] Exemplary device

[0109] The above text combined Figures 2 to 7 The following describes in detail the embodiments of the translation method of this application, in conjunction with... Figure 8 This application provides a detailed description of embodiments of the translation apparatus. It should be understood that the descriptions of the translation method embodiments correspond to the descriptions of the translation apparatus embodiments; therefore, any parts not described in detail can be found in the preceding method embodiments.

[0110] Figure 8 The diagram shown is a structural schematic of a translation apparatus provided in an exemplary embodiment of this application. Figure 8 As shown, the translation apparatus 80 provided in this application embodiment includes:

[0111] The acquisition module 810 is used to acquire an image to be translated, the image to be translated including at least one translatable object, and the translatable object contains translatable text;

[0112] The determination module 820 is used to determine the actual translation target from the translatable object and the translatable text;

[0113] Translation module 830 is used to translate the actual translation target and obtain the translation result.

[0114] In one embodiment of this application, the determining module 820 is further configured to determine, based on the image information of the image to be translated, a contour-complete object and a contour-complete text from the translatable object and the translatable text; if the number of contour-complete objects and contour-complete texts is 1, then based on the position information of each contour-complete object and contour-complete text in the image to be translated, determine the actual translation target from the contour-complete object and contour-complete text.

[0115] In one embodiment of this application, the determining module 820 is further configured to: determine objects with complete outlines and text with complete outlines that have a proportion greater than a preset proportion threshold as actual translation targets; or determine objects with complete outlines and text with complete outlines that have a distance less than a preset distance threshold from the center position of the image to be translated as actual translation targets.

[0116] In one embodiment of this application, the determining module 820 is further configured to: calculate a translation score for the outline-complete object based on at least one of the proportion of the outline-complete object in the image to be translated, the distance between the outline-complete object and the center position of the image to be translated, and the distance between the outline-complete text and the outline-complete object; calculate a translation score for the outline-complete text based on at least one of the proportion of the outline-complete text in the image to be translated, the distance between the outline-complete text and the center position of the image to be translated, and the distance between the outline-complete object and the outline-complete text; and determine the actual translation target from the outline-complete object and the outline-complete text based on the translation scores of the outline-complete object and the outline-complete text.

[0117] In one embodiment of this application, the determining module 820 is further configured to: determine a contour-complete object and a contour-complete text from translatable objects and translatable text based on image information of the image to be translated; and determine the actual translation target from the contour-complete object and the contour-complete text based on user information.

[0118] In one embodiment of this application, the determining module 820 is further configured to: determine the object translation word corresponding to the outline complete object and the text translation word corresponding to the outline complete text; determine whether the object translation word and the text translation word are within the translation vocabulary information; and determine the outline complete object or outline complete text within the translation vocabulary information as the actual translation target.

[0119] In one embodiment of this application, the determining module 820 is further configured to send translation selection prompts to the user so that the user can select an actual translation target from the translatable objects and translatable text; and in response to the user's translation selection operation, determine the actual translation target from the translatable objects and translatable text.

[0120] Below, for reference Figure 9 This describes an electronic device according to embodiments of the present application. Figure 9 The diagram shown is a structural schematic of an electronic device provided in an exemplary embodiment of this application.

[0121] like Figure 9 As shown, the electronic device 90 includes one or more processors 901 and memory 902.

[0122] The processor 901 may be a central processing unit (CPU) or other form of processing unit with data processing and / or instruction execution capabilities, and may control other components in the electronic device 90 to perform desired functions.

[0123] The memory 902 may include one or more computer program products, which may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. The volatile memory may include, for example, random access memory (RAM) and / or cache memory. The non-volatile memory may include, for example, read-only memory (ROM), hard disk, flash memory, etc. One or more computer program instructions may be stored on the computer-readable storage medium, and the processor 901 may execute the program instructions to implement the translation methods of the various embodiments of this application described above and / or other desired functions. The computer-readable storage medium may also store various content such as the image to be translated, the actual translation target, and the translation result.

[0124] In one example, the electronic device 90 may also include an input device 903 and an output device 904, which are interconnected via a bus system and / or other forms of connection mechanism (not shown).

[0125] The input device 903 may include, for example, a keyboard, a mouse, etc.

[0126] The output device 904 can output various information to the outside, including the image to be translated, the actual translation target, and the translation result. The output device 904 may include, for example, a display, a speaker, a printer, and a communication network and its connected remote output devices, etc.

[0127] Of course, for the sake of simplicity, Figure 9 Only some of the components of the electronic device 90 relevant to this application are shown in this illustration; components such as buses, input / output interfaces, etc., are omitted. In addition, the electronic device 90 may include any other suitable components depending on the specific application.

[0128] In addition to the methods and apparatus described above, embodiments of this application may also be computer program products, which include computer program instructions that, when executed by a processor, cause the processor to perform the steps of the translation methods according to the various embodiments of this application described above.

[0129] The computer program product can be written in any combination of one or more programming languages ​​to perform the operations of the embodiments of this application. The programming languages ​​include object-oriented programming languages ​​such as Java and C++, as well as conventional procedural programming languages ​​such as C or similar languages. The program code can be executed entirely on the user's computing device, partially on the user's computing device, as a standalone software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server.

[0130] Furthermore, embodiments of this application may also be computer-readable storage media storing computer program instructions that, when executed by a processor, cause the processor to perform the steps in the translation methods according to the various embodiments of this application described above.

[0131] The computer-readable storage medium may be any combination of one or more readable media. A readable medium may be a readable signal medium or a readable storage medium. A readable storage medium may, for example, include, but is not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatuses, or devices, or any combination thereof. More specific examples of readable storage media (a non-exhaustive list) include: electrical connections having one or more wires, portable disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.

[0132] The basic principles of this application have been described above with reference to specific embodiments. However, it should be noted that the advantages, benefits, and effects mentioned in this application are merely examples and not limitations, and should not be considered as essential features of each embodiment of this application. Furthermore, the specific details disclosed above are for illustrative and facilitative purposes only, and are not limitations. These details do not limit the application to the necessity of employing the aforementioned specific details for implementation.

[0133] The block diagrams of devices, apparatuses, devices, and systems involved in this application are merely illustrative examples and are not intended to require or imply that they must be connected, arranged, or configured in the manner shown in the block diagrams. As those skilled in the art will recognize, these devices, apparatuses, devices, and systems can be connected, arranged, and configured in any manner. Words such as “comprising,” “including,” “having,” etc., are open-ended terms meaning “including but not limited to,” and are used interchangeably with them. The terms “or” and “and” as used herein refer to the terms “and / or,” and are used interchangeably with them unless the context clearly indicates otherwise. The term “such as” as used herein refers to the phrase “such as but not limited to,” and is used interchangeably with it.

[0134] It should also be noted that in the apparatus, equipment, and methods of this application, the components or steps can be disassembled and / or recombined. These disassemblies and / or recombinations should be considered as equivalent solutions of this application.

[0135] The above description of the disclosed aspects is provided to enable any person skilled in the art to make or use this application. Various modifications to these aspects will be readily apparent to those skilled in the art, and the general principles defined herein can be applied to other aspects without departing from the scope of this application. Therefore, this application is not intended to be limited to the aspects shown herein, but rather to be accorded the widest scope consistent with the principles and novel features disclosed herein.

[0136] The above description has been given for purposes of illustration and description. Furthermore, this description is not intended to limit the embodiments of this application to the forms disclosed herein. Although numerous exemplary aspects and embodiments have been discussed above, those skilled in the art will recognize certain variations, modifications, alterations, additions, and sub-combinations thereof.

Claims

1. A translation method, characterized in that, include: Obtain an image to be translated, the image to be translated including at least one translatable object, and the translatable object containing translatable text; Determine the actual translation target from the translatable object and the translatable text; The actual translation target is translated to obtain the translation result; Determining the actual translation target from the translatable object and the translatable text includes: Based on the image information of the image to be translated, a contour-complete object and a contour-complete text are determined from the translatable object and the translatable text; If the number of both the outlined complete object and the outlined complete text is 1, then the actual translation target is determined from the outlined complete object and the outlined complete text based on their respective position information in the image to be translated.

2. The translation method according to claim 1, characterized in that, The location information includes a percentage or the distance between the object and the center of the image to be translated. The step of determining the actual translation target from the complete outline object and the complete outline text within the image to be translated, based on their respective location information, includes: The objects with complete outlines and the text with complete outlines whose proportions are greater than a preset proportion threshold are identified as the actual translation targets. Alternatively, the actual translation target can be determined as the object with the complete outline and the text with the complete outline that are less than a preset distance threshold from the center position of the image to be translated.

3. The translation method according to claim 1, characterized in that, The location information includes proportion, distance to the center of the image to be translated, and distance between the outlined object and the outlined text. The step of determining the actual translation target from the outlined object and the outlined text based on their respective location information in the image to be translated includes: The translation score of the outline-complete object is calculated based on at least one of the proportion of the outline-complete object in the image to be translated, the distance between the outline-complete object and the center position of the image to be translated, and the distance between the outline-complete text. The translation score of the outline-complete text is calculated based on at least one of the following: the proportion of the outline-complete text in the image to be translated, the distance between the outline-complete text and the center position of the image to be translated, and the distance between the outline-complete text and the outline-complete object. Based on the translation scores of the outline-complete object and the outline-complete text, the actual translation target is determined from the outline-complete object and the outline-complete text.

4. The translation method according to claim 1, characterized in that, Determining the actual translation target from the translatable object and the translatable text includes: Based on the image information of the image to be translated, a contour-complete object and a contour-complete text are determined from the translatable object and the translatable text; Based on user information, the actual translation target is determined from the outlined object and the outlined text.

5. The translation method according to claim 4, characterized in that, User information includes translation vocabulary information. The step of determining the actual translation target based on the user information from the outlined complete object and the outlined complete text includes: Determine the object translation term corresponding to the complete outline object and the text translation term corresponding to the complete outline text; Determine whether the object translation word and the text translation word are within the translated vocabulary information; The outlined object or outlined text within the translated vocabulary information is identified as the actual translation target.

6. The translation method according to claim 1, characterized in that, Determining the actual translation target from the translatable object and the translatable text includes: Send a translation selection prompt to the user so that the user can select the actual translation target from the translatable object and the translatable text; In response to the user's translation selection action, the translatable object and the actual translation target in the translatable text are determined.

7. A translation device, characterized in that, include: An acquisition module is used to acquire an image to be translated, the image to be translated including at least one translatable object, and the translatable object contains translatable text; A determination module is used to determine the actual translation target from the translatable object and the translatable text; The translation module is used to translate the actual translation target and obtain the translation result; Determining the actual translation target from the translatable object and the translatable text includes: Based on the image information of the image to be translated, a contour-complete object and a contour-complete text are determined from the translatable object and the translatable text; If the number of both the outlined complete object and the outlined complete text is 1, then the actual translation target is determined from the outlined complete object and the outlined complete text based on their respective position information in the image to be translated.

8. A computer-readable storage medium, characterized in that, The storage medium stores a computer program for executing the translation method according to any one of claims 1 to 6.

9. An electronic device, characterized in that, include: processor; Memory used to store the processor's executable instructions; The processor is used to execute the translation method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Translation display method and device based on augmented reality, computing equipment and medium

    CN108681393A

  • Corpus generation method and device, translation model training method, translation model translation method, equipment and medium

    CN111881900A