Text translation method and device, electronic equipment and storage medium

By acquiring text line information and automatically segmenting image regions using OCR and convolutional neural networks, the problem of text translation flexibility for electronic devices when dealing with disabled users or users whose hands are occupied is solved, achieving efficient text translation results.

CN115481644BActive Publication Date: 2025-12-16VIVO MOBILE COMM CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202211204015.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-09-29
Publication Date
2025-12-16
Estimated Expiration
2042-09-29

AI Technical Summary

Technical Problem

In existing technologies, electronic devices cannot flexibly translate the text content required by users when faced with disabled persons or when the user's hands are occupied, resulting in poor text translation flexibility.

Method used

By acquiring text line information from the text region of the image to be translated, and using OCR algorithms and convolutional neural networks to segment the image, the region to be translated is automatically determined without the need for complex user gestures or dragging operations.

Benefits of technology

It improves the efficiency and flexibility of electronic devices in translating text, accurately identifies and translates text paragraphs in images, and enhances the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115481644B_ABST
    Figure CN115481644B_ABST
Patent Text Reader

Abstract

The application discloses a text translation method and device, electronic equipment and a storage medium, and belongs to the technical field of images. The method comprises the following steps: acquiring text line information in M text regions of a to-be-translated image, wherein each text region comprises one text line; and dividing the to-be-translated image based on the text line information, so as to obtain N to-be-translated regions; wherein the text line information corresponding to any text region comprises at least one of the following: position information of any text region, layout information of a text line in any text region, and text information in any text region.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the technical field of images, and particularly relates to a text translation method and device, electronic equipment and a storage medium. BACKGROUND

[0002] Currently, when a user encounters a strange word or a text paragraph that cannot be smoothly translated during daily foreign language learning or reading foreign language documents, the user can recognize the strange word or the text paragraph that cannot be smoothly translated through an electronic device, and then the electronic device can translate the strange word or the text paragraph that cannot be smoothly translated.

[0003] In the prior art, the electronic device can obtain a text picture to be translated through a camera, and then the user can operate the text picture (for example, crop the text picture, perform smearing on the text picture, or click a certain region in the text picture) so that the electronic device can translate the content required by the user in the text picture according to the operation of the user.

[0004] However, in the above method, when the user is a disabled person (for example, the user has a disabled hand) or is inconvenient to operate the text picture (for example, the user's hands are occupied), the electronic device cannot translate the content required by the user in the text picture according to the operation of the user, and thus the flexibility of text translation of the electronic device is poor. SUMMARY

[0005] The embodiments of the application aim to provide a text translation method and device, electronic equipment and a storage medium, which can solve the problem of poor flexibility of text translation of the electronic device.

[0006] In a first aspect, the embodiments of the application provide a text translation method, which comprises: obtaining text line information in M text regions of an image to be translated, each text region containing one line of text; dividing the image to be translated based on the text line information to obtain N regions to be translated; wherein the text line information corresponding to any text region comprises at least one of the following: position information of any text region, layout information of a text line in any text region, and text information in any text region.

[0007] In a second aspect, the embodiments of the application provide a text translation device, which comprises: an acquisition module and a processing module. The acquisition module is configured to obtain text line information in M text regions of an image to be translated, each text region containing one line of text. The processing module is configured to divide the image to be translated based on the text line information to obtain N regions to be translated; wherein the text line information corresponding to any text region comprises at least one of the following: position information of any text region, layout information of a text line in any text region, and text information in any text region.

[0008] In a third aspect, an electronic device is provided, which includes a processor and a memory. The memory stores programs or instructions executable on the processor. The programs or instructions, when executed by the processor, implement the steps of the method according to the first aspect.

[0009] In a fourth aspect, a readable storage medium is provided, which stores programs or instructions. The programs or instructions, when executed by a processor, implement the steps of the method according to the first aspect.

[0010] In a fifth aspect, a chip is provided, which includes a processor and a communication interface. The communication interface is coupled to the processor. The processor is configured to execute programs or instructions to implement the method according to the first aspect.

[0011] In a sixth aspect, a computer program product is provided, which is stored in a storage medium. The computer program product is executed by at least one processor to implement the method according to the first aspect.

[0012] In the embodiments of the present application, the electronic device can obtain text line information in the M text regions of the image to be translated, and then divide the image to be translated according to the text line information to obtain N regions to be translated. In this scheme, the electronic device can determine the N regions to be translated in the division of the image to be translated according to the text line information, without the need for the user to perform complex gesture operations or drag operations on the paragraph to be translated. Therefore, the electronic device can more efficiently translate the text paragraphs in the N regions to be translated in the image to be translated, thereby improving the efficiency and flexibility of the electronic device in translating text. BRIEF DESCRIPTION OF DRAWINGS

[0013] Figure 1 is a flowchart of a text translation method provided by the embodiments of the present application;

[0014] Figure 2 is a structural schematic diagram of a text translation device provided by the embodiments of the present application;

[0015] Figure 3 is one of hardware structural schematic diagrams of an electronic device provided by the embodiments of the present application;

[0016] Figure 4 is another hardware structural schematic diagram of an electronic device provided by the embodiments of the present application. DETAILED DESCRIPTION

[0017] With reference to the drawings and the embodiments of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described. Obviously, the described embodiments are some but not all of the embodiments of the present application. Based on the embodiments of the present application, all other embodiments obtained by those of ordinary skill in the art without creative effort should fall within the scope of the present application.

[0018] The terms "first", "second", and the like in the specification and claims of the present application are used to distinguish similar objects, and are not used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of the present application can be implemented in an order other than that illustrated or described herein, and the objects distinguished by "first", "second", etc. are generally a category and do not limit the number of objects, for example, the first object can be one or more. In addition, "and / or" in the specification and claims indicates at least one of the connected objects, and the character " / ", generally indicates that the front and rear associated objects are in a "or" relationship.

[0019] The text translation method provided by the embodiments of the present application will be described in detail below with reference to the drawings, specific embodiments and application scenarios.

[0020] At present, with the development of communication technology, the functions in electronic devices are also increasing, for example, when a user learns foreign languages or reads foreign language documents daily, if a strange word or a text paragraph that cannot be translated smoothly is encountered, the electronic device can recognize the strange word or the text paragraph that cannot be translated smoothly, and then the electronic device can translate the strange word or the text paragraph that cannot be translated smoothly. Generally, the electronic device can collect an image of the text to be translated, thereby recognizing the image and then translating the text content in the image. In related technologies, the user performs gesture operations such as circle selection, clicking, dragging, etc. on the image to be translated, so that the electronic device can determine the translation range, thereby translating the selected text by the user. For example, the electronic device can display a crop frame to the user, so that the user can select the text content to be translated in the image to be translated according to the crop frame; or the user can perform a smearing operation on the image to be translated, so that the electronic device can translate the text content corresponding to the smearing area according to the smearing area; or the user can directly click the text content in the image to be translated, so that the electronic device can translate the text content corresponding to the position information input by the user according to the position information. However, in the above method, when the user is a disabled person (for example, the user has a disabled hand) or is not convenient to operate the text picture (for example, the user's hands are occupied), the electronic device cannot translate the content required by the user in the text picture according to the operation of the user, so the flexibility of the text translation of the electronic device is poor.

[0021] In the embodiment of the present application, the electronic device can obtain text line information in the M text regions of the image to be translated, and then divide the image to be translated according to the text line information to obtain N regions to be translated. In the present scheme, the electronic device can determine the N regions to be translated in the division of the image to be translated according to the text line information, without the user performing complex gesture operations or drag operations on the paragraph to be translated, so that the electronic device can more efficiently translate the text paragraph in the N regions to be translated in the image to be translated, and thus the efficiency and flexibility of the electronic device in translating text are improved.

[0022] The text translation method provided in the embodiment of the present application can be an electronic device or a functional module in the electronic device. The technical solutions provided in the embodiments of the present application are described below with the electronic device as an example.

[0023] The text translation method provided in the embodiment of the present application, Figure 1 A flowchart of a text translation method provided in the embodiment of the present application is shown. As shown in Figure 1 The text translation method provided in the embodiment of the present application can include the following steps 201 and 202.

[0024] Step 201, the electronic device obtains text line information in M text regions of an image to be translated.

[0025] In the embodiment of the present application, each of the M text regions contains a line of text.

[0026] In the embodiment of the present application, the text line information corresponding to any of the M text regions includes at least one of the following: position information of any text region, layout information of a text line in any text region, and text information in any text region.

[0027] It can be understood that the text line contains one line of text.

[0028] In the embodiment of the present application, the electronic device can obtain the text line information in the M text regions by using an optical character recognition (OCR) algorithm, so that the electronic device can divide the image to be translated according to the text line information to obtain N regions to be translated.

[0029] Optionally, in the embodiment of the present application, the text information in any text region can include at least one of the following: text font information, text size information, and text color information.

[0030] Optionally, in the embodiments of the present application, the position information of any of the text regions can be the vertex coordinates of any of the text regions, or the vertex coordinates of the text lines in the text region.

[0031] Optionally, in the embodiments of the present application, the image to be translated can be an image collected by the electronic device through a camera, or an image selected by the user in a target application program (for example, a photo album application program).

[0032] Optionally, in the embodiments of the present application, the camera can be any of the following: a wide-angle camera, an ultra-wide-angle camera, a long-focus camera, or a macro camera, etc. The specific determination can be made according to the use case, and the embodiments of the present application do not make any limitation.

[0033] Optionally, the text in the text line can be Chinese or English, or a combination of Chinese and English.

[0034] It can be understood that, in the case of a combination of Chinese and English in the text behavior, the electronic device can translate the text line according to the user-selected translation language.

[0035] It should be noted that the image to be translated is a broad sense of image, that is, the image to be translated can be a picture to be translated or a video to be translated.

[0036] Step 202, the electronic device divides the image to be translated based on the text line information, and obtains N image regions to be translated.

[0037] Optionally, in the embodiments of the present application, after obtaining the text line information, the electronic device can divide the image to be translated according to the information of each two text lines, and then obtain N image regions to be translated.

[0038] Optionally, in the embodiments of the present application, the electronic device can obtain N image regions to be translated according to the text position information, the text layout information, and the text information of each text line in the image to be translated, and identify the N image regions to be translated through a target identifier and display them on the screen in the electronic device.

[0039] Optionally, in the embodiments of the present application, the target identifier can include at least one of the following: a color identifier, a border identifier (for example, the electronic device can identify the N image regions to be translated through a dashed line or a solid line), a special symbol identifier, and a number identifier (for example, the electronic device can identify the N image regions to be translated through different numbers).

[0040] Optionally, in the embodiments of the present application, the electronic device can save the text line information, so that in the case of the same image to be translated, the electronic device can directly call the saved text line information to divide the image to be translated, and obtain N image regions to be translated, thereby improving the efficiency of the electronic device in obtaining N image regions to be translated.

[0041] Optionally, the step 202 can be implemented by steps 202a-202c in the embodiments of the present application.

[0042] In step 202a, the electronic device performs convolution processing on the text lines in the M text regions to obtain visual feature information of the text lines in the M text regions.

[0043] In the embodiments of the present application, the electronic device can collect an original image of the M text lines, and then perform convolution processing on the original image of the M text lines to obtain visual feature information of the M text lines.

[0044] For example, the electronic device can perform convolution processing on the original image of the M text lines by using a target convolution neural network to obtain visual feature information of the M text lines.

[0045] Optionally, the target convolution neural network can be any one of a convolution neural network, a back propagation (BP) neural network, a radial basis function (RBF) neural network, a linear neural network, or a self-organizing neural network.

[0046] It should be noted that the visual feature information is information about the layout of the text in the image to be translated (for example, horizontal layout or vertical layout).

[0047] In step 202b, the electronic device performs word segmentation processing on the text lines in the M text regions to obtain semantic feature information of the text lines in the M text regions.

[0048] In the embodiments of the present application, the electronic device can perform word segmentation processing on the text content of the M text lines recognized by the OCR algorithm to obtain at least one word group in each text line of the M text lines, and insert a start mark at the beginning of each text line and an end mark at the end of each text line. Then, the electronic device encodes the at least one word group in each text line by using a pre-trained language model (for example, a BERT model) to obtain semantic feature information of each text line.

[0049] Optionally, the start mark and the end mark can be the same or different.

[0050] Specifically, in the case where the start mark and the end mark are the same, the start mark and the end mark corresponding to each text line are different, and in the case where the start mark and the end mark are different, the electronic device can mark the M text lines by using the same start mark and end mark.

[0051] Optionally, in the embodiments of the present application, the start mark and the end mark can be any one of the following: a numerical identifier, a letter identifier, or a characteristic symbol identifier.

[0052] Optionally, in the embodiments of the present application, after obtaining the visual feature information and the semantic feature information, the electronic device can convert the visual feature information into a visual feature array or a visual feature vector, and convert the semantic feature information into a semantic feature array or a semantic feature vector, so that the electronic device can perform text paragraph division on the text in the image to be translated through the visual feature vector and the semantic feature vector, and obtain N image regions to be translated.

[0053] Step 202c, the electronic device divides the image to be translated according to the text line information, the visual feature information, and the semantic feature information, and obtains N image regions to be translated.

[0054] In the embodiments of the present application, the electronic device can input the text line information, the visual feature information, and the semantic feature information into the full connection layer, so as to divide the image to be translated according to the output result of the full connection layer, and obtain N image regions to be translated.

[0055] In the embodiments of the present application, the electronic device can accurately identify the context relationship of the original text in the image to be translated according to the position information of the text, the visual feature information of the text, and the semantic information of the text, and then merge each line of text through the context relationship of the original text to obtain N image regions to be translated, so as to improve the accuracy of the electronic device in determining the image regions to be translated.

[0056] Optionally, in the embodiments of the present application, the step 202b can be implemented through the following steps 301 and 302.

[0057] Step 301, the electronic device determines a text sequence according to the text of the text line in the M text regions.

[0058] In the embodiments of the present application, the text sequence is a sequence obtained by performing word segmentation on the text of the text line in the M text regions.

[0059] In the embodiments of the present application, the electronic device can recognize the text content of the text line in the M text regions through the OCR algorithm, then perform word segmentation processing on the text content to obtain a word group corresponding to each line of text in the M text regions, and then fuse the word group according to the visual feature information and the original text arrangement order of each line of text to obtain the text sequence.

[0060] Optionally, in the embodiment of the present application, the electronic device can perform the above operation on the text of the text lines in each two text regions to obtain a text sequence corresponding to each two text regions, and then the electronic device can perform recognition on the text sequence corresponding to each two text regions to obtain the semantic feature information of the text lines in the M text regions.

[0061] In step 302, the electronic device performs encoding processing on the text in the text sequence to obtain the semantic feature information of the text lines in the M text regions.

[0062] In the embodiment of the present application, the electronic device can perform encoding processing on the text in the text sequence through the encoding module to obtain a semantic vector matrix of the text lines in the M text regions, and then the electronic device can input the semantic vector matrix into the convolutional neural network to obtain the semantic feature information of the text lines in the M text regions.

[0063] In the embodiment of the present application, the electronic device can obtain the semantic feature information of the text lines in the M text regions through the OCR algorithm and the convolutional neural network, and then the electronic device can merge each line of text through the context relationship of the original text to obtain N translation regions, thereby improving the accuracy of the electronic device in determining the translation region.

[0064] Optionally, in the embodiment of the present application, the above step 202c can be implemented through the following step 401.

[0065] In step 401, the electronic device performs transformation processing on the text line information, the visual feature information and the semantic feature information to obtain a target probability value.

[0066] In the embodiment of the present application, the above target probability value is used to represent the probability of whether the M text regions in the image to be translated are merged.

[0067] Exemplarily, taking two lines of text as an example, when judging whether the two lines of text should be merged, the electronic device can first extract the 4 vertex coordinate information of text line 1, the color information of the text in the text line, the font information of the text in the text line, and the size information of the text in the text line, and the 4 vertex coordinate information of text line 2, the color information of the text in the text line, the font information of the text in the text line, and the size information of the text in the text line, and then represent the 8 information as an OCR feature vector to obtain Sentence1 OCR (OCR feature vector of text line 1) and Sentence2 OCR (OCR feature vector of text line 2). Secondly, the original text image of the two lines of text is obtained, and the original text image of the two lines of text is respectively subjected to convolution processing to obtain Sentence1 picture and Sentence2 picture. The Sentence1 picture is used to represent the visual feature information of text line 1, and the Sentence2 picture is used to represent the visual feature information of text line 2. Then, the electronic device can perform word segmentation processing on text line 1 and text line 2 to obtain each word group in text line 1 and each word group in text line 2, and add a start mark and an end mark to each text line after dividing the word groups. Each text line after adding the mark is input into a BERT model to encode each word group divided in the text to obtain semantic feature information of each line of text. The above three types of feature information are input into a neural network after matrix fusion to obtain feature information after fusion of the above three types of feature information. Then, the probability value of whether the two lines of text should be merged is calculated through a softmax strategy, so that the electronic device can merge the two lines of text according to the probability value.

[0068] The embodiment of the present application provides a text translation method. An electronic device can obtain text line information in M text regions of a to-be-translated image, and then divide the to-be-translated image according to the text line information to obtain N to-be-translated regions. In this solution, the electronic device can determine the N to-be-translated regions in the to-be-translated image division according to the text line information, without requiring the user to perform a complex gesture operation or a drag operation on a paragraph to be translated, so that the electronic device can more efficiently translate the text paragraph in the N to-be-translated regions of the to-be-translated image, thereby improving the efficiency and flexibility of the electronic device in translating text.

[0069] Optionally, after obtaining M, the electronic device can detect the number of paragraphs in the to-be-translated image. In the case that the number of paragraphs is 1, the electronic device can directly translate the paragraph through a translation model and display the translation result. In the case that the number of paragraphs is greater than 1, the electronic device can determine a target translation region through the gaze point information of the user's eyes on the screen.

[0070] Optionally, after step 202, the text translation method provided in the embodiment of the present application further includes steps 501 and 502.

[0071] In step 501, the electronic device determines position information of the gaze point of the user on the screen based on the collected gaze information of the user gazing at the image to be translated.

[0072] Optionally, in the embodiment of the present application, the gaze information includes at least one of the following: face information of the user, eyeball information of the user, eyeball movement information of the user within a predetermined time period, and distance information between the user and the screen.

[0073] Optionally, before step 501, the electronic device can display the image to be translated, so that the electronic device can determine the position information of the gaze point of the user on the screen based on the collected gaze information of the user gazing at the image to be translated.

[0074] In the embodiment of the present application, the electronic device can capture the face of the user, thereby collecting the gaze information of the user gazing at the image to be translated.

[0075] Specifically, for the face information of the user and the eyeball information of the user, after capturing the face of the user, the electronic device can identify the key point information of the face and the position information of the face relative to the screen through a face key point detection network, and the electronic device can extract the eyeball information of the user through the face key point detection network.

[0076] Specifically, the electronic device can determine the distance information between the user and the screen through infrared light.

[0077] For example, the electronic device can emit infrared light to irradiate the eyes of the user, and a reflection image is formed on the cornea and pupil of the eyes of the user, and a front camera can capture a sequence of reflection images of the eyes at a high speed. At this time, the distance sequence L1 and L2 between the left and right eyes of the user and the camera of the mobile phone are detected through infrared sensing.

[0078] Optionally, in the embodiment of the present application, if the electronic device does not identify the eyeball of the user, the electronic device can sample a random vector to represent the position of the eyeball of the user in the screen.

[0079] Specifically, the electronic device can capture a video of the face of the user, and then input the video into a convolutional neural network to extract the feature information of the eyes of the user in each frame of the video at a target time, that is, the sliding time sequence relationship of the corresponding screen area of the continuous frame of the eyes can reflect the motion trajectory information of the eyes of the user.

[0080] Optionally, in the embodiments of the present application, the step 501 can be implemented by the following steps 501a and 501b.

[0081] In the step 501a, the electronic device extracts first gaze feature information corresponding to the gaze information.

[0082] In the embodiments of the present application, the electronic device can extract the first gaze feature information corresponding to the gaze information through a convolutional neural network.

[0083] In the step 501b, the electronic device inputs the first gaze feature information into an image prediction model to determine the position information of the gaze point of the user on the screen.

[0084] In the embodiments of the present application, after obtaining the first gaze feature information, the electronic device can determine the position information of the gaze point of the user on the screen through a deep image prediction model.

[0085] It can be understood that the image prediction model integrates the overall image information, the local eye information, the sequence relationship information between multiple frames of images, and the distance position information, and can accurately predict the coordinate position of the eye focus on the screen according to a continuous multiple frames of information.

[0086] In the step 502, the electronic device determines a target translation region from the N translation regions based on the position information, and translates a text paragraph in the target translation region.

[0087] In the embodiments of the present application, the electronic device can determine the target translation region from the N translation regions based on the gaze point information of the eye on the screen, and translate the text paragraph in the target translation region.

[0088] Optionally, in the embodiments of the present application, the electronic device can translate the text paragraph in the target translation region through a machine translation model.

[0089] Optionally, in the embodiments of the present application, after obtaining the translation result of the text paragraph, the electronic device can post the text paragraph in the original text size, color, and coordinate position to the translation image.

[0090] Optionally, in the embodiments of the present application, after obtaining the translation result of the text paragraph, the electronic device can save the translation result, so as to facilitate subsequent translation of the same text paragraph.

[0091] In the embodiments of the present application, the electronic device can determine the target translation region from the N translation regions based on the position information, without the need for the user to manually select the translation region, thereby improving the flexibility of the electronic device in determining the target translation region.

[0092] Optionally, in the embodiments of the present application, the "determining, by the electronic device, the target to-be-translated region from the N to-be-translated regions based on the position information" in the step 502 can be implemented through the following step 502a or step 502b.

[0093] The step 502a: in the case that the position information of the gaze point matches the position information of the first to-be-translated region, the electronic device determines the first to-be-translated region as the target to-be-translated region.

[0094] In the embodiments of the present application, in the case that the position of the gaze point is the same as the position of the first to-be-translated region, the electronic device determines the first to-be-translated region as the target to-be-translated region.

[0095] The step 502b: in the case that the position information of the gaze point does not match the position information of the first to-be-translated region, the electronic device determines a second to-be-translated region as the target to-be-translated region according to the position information of the gaze point.

[0096] In the embodiments of the present application, the first to-be-translated region is any one of the N to-be-translated regions, and the second to-be-translated region is a to-be-translated region in a predetermined range centered on the gaze point in the N to-be-translated images.

[0097] In the embodiments of the present application, the electronic device can determine the target to-be-translated region according to the position information of the gaze point, thereby improving the flexibility and efficiency of the electronic device in translating text.

[0098] Optionally, in the embodiments of the present application, the step 502b can be implemented through the following steps 601 and 602.

[0099] The step 601: in the case that there are multiple to-be-translated regions in the predetermined range, the electronic device determines the distance between each to-be-translated region in the multiple to-be-translated regions and the gaze point.

[0100] For example, in the case that the position information of the electronic device at the gaze point does not match the position information of the first to-be-translated region, the electronic device can search for to-be-translated regions with the gaze point as the center and 1 / 4 of the width of the electronic device as the radius, and in the case that multiple to-be-translated regions are found, the electronic device can calculate the distance between each to-be-translated region in the multiple to-be-translated regions and the gaze point.

[0101] The step 602: the electronic device determines the target to-be-translated region from the multiple to-be-translated regions based on the distance.

[0102] In the embodiments of the present application, the electronic device determines the to-be-translated region with the shortest distance to the gaze point from the multiple to-be-translated regions as the target to-be-translated region.

[0103] In the embodiment of the present application, in the case that the electronic device finds a plurality of to-be-translated regions, the electronic device can determine the to-be-translated region with the shortest distance to the gaze point as the target to-be-translated region, and then the electronic device can translate the target to-be-translated region.

[0104] In the embodiment of the present application, the electronic device can determine the target to-be-translated region according to the distance between the plurality of to-be-translated regions and the gaze point, thereby improving the flexibility of determining the target to-be-translated region.

[0105] It should be noted that the text translation method provided in the embodiment of the present application can be executed by a text translation device, or an electronic device, or a functional module or entity in the electronic device. In the embodiment of the present application, the text translation device executing the text translation method is taken as an example to illustrate the text translation device provided in the embodiment of the present application.

[0106] Figure 2 A possible structural schematic diagram of the text translation device involved in the embodiment of the present application is shown. As shown in the figure, Figure 2 The text translation device 70 can include an acquisition module 71 and a processing module 72.

[0107] The acquisition module 71 is configured to acquire text line information in M text regions of a to-be-translated image, each text region containing one text line. The processing module 72 is configured to divide the to-be-translated image based on the text line information to obtain N to-be-translated regions. The text line information corresponding to any text region includes at least one of the following: position information of any text region, layout information of a text line in any text region, and text information in any text region.

[0108] In a possible implementation, the processing module 72 is specifically configured to perform convolution processing on the text lines in the M text regions to obtain visual feature information of the text lines in the M text regions, perform word segmentation processing on the text lines in the M text regions to obtain semantic feature information of the text lines in the M text regions, and divide the to-be-translated image according to the text line information, the visual feature information and the semantic feature information to obtain the N to-be-translated regions.

[0109] In a possible implementation, the processing module 72 is specifically configured to determine a text sequence according to the text of the text lines in the M text regions, the text sequence being a sequence obtained by performing word segmentation on the text of the text lines in the M text regions, and perform encoding processing on the text in the text sequence to obtain the semantic feature information of the text lines in the M text regions.

[0110] In a possible implementation, the processing module 72 is specifically configured to perform conversion processing on the text line information, the visual feature information, and the semantic feature information to obtain a target probability value, where the target probability value is used to represent a probability of whether the M text regions in the image to be translated are merged.

[0111] In a possible implementation, the text translation apparatus provided by the embodiment of the present application further includes a determining module and a translation module. The determining module is configured to, after the processing module 72 divides the image to be translated based on the text line information to obtain the N text regions to be translated, determine position information of a gaze point of a user on a screen based on gaze information collected by the user gazing at the image to be translated. The translation module is configured to determine a target text region to be translated from the N text regions to be translated based on the position information determined by the determining module, and translate a text paragraph in the target text region to be translated.

[0112] The text translation apparatus provided by the embodiment of the present application can determine the N text regions to be translated in the image to be translated based on the text line information, without the user performing a complex gesture operation or a drag operation on a text paragraph to be translated, so that the text translation apparatus can more efficiently translate the text paragraph in the N text regions to be translated in the image to be translated, thereby improving the efficiency and flexibility of the electronic device in translating text.

[0113] The text translation apparatus in the embodiment of the present application can be an apparatus, or a component, an integrated circuit, or a chip in an electronic device. The apparatus can be a mobile electronic device or a non-mobile electronic device. Illustratively, the mobile electronic device can be a mobile phone, a tablet computer, a notebook computer, a palm computer, a vehicle-mounted electronic device, a Mobile Internet Device (MID), an augmented reality (AR) / virtual reality (VR) device, a robot, a wearable device, an ultra-mobile personal computer (UMPC), a netbook, or a personal digital assistant (PDA), etc. The apparatus can also be a server, a Network Attached Storage (NAS), a personal computer (PC), a television (TV), a teller machine, or a self-service machine, etc. The embodiment of the present application is not limited in this regard.

[0114] The text translation apparatus in the embodiments of the present application can be an apparatus with an operating system. The operating system can be an Android operating system, an iOS operating system, or other possible operating systems, which are not limited in the embodiments of the present application.

[0115] The text translation apparatus provided in the embodiments of the present application can implement Figure 1 The method embodiments implement various processes, which will not be repeated here to avoid repetition.

[0116] Optionally, as Figure 3 indicated, the embodiments of the present application further provide an electronic device 90, which includes a processor 91 and a memory 92, and the memory 92 stores programs or instructions executable on the processor 91, which implement various steps of the above text translation method embodiments and achieve the same technical effects when executed by the processor 91, and will not be repeated here to avoid repetition.

[0117] It should be noted that the electronic device in the embodiments of the present application includes the mobile electronic device and the non-mobile electronic device described above.

[0118] Figure 4 To implement the hardware structure of an electronic device in the embodiments of the present application.

[0119] The electronic device 100 includes but is not limited to the radio frequency unit 101, the network module 102, the audio output unit 103, the input unit 104, the sensor 105, the display unit 106, the user input unit 107, the interface unit 108, the memory 109, and the processor 110, and the like.

[0120] Those skilled in the art can understand that the electronic device 100 can further include a power supply (such as a battery) for powering various components, and the power supply can be logically connected to the processor 110 through a power management system, so as to realize functions such as management of charging, discharging, and power consumption management through the power management system. Figure 4 The electronic device structure shown in the embodiments of the present application does not constitute a limitation on the electronic device, and the electronic device can include more or fewer components than those shown, or combine certain components, or different component arrangements, which will not be repeated here.

[0121] The processor 110 is configured to acquire text line information in M text regions of a to-be-translated image, each text region containing one text line, and divide the to-be-translated image based on the text line information to obtain N to-be-translated regions; and the text line information of any text region of the M text regions includes at least one of the following: position information of any text region, layout information of a text line in any text region, and text information in any text region.

[0122] The electronic device provided in the embodiments of the present application can determine the N to-be-translated regions in the to-be-translated image according to the text line information, without requiring the user to perform a complex gesture operation or a drag operation on the paragraph to be translated, so that the electronic device can more efficiently translate the text paragraph in the N to-be-translated regions in the to-be-translated image, thereby improving the efficiency and flexibility of the electronic device in translating text.

[0123] Optionally, in the embodiments of the present application, the processor 110 is specifically configured to perform convolution processing on the text lines in the M text regions to obtain visual feature information of the text lines in the M text regions, perform word segmentation processing on the text lines in the M text regions to obtain semantic feature information of the text lines in the M text regions, and divide the to-be-translated image according to the text line information, the visual feature information and the semantic feature information to obtain the N to-be-translated regions.

[0124] Optionally, in the embodiments of the present application, the processor 110 is specifically configured to determine a text sequence according to the text of the text lines in the M text regions, the text sequence being a sequence obtained by performing word segmentation on the text of the text lines in the M text regions, and perform encoding processing on the text in the text sequence to obtain the semantic feature information of the text lines in the M text regions.

[0125] Optionally, in the embodiments of the present application, the processor 110 is specifically configured to perform transformation processing on the text line information, the visual feature information and the semantic feature information to obtain a target probability value, the target probability value being used to represent a probability of whether the M text regions in the to-be-translated image are to be merged.

[0126] Optionally, in the embodiments of the present application, the processor 110 is further configured to, after dividing the to-be-translated image to obtain the N to-be-translated regions based on the text line information, determine position information of a gaze point of the user on the screen based on the gaze information of the user gazing at the to-be-translated image, and determine a target to-be-translated region from the N to-be-translated regions based on the position information and translate a text paragraph in the target to-be-translated region.

[0127] The electronic device provided in the embodiments of the present application can implement each process achieved by the method embodiments, and achieve the same technical effects. To avoid repetition, details are not described herein.

[0128] The beneficial effects of various implementation manners in the embodiments can refer to the beneficial effects of the corresponding implementation manners in the method embodiments, and details are not described herein to avoid repetition.

[0129] It should be understood that in the embodiments of the present application, the input unit 104 can include a graphics processing unit (GPU) 1041 and a microphone 1042. The graphics processing unit 1041 processes image data of a still picture or a video obtained by an image capture device (such as a camera) in a video capture mode or an image capture mode. The display unit 106 can include a display panel 1061, which can be configured in the form of a liquid crystal display, an organic light-emitting diode, etc. The user input unit 107 includes at least one of a touch panel 1071 and other input devices 1072. The touch panel 1071 is also referred to as a touch screen. The touch panel 1071 can include two parts of a touch detection device and a touch controller. The other input devices 1072 can include, but are not limited to, a physical keyboard, function keys (such as volume control keys, on-off keys, etc.), a trackball, a mouse, a joystick, and the like, which will not be described here.

[0130] The memory 109 can be used to store software programs and various data. The memory 109 can mainly include a first storage area storing programs or instructions and a second storage area storing data, wherein the first storage area can store an operating system, application programs or instructions required by at least one function (such as a sound playing function, an image playing function, etc.), and the like. In addition, the memory 109 can include a volatile memory or a non-volatile memory, or the memory 109 can include both volatile and non-volatile memories. The non-volatile memory can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically EPROM (EEPROM), or a flash memory. The volatile memory can be a random access memory (RAM), a static random access memory (SRAM), a dynamic random access memory (DRAM), a synchronous dynamic random access memory (SDRAM), a double data rate synchronous dynamic random access memory (DDR SDRAM), an enhanced synchronous dynamic random access memory (ESDRAM), a synch link DRAM (SLDRAM), and a direct memory bus random access memory (Direct Rambus RAM, DRRAM). The memory 109 in the embodiments of the present application includes but is not limited to these and any other suitable types of memory.

[0131] The processor 110 can include one or more processing units; optionally, the processor 110 integrates an application processor and a modem processor, wherein the application processor mainly processes operations related to operating systems, user interfaces, and application programs, and the modem processor mainly processes wireless communication signals, such as a baseband processor. It can be understood that the above-mentioned modem processor can also not be integrated into the processor 110.

[0132] The embodiment of the present application further provides a readable storage medium, and the readable storage medium stores a program or instructions, the program or instructions are executed by a processor to realize various processes of the above-mentioned method embodiments, and the same technical effects can be achieved, and details are not repeated here to avoid repetition.

[0133] The processor is the processor in the electronic device in the above-mentioned embodiments. The readable storage medium includes a computer readable storage medium, such as a computer readable only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, and the like.

[0134] The embodiment of the present application further provides a chip, and the chip includes a processor and a communication interface, the communication interface is coupled with the processor, and the processor is used to run a program or instructions to realize various processes of the above-mentioned method embodiments, and the same technical effects can be achieved, and details are not repeated here to avoid repetition.

[0135] It should be understood that the chip mentioned in the embodiment of the present application can also be referred to as a system level chip, a system chip, a chip system, or a system on chip, and the like.

[0136] The embodiment of the present application provides a computer program product, and the program product is stored in a storage medium, and the program product is executed by at least one processor to realize various processes of the above-mentioned text translation method embodiments, and the same technical effects can be achieved, and details are not repeated here to avoid repetition.

[0137] It should be noted that, in the present document, the terms "comprises", "comprising", or any other variations thereof, are intended to cover a non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements does not include only those elements but can include other elements not expressly listed or inherent to such process, method, article, or apparatus. An element proceeded by "comprises... a" does not, without more constraints, exclude the presence of additional identical elements in the process, method, article, or apparatus that comprises the element. Furthermore, it is to be understood that the method and apparatus of the present application can be carried out by more than one process, method, article, or apparatus either simultaneously, concurrently, or with intervening action that are carried out at the same time, either in a simultaneous fashion or in a fashion that is interleaved in time. For example, the described methods can be performed in a different order from that described, and / or various steps can be combined or omitted, and / or additional steps can be added, without departing from the scope of the present application. Also, features described with respect to certain examples can be combined in other examples.

[0138] From the above description of the embodiments, it is apparent that the above-mentioned method can be realized by means of software and necessary universal hardware platform, of course, it can also be realized by hardware, but in many cases, the former is a better embodiment. Based on such understanding, the technical solution of the present application can be embodied in the form of computer software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes a plurality of instructions for making a terminal (which can be a mobile phone, computer, server, or network equipment, etc.) execute the method described in various embodiments of the present application.

[0139] The embodiments of the present application are described above in conjunction with the drawings, but the present application is not limited to the above-described specific embodiments, and the above-described specific embodiments are merely illustrative, rather than limiting, and those skilled in the art can make many forms under the inspiration of the present application without departing from the scope of the present application and the scope protected by the claims.

Claims

1. A method of translating text, characterized by, The method comprises: obtaining text line information in M text regions of a to-be-translated image, each text region containing one line of text; based on the text line information, dividing the to-be-translated image to obtain N to-be-translated regions; wherein the text line information corresponding to any text region comprises the layout information of the text line in any text region; based on the text line information, dividing the to-be-translated image to obtain N to-be-translated regions, comprising: performing convolution processing on the text lines in the M text regions to obtain visual feature information of the text lines in the M text regions, the visual feature information being layout information of the text in the to-be-translated image; performing word segmentation processing on the text lines in the M text regions to obtain semantic feature information of the text lines in the M text regions; according to the text line information, the visual feature information and the semantic feature information, dividing the to-be-translated image to obtain the N to-be-translated regions; after the to-be-translated image is divided based on the text line information to obtain N to-be-translated regions, the method further comprises: extracting first gaze feature information corresponding to gaze information of a user gazing at the to-be-translated image; inputting the first gaze feature information into an image prediction model to determine position information of a gaze point of the user on a screen, the image prediction model fusing image overall information, eye local information, sequence relationship information and distance position information between multiple frames of images, the image prediction model predicting the position information according to continuous multiple frames of information; based on the position information, determining a target to-be-translated region from the N to-be-translated regions, and translating a text paragraph in the target to-be-translated region.

2. The method of claim 1, wherein, The word segmentation processing on the text lines in the M text regions to obtain semantic feature information of the text lines in the M text regions comprises: determining a text sequence according to the text of the text lines in the M text regions, the text sequence being a sequence obtained by performing word segmentation on the text of the text lines in the M text regions; performing encoding processing on the text in the text sequence to obtain semantic feature information of the text lines in the M text regions.

3. The method of claim 1, wherein, The dividing the to-be-translated image according to the text line information, the visual feature information and the semantic feature information to obtain N to-be-translated regions comprises: performing transformation processing on the text line information, the visual feature information and the semantic feature information to obtain a target probability value, the target probability value being used to represent a probability of whether the M text regions in the to-be-translated image are merged.

4. A text translation apparatus characterized by comprising: The text translation device comprises an acquisition module, a processing module, a determination module and a translation module; the acquisition module is used to obtain text line information in M text regions of a to-be-translated image, each text region containing one line of text; the processing module is used to divide the to-be-translated image based on the text line information to obtain N to-be-translated regions; wherein the text line information corresponding to any text region comprises the layout information of the text line in any text region; The processing module is specifically configured to perform convolution processing on the text lines in the M text regions to obtain visual feature information of the text lines in the M text regions, the visual feature information being layout information of the text in the image to be translated; perform word segmentation processing on the text lines in the M text regions to obtain semantic feature information of the text lines in the M text regions; and divide the image to be translated according to the text line information, the visual feature information, and the semantic feature information to obtain the N translation regions. The determining module is configured to, after the processing module divides the image to be translated to obtain the N translation regions based on the text line information, extract first gaze feature information corresponding to gaze information of a user gazing at the image to be translated; and input the first gaze feature information into an image prediction model to determine position information of a gaze point of the user on a screen, the image prediction model fusing image overall information, eye local information, sequence relationship information and distance position information between multiple images, the image prediction model predicting the position information according to continuous multiple frames of information. The translation module is configured to determine a target translation region from the N translation regions based on the position information, and translate a text paragraph in the target translation region.

5. The apparatus of claim 4, wherein, The processing module is specifically configured to determine a text sequence according to text of the text lines in the M text regions, the text sequence being a sequence obtained by performing word segmentation on the text of the text lines in the M text regions; and perform encoding processing on the text in the text sequence to obtain the semantic feature information of the text lines in the M text regions.

6. The apparatus of claim 4, wherein, The processing module is specifically configured to perform conversion processing on the text line information, the visual feature information, and the semantic feature information to obtain a target probability value, the target probability value being used to represent a probability of whether the M text regions in the image to be translated are to be merged.

7. An electronic device, comprising: The processor, the memory, and a program or instructions stored on the memory and executable on the processor are included, and the program or instructions are executed by the processor to implement steps of the text translation method in any one of claims 1 to 3.

8. A readable storage medium, characterized by, A program or instructions are stored on the readable storage medium, and the program or instructions are executed by the processor to implement steps of the text translation method in any one of claims 1 to 3.

Citation Information

Patent Citations

  • Video processing method and device, electronic equipment and storage medium

    CN110276349A

  • Translation method and device, electronic equipment and storage medium

    CN113255377A

  • Text image textbox sorting method and text image textbox sorting device

    CN114332889A

  • Picture recognition translation method and device, terminal and medium

    CN114402354A