Translation method and device
By obtaining the target page and text image of the dictionary pen, the position of the target text in the page is determined, thereby extracting the translation results from the translated text, solving the problem of inaccurate translation results of the dictionary pen and achieving a translation effect that is closer to the original meaning.
Patent Information
- Application Number
- CN202510502053.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-21
- Publication Date
- 2025-08-12
AI Technical Summary
The translation results of existing dictionary pens are relatively accurate because they are separated from the context of the text to be translated, resulting in the translation results that do not conform to the original meaning.
By obtaining the image of the target page where the target text is located and the image of the target text, matching, determining the position of the target text in the page, thereby extracting the translation results from the translated text and translating it with the original content of the page.
It improves the accuracy of the translation results, makes the translation results closer to the original text meaning, and ensures the accuracy of the translation results.
Smart Images

Figure CN120471070A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of text translation, and in particular to a translation method and device. Background Art
[0002] As the demand for language learning grows, electronic aids such as dictionary pens and electronic dictionaries have gradually become common learning devices for users. Through these learning devices, users can directly query the translation results of the text to be translated.
[0003] Taking dictionary pens as an example, they typically use a camera to scan an image of the text to be translated, then perform text recognition on the image to obtain the text to be translated. This text is then translated to produce a translation result. However, this method only translates the text to be translated, resulting in low translation accuracy. Summary of the Invention
[0004] The present invention provides a translation method and device to solve the defects in the prior art.
[0005] The present invention provides a translation method, comprising the following steps: Acquire a first image and a second image; the first image is an image of a target page where a target text is located, and the second image is an image of the target text; matching the first image and the second image to determine a position of the target text in the target page; Based on the position of the target text in the target page, a translation result of the target text is determined from the translated text of the target page.
[0006] According to a translation method provided by the present invention, determining a translation result of the target text from the translated text of the target page based on the position of the target text in the target page includes: Determining the regional text corresponding to the position from the translated text; Translating the target text to determine multiple candidate translation results of the target text; Each candidate translation result is matched with the regional text, and a translation result of the target text is determined from each candidate translation result.
[0007] According to a translation method provided by the present invention, matching each candidate translation result with the regional text and determining the translation result of the target text from each candidate translation result includes: Performing character matching on each candidate translation result and the regional text to determine the edit distance between each candidate translation result and the regional text; Based on the edit distance between each candidate translation result and the region text, a translation result of the target text is determined from each candidate translation result.
[0008] According to a translation method provided by the present invention, matching each candidate translation result with the regional text and determining the translation result of the target text from each candidate translation result includes: Performing semantic matching between each candidate translation result and the regional text to determine the semantic similarity between each candidate translation result and the regional text; Based on the semantic similarity between each candidate translation result and the regional text, a translation result of the target text is determined from each candidate translation result.
[0009] According to a translation method provided by the present invention, matching the first image and the second image to determine the position of the target text in the target page includes: matching the first image with the second image to determine an image region in the first image that matches the second image; The position of the target text in the target page is determined based on the position of the image area in the first image.
[0010] According to a translation method provided by the present invention, matching the first image with the second image and determining an image region in the first image that matches the second image includes: performing block processing on the first image to obtain a plurality of sub-images; determining an image similarity between each sub-image and the second image; The sub-images having image similarity greater than a threshold are respectively matched with the second image at the pixel level to determine an image region in the first image that matches the second image.
[0011] The present invention also provides a translation device, comprising the following modules: an acquisition unit, configured to acquire a first image and a second image; the first image is an image of a target page where a target text is located, and the second image is an image of the target text; a determination unit, configured to match the first image and the second image to determine a position of the target text in the target page; The translation unit is configured to determine a translation result of the target text from the translated text of the target page based on the position of the target text in the target page.
[0012] The present invention also provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements any of the above-described translation methods when executing the program.
[0013] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which implements any of the above-mentioned translation methods when executed by a processor.
[0014] The present invention also provides a computer program product, comprising a computer program, wherein when the computer program is executed by a processor, the computer program implements any of the above-mentioned translation methods.
[0015] The translation method and device provided by the present invention match the first image and the second image to determine the position of the target text in the target page, so as to extract the translation result of the target text from the translated text in combination with the position. Since the translated text is translated in combination with the original content of the target page, and the translation result is derived from the original translated text, the translation result of the target text can more accurately understand and convey the true meaning of the original text, ensure that the translation result is closer to the original meaning, and improve the accuracy of the translation result. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] In order to more clearly illustrate the technical solutions in the present invention or the prior art, a brief introduction is given below to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0017] Figure 1 This is one of the flow charts of the translation method provided by the present invention.
[0018] Figure 2 This is the second flowchart of the translation method provided by the present invention.
[0019] Figure 3 This is the third flow chart of the translation method provided by the present invention.
[0020] Figure 4 This is the fourth flow chart of the translation method provided by the present invention.
[0021] Figure 5 This is the fifth flow chart of the translation method provided by the present invention.
[0022] Figure 6 This is the sixth flowchart of the translation method provided by the present invention.
[0023] Figure 7It is a structural diagram of the translation device provided by the present invention.
[0024] Figure 8 It is a structural schematic diagram of the electronic device provided by the present invention. DETAILED DESCRIPTION
[0025] To make the objectives, technical solutions, and advantages of the present invention more clear, the technical solutions of the present invention will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the embodiments described are only some of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the present invention.
[0026] Currently, most dictionary pens use a camera to scan an image of the text to be translated, then perform text recognition on the image to obtain the text to be translated. This is then translated using a translation engine to produce a translation result. However, these translation results only translate the text to be translated, resulting in low accuracy.
[0027] After analysis, it was found that the reason why the translation results of traditional dictionary pens are less accurate is that they are translated out of the context of the text to be translated, which makes the translation results deviate from the original meaning and reduces the accuracy of the translation results.
[0028] For example, the text to be translated is "mouse", and the context of the text to be translated is "She set the mouse on the table and moved the cursor to the icon". According to the context of the text to be translated, we can know that "mouse" refers to the mouse, but if translated out of context, "mouse" can mean either mouse or rat. If "mouse" is used as the translation result, it does not conform to the original meaning.
[0029] To this end, the present invention provides a translation method designed to extract a translation result of a text to be translated from a context-based translation text (the translation text being the translation text of the page to which the text to be translated belongs) to improve the accuracy of the translation result. The method can be performed by a learning device (hereinafter referred to as "device"), such as a dictionary pen, an electronic dictionary, etc. To facilitate understanding of the technical solution of the present invention, the following embodiments are described using a dictionary pen as an example.
[0030] in, Figure 1 This is one of the flow charts of the translation method provided by the present invention, such as Figure 1 As shown, the method includes step 110 , step 120 and step 130 .
[0031] Step 110 , obtaining a first image and a second image; the first image is an image corresponding to the target page where the target text is located, and the second image is an image of the target text.
[0032] Specifically, the target text is the text to be translated. The target text can be a single word, a phrase, a sentence, or a text containing multiple sentences. This is not specifically limited in the embodiments of the present invention. The target text can be obtained by performing text recognition on the second image. Here, text recognition can be achieved through optical character recognition (OCR) technology, template matching, sequence modeling technology, etc. Considering the cost and efficiency of implementation, in the embodiments of the present invention, OCR technology is preferably used to perform text recognition on the second image.
[0033] The target page is the page where the target text is located, that is, the target page includes the target text. For example, if the target text is "Hello", the target page may be a page including "Hello".
[0034] The first image refers to an image of the target page where the target text is located. The first image can be an image of the entire page where the target text is located. For example, if the target text is a short phrase on line 5 of page 1, the first image can be an image of the entire page of page 1. The first image can also be an image of a portion of the page where the target text is located. For example, if the target text is a short phrase on line 5 of page 1, the first image can be an image of the text area corresponding to lines 2 to 6 of page 1. The first image can be acquired by a first image acquisition component, which can be a camera, an image sensor, or the like.
[0035] The second image refers to an image of the target text, which is used to represent the visual information of the target text. For example, if the target text is "Hello", the second image is an image containing "Hello". The first image and the second image can be acquired by the same image acquisition component or by different image acquisition components. The image acquisition component here can be a camera, image sensor, etc.
[0036] The first image acquisition element and the second image acquisition element can be the same image acquisition element or different image acquisition elements. When the two are the same image acquisition element, the image acquisition element can be deployed at the tip of the dictionary pen; when the two are different image acquisition elements, the first acquisition element can be deployed at the end away from the tip of the dictionary pen, and the second image acquisition element can be deployed at the tip of the dictionary pen; when the two are different image acquisition elements, the first acquisition element and the second acquisition element can both be deployed near the tip of the dictionary pen. In addition, when the first image acquisition element and the second image acquisition element are different acquisition elements, the activation order of the first acquisition element and the second acquisition element may be different. For example, the second acquisition element can be activated when the tip of the dictionary pen touches the page after the first acquisition element is activated.
[0037] Taking the camera deployed at the tip of a dictionary pen as the first image acquisition element, the triggering condition for the first image acquisition can be any of the following: ① detecting that the vertical distance between the camera and the target page containing the target text is greater than 0 and less than or equal to a first threshold; ② detecting that the pen body of the dictionary pen changes from a horizontal orientation to an inclined state (the inclined state here means that the angle between the pen body and the horizontal direction is greater than 0°); ③ detecting a capture instruction sent by the user. This capture instruction can be generated by the user pressing the first button on the dictionary pen, or by detecting a specific user gesture (such as an "OK" gesture). The triggering conditions for the first image acquisition are not limited to the above examples, and the embodiments of the present invention do not specifically limit the triggering conditions for the first image acquisition.
[0038] In addition, when the triggering condition for capturing the first image is met, one image may be captured as the first image, or multiple images may be captured and an image with the highest definition may be selected from the multiple images as the first image.
[0039] The triggering conditions for capturing the second image can be any of the following: ① detecting contact between the tip of the dictionary pen and the page containing the target text; ② detecting a capture instruction sent by the user, which can be generated by the user pressing the second button on the dictionary pen. The triggering conditions for capturing the second image are not limited to the above examples, and the embodiments of the present invention do not specifically limit the triggering conditions for capturing the second image. The first button and the second button can be the same button or different buttons, and the embodiments of the present invention do not specifically limit this.
[0040] Step 120: Match the first image and the second image to determine the position of the target text in the target page.
[0041] Specifically, the position of the target text in the page refers to the specific position of the target text in the target page, and the position may include information such as the paragraph and line number of the target text.
[0042] Optionally, the first image and the second image may be matched to determine an image region in the first image that matches the second image, and based on the position of the image region in the first image, the position of the target text in the target page may be determined.
[0043] Optionally, text recognition can be performed on the first image to obtain the page text, and text recognition can be performed on the second image to obtain the target text. The target text and the page text can then be matched. The text that matches the target text can be determined from the page text, and the position of the matched text in the page text can be used as the position of the target text in the target page. When performing text matching, an exact matching algorithm, a fuzzy matching algorithm, a string comparison algorithm, etc. can be used to match the target text with the page text.
[0044] For example, when using a string matching algorithm, if the target text is a complete sentence, phrase, or word, the target text can be compared with the page text character by character, and the text closest to the target text can be determined from the page text. This text is the matching text.
[0045] After identifying the matching text from the page text, the position of the matching text (including paragraph, line number, character offset, and other information) extracted by OCR can be used to calibrate the position of the target text on the target page. For example, if the target text is located on line 5 and between columns 2 and 5, the position of the target text on the target page can be identified as "line 5, start position: column 2, end position: column 5".
[0046] The position accuracy of the target text in the target page can be specific to the paragraph level, the line level, or the character level. The position accuracy can be determined according to user needs and is not specifically limited in this embodiment of the present invention.
[0047] Step 130: Determine a translation result of the target text from the translated text of the target page based on the position of the target text in the target page.
[0048] Specifically, the translated text is a text obtained by translating the page text of the target page from the source language to the target language, where the source language may be English and the target language may be Chinese.
[0049] As an optional embodiment, when translating the page text in advance, the entire content of the page text may be translated at one time. In this case, the translation text obtained is a translation text obtained based on context information, and the translation text is consistent with the original meaning of the page text.
[0050] As another optional embodiment, considering the high efficiency of machine translation, since machine translation usually has a limit on the length of the input text, when the text length of the page text is long, it may not be possible to complete the translation in one go. In this case, the page text can be split into multiple subtexts based on the upper limit of the input length of the machine translation, and then each subtext is translated separately to obtain the corresponding translated text of each subtext. Finally, according to the order of the subtexts in the page text, the translated texts of each subtext are spliced to form the translated text of the complete page text.
[0051] When splitting a page's text into sub-texts, it's important to ensure that each sub-text maintains its grammatical and semantic integrity, such as a complete sentence or paragraph. For example, suppose a sub-text is split into "withcomputers. A computer mouse allows users to move a pointer on the screen and select objects." Clearly, "with computers" and "A computer mouse allows users to move a pointer on the screen and select objects" are two separate sentences. Combining these two parts into one sub-text for translation may result in inaccurate translation. Therefore, it's possible to merge the "withcomputers" part into the previous sub-text and translate "A computer mouse allows users to move a pointer on the screen and select objects" as a separate sub-text.
[0052] The position of the target text in the target page is used to represent the specific location information of the target text in the page text, such as the paragraph and line number of the target text in the page text. Based on this position, the text matching this position can be extracted from the translation text as the translation result of the target text.
[0053] In addition, the layout format of the translated text can be the same as the layout format of the page text in the target page, that is, the position of the translation result of the target text in the translated text is the same as the position of the target text in the target page.
[0054] For example, the target text is located at page 2, paragraph 3, line 4 on the target page. If the typesetting format of the translated text is the same as that of the page text, the translation result of the target text is also located at page 2, paragraph 3, line 4 on the translated text. In this case, the text at page 2, paragraph 3, line 4 on the translated text can be directly extracted as the translation result of the target text.
[0055] The typesetting format of the translated text may also be different from the typesetting format of the page text. In this case, during the translation of the page text, a position mapping relationship between each word in the translated text and each word in the page text of the target page can be established. Based on the position of the target text in the target page, the position of the translation result of the target text in the translated text can be determined, and the translation result of the target text can be directly extracted from the translated text based on the position.
[0056] For example, the target text is located at page 2, paragraph 3, line 4 in the target page. If the typesetting format of the translated text is different from that of the page text, and based on the position mapping relationship between each segmentation in the translated text and each segmentation in the page text, it can be determined that the translation result of the target text is located at page 2, paragraph 2, line 5 in the translated text. In this case, the text at page 2, paragraph 2, line 5 in the translated text can be extracted as the translation result of the target text.
[0057] The translation method provided by the embodiment of the present invention matches the first image and the second image to determine the position of the target text in the target page, so as to extract the translation result of the target text from the translated text in combination with the position. Since the translated text is translated in combination with the original content of the target page, and the translation result is derived from the original translated text, the translation result of the target text can more accurately understand and convey the true meaning of the original text, ensure that the translation result is closer to the original meaning, and improve the accuracy of the translation result.
[0058] Based on the above embodiments, Figure 2 This is the second flow chart of the translation method provided by the present invention, such as Figure 2 As shown, the method includes: Step 210: Acquire a first image and a second image; the first image is an image of the target page where the target text is located, and the second image is an image of the target text.
[0059] Specifically, the first image is acquired by the first image acquisition element, and the second image is acquired by the second image acquisition element. The first image acquisition element and the second image acquisition element may be the same image acquisition element or different image acquisition elements. When the two are the same image acquisition element, the image acquisition element may be deployed at the tip of the dictionary pen; when the two are different image acquisition elements, the first acquisition element may be deployed at the end away from the tip of the dictionary pen, and the second image acquisition element may be deployed at the tip of the dictionary pen; when the two are different image acquisition elements, the first acquisition element and the second acquisition element may both be deployed near the tip of the dictionary pen. In addition, when the first image acquisition element and the second image acquisition element are different acquisition elements, the activation order of the first acquisition element and the second acquisition element may be different. For example, the second acquisition element may be activated when the tip of the dictionary pen touches the page after the first acquisition element is activated.
[0060] Step 220: Match the first image and the second image to determine the position of the target text in the target page.
[0061] Optionally, the first image and the second image may be matched to determine an image region in the first image that matches the second image, and based on the position of the image region in the first image, the position of the target text in the target page may be determined.
[0062] Step 230: Determine the regional text corresponding to the position from the translation text; translate the target text to determine multiple candidate translation results of the target text; match each candidate translation result with the regional text, and determine the translation result of the target text from each candidate translation result.
[0063] Specifically, considering that in some cases, the position of the target text in the target page may not be able to accurately indicate the specific position of the target text in the target page, that is, the position may only locate the rough position of the target text in the target page, and cannot locate the fine position of the target text in the target page. In this case, the regional text corresponding to the position in the translated text may not only include the translation result of the target text, but also include the translation results of other texts. That is, the regional text here can be understood as the text block in the translated text corresponding to the above-mentioned position, which includes the translation results of the target text and the translation results of other texts.
[0064] For example, the position of the target text on the target page is the 4th line of the 3rd paragraph on the 2nd page. The target text is the word "mouse". The text on the 4th line of the 3rd paragraph on the 2nd page of the target page is "A computer mouse allows users to move a pointer on the screen and select objects". It can be seen that the text on the 4th line of the 3rd paragraph on the 2nd page not only includes the word "mouse", but also includes other words. Therefore, in the regional text extracted from the translated text based on the position "the 4th line of the 3rd paragraph on the 2nd page", not only the translation result of "mouse" is included, but also the translation results of other words in "the 4th line of the 3rd paragraph on the 2nd page" are included.
[0065] Therefore, after determining the regional text from the translated text based on the position, it is necessary to accurately locate the translated text corresponding to the target text (i.e., the translation result of the target text) from the regional text.
[0066] In this case, the embodiment of the present invention translates the target text (i.e., translates the target text alone without combining the context of the target text). Since there may be polysemous words in the target text (for example, "mouse" can refer to a mouse or a computer mouse), there may be multiple candidate translation results when translating the target text alone. And there is a candidate translation result in these candidate translation results that matches a certain text in the regional text. This matching certain text can be used as the translation result of the target text, or the candidate translation result that matches a certain text can be used as the translation result of the target text.
[0067] Among them, the above matching can be understood as that the candidate translation result is exactly the same as a certain text in the regional text in terms of characters, or although there are differences in characters between the translation result and a certain text in the regional text, their semantics are similar.
[0068] For example, assume that the target text is "mouse", and a certain text in the regional text is "小鼠" (mouse in Chinese). The candidate translation results obtained by translating the target text alone include "老鼠" (rat) and "鼠标" (computer mouse). Obviously, although there are differences in characters between "老鼠" and "小鼠", their semantics are similar. However, there are not only differences in characters but also different semantics between "老鼠" and "鼠标". Therefore, "小鼠" in the regional text can be used as the translation result of the target text, or the candidate translation result "老鼠" can be used as the translation result of the target text.
[0069] Based on any of the above embodiments, Figure 3 is the third flowchart of the translation method provided by the present invention. As Figure 3 shown, the method includes: Step 310: Obtain a first image and a second image; the first image is an image of a target page where the target text is located, and the second image is an image of the target text.
[0070] Specifically, the first image is obtained by a first image acquisition component, and the second image is obtained by a second image acquisition component. Among them, the first image acquisition component and the second image acquisition component can be the same image acquisition component or different image acquisition components. The image acquisition component here can be a camera, an image sensor, etc. In the embodiments of the present invention, the image acquisition component is preferably a camera.
[0071] Step 320: Match the first image and the second image to determine the position of the target text on the target page.
[0072] Optionally, the first image and the second image can be matched to determine an image area in the first image that matches the second image, and based on the position of the image area in the first image, determine the position of the target text on the target page.
[0073] Step 330: Determine the regional text corresponding to the position from the translated text; translate the target text to determine multiple candidate translation results of the target text; perform character matching between each candidate translation result and the regional text to determine the edit distance between each candidate translation result and the regional text; based on the edit distance between each candidate translation result and the regional text, determine the translation result of the target text from each candidate translation result.
[0074] Specifically, character matching means comparing each character of the regional text with each candidate translation result one by one. The edit distance refers to the number of edit operations required when converting each candidate translation result into the regional text. The edit operations here can include inserting characters, deleting characters, replacing characters, etc. The more identical characters there are between each candidate translation result and the regional text, the fewer edit operations are required, and thus the corresponding edit distance is smaller. That is to say, the edit distance is used to measure the character difference between each candidate translation result and the regional text.
[0075] Optionally, the candidate translation result corresponding to the minimum edit distance can be used as the translation result of the target text.
[0076] For example, assume that the target text is "mouse", and the candidate translation results obtained by translating the target text alone include "mouse" and "rat". The regional text is "The mouse is an animal". When performing character matching between the candidate translation result "mouse" and the regional text, 3 characters need to be inserted after the candidate translation result "mouse" (inserting "is an animal"), that is, a total of 3 insert character operations are performed, which can be recorded as the edit distance between the candidate translation result "mouse" and the regional text is 3.
[0077] Perform character matching between the candidate translation result "mouse" and the regional text. It is necessary to first replace the candidate translation result "mouse" with "rat" (i.e., replace 2 characters), and then insert "is an animal" after "rat" (i.e., insert 3 characters). That is, perform 2 character replacement operations and 3 character insertion operations, which means a total of 5 editing operations. It can be recorded that the edit distance between the candidate translation result "mouse" and the regional text is 5.
[0078] Thus, it can be seen that the edit distance between the candidate translation result "rat" and the regional text is less than the edit distance between the candidate translation result "mouse" and the regional text. Furthermore, the candidate translation result "rat" can be used as the translation result of the target text "mouse".
[0079] Based on any of the above embodiments, Figure 4 is the fourth flow diagram of the translation method provided by the present invention. As Figure 4 shown, the method includes: Step 4-10: Obtain a first image and a second image; the first image is an image of the target page where the target text is located, and the second image is an image of the target text.
[0080] Specifically, the first image is obtained by a first image acquisition component, and the second image is obtained by a second image acquisition component. Among them, the first image acquisition component and the second image acquisition component can be the same image acquisition component or different image acquisition components. The image acquisition component here can be a camera, an image sensor, etc. In the embodiments of the present invention, the preferred image acquisition component is a camera.
[0081] Step 4-20: Match the first image and the second image to determine the position of the target text on the target page.
[0082] Optionally, the first image and the second image can be matched to determine the image area in the first image that matches the second image, and based on the position of the image area in the first image, determine the position of the target text on the target page.
[0083] Step 4-30: Determine the regional text corresponding to the position from the translation text; translate the target text to determine multiple candidate translation results of the target text; perform semantic matching between each candidate translation result and the regional text to determine the semantic similarity between each candidate translation result and the regional text; based on the semantic similarity between each candidate translation result and the regional text, determine the translation result of the target text from each candidate translation result.
[0084] Specifically, since there may be polysemous words in the target text (e.g., "mouse" can refer to a mouse or a mouse device), when translating the target text alone, there may be multiple candidate translation results. Among these candidate translation results, there is a candidate translation result that matches a certain text in the regional text. This matching certain text can be used as the translation result of the target text, or the candidate translation result that matches a certain text can be used as the translation result of the target text.
[0085] Among them, the above matching can be understood as that the candidate translation result is exactly the same as a certain text in the regional text in terms of characters, or although there are character differences between the translation result and a certain text in the regional text, their semantics are similar.
[0086] Based on this, considering that semantic similarity is used to measure whether the semantics between two texts are similar, in the embodiments of the present invention, semantic matching is performed between each candidate translation result and the regional text, and based on the semantic similarity between each candidate translation result and the regional text, the translation result of the target text is determined from each candidate translation result. The semantic similarity here can be measured by the semantic distance between the candidate translation result and the regional text. For example, extract the semantic features of the candidate translation result and the semantic features of the regional text, and calculate the distance between the two semantic features. Take this distance as the semantic distance between the candidate translation result and the regional text. The smaller the semantic distance, the higher the semantic similarity between the two, and thus the higher the probability that the corresponding candidate translation result is used as the translation result of the target text.
[0087] For example, assume that the target text is "mouse", and the candidate translation results obtained by translating the target text alone include "mouse" (referring to the animal) and "mouse device". The regional text is "A mouse is an animal". The regional text is used to describe the category of the mouse, that is, the regional text has a high semantic similarity with the candidate translation result "mouse" (referring to the animal). Therefore, "mouse" (referring to the animal) is used as the translation result of the target text.
[0088] Based on any of the above embodiments, Figure 5 is the fifth flow schematic diagram of the translation method provided by the present invention. As Figure 5 shown, the method includes: Step 510, obtain a first image and a second image; the first image is an image of the target page where the target text is located, and the second image is an image of the target text.
[0089] Specifically, the first image is obtained by collecting through a first image acquisition component, and the second image is obtained by collecting through a second image acquisition component. Among them, the first image acquisition component and the second image acquisition component can be the same image acquisition component or different image acquisition components.
[0090] Step 520: Match the first image and the second image to determine an image region in the first image that matches the second image; and determine a position of the target text in the target page based on the position of the image region in the first image.
[0091] Specifically, the image region refers to the image region in the first image that is identical to the image region in the second image. For example, if the second image is the image corresponding to "mouse" and the first image is the image corresponding to the target page where "mouse" is located, then the image region can be the image region corresponding to "mouse" in the first image, i.e., the image region is identical to the image region in the second image.
[0092] Among them, the image area can be determined by image matching between the first image and the second image. Image matching refers to detecting similar areas between two images through computer vision technology. For example, feature point-based methods (such as SIFT, SURF, ORB and other algorithms to extract and match key points), region-based methods (such as template matching, cross-correlation calculation, SSIM structural similarity measurement), deep learning-based methods (such as Siamese networks, LoFTR, SuperGlue and other end-to-end matching models) can be used to perform image matching between the first image and the second image.
[0093] After the image area is determined, the position of the image area in the first image can be used as the position of the target text in the target page. The position of the image area in the first image can be determined based on the coordinates of the image area (such as the coordinates of the upper left corner and lower right corner of the image area, or the coordinates of the center point of the image area).
[0094] Step 530: Determine a translation result of the target text from the translated text of the target page based on the position of the target text in the target page.
[0095] Specifically, the translated text is the text obtained by translating the target page's text from the source language into the target language. The translated text may be a pre-translated version of the target page. The position of the target text on the target page represents its specific location within the page, such as the paragraph and line number of the target text. Based on this location, text matching that location can be extracted from the translated text as the target text's translation result.
[0096] Based on any of the above embodiments, Figure 6 This is the sixth flow chart of the translation method provided by the present invention, such as Figure 6 As shown, the method includes: Step 610: Acquire a first image and a second image; the first image is an image of the target page where the target text is located, and the second image is an image of the target text.
[0097] Specifically, the first image is acquired by a first image acquisition element, and the second image is acquired by a second image acquisition element. The first image acquisition element and the second image acquisition element may be the same image acquisition element or different image acquisition elements. The image acquisition element may be a camera, an image sensor, or the like. In an embodiment of the present invention, the image acquisition element is preferably a camera, and the image acquisition element is disposed on the camera at the tip of the dictionary pen.
[0098] After detecting that the user has picked up the dictionary pen, the camera at the pen tip begins capturing a first image. This process can be continuous, meaning that as the user brings the pen tip closer to the paper, multiple images are captured, with the best quality image being used as the first image. During this process, the image capture range decreases as the distance between the pen tip and the paper decreases, while the clarity increases.
[0099] In addition, the image quality may be measured by definition. For example, the definition of each image may be determined by using the gradient magnitude, variance, or Laplace operator of each image, and the image with the highest definition may be used as the first image.
[0100] When the pen tip is pressed, the camera on the dictionary pen tip begins capturing a second image and stops capturing the second image when the pen tip is no longer pressed. Typically, the user presses the pen tip and slides it across the paper, with the sliding area covering the target text area. The resulting second image is the image of the target text.
[0101] Step 620: Divide the first image into blocks to obtain multiple sub-images; determine the image similarity between each sub-image and the second image; perform pixel-level matching on the sub-images whose image similarity is greater than a threshold value with the second image to determine the image area in the first image that matches the second image; and determine the position of the target text in the target page based on the position of the image area in the first image.
[0102] Specifically, if pixel-level matching is performed on the first image and the second image, pixel-level matching will be performed on irrelevant images in the first image and the second image, which increases the amount of calculation and wastes computing resources.
[0103] To this end, an embodiment of the present invention preferentially determines sub-images with higher image similarity with the second image from the first image, and performs pixel-level matching on each sub-image with the second image, thereby avoiding the problem of increased computational complexity caused by performing pixel-level matching on irrelevant images with the second image.
[0104] When determining a sub-image having a high degree of image similarity with the second image, the first image is first divided into blocks to obtain a plurality of sub-images. When dividing the first image into blocks, the first image can be divided into blocks based on the size of the second image. For example, the size of each sub-image can be set to be greater than or equal to the size of the second image.
[0105] When multiple sub-images are obtained, an image similarity between each sub-image and the second image is determined. This image similarity is used to characterize the consistency of visual content between each sub-image and the second image. A higher image similarity indicates a higher consistency of visual content. The image similarity between each sub-image and the second image can be measured using a feature distance between image features of each sub-image and image features of the second image. The smaller the feature distance, the higher the image similarity between the corresponding sub-image and the second image.
[0106] Considering that the text in the second image may be scattered in multiple sub-images, sub-images with image similarity greater than a threshold can be selected as candidate images, and each candidate image is matched with the second image at the pixel level to determine the pixels in each candidate image that are the same as the second image, and the area composed of all the same pixels is used as the image area.
[0107] Pixel-level matching refers to the correspondence between individual pixels of each candidate image and the second image to determine the correspondence of identical content in the two images. Dense optical flow methods, phase correlation methods, deep learning matching methods, and other methods can be used to perform pixel-level matching between the candidate image and the second image. Furthermore, the threshold value can be set based on actual conditions and is not specifically limited in this embodiment of the present invention.
[0108] Step 630: Determine a translation result of the target text from the translated text of the target page based on the position of the target text in the target page.
[0109] Specifically, the translated text is the text obtained by translating the target page's text from the source language into the target language. The translated text may be a pre-translated version of the target page. The position of the target text on the target page represents its specific location within the page, such as the paragraph and line number of the target text. Based on this location, text matching that location can be extracted from the translated text as the target text's translation result.
[0110] The translation device provided by the present invention is described below. The translation device described below and the translation method described above can be referenced to each other.
[0111] Based on any of the above embodiments, Figure 7 Schematic diagram of the structure of the translation device provided by the present invention. Figure 7As shown, the device includes: An acquisition unit 710 is configured to acquire a first image and a second image; the first image is an image of a target page where the target text is located, and the second image is an image of the target text; A determination unit 720 is configured to match the first image and the second image to determine a position of the target text in the target page; The translation unit 730 is configured to determine a translation result of the target text from the translated text of the target page based on the position of the target text in the target page.
[0112] Based on any of the above embodiments, determining a translation result of the target text from the translated text of the target page based on the position of the target text in the target page includes: Determine the regional text corresponding to the position from the translated text; Translate the target text and determine multiple candidate translation results of the target text; Each candidate translation result is matched with the regional text, and a translation result of the target text is determined from each candidate translation result.
[0113] Based on any of the above embodiments, matching each candidate translation result with the regional text and determining the translation result of the target text from each candidate translation result includes: Perform character matching on each candidate translation result and the regional text to determine the edit distance between each candidate translation result and the regional text; Based on the edit distance between each candidate translation result and the regional text, a translation result of the target text is determined from each candidate translation result.
[0114] Based on any of the above embodiments, matching each candidate translation result with the regional text and determining the translation result of the target text from each candidate translation result includes: Perform semantic matching between each candidate translation result and the regional text to determine the semantic similarity between each candidate translation result and the regional text; Based on the semantic similarity between each candidate translation result and the regional text, a translation result of the target text is determined from each candidate translation result.
[0115] Based on any of the above embodiments, matching the first image and the second image to determine the position of the target text in the target page includes: Matching the first image and the second image to determine an image region in the first image that matches the second image; The position of the target text in the target page is determined based on the position of the image area in the first image.
[0116] Based on any of the foregoing embodiments, matching the first image with the second image to determine an image region in the first image that matches the second image includes: Performing block processing on the first image to obtain multiple sub-images; determining an image similarity between each sub-image and the second image; Sub-images with image similarity greater than a threshold are matched with the second image at the pixel level to determine an image region in the first image that matches the second image.
[0117] Figure 8 Schematic diagram of the structure of the electronic device provided by the present invention, such as Figure 8 As shown, the electronic device may include: a processor 810, a communications interface 820, a memory 830, and a communications bus 840. The processor 810, the communications interface 820, and the memory 830 communicate with each other via the communications bus 840. The processor 810 may invoke logic instructions in the memory 830 to execute a translation method, which includes: acquiring a first image and a second image; the first image being an image of a target page containing target text, and the second image being an image of the target text; matching the first image and the second image to determine the location of the target text on the target page; and determining a translation result of the target text from the translated text on the target page based on the location of the target text on the target page.
[0118] Furthermore, the logic instructions in the aforementioned memory 830 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the portion that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product, stored in a storage medium, includes instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to perform all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, mobile hard drives, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical disks.
[0119] On the other hand, the present invention also provides a computer program product, which includes a computer program, which can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the translation method provided by the above methods, which includes: obtaining a first image and a second image; the first image is an image of a target page where a target text is located, and the second image is an image of the target text; matching the first image and the second image to determine the position of the target text in the target page; and determining the translation result of the target text from the translated text of the target page based on the position of the target text in the target page.
[0120] On the other hand, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, is implemented to execute the translation method provided by the above-mentioned methods, the method comprising: acquiring a first image and a second image; the first image is an image of a target page where a target text is located, and the second image is an image of the target text; matching the first image and the second image to determine the position of the target text in the target page; and determining the translation result of the target text from the translated text of the target page based on the position of the target text in the target page.
[0121] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, i.e., they may be located in one location or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of the present embodiment. Persons of ordinary skill in the art will be able to understand and implement the present invention without inventive effort.
[0122] Through the above description of the embodiments, those skilled in the art will clearly understand that each embodiment can be implemented using software plus a necessary general-purpose hardware platform, or of course, hardware. Based on this understanding, the essence of the above technical solution, or the portion that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a magnetic disk, or an optical disk, and includes a number of instructions for causing a computer device (such as a personal computer, server, or network device) to execute the methods described in each embodiment or certain portions of the embodiments.
[0123] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.
Claims
1. A translation method, characterized in that: include: Acquire a first image and a second image; The first image is an image of a target page where the target text is located, and the second image is an image of the target text; matching the first image and the second image to determine a position of the target text in the target page; Based on the position of the target text in the target page, a translation result of the target text is determined from the translated text of the target page.
2. The translation method according to claim 1, wherein: The determining the translation result of the target text from the translated text of the target page based on the position of the target text in the target page includes: Determining the regional text corresponding to the position from the translated text; Translating the target text to determine multiple candidate translation results of the target text; Each candidate translation result is matched with the regional text, and a translation result of the target text is determined from each candidate translation result.
3. The translation method according to claim 2, characterized in that The matching of each candidate translation result with the regional text and determining the translation result of the target text from each candidate translation result includes: Performing character matching on each candidate translation result and the regional text to determine the edit distance between each candidate translation result and the regional text; Based on the edit distance between each candidate translation result and the region text, a translation result of the target text is determined from each candidate translation result.
4. The translation method according to claim 2, characterized in that The matching of each candidate translation result with the regional text and determining the translation result of the target text from each candidate translation result includes: Performing semantic matching between each candidate translation result and the regional text to determine the semantic similarity between each candidate translation result and the regional text; Based on the semantic similarity between each candidate translation result and the regional text, a translation result of the target text is determined from each candidate translation result.
5. The translation method according to any one of claims 1 to 4, characterized in that: The matching of the first image and the second image to determine the position of the target text in the target page includes: matching the first image with the second image to determine an image region in the first image that matches the second image; The position of the target text in the target page is determined based on the position of the image area in the first image.
6. The translation method according to claim 5, characterized in that The matching the first image and the second image to determine an image area in the first image that matches the second image includes: performing block processing on the first image to obtain a plurality of sub-images; determining an image similarity between each sub-image and the second image; The sub-images having image similarity greater than a threshold are respectively matched with the second image at the pixel level to determine an image region in the first image that matches the second image.
7. A translation device, characterized in that: include: an acquisition unit, configured to acquire a first image and a second image; The first image is an image of a target page where the target text is located, and the second image is an image of the target text; a determination unit, configured to match the first image and the second image to determine a position of the target text in the target page; The translation unit is configured to determine a translation result of the target text from the translated text of the target page based on the position of the target text in the target page.
8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that: When the processor executes the computer program, the translation method according to any one of claims 1 to 6 is implemented.
9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the translation method according to any one of claims 1 to 6 is implemented.
10. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the translation method according to any one of claims 1 to 6 is implemented.