Translation method and device
By obtaining the page feature information and text images of the dictionary pen, the document and layout position of the target text are determined, and the translation results are extracted from the preset text, which solves the problem of inaccurate translation results of the dictionary pen and achieves more accurate translation.
Patent Information
- Application Number
- CN202510502054.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-21
- Publication Date
- 2025-08-12
AI Technical Summary
The translation results of existing dictionary pens are relatively accurate because they are separated from the context of the text to be translated, resulting in the translation results that do not conform to the original meaning.
By obtaining the page feature information of the target page and the image of the target text, the target document to which the target text belongs, and extracting the translation results from the preset text based on the layout position, which is the translated text of the target document.
It improves the accuracy of the translation results, makes the translation results closer to the original text meaning, and ensures the accuracy of the translation results.
Smart Images

Figure CN120471071A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of text translation, and in particular to a translation method and device. Background Art
[0002] As the demand for language learning grows, electronic aids such as dictionary pens and electronic dictionaries have gradually become common learning devices for users. Through these learning devices, users can directly query the translation results of the text to be translated.
[0003] Taking dictionary pens as an example, they typically use a camera to scan an image of the text to be translated, then perform text recognition on the image to obtain the text to be translated. This text is then translated to produce a translation result. However, this method only translates the text to be translated, resulting in low translation accuracy. Summary of the Invention
[0004] The present invention provides a translation method and device to solve the defects in the prior art.
[0005] The present invention provides a translation method, comprising the following steps: Acquire a first image and a second image, and perform image recognition on each image to obtain page feature information of a target page and target text; the first image is an image of the target page where the target text is located, and the second image is an image of the target text; Determining the target document to which the target text belongs based on the page feature information, and determining the typeset position of the target text in the target document; Based on the typeset position of the target text in the target document, a translation result of the target text is determined from preset texts, where the preset texts are translated texts of the target document.
[0006] According to a translation method provided by the present invention, The step of determining the translation result of the target text from the preset text based on the typeset position of the target text in the target document includes: Determining an initial translation text from the preset text based on the typeset position; Translating the target text to determine multiple candidate translation results of the target text; Each candidate translation result is matched with the initial translation text, and a translation result of the target text is determined from each candidate translation result.
[0007] According to a translation method provided by the present invention, matching each candidate translation result with the initial translation text and determining the translation result of the target text from each candidate translation result includes: Performing character matching on each candidate translation result and the initial translation text to determine the edit distance between each candidate translation result and the initial translation text; Based on the edit distance between each candidate translation result and the initial translation text, a translation result of the target text is determined from each candidate translation result.
[0008] According to a translation method provided by the present invention, matching each candidate translation result with the initial translation text and determining the translation result of the target text from each candidate translation result includes: Performing semantic matching on each candidate translation result and the initial translation text to determine the semantic similarity between each candidate translation result and the initial translation text; Based on the semantic similarity between each candidate translation result and the initial translation text, a translation result of the target text is determined from each candidate translation result.
[0009] According to a translation method provided by the present invention, the page feature information includes page text information; The determining the target document to which the target text belongs based on the page feature information includes: Extracting page keywords from the page text information; The target document to which the target text belongs is determined based on the page keywords.
[0010] According to a translation method provided by the present invention, the page feature information includes page text information and page layout information; The determining the target document to which the target text belongs based on the page feature information includes: Determining the document type of the target document based on the page layout information; Based on the document type, determining a target document library from a plurality of document libraries, wherein the target document library stores a plurality of documents, and each document has the same document type as the target document; Based on the page text information, the target document is determined from the target document library.
[0011] According to a translation method provided by the present invention, determining the typeset position of the target text in the target document includes: matching the first image with the second image to determine an image region in the first image that matches the second image; The layout position is determined based on the position of the image area in the first image.
[0012] The present invention also provides a translation device, comprising the following modules: an acquisition unit, configured to acquire a first image and a second image, and perform image recognition on each image to obtain page feature information of a target page and target text; the first image is an image of the target page where the target text is located, and the second image is an image of the target text; a determining unit, configured to determine the target document to which the target text belongs based on the page feature information, and determine the typeset position of the target text in the target document; The translation unit is configured to determine a translation result of the target text from preset texts based on a typeset position of the target text in the target document, wherein the preset texts are translation texts of the target document.
[0013] The present invention also provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements any of the above-described translation methods when executing the program.
[0014] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which implements any of the above-mentioned translation methods when executed by a processor.
[0015] The present invention also provides a computer program product, comprising a computer program, wherein when the computer program is executed by a processor, the computer program implements any of the above-mentioned translation methods.
[0016] The translation method and device provided by the present invention determine the target document to which the target text belongs based on page feature information, and determine the typeset position of the target text in the target document, so that the translation result of the target text can be extracted from the preset text in combination with the typeset position. Since the preset text is translated in combination with the original content of the target document, and the translation result of the target text is derived from the original preset text, the translation result of the target text can more accurately understand and convey the true meaning of the original text, ensure that the translation result is closer to the original meaning, and improve the accuracy of the translation result. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] In order to more clearly illustrate the technical solutions in the present invention or the prior art, a brief introduction is given below to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0018] Figure 1 This is one of the flow charts of the translation method provided by the present invention.
[0019] Figure 2 This is the second flowchart of the translation method provided by the present invention.
[0020] Figure 3 This is the third flow chart of the translation method provided by the present invention.
[0021] Figure 4 This is the fourth flow chart of the translation method provided by the present invention.
[0022] Figure 5 This is the fifth flow chart of the translation method provided by the present invention.
[0023] Figure 6 This is the sixth flowchart of the translation method provided by the present invention.
[0024] Figure 7 This is the seventh flow chart of the translation method provided by the present invention.
[0025] Figure 8 It is a structural diagram of the translation device provided by the present invention.
[0026] Figure 9 It is a structural schematic diagram of the electronic device provided by the present invention. DETAILED DESCRIPTION
[0027] To make the objectives, technical solutions, and advantages of the present invention more clear, the technical solutions of the present invention will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the embodiments described are only some of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the present invention.
[0028] Currently, most dictionary pens use a camera to scan an image of the text to be translated, then perform text recognition on the image to obtain the text to be translated. This is then translated using a translation engine to produce a translation result. However, these translation results only translate the text to be translated, resulting in low accuracy.
[0029] After analysis, it was found that the reason why the translation results of traditional dictionary pens are less accurate is that they are translated out of the context of the text to be translated, which makes the translation results deviate from the original meaning and reduces the accuracy of the translation results.
[0030] For example, the text to be translated is "mouse", and the context of the text to be translated is "She set the mouse on the table and moved the cursor to the icon". According to the context of the text to be translated, we can know that "mouse" refers to the mouse, but if translated out of context, "mouse" can mean either mouse or rat. If "mouse" is used as the translation result, it does not conform to the original meaning.
[0031] To this end, the present invention provides a translation method designed to extract a translation result of a text to be translated from a pre-set text (the pre-set text is the translation text of the document to which the text to be translated belongs) based on contextual translation, thereby improving the accuracy of the translation result. The method can be performed by a learning device (hereinafter referred to as "device"), such as a dictionary pen, an electronic dictionary, etc. To facilitate understanding of the technical solution of the present invention, the following embodiments use a dictionary pen as an example.
[0032] in, Figure 1 This is one of the flow charts of the translation method provided by the present invention, such as Figure 1 As shown, the method includes step 110 , step 120 and step 130 .
[0033] Step 110 , obtain a first image and a second image, and perform image recognition on each image to obtain page feature information and target text of the target page; the first image is an image corresponding to the target page where the target text is located, and the second image is an image of the target text.
[0034] Specifically, the target text is the text to be translated. The target text can be a single word, a phrase, a sentence, or a text containing multiple sentences. This is not specifically limited in the embodiments of the present invention. The target text can be obtained by performing text recognition on the second image. Here, text recognition can be achieved through optical character recognition (OCR) technology, template matching, sequence modeling technology, etc. Considering the cost and efficiency of implementation, in the embodiments of the present invention, OCR technology is preferably used to perform text recognition on the second image.
[0035] The target page is the page where the target text is located, that is, the target page includes the target text. For example, if the target text is "Hello", the target page may be a page including "Hello".
[0036] Page feature information refers to information contained in the target page, which is used to describe the target page's page content and layout. For example, page feature information may include page keywords, page layout information, page title information, header information, footer information, book title information, etc. Optionally, feature extraction can be performed on the first image to obtain image features, which are used to describe the target page's layout. Text recognition can be performed on the first image to obtain page text, and feature extraction can be performed on the page text to obtain text features, which are used to describe the target page's page content. Combining image features and text features can determine the target page's page feature information.
[0037] The first image refers to an image of the target page where the target text is located. The first image can be an image of the entire page where the target text is located. For example, if the target text is a short phrase on line 5 of page 1, the first image can be an image of the entire page of page 1. The first image can also be an image of a portion of the page where the target text is located. For example, if the target text is a short phrase on line 5 of page 1, the first image can be an image of the text area corresponding to lines 2 to 6 of page 1. The first image can be acquired by a first image acquisition component, which can be a camera, an image sensor, or the like.
[0038] The second image refers to an image of the target text, which is used to represent the visual information of the target text. For example, if the target text is "Hello", the second image is an image containing "Hello". The first image and the second image can be acquired by the same image acquisition component or by different image acquisition components. The image acquisition component here can be a camera, image sensor, etc.
[0039] The first image acquisition element and the second image acquisition element can be the same image acquisition element or different image acquisition elements. When the two are the same image acquisition element, the image acquisition element can be deployed at the tip of the dictionary pen; when the two are different image acquisition elements, the first acquisition element can be deployed at the end away from the tip of the dictionary pen, and the second image acquisition element can be deployed at the tip of the dictionary pen; when the two are different image acquisition elements, the first acquisition element and the second acquisition element can both be deployed near the tip of the dictionary pen. In addition, when the first image acquisition element and the second image acquisition element are different acquisition elements, the activation order of the first acquisition element and the second acquisition element may be different. For example, the second acquisition element can be activated when the tip of the dictionary pen touches the page after the first acquisition element is activated.
[0040] Taking the camera deployed at the tip of a dictionary pen as the first image acquisition element, the triggering condition for the first image acquisition can be any of the following: ① detecting that the vertical distance between the camera and the target page containing the target text is greater than 0 and less than or equal to a first threshold; ② detecting that the pen body of the dictionary pen changes from a horizontal orientation to an inclined state (the inclined state here means that the angle between the pen body and the horizontal direction is greater than 0°); ③ detecting a capture instruction sent by the user. This capture instruction can be generated by the user pressing the first button on the dictionary pen, or by detecting a specific user gesture (such as an "OK" gesture). The triggering conditions for the first image acquisition are not limited to the above examples, and the embodiments of the present invention do not specifically limit the triggering conditions for the first image acquisition.
[0041] In addition, when the triggering condition for capturing the first image is met, one image may be captured as the first image, or multiple images may be captured and an image with the highest definition may be selected from the multiple images as the first image.
[0042] The triggering conditions for capturing the second image can be any of the following: ① detecting contact between the tip of the dictionary pen and the page containing the target text; ② detecting a capture instruction sent by the user, which can be generated by the user pressing the second button on the dictionary pen. The triggering conditions for capturing the second image are not limited to the above examples, and the embodiments of the present invention do not specifically limit the triggering conditions for capturing the second image. The first button and the second button can be the same button or different buttons, and the embodiments of the present invention do not specifically limit this.
[0043] Step 120: Determine the target document to which the target text belongs based on the page feature information, and determine the layout position of the target text in the target document.
[0044] Specifically, the target document to which the target text belongs refers to the book, article, paper or other text material in printed or digital format that contains the target text, that is, the target document is used to identify the specific source of the target text.
[0045] Page feature information is used to describe the target page's content and layout. For example, page feature information may include page keywords, page layout information, page title information, header information, footer information, and book title information. Because each document contains different pages, the corresponding page content and / or page layout of each document also differ. This means that each page has different page feature information. Based on the target page's page feature information, the target document to which the target page belongs can be inferred.
[0046] For example, if the page feature information includes page keywords, since page keywords represent the core theme of a document and different documents have different core themes, the corresponding target document can be determined based on the page keywords. Alternatively, the corresponding keywords for different documents can be pre-determined, and the page keywords can be matched with the corresponding keywords for each document, with the matched document being used as the target document.
[0047] For another example, if the page feature information includes book title information, the target document can be directly determined based on the book title information, wherein the book title information can be located in the page title, the page header, or the page footer.
[0048] In addition, the typesetting position refers to the specific position of the target text in the target document. The typesetting position may include information such as the target page number where the target text is located, the paragraph where the target text is located, and the number of lines.
[0049] Optionally, the first image and the second image can be matched to determine an image region in the first image that matches the second image. This image region is the region in the first image where the target text is located, and the position of this image region in the first image is used as the layout position of the target text in the target page. Furthermore, the page feature information of the target page can include footer information, which typically includes the target page number. Based on this number, the layout position of the target page in the target document can be determined. Combined with the layout position of the target document in the target page, the layout position of the target text in the target document can be determined.
[0050] For example, if the target text is located at the third line of the second paragraph of the target page, and the target page is numbered page 4, then the target text's layout position in the target document is the third line of the second paragraph of page 4.
[0051] Step 130: Based on the typeset position of the target text in the target document, a translation result of the target text is determined from the preset text, where the preset text is the translated text of the target document.
[0052] Specifically, the preset text is the translated text of the target document, that is, it can be understood that the preset text is the text obtained by translating the target document in advance.
[0053] As an optional embodiment, when translating the target document in advance, the entire content of the target document may be translated at one time. In this case, the translation text obtained is a translation text obtained based on context information, and the translation text is consistent with the original meaning of the target document.
[0054] As another optional embodiment, considering the high efficiency of machine translation, since machine translation is usually limited by the length of the input text, when the text of the target document is long, it may not be possible to complete the translation in one go. In this case, the target document can be split into multiple subtexts based on the upper limit of the input length of machine translation, and then each subtext is translated separately to obtain the corresponding translated text. Finally, the translated texts of each subtext are spliced together according to the order of the subtexts in the target document to form the complete translated text of the target document, that is, the preset text.
[0055] When splitting the target document into subtexts, it is necessary to ensure that each subtext maintains grammatical and semantic integrity as much as possible, such as a complete sentence or paragraph. For example, suppose a subtext is split into "withcomputers. A computer mouse allows users to move a pointer on the screen and select objects." Obviously, "with computers" and "A computer mouse allows users to move a pointer on the screen and select objects" are two independent sentences. If these two parts are merged into one subtext for translation, it may result in inaccurate translation. Therefore, it is possible to merge the "withcomputers" part into the previous subtext and translate "A computer mouse allows users to move a pointer on the screen and select objects" as an independent subtext.
[0056] After determining the layout position, the specific location information of the target text in the target document can be obtained, such as the target page number, paragraph, and line number of the target text in the target document. Based on the layout position, the text that matches the layout position can be extracted from the preset text as the translation result of the target text.
[0057] In addition, the layout format of the preset text can be the same as that of the target document, that is, the layout position of the translation result of the target text in the preset text is the same as the layout position of the target text in the target document.
[0058] For example, the target text is typeset at page 2, paragraph 3, line 4 in the target document. If the typesetting format of the preset text is the same as that of the target document, the translation result of the target text is typeset at page 2, paragraph 3, line 4 in the preset text. In this case, the text on page 2, paragraph 3, line 4 in the preset text can be directly extracted as the translation result of the target text.
[0059] The typesetting format of the preset text may also be different from the typesetting format of the target document. In this case, during the translation of the target document, a typesetting position mapping relationship between each word in the preset text and each word in the target document can be established. Based on the typesetting position of the target text in the target document, the typesetting position of the translation result of the target text in the preset text can be determined, and the translation result of the target text can be directly extracted from the preset text based on the typesetting position.
[0060] For example, the target text is typeset at page 2, paragraph 3, line 4 in the target document. If the typesetting format of the preset text is different from that of the target document, and based on the typesetting position mapping relationship between each segmentation word in the preset text and each segmentation word in the target document, it can be determined that the typesetting position of the translation result of the target text in the preset text is page 2, paragraph 2, line 5. In this case, the text of page 2, paragraph 2, line 5 in the preset text can be extracted as the translation result of the target text.
[0061] The translation method provided by the embodiment of the present invention determines the target document to which the target text belongs based on page feature information, and determines the typeset position of the target text in the target document, so that the translation result of the target text can be extracted from the preset text in combination with the typeset position. Since the preset text is translated in combination with the original content of the target document, and the translation result is derived from the original preset text, the translation result of the target text can more accurately understand and convey the true meaning of the original text, ensure that the translation result is closer to the original meaning, and improve the accuracy of the translation result.
[0062] Based on the above embodiments, Figure 2 This is the second flow chart of the translation method provided by the present invention, such as Figure 2 As shown, the method includes: Step 210: Acquire a first image and a second image, and perform image recognition on each to obtain page feature information and target text of the target page; the first image is an image corresponding to the target page where the target text is located, and the second image is an image of the target text.
[0063] Specifically, the first image is acquired by the first image acquisition element, and the second image is acquired by the second image acquisition element. The first image acquisition element and the second image acquisition element may be the same image acquisition element or different image acquisition elements. When the two are the same image acquisition element, the image acquisition element may be deployed at the tip of the dictionary pen; when the two are different image acquisition elements, the first acquisition element may be deployed at the end away from the tip of the dictionary pen, and the second image acquisition element may be deployed at the tip of the dictionary pen; when the two are different image acquisition elements, the first acquisition element and the second acquisition element may both be deployed near the tip of the dictionary pen. In addition, when the first image acquisition element and the second image acquisition element are different acquisition elements, the activation order of the first acquisition element and the second acquisition element may be different. For example, the second acquisition element may be activated when the tip of the dictionary pen touches the page after the first acquisition element is activated.
[0064] Page feature information refers to the information contained in the target page, which is used to describe the page content and page layout of the target page. For example, page feature information may include page keywords, page layout information, page title information, header information, footer information, book title information, etc.
[0065] Step 220: Determine the target document to which the target text belongs based on the page feature information, and determine the layout position of the target text in the target document.
[0066] Specifically, since the pages in each document are different, the page content and / or page layout corresponding to the pages in each document are also different, that is, the page feature information of each page is different, and then based on the page feature information of the target page, the target document to which the target page belongs can be reversely inferred.
[0067] In addition, the typesetting position refers to the specific position of the target text in the target document. The typesetting position may include information such as the target page number where the target text is located, the paragraph where the target text is located, and the number of lines.
[0068] Step 230: Based on the typeset position, determine an initial translation text from the preset texts; translate the target text to determine multiple candidate translation results of the target text; match each candidate translation result with the initial translation text, and determine a translation result of the target text from each candidate translation result.
[0069] Specifically, considering that in some cases, the typeset position of the target text in the target document may not be able to accurately indicate the specific position of the target text in the target document, that is, the typeset position can only locate the rough position of the target text in the target document, but cannot locate the fine position of the target text in the target document. In this case, the initial translation text determined from the preset text based on the typeset position may include not only the translation results of the target text, but also the translation results of other texts. That is, the initial translation text here can be understood as the text corresponding to the typeset position in the preset text, which includes the translation results of the target text and the translation results of other texts.
[0070] For example, the target text is typeset at line 4, paragraph 3, page 2 in the target document, but the target text is the word "mouse." The text on line 4, paragraph 3, page 2 in the target document is "A computer mouse allows users to move a pointer on the screen and select objects." Therefore, it can be seen that the text on line 4, paragraph 3, page 2 includes not only the word "mouse," but also other words. Therefore, the initial translation text extracted from the preset text based on the typeset position "line 4, paragraph 3, page 2" includes not only the translation result for "mouse," but also the translation results for other words in "line 4, paragraph 3, page 2."
[0071] Therefore, after the initial translation text is determined from the preset text based on the typesetting position, it is necessary to accurately locate the translation text corresponding to the target text (ie, the translation result of the target text) from the initial translation text.
[0072] In this case, the embodiment of the present invention translates the target text (i.e., the target text is translated independently without considering the context of the target text). Since the target text may contain polysemous words (e.g., "mouse" can refer to both rat and mouse), when the target text is translated independently, multiple candidate translation results may exist, and one of these candidate translation results matches a certain text in the initial translation text. The matching text can be used as the translation result of the target text, and the candidate translation result that matches a certain text can also be used as the translation result of the target text.
[0073] The above-mentioned matching can be understood as that the candidate translation result is completely identical to a certain text in the initial translation text in terms of characters, or that the candidate translation result and a certain text in the initial translation text are different in terms of characters but have similar semantics.
[0074] For example, assume that the target text is "mouse", and in the initial translated text, a certain text is "mouse". The candidate translation results obtained by separately translating the target text include "rat" and "mouse". Obviously, although there are differences in characters between "rat" and "mouse", their semantics are similar, but there are not only differences in characters but also differences in semantics between "rat" and "mouse". Therefore, a certain text "mouse" in the initial translated text can be used as the translation result of the target text, and "rat" can also be used as the translation result of the target text.
[0075] Based on any of the above embodiments, Figure 3 is the third schematic flowchart of the translation method provided by the present invention. As Figure 3 shown, the method includes: Step 310: Obtain a first image and a second image, and perform image recognition on each to obtain the page feature information and the target text of the target page; the first image is the image corresponding to the target page where the target text is located, and the second image is the image of the target text.
[0076] Specifically, the first image is obtained by a first image acquisition component, and the second image is obtained by a second image acquisition component. Among them, the first image acquisition component and the second image acquisition component can be the same image acquisition component or different image acquisition components. The image acquisition component here can be a camera, an image sensor, etc. In the embodiments of the present invention, the image acquisition component is preferably a camera.
[0077] Step 320: Determine the target document to which the target text belongs based on the page feature information, and determine the typesetting position of the target text in the target document.
[0078] Specifically, since the pages in each document are different, and thus the page content and / or page layout corresponding to the pages in each document are also different, that is, the page feature information of each page is different. Therefore, according to the page feature information of the target page, the target document to which the target page belongs can be inferred in reverse.
[0079] In addition, the typesetting position refers to the specific position of the target text in the target document, and the typesetting position can include information such as the number of the target page where the target text is located, the paragraph where the target text is located, and the line number.
[0080] Step 330: Determine the initial translated text from the preset text based on the typesetting position; translate the target text to determine multiple candidate translation results of the target text; perform character matching between each candidate translation result and the initial translated text to determine the edit distance between each candidate translation result and the initial translated text; and determine the translation result of the target text from each candidate translation result based on the edit distance between each candidate translation result and the initial translated text.
[0081] Specifically, character matching refers to comparing the initial translation text character by character with each candidate translation result. The edit distance refers to the number of editing operations required when converting each candidate translation result into the initial translation text. The editing operations here can include inserting characters, deleting characters, replacing characters, etc. The more identical characters there are between each candidate translation result and the initial translation text, the fewer the number of editing operations required, and thus the smaller the corresponding edit distance. That is to say, the edit distance is used to measure the character differences between each candidate translation result and the initial translation text.
[0082] Optionally, the candidate translation result corresponding to the minimum edit distance can be used as the translation result of the target text.
[0083] For example, assume that the target text is "mouse". The candidate translation results obtained by separately translating the target text include "mouse" and "rat". The initial translation text is "A mouse is an animal". When performing character matching between the candidate translation result "mouse" and the initial translation text, 3 characters need to be inserted after the candidate translation result "mouse" (inserting "is an animal"), that is, a total of 3 character insertion operations are performed. It can be recorded that the edit distance between the candidate translation result "mouse" and the initial translation text is 3.
[0084] When performing character matching between the candidate translation result "rat" and the initial translation text, the candidate translation result "rat" needs to be first replaced with "mouse" (i.e., replacing 2 characters), and then "is an animal" needs to be inserted after "mouse" (i.e., inserting 3 characters), that is, 2 character replacement operations and 3 character insertion operations are performed, that is, a total of 5 editing operations are performed. It can be recorded that the edit distance between the candidate translation result "rat" and the initial translation text is 5.
[0085] It can be seen from this that the edit distance between the candidate translation result "mouse" and the initial translation text is less than the edit distance between the candidate translation result "rat" and the initial translation text. Therefore, the candidate translation result "mouse" can be used as the translation result of the target text "mouse".
[0086] Based on any of the above embodiments, Figure 4 is the fourth flowchart of the translation method provided by the present invention. As Figure 4 shown, the method includes: Step 410: Obtain a first image and a second image, and perform image recognition on each to obtain the page feature information of the target page and the target text; the first image is the image corresponding to the target page where the target text is located, and the second image is the image of the target text.
[0087] Specifically, the first image is acquired by a first image acquisition element, and the second image is acquired by a second image acquisition element. The first image acquisition element and the second image acquisition element may be the same image acquisition element or different image acquisition elements. The image acquisition element may be a camera, an image sensor, or the like. In the embodiment of the present invention, the image acquisition element is preferably a camera.
[0088] Step 420: Determine the target document to which the target text belongs based on the page feature information, and determine the layout position of the target text in the target document.
[0089] Specifically, since the pages in each document are different, the page content and / or page layout corresponding to the pages in each document are also different, that is, the page feature information of each page is different, and then based on the page feature information of the target page, the target document to which the target page belongs can be reversely inferred.
[0090] In addition, the typesetting position refers to the specific position of the target text in the target document. The typesetting position may include information such as the target page number where the target text is located, the paragraph where the target text is located, and the number of lines.
[0091] Step 430: Determine an initial translation text from the preset text based on the typeset position; translate the target text to determine multiple candidate translation results of the target text; semantically match each candidate translation result with the initial translation text to determine the semantic similarity between each candidate translation result and the initial translation text; and determine a translation result of the target text from each candidate translation result based on the semantic similarity between each candidate translation result and the initial translation text.
[0092] Specifically, since the target text may contain polysemous words (for example, "mouse" can refer to both rat and mouse), when translating the target text alone, there may be multiple candidate translation results, and among these candidate translation results, there is a candidate translation result that matches a text in the initial translation text. The matching text can be used as the translation result of the target text, and the candidate translation result that matches a text can also be used as the translation result of the target text.
[0093] The above-mentioned matching can be understood as that the candidate translation result is completely identical to a certain text in the initial translation text in terms of characters, or that the candidate translation result and a certain text in the initial translation text are different in terms of characters but have similar semantics.
[0094] Based on this, considering that semantic similarity is used to measure whether the semantics between two texts are similar, in the embodiments of the present invention, semantic matching is performed between each candidate translation result and the initial translation text, and based on the semantic similarity between each candidate translation result and the initial translation text, the translation result of the target text is determined from each candidate translation result. The semantic similarity here can be measured by the semantic distance between the candidate translation result and the initial translation text. For example, the semantic features of the candidate translation result and the semantic features of the initial translation text are extracted, and the distance between the two semantic features is calculated. This distance is used as the semantic distance between the candidate translation result and the initial translation text. The smaller the semantic distance, the higher the semantic similarity between the two, and thus the higher the probability that the corresponding candidate translation result is the translation result of the target text.
[0095] For example, assume that the target text is "mouse". The candidate translation results obtained by translating the target text alone include "rat" and "mouse". The initial translation text is "A rat is an animal". The initial translation text is used to describe the category of a rat, that is, the semantic similarity between the initial translation text and the candidate translation result "rat" is relatively high. Therefore, "rat" is used as the translation result of the target text.
[0096] Based on any of the above embodiments, Figure 5 is the fifth flowchart of the translation method provided by the present invention. As Figure 5 shown, the method includes: Step 510: Obtain a first image and a second image, and perform image recognition on each of them to obtain the page feature information of the target page and the target text; the first image is the image corresponding to the target page where the target text is located, and the second image is the image of the target text.
[0097] Specifically, the first image is collected by a first image acquisition component, and the second image is collected by a second image acquisition component. Among them, the first image acquisition component and the second image acquisition component can be the same image acquisition component or different image acquisition components.
[0098] Step 520: Extract the page topic words from the page text information; determine the target document to which the target text belongs based on the page topic words. The page feature information includes the page text information. Determine the layout position of the target text in the target document.
[0099] Specifically, the page feature information includes the page text information. The page text information refers to all the text contents in the target page. The page text information may include information such as the page title text, header text, footer text, and page body text in the target page.
[0100] Page keywords refer to words that represent the core theme of a target page. These keywords can be obtained by performing natural language processing or keyword extraction on the page text. For example, they can be obtained by performing a word frequency analysis on the page text.
[0101] Considering that different target documents express different subject information, and the page keywords are used to represent the subject information of the target page, we can determine the target document to which the target text belongs by comparing the similarity between the page keywords and the keywords of each document.
[0102] For example, the subject words corresponding to each document may be predetermined, and the page subject words may be matched with the subject words corresponding to each document, and the matched document may be used as the target document.
[0103] Step 530: Based on the typeset position of the target text in the target document, determine the translation result of the target text from the preset text, where the preset text is the translated text of the target document.
[0104] Specifically, after determining the typesetting position, the specific location information of the target text in the target document can be obtained, such as the target page number, paragraph, and line number of the target text in the target document. Based on the typesetting position, the text that matches the typesetting position can be extracted from the preset text as the translation result of the target text.
[0105] Based on any of the above embodiments, Figure 6 This is the sixth flow chart of the translation method provided by the present invention, such as Figure 6 As shown, the method includes: Step 610: Acquire a first image and a second image, and perform image recognition on each to obtain page feature information and target text of the target page; the first image is an image corresponding to the target page where the target text is located, and the second image is an image of the target text.
[0106] Specifically, the first image is acquired by a first image acquisition component, and the second image is acquired by a second image acquisition component. The first image acquisition component and the second image acquisition component can be the same image acquisition component or different image acquisition components.
[0107] Step 620: Determine the document type of the target document based on the page layout information. Based on the document type, determine a target document library from multiple document libraries, where the target document library stores multiple documents, each of which has the same document type as the target document. Determine the target document from the target document library based on the page text information, where the page feature information includes page text information and page layout information. Determine the layout position of the target text within the target document.
[0108] Specifically, page layout information refers to the layout of elements on the target page, including the arrangement and relative positions of paragraphs, titles, charts, images, etc. The document type of the target document refers to the category of the target document, such as papers, newspapers, books, technical reports, etc.
[0109] Different document types have different corresponding page layout information. For example, for a paper, its page layout is usually highly structured, usually including the title, abstract, main text, and references. For a newspaper or journal, its page layout is usually mainly based on the title and content blocks, and may include pictures and charts, with a relatively compact and flexible layout.
[0110] If you match the target document to the target text across all documents, you might waste computing resources due to type mismatches. For example, if the target document is a thesis, then if you match it against newspapers and periodicals, you won't find any matching documents because they are of different types. In this case, you only need to match against thesis documents, saving both matching time and computing resources.
[0111] Based on this, the embodiment of the present invention determines a target document library from multiple document libraries based on the document type of the target document. The target document library stores multiple documents of the same document type as the target document. It is understood that multiple document libraries can be pre-set, each corresponding to a different type of document, such as document library 1 storing papers, document library 2 storing newspapers and periodicals, document library 3 storing books, etc.
[0112] After determining the target document library, the method of the above embodiment can be referred to to determine the target document to which the target text belongs based on the page text information, such as extracting page keywords from the page text information; and determining the target document to which the target text belongs based on the page keywords.
[0113] Step 630: Based on the typeset position of the target text in the target document, determine the translation result of the target text from the preset text, where the preset text is the translated text of the target document.
[0114] Specifically, after determining the typesetting position, the specific location information of the target text in the target document can be obtained, such as the target page number, paragraph, and line number of the target text in the target document. Based on the typesetting position, the text that matches the typesetting position can be extracted from the preset text as the translation result of the target text.
[0115] Based on any of the above embodiments, Figure 7 This is the seventh flow chart of the translation method provided by the present invention, such as Figure 7 As shown, the method includes: Step 710: Acquire a first image and a second image, and perform image recognition on each to obtain page feature information of a target page and target text; the first image is an image corresponding to the target page where the target text is located, and the second image is an image of the target text.
[0116] Specifically, the first image is acquired by a first image acquisition element, and the second image is acquired by a second image acquisition element. The first image acquisition element and the second image acquisition element may be the same image acquisition element or different image acquisition elements. The image acquisition element may be a camera, an image sensor, or the like. In an embodiment of the present invention, the image acquisition element is preferably a camera, and the image acquisition element is disposed on the camera at the tip of the dictionary pen.
[0117] After detecting that the user has picked up the dictionary pen, the camera at the pen tip begins capturing a first image. This process can be continuous, meaning that as the user brings the pen tip closer to the paper, multiple images are captured, with the best quality image being used as the first image. During this process, the image capture range decreases as the distance between the pen tip and the paper decreases, while the clarity increases.
[0118] In addition, the image quality may be measured by definition. For example, the definition of each image may be determined by using the gradient magnitude, variance, or Laplace operator of each image, and the image with the highest definition may be used as the first image.
[0119] When the pen tip is pressed, the camera on the dictionary pen tip begins capturing a second image and stops capturing the second image when the pen tip is no longer pressed. Typically, the user presses the pen tip and slides it across the paper, with the sliding area covering the target text area. The resulting second image is the image of the target text.
[0120] Step 720: Determine the target document to which the target text belongs based on the page feature information. Match the first image with the second image to determine the image region in the first image that matches the second image; and determine the typesetting position based on the position of the image region in the first image.
[0121] The first image contains an image of the target page where the target text is located, and the second image contains an image of the target text. This means that the first image contains the target text in the second image. Based on this, the first and second images are matched to determine which region in the first image has the highest similarity to the second image. This region with the highest similarity is then used as the image region of the target text in the first image. This image region refers to the region in the first image where the target text is located.
[0122] After determining the image area, the position of the image area in the first image can be used as the layout position of the target text in the target page. Combined with the layout position of the target page in the document, the layout position of the target text in the target document can be determined. The position of the image area in the first image can be determined based on the coordinates of the image area (e.g., the coordinates of the upper left and lower right corners of the image area, or the coordinates of the center point of the image area).
[0123] Step 730: Based on the typeset position of the target text in the target document, determine the translation result of the target text from the preset text, where the preset text is the translated text of the target document.
[0124] Specifically, after determining the typesetting position, the specific location information of the target text in the target document can be obtained, such as the target page number, paragraph, and line number of the target text in the target document. Based on the typesetting position, the text that matches the typesetting position can be extracted from the preset text as the translation result of the target text.
[0125] The translation device provided by the present invention is described below. The translation device described below and the translation method described above can be referenced to each other.
[0126] Based on any of the above embodiments, Figure 8 Schematic diagram of the structure of the translation device provided by the present invention. Figure 8 As shown, the device includes: The acquisition unit 810 is configured to acquire a first image and a second image, and perform image recognition on each image to obtain page feature information of a target page and target text; the first image is an image of the target page where the target text is located, and the second image is an image of the target text; A determination unit 820 is configured to determine the target document to which the target text belongs based on the page feature information, and to determine the typesetting position of the target text in the target document; The translation unit 830 is configured to determine a translation result of the target text from a preset text based on the typeset position of the target text in the target document, where the preset text is the translated text of the target document.
[0127] Based on any of the above embodiments, determining the translation result of the target text from the preset text based on the typeset position of the target text in the target document includes: Determine the initial translation text from the preset text based on the typesetting position; Translate the target text and determine multiple candidate translation results of the target text; Each candidate translation result is matched with the initial translation text, and a translation result of the target text is determined from each candidate translation result.
[0128] Based on any of the above embodiments, matching each candidate translation result with the initial translation text, and determining the translation result of the target text from each candidate translation result includes: Perform character matching on each candidate translation result and the initial translation text to determine the edit distance between each candidate translation result and the initial translation text; Based on the edit distance between each candidate translation result and the initial translation text, a translation result of the target text is determined from each candidate translation result.
[0129] Based on any of the above embodiments, matching each candidate translation result with the initial translation text, and determining the translation result of the target text from each candidate translation result includes: Perform semantic matching between each candidate translation result and the initial translation text to determine the semantic similarity between each candidate translation result and the initial translation text; Based on the semantic similarity between each candidate translation result and the initial translation text, a translation result of the target text is determined from each candidate translation result.
[0130] Based on any of the above embodiments, the page feature information includes page text information; Determine the target document to which the target text belongs based on page feature information, including: Extract page keywords from page text information; Determine the target document to which the target text belongs based on the page keywords.
[0131] Based on any of the above embodiments, the page feature information includes page text information and page layout information; Determine the target document to which the target text belongs based on page feature information, including: Determine the document type of the target document based on the page layout information; Based on the document type, a target document library is determined from multiple document libraries, where the target document library stores multiple documents, each of which has the same document type as the target document. Based on the page text information, the target document is determined from the target document library.
[0132] Based on any of the above embodiments, determining the typeset position of the target text in the target document includes: Matching the first image and the second image to determine an image region in the first image that matches the second image; A layout position is determined based on a position of the image area in the first image.
[0133] Figure 9 Schematic diagram of the structure of the electronic device provided by the present invention, such as Figure 9As shown, the electronic device may include: a processor 910, a communications interface 920, a memory 930, and a communications bus 940, wherein the processor 910, the communications interface 920, and the memory 930 communicate with each other via the communications bus 940. The processor 910 may invoke logic instructions in the memory 930 to execute a translation method, which includes: acquiring a first image and a second image, and performing image recognition on each image to obtain page feature information and target text of a target page; the first image is an image of the target page where the target text is located, and the second image is an image of the target text; determining the target document to which the target text belongs based on the page feature information, and determining the layout position of the target text in the target document; and determining a translation result of the target text from preset text based on the layout position of the target text in the target document, wherein the preset text is the translated text of the target document.
[0134] Furthermore, the logic instructions in the aforementioned memory 930 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the portion that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product, stored in a storage medium, includes instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to perform all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as a USB flash drive, a mobile hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.
[0135] On the other hand, the present invention also provides a computer program product, which includes a computer program, which can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the translation method provided by the above methods, which includes: obtaining a first image and a second image, and performing image recognition respectively to obtain page feature information and target text of the target page; the first image is an image of the target page where the target text is located, and the second image is an image of the target text; based on the page feature information, determining the target document to which the target text belongs, and determining the typeset position of the target text in the target document; based on the typeset position of the target text in the target document, determining the translation result of the target text from a preset text, wherein the preset text is the translated text of the target document.
[0136] On the other hand, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, is implemented to execute the translation method provided by the above-mentioned methods, the method comprising: acquiring a first image and a second image, and performing image recognition respectively to obtain page feature information and target text of the target page; the first image is an image of the target page where the target text is located, and the second image is an image of the target text; determining the target document to which the target text belongs based on the page feature information, and determining the typeset position of the target text in the target document; determining the translation result of the target text from a preset text based on the typeset position of the target text in the target document, the preset text being the translated text of the target document.
[0137] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, i.e., they may be located in one location or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of the present embodiment. Persons of ordinary skill in the art will be able to understand and implement the present invention without inventive effort.
[0138] Through the above description of the embodiments, those skilled in the art will clearly understand that each embodiment can be implemented using software plus a necessary general-purpose hardware platform, or of course, hardware. Based on this understanding, the essence of the above technical solution, or the portion that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a magnetic disk, or an optical disk, and includes a number of instructions for causing a computer device (such as a personal computer, server, or network device) to execute the methods described in each embodiment or certain portions of the embodiments.
[0139] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.
Claims
1. A translation method, characterized in that: include: Acquire the first image and the second image, and perform image recognition on each image to obtain page feature information and target text of the target page; The first image is an image of the target page where the target text is located, and the second image is an image of the target text; Determining the target document to which the target text belongs based on the page feature information, and determining the typeset position of the target text in the target document; Based on the typeset position of the target text in the target document, a translation result of the target text is determined from preset texts, where the preset texts are translated texts of the target document.
2. The translation method according to claim 1, wherein: The step of determining the translation result of the target text from the preset text based on the typeset position of the target text in the target document includes: Determining an initial translation text from the preset text based on the typeset position; Translating the target text to determine multiple candidate translation results of the target text; Each candidate translation result is matched with the initial translation text, and a translation result of the target text is determined from each candidate translation result.
3. The translation method according to claim 2, characterized in that The step of matching each candidate translation result with the initial translation text and determining the translation result of the target text from each candidate translation result includes: Performing character matching on each candidate translation result and the initial translation text to determine the edit distance between each candidate translation result and the initial translation text; Based on the edit distance between each candidate translation result and the initial translation text, a translation result of the target text is determined from each candidate translation result.
4. The translation method according to claim 2, characterized in that The step of matching each candidate translation result with the initial translation text and determining the translation result of the target text from each candidate translation result includes: Performing semantic matching on each candidate translation result and the initial translation text to determine the semantic similarity between each candidate translation result and the initial translation text; Based on the semantic similarity between each candidate translation result and the initial translation text, a translation result of the target text is determined from each candidate translation result.
5. The translation method according to any one of claims 1 to 4, characterized in that: The page feature information includes page text information; The determining the target document to which the target text belongs based on the page feature information includes: Extracting page keywords from the page text information; The target document to which the target text belongs is determined based on the page keywords.
6. The translation method according to any one of claims 1 to 4, characterized in that: The page feature information includes page text information and page layout information; The determining the target document to which the target text belongs based on the page feature information includes: Determining the document type of the target document based on the page layout information; Based on the document type, determining a target document library from a plurality of document libraries, wherein the target document library stores a plurality of documents, and each document has the same document type as the target document; Based on the page text information, the target document is determined from the target document library.
7. The translation method according to any one of claims 1 to 4, characterized in that: Determining the typeset position of the target text in the target document includes: matching the first image with the second image to determine an image region in the first image that matches the second image; The layout position is determined based on the position of the image area in the first image.
8. A translation device, characterized in that: include: an acquisition unit, configured to acquire the first image and the second image, and perform image recognition on each image to obtain page feature information and target text of the target page; The first image is an image of the target page where the target text is located, and the second image is an image of the target text; a determining unit, configured to determine the target document to which the target text belongs based on the page feature information, and determine the typeset position of the target text in the target document; The translation unit is configured to determine a translation result of the target text from preset texts based on a typeset position of the target text in the target document, wherein the preset texts are translation texts of the target document.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that: When the processor executes the computer program, the translation method according to any one of claims 1 to 7 is implemented.
10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the translation method according to any one of claims 1 to 7 is implemented.
11. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the translation method according to any one of claims 1 to 7 is implemented.