Translation method and device

By obtaining the position of the target text in the page text in the dictionary pen and translating it with the context information of the page text, the problem of inaccurate translation results of the dictionary pen is solved, and more accurate translation results are achieved.

CN120471069APending Publication Date: 2025-08-12HEFEI IFLYTEK TOYCLOUD TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510502052.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-21
Publication Date
2025-08-12

AI Technical Summary

Technical Problem

The translation results of existing dictionary pens are relatively accurate because they are separated from the context of the text to be translated, resulting in the translation results that do not conform to the original meaning.

Method used

By obtaining the position of the target text in the page text, and combining the context text in the page text for translation, including paragraph text and candidate text, OCR technology is used to identify the text and obtain the semantic information of the attached figures in combination with semantic understanding, and determine the translation results.

Benefits of technology

It improves the accuracy of the translation results, makes the translation results closer to the original text meaning, and enhances the accuracy of the translation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120471069A_ABST
    Figure CN120471069A_ABST
Patent Text Reader

Abstract

The invention provides a translation method and device, and the method comprises the steps: obtaining a first image and a second image, and carrying out the text recognition of the first image and the second image, and obtaining a page text and a target text; the first image is an image of a page where the target text is located, and the second image is an image of the target text; determining the position of the target text in the page text; based on the position of the target text in the page text, extracting a context text of the target text from the page text; and based on the target text and the context text, determining a translation result of the target text. According to the translation method and device, the target text and the context text are combined for translation, the real meaning of the original text can be more accurately understood and conveyed, it is ensured that the translation result is closer to the meaning of the original text, and the accuracy of the translation result is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of text translation, and in particular to a translation method and device. Background Art

[0002] As the demand for language learning grows, electronic aids such as dictionary pens and electronic dictionaries have gradually become common learning devices for users. Through these learning devices, users can directly query the translation results of the text to be translated.

[0003] Taking dictionary pens as an example, they typically use a camera to scan an image of the text to be translated, then perform text recognition on the image to obtain the text to be translated. This text is then translated to produce a translation result. However, this method only translates the text to be translated, resulting in low translation accuracy. Summary of the Invention

[0004] The present invention provides a translation method and device to solve the defects in the prior art.

[0005] The present invention provides a translation method, comprising the following steps: Acquire a first image and a second image, and perform text recognition on each image to obtain page text and target text; the first image is an image of the page where the target text is located, and the second image is an image of the target text; Determining the position of the target text in the page text; extracting context text of the target text from the page text based on the position of the target text in the page text; A translation result of the target text is determined based on the target text and the context text.

[0006] According to a translation method provided by the present invention, extracting the context text of the target text from the page text based on the position of the target text in the page text includes: Based on the position, determining paragraph text from the page text, the paragraph text being the text corresponding to the paragraph where the target text is located; extracting a plurality of candidate texts from the page text based on the keywords of the paragraph text, wherein the semantic relevance between the candidate texts and the keywords is greater than a first threshold; The context text is determined based on the paragraph text and the plurality of candidate texts.

[0007] According to a translation method provided by the present invention, determining the context text based on the paragraph text and the multiple candidate texts includes: In a case where the reference text includes drawing number information, acquiring a target image from the first image, the target image being an image of the drawing corresponding to the drawing number information, and the reference text including the paragraph text and / or the candidate text; Performing semantic understanding on the target image to obtain semantic text of the target image; The context text is determined from the reference text and the semantic text.

[0008] According to a translation method provided by the present invention, obtaining a target image from the first image includes: Extracting the description text of the drawing number information from the reference text; extracting all accompanying images from the first image; The target image is determined from each of the accompanying images based on the semantic information of the description text and the semantic information of each of the accompanying images.

[0009] According to a translation method provided by the present invention, determining the context text based on the paragraph text and the multiple candidate texts includes: Determining a relevant text from each candidate text based on the semantic relevance between the paragraph text and each candidate text; the semantic relevance between the relevant text and the paragraph text is greater than a second threshold; The paragraph text and the related text are used as the context text.

[0010] According to a translation method provided by the present invention, determining the position of the target text in the page text includes: Performing text matching on the target text and the page text, and determining matching text from the page text, wherein the matching text is text in the page text that is identical to the target text; The position of the matching text in the page text is used as the position of the target text in the page text.

[0011] According to a translation method provided by the present invention, determining a translation result of the target text based on the target text and the context text includes: Merging the target text and the context text to obtain a composite text, and determining a position of the target text in the composite text; translating the synthesized text to obtain a translated text; A translation result of the target text is extracted from the translation text based on the position of the target text in the synthesized text.

[0012] The present invention also provides a translation device, comprising the following modules: an acquisition unit, configured to acquire a first image and a second image, and perform text recognition on each image to obtain page text and target text; the first image is an image of the page where the target text is located, and the second image is an image of the target text; A determination unit, configured to determine a position of the target text in the page text; an extraction unit, configured to extract context text of the target text from the page text based on a position of the target text in the page text; The translation unit is configured to determine a translation result of the target text based on the target text and the context text.

[0013] The present invention also provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements any of the above-described translation methods when executing the program.

[0014] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which implements any of the above-mentioned translation methods when executed by a processor.

[0015] The present invention also provides a computer program product, comprising a computer program, wherein when the computer program is executed by a processor, the computer program implements any of the above-mentioned translation methods.

[0016] The translation method and device provided by the present invention are based on the position of the target text in the page text, so that the context text of the target text can be extracted from the page text in combination with the position. Since the context text is closely related to the semantics of the target text, the target text and the context text are combined for translation, which can more accurately understand and convey the true meaning of the original text, ensure that the translation result is closer to the original meaning, and improve the accuracy of the translation result. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] In order to more clearly illustrate the technical solutions in the present invention or the prior art, a brief introduction is given below to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0018] Figure 1 This is one of the flow charts of the translation method provided by the present invention.

[0019] Figure 2 This is the second flowchart of the translation method provided by the present invention.

[0020] Figure 3This is the third flow chart of the translation method provided by the present invention.

[0021] Figure 4 This is the fourth flow chart of the translation method provided by the present invention.

[0022] Figure 5 This is the fifth flow chart of the translation method provided by the present invention.

[0023] Figure 6 This is the sixth flowchart of the translation method provided by the present invention.

[0024] Figure 7 This is the seventh flow chart of the translation method provided by the present invention.

[0025] Figure 8 It is a structural diagram of the translation device provided by the present invention.

[0026] Figure 9 It is a structural schematic diagram of the electronic device provided by the present invention. DETAILED DESCRIPTION

[0027] To make the objectives, technical solutions, and advantages of the present invention more clear, the technical solutions of the present invention will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the embodiments described are only some of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the present invention.

[0028] Currently, most dictionary pens use a camera to scan an image of the text to be translated, then perform text recognition on the image to obtain the text to be translated. This is then translated using a translation engine to produce a translation result. However, these translation results only translate the text to be translated, resulting in low accuracy.

[0029] After analysis, it was found that the reason why the translation results of traditional dictionary pens are less accurate is that they are translated out of the context of the text to be translated, which makes the translation results deviate from the original meaning and reduces the accuracy of the translation results.

[0030] For example, the text to be translated is "mouse", and the context of the text to be translated is "She set the mouse on the table and moved the cursor to the icon". According to the context of the text to be translated, we can know that "mouse" refers to the mouse, but if translated out of context, "mouse" can mean either mouse or rat. If "mouse" is used as the translation result, it does not conform to the original meaning.

[0031] To address this issue, the present invention provides a translation method designed to improve the accuracy of translation results by integrating contextual information of the text to be translated. This method can be performed by a learning device (hereinafter referred to as the "device"), such as a dictionary pen or electronic dictionary. To facilitate understanding of the present invention's technical solution, the following embodiments utilize a dictionary pen as an example.

[0032] in, Figure 1 This is one of the flow charts of the translation method provided by the present invention, such as Figure 1 As shown, the method includes step 110 , step 120 , step 130 and step 140 .

[0033] Step 110 , obtaining a first image and a second image, and performing text recognition on each of them to obtain page text and target text; the first image is an image of the page where the target text is located, and the second image is an image of the target text.

[0034] Specifically, the target text is the text to be translated. The target text can be a single word, a phrase, a sentence, or a text containing multiple sentences. This is not specifically limited in the embodiments of the present invention. The target text can be obtained by performing text recognition on the second image. Here, text recognition can be achieved through optical character recognition (OCR) technology, template matching, sequence modeling technology, etc. Considering the cost and efficiency of implementation, in the embodiments of the present invention, OCR technology is preferably used to perform text recognition on the second image.

[0035] The page text is all textual content on the page where the target text resides. This includes the target text and other textual information. Page text can also be recognized using techniques such as optical character recognition (OCR), template matching, and sequence modeling. Considering implementation costs and efficiency, in this embodiment of the present invention, OCR is preferred for text recognition in the first image.

[0036] The first image refers to an image of the page where the target text is located. The first image can be an image of the entire page where the target text is located. For example, if the target text is a short phrase on line 5 of page 1, the first image can be an image of the entire page of page 1. The first image can also be an image of a portion of the page where the target text is located. For example, if the target text is a short phrase on line 5 of page 1, the first image can be an image of the text area corresponding to lines 2 to 6 of page 1. The first image can be acquired by a first image acquisition component, which can be a camera, an image sensor, or the like.

[0037] The second image refers to an image of the target text, which is used to represent the visual information of the target text. For example, if the target text is "Hello", the second image is an image containing "Hello". The first image and the second image can be acquired by the same image acquisition component or by different image acquisition components. The image acquisition component here can be a camera, image sensor, etc.

[0038] The first image acquisition element and the second image acquisition element can be the same image acquisition element or different image acquisition elements. When the two are the same image acquisition element, the image acquisition element can be deployed at the tip of the dictionary pen; when the two are different image acquisition elements, the first acquisition element can be deployed at the end away from the tip of the dictionary pen, and the second image acquisition element can be deployed at the tip of the dictionary pen; when the two are different image acquisition elements, the first acquisition element and the second acquisition element can both be deployed near the tip of the dictionary pen. In addition, when the first image acquisition element and the second image acquisition element are different acquisition elements, the activation order of the first acquisition element and the second acquisition element may be different. For example, the second acquisition element can be activated when the tip of the dictionary pen touches the page after the first acquisition element is activated.

[0039] Taking the camera deployed at the tip of a dictionary pen as the first image acquisition element, the triggering condition for the first image acquisition can be any of the following: ① detecting that the vertical distance between the camera and the target page containing the target text is greater than 0 and less than or equal to a first threshold; ② detecting that the pen body of the dictionary pen changes from a horizontal orientation to an inclined state (the inclined state here means that the angle between the pen body and the horizontal direction is greater than 0°); ③ detecting a capture instruction sent by the user. This capture instruction can be generated by the user pressing the first button on the dictionary pen, or by detecting a specific user gesture (such as an "OK" gesture). The triggering conditions for the first image acquisition are not limited to the above examples, and the embodiments of the present invention do not specifically limit the triggering conditions for the first image acquisition.

[0040] In addition, when the triggering condition for capturing the first image is met, one image may be captured as the first image, or multiple images may be captured and an image with the highest definition may be selected from the multiple images as the first image.

[0041] The triggering conditions for capturing the second image can be any of the following: ① detecting contact between the tip of the dictionary pen and the page containing the target text; ② detecting a capture instruction sent by the user, which can be generated by the user pressing the second button on the dictionary pen. The triggering conditions for capturing the second image are not limited to the above examples, and the embodiments of the present invention do not specifically limit the triggering conditions for capturing the second image. The first button and the second button can be the same button or different buttons, and the embodiments of the present invention do not specifically limit this.

[0042] Step 120: Determine the position of the target text in the page text.

[0043] Specifically, the position of the target text in the page refers to the specific position of the target text in the page text, and the position may include information such as the paragraph and line number of the target text.

[0044] Optionally, text matching can be performed between the target text and the page text, and text matching the target text can be determined from the page text, and the position of the matching text in the page text can be used as the position of the target text in the page. When performing text matching, an exact matching algorithm, a fuzzy matching algorithm, a string comparison algorithm, etc. can be used to perform text matching between the target text and the page text.

[0045] For example, when using a string matching algorithm, if the target text is a complete sentence, phrase, or word, the target text can be compared with the page text character by character, and the text closest to the target text can be determined from the page text. This text is the matching text.

[0046] After identifying the matching text in the page text, the position of the matching text (including paragraph, line number, character offset, and other information) extracted by OCR can be used to calibrate the position of the target text in the page text. For example, if the target text is located on line 5 and between columns 2 and 5, the position of the target text in the page text can be identified as "line 5, start position: column 2, end position: column 5".

[0047] The position accuracy of the target text in the page text can be specific to the paragraph level, the line level, or the character level. The position accuracy can be determined according to user needs and is not specifically limited in this embodiment of the present invention.

[0048] Step 130: Extract the context text of the target text from the page text based on the position of the target text in the page text.

[0049] Specifically, the context text of the target text refers to other texts that are closely related to the target text. The close relevance here can be reflected in the close semantic relevance between the context text and the target text. Among them, the context text can be the paragraph text corresponding to the paragraph where the target text is located, or it can be a candidate text (here, the candidate text refers to the text in the page text that is semantically similar to the keyword of the paragraph text, that is, the subject information expressed by the candidate text and the paragraph text is the same or similar), or it can be a combination of the paragraph text and the candidate text. Among them, the combination text can be obtained by splicing the paragraph text and the candidate text according to the order of the paragraph text and the candidate text in the page text. When splicing, the paragraph text and the candidate text can be connected by special symbols (such as "&"), or by preset participles (such as "and").

[0050] For example, the target text is "mouse", the paragraph text is "The mouse is a common input device used to interact with computers", and the candidate text is "A computer mouse allows users to move a pointer on the screen and select objects". The context text can be the paragraph text "Themouse is a common input device used to interact with computers", the candidate text "A computer mouse allows users to move a pointer on the screen and selectobjects", or the combined text "The mouse is a common input device used to interactwith computers, and A computer mouse allows users to move a pointer on thescreen and select objects".

[0051] Step 140: Determine a translation result of the target text based on the target text and the context text.

[0052] Specifically, given the close semantic relationship between the context text and the target text, it can provide background information about the target text and help understand the context of the target text. Therefore, embodiments of the present invention combine the target text and the context text to determine the translation result of the target text. This can determine the translation result even in the presence of polysemous words, ensuring that the translation result is consistent with the original meaning of the target text, thereby improving the accuracy of the translation result.

[0053] For example, when the target text is "mouse", "mouse" is a polysemous word and can be translated as "mouse" or "mouse". If combined with the context of the target text "She set the mouse on the table and moved the cursor to the icon", it can be determined that the corresponding translation result of "mouse" is "mouse".

[0054] Optionally, after obtaining the context text, the target text and the context text can be combined into a paragraph, and the paragraph can be translated to obtain a translation result of the paragraph. Based on the position of the target text in the paragraph, the translation result of the target text is extracted from the translation result of the paragraph. Since the paragraph contains the context information of the target text, the translation result corresponding to the paragraph fits the actual context of the target text, and the translation result of the target text is extracted from the translation result of the paragraph, that is, the translation result of the target text comes from the original translation result of the paragraph, and the extracted translation result of the target text can more accurately convey the original meaning content.

[0055] The translation method provided by the embodiment of the present invention is based on the position of the target text in the page text, so that the context text of the target text can be extracted from the page text in combination with the position. Since the context text is closely related to the semantics of the target text, the target text and the context text are combined for translation, which can more accurately understand and convey the true meaning of the original text, ensure that the translation result is closer to the original meaning, and improve the accuracy of the translation result.

[0056] Based on the above embodiments, Figure 2 This is the second flow chart of the translation method provided by the present invention, such as Figure 2 As shown, the method includes: Step 210 , obtaining a first image and a second image, and performing text recognition on each of them to obtain page text and target text; the first image is an image of the page where the target text is located, and the second image is an image of the target text.

[0057] Specifically, the first image is acquired by the first image acquisition element, and the second image is acquired by the second image acquisition element. The first image acquisition element and the second image acquisition element may be the same image acquisition element or different image acquisition elements. When the two are the same image acquisition element, the image acquisition element may be deployed at the tip of the dictionary pen; when the two are different image acquisition elements, the first acquisition element may be deployed at the end away from the tip of the dictionary pen, and the second image acquisition element may be deployed at the tip of the dictionary pen; when the two are different image acquisition elements, the first acquisition element and the second acquisition element may both be deployed near the tip of the dictionary pen. In addition, when the first image acquisition element and the second image acquisition element are different acquisition elements, the activation order of the first acquisition element and the second acquisition element may be different. For example, the second acquisition element may be activated when the tip of the dictionary pen touches the page after the first acquisition element is activated.

[0058] Step 220: Determine the position of the target text in the page text.

[0059] Optionally, text matching may be performed between the target text and the page text, and text matching the target text may be determined from the page text, and the position of the matching text in the page text may be used as the position of the target text in the page.

[0060] Step 230: Based on the position, determine the paragraph text from the page text, where the paragraph text is the text corresponding to the paragraph where the target text is located; based on the keywords in the paragraph text, extract multiple candidate texts from the page text, where the semantic relevance between the candidate texts and the keywords is greater than a first threshold; and determine the context text based on the paragraph text and the multiple candidate texts.

[0061] Specifically, the paragraph text is the text corresponding to the paragraph containing the target text, which can be understood as the complete content of the paragraph containing the target text. Since the position of the target text in the page is the specific location of the target text in the page text, such as the paragraph and line number of the target text, the position of the paragraph containing the target text in the page text can be determined based on the position of the target text in the page text, and the corresponding paragraph text can be extracted from the page text based on this position.

[0062] For example, if the target text is located in the 4th line of the 3rd paragraph, then the location of the paragraph containing the target text in the page text can be determined to be the 3rd paragraph of the 2nd page, and the paragraph text can be extracted based on the location of the paragraph in the page text.

[0063] Furthermore, keywords in a paragraph are words or phrases that represent the core content of the paragraph and are used to represent the main information of the paragraph. Multiple candidate texts refer to text in the page text whose semantic relevance to the keywords is greater than a first threshold. These candidate texts can be a single word, a phrase, or a paragraph. The first threshold can be set based on actual needs and is not specifically limited in this embodiment of the present invention.

[0064] The semantic relevance here can be understood as the strength of the relationship between each candidate text and the keyword in the same context. The stronger the relationship, the more background information the text provides, which can help understand the original meaning of the target text.

[0065] Optionally, a word frequency statistics can be performed on each word in the paragraph text to determine the word frequency of each word (that is, the number of times each word appears in the paragraph text). When the word frequency is greater than a threshold, it indicates that the corresponding word is closely related to the subject content of the paragraph text. In this case, the word can be used as a keyword for the paragraph text.

[0066] Because the paragraph text corresponds to the target text's paragraph, it provides detailed context for the target text, making it closely semantically related to the target text. Furthermore, multiple candidate texts are extracted from the page text based on the paragraph text's keywords. This effectively expands the target text's context, making it closely semantically related to the target text.

[0067] Furthermore, the paragraph text and the multiple candidate texts may be used as the context text of the target text, or the context text of the target text may be obtained by screening the paragraph text and the multiple candidate texts based on semantic relevance.

[0068] Step 240: Determine a translation result of the target text based on the target text and the context text.

[0069] Optionally, after obtaining the context text, the target text and the context text can be combined into a paragraph, and the paragraph can be translated to obtain a translation result of the paragraph. Based on the position of the target text in the paragraph, the translation result of the target text is extracted from the translation result of the paragraph. Since the paragraph contains the context information of the target text, the translation result corresponding to the paragraph fits the actual context of the target text, and the translation result of the target text is extracted from the translation result of the paragraph, that is, the translation result of the target text comes from the original translation result of the paragraph, and the extracted translation result of the target text can more accurately convey the original meaning content.

[0070] Based on any of the above embodiments, Figure 3 This is the third flow chart of the translation method provided by the present invention, as shown in FIG. Figure 3As shown, the method includes: Step 310: Acquire a first image and a second image, and perform text recognition on each image to obtain page text and target text; the first image is an image of the page where the target text is located, and the second image is an image of the target text.

[0071] Specifically, the first image is acquired by a first image acquisition element, and the second image is acquired by a second image acquisition element. The first image acquisition element and the second image acquisition element may be the same image acquisition element or different image acquisition elements. The image acquisition element may be a camera, an image sensor, or the like. In the embodiment of the present invention, the image acquisition element is preferably a camera.

[0072] Step 320: Determine the position of the target text in the page text.

[0073] Optionally, text matching may be performed between the target text and the page text, and text matching the target text may be determined from the page text, and the position of the matching text in the page text may be used as the position of the target text in the page.

[0074] Step 330: Based on the position, determine the paragraph text from the page text, where the paragraph text refers to the text corresponding to the paragraph where the target text is located; extract multiple candidate texts from the page text based on the keywords of the paragraph text; when the reference text includes figure number information, obtain the target image from the first image, where the target image is the illustration image corresponding to the figure number information, and the reference text includes the paragraph text and / or candidate text; perform semantic understanding on the target image to obtain the semantic text of the target image; determine the context text from the reference text and the semantic text.

[0075] Specifically, the figure number information is used to indicate a specific figure on a page, such as " Figure 2 As shown in the figure, Figure 2 " is the figure number information. When there is figure information in the reference text, it means that there may be a figure corresponding to the figure number information on the page, that is, there may be a figure image corresponding to the figure number information (i.e., the target image) in the first image. Since the target image exists on the page, the visual content expressed by the target image may be related to the theme, background, and other information of the page text, and this information can help understand the actual context of the target text.

[0076] To address this issue, when the reference text contains an image number, the embodiment of the present invention retrieves the image corresponding to the image number from the first image as the target image. The target image is then semantically interpreted to generate semantic text that describes the visual content of the target image. For example, if the target image is a picture of a mouse, the corresponding semantic text might be "This is a little mouse."

[0077] When acquiring a target image, an image segmentation algorithm can be used to segment the first image to obtain the target image. For example, a convolutional neural network (CNN) can be used to segment the first image to obtain the target image. After obtaining the target image, image description generation technology can be used to perform semantic understanding of the target image to generate semantic text. This image description generation technology can include deep learning-based image description models, such as the Show and Tell model or the Transformers model.

[0078] Considering that the semantic text may have a semantic association with the target text, both the reference text and the semantic text can be used as context text, so that the actual context of the target text can be understood based on the context text. Alternatively, based on a correlation analysis of the reference text and the semantic text, the context text can be screened from the reference text and the semantic text. For example, the semantic correlation between the semantic text and the reference text can be calculated. If the semantic correlation is greater than a threshold, it indicates that the target image is highly correlated with the reference text, that is, the subject information expressed by the target image and the reference text may be the same or similar. In this case, the semantic text and the reference text can be used as context text. If the semantic correlation is less than or equal to the threshold, it indicates that the target image is less correlated with the reference text, that is, the subject information expressed by the target image may be different from that of the reference text, that is, the target image is not helpful in understanding the actual context of the target text. In this case, the reference text can be used as context text.

[0079] For example, the target text is "mouse" (which can be understood as a mouse or a mouse), and the reference text is " Figure 2 As shown, a small animal is shown here", " Figure 2 "The corresponding attached image is "a picture of a mouse", and then the attached image is semantically understood, and the semantic text obtained is "This is a mouse". According to the semantic text, it can be known that the actual semantics of the target text "mouse" should be "mouse".

[0080] Step 340: Determine a translation result of the target text based on the target text and the context text.

[0081] Optionally, after obtaining the context text, the target text and the context text can be combined into a paragraph, and the paragraph can be translated to obtain a translation result of the paragraph. Based on the position of the target text in the paragraph, the translation result of the target text is extracted from the translation result of the paragraph. Since the paragraph contains the context information of the target text, the translation result corresponding to the paragraph fits the actual context of the target text, and the translation result of the target text is extracted from the translation result of the paragraph, that is, the translation result of the target text comes from the original translation result of the paragraph, and the extracted translation result of the target text can more accurately convey the original meaning content.

[0082] Based on any of the above embodiments, Figure 4 This is the fourth flow chart of the translation method provided by the present invention, such as Figure 4 As shown, the method includes: Step 410: Acquire a first image and a second image, and perform text recognition on each image to obtain page text and target text; the first image is an image of the page where the target text is located, and the second image is an image of the target text.

[0083] Specifically, the first image is acquired by a first image acquisition component, and the second image is acquired by a second image acquisition component. The first image acquisition component and the second image acquisition component can be the same image acquisition component or different image acquisition components.

[0084] Step 420: Determine the position of the target text in the page text.

[0085] Optionally, text matching may be performed between the target text and the page text, and text matching the target text may be determined from the page text, and the position of the matching text in the page text may be used as the position of the target text in the page.

[0086] Step 430: Based on the position, determine the paragraph text from the page text, where the paragraph text refers to the text corresponding to the paragraph where the target text is located; extract multiple candidate texts from the page text based on the keywords of the paragraph text; if the reference text includes figure number information, extract the description text of the figure number information from the reference text, extract all the illustration images from the first image, and determine the target image from each illustration image based on the semantic information of the description text and the semantic information of each illustration image; perform semantic understanding on the target image to obtain the semantic text of the target image; and determine the context text from the reference text and the semantic text.

[0087] Specifically, the description text of the figure number information refers to the specific explanation or description of the image or illustration content, which usually appears immediately after the figure number information. For example, the reference text is " Figure 2 As shown, it shows that... ", then " Figure 2" is the figure number information, and "shows..." is the corresponding description text, which is used to describe Figure 2 The specific content.

[0088] In addition, considering that there may be multiple illustration images in the first image, the illustration image corresponding to the figure number information in the reference text (i.e., the target image) may be one of the multiple illustration images, and the semantic information of the target image is the same or similar to the semantic information of the description text.

[0089] To this end, embodiments of the present invention determine a target image from each of the accompanying images based on the semantic information of the descriptive text and the semantic information of each of the accompanying images. Optionally, the semantic information can be represented by the semantic features of the text. For example, the semantic features of the descriptive text and the semantic features of each of the accompanying images can be extracted, and the distance between the semantic features of the descriptive text and the semantic features of each of the accompanying images can be calculated. The accompanying image with the smallest distance is selected as the candidate image.

[0090] Furthermore, considering that the target image may not be located in the first image (for example, the target image may be located on the page next to or previous to the page containing the target text), if the minimum distance is greater than the threshold, it indicates that the target image is most likely not in the first image. In this case, the target image does not exist on the page containing the target text, and the context of the target text can be determined from the reference text. If the minimum distance is less than or equal to the threshold, it indicates that the target image exists in the first image, and the candidate image is selected as the target image.

[0091] Among them, when extracting the descriptive text of the figure number information from the reference text, the figure number information can be used as the starting position, and specific punctuation marks (such as a period, exclamation mark, etc.) can be searched backward, and the position of the specific punctuation mark can be used as the ending position, and the text between the starting position and the ending position can be used as the descriptive text.

[0092] Furthermore, when extracting all accompanying images from the first image, the first image may be segmented to extract all accompanying images from the first image. Image segmentation methods may include semantic segmentation, threshold segmentation, edge detection, and other techniques. Considering the distribution characteristics of the accompanying images and the complexity of the first image, a semantic segmentation method based on deep learning is preferably used for image segmentation.

[0093] Step 440: Determine a translation result of the target text based on the target text and the context text.

[0094] Optionally, after obtaining the context text, the target text and the context text can be combined into a paragraph, and the paragraph can be translated to obtain a translation result of the paragraph. Based on the position of the target text in the paragraph, the translation result of the target text is extracted from the translation result of the paragraph. Since the paragraph contains the context information of the target text, the translation result corresponding to the paragraph fits the actual context of the target text, and the translation result of the target text is extracted from the translation result of the paragraph, that is, the translation result of the target text comes from the original translation result of the paragraph, and the extracted translation result of the target text can more accurately convey the original meaning content.

[0095] Based on any of the above embodiments, Figure 5 This is the fifth flow chart of the translation method provided by the present invention, such as Figure 5 As shown, the method includes: Step 510: Acquire a first image and a second image, and perform text recognition on each image to obtain page text and target text; the first image is an image of the page where the target text is located, and the second image is an image of the target text.

[0096] Specifically, the first image is acquired by a first image acquisition component, and the second image is acquired by a second image acquisition component. The first image acquisition component and the second image acquisition component can be the same image acquisition component or different image acquisition components.

[0097] Step 520: Determine the position of the target text in the page text.

[0098] Optionally, text matching may be performed between the target text and the page text, and text matching the target text may be determined from the page text, and the position of the matching text in the page text may be used as the position of the target text in the page.

[0099] Step 530: Based on the position, determine the paragraph text from the page text, where the paragraph text refers to the text corresponding to the paragraph where the target text is located; based on the keywords of the paragraph text, extract multiple candidate texts from the page text; based on the semantic relevance between the paragraph text and each candidate text, determine the relevant text from each candidate text, where the semantic relevance between the relevant text and the paragraph text is greater than a second threshold; and use the paragraph text and the relevant text as context text.

[0100] Specifically, paragraph text refers to the text corresponding to the paragraph containing the target text, which can be understood as the complete content of the paragraph containing the target text. Keywords in paragraph text refer to words or phrases that can represent the core content of the paragraph and are used to represent the main information of the paragraph. Multiple candidate texts refer to other texts in the page text that are semantically similar or related to the keywords.

[0101] Considering that paragraph text is the text of the paragraph containing the target text, it directly contains detailed information about the target text in terms of structure, context, and linguistic context. Therefore, paragraph text is necessarily highly semantically related to the target text. However, multiple candidate texts are determined from the page text based on the keywords in the paragraph text. Due to the specific location of the text, contextual differences, or different semantic scopes, some candidate texts may have low semantic relevance to the target text.

[0102] For example, the target text is "image recognition" and the paragraph text is "Deep learning is a branch of machine learning that has been widely used in image recognition and has made significant progress." The keywords identified in the paragraph text are "deep learning, image recognition, application." Based on these keywords, candidate texts extracted from the page text include: Candidate 1: "The application of deep learning in the field of medical imaging has achieved breakthrough progress" and Candidate 2: "The development of image recognition technology has driven innovation in the field of computer vision." After semantic analysis, Candidate 1 is closely related to the target text, but Candidate 2 focuses on technological development rather than application scenarios, making it less semantically relevant to the target text.

[0103] Based on this, the embodiment of the present invention uses the semantics of the paragraph text as a screening criterion, that is, determines the semantic relevance between each candidate text and the paragraph text. The semantic relevance is used to characterize the strength of the relationship between the candidate text and the paragraph text in the same context. The greater the strength of the relationship, the more background information the candidate text provides, which can assist in understanding the original meaning of the target text. Optionally, the candidate text and the paragraph text with a semantic relevance greater than a second threshold can be used as context text. The second threshold here can be set according to actual needs. The second threshold can be the same as the first threshold above, or it can be different. The embodiment of the present invention does not make specific limitations on this.

[0104] The semantic features of each candidate text and the semantic features of the paragraph text may be extracted, and the semantic relevance between each candidate text and the paragraph text may be measured using the Euclidean distance or cosine similarity between the two features.

[0105] In addition, it should be noted that the embodiment of the present invention first determines multiple candidate texts from the page text based on the keywords of the paragraph text, that is, through preliminary screening, multiple candidate texts that may be closely related to the semantics of the target text are determined, so that within a small range (that is, within the range of multiple candidate texts), by calculating the semantic relevance between the paragraph text and each candidate text, relevant texts can be determined from multiple candidate texts, effectively avoiding the calculation of the semantic similarity between the paragraph text and other irrelevant texts in the page, thereby reducing unnecessary computing overhead.

[0106] Step 540: Determine a translation result of the target text based on the target text and the context text.

[0107] Optionally, after obtaining the context text, the target text and the context text can be combined into a paragraph, and the paragraph can be translated to obtain a translation result of the paragraph. Based on the position of the target text in the paragraph, the translation result of the target text is extracted from the translation result of the paragraph. Since the paragraph contains the context information of the target text, the translation result corresponding to the paragraph fits the actual context of the target text, and the translation result of the target text is extracted from the translation result of the paragraph, that is, the translation result of the target text comes from the original translation result of the paragraph, and the extracted translation result of the target text can more accurately convey the original meaning content.

[0108] Based on any of the above embodiments, Figure 6 This is the sixth flow chart of the translation method provided by the present invention, such as Figure 6 As shown, the method includes: Step 610: Acquire a first image and a second image, and perform text recognition on each image to obtain page text and target text; the first image is an image of the page where the target text is located, and the second image is an image of the target text.

[0109] Specifically, the first image is acquired by a first image acquisition component, and the second image is acquired by a second image acquisition component. The first image acquisition component and the second image acquisition component can be the same image acquisition component or different image acquisition components.

[0110] Step 620: Perform text matching on the target text and the page text, determine the matching text from the page text, where the matching text is the text in the page text that is identical to the target text; and use the position of the matching text in the page text as the position of the target text in the page text.

[0111] Specifically, matching text refers to text in the page text that is identical to the target text. For example, if the page text is "…The mouse is a common input device used to interact with computers…" and the target text is "mouse", then the matching text is "mouse" in "…The mouse is a common input device used to interact with computers…", meaning the matching text is identical to the target text. Matching text can be determined by performing text matching between the target text and the page text. Text matching refers to finding the portion of the page text that is closest to the target text based on a certain similarity or matching algorithm. Methods such as exact matching, fuzzy matching, and semantic matching can be used to match the target text and the page text.

[0112] The position of the matching text in the page text can be the specific starting and ending position of the matching text in the page text. Since the content of the matching text is completely consistent with that of the target text, the position of the matching text in the page text can be used as the position of the target text in the page text.

[0113] Among them, the position of the matching text in the page text can be determined based on the following steps: locate the starting position of the matching text in the page text through a search algorithm, then determine the ending position of the matching text, and use the starting position and the ending position as the position of the matching text in the page text.

[0114] Step 630: Extract the context text of the target text from the page text based on the position of the target text in the page text.

[0115] Specifically, the context of a target text refers to other text that is closely related to the target text. This close relevance can be reflected in the close semantic relationship between the context and the target text. The context can be the paragraph text corresponding to the paragraph containing the target text, or it can be a candidate text (a candidate text is a text within the page that has semantically similar keywords to the paragraph text, meaning that the candidate text and the paragraph text convey the same or similar subject matter). It can also be a combination of a paragraph text and a candidate text.

[0116] Step 640: Determine a translation result of the target text based on the target text and the context text.

[0117] Optionally, after obtaining the context text, the target text and the context text can be combined into a paragraph, and the paragraph can be translated to obtain a translation result of the paragraph. Based on the position of the target text in the paragraph, the translation result of the target text is extracted from the translation result of the paragraph. Since the paragraph contains the context information of the target text, the translation result corresponding to the paragraph fits the actual context of the target text, and the translation result of the target text is extracted from the translation result of the paragraph, that is, the translation result of the target text comes from the original translation result of the paragraph, and the extracted translation result of the target text can more accurately convey the original meaning content.

[0118] Based on any of the above embodiments, Figure 7 This is the seventh flow chart of the translation method provided by the present invention, such as Figure 7 As shown, the method includes: Step 710: Acquire a first image and a second image, and perform text recognition on each image to obtain page text and target text; the first image is an image of the page where the target text is located, and the second image is an image of the target text.

[0119] Specifically, the first image is acquired by a first image acquisition element, and the second image is acquired by a second image acquisition element. The first image acquisition element and the second image acquisition element may be the same image acquisition element or different image acquisition elements. The image acquisition element may be a camera, an image sensor, or the like. In an embodiment of the present invention, the image acquisition element is preferably a camera, and the image acquisition element is disposed on the camera at the tip of the dictionary pen.

[0120] After detecting that the user has picked up the dictionary pen, the camera at the pen tip begins capturing a first image. This process can be continuous, meaning that as the user brings the pen tip closer to the paper, multiple images are captured, with the best quality image being used as the first image. During this process, the image capture range decreases as the distance between the pen tip and the paper decreases, while the clarity increases.

[0121] In addition, the image quality may be measured by definition. For example, the definition of each image may be determined by using the gradient magnitude, variance, or Laplace operator of each image, and the image with the highest definition may be used as the first image.

[0122] When the pen tip is pressed, the camera on the dictionary pen tip begins capturing a second image and stops capturing the second image when the pen tip is no longer pressed. Typically, the user presses the pen tip and slides it across the paper, with the sliding area covering the target text area. The resulting second image is the image of the target text.

[0123] Step 720: Determine the position of the target text in the page text.

[0124] Optionally, text matching may be performed between the target text and the page text, and text matching the target text may be determined from the page text, and the position of the matching text in the page text may be used as the position of the target text in the page.

[0125] Step 730: Extract the context text of the target text from the page text based on the position of the target text in the page text.

[0126] Specifically, the context of a target text refers to other text that is closely related to the target text. This close relevance can be reflected in the close semantic relationship between the context and the target text. The context can be the paragraph text corresponding to the paragraph containing the target text, or it can be a candidate text (a candidate text is a text within the page that has semantically similar keywords to the paragraph text, meaning that the candidate text and the paragraph text convey the same or similar subject matter). It can also be a combination of a paragraph text and a candidate text.

[0127] Step 740: Merge the target text and the context text to obtain a synthesized text, and determine the position of the target text in the synthesized text; translate the synthesized text to obtain a translated text; and extract a translation result of the target text from the translated text based on the position of the target text in the synthesized text.

[0128] Specifically, the synthesized text refers to a text obtained by combining a target text and a context text. For example, the target text and the context text may be combined according to their order in the page text to obtain the synthesized text.

[0129] After obtaining the synthesized text, the synthesized text is translated to obtain a translated text. Since the synthesized text contains the context of the target text, the translated text is generated after considering the context information, so that the translated text can be consistent with the original text content.

[0130] Based on the target text's position in the synthesized text, the target text's translation is extracted from the translated text. This means the target text's translation is derived from the original text. Because the translated text adheres to the original text's semantics, the translation derived from the original text also adheres to the original text's semantics. This means the target text's translation accurately reflects the original text's semantics, resulting in a high degree of translation accuracy.

[0131] When combining the target text and the context text, the positions of the target text and the context text in the synthesized text can be marked. After obtaining the synthesized text, the position of the target text in the synthesized text can be directly determined based on the position identifier.

[0132] The translation device provided by the present invention is described below. The translation device described below and the translation method described above can be referenced to each other.

[0133] Based on any of the above embodiments, Figure 8 Schematic diagram of the structure of the translation device provided by the present invention. Figure 8 As shown, the device includes: The acquisition unit 810 is configured to acquire a first image and a second image, and perform text recognition on each image to obtain page text and target text; the first image is an image of the page where the target text is located, and the second image is an image of the target text; A determination unit 820 is configured to determine a position of the target text in the page text; An extraction unit 830 is configured to extract context text of the target text from the page text based on a position of the target text in the page text; The translation unit 840 is configured to determine a translation result of the target text based on the target text and the context text.

[0134] Based on any of the above embodiments, extracting context text of the target text from the page text based on the position of the target text in the page text includes: Based on the position, determine the paragraph text from the page text, where the paragraph text is the text corresponding to the paragraph where the target text is located; Based on the keywords of the paragraph text, a plurality of candidate texts are extracted from the page text, wherein the semantic relevance between the candidate texts and the keywords is greater than a first threshold; A context text is determined based on the paragraph text and the plurality of candidate texts.

[0135] Based on any of the above embodiments, determining context text from the paragraph text and the plurality of candidate texts includes: In the case where the reference text includes drawing number information, obtaining a target image from the first image, where the target image is an image of the drawing corresponding to the drawing number information, and the reference text includes paragraph text and / or candidate text; Perform semantic understanding on the target image and obtain the semantic text of the target image; Determine contextual text from reference text and semantic text.

[0136] Based on any of the foregoing embodiments, obtaining a target image from a first image includes: Extract the descriptive text of the figure number information from the reference text; Extract all accompanying images from the first image; Based on the semantic information of the description text and the semantic information of each of the attached images, a target image is determined from each of the attached images.

[0137] Based on any of the above embodiments, determining context text from the paragraph text and the plurality of candidate texts includes: Determining a relevant text from each candidate text based on the semantic relevance between the paragraph text and each candidate text; the semantic relevance between the relevant text and the paragraph text is greater than a second threshold; Use paragraph text and related text as contextual text.

[0138] Based on any of the above embodiments, determining the position of the target text in the page text includes: Performing text matching on the target text and the page text, and determining matching text from the page text, where the matching text is the text in the page text that is identical to the target text; The position of the matching text in the page text is used as the position of the target text in the page text.

[0139] Based on any of the above embodiments, determining a translation result of the target text based on the target text and the context text includes: Merge the target text and the context text to obtain a synthesized text, and determine the position of the target text in the synthesized text; Translating the synthesized text to obtain a translated text; The translation result of the target text is extracted from the translated text based on the position of the target text in the synthesized text.

[0140] Figure 9 Schematic diagram of the structure of the electronic device provided by the present invention, such as Figure 9 As shown, the electronic device may include: a processor 910, a communications interface 920, a memory 930, and a communications bus 940, wherein the processor 910, the communications interface 920, and the memory 930 communicate with each other via the communications bus 940. The processor 910 may invoke logic instructions in the memory 930 to execute a translation method, which includes: acquiring a first image and a second image, and performing text recognition on each image to obtain page text and target text; the first image is an image of the page where the target text is located, and the second image is an image of the target text; determining the position of the target text within the page text; extracting context text of the target text from the page text based on the position of the target text within the page text; and determining a translation result of the target text based on the target text and the context text.

[0141] Furthermore, the logic instructions in the aforementioned memory 930 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the portion that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product, stored in a storage medium, includes instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to perform all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as a USB flash drive, a mobile hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.

[0142] On the other hand, the present invention also provides a computer program product, which includes a computer program, which can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the translation method provided by the above methods, which includes: obtaining a first image and a second image, and performing text recognition to obtain page text and target text respectively; the first image is an image of the page where the target text is located, and the second image is an image of the target text; determining the position of the target text in the page text; extracting the context text of the target text from the page text based on the position of the target text in the page text; and determining the translation result of the target text based on the target text and the context text.

[0143] On the other hand, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, is implemented to execute the translation method provided by the above-mentioned methods, the method comprising: acquiring a first image and a second image, and performing text recognition respectively to obtain page text and target text; the first image is an image of the page where the target text is located, and the second image is an image of the target text; determining the position of the target text in the page text; extracting the context text of the target text from the page text based on the position of the target text in the page text; and determining the translation result of the target text based on the target text and the context text.

[0144] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, i.e., they may be located in one location or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of the present embodiment. Persons of ordinary skill in the art will be able to understand and implement the present invention without inventive effort.

[0145] Through the above description of the embodiments, those skilled in the art will clearly understand that each embodiment can be implemented using software plus a necessary general-purpose hardware platform, or of course, hardware. Based on this understanding, the essence of the above technical solution, or the portion that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a magnetic disk, or an optical disk, and includes a number of instructions for causing a computer device (such as a personal computer, server, or network device) to execute the methods described in each embodiment or certain portions of the embodiments.

[0146] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.

Claims

1. A translation method, characterized in that: include: Acquire the first image and the second image, and perform text recognition on each image to obtain page text and target text; The first image is an image of the page where the target text is located, and the second image is an image of the target text; Determining the position of the target text in the page text; extracting context text of the target text from the page text based on the position of the target text in the page text; A translation result of the target text is determined based on the target text and the context text.

2. The translation method according to claim 1, wherein: The step of extracting the context text of the target text from the page text based on the position of the target text in the page text includes: Based on the position, determining paragraph text from the page text, the paragraph text being the text corresponding to the paragraph where the target text is located; extracting a plurality of candidate texts from the page text based on the keywords of the paragraph text, wherein the semantic relevance between the candidate texts and the keywords is greater than a first threshold; The context text is determined based on the paragraph text and the plurality of candidate texts.

3. The translation method according to claim 2, characterized in that The determining the context text based on the paragraph text and the plurality of candidate texts includes: In a case where the reference text includes drawing number information, acquiring a target image from the first image, the target image being an image of the drawing corresponding to the drawing number information, and the reference text including the paragraph text and / or the candidate text; Performing semantic understanding on the target image to obtain semantic text of the target image; The context text is determined from the reference text and the semantic text.

4. The translation method according to claim 3, characterized in that The acquiring a target image from the first image includes: Extracting the description text of the drawing number information from the reference text; extracting all accompanying images from the first image; The target image is determined from each of the accompanying images based on the semantic information of the description text and the semantic information of each of the accompanying images.

5. The translation method according to claim 2, characterized in that The determining the context text based on the paragraph text and the plurality of candidate texts includes: Determining a relevant text from each candidate text based on the semantic relevance between the paragraph text and each candidate text; the semantic relevance between the relevant text and the paragraph text is greater than a second threshold; The paragraph text and the related text are used as the context text.

6. The translation method according to any one of claims 1 to 5, characterized in that: Determining the position of the target text in the page text includes: Performing text matching on the target text and the page text, and determining matching text from the page text, wherein the matching text is text in the page text that is identical to the target text; The position of the matching text in the page text is used as the position of the target text in the page text.

7. The translation method according to any one of claims 1 to 5, characterized in that: The determining a translation result of the target text based on the target text and the context text includes: Merging the target text and the context text to obtain a composite text, and determining a position of the target text in the composite text; translating the synthesized text to obtain a translated text; A translation result of the target text is extracted from the translation text based on the position of the target text in the synthesized text.

8. A translation device, characterized in that: include: An acquisition unit, configured to acquire the first image and the second image, and perform text recognition on each image to obtain page text and target text; The first image is an image of the page where the target text is located, and the second image is an image of the target text; A determination unit, configured to determine a position of the target text in the page text; an extraction unit, configured to extract context text of the target text from the page text based on a position of the target text in the page text; The translation unit is configured to determine a translation result of the target text based on the target text and the context text.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that: When the processor executes the computer program, the translation method according to any one of claims 1 to 7 is implemented.

10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the translation method according to any one of claims 1 to 7 is implemented.

11. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the translation method according to any one of claims 1 to 7 is implemented.