Translation method and device
By obtaining the page feature information of the dictionary pen, the document type and layout position of the target text are determined, and the context text is extracted from the target literature, which solves the problem of inaccurate translation results of the dictionary pen and achieves more accurate translation results.
Patent Information
- Application Number
- CN202510502051.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-21
- Publication Date
- 2025-08-12
AI Technical Summary
The translation results of existing dictionary pens are relatively accurate because they are separated from the context of the text to be translated, resulting in the translation results that do not conform to the original meaning.
By obtaining the page feature information of the target page, determining the target literature to which the target text belongs, and extracting the context text from the target literature based on the layout position, combining the target text and the context text for translation, improving the accuracy of the translation results.
By combining context information, the translation results are closer to the original text meaning and improve the accuracy of the translation results.
Smart Images

Figure CN120471068A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of text translation, and in particular to a translation method and device. Background Art
[0002] As the demand for language learning grows, electronic aids such as dictionary pens and electronic dictionaries have gradually become common learning devices for users. Through these learning devices, users can directly query the translation results of the text to be translated.
[0003] Taking dictionary pens as an example, they typically use a camera to scan an image of the text to be translated, then perform text recognition on the image to obtain the text to be translated. This text is then translated to produce a translation result. However, this method only translates the text to be translated, resulting in low translation accuracy. Summary of the Invention
[0004] The present invention provides a translation method and device to solve the defects in the prior art.
[0005] The present invention provides a translation method, comprising the following steps: Acquire a first image and a second image, and perform image recognition on each image to obtain page feature information and target text of a target page; the first image is an image of the target page where the target text is located, and the second image is an image of the target text; Determining the target document to which the target text belongs based on the page feature information, and determining the typeset position of the target text in the target document; extracting context text of the target text from the target document based on the typeset position of the target text in the target document; A translation result of the target text is determined based on the target text and the context text.
[0006] According to a translation method provided by the present invention, extracting the context text of the target text from the target document based on the typeset position of the target text in the target document includes: Determining a paragraph text from the target document based on the typeset position, wherein the paragraph text refers to the text corresponding to the paragraph where the target text is located; extracting a plurality of candidate texts from the document based on keywords of the paragraph text; The context text is determined from the paragraph text and the plurality of candidate texts.
[0007] According to a translation method provided by the present invention, determining the context text from the paragraph text and the multiple candidate texts includes: Determining relevant texts from among the candidate texts based on the semantic relevance between the paragraph text and the candidate texts; The paragraph text and the related text are used as the context text.
[0008] According to a translation method provided by the present invention, the page feature information includes page text information; The determining the target document to which the target text belongs based on the page feature information includes: Extracting page keywords from the page text information; The target document to which the target text belongs is determined based on the page keywords.
[0009] According to a translation method provided by the present invention, the page feature information includes page text information and page layout information; The determining the target document to which the target text belongs based on the page feature information includes: Determining the document type of the target document based on the page layout information; Based on the document type, determining a target document library from a plurality of document libraries, wherein the target document library stores a plurality of documents, and each document has the same document type as the target document; Based on the page text information, the target document is determined from the target document library.
[0010] According to a translation method provided by the present invention, determining the typeset position of the target text in the target document includes: matching the first image with the second image to determine an image region in the first image that matches the second image; The layout position is determined based on the position of the image area in the first image.
[0011] According to a translation method provided by the present invention, determining a translation result of the target text based on the target text and the context text includes: Merging the target text and the context text to obtain a composite text, and determining a position of the target text in the composite text; translating the synthesized text to obtain a translated text; A translation result of the target text is extracted from the translation text based on the position of the target text in the synthesized text.
[0012] The present invention also provides a translation device, comprising the following modules: an acquisition unit, configured to acquire a first image and a second image, and perform image recognition on each image to obtain page feature information of a target page and target text; the first image is an image of the target page where the target text is located, and the second image is an image of the target text; a determining unit, configured to determine the target document to which the target text belongs based on the page feature information, and determine the typeset position of the target text in the target document; an extraction unit, configured to extract a context text of the target text from the target document based on a typeset position of the target text in the target document; The translation unit is configured to determine a translation result of the target text based on the target text and the context text.
[0013] The present invention also provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements any of the above-described translation methods when executing the program.
[0014] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which implements any of the above-mentioned translation methods when executed by a processor.
[0015] The present invention also provides a computer program product, comprising a computer program, wherein when the computer program is executed by a processor, the computer program implements any of the above-mentioned translation methods.
[0016] The translation method and device provided by the present invention determine the target document to which the target text belongs based on page feature information, and determine the typeset position of the target text in the target document, so as to extract the context text of the target text from the target document in combination with the typeset position. Since the context text is closely related to the semantics of the target text, the target text and the context text are combined for translation, which can more accurately understand and convey the true meaning of the original text, ensure that the translation result is closer to the original meaning, and improve the accuracy of the translation result. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] In order to more clearly illustrate the technical solutions in the present invention or the prior art, a brief introduction is given below to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0018] Figure 1 This is one of the flow charts of the translation method provided by the present invention.
[0019] Figure 2This is the second flowchart of the translation method provided by the present invention.
[0020] Figure 3 This is the third flow chart of the translation method provided by the present invention.
[0021] Figure 4 This is the fourth flow chart of the translation method provided by the present invention.
[0022] Figure 5 This is the fifth flow chart of the translation method provided by the present invention.
[0023] Figure 6 This is the sixth flowchart of the translation method provided by the present invention.
[0024] Figure 7 This is the seventh flow chart of the translation method provided by the present invention.
[0025] Figure 8 It is a structural diagram of the translation device provided by the present invention.
[0026] Figure 9 It is a structural schematic diagram of the electronic device provided by the present invention. DETAILED DESCRIPTION
[0027] To make the objectives, technical solutions, and advantages of the present invention more clear, the technical solutions of the present invention will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the embodiments described are only some of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the present invention.
[0028] Currently, most dictionary pens use a camera to scan an image of the text to be translated, then perform text recognition on the image to obtain the text to be translated. This text is then translated using a translation engine to produce a translation result. However, these translation results only translate the text to be translated, resulting in low accuracy.
[0029] After analysis, it was found that the reason why the translation results of traditional dictionary pens are less accurate is that they are translated out of the context of the text to be translated, which makes the translation results deviate from the original meaning and reduces the accuracy of the translation results.
[0030] For example, the text to be translated is "mouse", and the context of the text to be translated is "She set the mouse on the table and moved the cursor to the icon". According to the context of the text to be translated, we can know that "mouse" refers to the mouse, but if translated out of context, "mouse" can mean either mouse or rat. If "mouse" is used as the translation result, it does not conform to the original meaning.
[0031] To address this issue, the present invention provides a translation method designed to improve the accuracy of translation results by integrating contextual information of the text to be translated. This method can be performed by a learning device (hereinafter referred to as the "device"), such as a dictionary pen or electronic dictionary. To facilitate understanding of the present invention's technical solution, the following embodiments utilize a dictionary pen as an example.
[0032] in, Figure 1 This is one of the flow charts of the translation method provided by the present invention, such as Figure 1 As shown, the method includes step 110 , step 120 , step 130 and step 140 .
[0033] Step 110 , obtain a first image and a second image, and perform image recognition on each image to obtain page feature information and target text of the target page; the first image is an image corresponding to the target page where the target text is located, and the second image is an image of the target text.
[0034] Specifically, the target text is the text to be translated. The target text can be a single word, a phrase, a sentence, or a text containing multiple sentences. This is not specifically limited in the embodiments of the present invention. The target text can be obtained by performing text recognition on the second image. Here, text recognition can be achieved through optical character recognition (OCR) technology, template matching, sequence modeling technology, etc. Considering the cost and efficiency of implementation, in the embodiments of the present invention, OCR technology is preferably used to perform text recognition on the second image.
[0035] The target page is the page where the target text is located, that is, the target page includes the target text. For example, if the target text is "Hello", the target page may be a page including "Hello".
[0036] Page feature information refers to information contained in the target page, which is used to describe the target page's page content and layout. For example, page feature information may include page keywords, page layout information, page title information, header information, footer information, book title information, etc. Optionally, feature extraction can be performed on the first image to obtain image features, which are used to describe the target page's layout. Text recognition can be performed on the first image to obtain page text, and feature extraction can be performed on the page text to obtain text features, which are used to describe the target page's page content. Combining image features and text features can determine the target page's page feature information.
[0037] The first image refers to an image of the target page where the target text is located. The first image can be an image of the entire page where the target text is located. For example, if the target text is a short phrase on line 5 of page 1, the first image can be an image of the entire page of page 1. The first image can also be an image of a portion of the page where the target text is located. For example, if the target text is a short phrase on line 5 of page 1, the first image can be an image of the text area corresponding to lines 2 to 6 of page 1. The first image can be acquired by a first image acquisition component, which can be a camera, an image sensor, or the like.
[0038] The second image refers to an image of the target text, which is used to represent the visual information of the target text. For example, if the target text is "Hello", the second image is an image containing "Hello". The first image and the second image can be acquired by the same image acquisition component or by different image acquisition components. The image acquisition component here can be a camera, image sensor, etc.
[0039] The first image acquisition element and the second image acquisition element can be the same image acquisition element or different image acquisition elements. When the two are the same image acquisition element, the image acquisition element can be deployed at the tip of the dictionary pen; when the two are different image acquisition elements, the first acquisition element can be deployed at the end away from the tip of the dictionary pen, and the second image acquisition element can be deployed at the tip of the dictionary pen; when the two are different image acquisition elements, the first acquisition element and the second acquisition element can both be deployed near the tip of the dictionary pen. In addition, when the first image acquisition element and the second image acquisition element are different acquisition elements, the activation order of the first acquisition element and the second acquisition element may be different. For example, the second acquisition element can be activated when the tip of the dictionary pen touches the page after the first acquisition element is activated.
[0040] Taking the camera deployed at the tip of a dictionary pen as the first image acquisition element, the triggering condition for the first image acquisition can be any of the following: ① detecting that the vertical distance between the camera and the target page containing the target text is greater than 0 and less than or equal to a first threshold; ② detecting that the pen body of the dictionary pen changes from a horizontal orientation to an inclined state (the inclined state here means that the angle between the pen body and the horizontal direction is greater than 0°); ③ detecting a capture instruction sent by the user. This capture instruction can be generated by the user pressing the first button on the dictionary pen, or by detecting a specific user gesture (such as an "OK" gesture). The triggering conditions for the first image acquisition are not limited to the above examples, and the embodiments of the present invention do not specifically limit the triggering conditions for the first image acquisition.
[0041] In addition, when the triggering condition for capturing the first image is met, one image may be captured as the first image, or multiple images may be captured and an image with the highest definition may be selected from the multiple images as the first image.
[0042] The triggering conditions for capturing the second image can be any of the following: ① detecting contact between the tip of the dictionary pen and the page containing the target text; ② detecting a capture instruction sent by the user, which can be generated by the user pressing the second button on the dictionary pen. The triggering conditions for capturing the second image are not limited to the above examples, and the embodiments of the present invention do not specifically limit the triggering conditions for capturing the second image. The first button and the second button can be the same button or different buttons, and the embodiments of the present invention do not specifically limit this.
[0043] Step 120: Determine the target document to which the target text belongs based on the page feature information, and determine the layout position of the target text in the target document.
[0044] Specifically, the target document to which the target text belongs refers to the book, article, paper or other text material in printed or digital format that contains the target text, that is, the target document is used to identify the specific source of the target text.
[0045] Page feature information is used to describe the target page's content and layout. For example, page feature information may include page keywords, page layout information, page title information, header information, footer information, and book title information. Because each document contains different pages, the corresponding page content and / or page layout of each document also differ. This means that each page has different page feature information. Based on the target page's page feature information, the target document to which the target page belongs can be inferred.
[0046] For example, if the page feature information includes page keywords, since page keywords represent the core theme of a document and different documents have different core themes, the corresponding target document can be determined based on the page keywords. Alternatively, the corresponding keywords for different documents can be pre-determined, and the page keywords can be matched with the corresponding keywords for each document, with the matched document being used as the target document.
[0047] For another example, if the page feature information includes book title information, the target document can be directly determined based on the book title information, wherein the book title information can be located in the page title, the page header, or the page footer.
[0048] In addition, the typesetting position refers to the specific position of the target text in the target document. The typesetting position may include information such as the target page number where the target text is located, the paragraph where the target text is located, and the number of lines.
[0049] Optionally, the first image and the second image can be matched to determine an image region in the first image that matches the second image. This image region is the region in the first image where the target text is located, and the position of this image region in the first image is used as the layout position of the target text in the target page. Furthermore, the page feature information of the target page can include footer information, which typically includes the target page number. Based on this number, the layout position of the target page in the target document can be determined. Combined with the layout position of the target document in the target page, the layout position of the target text in the target document can be determined.
[0050] For example, if the target text is located at the third line of the second paragraph of the target page, and the target page is numbered page 4, then the target text's layout position in the target document is the third line of the second paragraph of page 4.
[0051] Step 130: Extract the context text of the target text from the target document based on the typeset position of the target text in the target document.
[0052] Specifically, the context text of the target text refers to other texts that are closely related to the target text. The close relevance here can be reflected in the close semantic relevance between the context text and the target text. Among them, the context text can be the paragraph text corresponding to the paragraph where the target text is located, or it can be a candidate text (here, the candidate text refers to the text in the target document that is semantically similar to the keyword of the paragraph text, that is, the subject information expressed by the candidate text and the paragraph text is the same or similar), or it can be a combination of the paragraph text and the candidate text. Among them, the combination text can be obtained by splicing the paragraph text and the candidate text according to the order of the paragraph text and the candidate text in the target document. When splicing, the paragraph text and the candidate text can be connected by special symbols (such as "&"), or by preset participles (such as "and").
[0053] For example, the target text is "mouse", the paragraph text is "The mouse is a common input device used to interact with computers", and the candidate text is "A computer mouse allows users to move a pointer on the screen and select objects". The context text can be the paragraph text "Themouse is a common input device used to interact with computers", the candidate text "A computer mouse allows users to move a pointer on the screen and selectobjects", or the combined text "The mouse is a common input device used to interactwith computers, and A computer mouse allows users to move a pointer on thescreen and select objects".
[0054] Optionally, after determining the typeset position of the target text in the target document, the paragraph text of the paragraph where the target text is located can be extracted from the target text based on the typeset position. Since the paragraph text usually contains information related to the target text, that is, the paragraph text is closely related to the semantics of the target text, the paragraph text can be used as context text.
[0055] Step 140: Determine a translation result of the target text based on the target text and the context text.
[0056] Specifically, given the close semantic relationship between the context text and the target text, it can provide background information about the target text and help understand the context of the target text. Therefore, embodiments of the present invention combine the target text and the context text to determine the translation result of the target text. This can determine the translation result even in the presence of polysemous words, ensuring that the translation result is consistent with the original meaning of the target text, thereby improving the accuracy of the translation result.
[0057] For example, when the target text is "mouse", "mouse" is a polysemous word and can be translated as "mouse" or "mouse". If combined with the context of the target text "She set the mouse on the table and moved the cursor to the icon", it can be determined that the corresponding translation result of "mouse" is "mouse".
[0058] Optionally, after obtaining the context text, the target text and the context text can be combined into a paragraph, and the paragraph can be translated to obtain a translation result of the paragraph. Based on the position of the target text in the paragraph, the translation result of the target text is extracted from the translation result of the paragraph. Since the paragraph contains the context information of the target text, the translation result corresponding to the paragraph fits the actual context of the target text, and the translation result of the target text is extracted from the translation result of the paragraph, that is, the translation result of the target text comes from the original translation result of the paragraph, and the extracted translation result of the target text can more accurately convey the original meaning content.
[0059] The translation method provided by the embodiment of the present invention determines the target document to which the target text belongs based on page feature information, and determines the typeset position of the target text in the target document, so that the context text of the target text can be extracted from the target document in combination with the typeset position. Since the context text is closely related to the semantics of the target text, the target text and the context text are combined for translation, which can more accurately understand and convey the true meaning of the original text, ensure that the translation result is closer to the original meaning, and improve the accuracy of the translation result.
[0060] Based on the above embodiments, Figure 2 This is the second flow chart of the translation method provided by the present invention, such as Figure 2 As shown, the method includes: Step 210: Acquire a first image and a second image, and perform image recognition on each to obtain page feature information and target text of the target page; the first image is an image corresponding to the target page where the target text is located, and the second image is an image of the target text.
[0061] Specifically, the first image is acquired by the first image acquisition element, and the second image is acquired by the second image acquisition element. The first image acquisition element and the second image acquisition element may be the same image acquisition element or different image acquisition elements. When the two are the same image acquisition element, the image acquisition element may be deployed at the tip of the dictionary pen; when the two are different image acquisition elements, the first acquisition element may be deployed at the end away from the tip of the dictionary pen, and the second image acquisition element may be deployed at the tip of the dictionary pen; when the two are different image acquisition elements, the first acquisition element and the second acquisition element may both be deployed near the tip of the dictionary pen. In addition, when the first image acquisition element and the second image acquisition element are different acquisition elements, the activation order of the first acquisition element and the second acquisition element may be different. For example, the second acquisition element may be activated when the tip of the dictionary pen touches the page after the first acquisition element is activated.
[0062] Page feature information refers to the information contained in the target page, which is used to describe the page content and page layout of the target page. For example, page feature information may include page keywords, page layout information, page title information, header information, footer information, book title information, etc.
[0063] Step 220: Determine the target document to which the target text belongs based on the page feature information, and determine the layout position of the target text in the target document.
[0064] Specifically, since the pages in each document are different, the page content and / or page layout corresponding to the pages in each document are also different, that is, the page feature information of each page is different, and then based on the page feature information of the target page, the target document to which the target page belongs can be reversely inferred.
[0065] In addition, the typesetting position refers to the specific position of the target text in the target document. The typesetting position may include information such as the target page number where the target text is located, the paragraph where the target text is located, and the number of lines.
[0066] Step 230: Based on the typesetting position, determine the paragraph text from the target document, where the paragraph text refers to the text corresponding to the paragraph where the target text is located; extract multiple candidate texts from the document based on the keywords of the paragraph text; and determine the context text from the paragraph text and the multiple candidate texts.
[0067] Specifically, paragraph text refers to the text corresponding to the target text's paragraph, which can be understood as the complete content of the target text's paragraph. Since the target text's layout position refers to its specific location within the target document, such as the target page number, paragraph, and line number, the target text's layout position within the target document can be determined based on the target text's layout position, and the corresponding paragraph text can be extracted from the target document based on this layout position.
[0068] For example, if the target text is typeset at line 4, paragraph 3, page 2, then the typeset position of the paragraph containing the target text in the target document can be determined to be paragraph 3, page 2, and the paragraph text can be extracted based on the typeset position of the paragraph in the target document.
[0069] In addition, the keywords of a paragraph text refer to words or phrases that can represent the core content of the paragraph, which are used to represent the main information of the paragraph. Multiple candidate texts refer to other texts in the target document that are semantically similar or related to the keywords. These candidate texts can be a word, a phrase, or a paragraph, etc.
[0070] Optionally, a word frequency statistics can be performed on each word in the paragraph text to determine the word frequency of each word (that is, the number of times each word appears in the paragraph text). When the word frequency is greater than a threshold, it indicates that the corresponding word is closely related to the subject content of the paragraph text. In this case, the word can be used as a keyword for the paragraph text.
[0071] Because the paragraph text corresponds to the target text's paragraph, it provides detailed context for the target text, making it closely related to the target text's semantics. Furthermore, multiple candidate texts are extracted from the literature based on the keywords in the paragraph text, effectively expanding the target text's context and further strengthening its semantics.
[0072] Furthermore, the paragraph text and the multiple candidate texts may be used as the context text of the target text, or the context text of the target text may be obtained by screening the paragraph text and the multiple candidate texts based on semantic relevance.
[0073] Step 240: Determine a translation result of the target text based on the target text and the context text.
[0074] Optionally, after obtaining the context text, the target text and the context text can be combined into a paragraph, and the paragraph can be translated to obtain a translation result of the paragraph. Based on the position of the target text in the paragraph, the translation result of the target text is extracted from the translation result of the paragraph. Since the paragraph contains the context information of the target text, the translation result corresponding to the paragraph fits the actual context of the target text, and the translation result of the target text is extracted from the translation result of the paragraph, that is, the translation result of the target text comes from the original translation result of the paragraph, and the extracted translation result of the target text can more accurately convey the original meaning content.
[0075] Based on any of the above embodiments, Figure 3 This is the third flow chart of the translation method provided by the present invention, as shown in FIG. Figure 3 As shown, the method includes: Step 310: Acquire a first image and a second image, and perform image recognition on each image to obtain page feature information and target text of the target page; the first image is an image corresponding to the target page where the target text is located, and the second image is an image of the target text.
[0076] Specifically, the first image is acquired by a first image acquisition element, and the second image is acquired by a second image acquisition element. The first image acquisition element and the second image acquisition element may be the same image acquisition element or different image acquisition elements. The image acquisition element may be a camera, an image sensor, or the like. In the embodiment of the present invention, the image acquisition element is preferably a camera.
[0077] Step 320: Determine the target document to which the target text belongs based on the page feature information, and determine the layout position of the target text in the target document.
[0078] Specifically, since the pages in each document are different, the page content and / or page layout corresponding to the pages in each document are also different, that is, the page feature information of each page is different, and then based on the page feature information of the target page, the target document to which the target page belongs can be reversely inferred.
[0079] In addition, the typesetting position refers to the specific position of the target text in the target document. The typesetting position may include information such as the target page number where the target text is located, the paragraph where the target text is located, and the number of lines.
[0080] Step 330: Based on the typesetting position, determine the paragraph text from the target document, where the paragraph text refers to the text corresponding to the paragraph where the target text is located; based on the keywords of the paragraph text, extract multiple candidate texts from the document; based on the semantic relevance between the paragraph text and each candidate text, determine the related text from each candidate text; and use the paragraph text and the related text as the context text.
[0081] Specifically, the paragraph text refers to the text corresponding to the paragraph of the target document, which can be understood as the complete content of the paragraph. The keywords of the paragraph text are words or phrases that can represent the core content of the paragraph, and are used to represent the main information of the paragraph. Multiple candidate texts refer to other texts in the target document that are semantically similar or related to the keywords.
[0082] Considering that the paragraph text is the text of the paragraph in which the target text is located, it directly contains detailed information about the target text in terms of structure, context, and linguistic context. Therefore, the paragraph text is necessarily highly semantically related to the target text. However, multiple candidate texts are identified from the target document based on keywords in the paragraph text. Due to the specific location of the text, contextual differences, or different semantic scopes, some candidate texts may have low semantic relevance to the target text.
[0083] For example, let's say the target text is "image recognition" and the paragraph text is "Deep learning is a branch of machine learning that has been widely used in image recognition and has made significant progress." The keywords for the paragraph text are "deep learning, image recognition, application." Based on these keywords, candidate texts extracted from the target document include: Candidate 1, "The application of deep learning in medical imaging has achieved breakthrough progress," and Candidate 2, "The development of image recognition technology has driven innovation in computer vision." Semantic analysis shows that Candidate 1 is closely related to the target text, but Candidate 2 focuses on technological development rather than application scenarios, making it less semantically relevant to the target text.
[0084] Based on this, the embodiment of the present invention uses the semantics of paragraph text as a screening criterion, that is, determining the semantic relevance between each candidate text and the paragraph text. This semantic relevance is used to characterize the strength of the relationship between the candidate text and the paragraph text in the same context. The stronger the relationship, the more background information the candidate text provides, which in turn can assist in understanding the original meaning of the target text. Optionally, candidate texts and paragraph texts with a semantic relevance greater than a threshold can be used as context text.
[0085] The semantic features of each candidate text and the semantic features of the paragraph text may be extracted, and the semantic relevance between each candidate text and the paragraph text may be measured using the Euclidean distance or pre-similarity between the two features.
[0086] Step 340: Determine a translation result of the target text based on the target text and the context text.
[0087] Optionally, after obtaining the context text, the target text and the context text can be combined into a paragraph, and the paragraph can be translated to obtain a translation result of the paragraph. Based on the position of the target text in the paragraph, the translation result of the target text is extracted from the translation result of the paragraph. Since the paragraph contains the context information of the target text, the translation result corresponding to the paragraph fits the actual context of the target text, and the translation result of the target text is extracted from the translation result of the paragraph, that is, the translation result of the target text comes from the original translation result of the paragraph, and the extracted translation result of the target text can more accurately convey the original meaning content.
[0088] Based on any of the above embodiments, Figure 4 This is the fourth flow chart of the translation method provided by the present invention, such as Figure 4 As shown, the method includes: Step 410: Acquire a first image and a second image, and perform image recognition on each to obtain page feature information of a target page and target text; the first image is an image corresponding to the target page where the target text is located, and the second image is an image of the target text.
[0089] Specifically, the first image is acquired by a first image acquisition component, and the second image is acquired by a second image acquisition component. The first image acquisition component and the second image acquisition component can be the same image acquisition component or different image acquisition components.
[0090] Step 420: extract page keywords from the page text information; determine the target document to which the target text belongs based on the page keywords, the page feature information includes the page text information; and determine the layout position of the target text in the target document.
[0091] Specifically, the page feature information includes page text information. The page text information refers to all text content in the target page. The page text information may include page title text, header text, footer text, page body text and other information in the target page.
[0092] Page keywords refer to words that represent the core theme of a target page. These keywords can be obtained by performing natural language processing or keyword extraction on the page text. For example, they can be obtained by performing a word frequency analysis on the page text.
[0093] Considering that different target documents express different subject information, and the page keywords are used to represent the subject information of the target page, we can determine the target document to which the target text belongs by comparing the similarity between the page keywords and the keywords of each document.
[0094] For example, the subject words corresponding to each document may be predetermined, and the page subject words may be matched with the subject words corresponding to each document, and the matched document may be used as the target document.
[0095] Step 430: Extract the context text of the target text from the target document based on the typeset position of the target text in the target document.
[0096] Optionally, after determining the typeset position of the target text in the target document, the paragraph text of the paragraph where the target text is located can be extracted from the target text based on the typeset position. Since the paragraph text usually contains information related to the target text, that is, the paragraph text is closely related to the semantics of the target text, the paragraph text can be used as context text.
[0097] Step 440: Determine a translation result of the target text based on the target text and the context text.
[0098] Optionally, after obtaining the context text, the target text and the context text can be combined into a paragraph, and the paragraph can be translated to obtain a translation result of the paragraph. Based on the position of the target text in the paragraph, the translation result of the target text is extracted from the translation result of the paragraph. Since the paragraph contains the context information of the target text, the translation result corresponding to the paragraph fits the actual context of the target text, and the translation result of the target text is extracted from the translation result of the paragraph, that is, the translation result of the target text comes from the original translation result of the paragraph, and the extracted translation result of the target text can more accurately convey the original meaning content.
[0099] Based on any of the above embodiments, Figure 5 This is the fifth flow chart of the translation method provided by the present invention, such as Figure 5 As shown, the method includes: Step 510: Acquire a first image and a second image, and perform image recognition on each image to obtain page feature information and target text of the target page; the first image is an image corresponding to the target page where the target text is located, and the second image is an image of the target text.
[0100] Specifically, the first image is acquired by a first image acquisition component, and the second image is acquired by a second image acquisition component. The first image acquisition component and the second image acquisition component can be the same image acquisition component or different image acquisition components.
[0101] Step 520: Determine the document type of the target document based on the page layout information. Based on the document type, determine a target document library from multiple document libraries, where the target document library contains multiple documents, each of which has the same document type as the target document. Determine the target document from the target document library based on the page text information, where the page feature information includes page text information and page layout information. Determine the layout position of the target text within the target document.
[0102] Specifically, page layout information refers to the layout of elements on the target page, including the arrangement and relative positions of paragraphs, titles, charts, images, etc. The document type of the target document refers to the category of the target document, such as papers, newspapers, books, technical reports, etc.
[0103] Different document types have different corresponding page layout information. For example, for a paper, its page layout is usually highly structured, usually including the title, abstract, main text, and references. For a newspaper or journal, its page layout is usually mainly based on the title and content blocks, and may include pictures and charts, with a relatively compact and flexible layout.
[0104] If you match the target document to the target text across all documents, you might waste computing resources due to type mismatches. For example, if the target document is a thesis, then if you match it against newspapers and periodicals, you won't find any matching documents because they are of different types. In this case, you only need to match against thesis documents, saving both matching time and computing resources.
[0105] Based on this, the embodiment of the present invention determines a target document library from multiple document libraries based on the document type of the target document. The target document library stores multiple documents of the same document type as the target document. It is understood that multiple document libraries can be pre-set, each corresponding to a different type of document, such as document library 1 storing papers, document library 2 storing newspapers and periodicals, document library 3 storing books, etc.
[0106] After determining the target document library, the method of the above embodiment can be referred to to determine the target document to which the target text belongs based on the page text information, such as extracting page keywords from the page text information; and determining the target document to which the target text belongs based on the page keywords.
[0107] Step 530: Extract the context text of the target text from the target document based on the typeset position of the target text in the target document.
[0108] Optionally, after determining the typeset position of the target text in the target document, the paragraph text of the paragraph where the target text is located can be extracted from the target text based on the typeset position. Since the paragraph text usually contains information related to the target text, that is, the paragraph text is closely related to the semantics of the target text, the paragraph text can be used as context text.
[0109] Step 540: Determine a translation result of the target text based on the target text and the context text.
[0110] Optionally, after obtaining the context text, the target text and the context text can be combined into a paragraph, and the paragraph can be translated to obtain a translation result of the paragraph. Based on the position of the target text in the paragraph, the translation result of the target text is extracted from the translation result of the paragraph. Since the paragraph contains the context information of the target text, the translation result corresponding to the paragraph fits the actual context of the target text, and the translation result of the target text is extracted from the translation result of the paragraph, that is, the translation result of the target text comes from the original translation result of the paragraph, and the extracted translation result of the target text can more accurately convey the original meaning content.
[0111] Based on any of the above embodiments, Figure 6 This is the sixth flow chart of the translation method provided by the present invention, such as Figure 6 As shown, the method includes: Step 610: Acquire a first image and a second image, and perform image recognition on each to obtain page feature information and target text of the target page; the first image is an image corresponding to the target page where the target text is located, and the second image is an image of the target text.
[0112] Specifically, the first image is acquired by a first image acquisition component, and the second image is acquired by a second image acquisition component. The first image acquisition component and the second image acquisition component can be the same image acquisition component or different image acquisition components.
[0113] Step 620: Determine the target document to which the target text belongs based on the page feature information. Match the first image with the second image to determine the image region in the first image that matches the second image; and determine the typesetting position based on the position of the image region in the first image.
[0114] The first image contains an image of the target page where the target text is located, and the second image contains an image of the target text. This means that the first image contains the target text in the second image. Based on this, the first and second images are matched to determine which region in the first image has the highest similarity to the second image. This region with the highest similarity is then used as the image region of the target text in the first image. This image region refers to the region in the first image where the target text is located.
[0115] After determining the image area, the position of the image area in the first image can be used as the layout position of the target text in the target page. Combined with the layout position of the target page in the document, the layout position of the target text in the target document can be determined. The position of the image area in the first image can be determined based on the coordinates of the image area (e.g., the coordinates of the upper left and lower right corners of the image area, or the coordinates of the center point of the image area).
[0116] Step 630: Extract the context text of the target text from the target document based on the typeset position of the target text in the target document.
[0117] Optionally, after determining the typeset position of the target text in the target document, the paragraph text of the paragraph where the target text is located can be extracted from the target text based on the typeset position. Since the paragraph text usually contains information related to the target text, that is, the paragraph text is closely related to the semantics of the target text, the paragraph text can be used as context text.
[0118] Step 640: Determine a translation result of the target text based on the target text and the context text.
[0119] Optionally, after obtaining the context text, the target text and the context text can be combined into a paragraph, and the paragraph can be translated to obtain a translation result of the paragraph. Based on the position of the target text in the paragraph, the translation result of the target text is extracted from the translation result of the paragraph. Since the paragraph contains the context information of the target text, the translation result corresponding to the paragraph fits the actual context of the target text, and the translation result of the target text is extracted from the translation result of the paragraph, that is, the translation result of the target text comes from the original translation result of the paragraph, and the extracted translation result of the target text can more accurately convey the original meaning content.
[0120] Based on any of the above embodiments, Figure 7 This is the seventh flow chart of the translation method provided by the present invention, such as Figure 7 As shown, the method includes: Step 710: Acquire a first image and a second image, and perform image recognition on each to obtain page feature information of a target page and target text; the first image is an image corresponding to the target page where the target text is located, and the second image is an image of the target text.
[0121] Specifically, the first image is acquired by a first image acquisition element, and the second image is acquired by a second image acquisition element. The first image acquisition element and the second image acquisition element may be the same image acquisition element or different image acquisition elements. The image acquisition element may be a camera, an image sensor, or the like. In an embodiment of the present invention, the image acquisition element is preferably a camera, and the image acquisition element is disposed on the camera at the tip of the dictionary pen.
[0122] After detecting that the user has picked up the dictionary pen, the camera at the pen tip begins capturing a first image. This process can be continuous, meaning that as the user brings the pen tip closer to the paper, multiple images are captured, with the best quality image being used as the first image. During this process, the image capture range decreases as the distance between the pen tip and the paper decreases, while the clarity increases.
[0123] In addition, the image quality may be measured by definition. For example, the definition of each image may be determined by using the gradient magnitude, variance, or Laplace operator of each image, and the image with the highest definition may be used as the first image.
[0124] When the pen tip is pressed, the camera on the dictionary pen tip begins capturing a second image and stops capturing the second image when the pen tip is no longer pressed. Typically, the user presses the pen tip and slides it across the paper, with the sliding area covering the target text area. The resulting second image is the image of the target text.
[0125] Step 720: Determine the target document to which the target text belongs based on the page feature information, and determine the layout position of the target text in the target document.
[0126] Specifically, since the pages in each document are different, the page content and / or page layout corresponding to the pages in each document are also different, that is, the page feature information of each page is different, and then based on the page feature information of the target page, the target document to which the target page belongs can be reversely inferred.
[0127] Step 730: Extract the context text of the target text from the target document based on the typeset position of the target text in the target document.
[0128] Optionally, after determining the typeset position of the target text in the target document, the paragraph text of the paragraph where the target text is located can be extracted from the target text based on the typeset position. Since the paragraph text usually contains information related to the target text, that is, the paragraph text is closely related to the semantics of the target text, the paragraph text can be used as context text.
[0129] Step 740: Merge the target text and the context text to obtain a synthesized text, and determine the position of the target text in the synthesized text; translate the synthesized text to obtain a translated text; and extract a translation result of the target text from the translated text based on the position of the target text in the synthesized text.
[0130] Specifically, the synthetic text refers to a text obtained by combining the target text and the context text. For example, the target text and the context text may be combined according to their order in the target document to obtain the synthetic text.
[0131] After obtaining the synthesized text, the synthesized text is translated to obtain a translated text. Since the synthesized text contains the context of the target text, the translated text is generated after considering the context information, so that the translated text can be consistent with the original text content.
[0132] Based on the target text's position in the synthesized text, the target text's translation is extracted from the translated text. This means the target text's translation is derived from the original text. Because the translated text adheres to the original text's semantics, the translation derived from the original text also adheres to the original text's semantics. This means the target text's translation accurately reflects the original text's semantics, resulting in a high degree of translation accuracy.
[0133] When combining the target text and the context text, the positions of the target text and the context text in the synthesized text can be marked. After obtaining the synthesized text, the position of the target text in the synthesized text can be directly determined based on the position identifier.
[0134] The translation device provided by the present invention is described below. The translation device described below and the translation method described above can be referenced to each other.
[0135] Based on any of the above embodiments, Figure 8 Schematic diagram of the structure of the translation device provided by the present invention. Figure 8 As shown, the device includes: The acquisition unit 810 is configured to acquire a first image and a second image, and perform image recognition on each image to obtain page feature information of a target page and target text; the first image is an image of the target page where the target text is located, and the second image is an image of the target text; A determination unit 820 is configured to determine the target document to which the target text belongs based on the page feature information, and to determine the typesetting position of the target text in the target document; An extraction unit 830 is configured to extract context text of the target text from the target document based on the typeset position of the target text in the target document; The translation unit 840 is configured to determine a translation result of the target text based on the target text and the context text.
[0136] Based on any of the above embodiments, extracting context text of the target text from the target document based on the typeset position of the target text in the target document includes: Based on the typesetting position, determine the paragraph text from the target document. The paragraph text refers to the text corresponding to the paragraph where the target text is located. Extract multiple candidate texts from the document based on the keywords of the paragraph text; A context text is determined from the paragraph text and a plurality of candidate texts.
[0137] Based on any of the above embodiments, determining context text from the paragraph text and the plurality of candidate texts includes: Determining relevant texts from among the candidate texts based on the semantic relevance between the paragraph text and the candidate texts; Use paragraph text and related text as contextual text.
[0138] Based on any of the above embodiments, the page feature information includes page text information; Determine the target document to which the target text belongs based on page feature information, including: Extract page keywords from page text information; Determine the target document to which the target text belongs based on the page keywords.
[0139] Based on any of the above embodiments, the page feature information includes page text information and page layout information; Determine the target document to which the target text belongs based on page feature information, including: Determine the document type of the target document based on the page layout information; Based on the document type, a target document library is determined from multiple document libraries, where the target document library stores multiple documents, each of which has the same document type as the target document. Based on the page text information, the target document is determined from the target document library.
[0140] Based on any of the above embodiments, determining the typeset position of the target text in the target document includes: Matching the first image and the second image to determine an image region in the first image that matches the second image; A layout position is determined based on a position of the image area in the first image.
[0141] Based on any of the above embodiments, determining a translation result of the target text based on the target text and the context text includes: Merge the target text and the context text to obtain a synthesized text, and determine the position of the target text in the synthesized text; Translating the synthesized text to obtain a translated text; The translation result of the target text is extracted from the translated text based on the position of the target text in the synthesized text.
[0142] Figure 9 Schematic diagram of the structure of the electronic device provided by the present invention, such as Figure 9 As shown, the electronic device may include: a processor 910, a communication interface 920, a memory 930, and a communication bus 940, wherein the processor 910, the communication interface 920, and the memory 930 communicate with each other via the communication bus 940. The processor 910 may call logic instructions in the memory 930 to execute a translation method, which includes: acquiring a first image and a second image, and performing image recognition on each image to obtain page feature information and target text of a target page; the first image is an image of the target page where the target text is located, and the second image is an image of the target text; determining the target document to which the target text belongs based on the page feature information, and determining the layout position of the target text in the target document; extracting context text of the target text from the target document based on the layout position of the target text in the target document; and determining a translation result of the target text based on the target text and the context text.
[0143] Furthermore, the logic instructions in the aforementioned memory 930 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the portion that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product, stored in a storage medium, includes instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to perform all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as a USB flash drive, a mobile hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.
[0144] On the other hand, the present invention also provides a computer program product, which includes a computer program, which can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the translation method provided by the above methods, which includes: obtaining a first image and a second image, and performing image recognition respectively to obtain page feature information and target text of the target page; the first image is an image of the target page where the target text is located, and the second image is an image of the target text; determining the target document to which the target text belongs based on the page feature information, and determining the typeset position of the target text in the target document; extracting the context text of the target text from the target document based on the typeset position of the target text in the target document; and determining the translation result of the target text based on the target text and the context text.
[0145] On the other hand, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, is implemented to execute the translation method provided by the above-mentioned methods, the method comprising: acquiring a first image and a second image, and performing image recognition respectively to obtain page feature information and target text of the target page; the first image is an image of the target page where the target text is located, and the second image is an image of the target text; determining the target document to which the target text belongs based on the page feature information, and determining the typeset position of the target text in the target document; extracting the context text of the target text from the target document based on the typeset position of the target text in the target document; and determining the translation result of the target text based on the target text and the context text.
[0146] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, i.e., they may be located in one location or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of the present embodiment. Persons of ordinary skill in the art will be able to understand and implement the present invention without inventive effort.
[0147] Through the above description of the embodiments, those skilled in the art will clearly understand that each embodiment can be implemented using software plus a necessary general-purpose hardware platform, or of course, hardware. Based on this understanding, the essence of the above technical solution, or the portion that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a magnetic disk, or an optical disk, and includes a number of instructions for causing a computer device (such as a personal computer, server, or network device) to execute the methods described in each embodiment or certain portions of the embodiments.
[0148] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.
Claims
1. A translation method, characterized in that: include: Acquire the first image and the second image, and perform image recognition on each image to obtain page feature information and target text of the target page; The first image is an image of the target page where the target text is located, and the second image is an image of the target text; Determining the target document to which the target text belongs based on the page feature information, and determining the typeset position of the target text in the target document; extracting context text of the target text from the target document based on the typeset position of the target text in the target document; A translation result of the target text is determined based on the target text and the context text.
2. The translation method according to claim 1, wherein: The step of extracting the context text of the target text from the target document based on the typeset position of the target text in the target document includes: Determining a paragraph text from the target document based on the typeset position, wherein the paragraph text refers to the text corresponding to the paragraph where the target text is located; extracting a plurality of candidate texts from the document based on keywords of the paragraph text; The context text is determined from the paragraph text and the plurality of candidate texts.
3. The translation method according to claim 2, characterized in that The determining the context text from the paragraph text and the plurality of candidate texts includes: Determining relevant texts from among the candidate texts based on the semantic relevance between the paragraph text and the candidate texts; The paragraph text and the related text are used as the context text.
4. The translation method according to any one of claims 1 to 3, characterized in that The page feature information includes page text information; The determining the target document to which the target text belongs based on the page feature information includes: Extracting page keywords from the page text information; The target document to which the target text belongs is determined based on the page keywords.
5. The translation method according to any one of claims 1 to 3, characterized in that: The page feature information includes page text information and page layout information; The determining the target document to which the target text belongs based on the page feature information includes: Determining the document type of the target document based on the page layout information; Based on the document type, determining a target document library from a plurality of document libraries, wherein the target document library stores a plurality of documents, and each document has the same document type as the target document; Based on the page text information, the target document is determined from the target document library.
6. The translation method according to any one of claims 1 to 3, characterized in that: Determining the typeset position of the target text in the target document includes: matching the first image with the second image to determine an image region in the first image that matches the second image; The layout position is determined based on the position of the image area in the first image.
7. The translation method according to any one of claims 1 to 3, characterized in that: The determining a translation result of the target text based on the target text and the context text includes: Merging the target text and the context text to obtain a composite text, and determining a position of the target text in the composite text; translating the synthesized text to obtain a translated text; A translation result of the target text is extracted from the translation text based on the position of the target text in the synthesized text.
8. A translation device, characterized in that: include: an acquisition unit, configured to acquire the first image and the second image, and perform image recognition on each image to obtain page feature information and target text of the target page; The first image is an image of the target page where the target text is located, and the second image is an image of the target text; a determining unit, configured to determine the target document to which the target text belongs based on the page feature information, and determine the typeset position of the target text in the target document; an extraction unit, configured to extract a context text of the target text from the target document based on a typeset position of the target text in the target document; The translation unit is configured to determine a translation result of the target text based on the target text and the context text.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that: When the processor executes the computer program, the translation method according to any one of claims 1 to 7 is implemented.
10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the translation method according to any one of claims 1 to 7 is implemented.
11. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the translation method according to any one of claims 1 to 7 is implemented.