Picture processing method and apparatus, and electronic device
By merging incomplete text that satisfies semantic information to generate complete merged text, the problem of insufficient processing power of electronic devices when dealing with complex text is solved, and more efficient text processing and output are achieved.
Patent Information
- Application Number
- CN202111509057.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-10
- Publication Date
- 2025-12-30
- Estimated Expiration
- 2041-12-10
AI Technical Summary
Existing technologies have poor text processing capabilities for electronic devices when dealing with complex text, such as text in columns, pages, or irregularly shaped text.
By acquiring multiple texts from the target image and their completeness and translation completeness information, incomplete texts that satisfy semantic information are merged to generate a complete merged text, which is then output when both the merged text and the translation are complete.
It improves the ability to process complex text in images, making the semantics of merged text more fluent and enhancing the accuracy and efficiency of text processing.
Smart Images

Figure CN114299525B_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of communication technology, specifically relating to an image processing method, apparatus, and electronic device. Background Technology
[0002] With the development of electronic device technology, electronic devices are being used more and more widely. For example, electronic devices can recognize and process text in images.
[0003] Currently, when an image contains multiple lines of text, during the process of electronic devices recognizing the image, the electronic devices can merge the multiple lines of text in the image based on the physical location coordinates of the text lines and the text layout.
[0004] However, based on the above method, when the text in an image includes complex text such as column text, paginated text, or irregularly shaped text, the electronic device may be unable to merge the text in the image according to the physical coordinates of the text lines and the text layout. This results in the electronic device having poor text processing capabilities for images. Summary of the Invention
[0005] The purpose of this application is to provide an image processing method, apparatus, and electronic device that can solve the problem of poor text processing capabilities of electronic devices in images.
[0006] In a first aspect, embodiments of this application provide an image processing method, the method comprising: acquiring N texts and target information included in a target image, the target information including at least one of the following: a first completeness of the N texts, a second completeness of a first translation corresponding to the N texts, where N is an integer greater than 1; merging S texts from P texts that satisfy the first semantic information to obtain a first text, wherein the P texts are incomplete texts determined from the N texts based on the first completeness, where P and S are both integers greater than 1; and outputting the first text when both the first text and the second translation are complete texts, wherein the second translation is the text obtained by merging the translations corresponding to the S texts in the third translation, and the third translation is an incomplete translation determined from the first translation based on the second completeness.
[0007] Secondly, embodiments of this application provide an image processing apparatus, comprising an acquisition module, a processing module, and an output module. The acquisition module is configured to acquire N texts and target information included in a target image. The target information includes at least one of the following: a first completeness of the N texts, and a second completeness of a first translation corresponding to the N texts, where N is an integer greater than 1. The processing module is configured to merge S texts from P texts that satisfy the first semantic information to obtain a first text. The P texts are incomplete texts determined from the N texts based on the first completeness, where P and S are both integers greater than 1. The output module is configured to output the first text if the first text and the second translation are complete texts. The second translation is the text obtained by merging the translations corresponding to the S texts in the third translation, where the third translation is an incomplete translation determined from the first translation based on the second completeness.
[0008] Thirdly, embodiments of this application provide an electronic device including a processor and a memory, wherein the memory stores a program or instructions executable on the processor, and the program or instructions, when executed by the processor, implement the steps of the method as described in the first aspect above.
[0009] Fourthly, embodiments of this application provide a readable storage medium on which a program or instructions are stored, which, when executed by a processor, implement the steps of the method as described in the first aspect above.
[0010] Fifthly, embodiments of this application provide a chip including a processor and a communication interface, the communication interface being coupled to the processor, the processor being used to run programs or instructions to implement the method as described in the first aspect above.
[0011] In a sixth aspect, embodiments of this application provide a computer program product stored in a storage medium, which is executed by at least one processor to implement the method described in the first aspect.
[0012] In this embodiment, N texts and target information included in the target image are obtained. The target information includes at least one of the following: a first completeness of the N texts, a second completeness of the first translation corresponding to the N texts, where N is an integer greater than 1; S texts satisfying the first semantic information from P texts are merged to obtain a first text, where the P texts are incomplete texts determined from the N texts based on the first completeness, where P and S are both integers greater than 1; if both the first text and the second translation are complete texts, the first text is output; the second translation is the text obtained by merging the translations corresponding to the S texts in the third translation, where the third translation is an incomplete translation determined from the first translation based on the second completeness. Through this scheme, after obtaining multiple texts and target information in the target image, since at least one text satisfying the semantic information from the incomplete texts determined based on the target information can be merged to obtain a merged text, when the text in the image includes complex texts such as column texts, paginated texts, or irregularly shaped texts, these complex texts can be merged based on semantic information. Furthermore, since the merged text is only output when both the merged text and its corresponding translation are complete texts, the resulting merged text has a more fluent semantic meaning. This improves the ability to process text within images. Attached Figure Description
[0013] Figure 1 A schematic diagram illustrating an image processing method provided in an embodiment of this application;
[0014] Figure 2(a) is one of the schematic diagrams of an image processing interface provided in an embodiment of this application;
[0015] Figure 2(b) is a second schematic diagram of an image processing interface provided in an embodiment of this application;
[0016] Figure 3 This is a schematic diagram of the structure of the image processing apparatus provided in the embodiments of this application;
[0017] Figure 4 This is a schematic diagram of the structure of the electronic device provided in the embodiments of this application;
[0018] Figure 5 A hardware schematic diagram of the electronic device provided in the embodiments of this application. Detailed Implementation
[0019] The technical solutions of the embodiments of this application will be clearly described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application are within the scope of protection of this application.
[0020] The terms "first," "second," etc., used in the specification and claims of this application are used to distinguish similar objects and not to describe a specific order or sequence. It should be understood that such use of data can be interchanged where appropriate so that embodiments of this application can be implemented in orders other than those illustrated or described herein, and the objects distinguished by "first," "second," etc., are generally of the same class and the number of objects is not limited; for example, a first object can be one or more. Furthermore, in the specification and claims, "and / or" indicates at least one of the connected objects, and the character " / " generally indicates that the preceding and following objects are in an "or" relationship.
[0021] The image processing method, apparatus, and electronic device provided in this application will be described in detail below with reference to the accompanying drawings and through specific embodiments and application scenarios.
[0022] like Figure 1 As shown in the figure, this application provides an image processing method, which includes the following steps S101 to S103.
[0023] S101, The image processing device acquires N texts and target information included in the target image.
[0024] The target information mentioned above includes at least one of the following: the first completeness of N texts, the second completeness of the first translation corresponding to the N texts, where N is an integer greater than 1.
[0025] Optionally, the target image can be any of the following: an image taken by an electronic device, a screenshot saved by an electronic device, or an online image acquired by an electronic device.
[0026] Optionally, in this embodiment, the target image may include multiple texts. The N texts are texts selected from these multiple texts.
[0027] Optionally, the language type of the above N texts can be Chinese, English, Korean, Japanese, etc.
[0028] Furthermore, each of the aforementioned N texts can be text in the traditional sense or a line of text. The specific choice depends on the actual usage, and this application does not impose any limitations on this.
[0029] Furthermore, when one of the N texts is a text line, that text line can be an independent single text line, in which case the single text line can be considered as a text; or, the text line can be a text line within a certain text paragraph.
[0030] Optionally, in this embodiment of the application, image text recognition technology can be used to identify the text content included in the target image. The text content may specifically include: the text included in the target image and the coordinates of the text.
[0031] Optionally, the aforementioned first translation may include a translation in one language or multiple languages. The specific choice can be determined based on actual usage, and this application embodiment does not impose such a limitation.
[0032] Optionally, the first completeness and the second completeness mentioned above are determined based on semantic information. For details, please refer to the detailed description of the following embodiments. The embodiments of this application will not be repeated here.
[0033] S102. The image processing device merges S texts from P texts that satisfy the first semantic information to obtain the first text.
[0034] Here, the P texts mentioned above are incomplete texts determined from N texts based on a first completeness score. P and S are both integers greater than 1.
[0035] Optionally, determining whether P out of N texts are incomplete texts can include the following scenarios:
[0036] Scenario 1: The semantics of P texts are incomplete.
[0037] Scenario 2: The sentence structure of the first or last sentence of each of the P texts is missing.
[0038] Scenario 3: The sentence-ending words in each of the P texts cannot stand alone as words.
[0039] It should be noted that the above three scenarios all use semantic information to determine that P out of N texts are incomplete texts. These three scenarios are merely illustrative examples provided by the embodiments of this application. Of course, determining that P out of N texts are incomplete texts using semantic information can also include other implementation methods, which are not limited in this application.
[0040] Optionally, the aforementioned first semantic information may include at least one of the following: sentence structure information, sentence component information, and phrase composition information.
[0041] For example, taking the first semantic information as sentence structure information. Sentence structure information may include at least one of the following: subject-predicate structure, verb-object structure, subject-predicate-object structure, subject-predicate-object attributive-adverbial-complement structure, etc.
[0042] For example, taking the first semantic information as sentence component information. Sentence component information may include at least one of the following: subject, predicate, object, attributive, adverbial, complement, etc.
[0043] For example, let's take phrase composition information as the first semantic information. Phrase composition information may include at least one of the following: sentence-initial word, sentence-ending word, common words, phrases, etc.
[0044] It should be noted that the above embodiments are merely exemplary descriptions of the first semantic information. Of course, the first semantic information may also include other information related to semantics, and the embodiments of this application do not limit this.
[0045] Furthermore, the above description of the first semantic information is merely an example of possible scenarios when the N texts are Chinese texts. When the N texts are in other language types, the semantic information can be interpreted according to the semantic rules or grammar of those other language types. This application does not limit this.
[0046] Optionally, in the embodiments of this application, one possible case is that the P texts only include a set of texts that conform to the first semantic information, that is, the S texts are this set of texts; another possible case is that the P texts include multiple sets of texts that conform to the first semantic information, and the S texts are any set of texts among the multiple sets of texts.
[0047] Furthermore, in the case where P texts include multiple sets of texts that conform to the first semantic information, the implementation method for merging the multiple sets of texts can be referred to the detailed description of S texts, which will not be repeated in the embodiments of this application.
[0048] S103. If both the first text and the second translation are complete texts, the image processing device outputs the first text.
[0049] Among them, the second translation is the text obtained by merging the translations corresponding to the S texts in the third translation, and the third translation is the incomplete translation determined from the first translation based on the second completeness.
[0050] Example 1: Taking a mobile phone as the image processing device. As shown in Figure 2(a), when the mobile phone displays an image, the image includes text S1 to S8; as shown in Figure 2(b), the translations corresponding to S1 to S8 are 01 to 08, i.e., the first translation. The mobile phone can obtain the text S1 to S8 included in the image, as well as the completeness of S1 to S8 and the completeness of the first translation. Since S2 to S8 are incomplete texts, the mobile phone can merge S2, S3, and S4 from S2 to S8 that satisfy the first semantic information to obtain S9 after merging S2, S3, and S4. Then, if S9 and the translations obtained by merging 02 to 08 corresponding to S2 to S8 are both complete texts, S9 is output.
[0051] Furthermore, for S5 and S6, and S7 and S8, which satisfy the first semantic information among these 8 texts, the above process can be repeated to merge S5 and S6, and S7 and S8 respectively. Finally, output S10 after merging S5 and S6, and S11 after merging S7 and S8.
[0052] In this way, through the above process, multiple sets of text that satisfy semantic information in an image can be merged, thereby completing the text processing in the image.
[0053] Optionally, after S102 and before S103, the image processing method provided in this application embodiment may further include: the image processing device determining that the first text is complete text based on the first semantic information.
[0054] Furthermore, for determining whether the first text is complete (S2 to S8), please refer to the detailed description of determining the completeness of N texts in the following embodiments, which will not be repeated here.
[0055] This application provides an image processing method. After acquiring multiple texts and target information from a target image, it merges at least one text that satisfies semantic information from the incomplete texts determined based on the target information to obtain a merged text. Therefore, when the text in the image includes complex text such as column text, paginated text, or irregularly shaped text, these complex texts can be merged according to semantic information. Furthermore, since the merged text and its corresponding translation are both complete texts before being output, the semantics of the resulting merged text are more fluent. This improves the ability to process text in images.
[0056] Optionally, the first completeness includes the sentence completeness of the first target sentence of each of the N texts, and the second completeness includes the sentence completeness of the second target sentence of each of the first translations; accordingly, the above S101 may specifically include implementation through the following S101A to S101C.
[0057] S101A, The image processing device extracts the text contained in the target image to obtain N texts.
[0058] S101B, the image processing device analyzes the sentence completeness of the first target sentence based on the first semantic information.
[0059] The aforementioned first semantic information may include at least one of the following: sentence structure information, sentence component information, and phrase composition information. The first target clause may include at least one of the following: the first sentence in the text, or the last sentence in the text.
[0060] Optionally, based on the first semantic information, analyzing the sentence integrity of the first target sentence may include the following two possible implementation manners:
[0061] Implementation manner one: Using the first semantic information as a preset rule to analyze the sentence integrity of the first target sentence.
[0062] Exemplarily, judging from the first semantic information such as sentence pattern structure, phrase composition, and sentence components, and screening out the texts that may be incomplete. Taking Figure 2(a) as an example. The first sentence of S6 is "A newly opened store of the Tao family". By analyzing the sentence component information, it can be known that the first sentence of S6 lacks a subject, so S6 is considered incomplete.
[0063] Exemplarily, judging based on the first semantic information such as the first word, the last word, and phrases of the sentence. Word lists, phrase lists, and end-word lists of different types of languages can be constructed, and a weight is set for each word in the word list; among them, the weight can be set according to the usage frequency of a word. In this way, based on the phrase composition information, it can be judged whether the last word of the last sentence of the text is a common end-word, or whether the first word and the last word of the text can form a word or a phrase alone to determine whether the text line paragraph is complete. As shown in Figure 2(a), the last word of S2 is "Peng", and through the word list, it can be known that the probability that "Peng" can form a word alone and be used as an end-word is very low, so S2 is considered incomplete.
[0064] Implementation manner two: Constructing a semantic model corresponding to the first semantic information, inputting N texts into the semantic model, and analyzing the sentence integrity of the first target sentence.
[0065] Specifically, text data with features such as morphology, syntax structure, and end-words can be used to train the semantic model, and through the semantic model, the first semantic information such as the词性 (should be 'part-of-speech' in English) and syntax structure of different types of languages can be set. In this way, the semantic model can be directly used to judge whether the target sentence of the current text is complete.
[0066] It should be noted that when constructing the semantic model, information such as the morphology, syntax structure, and possible missing sentence components of incomplete sentences is output at the same time. The specific algorithm of the semantic model in this application embodiment is not limited, as long as the corresponding model training data is constructed according to different types of languages.
[0067] S101C. The image processing device analyzes the sentence integrity of the second target sentence based on the second semantic information.
[0068] Among them, the above second semantic information respectively includes at least one of the following: sentence pattern structure information, sentence component information, phrase composition information.
[0069] Optionally, for the specific implementation of analyzing the sentence integrity of the second target sentence based on the second semantic information, reference may be made to the detailed description of S101B in the foregoing embodiments, and details thereof are not elaborated in the embodiments of the present application.
[0070] Optionally, after the above S101B and before S102, the image processing method provided by the embodiments of the present application may further include the following S104.
[0071] S104. The image processing device determines P texts from the N texts according to the sentence integrity of the first target sentence.
[0072] Among them, the above P texts correspond to the third translation.
[0073] Example 2. Combining with the above Example 1, according to the first semantic information, since the last word of the last sentence of S2 is "friend", which cannot be used as an end word of the sentence, the last sentence is incomplete, that is, S2 is incomplete; the first word of the first sentence of S3 is "friend", which cannot be used as a start word of the sentence, so the first sentence is incomplete, that is, S3 is incomplete; the first sentence of S4 is "a lot of things", which lacks the subject and predicate, so the first sentence is incomplete, that is, S4 is incomplete; the last sentence of S5 is "I know", which lacks the object, so the last sentence is incomplete, that is, S5 is incomplete; the first sentence of S6 is "know...", which lacks the subject, so the first sentence is incomplete, that is, S6 is incomplete; the last sentence of S7 is "can't", which cannot be used as an end word of the sentence, so the last sentence is incomplete, that is, S7 is incomplete; the first sentence of S8 is "and", which cannot be used as a start word of the sentence, so the first sentence is incomplete, that is, S8 is incomplete. Thus, incomplete paragraphs S2 to S8 can be determined from S1 to S8.
[0074] The image processing method provided by the embodiments of the present application can extract the texts included in the target image to obtain N texts, analyze the sentence integrity of the first target sentence based on the first semantic information, and analyze the sentence integrity of the second target sentence based on the second semantic information, that is, the integrity of the N texts and the integrity of the translations of the N texts can be determined.
[0075] Furthermore, since P texts can be determined from the N texts according to the sentence integrity of the first target sentence, it is convenient to select texts that meet the first semantic information from the P texts for merging later.
[0076] Optionally, after the above S101 and before S102, the image processing method provided by the embodiments of the present application may further include the following S105 to S108.
[0077] S105. The image processing device obtains at least two texts from P texts that match the second text in the P texts, based on the first semantic information.
[0078] Optionally, the description of the first semantic information can be referred to the detailed description in the above embodiments, and will not be repeated in this application embodiment.
[0079] Optionally, the second text can be any one of the P texts. For example, the second text can be the text that appears first in the P texts.
[0080] Optionally, S105 may specifically include: the image processing device determining, based on the first semantic information, whether the second text in the P texts can be merged with any other text in the P texts other than the second text, thereby obtaining at least two texts that match the second text.
[0081] Furthermore, in S105 above, "at least two texts that match the second text among P texts" means that the second text and at least two other texts satisfy the first semantic information.
[0082] S106. The image processing device merges the second text with the at least two texts respectively to obtain at least two merged texts.
[0083] Alternatively, the above S106 may include the following two specific possible implementations:
[0084] (1) Directly merge the second text with the at least two texts respectively to obtain at least two merged texts.
[0085] (2) Combine the last sentence of the second text with the first sentence of each of the at least two texts to obtain at least two combined sentences.
[0086] S107. The image processing device determines the sentence perplexity of the at least two merged texts.
[0087] Sentence perplexity is used to indicate the fluency of sentences in the merged text.
[0088] Understandably, determining the sentence perplexity of at least two merged texts essentially means determining the sentence perplexity of the merged sentences included in at least two merged texts.
[0089] It should be noted that the lower the sentence perplexity, the higher the sentence fluency, and thus the higher the semantic accuracy; conversely, the higher the sentence perplexity, the lower the sentence fluency, and thus the lower the semantic accuracy.
[0090] S108. The image processing device determines the third text corresponding to the target merged text as the text to be merged corresponding to the second text.
[0091] The target merged text is the text with the lowest sentence perplexity among at least two merged texts. The S texts include the second text and the third text.
[0092] It should be noted that since the lower the sentence perplexity, the higher the sentence fluency, after identifying the third text with the lowest sentence perplexity as the text to be merged corresponding to the second text, the second text and the third text can be merged.
[0093] The image processing method provided in this application, after obtaining at least two texts that match the second text in P texts based on the first semantic information, can merge the second text with the at least two texts respectively to obtain at least two merged texts, and determine the sentence perplexity of the at least two merged texts. Therefore, it can select the text to be merged that is more matched with the second text from the at least two texts based on the perplexity of the two merged texts, thereby improving the correctness of text merging.
[0094] Optionally, after S101 and before S102, the image processing method provided in this application embodiment may further include S109 and S110. That is, the above-mentioned specific implementation can be achieved through S110 to S112.
[0095] S109. The image processing device determines Q adjacent texts from the P texts based on the distribution position of each text in the P texts.
[0096] Where Q is an integer greater than or equal to S.
[0097] It should be noted that by analyzing the distribution of each text in the P texts, the distribution of two texts that can be merged is determined, thereby eliminating some texts that obviously cannot be merged to form the same paragraph. Thus, Q texts with adjacent distribution positions are identified from the P texts.
[0098] Specifically, if two texts contain other text in between, then the two texts cannot be merged; that is, they cannot be merged across lines. Alternatively, you can record the numbers of the texts that cannot be merged.
[0099] For example, as shown in Figure 2(a), the image includes texts S1, S2, ..., and S8. Based on the distribution of these 8 texts, since there are multiple text lines between two text segments, it can be determined that S2 and S7, S2 and S6, and S2 and S8 cannot be merged. Therefore, the numbers of the texts that cannot be merged can be recorded in the non-merge list not_merge_list = [S2_S7, S2_S6, S2_S8].
[0100] Understandably, since the merging of two texts involves a sequential relationship, the order of the numbers in the non-mergeable list can represent the actual merging order. For example, in the non-mergeable list, S2_S7 means that the next line after S2 is not S7, but it does not mean that the next line after S7 cannot be S2.
[0101] S110, the image processing device determines the S texts that satisfy the first semantic information from the Q texts as the texts to be merged.
[0102] Optionally, the description of the first semantic information can be referred to the detailed description in the above embodiments, and will not be repeated in this application embodiment.
[0103] Optionally, for the implementation of determining S texts that satisfy the first semantic information from Q texts, refer to the detailed description in S105 to S108 of the above embodiments. Specifically, it may include:
[0104] (1) Based on the first semantic information, obtain at least two texts that match text 1 in the Q texts from the Q texts.
[0105] (2) Merge text 1 with the at least two texts respectively to obtain at least two merged texts.
[0106] (3) Determine the sentence perplexity of the at least two merged texts.
[0107] (4) Determine the text 2 corresponding to merged text 1 as the text to be merged. Merged text 1 is the text with the lowest sentence perplexity among at least two merged texts. The S texts include text 1 and text 2.
[0108] It should be noted that if no text matching text 2 is obtained according to the first semantic information, then the S texts only include text 1 and text 2, so that by using (1) to (4) in the above embodiment, the S texts that satisfy the semantic information can be determined from the Q texts.
[0109] If other texts matching text 2 are obtained based on the first semantic information, it means that the S texts also include other texts besides text 1 and text 2, so that (1) to (4) in the above embodiment can continue to be executed in a loop to determine other texts matching text 1 and text 2.
[0110] Thus, through the above implementation method, S texts that satisfy the first semantic information can be obtained from Q texts, and these texts can be identified as the texts to be merged.
[0111] Understandably, since Q texts with adjacent distribution positions can be determined from the P texts based on their distribution positions, some texts that cannot be merged at their distribution positions can be excluded, thus reducing unnecessary text merging operations on electronic devices. Furthermore, since S texts satisfying the first semantic information can be identified as the texts to be merged from the Q texts, after a rough screening based on distribution positions, the texts to be merged are determined from the Q texts based on the first semantic information, resulting in a higher semantic fluency of the merged text.
[0112] Optionally, after S110 and before S102, the image processing method provided in this application embodiment may further include S111 as described below. Accordingly, S102 can be specifically implemented through S102A as described below.
[0113] S111, The image processing device determines the target arrangement order of S texts based on the first semantic information.
[0114] It is understandable that, since the first semantic information includes sentence structure information, sentence component information, and phrase composition information, the text order can be determined based on the sentence component information and phrase composition information.
[0115] S102A, the image processing device merges the S texts according to the target arrangement order to obtain the first text.
[0116] It should be noted that merging S texts according to the target order is essentially: merging the last sentence of one text and the first sentence of the other text from two adjacent texts in the S texts, and repeating this process until all S texts are merged to obtain the first text.
[0117] Exemplarily, taking the first semantic information as the sentence pattern structure information and the sentence component information as an example. Suppose the last sentence of text A is "I know", which is a subject-predicate structure; the first sentence of text B is "know a newly opened store", which is a verb-object structure. According to the sentence pattern structure information, the sentence component information, and the phrase composition information, it can be known that text A lacks an object, text B lacks a subject, and "know" and "know" conform to the phrase composition information, so the arrangement order of text A and text B can be determined as A_B. That is, text B is merged at the end of text A.
[0118] Exemplarily, taking the first semantic information as the phrase composition information as an example. Suppose the last word of the last sentence of text C is "friend"; the first word of the first sentence of text D is "companion". According to the phrase composition information, it can be known that "friend" in text C and "companion" in text D conform to the phrase composition information, so the arrangement order of text C and text D can be determined as C_D. That is, text D is merged at the end of text C.
[0119] The image processing method provided by the embodiments of the present application can determine the target arrangement order of S texts according to the first semantic information. Therefore, after merging the S texts according to the target arrangement order to obtain the first text, the semantics of the first text are more complete, and the problem of semantic contradiction is not likely to occur.
[0120] Optionally, the image processing method provided by the embodiments of the present application may further include another possible implementation manner. The method may further include the following S112 to S115.
[0121] S112. Obtain M texts in the target image.
[0122] S113. When T text paragraphs among the M texts are incomplete texts, merge L texts that meet the third semantic information among the T texts to obtain the fourth text.
[0123] Wherein, M, T, and L are all integers greater than 1;
[0124] Optionally, for the description of the third semantic information, reference may be made to the relevant description of the first semantic information in the above embodiments, and the embodiments of the present application will not elaborate on this.
[0125] S114. When the fourth text is a complete text, the image processing device translates the fourth text to obtain the fourth translation.
[0126] Optionally, the above fourth translation may include a translation of one language type, or include translations of multiple language types. The embodiments of the present application do not limit the quantity and language type of the fourth translation.
[0127] For example, the fourth text is Chinese text and the fourth translation is English translation; or, the fourth text is English text and the fourth translation includes Chinese translation and Korean translation.
[0128] S115. When both the fourth text and the fourth translation are complete texts, the image processing device outputs the fourth text and the fourth translation.
[0129] For example, suppose the second text is Chinese text. If the Chinese text is determined to be complete text, it is translated to obtain an English translation. If the English translation is also complete text, the image processing device can output both the Chinese text and the English translation.
[0130] The image processing method provided in this application, after acquiring M texts from a target image, can merge L texts from T texts that satisfy the third semantic information to obtain a fourth text, and translate the fourth text to obtain a fourth translation. Therefore, the first text and the first translation are only output when the fourth text is a complete text and the fourth translation is a complete paragraph. This allows for determining whether to output the fourth text based on whether the merged fourth text is complete, combined with a judgment on the completeness of the fourth translation, thereby improving the accuracy of paragraph merging. Furthermore, since a fourth translation can also be output, in scenarios where text in a target image needs to be translated, a translation with higher accuracy can be output.
[0131] Optionally, after S114 above, the image processing method provided in this application embodiment may further include S116 and S117 below.
[0132] S116. In the case that the fourth translation is an incomplete text, the image processing device merges R texts out of T texts to obtain the fifth text.
[0133] Among them, the above R texts include paragraphs determined based on the semantic information of the fourth translation, where R is an integer greater than 1.
[0134] Optionally, the aforementioned R texts may include all texts from the L texts, or may include a portion of the L texts, depending on the actual situation. This application embodiment does not impose any limitations on this.
[0135] It should be noted that R texts are texts that satisfy the semantic information among T texts.
[0136] Furthermore, when the fourth translation is an incomplete text, based on the semantic information of the fourth translation, other texts that satisfy the semantic information can be obtained from the T texts, and the fourth text can be merged with these other texts. It is understandable that the merged position of the fourth text with these other texts corresponds to the position of the semantically incomplete text in the fourth translation.
[0137] S117. If both the fifth text and the fifth translation are complete paragraphs, the image processing device outputs the third text and the fifth translation.
[0138] The fifth translation mentioned above is the translation corresponding to the fifth text.
[0139] Optionally, for the explanation of determining whether the fifth text and the fifth translation are complete texts, refer to the explanation of the first text in the above embodiments, and this application embodiment will not repeat it again.
[0140] It's important to note that since the purpose of translating text within images is to obtain a semantically accurate translation, the completeness of the translation is paramount in image translation. If the translation is incomplete, even if the original text paragraph (also known as the source text paragraph) is complete, it's necessary to merge the semantically incomplete text into the corresponding positions in the source text paragraph, based on the locations of the incomplete text in the translation. This process requires a second translation to assess the completeness of the final output text.
[0141] It is understandable that merging from the original text paragraphs can ensure the integrity of the original paragraphs. Only when the original paragraphs are complete can an effective translation be obtained after the original paragraphs are translated by the translation model. Conversely, if only merging is done on the translation side, it is difficult to obtain a translation that satisfies the semantic information.
[0142] Optionally, after S116 and before S117, the image processing method provided in this application embodiment may further include: translating the fifth text to obtain a fifth translation if the fifth text is a complete text. Thus, the translation process is only performed when the merged fifth text is a complete text, thereby avoiding invalid translation operations when the merged text is an incomplete text, and also saving the operating resources of the electronic device.
[0143] The image processing method provided in this application improves the accuracy of text merging by merging R texts from T texts to obtain a fifth text when the fourth translation is incomplete. This allows for the re-merging of R texts from T texts that satisfy semantic information based on the incomplete fourth translation. Furthermore, since the fifth text and fifth translation are only output when both are complete texts, a highly accurate translation can be guaranteed.
[0144] The image processing method provided in this application embodiment can be executed by an image processing device. This application embodiment uses an image processing device executing the image processing method as an example to illustrate the method provided in this application embodiment. Figure 3 As shown in the figure, this application provides an image processing apparatus 200, which may include an acquisition module 201, a processing module 202, and an output module 203. The acquisition module 201 is used to acquire N texts and target information included in a target image. The target information includes at least one of the following: a first completeness of the N texts, and a second completeness of the first translation corresponding to the N texts, where N is an integer greater than 1. The processing module 202 is used to merge S texts from P texts that satisfy the first semantic information to obtain a first text. The P texts are incomplete texts determined from the N texts based on the first completeness, where P and S are both integers greater than 1. The output module 203 is used to output the first text when the first text and the second translation are complete texts. The second translation is the text obtained by merging the translations corresponding to the S texts in the third translation, where the third translation is an incomplete translation determined from the first translation based on the second completeness.
[0145] Optionally, the first completeness includes the sentence completeness of the first target sentence in each of the N texts, and the second completeness includes the sentence completeness of the second target sentence in each of the first translations. The acquisition module 201 is specifically used to extract the text included in the target image, obtaining N texts; and to analyze the sentence completeness of the first target sentence based on first semantic information; and to analyze the sentence completeness of the second target sentence based on second semantic information; wherein the first target sentence and the second target sentence each include at least one of the following: the first sentence in the text, the last sentence in the text; the first semantic information and the second semantic information each include at least one of the following: sentence structure information, sentence component information, and phrase composition information;
[0146] Optionally, the image processing apparatus may further include a determining module. The determining module may be used to determine P texts from N texts based on the sentence completeness of the first target sentence, wherein the P texts correspond to the third translation.
[0147] Optionally, the image processing apparatus may further include a determining module. The acquiring module 201 may also be configured to acquire at least two texts from the P texts that match the second text among the P texts, based on the first semantic information. The processing module 202 may also be configured to merge the second text with each of the at least two texts to obtain at least two merged texts. The determining module is configured to determine the third text corresponding to the target merged text as the text to be merged corresponding to the second text, wherein the target merged text is the text with the lowest sentence perplexity among the at least two merged texts; wherein the S texts include the second text and the third text.
[0148] Optionally, the image processing apparatus may further include a determining module. The determining module can be configured to determine Q adjacent texts from the P texts based on the distribution position of each text in the P texts, where Q is an integer greater than or equal to S; and to determine the S texts among the Q texts that satisfy the first semantic information as the texts to be merged.
[0149] Optionally, the determining module can also be used to determine the target arrangement order of the S texts based on the first semantic information. The processing module can specifically be used to merge the S texts according to the target arrangement order to obtain the first text.
[0150] This application provides an image processing apparatus. After acquiring multiple texts and target information from a target image, it can merge at least one text that satisfies semantic information from the incomplete texts determined based on the target information to obtain a merged text. Therefore, when the text in the image includes complex text such as column text, paginated text, or irregularly shaped text, these complex texts can be merged according to semantic information. Furthermore, since the merged text and its corresponding translation are both complete texts before being output, the semantics of the resulting merged text are more fluent. Thus, the processing capability of text in images is improved.
[0151] The image processing device in this application embodiment can be an electronic device or a component within an electronic device, such as an integrated circuit or a chip. The electronic device can be a terminal or other devices besides a terminal. For example, the electronic device can be a mobile phone, tablet computer, laptop computer, PDA, in-vehicle electronic device, mobile internet device (MID), augmented reality (AR) / virtual reality (VR) device, robot, wearable device, ultra-mobile personal computer (UMPC), netbook, or personal digital assistant (PDA), etc. It can also be a server, network attached storage (NAS), personal computer (PC), television set (TV), ATM, or self-service machine, etc. This application embodiment does not specifically limit the device.
[0152] The image processing device in this application embodiment can be a device with an operating system. This operating system can be Android, iOS, or other possible operating systems; this application embodiment does not specifically limit the specific operating system used.
[0153] The image processing device provided in this application embodiment can achieve... Figure 1 The various processes implemented in the method embodiment shown in Figure 2 will not be described again here to avoid repetition.
[0154] Optional, such as Figure 4 As shown, this application embodiment also provides an electronic device 300, including a processor 301 and a memory 302. The memory 302 stores a program or instructions that can run on the processor 301. When the program or instructions are executed by the processor 301, they implement the various steps of the above-described image processing method embodiment and can achieve the same technical effect. To avoid repetition, they will not be described again here.
[0155] It should be noted that the electronic devices in the embodiments of this application include the mobile electronic devices and non-mobile electronic devices described above.
[0156] Figure 5 A schematic diagram of the hardware structure of an electronic device to implement an embodiment of this application.
[0157] The electronic device 400 includes, but is not limited to, components such as: radio frequency unit 401, network module 402, audio output unit 403, input unit 404, sensor 405, display unit 406, user input unit 407, interface unit 408, memory 409, and processor 410.
[0158] Those skilled in the art will understand that the electronic device 400 may also include a power supply (such as a battery) for supplying power to various components. The power supply may be logically connected to the processor 410 through a power management system, thereby enabling functions such as managing charging, discharging, and power consumption through the power management system. Figure 5 The electronic device structure shown does not constitute a limitation on the electronic device. The electronic device may include more or fewer components than shown, or combine certain components, or have different component arrangements, which will not be elaborated here.
[0159] The processor 410 can be used to acquire N texts and target information included in the target image, wherein the target information includes at least one of the following: a first completeness of the N texts, a second completeness of the first translation corresponding to the N texts, where N is an integer greater than 1; and to merge S texts from P texts that satisfy the first semantic information to obtain a first text, wherein the P texts are incomplete texts determined from the N texts based on the first completeness, where P and S are both integers greater than 1; and to output the first text when the first text and the second translation are complete texts, wherein the second translation is the text obtained by merging the translations corresponding to the S texts in the third translation, and the third translation is an incomplete translation determined from the first translation based on the second completeness.
[0160] Optionally, the first completeness includes the sentence completeness of the first target sentence in each of the N texts, and the second completeness includes the sentence completeness of the second target sentence in each of the first translations. The processor 410 is specifically configured to extract the text included in the target image to obtain N texts; and to analyze the sentence completeness of the first target sentence based on first semantic information; and to analyze the sentence completeness of the second target sentence based on second semantic information; wherein the first target sentence and the second target sentence each include at least one of the following: the first sentence in the text, the last sentence in the text; the first semantic information and the second semantic information each include at least one of the following: sentence structure information, sentence component information, and phrase composition information.
[0161] Optionally, the processor 410 can be used to determine P texts from N texts based on the sentence completeness of the first target sentence, wherein the P texts correspond to the third translation.
[0162] Optionally, the processor 410 can also be used to obtain at least two texts that match the second text in the P texts from the P texts based on the first semantic information; and to merge the second text with the at least two texts respectively to obtain at least two merged texts; and to determine the third text corresponding to the target merged text as the text to be merged corresponding to the second text, wherein the target merged text is the text with the lowest sentence perplexity among the at least two merged texts; wherein the S texts include the second text and the third text.
[0163] Optionally, the processor 410 can be used to determine Q adjacent texts from the P texts based on the distribution position of each text in the P texts, where Q is an integer greater than or equal to S; and to determine S texts from the Q texts that satisfy the first semantic information as texts to be merged.
[0164] Optionally, the processor 410 can also be used to determine the target arrangement order of the S texts based on the first semantic information; and to merge the S texts according to the target arrangement order to obtain the first text.
[0165] This application provides an electronic device that, after acquiring multiple texts and target information from a target image, can merge at least one text that satisfies semantic information from the incomplete texts determined based on the target information to obtain a merged text. Therefore, when the text in the image includes complex text such as column text, paginated text, or irregularly shaped text, these complex texts can be merged according to semantic information. Furthermore, since the merged text and its corresponding translation are both complete texts before being output, the resulting merged text has a more fluent semantic meaning. This improves the processing capability of text in images.
[0166] It should be understood that, in this embodiment, the input unit 404 may include a graphics processing unit (GPU) 4041 and a microphone 4042. The GPU 4041 processes image data of still images or videos obtained by an image capture device (such as a camera) in video capture mode or image capture mode. The display unit 406 may include a display panel 4061, which may be configured in the form of a liquid crystal display, an organic light-emitting diode, or the like. The user input unit 407 includes at least one of a touch panel 4071 and other input devices 4072. The touch panel 4071 is also called a touch screen. The touch panel 4071 may include a touch detection device and a touch controller. Other input devices 4072 may include, but are not limited to, physical keyboards, function keys (such as volume control buttons, power buttons, etc.), trackballs, mice, and joysticks, which will not be described in detail here.
[0167] Memory 409 can be used to store software programs and various data. Memory 109 may mainly include a first storage area for storing programs or instructions and a second storage area for storing data. The first storage area may store the operating system, application programs or instructions required for at least one function (such as sound playback function, image playback function, etc.). Furthermore, memory 109 may include volatile memory or non-volatile memory, or memory x09 may include both volatile and non-volatile memory. The non-volatile memory may be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. Volatile memory can be random access memory (RAM), static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDRSDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM), and direct memory bus RAM (DRRAM). The memory 109 in the embodiments of this application includes, but is not limited to, these and any other suitable types of memory.
[0168] Processor 110 may include one or more processing units; optionally, processor 110 integrates an application processor and a modem processor, wherein the application processor mainly handles operations involving the operating system, user interface, and applications, and the modem processor mainly handles wireless communication signals, such as a baseband processor. It is understood that the aforementioned modem processor may also not be integrated into processor 110.
[0169] This application also provides a readable storage medium storing a program or instructions. When the program or instructions are executed by a processor, they implement the various processes of the above-described image processing method embodiments and achieve the same technical effect. To avoid repetition, they will not be described again here.
[0170] The processor is the processor in the electronic device described in the above embodiments. The readable storage medium includes computer-readable storage media, such as computer read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disk.
[0171] This application embodiment also provides a chip, which includes a processor and a communication interface. The communication interface and the processor are coupled. The processor is used to run programs or instructions to implement the various processes of the above-described image processing method embodiments and can achieve the same technical effect. To avoid repetition, it will not be described again here.
[0172] It should be understood that the chip mentioned in the embodiments of this application may also be referred to as a system-on-a-chip, system chip, chip system, or system-on-a-chip, etc.
[0173] This application provides a computer program product, which is stored in a storage medium and executed by at least one processor to implement the various processes of the above-described image processing method embodiments, and can achieve the same technical effect. To avoid repetition, it will not be described again here.
[0174] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element. Furthermore, it should be noted that the scope of the methods and apparatuses in the embodiments of this application is not limited to performing functions in the order shown or discussed, but may also include performing functions substantially simultaneously or in the reverse order, depending on the functions involved. For example, the described methods may be performed in a different order than described, and various steps may be added, omitted, or combined. Additionally, features described with reference to certain examples may be combined in other examples.
[0175] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a computer software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal (which may be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods of the various embodiments of this application.
[0176] The embodiments of this application have been described above with reference to the accompanying drawings. However, this application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of this application without departing from the spirit and scope of the claims, and all of these forms are within the protection scope of this application.
Claims
1. A picture processing method, characterized by, The method comprises: obtaining N texts and target information included in a target picture, the target information comprising at least one of first completeness of the N texts and second completeness of first translations corresponding to the N texts, N being an integer greater than 1; merging S texts meeting first semantic information in P texts to obtain a first text, the P texts being non-complete texts determined from the N texts according to the first completeness, P and S both being integers greater than 1; in a case where the first text and a second translation are both complete texts, outputting the first text, the second translation being a text obtained by merging translations corresponding to the S texts in third translations, the third translations being non-complete translations determined from the first translations according to the second completeness; the first completeness comprising sentence completeness of a first target sentence of each text in the N texts, and the second completeness comprising sentence completeness of a second target sentence of each translation in the first translations; the obtaining the N texts and the target information included in the target picture comprises: extracting texts included in the target picture to obtain the N texts; analyzing sentence completeness of the first target sentence based on the first semantic information; analyzing sentence completeness of the second target sentence based on second semantic information; wherein the first target sentence and the second target sentence each comprise at least one of a first sentence in a text and a last sentence in a text; the first semantic information and the second semantic information each comprise at least one of sentence pattern structure information, sentence component information and word group composition information.
2. The method of claim 1, wherein, After analyzing the sentence completeness of the first target sentence based on the first semantic information, and before merging S texts meeting first semantic information in P texts to obtain a first text, the method further comprises: determining the P texts from the N texts according to the sentence completeness of the first target sentence, the P texts corresponding to the third translations.
3. The method of claim 1, wherein, Before merging the S texts meeting first semantic information in the P texts to obtain a first text, the method further comprises: obtaining at least two texts matching a second text in the P texts from the P texts according to the first semantic information; merging the second text with the at least two texts respectively to obtain at least two merged texts; determining sentence perplexity of the at least two merged texts, the sentence perplexity being used to indicate fluency of a sentence in a merged text; determining a third text corresponding to a target merged text as a to-be-merged text corresponding to the second text, the target merged text being a text with the lowest sentence perplexity in the at least two merged texts; wherein the S texts comprise the second text and the third text.
4. The method of claim 1, wherein, Before merging the S texts meeting first semantic information in the P texts to obtain a first text, the method further comprises: determining Q texts with adjacent distribution positions from the P texts according to distribution positions of each text in the P texts, Q being an integer greater than or equal to S. The S texts in the Q texts satisfying the first semantic information are determined as to-be-merged texts.
5. The method of claim 4, wherein, After the S texts in the Q texts satisfying the first semantic information are determined as to-be-merged texts, the method further includes: According to the first semantic information, a target arrangement order of the S texts is determined. The S texts in the P texts satisfying the first semantic information are merged to obtain a first text, including: The S texts are merged according to the target arrangement order to obtain the first text.
6. An image processing apparatus characterized by comprising: The picture processing apparatus includes an acquisition module, a processing module, and an output module. The acquisition module is configured to acquire N texts included in a target picture and target information, the target information including at least one of first completeness of the N texts and second completeness of a first translation corresponding to the N texts, N being an integer greater than 1. The processing module is configured to merge S texts in P texts satisfying first semantic information to obtain a first text, the P texts being incomplete texts determined from the N texts according to the first completeness, P and S both being integers greater than 1. The output module is configured to output the first text in a case where the first text and a second translation are complete texts, the second translation being a text obtained by merging translations corresponding to the S texts in a third translation, the third translation being an incomplete translation determined from the first translation according to the second completeness. The first completeness includes sentence completeness of a first target sentence of each text in the N texts, and the second completeness includes sentence completeness of a second target sentence of each translation in the first translation. The acquisition module is specifically configured to extract texts included in the target picture to obtain the N texts, analyze sentence completeness of the first target sentence based on the first semantic information, and analyze sentence completeness of the second target sentence based on second semantic information. The first target sentence and the second target sentence each include at least one of a first sentence in a text and a last sentence in a text. The first semantic information and the second semantic information each include at least one of sentence pattern structure information, sentence component information, and word group composition information. The picture processing apparatus further includes a determination module.
7. The apparatus of claim 6, wherein, The determination module is configured to determine the P texts from the N texts according to the sentence completeness of the first target sentence, the P texts corresponding to the third translation. The picture processing apparatus further includes a determination module.
8. The apparatus of claim 6, wherein, The acquisition module is further configured to acquire at least two texts matching a second text in the P texts from the P texts according to the first semantic information. The processing module is further configured to merge the second text with the at least two texts respectively to obtain at least two merged texts. The determination module is configured to determine a third text corresponding to a target merged text as a to-be-merged text corresponding to the second text, the target merged text being a text with the lowest sentence perplexity in the at least two merged texts. The S texts include the second text and the third text.
9. The apparatus of claim 6, wherein, The picture processing apparatus further includes a determination module; The determination module is configured to determine, from the P texts, Q texts with adjacent distribution positions according to the distribution positions of each of the P texts, Q being an integer greater than or equal to S; and determine, as the to-be-merged texts, S texts from the Q texts that satisfy the first semantic information.
10. The apparatus of claim 9, wherein, The determination module is further configured to determine a target arrangement order of the S texts according to the first semantic information. The processing module is specifically configured to merge the S texts according to the target arrangement order to obtain the first text.
11. An electronic device, comprising: A processor and a memory are included, the memory stores programs or instructions executable on the processor, and the programs or instructions, when executed by the processor, implement the steps of the picture processing method according to any one of claims 1-5.
12. A readable storage medium, characterized by, The programs or instructions are stored on the readable storage medium, and when executed by the processor, implement the steps of the picture processing method according to any one of claims 1-5.
Citation Information
Patent Citations
Translation method and device, electronic equipment and storage medium
CN112711954A
Subtitle translation method and device and device for subtitle translation
CN113343720A
Method for translating words in a picture, electronic device, and storage medium
US20210272342A1