Image processing method and device, storage medium and program product
By training the recognition model by considering the contextual relationship in the image during the image translation process, the problem of mistranslation in image translation is solved, and higher translation accuracy and user experience are achieved.
Patent Information
- Application Number
- CN202410289287.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-03-13
- Publication Date
- 2025-09-16
Smart Images

Figure CN120656182A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of image processing technology, and in particular to an image processing method, device, storage medium, and program product. Background Art
[0002] In the image translation scenario, the service provider can perform text recognition on the image provided by the user and translate the text in the image into the language selected by the user. However, in some scenarios, translation errors may occur, which reduces the user experience. Summary of the Invention
[0003] To overcome the problems existing in the related art, the present disclosure provides an image processing method, apparatus, storage medium and program product.
[0004] According to a first aspect of an embodiment of the present disclosure, there is provided an image processing method, including:
[0005] Obtain the character string in the target image to obtain the first character string;
[0006] Inputting the first character string into a recognition model to obtain a second character string output by the recognition model and obtained by translating the first character string;
[0007] The recognition model is trained by taking training data as input and using translation results of training strings in the training data as at least part of the output, wherein the training data includes first sample data, the first sample data is generated based on a sample image, the first sample data includes a sample string in the sample image, and a contextual relationship of the sample string in the sample image.
[0008] Optionally include:
[0009] determining a plurality of regions in the image to be processed;
[0010] For each of the plurality of regions, obtaining a character string in the region to obtain a third character string;
[0011] For each of the regions, when the third character string in any of the regions includes a set ambiguous character string, the image to be processed is used as the target image.
[0012] Optionally, when the third character string in any of the regions includes a set ambiguous character string, taking the image to be processed as the target image includes:
[0013] When the third character string in any of the regions includes a set ambiguous character string, obtaining the number of character strings in the region;
[0014] When the number of the character strings is less than a first threshold, the image to be processed is used as the target image.
[0015] Optionally, the first threshold is greater than 1, and when the number of the character strings is less than the first threshold, taking the image to be processed as the target image includes:
[0016] When the number of the character strings is 1, the image to be processed is used as the target image.
[0017] Optionally include:
[0018] Get the character string in the sample image to obtain the sample character string;
[0019] Obtaining a contextual relationship of the sample character string in the sample image;
[0020] determining a translation result of the sample character string, wherein the translation result is obtained by translating part or all of the sample character string into a target language;
[0021] generating a first instruction, wherein the first instruction is used to instruct to translate the sample character string into the target language;
[0022] The first sample data is generated according to the sample character string, the context relationship, the translation result, and the first instruction.
[0023] Optionally, determining the translation result of the sample character string includes:
[0024] generating a second instruction, wherein the second instruction is used to instruct to translate part or all of the sample character string into a target language based on the sample character string and a contextual relationship of the sample character string in the sample image;
[0025] The sample character string, the context relationship, and the first instruction are input into a translation model to obtain the translation result output by the translation model.
[0026] Optionally include:
[0027] Obtaining a string pair, the string pair comprising an original string and a target string obtained by translating the original string into a target language;
[0028] generating a third instruction, wherein the third instruction is used to instruct to translate the original character string into the target language;
[0029] Second sample data is generated according to the string pair and the third instruction, and the training data includes the second sample data.
[0030] Optionally include:
[0031] Obtaining a string pair, the string pair comprising an original string and a target string obtained by translating the original string into a target language;
[0032] Obtaining a prompt string pair, the prompt string pair including an original prompt string and a target prompt string obtained by translating the original prompt string into the target language, wherein a similarity between the original prompt string and the original string is greater than a second threshold;
[0033] generating a fourth instruction, the fourth instruction being used to instruct translation of the original character string into a target language with reference to the prompt character string pair;
[0034] Third sample data is generated according to the string pair, the prompt string pair, and the fourth instruction, and the training data includes the third sample data.
[0035] According to a second aspect of an embodiment of the present disclosure, there is provided an image processing apparatus, including:
[0036] The first module is configured to obtain a character string in a target image to obtain a first character string;
[0037] A second module is configured to input the first character string into a recognition model, and obtain a second character string output by the recognition model and obtained by translating the first character string;
[0038] The recognition model is trained by taking training data as input and using translation results of training strings in the training data as at least part of the output, wherein the training data includes first sample data, the first sample data is generated based on a sample image, the first sample data includes a sample string in the sample image, and a contextual relationship of the sample string in the sample image.
[0039] Optionally, the device comprises:
[0040] A third module is configured to determine a plurality of regions in the image to be processed;
[0041] A fourth module is configured to obtain a character string in each of the plurality of regions to obtain a third character string;
[0042] The fifth module is configured to, for each of the regions, use the image to be processed as the target image when the third character string in any of the regions includes a set ambiguous character string.
[0043] Optionally, the fifth module includes:
[0044] The first submodule is configured to obtain the number of character strings in the area when a third character string in any area includes a set ambiguous character string;
[0045] The second submodule is configured to use the image to be processed as the target image when the number of the character strings is less than a first threshold.
[0046] Optionally, the first threshold is greater than 1, and the second submodule is configured to:
[0047] When the number of the character strings is 1, the image to be processed is used as the target image.
[0048] Optionally, the device comprises:
[0049] A sixth module is configured to obtain a character string in a sample image to obtain a sample character string;
[0050] A seventh module is configured to obtain a contextual relationship of the sample character string in the sample image;
[0051] an eighth module, configured to determine a translation result of the sample character string, wherein the translation result is obtained by translating part or all of the sample character string into a target language;
[0052] A ninth module is configured to generate a first instruction, wherein the first instruction is used to instruct to translate the sample character string into the target language;
[0053] A tenth module is configured to generate the first sample data according to the sample character string, the context relationship, the translation result and the first instruction.
[0054] Optionally, the eighth module includes:
[0055] A third submodule is configured to generate a second instruction, wherein the second instruction is used to instruct to translate part or all of the sample character string into a target language based on the sample character string and a contextual relationship of the sample character string in the sample image;
[0056] The fourth submodule is configured to input the sample character string, the context relationship, and the first instruction into a translation model to obtain the translation result output by the translation model.
[0057] Optionally, the device comprises:
[0058] an eleventh module configured to obtain a string pair, wherein the string pair includes an original string and a target string obtained by translating the original string into a target language;
[0059] A twelfth module is configured to generate a third instruction, wherein the third instruction is used to instruct to translate the original character string into the target language;
[0060] The thirteenth module is configured to generate second sample data according to the string pair and the third instruction, and the training data includes the second sample data.
[0061] Optionally, the device comprises:
[0062] A fourteenth module is configured to obtain a string pair, wherein the string pair includes an original string and a target string obtained by translating the original string into a target language;
[0063] A fifteenth module is configured to obtain a prompt string pair, the prompt string pair comprising an original prompt string and a target prompt string obtained by translating the original prompt string into the target language, wherein a similarity between the original prompt string and the original string is greater than a second threshold;
[0064] A sixteenth module is configured to generate a fourth instruction, wherein the fourth instruction is used to instruct to translate the original character string into a target language with reference to the prompt character string pair;
[0065] The seventeenth module is configured to generate third sample data according to the string pair, the prompt string pair and the fourth instruction, and the training data includes the third sample data.
[0066] According to a third aspect of an embodiment of the present disclosure, there is provided an image processing apparatus, including:
[0067] processor;
[0068] a memory for storing processor-executable instructions;
[0069] The processor is configured to execute the method described in any one of the first aspects.
[0070] According to a fourth aspect of an embodiment of the present disclosure, a computer-readable storage medium is provided, on which computer program instructions are stored. When the program instructions are executed by a processor, the method according to any one of the first aspects is implemented.
[0071] According to a fifth aspect of an embodiment of the present disclosure, a computer program product is provided, comprising a computer program, which implements the steps of any one of the methods in the first aspect when executed by a processor.
[0072] In the above solution, a recognition model can be trained by using training data as input and using translation results of training strings in the training data as at least a portion of the output. The training data includes first sample data generated based on a sample image, including a sample string in the sample image and the context of the sample string in the sample image. Because the sample data includes the context of the sample string in the sample image, the recognition model trained based on the sample data can also take into account the context of the string in the image when performing translation tasks, thereby helping to improve translation accuracy.
[0073] In this way, after obtaining a first character string in a target image, the first character string can be input into a recognition model to obtain a second character string translated from the first character string as output by the recognition model. In this way, the second character string translated by the recognition model has a higher accuracy.
[0074] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0075] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present disclosure and, together with the description, serve to explain the principles of the present disclosure.
[0076] Figure 1 is a schematic diagram showing an image processing scenario according to an exemplary embodiment.
[0077] Figure 2 is a schematic diagram showing an image processing scenario according to an exemplary embodiment.
[0078] Figure 3 The figure is a flowchart of an image processing method according to an exemplary embodiment.
[0079] Figure 4 The figure is a flowchart of an image processing method according to an exemplary embodiment.
[0080] Figure 5 The figure is a flowchart of obtaining first sample data according to an exemplary embodiment.
[0081] Figure 6 The figure is a schematic diagram showing an image including an English character string according to an exemplary embodiment.
[0082] Figure 7 The figure is a flowchart of obtaining first sample data according to an exemplary embodiment.
[0083] Figure 8 is a schematic diagram showing an image processing scenario according to an exemplary embodiment.
[0084] Figure 9 is a block diagram of an image processing apparatus according to an exemplary embodiment.
[0085] Figure 10 The figure is a block diagram showing an apparatus for image processing according to an exemplary embodiment. DETAILED DESCRIPTION
[0086] Exemplary embodiments will be described in detail herein, with examples illustrated in the accompanying drawings. In the following description, when referring to the drawings, identical numerals in different figures represent identical or similar elements, unless otherwise indicated. The embodiments described in the following exemplary embodiments are not intended to represent all possible embodiments consistent with the present disclosure. Rather, they are merely examples of apparatus and methods consistent with certain aspects of the present disclosure, as detailed in the appended claims.
[0087] The embodiments described in the following examples of the present disclosure do not represent all embodiments consistent with the present disclosure. Instead, they are merely examples of devices and methods consistent with some aspects of the present disclosure as detailed in the appended claims.
[0088] Before introducing the image processing method, device, storage medium, and program product of the present disclosure, relevant scenarios of the embodiments of the present disclosure are first introduced.
[0089] In image translation tasks, service providers can perform text recognition on user-provided images and translate the text into the user's selected language. Image translation tasks can be applied to scenarios such as tourism, education, business, and cultural exchange. For example, while traveling, tourists may encounter objects containing English text, such as signs, menus, and road signs. In these cases, tourists may need to translate the English text in these objects.
[0090] Figure 1 This is a schematic diagram of an image processing scenario shown in an exemplary embodiment of the present disclosure, in which a menu including English text is used as an example. Visitors can use mobile phones or other terminals to capture images of the menu. For the captured image 100, the service provider can process and translate the English text in the image 100. Figure 1 In some scenarios, the English text in image 100 can be erased and the translation result can be backfilled into the image after the English text is erased, thereby obtaining image 101 including the Chinese translation. In this way, tourists can be assisted in understanding the menu.
[0091] It is worth noting that English words may have multiple meanings, so mistranslation may occur in some scenarios. For example, Figure 2 FIG. 1 is a schematic diagram of an image processing scenario according to an exemplary embodiment of the present disclosure. Figure 2 , wherein during translation, "Quarter" in the identification box 201 is translated as "a quarter", and "Share" in the identification box 202 is translated as "sharing". Such incorrect translation may reduce the user experience.
[0092] To this end, embodiments of the present disclosure provide an image processing method that can be applied to various stationary or mobile computing devices, such as mobile phones, tablet devices, wearable devices, and the like. Figure 3 is a flowchart of an image processing method shown in an exemplary embodiment of the present disclosure, with reference to Figure 3 , the method comprising:
[0093] In step S31, a character string in a target image is obtained to obtain a first character string;
[0094] In step S32, the first character string is input into a recognition model, and a second character string is obtained by translating the first character string and output by the recognition model.
[0095] The recognition model is trained by taking training data as input and using translation results of training strings in the training data as at least part of the output, wherein the training data includes first sample data, the first sample data is generated based on a sample image, the first sample data includes a sample string in the sample image, and a contextual relationship of the sample string in the sample image.
[0096] For example, in a possible implementation, an image captured by a terminal may be used as the target image.
[0097] In one implementation, an image selected by the user may also be used as the target image.
[0098] In one embodiment, considering that relatively more resources are required to obtain the second character string through the recognition model, the image to be processed may be judged to determine whether the image to be processed needs to be used as the target image.
[0099] Figure 4 is a flowchart of an image processing method shown in an exemplary embodiment of the present disclosure. In one embodiment, the method Figure 3 On the basis of, before step S31, it also includes:
[0100] In step S301 , a plurality of regions in an image to be processed are determined.
[0101] As an example, in a scenario where an image is processed line by line, each line in the image to be processed can be considered as a region. For example, an image can be processed using OCR (Optical Character Recognition) technology. When processing an image using OCR technology, the image can be processed line by line. Therefore, each line can be considered as a region.
[0102] In addition, in an optional implementation, the image may be divided into multiple regions based on needs. The embodiment of the present disclosure does not limit the form and number of the regions.
[0103] In step S302 , for each of the multiple regions, a character string in the region is obtained to obtain a third character string.
[0104] Continuing with the above example, the character string in each of the regions may be extracted using OCR technology to obtain a third character string.
[0105] In step S303 , for each region, if the third character string in any region includes a set ambiguous character string, the image to be processed is used as a target image.
[0106] For example, in one embodiment, an ambiguous string set may be provided, which may include one or more ambiguous strings. Ambiguous strings may include potentially ambiguous strings, such as potentially ambiguous words (e.g., the English words cover, share, discount, etc.) and symbol abbreviations (e.g., hp, CNG, etc.). Ambiguous strings may be manually set or may be strings whose translation error rate, as determined by machine statistics, is greater than a set threshold.
[0107] Thus, in step S303, for each region, it is possible to detect whether the third character string in the region includes any ambiguous character string. If the third character string in any region includes an ambiguous character string, it can be determined that the ambiguous character string in the region is prone to mistranslation. Therefore, the image to be processed can be used as the target image and image processing can be performed using the recognition model.
[0108] It should be noted that, in the case where the region includes more context information, ambiguous character strings may also be correctly translated. Therefore, in the scenario where the region includes more context information, calling the recognition model for translation may lead to increased resource consumption.
[0109] To this end, in a possible implementation, when the third character string in any of the regions includes a set ambiguous character string, using the image to be processed as the target image includes:
[0110] When the third character string in any of the regions includes a set ambiguous character string, obtaining the number of character strings in the region;
[0111] If the number of the character strings is less than a first threshold, the image to be processed is used as the target image. The value of the first threshold can be set according to requirements and is not limited in the embodiment of the present disclosure.
[0112] The above solution can obtain the number of strings in any region if the third string in any region includes a predetermined ambiguous string. If the number of strings is less than a first threshold, it can be determined that the region contains insufficient contextual information, which may result in an incorrect translation of the ambiguous string. Therefore, the image to be processed can be used as the target image.
[0113] As an example, the first threshold is greater than 1, and when the number of the character strings is less than the first threshold, the image to be processed is used as the target image, including:
[0114] When the number of the character strings is 1, the image to be processed is used as the target image.
[0115] Reference Figure 2 To illustrate, when using OCR technology to process an original image, text within the image can be recognized row by row, meaning each row can be considered a region. The region containing identification box 202 contains the string "Share," and the number of strings is 1, which is less than the first threshold. Thus, without contextual information, the string "Share" may be mistranslated. Therefore, this image can be used as the target image.
[0116] For a target image, a first character string in the target image may be obtained, and the first character string is input into a recognition model to obtain a second character string output by the recognition model, which is obtained by translating the first character string.
[0117] Because the sample data includes the contextual relationship of the sample character string within the sample image, the recognition model trained based on the sample data can also take the contextual relationship of the character string within the image into account when performing translation tasks, thereby helping to improve translation accuracy. As a result, the second character string translated using the recognition model also has higher accuracy.
[0118] The following is an exemplary description of the training method of the recognition model.
[0119] In one embodiment, when training the recognition model, training data of the recognition model can be obtained, such as the first sample data in the training data.
[0120] Figure 5 is a flowchart for obtaining the first sample data shown in an exemplary embodiment of the present disclosure. Referring to Figure 5 , the first sample data can be obtained through the following method:
[0121] In step S51, the string in the sample image is obtained to get the sample string.
[0122] In one embodiment, the string in the sample image can be recognized by OCR technology to get the sample string.
[0123] In one embodiment, the string in the sample image can also be manually extracted to get the sample string.
[0124] In step S52, the context relationship of the sample string in the sample image is obtained.
[0125] The context relationship of the sample string in the sample image can include the sequence relationship, up and down position relationship, line number order relationship, etc. of different sample strings in the image. In specific implementation, the description method of the context relationship of the sample string in the sample image can be determined based on requirements. As an example, the context relationship of the sample string in the sample image can be described by line break characters.
[0126] Exemplarily, Figure 6 is a schematic diagram of an image including English strings shown in an exemplary embodiment of the present disclosure. When recognizing by OCR, each line includes a string. It is likely to cause translation errors. For example, for "Power", due to the lack of context information, it may be translated as "power supply".
[0127] In the above solution, the context relationship of the sample string in the sample image can be obtained. For example, the context relationship of the sample string in the sample image can be described by line break characters. In this way, after adding the context relationship, Figure 6 the sample string in
[0128] can be expressed as:
[0129] In this way, context information can be provided for the string "Power", which helps to obtain more accurate translation results.
[0130] Reference Figure 5 In step S53, the translation result of the sample character string is determined, and the translation result is obtained by translating part or all of the sample character string into the target language.
[0131] As an example, determining the translation result of the sample character string includes:
[0132] generating a second instruction, wherein the second instruction is used to instruct to translate part or all of the sample character string into a target language based on the sample character string and a contextual relationship of the sample character string in the sample image;
[0133] The sample character string, the context relationship, and the first instruction are input into a translation model to obtain the translation result output by the translation model.
[0134] Continued use Figure 6 For example, according to the sample character string and the context relationship, the sample character string can be represented as:
[0135] "\nDamage\nHealth\nEpic\nAttack\nStack\nCastle\nPRIEST\nSpeed\nRanged\nBlessing\nFighter\nLevel\nDefense\nPower\n".
[0136] In addition, a second instruction may be generated, which may be, for example, "translate the word power according to the context", "translate all the words according to the context", etc.
[0137] In this way, the second instruction and "\nDamage\nHealth\nEpic\nAttack\nStack\nCastle\nPRIEST\nSpeed\nRanged\nBlessing\nFighter\nLevel\nDefense\nPower\n" generated according to the sample character string can be input into the translation model to obtain the translation result output by the translation model.
[0138] In one embodiment, the input of the translation model can also be generated in JSON format. Continuing with the above example, the input data can be generated:
[0139] {"instruction":"The following text comes from the same image:\nDamage\nHealth\nEpic\nAttack\nStack\nCastle\nPRIEST\nSpeed\nRanged\nBl essing\nFighter\nLevel\nDefense\nPower\n. Translate the power into Chinese.","input":"","output":""}
[0140] In this way, the input data can be input into the translation model to obtain the translation result of "power" output by the translation model. The translation model can be various models, such as various general large models.
[0141] In this way, after extracting sample strings from sample images, the above solution can quickly obtain translation results for part or all of the sample strings with the help of the translation model. This helps to quickly build training data for the model.
[0142] In addition, in some implementations, the sample character strings may also be translated manually to obtain translation results.
[0143] In step S54 , a first instruction is generated, where the first instruction is used to instruct the translation of the sample character string into the target language.
[0144] As an example, the first instruction may be an instruction based on natural language, such as “Please translate the text content in the following image: XXXXXXX, and translate YYY into Chinese.”
[0145] In step S55 , first sample data is generated according to the sample character string, the context, the translation result, and the first instruction.
[0146] Continued use Figure 6 For example, the first sample data may be:
[0147] {"instruction":"Please translate the text in the following images:\nDamage\nHealth\nEpic\nAttack\nStack\nCastle\nPRIEST\nSpeed\nRanged\nBl essing\nFighter\nLevel\nDefense\nPower\n, and translate the word "power" into Chinese.","input":"","output":"Power"}
[0148] In addition, the above Figure 6The step of processing to obtain the first sample data can be repeated. Therefore, the number of first sample data can also be multiple. Figure 7 This is a flowchart of obtaining first sample data shown in an exemplary embodiment of the present disclosure, referring to Figure 7 , you can select an image from the image library. For the selected image, you can use OCR to recognize the text in the image. If each area in the image does not include an ambiguous string, you can reselect the image. In the case that an area in the image includes an ambiguous string, you can use the translation model to obtain the translation of the text in the image and generate the first sample data. The method of generating the first sample data can refer to the method about Figure 5 For the sake of brevity, the present disclosure does not elaborate on this embodiment. In addition, the first sample data can also be added to the sample pool. In this way, the image can be reselected to generate the first sample data until the first sample data exceeds the sample quantity threshold (23K is used as an example in the figure).
[0149] After obtaining the first sample data, the recognition model can be trained based on the first sample data. As an example, a pre-trained model can be obtained and fine-tuned using the first sample data to obtain the recognition model.
[0150] For example, the BLOOM-7b1-mt multi-language translation model architecture can be used as a base model and fine-tuned using the first sample data. For example, the base model can be fine-tuned using techniques such as full-shard data parallelism and mixed-precision training to obtain the recognition model.
[0151] In the above solution, a recognition model can be trained by using training data as input and using translation results of training strings in the training data as at least a portion of the output. The training data includes first sample data generated based on a sample image, including a sample string in the sample image and the context of the sample string in the sample image. Because the sample data includes the context of the sample string in the sample image, the recognition model trained based on the sample data can also take into account the context of the string in the image when performing translation tasks, thereby helping to improve translation accuracy.
[0152] In some implementations, other sample data may be added to the training data to improve the translation capability of the recognition model. For example, in one possible implementation, the method includes:
[0153] Obtaining a string pair, the string pair comprising an original string and a target string obtained by translating the original string into a target language;
[0154] generating a third instruction, wherein the third instruction is used to instruct to translate the original character string into the target language;
[0155] Second sample data is generated according to the string pair and the third instruction, and the training data includes the second sample data.
[0156] For example, if the source language is Chinese and the target language is English, we can collect bilingual parallel sentence pairs. For example, the source string is "On June 1, 2018, in Cardiff, the United States, the corporate identity on XX electric cars." The target string is "On June 1, 2018, in Cardiff, the United States, the corporate identity on XX electric cars."
[0157] Furthermore, a third instruction may be generated, instructing the translation of the original string into the target language. For example, the third instruction may be "Translate the following sentence from English to Chinese: XXX." Thus, second sample data may be generated based on the string pair and the third instruction, and the training data may include the second sample data.
[0158] Continuing with the above example, the second sample data may be:
[0159] {"instruction":"Translate the following sentence from English to Chinese: On June 1, 2018, in Cardiff, the United States, the corporate identity on XX electric cars.","input":"","output":"On June 1, 2018, in Cardiff, the United States, the corporate identity on XX electric cars."}
[0160] The number of second sample data can be set based on demand. For example, 165K second sample data can be generated. The second sample data can be used as zero-shot to help achieve zero-shot learning of the model.
[0161] In some implementations, other sample data may be added to the training data to improve the translation capability of the recognition model. For example, in one possible implementation, the method includes:
[0162] Obtain a string pair, the string pair including an original string and a target string obtained by translating the original string into a target language. For example, if the original language is Chinese and the target language is English, Chinese-English parallel sentence pairs can be collected.
[0163] In addition, a prompt string pair may be obtained, the prompt string pair including an original prompt string and a target prompt string obtained by translating the original prompt string into the target language, wherein the similarity between the original prompt string and the original string is greater than a second threshold.
[0164] For example, the original prompt string may be “To push the toilet revolution deeper is to think more about what the masses think”, and the target prompt string may be “To push the toilet revolution deeper is to think more about what the masses think”.
[0165] In addition, a fourth instruction may be generated, wherein the fourth instruction is used to instruct to translate the original character string into a target language with reference to the prompt character string pair.
[0166] Continuing with the above example, the fourth instruction may be:
[0167] "To push the toilet revolution deeper is to think more about what the masses think." is a Chinese translation of "To push the toilet revolution deeper is to think more about what the masses think." Please translate: XXXXXYXYYYX.
[0168] In this way, third sample data may be generated according to the string pair, the prompt string pair, and the fourth instruction, and the training data includes the third sample data.
[0169] As an example, the third sample data may be:
[0170] {"instruction":"\"To push the toilet revolution deeper is to think more about what the masses think." is a Chinese translation of \"To push the toilet revolution deeper is to think more about what the masses think.". The English translation of \"input":"","output":"and to push the revolution in toilets further,we should think more of what the masses think."}
[0171] The number of third sample data can be set based on demand. For example, 200K third sample data can be generated. The third sample data (i.e., one-shot) helps to achieve few-sample learning of the model.
[0172] In the above solution, a recognition model can be trained by using training data as input and using translation results of training strings in the training data as at least a portion of the output. The training data includes first sample data generated based on a sample image, including a sample string in the sample image and the context of the sample string in the sample image. Because the sample data includes the context of the sample string in the sample image, the recognition model trained based on the sample data can also take into account the context of the string in the image when performing translation tasks, thereby helping to improve translation accuracy.
[0173] After the recognition model is trained, it can be tested. For example, a test set can be constructed, which includes 100 images, each of which includes at least one region containing only ambiguous character strings.
[0174] This approach allows for the extraction of character strings from images in the test set, which are then fed into the recognition model to generate translations of the strings. Testing has shown that this approach achieves a translation accuracy of 72.85%, an improvement of approximately 27% compared to direct translation in related scenarios.
[0175] For example, refer to Figure 8 The schematic diagram of an image processing scenario shown above can translate the text in the identification boxes 201 and 202 by combining the context of the text in the image. Figure 2 In the example above, "Quarter" in the marker box 201 is translated as "quarter" and "Share" in the marker box 202 is translated as "share". This approach can translate "Quarter" in the marker box 201 as "quarter" and "Share" in the marker box 202 as "share" based on the context. This approach achieves higher translation accuracy.
[0176] Based on the same inventive concept, an embodiment of the present disclosure provides an image processing device. Figure 9 is a block diagram of an image processing device shown in an exemplary embodiment of the present disclosure, with reference to Figure 9 , the image processing device includes:
[0177] The first module 901 is configured to obtain a character string in a target image to obtain a first character string;
[0178] The second module 902 is configured to input the first character string into a recognition model, and obtain a second character string output by the recognition model, which is obtained by translating the first character string;
[0179] The recognition model is trained by taking training data as input and using translation results of training strings in the training data as at least part of the output, wherein the training data includes first sample data, the first sample data is generated based on a sample image, the first sample data includes a sample string in the sample image, and a contextual relationship of the sample string in the sample image.
[0180] In the above solution, a recognition model can be trained by using training data as input and using translation results of training strings in the training data as at least a portion of the output. The training data includes first sample data generated based on a sample image, including a sample string in the sample image and the context of the sample string in the sample image. Because the sample data includes the context of the sample string in the sample image, the recognition model trained based on the sample data can also take into account the context of the string in the image when performing translation tasks, thereby helping to improve translation accuracy.
[0181] In this way, after obtaining a first character string in a target image, the first character string can be input into a recognition model to obtain a second character string translated from the first character string as output by the recognition model. In this way, the second character string translated by the recognition model has a higher accuracy.
[0182] Optionally, the device comprises:
[0183] A third module is configured to determine a plurality of regions in the image to be processed;
[0184] A fourth module is configured to obtain a character string in each of the plurality of regions to obtain a third character string;
[0185] The fifth module is configured to, for each of the regions, use the image to be processed as the target image when the third character string in any of the regions includes a set ambiguous character string.
[0186] Optionally, the fifth module includes:
[0187] The first submodule is configured to obtain the number of character strings in the area when a third character string in any area includes a set ambiguous character string;
[0188] The second submodule is configured to use the image to be processed as the target image when the number of the character strings is less than a first threshold.
[0189] Optionally, the first threshold is greater than 1, and the second submodule is configured to:
[0190] When the number of the character strings is 1, the image to be processed is used as the target image.
[0191] Optionally, the device comprises:
[0192] A sixth module is configured to obtain a character string in a sample image to obtain a sample character string;
[0193] A seventh module is configured to obtain a contextual relationship of the sample character string in the sample image;
[0194] an eighth module, configured to determine a translation result of the sample character string, wherein the translation result is obtained by translating part or all of the sample character string into a target language;
[0195] A ninth module is configured to generate a first instruction, wherein the first instruction is used to instruct to translate the sample character string into the target language;
[0196] A tenth module is configured to generate the first sample data according to the sample character string, the context relationship, the translation result and the first instruction.
[0197] Optionally, the eighth module includes:
[0198] A third submodule is configured to generate a second instruction, wherein the second instruction is used to instruct to translate part or all of the sample character string into a target language based on the sample character string and a contextual relationship of the sample character string in the sample image;
[0199] The fourth submodule is configured to input the sample character string, the context relationship, and the first instruction into a translation model to obtain the translation result output by the translation model.
[0200] Optionally, the device comprises:
[0201] an eleventh module configured to obtain a string pair, wherein the string pair includes an original string and a target string obtained by translating the original string into a target language;
[0202] A twelfth module is configured to generate a third instruction, wherein the third instruction is used to instruct to translate the original character string into the target language;
[0203] The thirteenth module is configured to generate second sample data according to the string pair and the third instruction, and the training data includes the second sample data.
[0204] Optionally, the device comprises:
[0205] A fourteenth module is configured to obtain a string pair, wherein the string pair includes an original string and a target string obtained by translating the original string into a target language;
[0206] A fifteenth module is configured to obtain a prompt string pair, the prompt string pair comprising an original prompt string and a target prompt string obtained by translating the original prompt string into the target language, wherein a similarity between the original prompt string and the original string is greater than a second threshold;
[0207] A sixteenth module is configured to generate a fourth instruction, wherein the fourth instruction is used to instruct to translate the original character string into a target language with reference to the prompt character string pair;
[0208] The seventeenth module is configured to generate third sample data according to the string pair, the prompt string pair and the fourth instruction, and the training data includes the third sample data.
[0209] An embodiment of the present disclosure provides an image processing device, including:
[0210] processor;
[0211] a memory for storing processor-executable instructions;
[0212] The processor is configured to execute the image processing method provided in any practical example of the present disclosure.
[0213] An embodiment of the present disclosure provides a computer-readable storage medium having computer program instructions stored thereon. When the program instructions are executed by a processor, the image processing method provided in any practical example of the present disclosure is implemented.
[0214] Regarding the apparatus in the above embodiment, the specific manner in which each module performs operations has been described in detail in the embodiment of the method, and will not be elaborated here.
[0215] Figure 10 FIG. 8 is a block diagram of an apparatus 800 for image processing according to an exemplary embodiment. For example, the apparatus 800 may be a mobile phone, a tablet device, or the like.
[0216] Reference Figure 10 , the apparatus 800 may include one or more of the following components: a processing component 802 , a memory 804 , a power component 806 , a multimedia component 808 , an audio component 810 , an input / output interface 812 , a sensor component 814 , and a communication component 816 .
[0217] The processing component 802 generally controls the overall operation of the device 800, such as operations associated with display, phone calls, data communications, camera operation, and recording operations. The processing component 802 may include one or more processors 820 to execute instructions to perform all or part of the steps of the image processing method described above. In addition, the processing component 802 may include one or more modules to facilitate interaction between the processing component 802 and other components. For example, the processing component 802 may include a multimedia module to facilitate interaction between the multimedia component 808 and the processing component 802.
[0218] The memory 804 is configured to store various types of data to support the operations of the device 800. Examples of such data include instructions for any application or method operating on the device 800, contact data, phone book data, messages, pictures, videos, etc. The memory 804 can be implemented by any type of volatile or non-volatile storage device, or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk, or optical disk.
[0219] The power supply component 806 provides power to the various components of the device 800. The power supply component 806 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power to the device 800.
[0220] The multimedia component 808 includes a screen that provides an output interface between the device 800 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen can be implemented as a touch screen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, slides, and gestures on the touch panel. The touch sensor can not only sense the boundaries of the touch or slide action, but also detect the duration and pressure associated with the touch or slide operation. In some embodiments, the multimedia component 808 includes a front camera and / or a rear camera. When the device 800 is in an operating mode, such as a shooting mode or a video mode, the front camera and / or the rear camera can receive external multimedia data. Each front camera and rear camera can be a fixed optical lens system or have focal length and optical zoom capabilities.
[0221] The audio component 810 is configured to output and / or input audio signals. For example, the audio component 810 includes a microphone (MIC), which is configured to receive external audio signals when the device 800 is in an operating mode, such as a call mode, a recording mode, and a voice recognition mode. The received audio signal can be further stored in the memory 804 or transmitted via the communication component 816. In some embodiments, the audio component 810 also includes a speaker for outputting audio signals.
[0222] The input / output interface 812 provides an interface between the processing component 802 and peripheral interface modules, such as a keyboard, a click wheel, buttons, etc. These buttons may include but are not limited to: a home button, a volume button, a start button, and a lock button.
[0223] The sensor assembly 814 includes one or more sensors for providing various aspects of the status assessment of the device 800. For example, the sensor assembly 814 can detect the open / closed state of the device 800, the relative positioning of components, such as the display and keypad of the device 800. The sensor assembly 814 can also detect changes in the position of the device 800 or a component of the device 800, the presence or absence of user contact with the device 800, the orientation or acceleration / deceleration of the device 800, and temperature changes of the device 800. The sensor assembly 814 may include a proximity sensor configured to detect the presence of nearby objects without any physical contact. The sensor assembly 814 may also include an optical sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, the sensor assembly 814 may also include an accelerometer, a gyroscope, a magnetic sensor, a pressure sensor, or a temperature sensor.
[0224] The communication component 816 is configured to facilitate wired or wireless communication between the device 800 and other devices. The device 800 can access a wireless network based on a communication standard, such as WiFi, 4G or 5G, or a combination thereof. In an exemplary embodiment, the communication component 816 receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component 816 also includes a near field communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on radio frequency identification (RFID) technology, infrared data association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology and other technologies.
[0225] In an exemplary embodiment, the apparatus 800 may be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components to perform the above-described image processing method.
[0226] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions is also provided, such as a memory 804 including instructions. The instructions can be executed by the processor 820 of the apparatus 800 to perform the above-described image processing method. For example, the non-transitory computer-readable storage medium can be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, an optical data storage device, and the like.
[0227] In another exemplary embodiment, a computer program product is further provided. The computer program product includes a computer program executable by a programmable device, and has a code portion for performing the above-mentioned image processing method when executed by the programmable device.
[0228] In the above detailed description, reference is made to the accompanying drawings, which illustrate specific aspects of the present disclosure that can be practiced. It should be understood that other aspects can be utilized and structural or logical changes can be made without departing from the concepts of the present disclosure. Therefore, the following detailed description should not be regarded as limiting.
[0229] It should be understood that, unless otherwise specifically noted, the features of the various embodiments of the present disclosure described herein may be combined with each other. As used herein, the term "and / or" includes any one of the relevant listed items and any combination of any two or more thereof; similarly, "at least one of" includes any one of the relevant listed items and any combination of any two or more thereof.
[0230] It should be understood that, unless otherwise specified and limited, the terms used in the embodiments of the present disclosure should be understood in a broad sense. Unless otherwise clearly defined, those skilled in the art can understand the specific meanings of the above terms in this article according to the specific circumstances.
[0231] In addition, although terms such as "first" and "second" may be used herein to describe various modules, etc., these modules are not limited to these terms. On the contrary, these terms are only used to distinguish one module from another. Therefore, without departing from the teachings of the various examples, the first module mentioned in the examples described herein may also be referred to as the second module. In addition, the terms "first" and "second" are used for descriptive purposes only and cannot be understood as indicating or implying relative importance or implicitly indicating the number of technical features indicated. Thus, the features defined as "first" and "second" may explicitly or implicitly include at least one of the features. In the description herein, the meaning of "multiple" is at least two, such as two, three, etc., unless otherwise clearly and specifically defined.
[0232] In addition, the word "exemplary" is used herein to indicate serving as an example, instance, or diagram. Any aspect or design described herein as "exemplary" is not necessarily to be understood as advantageous over other aspects or designs. On the contrary, the use of the word exemplary is intended to present concepts in a concrete manner. As used herein, the term "or" is intended to mean an inclusive "or" rather than an exclusive "or." In addition, unless otherwise specified or clearly directed to a singular form from the context, the articles "a" and "an" as used in this application and the appended claims are generally understood to mean "one or more."
[0233] Likewise, although the present disclosure has been shown and described with respect to one or more implementations, equivalent variations and modifications will occur to those skilled in the art after reading and understanding the specification and drawings. The present disclosure includes all such modifications and variations and is limited only by the scope of the claims. In particular, with respect to the various functions performed by the components described above (e.g., modules, resources, etc.), unless otherwise indicated, the terms used to describe such components are intended to correspond to any component (functionally equivalent) that performs the specific functions of the described components, even if structurally not equivalent to the disclosed structures. In addition, although specific features of the present disclosure may have been disclosed with respect to only one of several implementations, such features may be combined with one or more other features of other implementations as may be desired and beneficial for any given or specific application. In addition, with respect to the terms "including," "having," "having," "having," or variations thereof used in the specific embodiments or claims, such terms are intended to be inclusive in a manner similar to the term "comprising."
[0234] Those skilled in the art will readily appreciate other embodiments of the present disclosure after considering the specification and practicing the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include common knowledge or customary techniques in the art not disclosed herein. The description and examples are to be considered as exemplary only, with the true scope and spirit of the present disclosure being indicated by the appended claims.
[0235] It should be understood that the present disclosure is not limited to the exact structures that have been described above and shown in the drawings, and that various modifications and changes can be made without departing from the scope thereof. The scope of the present disclosure is limited only by the appended claims.
Claims
1. An image processing method, characterized in that: include: Obtain the character string in the target image to obtain the first character string; Inputting the first character string into a recognition model to obtain a second character string output by the recognition model and obtained by translating the first character string; The recognition model is trained by taking training data as input and using translation results of training strings in the training data as at least part of the output, wherein the training data includes first sample data, the first sample data is generated based on a sample image, the first sample data includes a sample string in the sample image, and a contextual relationship of the sample string in the sample image.
2. The method according to claim 1, characterized in that include: determining a plurality of regions in the image to be processed; For each of the plurality of regions, obtaining a character string in the region to obtain a third character string; For each of the regions, when the third character string in any of the regions includes a set ambiguous character string, the image to be processed is used as the target image.
3. The method according to claim 2, characterized in that In the case where the third character string in any of the regions includes a set ambiguous character string, using the image to be processed as the target image comprises: When the third character string in any of the regions includes a set ambiguous character string, obtaining the number of character strings in the region; When the number of the character strings is less than a first threshold, the image to be processed is used as the target image.
4. The method according to claim 3, characterized in that The first threshold is greater than 1, and when the number of the character strings is less than the first threshold, the image to be processed is used as the target image, including: When the number of the character strings is 1, the image to be processed is used as the target image.
5. The method according to any one of claims 1 to 4, characterized in that include: Get the character string in the sample image to obtain the sample character string; Obtaining a contextual relationship of the sample character string in the sample image; determining a translation result of the sample character string, wherein the translation result is obtained by translating part or all of the sample character string into a target language; generating a first instruction, wherein the first instruction is used to instruct to translate the sample character string into the target language; The first sample data is generated according to the sample character string, the context relationship, the translation result, and the first instruction.
6. The method according to claim 5, characterized in that Determining the translation result of the sample character string includes: generating a second instruction, wherein the second instruction is used to instruct to translate part or all of the sample character string into a target language based on the sample character string and a contextual relationship of the sample character string in the sample image; The sample character string, the context relationship, and the first instruction are input into a translation model to obtain the translation result output by the translation model.
7. The method according to any one of claims 1 to 4, characterized in that include: Obtaining a string pair, the string pair comprising an original string and a target string obtained by translating the original string into a target language; generating a third instruction, wherein the third instruction is used to instruct to translate the original character string into the target language; Second sample data is generated according to the string pair and the third instruction, and the training data includes the second sample data.
8. The method according to any one of claims 1 to 4, characterized in that include: Obtaining a string pair, the string pair comprising an original string and a target string obtained by translating the original string into a target language; Obtaining a prompt string pair, the prompt string pair including an original prompt string and a target prompt string obtained by translating the original prompt string into the target language, wherein a similarity between the original prompt string and the original string is greater than a second threshold; generating a fourth instruction, the fourth instruction being used to instruct translation of the original character string into a target language with reference to the prompt character string pair; Third sample data is generated according to the string pair, the prompt string pair, and the fourth instruction, and the training data includes the third sample data.
9. An image processing device, characterized in that: include: The first module is configured to obtain a character string in a target image to obtain a first character string; A second module is configured to input the first character string into a recognition model, and obtain a second character string output by the recognition model and obtained by translating the first character string; The recognition model is trained by taking training data as input and using translation results of training strings in the training data as at least part of the output, wherein the training data includes first sample data, the first sample data is generated based on a sample image, the first sample data includes a sample string in the sample image, and a contextual relationship of the sample string in the sample image.
10. An image processing device, characterized in that: include: processor; a memory for storing processor-executable instructions; The processor is configured to execute the method according to any one of claims 1 to 8.
11. A computer-readable storage medium having computer program instructions stored thereon, characterized in that: When the program instructions are executed by a processor, the method according to any one of claims 1 to 8 is implemented.
12. A computer program product, characterized in that The invention comprises a computer program which, when executed by a processor, implements the steps of the method according to any one of claims 1 to 8.