Method and system for identifying content of image file
By extracting text areas in the image file and determining alternative characters using matching rule data, the problem of errors in image file content recognition in the prior art is solved, and a more accurate and adaptive recognition effect is achieved.
Patent Information
- Application Number
- CN202311486730.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-08
- Publication Date
- 2025-05-09
AI Technical Summary
The prior art is prone to recognition errors when extracting text content from image files containing text content, resulting in large semantic deviations from the real content, and characters with defects cannot be automatically corrected or replaced.
Content recognition and correction is performed by extracting text areas in the image file, obtaining the current matching rule data, including the content matching data set and the content substitution data set, and determining existing characters in each text area.
It improves the accuracy and adaptability of image file content recognition, and meets users' requirements for fast, accurate and adaptive recognition.
Smart Images

Figure CN119964182A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image processing, and in particular to a method and system for performing content recognition on an image file, as well as a storage medium and an electronic device. Background Art
[0002] At present, when extracting text content from an image file containing text content, recognition errors often occur. This situation results in a large semantic deviation between the recognized text content and the actual text content. When a large semantic deviation occurs, it may have a negative impact on the content recognition of the image file.
[0003] In other cases, there are some defective characters or characters that obviously do not meet the requirements in the image file. In this case, the recognized text content can only express the defective characters or characters that obviously do not meet the requirements in the original text, and cannot automatically correct or automatically replace the defective characters or characters that obviously do not meet the requirements. Therefore, the solutions in the prior art cannot meet the user's requirements for fast recognition, accurate recognition or adaptive recognition. Summary of the invention
[0004] In order to solve the problems in the prior art, the present application proposes a method and system for content recognition of image files, as well as a storage medium and an electronic device, which utilize at least one existing character and at least one alternative character to improve the results of content recognition to meet users' requirements for fast recognition, accurate recognition or adaptive recognition.
[0005] According to one aspect of the present invention, a method for performing content recognition on an image file is provided, the method comprising:
[0006] When receiving an image file to be recognized, extracting a plurality of text regions containing characters from the image file to be recognized;
[0007] Acquire current matching rule data from a server, wherein the matching rule data includes: a content matching data set and a content replacement data set;
[0008] determining at least one existing character included in each text region and determining at least one replacement character for each existing character based on the content replacement data set, such that each text region has at least one existing character and each text region is associated with at least one replacement character;
[0009] Performing content recognition on at least one existing character and at least one replacement character in each text region based on the content matching data set to determine whether a target character and / or a target character group exists in each text region; and
[0010] When it is determined that the target character and / or the target character group exists in the text area, the recognition type of the target character and / or the target character group is determined.
[0011] Preferably, extracting a plurality of text regions containing characters from the image file to be recognized comprises:
[0012] Using a pre-set character recognition tool, extracting a plurality of text regions containing characters from the image file to be recognized;
[0013] Each text region includes at least one existing character, and a character distance between adjacent characters in the text region is less than a distance threshold.
[0014] Preferably, the content matching data set includes a plurality of matching data subsets, each matching data subset is associated with a recognition type, and each matching data subset includes a plurality of matching data items.
[0015] Preferably, the content substitution data set includes a plurality of content substitution data items, each of which includes a plurality of characters with equivalent meanings.
[0016] Preferably, determining at least one replacement character for each existing character based on the content replacement data set comprises:
[0017] Using each existing character, searching in the content replacement data set to determine a content replacement data item associated with each existing character;
[0018] Each character in the content replacement data item, except the existing character, is determined as at least one replacement character for the existing character.
[0019] Preferably, performing content recognition on at least one existing character and at least one replacement character in each text region based on the content matching data set to determine whether a target character and / or a target character group exists in each text region comprises:
[0020] A content matching degree of a single existing character with multiple matching data items in each matching data subset is determined, and based on the content matching degree with multiple matching data items in each matching data subset, it is determined whether the existing character is a target character.
[0021] Preferably, performing content recognition on at least one existing character and at least one replacement character in each text region based on the content matching data set to determine whether a target character and / or a target character group exists in each text region comprises:
[0022] A content matching degree of the single replacement character with the plurality of matching data items in each matching data subset is determined, and based on the content matching degree with the plurality of matching data items in each matching data subset, it is determined whether the replacement character is a target character.
[0023] Preferably, performing content recognition on at least one existing character and at least one replacement character in each text region based on the content matching data set to determine whether a target character and / or a target character group exists in each text region comprises:
[0024] Determine the content matching degree between the character group formed by at least one existing character and the multiple matching data items in each matching data subset, and determine whether the character group formed by at least one existing character is a target character group based on the content matching degree with the multiple matching data items in each matching data subset.
[0025] Preferably, performing content recognition on at least one existing character and at least one replacement character in each text region based on the content matching data set to determine whether a target character and / or a target character group exists in each text region comprises:
[0026] Determine the content matching degree between the character group formed by the at least one replacement character and the multiple matching data items in each matching data subset, and determine whether the character group formed by the at least one replacement character is a target character group based on the content matching degree with the multiple matching data items in each matching data subset.
[0027] Preferably, performing content recognition on at least one existing character and at least one replacement character in each text region based on the content matching data set to determine whether a target character and / or a target character group exists in each text region comprises:
[0028] Determine the content matching degree of a character group consisting of at least one existing character and at least one replacement character with multiple matching data items in each matching data subset, and determine whether the character group consisting of at least one existing character and at least one replacement character is a target character group based on the content matching degree with multiple matching data items in each matching data subset.
[0029] Preferably, when it is determined that there are target characters and / or target character groups in the text area, determining the recognition type of the target characters and / or target character groups includes:
[0030] When it is determined that the target character and / or the target character group exists in the text area, determining a matching data subset corresponding to the target character and / or the target character group;
[0031] The recognition type associated with the matching data subset is determined as the recognition type of the target character and / or target character group.
[0032] According to another aspect of the present invention, a system for performing content recognition on an image file is provided, the system comprising:
[0033] An extraction unit, configured to extract a plurality of text regions containing characters from an image file to be identified when receiving the image file to be identified;
[0034] An acquisition unit, configured to acquire current matching rule data from a server, wherein the matching rule data includes: a content matching data set and a content replacement data set;
[0035] A first determining unit, configured to determine at least one existing character included in each text region, and determine at least one replacement character for each existing character based on the content replacement data set, so that each text region has at least one existing character and each text region is associated with at least one replacement character;
[0036] a recognition unit, configured to perform content recognition on at least one existing character and at least one replacement character in each text region based on the content matching data set, so as to determine whether a target character and / or a target character group exists in each text region; and
[0037] The second determining unit is configured to determine the recognition type of the target character and / or the target character group when it is determined that the target character and / or the target character group exists in the text area.
[0038] According to another aspect of the present invention, an electronic device is provided, comprising: a memory and a processor, wherein the memory and the processor are coupled; the memory stores program instructions, and when the program instructions are executed by the processor, the electronic device executes any one of the methods described above.
[0039] According to yet another aspect of the present invention, there is provided a computer-readable storage medium, comprising a computer program, which, when executed on an electronic device, enables the electronic device to execute any one of the above methods.
[0040] The technical solution of the present invention is further described in detail below through the accompanying drawings and embodiments. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] The above and other purposes, features and advantages of the present invention will become more apparent by describing the embodiments of the present invention in more detail in conjunction with the accompanying drawings. The accompanying drawings are used to provide a further understanding of the embodiments of the present invention and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention and do not constitute a limitation of the present invention. In the accompanying drawings, the same reference numerals generally represent the same components or steps.
[0042] Figure 1 is a flow chart of a method for performing content recognition on an image file provided by an exemplary embodiment of the present invention;
[0043] Figure 2 is a schematic diagram of content identification provided by an exemplary embodiment of the present invention; and
[0044] Figure 3 It is a schematic diagram of the structure of a system for performing content recognition on an image file provided by an exemplary embodiment of the present invention. DETAILED DESCRIPTION
[0045] Below, the exemplary embodiments according to the present invention will be described in detail with reference to the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all the embodiments of the present invention, and it should be understood that the present invention is not limited to the exemplary embodiments described here.
[0046] It should be noted that the relative arrangement of components and steps, the numerical expressions and numerical values set forth in these embodiments do not limit the scope of the present invention unless specifically stated otherwise.
[0047] Those skilled in the art can understand that the terms "first" and "second" in the embodiments of the present invention are only used to distinguish different steps, devices or modules, etc., and neither represent any specific technical meaning nor indicate the necessary logical order between them.
[0048] It should also be understood that, in the embodiments of the present invention, “plurality” may refer to two or more than two, and “at least one” may refer to one, two or more than two.
[0049] It should also be understood that any component, data or structure mentioned in the embodiments of the present invention can generally be understood as one or more, unless explicitly limited or otherwise indicated in the context.
[0050] Figure 1 FIG. 1 is a flow chart of a method for identifying content of an image file provided by an exemplary embodiment of the present invention. Figure 1 As shown, the method 100 includes:
[0051] Step 101, when receiving an image file to be identified, extract multiple text areas containing characters from the image file to be identified. Generally, the image file to be identified can be various types of image prints, image files of various formats, image files acquired by a camera, image files generated by design software, etc. Due to different display methods, the image file may have graphic content, character content or text content, etc. in different positions or areas. Therefore, in order to efficiently and accurately identify the content of the image file, the image file to be identified can be partitioned according to the graphic content, text content or character content, etc.
[0052] Preferably, extracting multiple text regions containing characters from the image file to be recognized comprises: extracting multiple text regions containing characters from the image file to be recognized using a pre-set character recognition tool. The character recognition tool may be any suitable recognition tool for content recognition, such as optical character recognition (OCR). Each text region includes at least one existing character and at least one existing text, and the character distance between adjacent characters in the text region is less than a distance threshold. Preferably, the distance threshold may be pre-set, and the distance threshold may be set according to the size of the characters or text in the image file to be recognized.
[0053] Figure 2 FIG. 1 is a schematic diagram of content identification provided by an exemplary embodiment of the present invention. Figure 2 As shown, using a preset character recognition tool, multiple text areas containing characters are extracted from the image file to be recognized: text area 1, text area 2, text area 3, text area 4, and text area 5. Each of text area 1, text area 2, text area 3, text area 4, and text area 5 may contain at least one existing character and at least one existing text, and the character distance between adjacent characters in the text area is less than a distance threshold.
[0054] Step 102, obtain current matching rule data from the server, the matching rule data includes: content matching data set and content replacement data set. It should be understood that the content matching data set is used to match text or characters, and the content replacement data set is used to match and replace defective characters or characters that obviously do not meet the requirements.
[0055] Preferably, the content matching data set includes a plurality of matching data subsets, each matching data subset is associated with a recognition type, and each matching data subset includes a plurality of matching data items. Preferably, the content substitution data set includes a plurality of content substitution data items, each content substitution data item includes a plurality of characters with equivalent meanings.
[0056] Step 103, determining at least one existing character included in each text region, and determining at least one replacement character for each existing character based on the content replacement data set, so that each text region has at least one existing character and each text region is associated with at least one replacement character. Preferably, in order to perform content correction or content replacement on characters such as those with defects or characters that obviously do not meet the requirements, the present application determines at least one replacement character for each existing character based on the content replacement data set. In this way, each text region has at least one existing character and each text region is associated with at least one replacement character.
[0057] Preferably, determining at least one substitute character for each existing character based on the content substitute data set comprises: using each existing character to search in the content substitute data set to determine a content substitute data item associated with each existing character, and determining each character in the content substitute data item other than the existing character as at least one substitute character for the existing character.
[0058] Step 104, performing content recognition on at least one existing character and at least one replacement character in each text region based on the content matching data set to determine whether a target character and / or a target character group exists in each text region. In order to perform content recognition on at least one existing character and at least one replacement character in each text region to determine whether a target character and / or a target character group exists in each text region, the present application uses a variety of replaceable implementations.
[0059] In one embodiment, content recognition is performed on at least one existing character and at least one replacement character in each text region based on a content matching data set to determine whether a target character and / or a target character group exists in each text region, including: determining a content matching degree between a single existing character and multiple matching data items in each matching data subset, and determining whether the existing character is a target character based on the content matching degree with multiple matching data items in each matching data subset.
[0060] In one embodiment, content recognition is performed on at least one existing character and at least one replacement character in each text region based on a content matching data set to determine whether a target character and / or a target character group exists in each text region, including: determining a content matching degree of a single replacement character with multiple matching data items in each matching data subset, and determining whether the replacement character is a target character based on the content matching degree with multiple matching data items in each matching data subset.
[0061] In one embodiment, content recognition is performed on at least one existing character and at least one replacement character in each text region based on a content matching data set to determine whether a target character and / or a target character group exists in each text region, including: determining a content matching degree between a character group consisting of at least one existing character and a plurality of matching data items in each matching data subset, and determining whether a character group consisting of at least one existing character is a target character group based on a content matching degree with a plurality of matching data items in each matching data subset.
[0062] In one embodiment, content recognition is performed on at least one existing character and at least one replacement character in each text region based on a content matching data set to determine whether a target character and / or a target character group exists in each text region, including: determining a content matching degree between a character group consisting of at least one replacement character and a plurality of matching data items in each matching data subset, and determining whether a character group consisting of at least one replacement character is a target character group based on a content matching degree with a plurality of matching data items in each matching data subset.
[0063] In one embodiment, content recognition is performed on at least one existing character and at least one replacement character in each text region based on a content matching data set to determine whether a target character and / or a target character group exists in each text region, including: determining a content matching degree between a character group consisting of at least one existing character and at least one replacement character and a plurality of matching data items in each matching data subset, and determining whether a character group consisting of at least one existing character and at least one replacement character is a target character group based on a content matching degree with a plurality of matching data items in each matching data subset.
[0064] Step 105, when it is determined that the target character and / or the target character group exists in the text area, the recognition type of the target character and / or the target character group is determined. Specifically, when it is determined that the target character and / or the target character group exists in the text area, the recognition type of the target character and / or the target character group is determined, including: when it is determined that the target character and / or the target character group exists in the text area, determining the matching data subset corresponding to the target character and / or the target character group; determining the recognition type associated with the matching data subset as the recognition type of the target character and / or the target character group.
[0065] Figure 3 FIG. 1 is a schematic diagram of a system for performing content recognition on an image file provided by an exemplary embodiment of the present invention. Figure 3 As shown, the system 300 includes: an extracting unit 301 , an acquiring unit 302 , a first determining unit 303 , an identifying unit 304 , and a second determining unit 305 .
[0066] Preferably, the extraction unit 301 is used to extract multiple text regions containing characters from the image file to be recognized when receiving the image file to be recognized. The extraction unit 301 is specifically used to extract multiple text regions containing characters from the image file to be recognized using a pre-set character recognition tool; wherein each text region includes at least one existing character, and the character distance between adjacent characters in the text region is less than a distance threshold.
[0067] Preferably, the acquisition unit 302 is used to acquire current matching rule data from the server, and the matching rule data includes: a content matching data set and a content substitution data set. The content matching data set includes multiple matching data subsets, each matching data subset is associated with a recognition type, and each matching data subset includes multiple matching data items. The content substitution data set includes multiple content substitution data items, and each content substitution data item includes multiple characters with equivalent meanings.
[0068] Preferably, the first determining unit 303 is used to determine at least one existing character included in each text region, and determine at least one replacement character for each existing character based on the content replacement data set, so that each text region has at least one existing character and each text region is associated with at least one replacement character.
[0069] The first determining unit 303 is specifically configured to use each existing character to search in the content replacement data set, thereby determining a content replacement data item associated with each existing character; and determining each character in the content replacement data item except the existing character as at least one replacement character for the existing character.
[0070] Preferably, the recognition unit 304 is used to perform content recognition on at least one existing character and at least one replacement character in each text region based on the content matching data set to determine whether there is a target character and / or a target character group in each text region.
[0071] The recognition unit 304 is specifically used to determine the content matching degree of a single existing character with multiple matching data items in each matching data subset, and determine whether the existing character is a target character based on the content matching degree with multiple matching data items in each matching data subset. Alternatively, the recognition unit 304 is specifically used to determine the content matching degree of a single replacement character with multiple matching data items in each matching data subset, and determine whether the replacement character is a target character based on the content matching degree with multiple matching data items in each matching data subset. Alternatively, the recognition unit 304 is specifically used to determine the content matching degree of a character group composed of at least one existing character with multiple matching data items in each matching data subset, and determine whether the character group composed of at least one existing character is a target character group based on the content matching degree with multiple matching data items in each matching data subset. Alternatively, the recognition unit 304 is specifically used to determine the content matching degree of a character group composed of at least one replacement character with multiple matching data items in each matching data subset, and determine whether the character group composed of at least one replacement character is a target character group based on the content matching degree with multiple matching data items in each matching data subset. Alternatively, the recognition unit 304 is specifically used to determine the content matching degree between a character group consisting of at least one existing character and at least one replacement character and multiple matching data items in each matching data subset, and based on the content matching degree with multiple matching data items in each matching data subset, determine whether the character group consisting of at least one existing character and at least one replacement character is a target character group.
[0072] Preferably, the second determining unit 305 is used to determine the recognition type of the target character and / or the target character group when it is determined that the target character and / or the target character group exists in the text area. The second determining unit 305 is specifically used to determine the matching data subset corresponding to the target character and / or the target character group when it is determined that the target character and / or the target character group exists in the text area; and determine the recognition type associated with the matching data subset as the recognition type of the target character and / or the target character group.
[0073] The basic principles of the present disclosure are described above in conjunction with specific embodiments. However, it should be noted that the advantages, strengths, effects, etc. mentioned in the present disclosure are only examples and not limitations, and it cannot be considered that these advantages, strengths, effects, etc. are required by each embodiment of the present disclosure. In addition, the specific details disclosed above are only for the purpose of illustration and ease of understanding, and are not limitations. The above details do not limit the present disclosure to the necessity of adopting the above specific details to be implemented.
[0074] Each embodiment in this specification is described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the embodiments can be referred to each other. For the system embodiment, since it basically corresponds to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.
[0075] The block diagrams of the devices, apparatuses, equipment, and systems involved in this disclosure are only illustrative examples and are not intended to require or imply that they must be connected, arranged, and configured in the manner shown in the block diagrams. As will be appreciated by those skilled in the art, these devices, apparatuses, equipment, and systems can be connected, arranged, and configured in any manner. Words such as "including," "comprising," "having," and the like are open words, referring to "including but not limited to," and can be used interchangeably therewith. The words "or" and "and" used herein refer to the words "and / or," and can be used interchangeably therewith, unless the context clearly indicates otherwise. The word "such as" used herein refers to the phrase "such as but not limited to," and can be used interchangeably therewith.
[0076] The method and apparatus of the present disclosure may be implemented in many ways. For example, the method and apparatus of the present disclosure may be implemented by software, hardware, firmware, or any combination of software, hardware, and firmware. The above order of steps for the method is for illustration only, and the steps of the method of the present disclosure are not limited to the order specifically described above, unless otherwise specifically stated. In addition, in some embodiments, the present disclosure may also be implemented as a program recorded in a recording medium, which includes machine-readable instructions for implementing the method according to the present disclosure. Therefore, the present disclosure also covers a recording medium storing a program for executing the method according to the present disclosure.
[0077] It should also be noted that in the apparatus, equipment and method of the present disclosure, each component or each step can be decomposed and / or recombined. These decompositions and / or recombinations should be regarded as equivalent schemes of the present disclosure. The above description of the disclosed aspects is provided to enable any technician in the field to make or use the present disclosure. Various modifications to these aspects are very obvious to those skilled in the art, and the general principles defined herein can be applied to other aspects without departing from the scope of the present disclosure. Therefore, the present disclosure is not intended to be limited to the aspects shown here, but to the widest scope consistent with the principles and novel features disclosed herein.
[0078] The above description has been given for the purpose of illustration and description. In addition, this description is not intended to limit the embodiments of the present disclosure to the forms disclosed herein. Although multiple example aspects and embodiments have been discussed above, those skilled in the art will recognize certain variations, modifications, changes, additions and sub-combinations thereof.
Claims
1. A method for content recognition of an image file, characterized in that: The method comprises: When receiving an image file to be recognized, extracting a plurality of text regions containing characters from the image file to be recognized; Acquire current matching rule data from a server, wherein the matching rule data includes: a content matching data set and a content replacement data set; determining at least one existing character included in each text region and determining at least one replacement character for each existing character based on the content replacement data set, such that each text region has at least one existing character and each text region is associated with at least one replacement character; Performing content recognition on at least one existing character and at least one replacement character in each text region based on the content matching dataset to determine whether a target character and / or a target character group exists in each text region; and When it is determined that the target character and / or the target character group exists in the text area, the recognition type of the target character and / or the target character group is determined.
2. The method according to claim 1, characterized in that Extracting a plurality of text regions containing characters from the image file to be recognized includes: Using a pre-set character recognition tool, extracting a plurality of text regions containing characters from the image file to be recognized; Each text region includes at least one existing character, and a character distance between adjacent characters in the text region is less than a distance threshold.
3. The method according to claim 1, characterized in that The content matching data set includes a plurality of matching data subsets, each matching data subset is associated with a recognition type, and each matching data subset includes a plurality of matching data items.
4. The method according to claim 1, characterized in that: The content substitution data set includes a plurality of content substitution data items, each of which includes a plurality of characters with equivalent meanings.
5. The method according to claim 4, characterized in that in, Determining at least one replacement character for each existing character based on the content replacement data set includes: Using each existing character, searching in the content replacement data set to determine a content replacement data item associated with each existing character; Each character in the content replacement data item, except the existing character, is determined as at least one replacement character for the existing character.
6. The method according to claim 3, characterized in that Performing content recognition on at least one existing character and at least one replacement character in each text region based on the content matching dataset to determine whether a target character and / or a target character group exists in each text region, including: A content matching degree of a single existing character with multiple matching data items in each matching data subset is determined, and based on the content matching degree with multiple matching data items in each matching data subset, it is determined whether the existing character is a target character.
7. The method according to claim 3, characterized in that Performing content recognition on at least one existing character and at least one replacement character in each text region based on the content matching dataset to determine whether a target character and / or a target character group exists in each text region, including: A content matching degree of the single replacement character with the plurality of matching data items in each matching data subset is determined, and based on the content matching degree with the plurality of matching data items in each matching data subset, it is determined whether the replacement character is a target character.
8. The method according to claim 3, characterized in that Performing content recognition on at least one existing character and at least one replacement character in each text region based on the content matching dataset to determine whether a target character and / or a target character group exists in each text region, including: Determine the content matching degree between the character group formed by at least one existing character and the multiple matching data items in each matching data subset, and determine whether the character group formed by at least one existing character is a target character group based on the content matching degree with the multiple matching data items in each matching data subset.
9. The method according to claim 3, characterized in that: Performing content recognition on at least one existing character and at least one replacement character in each text region based on the content matching dataset to determine whether a target character and / or a target character group exists in each text region, including: Determine the content matching degree between the character group formed by the at least one replacement character and the multiple matching data items in each matching data subset, and determine whether the character group formed by the at least one replacement character is a target character group based on the content matching degree with the multiple matching data items in each matching data subset.
10. The method according to claim 3, characterized in that: Performing content recognition on at least one existing character and at least one replacement character in each text region based on the content matching dataset to determine whether a target character and / or a target character group exists in each text region, including: Determine the content matching degree of a character group consisting of at least one existing character and at least one replacement character with multiple matching data items in each matching data subset, and determine whether the character group consisting of at least one existing character and at least one replacement character is a target character group based on the content matching degree with multiple matching data items in each matching data subset.
11. The method according to any one of claims 3 and 6-10, characterized in that: When it is determined that there is a target character and / or a target character group in the text area, determining the recognition type of the target character and / or the target character group includes: When it is determined that the target character and / or the target character group exists in the text area, determining a matching data subset corresponding to the target character and / or the target character group; The recognition type associated with the matching data subset is determined as the recognition type of the target character and / or target character group.
12. A system for content recognition of an image file, characterized in that: The system comprises: An extraction unit, configured to extract a plurality of text regions containing characters from an image file to be identified when receiving the image file to be identified; An acquisition unit, configured to acquire current matching rule data from a server, wherein the matching rule data includes: a content matching data set and a content replacement data set; A first determining unit, configured to determine at least one existing character included in each text region, and determine at least one replacement character for each existing character based on the content replacement data set, so that each text region has at least one existing character and each text region is associated with at least one replacement character; a recognition unit, configured to perform content recognition on at least one existing character and at least one replacement character in each text region based on the content matching data set, so as to determine whether a target character and / or a target character group exists in each text region; and The second determining unit is used to determine the recognition type of the target character and / or the target character group when it is determined that the target character and / or the target character group exists in the text area.
13. An electronic device, characterized in that: The electronic device comprises: a memory and a processor, wherein the memory and the processor are coupled; the memory stores program instructions, and when the program instructions are executed by the processor, the electronic device executes the method according to any one of claims 1 to 11.
14. A computer-readable storage medium, characterized in that: The method comprises a computer program, which, when executed on an electronic device, enables the electronic device to execute the method according to any one of claims 1 to 11.