Text recognition method and related electronic device

By caching un-graffitied images in electronic devices and using text recognition algorithms to extract the text before graffiti, the problem of inaccurate text recognition caused by graffiti occlusion is solved, achieving higher recognition accuracy and user-friendly operation.

CN118057486BActive Publication Date: 2026-01-09HONOR DEVICE CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202211458126.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-21
Publication Date
2026-01-09
Estimated Expiration
2042-11-21

AI Technical Summary

Technical Problem

The existing technology suffers from inaccurate character recognition when graffiti obscures the image.

Method used

The system displays an image of the area before graffiti on an electronic device and uses a text recognition algorithm to extract text from the ungraffitied areas. It then caches the ungraffitied image for recognition and extraction and displays the recognized text in text form according to preset display rules.

Benefits of technology

It improves the accuracy of text recognition, enabling accurate extraction and display of text in un-doodled areas even after graffiti, thus enhancing the reliability of user operation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118057486B_ABST
    Figure CN118057486B_ABST
Patent Text Reader

Abstract

The application provides a text recognition method and related electronic equipment, the method comprising: displaying a first image and a first control in a first interface, the text in the first image being displayed in the form of an image; in response to a scribbling operation, displaying a second image on the first interface, the second image being obtained based on the first image, the second image being covered with scribbling marks, and the text in the second image being displayed in the form of an image; in the case of displaying the second image, in response to a second operation on the first control, the electronic equipment extracts the text in the first image using a text recognition algorithm; and displaying a second interface, the second interface comprising a third image, the third image being obtained based on the second image, and all or part of the text in the third image being displayed in the form of text.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of text recognition, and in particular, to a text recognition method and related electronic device. BACKGROUND

[0002] Now, the OCR technology has been applied to the terminal device, and the user can recognize and extract the text in the picture by clicking the picture, and then can copy the extracted text. This greatly saves the user's typing time. Therefore, how to improve the accuracy of text recognition in the picture is a problem that the technical personnel increasingly concerns. SUMMARY

[0003] Embodiments of the present application provide a text recognition method and related electronic device, which solve the problem of inaccurate text recognition in the case that the image exists a scribble or the like.

[0004] In a first aspect, the embodiments of the present application provide a text recognition method applied to an electronic device, which includes: displaying a first image and a first control in a first interface, the text in the first image being displayed in the form of an image; in response to a scribble operation, displaying a second image on the first interface, the second image being obtained based on the first image, the second image being covered with scribble marks, and the text in the second image being displayed in the form of an image; in the case that the second image is displayed, in response to a second operation on the first control, the electronic device extracts the text in the first image using a text recognition algorithm; and displaying a second interface, the second interface including a third image, the third image being obtained based on the second image, and all or part of the text in the third image being displayed in the form of text.

[0005] In the above embodiment, the electronic device temporarily caches the image without scribble. After the user scribbles on the image, if the text recognition is to be performed on the scribbled image, the electronic device can recognize and extract the text in the image before scribbling. Then, the recognized and extracted text is displayed in the form of text according to a preset display rule, so that the user can perform operations such as copying and pasting on the text displayed in the form of text. Since the electronic device recognizes and extracts the text using the image without scribbling, compared with recognizing and extracting the text using the scribbled image, the text recognized and extracted by the electronic device in the embodiments of the present application is more accurate.

[0006] In combination with the first aspect, in a possible implementation manner, before the first image and the first control are displayed in the first interface, the method further includes: displaying a browsing interface; in response to a screenshot operation, performing screenshot on the browsing interface; and obtaining an image of the browsing interface, the image being the first image.

[0007] With reference to the first aspect, in a possible implementation manner, before the first image and the first control are displayed in the first interface, the method further includes: displaying a browsing interface; in response to the screenshot operation, performing screenshot on the first region in the browsing screenshot; and obtaining the image of the browsing interface, the image corresponding to the first region in the image of the browsing interface being the first image.

[0008] With reference to the first aspect, in a possible implementation manner, after the first image and the first control are displayed in the first interface, the method further includes: in the first interface, in response to a third operation on the first region in the image of the browsing interface; transforming the first region in the image of the browsing interface into a second region, the first region being different from the second region; the image in the second region being the updated first image.

[0009] With reference to the first aspect, in a possible implementation manner, the first interface further includes a second control, the second control being used to indicate a first shape, and after the first image and the first control are displayed in the first interface, the method further includes: detecting a fourth operation on the second control; in response to the fourth operation, transforming the shape of the first region into the first shape; the image in the first region after the shape is transformed being the updated first image.

[0010] With reference to the first aspect, in a possible implementation manner, after the scribbling operation is responded to, the position information of the scribbling mark in the first interface is obtained.

[0011] With reference to the first aspect, in a possible implementation manner, before the second interface is displayed, the method further includes: determining the text extracted from the first image as target text; the target text being the text displayed in the third image in the form of text.

[0012] With reference to the first aspect, in a possible implementation manner, before the second interface is displayed, the method further includes: determining, according to the position information of the text in the first image in the first interface and the position information of the scribbling mark in the first interface, overlapping text that has an overlapping region with the scribbling mark; calculating a proportion value of an overlapping region of each overlapping text relative to an overlapping text region; determining, as target text, overlapping text whose proportion value is less than or equal to a first proportion threshold and text that does not have an overlapping region with the scribbling mark; the target text being text displayed in the third image in the form of text from the text extracted by the electronic device.

[0013] With reference to the first aspect, in a possible implementation manner, before the second interface is displayed, the method further includes: extracting, by the electronic device, text in the first image after scribbling using a text recognition algorithm to obtain second text; comparing the second text with first text; the first text being text extracted by the electronic device from the first image; determining, as target text, text in the first text that is consistent with the second text; the target text being text displayed in the third image in the form of text from the text extracted by the electronic device.

[0014] In a second aspect, an electronic device is provided. The electronic device includes a display screen, one or more processors, the display screen, and a memory. The memory is coupled to the one or more processors. The memory is configured to store computer program code including computer instructions. The one or more processors are configured to invoke the computer instructions to cause the electronic device to perform the following operations: controlling the display screen to display a first image and a first control in a first interface. Text in the first image is displayed in the form of an image. In response to a scribbling operation, controlling the display screen to display a second image on the first interface. The second image is obtained based on the first image. The second image is overlaid with scribbling marks. Text in the second image is displayed in the form of an image. In a case where the display screen displays the second image, in response to a second operation on the first control, using a text recognition algorithm to extract the text in the first image. Controlling the display screen to display a second interface. The second interface includes a third image. The third image is obtained based on the second image. All or part of the text in the third image is displayed in the form of text.

[0015] With reference to the second aspect, in a possible implementation manner, the one or more processors invoke the computer instructions to cause the electronic device to perform the following operations: before controlling the display screen to display the first image and the first control in the first interface, the operations further include: controlling the display screen to display a browsing interface; in response to a screenshot operation, performing a screenshot on the browsing interface; obtaining an image of the browsing interface. The image is the first image.

[0016] With reference to the second aspect, in a possible implementation manner, the one or more processors invoke the computer instructions to cause the electronic device to perform the following operations: before displaying the first image and the first control in the first interface, the operations further include: controlling the display screen to display a browsing interface; in response to a screenshot operation, performing a screenshot on a first region in the browsing screenshot; obtaining an image of the browsing interface. The image corresponding to the first region in the image of the browsing interface is the first image.

[0017] With reference to the second aspect, in a possible implementation manner, the one or more processors invoke the computer instructions to cause the electronic device to perform the following operations: after controlling the display screen to display the first image and the first control in the first interface, the operations further include: in the first interface, in response to a third operation on a first region in the image of the browsing interface; transforming the first region in the image of the browsing interface into a second region. The first region is different from the second region. An image in the second region is an updated first image.

[0018] With reference to the second aspect, in a possible implementation manner, the one or more processors invoke the computer instructions to cause the electronic device to perform: the first interface further includes a second control, the second control being used to indicate the first shape, and the control display screen displays the first image and the first control in the first interface, and further includes: detecting a fourth operation on the second control; in response to the fourth operation, transforming the shape of the first region into the first shape; and transforming the image in the first region after the shape is transformed into the updated first image.

[0019] With reference to the second aspect, in a possible implementation manner, the one or more processors invoke the computer instructions to cause the electronic device to perform: obtaining position information of the graffiti mark in the first interface.

[0020] With reference to the second aspect, in a possible implementation manner, the one or more processors invoke the computer instructions to cause the electronic device to perform: determining the text extracted from the first image as the target text; and the target text is the text displayed in the third image in the form of text.

[0021] With reference to the second aspect, in a possible implementation manner, the one or more processors invoke the computer instructions to cause the electronic device to perform: determining, according to the position information of the text in the first image in the first interface and the position information of the graffiti mark in the first interface, the coincident text that has a coincident region with the graffiti mark; calculating a proportion value of the coincident region of each coincident text in the coincident text region; determining the coincident text with the proportion value less than or equal to the first proportion threshold and the text without the coincident region with the graffiti mark as the target text; and the target text is the text displayed in the third image in the form of text among the text extracted by the electronic device.

[0022] With reference to the second aspect, in a possible implementation manner, the one or more processors invoke the computer instructions to cause the electronic device to perform: extracting the text in the first image after the graffiti using a text recognition algorithm to obtain second text; comparing the second text with the first text; the first text is the text extracted from the first image; determining the text consistent with the second text in the first text as the target text; and the target text is the text displayed in the third image in the form of text among the text extracted by the electronic device.

[0023] In a third aspect, an electronic device is provided, including a touch screen, a camera, one or more processors, and one or more memories. The one or more processors are coupled to the touch screen, the camera, and the one or more memories. The one or more memories are configured to store computer program codes. The computer program codes include computer instructions. When the one or more processors execute the computer instructions, the electronic device performs the method according to the first aspect or any possible implementation of the first aspect.

[0024] In a fourth aspect, a chip system is provided. The chip system is applied to an electronic device. The chip system includes one or more processors. The processor is configured to invoke computer instructions to cause the electronic device to perform the method according to the first aspect or any possible implementation of the first aspect.

[0025] In a fifth aspect, a computer program product is provided. The computer program product includes instructions. When the computer program product is executed on an electronic device, the electronic device performs the method according to the first aspect or any possible implementation of the first aspect.

[0026] In a sixth aspect, a computer readable storage medium is provided. The computer readable storage medium includes instructions. When the instructions are executed on an electronic device, the electronic device performs the method according to the first aspect or any possible implementation of the first aspect. BRIEF DESCRIPTION OF DRAWINGS

[0027] Figures 1A-1L FIG. 1 is a set of exemplary user interface diagrams provided by embodiments of the present application;

[0028] Figures 2A-2P FIG. 2 is another set of exemplary user interface diagrams provided by embodiments of the present application;

[0029] Figure 3A FIG. 3 is a flowchart of a character recognition method provided by embodiments of the present application;

[0030] Figures 3B-3E FIG. 4 is another set of exemplary user interface diagrams provided by embodiments of the present application;

[0031] Figure 3F FIG. 5 is a feature information recognition algorithm flowchart provided by embodiments of the present application;

[0032] Figures 3G-3K FIG. 6 is another set of exemplary user interface diagrams provided by embodiments of the present application;

[0033] Figure 4 FIG. 7 is another character recognition method flowchart provided by embodiments of the present application;

[0034] Figure 5 is a flowchart of another character recognition method provided by an embodiment of the present application;

[0035] Figure 6 is a flowchart of another character recognition method provided by an embodiment of the present application;

[0036] Figure 7 is a hardware structure schematic diagram of an electronic device 100 provided by an embodiment of the present application;

[0037] Figure 8 is a software structure block diagram of an electronic device 100 of an embodiment of the present application. DETAILED DESCRIPTION

[0038] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. In this document, the phrase "embodiments" means that the specific features, structures or characteristics described in connection with the embodiments can be included in at least one embodiment of the present application. The appearance of this phrase at various places in the specification does not necessarily mean the same embodiments, nor does it mean mutually exclusive or alternative embodiments. It is obvious to those skilled in the art that the embodiments described herein can be combined with other embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor fall within the scope of protection of the present application.

[0039] The terms "first", "second", "third", etc. in the specification and claims of the present application and the drawings are to distinguish different objects, rather than to describe a specific order. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion. For example, a series of steps or units are included, or optionally, other steps or units not listed are also included, or optionally, other steps or units inherent to the process, method, product or equipment are also included.

[0040] Only parts related to the present application are shown in the drawings, rather than all. Before discussing the exemplary embodiments in more detail, it should be mentioned that some exemplary embodiments are described as processes or methods depicted by flowcharts. Although the flowcharts describe the operations (or steps) as sequential processes, many of the operations can be implemented in parallel, concurrently or simultaneously. In addition, the order of the operations can be rearranged. The process can be terminated when its operations are completed, but can also have additional steps not included in the drawings. The process can correspond to a method, function, procedure, subroutine, etc.

[0041] The terms "component," "module," "system," "unit," and the like used in the present specification are used to represent a computer-related entity, hardware, firmware, a combination of hardware and software, software, or software in execution. For example, a unit can be, but is not limited to, a process running on a processor, a processor, an object, an executable, an execution thread, a program, and / or distributed among two or more computers. In addition, these units can be executed from various computer-readable media having various data structures stored thereon. The units can communicate, for example, according to a signal having one or more data packets (e.g., from a second unit data to interact with another unit of a local system, a distributed system, and / or a network. For example, the Internet) by a local and / or remote process.

[0042] The electronic device / electronic device involved in the embodiments of the present application can be a mobile phone, a tablet computer, a desktop computer, a laptop computer, a notebook computer, an ultra-mobile personal computer (UMPC), a handheld computer, a netbook, a personal digital assistant (PDA), a virtual reality device, a PDA (Personal Digital Assistant, also known as a palm computer), a portable Internet device, a music player, a data storage device, a camera, or a wearable device (e.g., a smart watch, a smart bracelet, smart glasses, a head-mounted device (HMD), an electronic clothing, an electronic bracelet, an electronic necklace, an electronic accessory, an electronic tattoo, and a smart mirror), and the like.

[0043] With the development of text recognition technology, in recent years, more and more electronic devices support text recognition function. Text recognition is a process in which an electronic device detects, recognizes, and extracts text in a picture through a text recognition algorithm. The electronic device can recognize the picture and extract the text in the picture through a text recognition algorithm (e.g., an OCR algorithm), and the user can copy or paste the text extracted from the picture by the electronic device, and copy the extracted text to other documents, which greatly saves the user's typing time and greatly improves the user's use efficiency. In order to facilitate the understanding of the embodiments of the present application, the application scenarios corresponding to the embodiments of the present application and the technical problems that can be solved will be specifically analyzed in combination with the accompanying drawings. It should be understood that the following application scenarios indicate exemplary scenarios listed by the embodiments of the present application, which should not limit the protection scope of the embodiments of the present application.

[0044] Next, the application scenarios of the text recognition method provided by the embodiments of the present application will be exemplarily described in combination with the related drawings. Please refer to Figures 1A-1L 、 Figures 2A-2P , Figures 1A-1L ,Figures 2A-2P are two groups of exemplary user interfaces provided by embodiments of the present application.

[0045] As shown in Figure 1A Fig. 10 is a browsing interface 10 of the electronic device 100, which includes text information, such as time information "Time: Mon-Sun 10:00-22:00", address information "XX Province XX City XX District XX Street XX Shopping Center XX", and other text information. When the electronic device 100 detects a screenshot operation (e.g., the user taps the screen with knuckles for multiple times in succession) of the user on the screen, in response to the screenshot operation, the electronic device 100 performs screenshot processing on the browsing interface, and displays a screenshot image on a screenshot interface 11 as shown in Figure 1B Fig. 11.

[0046] It should be understood that the embodiments of the present application only exemplarily illustrate the way of the user to perform the screenshot, i.e., the embodiments of the present application take the way of the user to perform the screenshot by tapping the screen of the electronic device 100 with knuckles in succession. In fact, the screenshot can also be performed in other ways. For example, when the electronic device 100 detects a single tap operation of the user on the screenshot control, in response to the operation, the electronic device 100 performs the screenshot on the browsing interface 10. Alternatively, when the electronic device 100 detects an input operation of the user on the screen (three fingers slide downward along the screen at the same time), in response to the operation, the electronic device 100 performs the screenshot on the current browsing interface 10. The embodiments of the present application do not limit the way of the user to perform the screenshot, and the exemplarily illustration of the way of the screenshot by the embodiments of the present application should not constitute a limitation on the protection scope of the embodiments of the present application. As shown in Figure 1B Fig. 11 is a screenshot interface 11 of the electronic device 100, which is used to display Figure 1A the screenshot image 113 of the browsing interface 10 in Fig. 10, and further includes a text recognition icon 111. When the electronic device 100 detects an input operation (e.g., a single tap) of the user on the text recognition icon 111, in response to the operation, the electronic device 100 displays a text recognition interface 12 as shown in Figure 1C Fig. 12.

[0047] As shown in Figure 1C Fig. 12 is a text recognition interface 12 of the electronic device 100, which is used to display the text recognition result of the screenshot image 113 of the browsing interface 10 in Fig. 10, and further includes a text recognition result icon 121. When the electronic device 100 detects an input operation (e.g., a single tap) of the user on the text recognition result icon 121, in response to the operation, the electronic device 100 displays a text information interface 13 as shown in Figure 1BThe text recognized and extracted in the screenshot image 113. The text recognized and extracted in the screenshot image 113 by the electronic device 100 can be copied, can be selected all, can be shared, and can be searched. In order to distinguish the text that can be copied, selected all, shared, and searched by the user, the electronic device 100 can highlight the text recognized and extracted by the electronic device 100. For example, the text displayed in the gray rectangular bar is highlighted, and the method of highlighting the text by the electronic device 100 is not limited in the embodiments of the present application. In Figure 1C , the user can perform the operation of copying, selecting all, and the like on the highlighted text. For example, when the user selects the text in Figure 1C , the electronic device 100 displays the text recognition interface 13 as shown in Figure 1D .

[0048] As shown in Figure 1D , the text recognition interface 13 includes a text operation prompt box 131, and the text operation prompt box 131 includes a "copy" icon, a "select all" icon, a "translate" icon, a "share" icon, and a "search" icon. It should be understood that other function icons can also be included in the text operation prompt box 131, Figure 1D The embodiments are only illustrative. In Figure 1D , the text recognition interface 13 also includes a text selection icon 132, which is used to prompt the user that the text range has been selected. The text between the two text selection icons 132 is the selected text. In the above Figure 1D embodiment, the selected text is "Store name: Some grass shopping center Address: XX Province, XX City, XX District, XX Street, XX Shopping Center, XX, Time: Monday-Sunday 10:00-22:00 Transportation: Bus 346, XX station, Self-driving: The underground parking is relatively convenient, and if there is no parking space, the underground garage opposite XX office building can also be used". When the electronic device 100 detects an input operation (for example, a single click) on the "copy" icon in the text operation prompt box 131, the electronic device 100 copies the selected text in response to the operation.

[0049] As shown in Figure 1E , after the electronic device 100 copies the selected text in the text recognition interface 13, the electronic device 100 displays the message interface 14. The message interface 14 includes a text editing box 141. When the electronic device 100 detects an input operation (for example, a long press) on the text editing box 141, the electronic device 100 displays a text operation prompt box 142 in response to the operation. As shown inFigure 1E As shown, the text operation prompt box 142 includes function icons such as a "Paste" icon, a "Frequently Used Phrases" icon, and an "Auto Fill" icon. The text operation prompt box 142 may also include other function icons; this embodiment does not impose limitations. When the electronic device 100 detects an input operation (e.g., a click) on the "Paste" icon in the text operation prompt box 142, in response to this operation, the electronic device 100 will... Figure 1D The copied text in the Chinese character recognition interface 13 is displayed in the text editing box 141.

[0050] like Figure 1F As shown, electronic device 100 is in the above-mentioned Figure 1E In this embodiment, after detecting an input operation on the "Paste" icon in the text editing box 141, the SMS interface 15 is displayed. In the SMS interface 15, the text editing box 141 displays the text: "Store Name: [Name of Store] Grassland Shopping Center Address: [Address] No. [Number], [Street], [District], [City], [Province] Time: Monday-Sunday 10:00-22:00 Transportation: Bus 346, [Stop] Driving Directions: Underground parking is convenient; if there are no spaces, you can also use the underground parking garage of the [Building Name] office building across the street." In this way, users can directly copy and paste text information from a screenshot, thereby reducing editing time and improving the user experience.

[0051] In some embodiments, when recognizing text information in a screenshot, the electronic device 100 can extract feature information such as addresses and phone numbers from the screenshot and specially mark the feature information in the text recognition interface. After detecting a user's click operation on the feature information in the text recognition interface, the corresponding application can be launched to perform the target operation. For example, ... Figure 1G The above is shown. Figure 1B In this embodiment, another text recognition interface 16 corresponds to the screenshot interface 11. Similar to the above... Figure 1B In this embodiment, the operation of triggering the electronic device 100 to display recognized text is the same; when the electronic device 100 detects a target... Figure 1B After an input operation (e.g., clicking) is performed on the Chinese character recognition icon 111, the electronic device 100 displays the character in response to the operation. Figure 1GThe text recognition interface 16 is shown. Similar to the screenshot interface 11, the electronic device 100 highlights the recognized text, allowing the user to perform operations such as copying and pasting. In the text recognition interface 16, the electronic device 100 marks the address text information "Address: XX Province XX City XX District XX Street XX Shopping Center XX Number" and the telephone number "Telephone: 177XXX80XX9", that is, marking lines 161 and 162 are displayed below these two text segments. When the electronic device 100 detects an input operation (e.g., a click) on the text "Telephone: 177XXX80XX9", in response to the operation, the electronic device 100 displays as shown. Figure 1H The call prompt box 171 shown is an example. It should be understood that the electronic device 100 can mark the feature information in the text recognition interface 16 in various ways. Figure 1G The embodiments are merely illustrative examples of displaying marker lines below feature information and should not be construed as limiting the scope of protection of the embodiments of this application.

[0052] like Figure 1H As shown, the call prompt box 171 in the text recognition interface 17 is used to prompt the user whether to call the user with the phone number "177XXX80XX9". When the electronic device 100 detects an input operation (e.g., a click) on the call prompt box 171, in response to the operation, the electronic device 100 displays as shown in the image. Figure 1I The call interface 18 shown is as follows. Figure 1I As shown, the call interface 18 is the interface of the user call telephone number of the electronic device 100 as "177XXX80XX9".

[0053] In some embodiments, the user can crop the screenshot to obtain a cropped screenshot. Furthermore, the electronic device 100 can also highlight only the text within the cropped screenshot on the text recognition interface. For example... Figure 1J The above is shown. Figure 1Acorresponding to the browsing interface 10 in FIG. 1. The screenshot interface 19 includes a cropping function area 192 in which a plurality of icons of cropping shapes are included. For example, a freeform icon 1921, a rectangle icon 1922, an oval icon 1923, and a heart icon 1924. Selecting different cropping icons can crop the screenshot image into different shapes. For example, the electronic device 100 crops the screenshot image using the freeform icon 1921 to crop the screenshot image into an image of an arbitrary shape, the electronic device 100 crops the screenshot image using the rectangle icon 1922 to crop the screenshot image into a rectangular image, the electronic device 100 crops the screenshot image using the oval icon 1923 to crop the screenshot image into an oval image, and the electronic device 100 crops the screenshot image using the heart icon 1924 to crop the screenshot image into a heart-shaped image. Taking the electronic device 100 cropping the screenshot image using the oval icon 1923 as an example. When the electronic device 100 detects an input operation (e.g., a single tap) on the oval icon 1923, the electronic device 100 displays the screenshot area 191 on the screenshot interface 19 in response to the user input operation on the screenshot image (e.g., drawing an oval on the screenshot image). As shown in the screenshot interface 19, the screenshot image in the oval-shaped screenshot area 191 is the cropped screenshot image. After cropping the screenshot image, when the electronic device 100 detects an input operation (e.g., a single tap) on the text recognition icon 111, the electronic device 100 displays the text recognition interface 20 as shown in Figure 2A .

[0054] As shown in Figure 2A , only the text in the text recognition area 201 in the text recognition interface 20 is highlighted, that is, only the text in the text recognition area 201 can be copied, pasted, and the like. The text outside the text recognition area 201 is not highlighted. The size, shape, and position of the text recognition area 201 in the screenshot image are the same as those of the screenshot area 191 in FIG. 1. Figure 1J .

[0055] In some embodiments, after cropping the screenshot image, the electronic device 100 can display only the cropped screenshot image on the screenshot interface, or can display only the cropped screenshot image and the recognized text in the text recognition area. For example, in the above-mentioned Figure 1J embodiment, in response to the user input operation on the screenshot image (e.g., drawing an oval on the screenshot image), the electronic device 100 displays the screenshot interface 19A as shown in Figure 1K . Figure 1KAs shown, screenshot interface 19A only displays the screenshot image within screenshot area 191 of screenshot interface 19. When electronic device 100 detects an input operation on text recognition icon 111, in response to the operation, electronic device 100 displays as shown. Figure 1L The text recognition interface shown is 17A. (As shown in the image...) Figure 1L As shown, the text recognition interface 17A only displays the image corresponding to the screenshot interface 19A and the highlighted text "Mon-Sun 10:00-22:00 Bus 346 XXya Station Driving directions: Underground parking is convenient. If there are no parking spaces, you can also go to the underground parking garage of the XX office building diagonally opposite. Call: 177XXX80XX9". Compared to the above Figure 2A The text recognition interface 20 in the middle no longer displays the screenshot image outside the cropped screenshot image area.

[0056] In some embodiments, users can also take partial screenshots of the browsing interface, that is, screenshot a portion of the browsing interface. Users can take partial screenshots using gestures or knuckle gestures. For ease of understanding, this application embodiment illustrates the method of a user taking a partial screenshot of the browsing interface using their knuckle gesture. Figure 2B As shown, this is the browsing interface 21. When the electronic device 100 detects a user's knuckle input operation on the following video (e.g., drawing a screenshot area 212 on the browsing interface 21 with a knuckle), in response to this operation, the electronic device 100 displays as shown below. Figure 2C The screenshot shown is 19B. Figure 2C As shown, screenshot interface 19B is the screenshot interface corresponding to browsing interface 21. Screenshot interface 19B includes a target screenshot area 212, which is the area drawn by the user with their knuckle in browsing interface 21. When electronic device 100 detects a click operation on save control 202, it responds to the operation and can save an image containing only the screenshot image within the target screenshot area. If electronic device 100 detects an input operation (e.g., a click) on text recognition icon 211, it responds to the operation and displays as shown. Figure 2D The text recognition interface 20A is shown. In the text recognition interface 20A, only the text in the image of the target screenshot area in the screenshot interface 19B is highlighted.

[0057] In some embodiments, after a user takes a partial screenshot of the browsing interface, the screenshot interface can display only the screenshot of that portion, highlight only the text in the partial screenshot in the text recognition area, and display only the screenshot of the partial display area. For example, as... Figure 2E The screenshot interface 19C shown is the one mentioned above. Figure 2BIn this context, the electronic device 100 displays a screenshot interface in response to a user's knuckle input to the display screen (e.g., drawing a screenshot area 212 on the browsing interface 21 with a knuckle). This screenshot interface displays only the aforementioned... Figure 2B The image of the central screenshot area 212. When the electronic device 100 detects a target... Figure 2E After an input operation (e.g., clicking) is performed on the Chinese character recognition icon 211, the electronic device 100 displays the following in response to the operation: Figure 2F The text recognition interface 20B shown above. The text recognition interface 20B only displays the image corresponding to the screenshot interface 19C and the highlighted text "Mon-Sun 10:00-22:00 Bus 346 XXya Station towards: Underground parking is relatively convenient, if you hear it, the underground parking garage of the XX office building diagonally opposite is 177XXX80XX9". Compared to the above... Figure 2A The text recognition interface 20A no longer displays screenshots outside the cropped screenshot area.

[0058] In some embodiments, a user can draw on a screenshot image in the screenshot interface, thereby causing the electronic device 100 to highlight the text in the drawn screenshot image in the text recognition interface. For example... Figure 2G The screenshot shown is interface 22, which includes doodles 221 and 222. Doodle 221 covers the text "Shop Name: A Certain Grassland Shopping Center" in the browsing interface 10, and doodle 222 covers the text "Parking space, also available in the underground parking garage of the XX office building diagonally opposite" in the browsing interface 10. When the electronic device 100 detects an input operation (e.g., a click) on the text recognition icon 211, in response to the operation, the electronic device 100 displays as shown. Figure 2H The text recognition interface 23A is shown. In the text recognition interface 23A, the electronic device 100 highlights the text in the screenshot image that has not been drawn, and does not highlight the text that has been drawn; that is, the text that is not highlighted cannot be copied, selected, or otherwise manipulated. Furthermore, in the text recognition interface 23A, the electronic device 100 can also mark feature information in the highlighted text (for example, displaying an underline below the text). For example, underlines are displayed below the text "Address: XX Province XX City XX District XX Street XX Shopping Center XX Number" and the text "Telephone: 177XXX80XX9". Thus, when the electronic device 100 detects a click operation targeting the feature information, it can launch the relevant application. For example, as described above... Figures 1G-1H The example describes the dialing operation for the text "Telephone: 177XXX80XX9".

[0059] In some embodiments, the electronic device 100 can highlight all the text in the screenshot image after graffiti on the text recognition interface. For example... Figure 2JThe above is shown. Figure 2G The screenshot interface 22 corresponds to another text recognition interface 23B. In this text recognition interface, regardless of the... Figure 2G Whether the screenshot image has been graffitied, all the text in the text recognition interface 23A is highlighted.

[0060] In some embodiments, the electronic device 100 can highlight a portion of the text in a screenshot image after graffiti on a text recognition interface. For example... Figure 2I The above is shown. Figure 2G The screenshot interface 22 corresponds to another text recognition interface 23C. In this text recognition interface, all undrawn text is highlighted, and partially drawn text is also highlighted.

[0061] In some embodiments, the electronic device 100 performs text recognition on images stored in a gallery, thereby highlighting the recognized text in a text recognition interface. For example... Figure 2K The image shown is the gallery interface 24 of the electronic device 100, which includes multiple thumbnails. When the electronic device 100 detects an input operation (e.g., a click) on a thumbnail 241, in response to the operation, the electronic device 100 displays as shown below. Figure 2L The image browsing interface shown is 25.

[0062] like Figure 2L As shown, the gallery browsing interface 25 includes a preview image 251 and a text recognition icon 252. Users can draw on the preview image 251 to obtain a preview image with their drawings. Figure 2M As shown, the image browsing interface 26 displays preview images of the doodles, including doodles 261 and 262. When the electronic device 100 detects an input operation (e.g., a click) on the text recognition icon 252, in response to the operation, the electronic device 100 displays images as shown. Figure 2N The text recognition interface 27A is shown. In the text recognition interface 27A, the electronic device 100 highlights the text in the preview image that has not been drawn, and does not highlight the text that has been drawn; that is, the text that is not highlighted cannot be copied, selected, or otherwise manipulated. Furthermore, in the text recognition interface 27A, the electronic device 100 can also mark feature information in the highlighted text (for example, displaying an underline below the text). For example, underlines are displayed below the text "Address: XX Province XX City XX District XX Street XX Shopping Center XX Number" and the text "Telephone: 177XXX80XX9". Thus, when the electronic device 100 detects a click operation targeting the feature information, it can launch the relevant application. For example, the above... Figures 1G-1H The example describes the dialing operation for the text "Telephone: 177XXX80XX9".

[0063] In some embodiments, the electronic device 100 can highlight all the text in the preview image after graffiti on the text recognition interface 27B. For example... Figure 2O The above is shown. Figure 2M Another text recognition interface corresponding to the image browsing interface 26. In this text recognition interface 27B, regardless of... Figure 2O In the preview image, all text in the text recognition interface 27B is highlighted to check if it has been graffitied.

[0064] In some embodiments, the electronic device 100 can highlight a portion of the text in the preview image after graffiti in a text recognition interface. For example... Figure 2P The above is shown. Figure 2M Another text recognition interface 27C corresponds to the image browsing interface 26. In this text recognition interface, all undrawn text is highlighted, and partially drawn text is also highlighted.

[0065] The above Figures 1A-2P The application scenarios of a character recognition method provided in this application embodiment are illustrated below. The flow of a character recognition method provided in this application embodiment is described below with reference to the accompanying drawings. Please refer to... Figure 3A , Figure 3A This is a flowchart of a text recognition method provided in an embodiment of this application. The specific process is as follows:

[0066] Step 301A: In response to the first operation, the electronic device takes a screenshot of the browsing interface and displays a first interface, which includes a first screenshot image and a first control, and the text in the first screenshot image is displayed in image form.

[0067] Specifically, the first operation can be a screenshot operation. After responding to the user's screenshot operation, the electronic device can take a screenshot of the browsing interface. The browsing interface can be as described above. Figure 1A The browsing interface 10 in this embodiment. The screenshot operation can be a knuckle screenshot, a three-finger swipe screenshot, a single click on the screenshot control, or simultaneously pressing the power button and volume up / down buttons of an electronic device; this embodiment does not limit these methods. Furthermore, the screenshot type can be a long screenshot, a partial screenshot, a global screenshot, or a scrolling screenshot; this embodiment also does not limit the screenshot type. The first operation can be a global screenshot operation or a partial screenshot operation. If the first operation is a global screenshot operation, it can be one of the above... Figure 1A The user's screenshot operation in the embodiment (e.g., the user taps the screen multiple times with their knuckles). If the first operation is a partial screenshot operation, it can be as described above. Figure 2BThe knuckle of the user in the embodiment is for the input operation of the following video (for example, drawing a screenshot area 212 on the browsing interface 21 with the knuckle). The image obtained by the electronic device 100 is the first screenshot image. In the case that the first operation is a global screenshot, the first screenshot image can be the image in the browsing interface, which is the image in the browsing interface, for example, the image in the browsing interface in the above embodiment Figure 1B The screenshot image 113 in the screenshot interface 11 in the embodiment. In the case that the first operation is a local screenshot, the first screenshot image can be the image in the browsing interface, which is the image in the browsing interface, for example, the image in the browsing interface in the above embodiment Figure 2E The screenshot image displayed in the screenshot interface 19C in the embodiment; or, the first screenshot image in the above embodiment Figure 2C The image in the target screenshot area 212 displayed in the screenshot interface 19B.

[0068] Optionally, in some embodiments, the electronic device displays the images after the browsing interface is screenshot in the first interface in response to the first operation. The user can crop the images by the electronic device to obtain the cropped images, and the cropped images are the first screenshot image. For example, as described in the above embodiment Figure 2A As described in the embodiment, the electronic device 100 detects the input operation (for example, a single click) on the oval icon 1923, and in response to the input operation of the user on the screenshot image (for example, drawing an oval on the screenshot image), the electronic device 100 displays the screenshot area 191 on the screenshot interface 19, and the corresponding image in the screenshot area 191 is the first screenshot image.

[0069] The first interface is an interface for displaying the first screenshot image, and the first interface includes a first control for triggering the electronic device to display the text displayed in the first screenshot image in the form of text. For example, the first interface can be the above-mentioned Figure 1B The screenshot interface 11 in the embodiment, and the first control can be the above-mentioned Figure 1B The text recognition icon 211 in the embodiment.

[0070] In some embodiments, the electronic device can first perform text detection on the first screenshot image, and if the text is detected, the first control is displayed in the first interface, and if the text is not detected, the first control can not be displayed in the first interface.

[0071] Optionally, after obtaining the first screenshot image, the electronic device can temporarily store the first screenshot image in the buffer of the electronic device.

[0072] In some embodiments, after the browsing interface is screenshot, the electronic device displays the image of the browsing interface in the first interface, and the image of the browsing interface includes a screenshot area, and the corresponding image of the screenshot area is the first screenshot image. For example, as described in the above embodiment Figure 2CThe screenshot interface 19B is an interface in which the image 213 is an image of a browsing interface, and the image in the screenshot area 212 in the image 213 is a screenshot image of a partial screenshot of the browsing interface 21, i.e., an image of an actual screenshot.

[0073] Optionally, after the partial screenshot, the user can adjust the size and / or position of the screenshot area, so as to change the size of the first screenshot image and / or the position of the first screenshot image in the image of the browsing interface. For example, as shown in Figure 3B FIG. 19B is a screenshot interface 19B after a partial screenshot, when the electronic device 100 detects an input operation (e.g., dragging and stretching) of the user on the screenshot area 212, the electronic device 100 adjusts the size and position of the screenshot area in response to the operation, and displays a screenshot interface 30 as shown in Figure 3C FIG. 30. As shown in Figure 3C In the screenshot interface 30, the size and position of the screenshot area 212 are changed (the screenshot area 212 becomes larger and is positioned more forward) compared with those in the screenshot interface 19B, so that the size and position of the first screenshot image are changed.

[0074] Optionally, after the partial screenshot, the user can adjust the shape of the screenshot area. For example, in the above Figure 3B embodiment, the screenshot interface 19B has multiple cropping icons, including a rectangular cropping icon 2121, an oval cropping icon 2122, a heart-shaped cropping icon 2123, and a free-form cropping icon 2124. The user can change the shape of the screenshot area to the shape of the icon by clicking one of the multiple cropping icons. For example, in Figure 3A FIG. 19B, when the electronic device 100 detects that the user selects the screenshot area 212, and detects an input operation (e.g., clicking) on the rectangular cropping icon 2121, the electronic device 100 adjusts the shape of the screenshot area 212 to a rectangular screenshot area 212 as shown in Figure 3D FIG. 30 in response to the operation.

[0075] Step 302A: In response to the second operation, the electronic device displays a doodle on the first screenshot image to obtain a second screenshot image.

[0076] Specifically, the second operation can be a scribbling operation on the first screenshot image. The scribbling operation in the embodiments of the present application refers to drawing on the first screenshot image, covering the pixels in the first image, so that the user cannot visually distinguish the first image from the covered content. After the scribbling operation is performed on the first screenshot image, the electronic device 100 displays the scribbling on the first screenshot image. At this time, the first screenshot image with scribbling displayed on the first interface is referred to as a second screenshot image. For example, the second screenshot image can be the screenshot image displayed in the screenshot interface 22 in the embodiments of the present application, which includes the scribbling 221 and the scribbling 222. Figure 2H The screenshot image displayed in the screenshot interface 22 in the embodiments of the present application, which includes the scribbling 221 and the scribbling 222.

[0077] In the process of performing the scribbling on the first screenshot image, the electronic device records the position information of the scribbling in the first screenshot image, which can be coordinate data of the scribbling in the first screenshot image. The position information of the scribbling in the first screenshot image is scribbling data.

[0078] Optionally, after obtaining the second screenshot image, the electronic device can temporarily store the second screenshot image in the Buffer.

[0079] In some embodiments, before responding to the second operation, if the electronic device is a global screenshot, the screenshot interface can directly display a scribbling control, for example, the screenshot interface 22 in the embodiments of the present application Figure 2G includes the scribbling control 223. When the electronic device 100 detects an input operation on the scribbling control 223, the electronic device 100 enters a scribbling mode, and the user can perform the scribbling. If the electronic device is a local screenshot, the electronic device can display an editing control on the screenshot interface, and when the electronic device detects an input operation (for example, a single click) on the editing control, the electronic device can display a scribbling control on the screenshot interface, so that the user can perform the scribbling. For example, in the embodiments of the present application Figure 3B , the screenshot interface 19B includes the editing control 2125, and when the electronic device 100 detects an input operation (for example, a single click) on the editing control 2125, the electronic device 100 responds to the operation and displays the screenshot interface 32 as shown in Figure 3E , which includes the scribbling control 321. When the electronic device 100 detects an input operation on the scribbling control 321, the electronic device 100 enters a scribbling mode, and the user can perform the scribbling.

[0080] Step 303A: In response to a third operation on the first control, the electronic device extracts the text in the first screenshot image by a text recognition algorithm to obtain first data, the first data including position information of the extracted text.

[0081] Specifically, the third operation can be the text recognition operation described above Figure 2HIn an embodiment, the input operation is directed to the character recognition icon 211. The character recognition algorithm can be an Optical Character Recognition (OCR) algorithm. Through the character recognition algorithm, the electronic device can recognize and extract the characters in the first screenshot image, and obtain the position information of the extracted characters in the first interface, which is the first data. The position information of the extracted characters in the first screenshot image can be coordinate data of the extracted characters in the first interface. It should be understood that, in the case of global screenshot, the first screenshot image is the image of the entire browsing interface. At this time, the electronic device can extract all the characters in the entire browsing interface image, and record the position information (first data) of the extracted characters in the first interface. In the case of local screenshot, since the screenshot interface includes the image of the browsing interface and the screenshot region (the image corresponding to the screenshot region is the first screenshot image), the electronic device can only extract the characters of the image in the screenshot region, and obtain the position information (first data) of the extracted characters.

[0082] In some embodiments, for the case of local screenshot, the electronic device can also not recognize and extract the characters that do not completely fall into the screenshot region, or recognize and extract the part of the characters that fall into the screenshot region. For example, for the characters that do not completely fall into the screenshot region, the electronic device can calculate the proportion of the overlapping region of the characters in the screenshot region to the region occupied by the characters to determine whether to recognize the characters. If the proportion is greater than or equal to a preset proportion, the characters are completely recognized and extracted, and the position information of the characters is obtained. Otherwise, the characters are not recognized or only the part of the characters that fall into the screenshot region is recognized. Figure 3B In the screenshot interface 30 described above, the characters that partially fall into the screenshot region 212 are “Yu”, “He”, “Lai”, “Jie”, “Dao”, “X”, and “X”. Among these characters, the proportion of the overlapping region of “Yu” in the screenshot region 212 to “Yu” is 60%, the proportion of the overlapping region of “He” in the screenshot region 212 to “He” is 50%, the proportion of the overlapping region of “Lai” in the screenshot region 212 to “Lai” is 20%, the proportion of the overlapping region of “Jie” in the screenshot region 212 to “Jie” is 50%, the proportion of the overlapping region of “Dao” in the screenshot region 212 to “Dao” is 10%, and the proportion of the overlapping region of “X” in the screenshot region 212 to “X” is 60%. Assuming that the preset proportion is 40%, the electronic device 100 can completely recognize and extract “Yu”, “He”, “Jie”, and “X”, and obtain the position information of these characters. For the remaining characters, the electronic device 100 can not recognize and extract them, or partially recognize and extract them. In this way, the number of recognized and extracted characters can be reduced, and the computing resources of the electronic device can be saved.

[0083] In some embodiments, if part of the character information (for example, the character information can be: name, address, express number, phone number, ID number, age, etc.) is in the screenshot area, the electronic device can recognize the character information in the following three ways:

[0084] The first way: the degree of recognition y1 is identified by the formula y1 = Confidence * (S1 / S). Wherein y1 is the first degree of recognition of the character information, Confidence is the confidence of the character information, S1 is the overlapping area of the character information and the screenshot area, and S is the area of the character information. If y1 is greater than or equal to the first recognition threshold, all the characters in the character information are recognized. If y1 is less than the first recognition threshold, according to the local screenshot situation, the characters of the character information that do not completely fall into the screenshot area are recognized and extracted, or they can not be recognized and extracted, or part of them that falls into the screenshot area is recognized and extracted. The first recognition threshold can be obtained from historical values, empirical values, or experimental data, and the present application does not limit it.

[0085] The second way: the degree of recognition y2 is identified by the formula y2 = Confidence * (S2 / S). Wherein y2 is the second degree of recognition of the character information, Confidence is the confidence of the character information, S2 is the area of the characters in the character information that overlap with the screenshot area, and S is the area of the character information. If y2 is greater than or equal to the second recognition threshold, all the characters in the character information are recognized. If y2 is less than the second recognition threshold, according to the local screenshot situation, the characters of the character information that do not completely fall into the screenshot area are recognized and extracted, or they can not be recognized and extracted, or part of them that falls into the screenshot area is recognized and extracted. The second recognition threshold can be obtained from historical values, empirical values, or experimental data, and the present application does not limit it.

[0086] The third way: through the formula y3=Confidence*(M1 / M), identify the degree y3. Wherein, y3 is the third degree of identification of the feature information, Confidence is the confidence of the feature information, M1 is the number of characters in the feature information that have overlapping regions with the screenshot area, S is the number of characters in the feature information. If y3 is greater than or equal to the third identification threshold, all the characters in the feature information are identified. If y3 is less than the third identification threshold, according to the local screenshot situation, the characters of the feature information that do not completely fall into the screenshot area are identified and extracted, or these characters are not identified and extracted, or the part of these characters that falls into the screenshot area is identified and extracted. The third identification threshold can be obtained from historical values, or from empirical values, or from experimental data, and the embodiments of the present application do not limit.

[0087] Among the above three ways, in the feature information, the characters that are partially in the screenshot area are determined to have overlapping regions with the screenshot area only when the proportion of the overlapping part of the character with the screenshot area to the area occupied by the character is greater than or equal to the first overlap threshold.

[0088] In some embodiments, the electronic device can also directly input the first screenshot image after graffiti as the input of the character recognition algorithm, and perform character recognition and extraction, thereby obtaining the extracted characters and the first data.

[0089] The first control can be the character recognition icon 211 in the above Figure 2H embodiments.

[0090] Optionally, the electronic device can temporarily store the first data in the Buffer after obtaining the first data.

[0091] Step 304A: The electronic device determines the target character from the characters extracted from the first screenshot image according to the graffiti data and the first data.

[0092] Specifically, the electronic device can determine the target character from the characters extracted from the first screenshot image according to the graffiti data and the first data, and the target character is the character to be displayed in text form. The electronic device determines the target character in the following four ways:

[0093] The first way: the electronic device can directly extract the characters in the first screenshot image (without graffiti) as the target characters.

[0094] The second way: the electronic device can extract the characters from the first screenshot image after graffiti as the target characters.

[0095] The third mode: the electronic device can determine the text in the first screenshot image that has an overlapping area with the scribble according to the scribble data and the first data. Then, among the text that has an overlapping area with the scribble, the electronic device calculates whether the proportion of the overlapping area of each text with the scribble to the area of the text exceeds a first intersection threshold. In the first screenshot image, the electronic device determines the text whose proportion exceeds the first intersection threshold as the target text, and determines the other text as the target text. The first intersection threshold can be obtained by an experience value, a historical value, or experimental data, which is not limited in the embodiments of the present application.

[0096] The fourth mode: the electronic device can compare the text extracted from the first screenshot image after the scribble with the text extracted from the first screenshot image (without scribble). The text that is not consistent is determined as the non-target text, and the other extracted text is determined as the target text.

[0097] Step 305A: the electronic device displays a second interface in which a third screenshot image is displayed; wherein in the second interface, the target text is displayed in the form of text on the third screenshot image, and the non-target text is displayed in the form of an image on the third screenshot image.

[0098] Specifically, after determining the target text, the electronic device can display a second interface. The second interface is a text recognition interface of the electronic device. In the second interface, a third screenshot image is displayed, and the target text is displayed in the form of text in the third screenshot image, and the non-target text is displayed in the form of an image in the second screenshot image. The electronic device can perform editing operations such as copying and selecting all on the text displayed in the form of text.

[0099] For example, the second interface can be the text recognition interface 23A in the above Figure 2H , and the third screenshot image can be the screenshot image displayed in the text recognition interface 23A in the above Figure 2H .

[0100] In a possible implementation, if there is feature information in the target text, the electronic device can display mark information of the feature information in the second interface. For example, the mark information can be the mark line 161 and the mark line 162 in the text recognition interface 16 in the above Figure 1H . The electronic device detects an input operation for the target feature information, and in response to the operation, the electronic device starts a target application. The target application is an application corresponding to the target feature information. For example, when the target feature information is "Phone: 177XXX80XX9" in the text recognition interface 16 in the above Figure 1G , the target application is a call application.

[0101] In some embodiments, the electronic device can also adaptively determine the target text according to the sensitivity of the text and the intersection degree of the graffiti and the text, so as to display the recognized text in the form of text on the second screenshot image of the second interface. Specifically, the electronic device can extract the feature information of the text in the extracted first screenshot image by a feature information recognition algorithm, and determine the sensitivity of the extracted feature information. The feature information recognition algorithm includes a plurality of feature information recognition models. The feature information can be an ID number, a name, a courier number, an address, a telephone number, etc. Each type of feature information corresponds to a feature information recognition model. For example, if the electronic device needs to recognize five types of feature information, i.e. ID number, name, courier number, address, and telephone number, it needs an ID number corresponding feature information recognition model to recognize the ID number in the text extracted by the electronic device, a name corresponding feature information recognition model to recognize the name in the text extracted by the electronic device, a courier number corresponding feature information recognition model to recognize the courier number in the text extracted by the electronic device, an address corresponding feature information recognition model to recognize the address information in the text extracted by the electronic device, and a telephone number corresponding feature information recognition model to recognize the telephone number in the text extracted by the electronic device.

[0102] For ease of understanding, the electronic device extracts the text and identifies the address information by a feature information recognition model in the embodiments of the present application. As shown in Figure 3F It is a flow chart of a feature information recognition algorithm provided by the embodiments of the present application. In Figure 3F , the electronic device takes the text information (extracted text) as the input of the address feature information recognition model. Then, the address feature information recognition model processes the input text through a hot repair channel, a regular entity recognition module, a POI analysis model, and a standard address analysis module respectively, and outputs a first recognition result, a second recognition result, a third recognition result, and a fourth recognition result. Then, the address feature information recognition model corrects the second recognition result, the third recognition result, and the fourth recognition result entity error correction module, and outputs a first correction result. Finally, the address feature information fuses the first correction result and the first recognition result, and outputs the address information in the text information.

[0103] The electronic device can determine the intersection degree of the text and the scribble based on a proportion value of an area of the intersection of the feature information and the scribble to an area of the feature information. The intersection degree of the feature information and the scribble can be classified from small to large as a first intersection degree, a second intersection degree, a third intersection degree, and a fourth intersection degree. When the intersection degree of the text and the scribble is determined based on a proportion value of an area of the intersection of the text and the scribble to an area of the text, the first intersection degree is when the proportion value is greater than 0 and less than or equal to a first intersection threshold value (e.g., 30%). The second intersection degree is when the proportion value is greater than the first intersection threshold value and less than or equal to a second intersection threshold value (e.g., 50%). The third intersection degree is when the proportion value is greater than the second intersection threshold value and less than or equal to a third intersection threshold value. The fourth intersection degree is when the proportion value is greater than or equal to the third intersection threshold value. The first intersection threshold value is less than the second intersection threshold value, and the second intersection threshold value is less than the third intersection threshold value. The first intersection threshold value, the third intersection threshold value, and the second intersection threshold value can be obtained from historical values, can be obtained from empirical values, and can be obtained from experimental data, which are not limited by the embodiments of the present application. Preferably, the first intersection threshold value can be 30%, the second intersection threshold value can be 50%, and the third intersection threshold value can be 60%.

[0104] After the electronic device extracts the feature information by the feature information recognition algorithm, the electronic device can classify the feature information according to a preset sensitivity degree. For example, the feature information can be classified as a first sensitivity degree, a second sensitivity degree, and a third sensitivity degree from high to low. For example, the address and the ID number can be determined as the first sensitivity degree, the express number and the phone number can be determined as the second sensitivity degree, and the name and other text can be determined as the third sensitivity degree.

[0105] Then, the electronic device can determine the target text in the text extracted from the first screenshot according to the target text determination table. For example, the target text determination table can be shown in Table 1 as follows:

[0106] Table 1

[0107]

[0108] The electronic device can determine whether the text is the target text according to Table 1. The first way of determination means that the feature information corresponding to this way is determined as the target text according to the first way in step 304A. The second way of determination means that the feature information corresponding to this way is determined as the target text according to the second way in step 304A. The third way of determination means that the feature information corresponding to this way is determined as the target text according to the third way in step 304A. The fourth way of determination means that the feature information corresponding to this way is determined as the target text according to the fourth way in step 304A. For example, as shown in Table 1, the first way of determination is that the address is determined as the target text according to the first way in step 304A. The second way of determination is that the ID number is determined as the target text according to the second way in step 304A. The third way of determination is that the express number is determined as the target text according to the third way in step 304A. The fourth way of determination is that the phone number is determined as the target text according to the fourth way in step 304A.Figure 3G As shown in FIG. 1, the text scribbled in the image 1 is "XX sugar shop address: XX City XX District XX Street XX", assuming that the intersection degree of this text and the scribble is a three-level intersection degree, the sensitive degree of the address "XX City XX District XX Street XX" is a first-level sensitive degree, and the sensitive degree of the other text is a three-level sensitive degree. Then, according to the above-mentioned text determination method in Table 1, "XX City XX District XX Street XX" is displayed in the form of an image (not highlighted) in the image 2, and "XX sugar shop address:" is displayed in the form of text (highlighted).

[0109] In some embodiments, the target text can also be determined by a formula, for example, the display confidence y of the text can be calculated by the formula y = Sscale*((Sarea∩Tarea) / Tarea). Wherein, Sscale is the sensitive degree of the text, the first-level sensitive degree is a (for example, 1), the second-level sensitive degree is b (for example, 0.6), and the third-level sensitive degree is c (for example, 0.3). Wherein, a > b > c. Sarea is the scribble area, Tarea is the text area, Sarea∩Tarea is the intersection of the scribble area and the text area; if y > m, the text is not determined as the target text, otherwise, the text is determined as the target text. Wherein, m can be obtained by experimental value, or can be obtained by experience value, or can be obtained by experimental data, and the embodiments of the present application do not make any limitation.

[0110] In some embodiments, in the screenshot interface, the change of the screenshot area will also cause the change of the range of the text displayed in the form of text in the second interface. For example, in the above-mentioned Figure 3C embodiment, since the screenshot area 212 in the screenshot interface 30 has changed, the range of the text displayed in the form of text in the corresponding text recognition interface has also changed. For example, as Figure 3H shown in FIG. 1, in the above-mentioned Figure 3B , the text recognition interface 30 after the screenshot interface 19B is scribbled. As Figure 3I shown in FIG. 1, in the above-mentioned Figure 3C , the text recognition interface 31 after the screenshot interface 30 is scribbled. As Figure 3H and Figure 3I can be seen, since the screenshot area 212 has changed, the range of the text (highlighted) displayed in the form of text in the image of the browsing interface has also changed.

[0111] In some embodiments, in the text recognition interface, the electronic device can also adjust the size of the screenshot area, so that the text displayed in the form of text in the screenshot area will also change with the adjustment of the size of the screenshot area. For example, Figure 3JThe screenshot area of the illustrated character recognition interface 32 is 212. When an input operation (for example, zooming in the screenshot area 212) directed to the screenshot area is detected, the electronic device 100 zooms in the screenshot area 212 in response to the operation, and displays the character recognition interface 33 as illustrated. Figure 3K As illustrated, with the zooming in of the screenshot area 212, the characters displayed in text form (highlighted characters) also increase. Figure 3K As illustrated, with the zooming in of the screenshot area 212, the characters displayed in text form (highlighted characters) also increase.

[0112] The character recognition method provided in the embodiments of the present application temporarily caches the image after the user performs the screenshot on the browsing interface. After the user performs the doodling on the screenshot image, if the character recognition is to be performed on the screenshot image after the doodling, the electronic device can recognize and extract the characters in the screenshot image before the doodling. Then, the recognized and extracted characters are displayed in text form according to the preset display rule, so that the user can perform the copying, pasting and other operations on the characters displayed in text form. Since the electronic device uses the image that is not doodled to recognize and extract the characters, compared with the recognition and extraction of the characters using the image after the doodling, the characters recognized and extracted by the electronic device in the embodiments of the present application are more accurate.

[0113] The above Figure 3A The embodiments introduce the flow of the character recognition method provided in the embodiments of the present application. Next, the interaction flow of the modules of the electronic device in the above Figure 4 embodiments is described. Please refer to Figure 3A , Figure 4 , Figure 4 is another flowchart of a character recognition method provided in the embodiments of the present application. In the flowchart, the electronic device includes a screenshot application, an information processing module, a graphic-text recognition algorithm engine, a doodling recognition module, an information sensitivity degree judgment engine and an entity recognition algorithm engine. The specific flow is as follows:

[0114] Step 401: The screenshot application responds to the first operation to perform the screenshot on the browsing interface, displays the first interface, the first interface includes the first screenshot image and the first control, and the characters in the first screenshot image are displayed in image form.

[0115] Step 401 can refer to the related description in the above step 301A, and will not be described here.

[0116] Step 402: The screenshot application responds to the second operation to send the first indication information to the doodling recognition algorithm engine.

[0117] Specifically, the second operation can be a scribbling operation on the first screenshot image. After detecting the second operation, the screenshot application sends first indication information to the scribble recognition algorithm engine. The first indication information is used to instruct the scribble recognition algorithm engine to record the position information of the scribble in the first screenshot image. The scribbling operation in the embodiments of the present application refers to scribbling on the first screenshot image, covering the pixels in the first image, so that the user cannot visually distinguish the covered content of the first image.

[0118] Step 403: The scribble recognition algorithm engine calculates the scribble data.

[0119] Specifically, after receiving the first indication information, the scribble recognition algorithm engine calculates the position information of the scribble on the first screenshot image, which can be coordinate data. The position information calculated by the scribble recognition algorithm engine is the scribble data.

[0120] Step 404: The scribble recognition algorithm engine sends the scribble data to the screenshot application.

[0121] Specifically, after calculating the scribble data, the scribble recognition algorithm engine sends the scribble data to the screenshot application.

[0122] Step 405: The screenshot application controls the display screen to display the scribble on the first screenshot image according to the scribble data, to obtain a second screenshot image.

[0123] Specifically, after receiving the scribble data sent by the scribble recognition algorithm engine, the screenshot application controls the display screen to display the scribble on the first screenshot image according to the scribble data, to obtain a second screenshot image.

[0124] Step 406: The screenshot application sends a first character recognition instruction to the information processing module in response to a third operation on the first control.

[0125] Specifically, the third operation can be the input operation on the character recognition icon 211 in the above Figure 2H In the embodiments, the input operation on the character recognition icon 211. After detecting the third operation on the first control, the screenshot application sends a first character recognition instruction to the information processing module. The first character recognition instruction is used to instruct the information processing module to recognize and extract the characters in the first screenshot image.

[0126] The first character recognition instruction can include address information of the first screenshot image.

[0127] Step 407: The information processing module sends the first recognition instruction to the graphic recognition algorithm engine.

[0128] Specifically, after receiving the first text recognition instruction sent by the screenshot application, the information processing module can send a first recognition instruction to the graphic-text recognition algorithm engine. The first recognition instruction is used to instruct the graphic-text recognition algorithm engine to recognize and extract the text in the first screenshot image. The first recognition instruction can include address information of the first screenshot image.

[0129] Optionally, the first recognition instruction is also used to instruct the graphic-text recognition algorithm engine to send the text information (the text recognized and extracted from the first screenshot image) to the feature information recognition engine.

[0130] Step 408: The graphic-text recognition algorithm engine extracts the text in the first screenshot image according to the text recognition algorithm to obtain first data, and the first data includes position information of the extracted text.

[0131] Specifically, after receiving the first recognition instruction, the graphic-text recognition algorithm engine can obtain the first screenshot image from the buffer according to the address information of the first screenshot image in the first recognition instruction, extract the text in the first screenshot image by the text recognition algorithm, and obtain first data including position information of the extracted text. The position information is the first data, and the position information of the extracted text in the first screenshot image can be coordinate data of the extracted text in the first screenshot image.

[0132] In some embodiments, the screenshot application can crop the screenshot image according to the instruction, and display the cropped region (the first screenshot image) and the uncropped region of the screenshot image on the first interface after cropping the screenshot image. The first text recognition instruction sent by the screenshot application to the information processing module can also include position information of the cropped region. The first recognition instruction sent by the information processing module to the graphic-text recognition algorithm engine can also include the position information of the cropped region. In this way, after the graphic-text recognition algorithm engine receives the first recognition instruction, it can only recognize and extract the text in the cropped region of the first screenshot image according to the position information of the cropped region in the first recognition instruction.

[0133] It should be understood that, in the case of global screenshot, the first screenshot image is an image of the entire browsing interface. At this time, the graphic-text recognition algorithm engine can extract all the text in the entire browsing interface image and record the position information (first data) of the extracted text in the first interface. In the case of local screenshot, since the screenshot interface includes the image of the browsing interface and the screenshot region (the image corresponding to the screenshot region is the first screenshot image), the graphic-text recognition algorithm engine can only extract the text of the image in the screenshot region and obtain the position information (first data) of the extracted text.

[0134] In some embodiments, for the case of partial screenshot, the text recognition algorithm engine can recognize and extract the text whose edges do not completely fall within the screenshot area, or can also choose not to recognize and extract such text, or recognize and extract only the part of such text that falls within the screenshot area. Exemplarily, for the text that does not completely fall within the screenshot area, the text recognition algorithm engine can calculate the proportion of the overlapping area of the text and the screenshot area to the area occupied by the text to determine whether to recognize the text. If the proportion value is greater than or equal to the preset proportion value, the text is completely recognized and extracted, and the position information of the text is obtained. Otherwise, the text is not recognized or only the part of the text that falls within the screenshot area is recognized. Exemplarily, in the screenshot interface 30 as described above Figure 3B , the words that partially fall within the screenshot area 212 are respectively "于", "和", "来", "街", "道", "X", "X". Among these words, the proportion of the overlapping area of "于" and the screenshot area 212 to "于" is 60%, the proportion of the overlapping area of "和" and the screenshot area 212 to "和" is 50%, the proportion of the overlapping area of "来" and the screenshot area 212 to "来" is 20%, the proportion of the overlapping area of "街" and the screenshot area 212 to "街" is 50%, the proportion of the overlapping area of "道" and the screenshot area 212 to "道" is 10%, and the proportion of the overlapping area of "X" and the screenshot area 212 to "X" is 60%. Assuming that the preset proportion value is 40%, for the words "于", "和", "街", "X", the text recognition algorithm engine 100 can completely recognize, extract, and obtain their position information. For the remaining words, the text recognition algorithm engine 100 can choose not to recognize, extract, or partially recognize and extract. In this way, the number of recognized and extracted words can be reduced, saving the computing resources of the text recognition algorithm engine.

[0135] In some embodiments, the text recognition algorithm engine can also directly use the first screenshot image after scribbling as the input of the text recognition algorithm to perform text recognition and extraction, so as to obtain the extracted text and the first data.

[0136] The text recognition algorithm can be an Optical Character Recognition (OCR) algorithm. Through the text recognition algorithm, the graphic algorithm recognition engine can recognize and extract the text in the first screenshot image, and obtain the position information of the extracted text in the first screenshot image, and this position information is the first data. The position information of the extracted text in the first screenshot image can be the coordinate data of the extracted text in the first screenshot image.

[0137] Step 409: The text recognition algorithm engine sends the extracted text and the first data to the information processing module.

[0138] Specifically, after recognizing and extracting the text in the first screenshot image, the image-text recognition algorithm engine can send the first data and the extracted text to the information processing module.

[0139] Step 410: The image-text recognition algorithm engine sends the text information to the feature information recognition algorithm engine.

[0140] Specifically, after extracting the text in the first screenshot image, the image-text recognition algorithm engine can also send the extracted text as text information to the feature information recognition algorithm module.

[0141] It should be understood that step 409 can be executed before step 410, after step 410, or simultaneously with step 410, and the embodiments of the present application do not limit this.

[0142] Step 411: The feature information recognition algorithm engine extracts feature information according to the text information.

[0143] Specifically, the feature information recognition algorithm engine extracts feature information according to the text information sent by the image-text recognition algorithm engine. The feature information recognition algorithm engine can extract feature information by passing the extracted text in the first screenshot image through the feature information recognition algorithm. The feature information recognition algorithm includes a plurality of feature information recognition models. The feature information can be an ID number, a name, a courier number, an address, a telephone number, etc. Each type of feature information corresponds to a feature information recognition model. For example, if the feature information recognition algorithm engine needs to recognize five types of feature information, i.e., ID number, name, courier number, address, and telephone number, it needs to have an ID number corresponding feature information recognition model to recognize the ID number in the extracted text, a name corresponding feature information recognition model to recognize the name in the extracted text, a courier number corresponding feature information recognition model to recognize the courier number in the extracted text, an address corresponding feature information recognition model to recognize the address information in the extracted text, and a telephone number corresponding feature information recognition model to recognize the telephone number in the extracted text.

[0144] The feature information recognition algorithm engine can extract the feature information by inputting the extracted text in the first screenshot image into the feature information recognition algorithm, and determine the sensitivity of the extracted feature information. The feature information recognition algorithm includes a plurality of feature information recognition models. The feature information can be an ID number, a name, a courier number, an address, a telephone number, etc. Each type of feature information corresponds to a feature information recognition model. For example, if the feature information recognition algorithm engine needs to identify five types of feature information, i.e., ID number, name, courier number, address, and telephone number, it needs an ID number corresponding feature information recognition model to identify the ID number in the text extracted by the feature information recognition algorithm engine, a name corresponding feature information recognition model to identify the name in the text extracted by the feature information recognition algorithm engine, a courier number corresponding feature information recognition model to identify the courier number in the text extracted by the feature information recognition algorithm engine, an address corresponding feature information recognition model to identify the address information in the text extracted by the feature information recognition algorithm engine, and a telephone number corresponding feature information recognition model to identify the telephone number in the text extracted by the feature information recognition algorithm engine.

[0145] For ease of understanding, the feature information recognition algorithm engine will be taken as an example to identify the address information by inputting the extracted text into the feature information recognition model, and the embodiments of the present application will be described. As shown in Figure 3F Fig. 1 is a feature information recognition algorithm flowchart provided by the embodiments of the present application. In Figure 3F , the feature information recognition algorithm engine inputs the text information (extracted text) as the input of the address feature information recognition model. Then, the address feature information recognition model processes the input text through the hot repair channel, the regular entity recognition module, the POI analysis model, and the standard address analysis module respectively, and outputs the first recognition result, the second recognition result, the third recognition result, and the fourth recognition result. Then, the address feature information recognition model corrects the second recognition result, the third recognition result, and the fourth recognition result through the entity error correction module, and outputs the first correction result. Finally, the address feature information fuses the first correction result and the first recognition result, and outputs the address information in the text information.

[0146] The feature information recognition algorithm engine can extract the feature information according to the text information, which can refer to the related description of step 304A in the above Figure 3A embodiment, which will not be described here.

[0147] Step 412: The feature information recognition algorithm engine sends the feature information to the information processing module.

[0148] Step 412A: The feature information recognition algorithm engine sends the feature information to the information sensitivity judgment engine.

[0149] Step 412B: The information sensitivity judgment engine judges the sensitivity of the feature information.

[0150] Specifically, after receiving the feature information, the information sensitivity judgment engine can classify the feature information according to the preset sensitivity. For example, the feature information can be classified into first-level sensitivity, second-level sensitivity, and third-level sensitivity from high to low. For example, the address and ID number can be determined as first-level sensitivity, the express number and telephone number can be determined as second-level sensitivity, and the name and other characters can be determined as third-level sensitivity.

[0151] Step 412C: The information sensitivity judgment engine sends the sensitivity information of the feature information to the information processing module.

[0152] Steps 412A-412C are optional steps.

[0153] Step 413: The information processing module determines the target text according to the graffiti data and the text extracted from the first screenshot image.

[0154] Specifically, the information processing module can determine the target text according to the graffiti data and the text extracted from the first screenshot image. The target text is the text to be displayed in text form. The information processing module determines the target text in the following four ways:

[0155] The first way: The information processing module can directly extract the text in the first screenshot image (without graffiti) as the target text.

[0156] The second way: The information processing module can extract the text from the first screenshot image after graffiti as the target text.

[0157] The third way: The information processing module can determine the text in the first screenshot image that has an overlapping area with the graffiti according to the graffiti data and the first data. Then, in the text that has an overlapping area with the graffiti, it is determined whether the proportion of the overlapping area of the text with the graffiti to the area of the text exceeds a first intersection threshold value. In the first screenshot image, the information processing module does not determine the text whose proportion value exceeds the first intersection threshold value as the target text, and determines the other text as the target text. The first intersection threshold value can be obtained from experience value, historical value, or experimental data, which is not limited by the embodiments of the present application.

[0158] The fourth way: The information processing module can compare the text extracted from the first screenshot image after graffiti with the text extracted from the first screenshot image (without graffiti). The inconsistent text is determined as non-target text, and the other extracted text is determined as target text.

[0159] In some embodiments, the information processing module can further determine the target text adaptively according to the sensitivity of the text and the intersection degree of the text and the graffiti, and display the recognized text in text form on the second screenshot image of the second interface. Specifically:

[0160] The information processing module can determine the intersection degree of the text and the graffiti based on the proportion value of the area of the intersection of the feature information and the graffiti to the area of the feature information. The intersection degree of the feature information and the graffiti can be divided into first, second, third and fourth intersection degrees from small to large. When the intersection degree of the text and the graffiti is determined based on the proportion value of the area of the intersection of the text and the graffiti to the area of the text, when the proportion value is greater than 0 and less than or equal to a first intersection threshold (for example, 30%), it is a first intersection degree. When the proportion value is greater than the first intersection threshold and less than or equal to a second intersection threshold (for example, 50%), it is a second intersection degree. When the proportion value is greater than the second intersection threshold and less than or equal to a third intersection threshold, it is a third intersection degree. When the proportion value is greater than or equal to the third intersection threshold, it is a fourth intersection degree. The first intersection threshold is less than the second intersection threshold, and the second intersection threshold is less than the third intersection threshold. The first intersection threshold, the third intersection threshold and the second intersection threshold can be obtained from historical values, can also be obtained from empirical values, and can also be obtained from experimental data. The present application does not limit. Preferably, the first intersection threshold can be 30%, the second intersection threshold can be 50%, and the third intersection threshold can be 60%.

[0161] Then, the information processing module can determine the target text from the text extracted from the first screenshot image according to the target text determination table. For example, the target text determination table can be shown in Table 2 as follows:

[0162] Table 2

[0163]

[0164]

[0165] The information processing module can determine whether the text is the target text according to the above Table 2. Wherein, the first way means that the corresponding feature information is determined to be the target text according to the first way in the above step 304A. The second way means that the corresponding feature information is determined to be the target text according to the second way in the above step 304A. The third way means that the corresponding feature information is determined to be the target text according to the third way in the above step 304A. The fourth way means that the corresponding feature information is determined to be the target text according to the fourth way in the above step 304A.

[0166] Step 414: The information processing module controls the display screen to display a second interface, and a third screenshot image is displayed in the second interface; wherein in the second interface, the target text is displayed in the form of text on the third screenshot image, and the non-target text is displayed in the form of a picture on the third screenshot image.

[0167] The above Figures 3A-4 The above Figure 5 The above Figure 5 The above Figure 5 is a flowchart of another text recognition method provided by the embodiments of the present application. The specific process is as follows:

[0168] Step 501: In response to a first input operation, the electronic device displays a first browsing interface, the first browsing interface includes a first image and a first control, and the text in the first image is displayed in the form of an image.

[0169] For example, the first input operation can be the input operation (for example, a single click) on the thumbnail 241 in the above Figure 2K The first browsing interface can be the gallery browsing interface 25 shown in the above Figure 2L The first image can be the preview image 251, and the first control can be the text recognition icon 252.

[0170] Step 502: In response to a second input operation, the electronic device displays a scribble on the first screenshot image to obtain a second image.

[0171] Specifically, the second input operation can be a scribbling operation on the first image. After the scribbling operation on the first image, the electronic device 100 displays the scribble on the first image. At this time, the first image with the scribble displayed on the first browsing interface is called the second image. For example, the second image can be the preview image displayed in the image browsing interface 26 in the above Figure 2M The preview image includes the scribble 261 and the scribble 262.

[0172] During the scribbling on the first image, the electronic device records the position information of the scribble in the first image, which can be coordinate data of the scribble in the first image. The position information of the scribble in the first image is the scribble data.

[0173] Step 503: In response to a third input operation on the first control, the electronic device extracts the text in the first image by a text recognition algorithm to obtain first data, and the first data includes the position information of the extracted text.

[0174] Specifically, the third input operation can be the input operation on the character recognition icon 252. Figure 2M In an embodiment, the input operation is on the character recognition icon 252. The character recognition algorithm can be an optical character recognition (OCR) algorithm. Through the character recognition algorithm, the electronic device can recognize and extract the character in the first image, and obtain the position information of the extracted character in the first image as the first data. The position information of the extracted character in the first image can be coordinate data of the extracted character in the first image.

[0175] Step 504: The electronic device determines a target character from the scribble data and the character extracted in the first image according to the first data.

[0176] Step 504 can refer to the related description of the electronic device determining a target character from the scribble data and the character extracted in the first screenshot image according to the first data in step 304A, and will not be repeated here.

[0177] Step 505: The electronic device displays a second browsing interface, and displays a third image in the second browsing interface; wherein in the second browsing interface, the target character is displayed in text form on the third image, and the non-target character is displayed in picture form on the third image.

[0178] The character recognition method provided in the embodiments of the present application can be used to recognize and extract the character in the preview image before the user scribbles on the preview image. Then, according to the preset display rule, the recognized and extracted character is displayed in text form, so that the user can perform copy, paste and other operations on the character displayed in text form. Since the electronic device uses the image before the scribbling to recognize and extract the character, compared with using the image after the scribbling to recognize and extract the character, the character recognized and extracted by the electronic device in the embodiments of the present application is more accurate.

[0179] The above Figure 5 The embodiments introduce the process of the character recognition method provided in the embodiments of the present application. Next, the character recognition method provided in the embodiments of the present application will be described in combination with Figure 6 The above Figure 5 In the embodiments, the interaction process of each module of the electronic device will be described. Please refer to Figure 6 , Figure 6 is another process flowchart of the character recognition method provided in the embodiments of the present application. In the flowchart, the electronic device includes a first application, an information processing module, a graphic-text recognition algorithm engine, a scribble recognition module and an entity recognition algorithm engine. The specific process is as follows:

[0180] Step 601: The first application displays a first browsing interface in response to a first input operation, the first browsing interface including a first image and a first control, and text in the first image being displayed in the form of an image.

[0181] Specifically, the first application can be a gallery application or other application for browsing pictures.

[0182] Step 602: The first application sends first indication information to a graffiti recognition algorithm engine in response to a second input operation.

[0183] Step 603: The graffiti recognition algorithm engine calculates graffiti data.

[0184] Step 604: The graffiti recognition algorithm engine sends the graffiti data to the first application.

[0185] Step 605: The first application controls the display screen to display graffiti on the first image according to the graffiti data, obtaining a second image.

[0186] Step 606: The first application sends a first text recognition instruction to an information processing module in response to a third input operation on the first control.

[0187] Step 607: The information processing module sends the first recognition instruction to a graphic recognition algorithm engine.

[0188] Step 608: The graphic recognition algorithm engine extracts text in the first image according to a text recognition algorithm, obtaining first data, the first data including position information of the extracted text.

[0189] Step 609: The graphic recognition algorithm engine sends the extracted text and the first data to the information processing module.

[0190] Step 610: The graphic recognition algorithm engine sends text information to a feature information recognition algorithm engine.

[0191] Step 611: The feature information recognition algorithm engine extracts feature information according to the text information.

[0192] Step 612: The feature information recognition algorithm engine sends the feature information to the information processing module.

[0193] Step 612A: The feature information recognition algorithm engine sends the feature information to an information sensitivity degree judgment engine.

[0194] Step 612B: The information sensitivity degree judgment engine judges the sensitivity degree of the feature information.

[0195] Specifically, after receiving the feature information, the information sensitivity degree judgment engine can classify the feature information according to the preset sensitivity degree. For example, the feature information can be classified into first-level sensitivity, second-level sensitivity and third-level sensitivity from high to low according to the sensitivity degree. For example, the address and the ID number can be determined as the first-level sensitivity, the express number and the telephone number can be determined as the second-level sensitivity, and the name and other characters can be determined as the third-level sensitivity.

[0196] Step 612C: The information sensitivity degree judgment engine sends the sensitivity degree information of the feature information to the information processing module.

[0197] Steps 612A-612C are optional steps.

[0198] Step 613: The information processing module determines the target text according to the graffiti data and the text extracted from the first image according to the first data.

[0199] Step 614: The information processing module controls the display screen to display a second browsing interface, and displays a third screenshot image in the second browsing interface; wherein in the second browsing interface, the target text is displayed in the form of text on the third screenshot image, and the non-target text is displayed in the form of picture on the third screenshot image.

[0200] Steps 602-614 can refer to the above Figure 4 related descriptions of steps 402-414 in the embodiment, which will not be repeated here.

[0201] The structure of the electronic device 100 will be introduced below. Please refer to Figure 7 , Figure 7 is a hardware structure schematic diagram of the electronic device 100 provided in the embodiment of the present application.

[0202] The electronic device 100 can include a processor 110, an external memory interface 120, an internal memory 121, a universal serial bus (USB) interface 130, a charging management module 140, a power management module 141, a battery 142, an antenna 1, an antenna 2, a mobile communication module 150, a wireless communication module 160, an audio module 170, a speaker 170A, a receiver 170B, a microphone 170C, a headset interface 170D, a sensor module 180, a key 190, a motor 191, an indicator 192, a camera 193, a display screen 194, and a subscriber identification module (SIM) card interface 195, etc. The sensor module 180 can include a pressure sensor 180A, a gyroscope sensor 180B, a barometric pressure sensor 180C, a magnetic sensor 180D, an acceleration sensor 180E, a distance sensor 180F, a proximity light sensor 180G, a fingerprint sensor 180H, a temperature sensor 180J, a touch sensor 180K, an ambient light sensor 180L, a bone conduction sensor 180M, etc.

[0203] It can be understood that the structure shown in the embodiments of the present application does not constitute a specific limitation on the electronic device 100. In other embodiments of the present application, the electronic device 100 can include more or fewer components, or combine certain components, or split certain components, or different component arrangements. Figure 7 The components shown can be implemented in hardware, software, or a combination of software and hardware. Figure 7 The components shown can be implemented in hardware, software, or a combination of software and hardware.

[0204] The processor 110 can include one or more processing units, for example: the processor 110 can include an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a memory, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU), etc. Different processing units can be independent devices, or can be integrated in one or more processors.

[0205] The wireless communication function of the electronic device 100 can be implemented through the antenna 1, the antenna 2, the mobile communication module 150, the wireless communication module 160, the modem processor, and the baseband processor, etc.

[0206] Antennas 1 and 2 are used for transmitting and receiving electromagnetic wave signals. Each antenna in electronic device 100 can be used to cover a single or multiple communication frequency bands. Different antennas can also be multiplexed to improve the utilization of antennas. For example, antenna 1 can be multiplexed as a diversity antenna for wireless local area networks. In some other embodiments, antennas can be used in combination with tuning switches.

[0207] Mobile communication module 150 can provide solutions for wireless communication including 2G / 3G / 4G / 5G, etc. applied on electronic device 100. Mobile communication module 150 can include at least one filter, switch, power amplifier, low noise amplifier (LNA), etc. Mobile communication module 150 can receive electromagnetic waves by antenna 1, and perform filtering, amplification, etc. on the received electromagnetic waves, and transmit the processed signals to a modem processor for demodulation. Mobile communication module 150 can also amplify signals modulated by the modem processor, and convert the signals into electromagnetic waves radiated by antenna 1. In some embodiments, at least part of the functional modules of mobile communication module 150 can be arranged in processor 110. In some embodiments, at least part of the functional modules of mobile communication module 150 can be arranged in the same device as at least part of the modules of processor 110.

[0208] Wireless communication module 160 can provide solutions for wireless communication including wireless local area networks (WLAN) (such as Wi-Fi networks), Bluetooth (BT), BLE broadcasting, global navigation satellite system (GNSS), frequency modulation (FM), near field communication (NFC), infrared technology (IR), etc. applied on electronic device 100. Wireless communication module 160 can be one or more devices integrated with at least one communication processing module. Wireless communication module 160 receives electromagnetic waves via antenna 2, performs frequency modulation and filtering processing on the electromagnetic wave signals, and transmits the processed signals to processor 110. Wireless communication module 160 can also receive signals to be transmitted from processor 110, perform frequency modulation and amplification on the signals, and convert the signals into electromagnetic waves radiated by antenna 2.

[0209] The electronic device 100 implements a display function through a GPU, a display screen 194, and an application processor, etc. The GPU is a microprocessor for image processing, connected to the display screen 194 and the application processor. The GPU is used to perform mathematical and geometric calculations for graphics rendering. The processor 110 can include one or more GPUs that execute program instructions to generate or change display information.

[0210] The display screen 194 is used to display images, videos, etc. The display screen 194 includes a display panel. The display panel can adopt a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active-matrix organic light-emitting diode (AMOLED), a flex light-emitting diode (FLED), a Miniled, a MicroLed, a Micro-oLed, a quantum dot light emitting diode (QLED), etc. In some embodiments, the electronic device 100 can include 1 or N display screens 194, N being a positive integer greater than 1.

[0211] The electronic device 100 can implement a shooting function through an ISP, a camera 193, a video codec, a GPU, a display screen 194, and an application processor, etc.

[0212] The ISP is used to process data fed back by the camera 193. For example, when taking a photo, the shutter is opened, the light is transmitted to the camera photosensitive element through the lens, the light signal is converted into an electrical signal, and the camera photosensitive element transmits the electrical signal to the ISP for processing to convert it into an image visible to the naked eye. The ISP can also optimize algorithms for image noise, brightness, and skin color. The ISP can also optimize parameters such as exposure and color temperature of the shooting scene. In some embodiments, the ISP can be provided in the camera 193.

[0213] The digital signal processor is used to process digital signals, in addition to being able to process digital image signals, it can also process other digital signals. For example, when the electronic device 100 selects a frequency point, the digital signal processor is used to perform Fourier transform on the frequency point energy, etc.

[0214] The NPU is a neural-network (NN) computing processor that quickly processes input information by drawing on the structure of a biological neural network, such as the transmission mode between human brain neurons, and can also continuously self-learn. Through the NPU, the electronic device 100 can implement intelligent cognitive applications such as image recognition, face recognition, voice recognition, text understanding, and the like.

[0215] The electronic device 100 can implement audio functions through the audio module 170, the speaker 170A, the receiver 170B, the microphone 170C, the earphone interface 170D, and the application processor, and the like. For example, music playing, recording, and the like.

[0216] The audio module 170 is configured to convert digital audio information into an analog audio signal for output, and is also configured to convert an analog audio input into a digital audio signal. The audio module 170 can also be configured to encode and decode audio signals. In some embodiments, the audio module 170 can be disposed in the processor 110, or some functional modules of the audio module 170 can be disposed in the processor 110.

[0217] The speaker 170A, also referred to as a "loudspeaker", is configured to convert an audio electrical signal into a sound signal. The electronic device 100 can listen to music or listen to a hands-free call through the speaker 170A.

[0218] The receiver 170B, also referred to as an "earpiece", is configured to convert an audio electrical signal into a sound signal. When the electronic device 100 is on a call or receiving a voice message, the user can listen to the voice by holding the receiver 170B close to the ear.

[0219] The microphone 170C, also referred to as a "microphone", "sound transducer", is configured to convert a sound signal into an electrical signal. When making a call or sending a voice message, the user can make a sound by holding the mouth close to the microphone 170C, and input the sound signal into the microphone 170C. The electronic device 100 can be provided with at least one microphone 170C. In other embodiments, the electronic device 100 can be provided with two microphones 170C, in addition to collecting sound signals, the electronic device 100 can also implement a noise reduction function. In other embodiments, the electronic device 100 can also be provided with three, four or more microphones 170C, to collect sound signals, reduce noise, and also identify the source of the sound, to implement a directional recording function, and the like.

[0220] The pressure sensor 180A is configured to sense a pressure signal and convert the pressure signal into an electrical signal. In some embodiments, the pressure sensor 180A can be disposed on the display screen 194.

[0221] Touch sensor 180K, also referred to as a "touch panel." Touch sensor 180K can be disposed on display screen 194, with touch sensor 180K and display screen 194 forming a touch screen, also referred to as a "touch screen." Touch sensor 180K is configured to detect touch operations applied to or near touch sensor 180K. Touch sensor 180K can pass detected touch operations to the application processor to determine the type of touch event. Visual output related to the touch operation can be provided via display screen 194. In other embodiments, touch sensor 180K can be disposed on a surface of electronic device 100, in a location different from where display screen 194 is disposed.

[0222] The software system of electronic device 100 can employ a layered architecture, an event-driven architecture, a microkernel architecture, a microservices architecture, or a cloud architecture. Embodiments of the present application take an Android system with a layered architecture as an example to illustrate the software structure of electronic device 100. Figure 8 is a software structure block diagram of electronic device 100 according to an embodiment of the present application. A layered architecture divides software into several layers, each of which has a clear role and division of labor. Layers communicate with each other through software interfaces. In some embodiments, an Android system is divided into four layers, from top to bottom, an application layer, an application framework layer, an algorithm engine layer, an Android runtime and system library, and a kernel layer.

[0223] The application layer can include a series of application packages. As shown in Figure 8 , the application packages can include camera, first application, calendar, call, map, navigation, WLAN, information processing module, screenshot application, and the like.

[0224] The application framework layer provides application programming interfaces (APIs) and programming frameworks for the applications of the application layer. The application framework layer includes a number of pre-defined functions. As shown in Figure 8 , the application framework layer can include a window manager, a content provider, a view system, a phone manager, a resource manager, a notification manager, and the like.

[0225] The window manager is used to manage window programs. The window manager can obtain the size of the display screen, determine whether there is a status bar, lock the screen, and take a screenshot, and the like.

[0226] The content provider is used to store and obtain data, and make the data accessible to applications. The data can include videos, images, audio, dialed and received calls, browsing history and bookmarks, phonebook, and the like.

[0227] The view system includes visual controls, such as controls that display text, controls that display pictures, and the like. The view system can be used to build an application. A display interface can be composed of one or more views. For example, a display interface that includes a short message notification icon can include a view that displays text and a view that displays a picture.

[0228] The phone manager is used to provide the communication function of the electronic device 100. For example, the management of the call state (including the connection, hang-up, and the like).

[0229] The resource manager provides various resources for the application, such as localized strings, icons, pictures, layout files, video files, and the like.

[0230] The notification manager enables the application to display notification information in the status bar, which can be used to convey a notification type of message that can automatically disappear after a short stay without user interaction. For example, the notification manager is used to notify the completion of the download, the message reminder, and the like. The notification manager can also be a notification that appears in the top status bar of the system in the form of a chart or a scroll bar text, such as a notification of an application running in the background, and can also be a notification that appears on the screen in the form of a dialog window. For example, the text information is prompted in the status bar, a prompt sound is emitted, the electronic device is vibrated, the indicator light is blinked, and the like.

[0231] The algorithm engine layer includes a graphic-text recognition algorithm engine, a scribble recognition algorithm engine, a feature information recognition algorithm engine, and an information sensitivity judgment engine. The graphic-text recognition algorithm engine is used to recognize and extract image text. The scribble recognition algorithm engine is used to obtain the position information of the scribble in the image. The feature information recognition algorithm engine is used to extract the feature information in the text information. The information sensitivity judgment engine is used to judge the sensitivity of the feature information and send the sensitivity information of the feature information to the information processing module.

[0232] The Android runtime includes a core library and a virtual machine. The Android runtime is responsible for the scheduling and management of the Android system.

[0233] The core library includes two parts: one part is the function function that the java language needs to call, and the other part is the core library of Android.

[0234] The application layer and the application framework layer run in the virtual machine. The virtual machine executes the java file of the application layer and the application framework layer into a binary file. The virtual machine is used to perform the management of the object life cycle, the stack management, the thread management, the security and exception management, and the garbage collection, and the like.

[0235] The system library can include a plurality of functional modules. For example, a surface manager, media libraries, a three-dimensional graphics processing library (e.g., OpenGL ES), a 2D graphics engine (e.g., SGL), and the like.

[0236] The surface manager is used to manage a display subsystem and provides fusion of 2D and 3D layers for a plurality of applications.

[0237] The media libraries support playback and recording of a plurality of commonly used audio, video formats, and still image files. The media libraries can support a plurality of audio and video encoding formats, such as MPEG4, H.264, MP3, AAC, AMR, JPG, PNG, and the like.

[0238] The three-dimensional graphics processing library is used to implement three-dimensional graphics drawing, image rendering, composition, and layer processing, and the like.

[0239] The 2D graphics engine is a drawing engine for 2D drawing.

[0240] The kernel layer is a layer between hardware and software. The kernel layer at least includes display drivers, camera drivers, audio drivers, and sensor drivers.

[0241] In the above embodiments, all or part of the embodiments can be implemented by software, hardware, firmware, or any combination thereof. When implemented by software, all or part of the embodiments can be implemented in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the present application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transferred from one computer-readable storage medium to another computer-readable storage medium, for example, the computer instructions can be transferred from one website, computer, server, or data center to another website, computer, server, or data center through wired (such as coaxial cable, optical fiber, digital subscriber line) or wireless (such as infrared, wireless, microwave, etc.) mode. The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server, data center, etc. that includes one or more available media sets. The available media can be a magnetic medium (e.g., a floppy disk, a hard disk, a magnetic tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid state disk), etc.

[0242] The steps in the method of the embodiments of the present application can be sequentially adjusted, combined, and deleted according to actual needs.

[0243] The modules in the device embodiments of the present application can be combined, divided, and deleted according to actual needs.

[0244] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be instructed by a computer program to relevant hardware, and the program can be stored in a computer readable storage medium. When the program is executed, the program can include the processes of the above-mentioned method embodiments. The aforementioned storage medium includes ROM or random storage memory RAM, magnetic disc or optical disc, and various storage program codes.

[0245] In summary, the above only describes the embodiments of the technical solutions of the present application, and is not used to limit the protection scope of the present application. Any modification, equivalent replacement, improvement, etc. made according to the disclosure of the present application shall be included in the protection scope of the present application.

Claims

1. A text recognition method, characterized by, The method is applied to an electronic device, and the method comprises: displaying a first image and a first control in a first interface, wherein text in the first image is displayed in the form of an image; in response to a scribbling operation, displaying a second image on the first interface, wherein the second image is obtained by scribbling on the first image, scribbling marks are overlaid on the second image, and text in the second image is displayed in the form of an image; in a case where the second image is displayed, in response to a second operation on the first control, the electronic device extracts the text in the first image by using a text recognition algorithm; displaying a second interface, wherein the second interface comprises a third image, the third image is obtained based on the second image, and all or part of the text in the third image is displayed in the form of text; after the scribbling operation, obtaining position information of the scribbling marks in the first interface; before the second interface is displayed, the method further comprises: adaptively determining target text based on the position information of the scribbling marks in the first interface, a sensitive degree of the text, and an intersection degree of the scribbling and the text; the target text is text displayed in the form of text in the third image among the text extracted by the electronic device; wherein the intersection degree of the scribbling and the text is determined by a proportion value of an area of a region of the feature information that overlaps with the scribbling relative to the region of the feature information; the feature information is extracted from the extracted text in the first image by using a feature information recognition algorithm; and the sensitive degree of the text is obtained by classifying the feature information according to a preset sensitive degree.

2. The method of claim 1, wherein, before the first image and the first control are displayed in the first interface, the method further comprises: displaying a browsing interface; in response to a screenshot operation, performing a screenshot on the browsing interface; obtaining an image of the browsing interface, wherein the image is the first image.

3. The method of claim 1, wherein, before the first image and the first control are displayed in the first interface, the method further comprises: displaying a browsing interface; in response to a screenshot operation, performing a screenshot on a first region in the browsing interface; obtaining an image of the browsing interface, wherein a corresponding image of the first region in the image of the browsing interface is the first image.

4. The method of claim 3, wherein, after the first image and the first control are displayed in the first interface, the method further comprises: in the first interface, in response to a third operation on the first region in the image of the browsing interface; transforming the first region in the image of the browsing interface into a second region, wherein the first region is different from the second region; and an image in the second region is an updated first image.

5. The method of claim 3, wherein, the first interface further comprises a second control for indicating a first shape; after the first image and the first control are displayed in the first interface, the method further comprises: detecting a fourth operation on the second control; in response to the fourth operation, transforming a shape of the first region into the first shape; and an image in the first region after the shape is transformed is an updated first image.

6. An electronic device, comprising: comprise: a memory, a processor, and a touch screen; wherein: the touch screen is used to display content; the memory is used to store a computer program, and the computer program comprises program instructions; The processor is configured to invoke the program instructions to cause the electronic device to perform the method of any one of claims 1-5.

7. A computer readable storage medium characterized by The computer readable storage medium stores a computer program, which, when executed by a processor, implements the method of any one of claims 1-5.

Citation Information

Patent Citations

  • Content identification method and device and mobile terminal

    CN109085982A

  • Image processing method and device, computer equipment and storage medium

    CN111126301A

  • Information processing apparatus and recording medium

    CN112463010A

  • Hand writing input device, program and hand-writing input method system

    CN1493961A