Image processing method, device, computer equipment and storage medium

By identifying and generating editable target images in instant messaging applications, the problem of image text cannot be edited in chat sessions is solved, and the user can directly edit image text in chat sessions is realized.

CN114332887BActive Publication Date: 2025-08-29TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202210003009.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2019-12-26
Publication Date
2025-08-29
Estimated Expiration
2039-12-26

Smart Images

  • Figure CN114332887B_ABST
    Figure CN114332887B_ABST
Patent Text Reader

Abstract

An embodiment of the present invention discloses an image processing method, apparatus, computer equipment and storage medium, which can display a chat session page of an instant messaging client, wherein the chat session page includes an original image sent by a chat session user; based on an image text recognition operation on the original image, a recognition result page of the original image is displayed, and the recognition result page includes a target image, and the target image includes: text recognized from the original image, and background content corresponding to the text, the text is editable text, and the background content is content other than the text in the original image; when an editing operation on the text in the target image is detected, the editing result of the text is displayed, thereby, the original image sent by the chat session user in the chat session can be identified as a target image with editable text, and the user can directly edit the text in the target image to obtain the required editing result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of Internet technology, and in particular to an image processing method, apparatus, computer equipment, and storage medium. Background Art

[0002] IM (Instant Messaging) applications are software that enables online chatting and communication based on instant messaging technology. In addition, instant messaging applications also provide image recognition functions for images sent by users in chat session pages. This image recognition function can perform text recognition on images sent by users, making it easier for users to use the text recognition results corresponding to the images. Summary of the Invention

[0003] Embodiments of the present invention provide an image processing method, apparatus, computer device, and storage medium that can recognize an original image sent by a user in a chat session as a target image with editable text, so that the user can edit the text recognition result of the original image on the target image.

[0004] An embodiment of the present invention provides an image processing method, the method comprising:

[0005] Displaying a chat session page of an instant messaging client, wherein the chat session page includes an original image sent by a user of the chat session;

[0006] Based on the image text recognition operation on the original image, displaying a recognition result page of the original image, the recognition result page including a target image, the target image including: text recognized from the original image, and background content corresponding to the text, the text being editable text, and the background content being content in the original image other than the text;

[0007] When an editing operation on the text in the target image is detected, the editing result of the text is displayed.

[0008] This embodiment further provides an image processing device, comprising:

[0009] a conversation page display unit, configured to display a chat conversation page of an instant messaging client, wherein the chat conversation page includes an original image sent by a chat conversation user;

[0010] a recognition result display unit, configured to display a recognition result page of the original image based on the image text recognition operation on the original image, wherein the recognition result page includes a target image, and the target image includes: text recognized from the original image and background content corresponding to the text, wherein the text is editable text and the background content is content in the original image other than the text;

[0011] The editing result display unit is configured to display the editing result of the text when an editing operation on the text in the target image is detected.

[0012] This embodiment further provides a storage medium having a computer program stored thereon, wherein when the computer program is executed by a processor, the steps of the image processing method shown in the embodiment of the present invention are implemented.

[0013] This embodiment also provides a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the steps of the image processing method shown in the embodiment of the present invention are implemented.

[0014] An embodiment of the present invention provides an image processing method, apparatus, computer device and storage medium, which can display a chat session page of an instant messaging client, wherein the chat session page includes an original image sent by a chat session user; based on an image text recognition operation on the original image, a recognition result page of the original image is displayed, and the recognition result page includes a target image, and the target image includes: text recognized from the original image, and background content corresponding to the text, the text is editable text, and the background content is content other than the text in the original image; when an editing operation on the text in the target image is detected, the editing result of the text is displayed, thereby, the original image sent by the chat session user in the chat session can be identified as a target image with editable text, and the user can directly edit the text in the target image, obtaining an editing experience similar to editing text in the original image. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative work.

[0016] Figure 1a is a schematic diagram of a scenario of an image processing method provided by an embodiment of the present invention;

[0017] Figure 1b is a flowchart of an image processing method provided by an embodiment of the present invention;

[0018] Figure 2a This is a schematic diagram of displaying a recognition result page provided by an embodiment of the present invention;

[0019] Figure 2b is a schematic diagram of displaying another recognition result page provided by an embodiment of the present invention;

[0020] Figure 2c This is a schematic diagram of another display of a recognition result page provided by an embodiment of the present invention;

[0021] Figure 2d This is a schematic diagram showing a quick operation on an image provided by an embodiment of the present invention;

[0022] Figure 2e This is a schematic diagram showing a quick operation on an image provided by an embodiment of the present invention;

[0023] Figure 2f This is a schematic diagram showing a quick operation on an image provided by an embodiment of the present invention;

[0024] Figure 2g This is a schematic diagram showing a quick operation on an image provided by an embodiment of the present invention;

[0025] Figure 2h This is a schematic diagram showing a quick operation on an image provided by an embodiment of the present invention;

[0026] Figure 3a is a schematic diagram of text modification of a target image provided by an embodiment of the present invention;

[0027] Figure 3b This is a schematic diagram of sharing part of the text of a target image provided by an embodiment of the present invention;

[0028] Figure 3c is a schematic diagram of text sharing of a target image provided by an embodiment of the present invention;

[0029] Figure 3d is a schematic diagram of image sharing of a target image provided by an embodiment of the present invention;

[0030] Figure 3e This is a schematic diagram of image sharing based on translation of a target image provided by an embodiment of the present invention;

[0031] Figure 3f This is an optional schematic diagram of a sharing settings page provided by an embodiment of the present invention;

[0032] Figure 3g is another optional schematic diagram of a sharing settings page provided by an embodiment of the present invention;

[0033] Figure 4a 2 is a schematic diagram showing a text extraction result page for a target image provided by an embodiment of the present invention;

[0034] Figure 4b 2 is a schematic diagram showing a display of a comparison page corresponding to a text extraction result page provided by an embodiment of the present invention;

[0035] Figure 5a This is a flow chart of an image processing method provided by an embodiment of the present invention;

[0036] Figure 5b is another flowchart of the image processing method provided by an embodiment of the present invention;

[0037] Figure 5c This is a schematic diagram of an optional process for performing coarse classification on an original image in an embodiment of the present invention;

[0038] Figure 6 is a structural diagram of an image processing device provided by an embodiment of the present invention;

[0039] Figure 7 is a schematic structural diagram of a computer device provided by an embodiment of the present invention;

[0040] Figure 8 This is an optional structural diagram of the distributed system 800 provided in an embodiment of the present invention applied to a blockchain system;

[0041] Figure 9 This is an optional schematic diagram of a block structure provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0042] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without making any creative efforts shall fall within the scope of protection of the present invention.

[0043] Embodiments of the present invention provide an image processing method, apparatus, computer device, and storage medium. Specifically, embodiments of the present invention provide an image processing apparatus suitable for a first computer device (for distinction, referred to as a first image processing apparatus), wherein the first computer device may be a terminal or other device, wherein the terminal may be a mobile phone, a tablet computer, a laptop computer, or other device. Embodiments of the present invention also provide an image processing apparatus suitable for a second computer device (for distinction, referred to as a second image processing apparatus), wherein the second computer device may be a network-side device such as a server, wherein the server may be a single server or a server cluster composed of multiple servers, and may be a physical server or a virtual server.

[0044] For example, the first image processing device may be integrated into a terminal, and the second image processing device may be integrated into a server.

[0045] The embodiment of the present invention takes the first computer device as a terminal and the second computer device as a server as an example to introduce the image processing method.

[0046] refer to Figure 1a An embodiment of the present invention provides an image processing system including a terminal 10 and a server 20, etc.; the terminal 10 and the server 20 are connected via a network, such as a wired or wireless network connection, etc., wherein the first image processing device is integrated in the terminal, for example, integrated in the terminal in the form of a client.

[0047] Among them, the terminal 10 can be used to display a chat session page of an instant messaging client, wherein the chat session page includes an original image sent by a chat session user; based on an image text recognition operation on the original image, a recognition result page of the original image is displayed, and the recognition result page includes a target image, and the target image includes: text recognized from the original image, and background content corresponding to the text, the text is editable text, and the background content is the content in the original image other than the text; when an editing operation on the text in the target image is detected, the editing result of the text is displayed.

[0048] Among them, the target image corresponding to the original image can be generated by the server 20, and the terminal can obtain the target image by sending an image recognition request carrying the original image to the server 20 when it needs to obtain the target image; the server 20 can be specifically used to: receive the image recognition request sent by the terminal; obtain the original image sent by the terminal based on the image recognition request, perform text recognition on the original image, and obtain the text recognition result of the original image, wherein the text recognition result includes the text recognized from the original image and the text position of the recognized text in the original image, and the recognized text is used to replace the text at the corresponding text position in the original image in the form of editable text to obtain the target image corresponding to the original image, and send the target image to the terminal 10.

[0049] After receiving the target image, the terminal 10 may display a recognition result page, wherein the recognition result page includes the target image.

[0050] In one embodiment, after obtaining the text recognition result, the server may send the text recognition result to the terminal, and the terminal generates a target image of the original image based on the text recognition result and the original image.

[0051] It should be noted that the order of description of the following embodiments is not intended to limit the preferred order of the embodiments.

[0052] The embodiment of the present invention will be described from the perspective of a first image processing device, which may be integrated into a terminal.

[0053] An embodiment of the present invention provides an image processing method, which can be executed by a processor of a terminal, such as Figure 1b As shown, the process of the image processing method can be as follows:

[0054] 101. Displaying a chat session page of an instant messaging client, wherein the chat session page includes an original image sent by a user of the chat session;

[0055] To facilitate understanding of the content of this embodiment, some technical terms appearing in this embodiment are explained:

[0056] Instant messaging: A terminal service that allows two or more people to instantly exchange text messages, files, voice, and video messages over the Internet. Typical examples include mobile QQ, WeChat, WhatsApp, and other instant messaging tools.

[0057] Image OCR: The full name is Optical Character Recognition, which refers to the process by which electronic devices use character recognition methods to translate the shapes on an image into computer text.

[0058] In the embodiment of the present invention, the chat session page of the instant messaging client can be a single chat session page, a group chat session page, or a chat session page with a public account, and this embodiment does not limit this. The chat session user who sends the original image can be the current user of the terminal, that is, the user currently logged into the terminal, or another user in a chat session with the current user on the chat session page, and this embodiment does not limit this.

[0059] In this embodiment, the original image can be any type of image, such as a JGP image, an emoticon image, etc.; the content carried in the original image is not limited, and may include content in the form of tables, text, pictures, etc. The source of the original image is not limited, and may be an image obtained by screenshot or by photography, etc.

[0060] For example, in one embodiment, the original image may be a screenshot image obtained by a chat session user performing a screenshot operation on the chat session page, or the original image may be an image captured by a camera of a terminal during the chat session.

[0061] 102. Based on the image text recognition operation for the original image, a recognition result page of the original image is displayed, the recognition result page includes a target image, and the target image includes: text recognized from the original image, and background content corresponding to the text, the text is editable text, and the background content is the content in the original image other than the text.

[0062] It can be understood that in this embodiment, the recognition result page of the original image will only be displayed when there is text in the original image and the text is recognized from the original image. If no text can be recognized from the original image (such as there is no text in the original image, or the text recognition of the original image fails), the recognition result page will not be displayed.

[0063] In this embodiment, with respect to the case where text can be recognized from the original image, given that some text in the original image may be difficult to recognize, the text recognized from the original image is not necessarily completely identical to the original text in the original image. However, it is understandable that the distribution of the editable text and background content in the target image in this embodiment on the target image is similar to the distribution of the original text and other content other than the original text in the original image on the original image. In one embodiment, the target image can be understood as being obtained by replacing the original text in the original image with the text recognized from the original image on the basis of the original image. One difference between the original image and the target image is that the text in the target image recognized from the original image is editable, while the original text in the original image is not editable.

[0064] In this embodiment, the image text recognition operation may be a specific touch operation, such as a long press operation, a double-click operation, a sliding operation, etc. Optionally, the image text recognition operation may also be a combination of a series of operations, which is not limited in this embodiment.

[0065] For example, reference Figure 2a The recognition result page shows a schematic diagram, Figure 2a In the chat session page shown in 201, the friend A of the current user of the terminal sends image A to the current user. The image A is the original image mentioned above. In the page shown in 201, the original image A is displayed in a thumbnail state. The user can perform an image text recognition operation on the original image. Based on the image text recognition operation on the original image, a recognition result page as shown in 202 can be displayed. The recognition result page contains text and illustrations. The text is the text recognized from the original image A, and the text has an editable feature.

[0066] Optionally, in this embodiment, the recognition result page may also include text markers, each of which corresponds to the text of a text area in the target image. The text marker may be an underline, a color mark, a text box, etc. The text area may be divided into text rows or text columns (depending on the arrangement of the text in the original image). Figure 2a In 202 , each line of text corresponds to a text box including the text line, and the text in the text box can be edited as a whole, such as forwarding, copying, modifying, etc.

[0067] Optionally, in this embodiment, the step of “displaying a recognition result page of the original image based on the image text recognition operation on the original image” may include:

[0068] Displaying an image text recognition control based on a control display operation for the original image;

[0069] When a trigger operation for the image text recognition control is detected, a recognition result page of the original image is displayed.

[0070] The control in this embodiment may be in the form of an icon, an input box, a button, or the like.

[0071] In this embodiment, the control display operation can be a touch operation on the original image, such as double-clicking, long pressing, etc. The control display operation can also be triggered by voice.

[0072] Optionally, based on the control display operation for the original image, in addition to displaying the image text recognition control, other controls for the original image may also be displayed, such as a forwarding control for forwarding the original image when triggered, an editing control for editing the original image when triggered, etc. This embodiment does not impose any restrictions on this.

[0073] Optionally, there are multiple ways to display the image text recognition control.

[0074] (1) Based on the operation of the original image in the full-screen display state;

[0075] Optionally, the step of “displaying an image text recognition control based on a control display operation for the original image” may include:

[0076] When a display operation for the original image is detected, displaying an image magnification page of the original image, the image magnification page including the original image in a full-screen display state;

[0077] When a control display triggering operation for the original image is detected on the image magnification page, an image text recognition control is displayed.

[0078] In this embodiment, when a control display trigger operation for the original image is detected on the image magnification page, in addition to displaying the image text recognition control, other controls may also be displayed, such as a collection control for adding the original image to an image collection set.

[0079] Among them, when a control display trigger operation for the original image is detected on the image magnification page, displaying the image text recognition control may include: when a control display trigger operation for the original image is detected on the image magnification page, displaying a sub-page on the image magnification page, and the sub-page includes displaying the image text recognition control.

[0080] For example, reference Figure 2b The recognition result page shown in the figure is displayed. Figure 2b In the chat session page shown in 201, the friend A of the current user of the terminal sends image A to the current user. The image A is the original image mentioned above. In the page shown in 201, the original image A is displayed in a thumbnail state. When a display operation such as a click operation on the original image A is detected on the page shown in 210, the image A is displayed as shown in FIG. Figure 2b The image magnification page shown in the middle reference numeral 203 includes the original image in full-screen display. When a control display trigger operation, such as a long press operation, is detected on the image magnification page, an image text recognition control, such as the "Extract Text from Image" control on the page shown in 204, is displayed. When a trigger operation, such as a click operation, is detected on the "Extract Text from Image" control, the recognition results page shown in 202 is displayed.

[0081] The original image display operation can also include double-clicking, long-pressing, and other operations, which are not limited in this embodiment. The image recognition control can be displayed in a small window or, as shown in 204, in a sub-page. It is understood that other functional controls can also be displayed in this sub-page, such as a "Friends" control for sharing the original image with associated users. In this embodiment, the associated users are users in the current user's instant messaging client address book.

[0082] In one embodiment, when a trigger operation for the image text recognition control is detected, displaying a recognition result page of the original image includes:

[0083] When a trigger operation for the image text recognition control is detected, a recognition waiting page of the original image is displayed, wherein the recognition waiting page includes the original image and a recognition result loading icon;

[0084] When the recognition of the original image is successful, a recognition result page of the original image is displayed.

[0085] (2) Based on the operation of the original image of the chat reply page;

[0086] Optionally, the step of “displaying an image text recognition control based on a control display operation for the original image” may include:

[0087] When a control display operation for the original image is detected, a function control list corresponding to the original image is displayed on the chat session page, and the function control list includes an image text recognition control.

[0088] In this embodiment, the control display operation for the original image can be a long press, circle, or other operation on the original image. In addition to the image text recognition control, the function list can also include other controls, such as a forwarding control for forwarding the original image, etc.

[0089] For example, reference Figure 2c The recognition result page shown in the figure is displayed. Figure 2c In the chat session page shown in 201, friend A of the current user of the terminal sends image A to the current user. When a control display operation, such as a long press or double-click, is detected on the page shown in 201 for the original image A, a function control list 2011 is displayed on the chat session page shown in 201. This function control list includes an image text recognition control, such as a control named "Text Recognition". When a trigger operation, such as a click operation, is detected for the "Text Recognition" control in function control list 2011, a recognition waiting page for the original image, shown in 205, is displayed. The recognition waiting page includes the original image and a recognition result loading icon, such as an "Extracting Text" icon. When recognition of the original image is successful, the recognition result page, shown in 202, is displayed.

[0090] In one example, when a trigger operation such as a click operation is detected for the "text recognition" control in the function control list 2011, the recognition waiting page of the original image shown in 205 may not be displayed. Instead, when the recognition of the original image is successful, the recognition result page shown in 202 may be directly displayed.

[0091] This embodiment takes into account that users may have similar preferences for certain image content. For example, for ID card photos, users may prefer to extract ID card numbers, while for bank card photos, users may prefer to extract bank card numbers. Inspired by these situations, this embodiment provides quick operations for original images, reducing user image manipulation time and allowing users to quickly obtain desired results.

[0092] Optionally, in this embodiment, the image magnification page further includes: a quick operation control corresponding to the target content in the original image, wherein the quick operation control is used to perform an operation indicated by the quick operation control on the target content when triggered.

[0093] The target content can be set by the user or by the developer of the instant messaging client, and this embodiment does not limit this. The target content can include: various ID cards such as bank cards, ID cards, driver's licenses, or code images such as QR codes and bar codes, or documents with specific formats such as airline tickets, express delivery orders, and tax bills.

[0094] In one embodiment, the original image may be classified, and based on the image type of the original image, a quick operation control corresponding to the image type of the original image may be determined. In this embodiment, each quick operation control may be provided with corresponding target content.

[0095] Optionally, when a display operation on the original image is detected, displaying an image magnification page of the original image may include:

[0096] When a display operation for the original image is detected, triggering image type recognition of the original image to obtain the image type of the original image;

[0097] An image magnification page of the original image is displayed, wherein the image magnification page includes a quick operation control corresponding to the image type, and the quick operation control is used to perform the operation indicated by the quick operation control on the target content in the original image when triggered.

[0098] For example, reference Figure 2dAssuming the original image is a photo of China XXX Bank, the image magnification page 203 of the original image displays a quick operation control corresponding to the bank photo, such as a number extraction control named "Extract Number." When a trigger operation is detected for the number extraction control, a number extraction result page for the original image is displayed. The number extraction result page includes a number extraction result image. The number extraction result image includes: the number recognized from the original image and background content corresponding to the number. The number is editable text, and the background content is the content of the original image other than the number. In addition to the extracted number, a text box corresponding to the number is displayed on the number extraction result page. When a trigger operation is detected for the text box, such as a click operation, a list of function controls for the text box is displayed. The function control list includes function controls such as a copy control, a forward control, and an edit control. When a control in the function control list is triggered, it operates on the content in the clicked text box. For example, clicking the copy control in the function list will add the number in the text box, such as 6224XXXXXXXXXXXXXXX, to the copied content collection for subsequent use.

[0099] For example, reference Figure 2e , assuming that the original image includes text content, such as English text, the quick operation control displayed in the image magnification page can be a translation control, such as Figure 2e The control in the dialog is called "Translate Text in Image".

[0100] For example, refer to Figure 2f , assuming that the original image includes text content, the quick operation control displayed in the image magnification page can be an image text recognition control, such as Figure 2f The control in it is called "Recognize Text in Images".

[0101] For example, refer to Figure 2g , assuming that the original image includes a QR code, the quick operation control displayed in the image magnification page can be a QR code recognition control, such as Figure 2g The control in it is called "Recognize QR Code".

[0102] For example, refer to Figure 2h , assuming that the original image includes a barcode, the quick operation control displayed in the image magnification page can be a barcode recognition control, such as Figure 2h The control in the tool is called "Recognize Barcode".

[0103] 103. When an editing operation on the text in the target image is detected, display the editing result of the text.

[0104] In this embodiment, the editing operation on the text in the target image may be any type of text editing operation on text in the prior art, such as modification, copying, forwarding, cutting, and the like.

[0105] Optionally, the step of “when an editing operation on the text in the target image is detected, displaying the editing result of the text” may include:

[0106] When a modification trigger operation for a target text in the text is detected, displaying a text input control;

[0107] Determining a modified text corresponding to the target text based on a text input operation on the text input control;

[0108] When a text input completion operation for the text input control is detected, a modified target image is displayed, and the target text in the modified target image is replaced by the modified text.

[0109] In this embodiment, the target text may be all editable text in the target image, or may be obtained based on a text selection operation.

[0110] Optionally, the step of "displaying a text input control when a modification trigger operation for the target text in the text is detected" includes: determining the selected target text in the target image based on a selection operation for the text in the target image, and displaying a text input control.

[0111] In one embodiment, a text box is displayed around each line of text. The selection operation on the target image may be a selection operation on the text box, and the text in the selected text box is the target text.

[0112] In one embodiment, a text input control includes an input box and an input sub-control, wherein the selected target text is displayed in the input box. The target text in the input box can be modified based on the text input operation on the input sub-control. When the text input end operation on the text input control is detected, the text in the input box is used as the modified text of the target text, and the target text in the target image is replaced with the modified text, and the replaced target image is displayed.

[0113] The input sub-control may be an empty control such as a keyboard.

[0114] refer to Figure 3a , when detecting Figure 3aWhen a trigger operation is performed on the text box in page 301, an editing function control list is displayed, wherein the editing function control list includes controls such as copy, forward, and edit. When a trigger operation is detected for the "edit" control, an image editing page shown in 302 is displayed, wherein the text in the text box corresponding to the trigger operation is the target text. The image editing page 302 includes a text input control, which includes an input box 3021 and an input sub-control 3022. The target text "He adds coal for the eyes and buttons. In" is displayed in the input box. Based on the text input operation on the input sub-control, a modified text corresponding to the target text is determined. When a text input end operation is detected for the text input control, the text "He adds tone for the eyes and buttons. In" in the input box is used as the modified text, and a modified target image (as shown in 304) is displayed. The original "He adds coal for the eyes and buttons. In" in the modified target image is replaced by "He adds tone for the eyes and buttons. In".

[0115] For example, reference Figure 3b , more than one text box can be selected. When the Figure 3a When a trigger operation is performed on a text box in the 301 page, a list of edit function controls is displayed, wherein the text box corresponding to the trigger operation can be identified, for example, a gray text box is used to represent the text box corresponding to the trigger operation, that is, the text box selected by the user, and the edit control list includes controls such as copy, forward, and edit. When a trigger operation is detected for the "forward" control, the forwarding destination selection page shown in 305 is displayed, and based on the selection operation on the forwarding destination selection page, the forwarding user selection page is displayed. When a user selection operation is detected on the forwarding user selection page, the text in the text box selected by the user is forwarded to the user corresponding to the user selection operation. For example, the contents of the two gray text boxes are forwarded to friend B corresponding to the user selection operation (reference Figure 3b 307 in the page).

[0116] In this embodiment, the content in the target image can be shared. The sharing can be in the form of pure text or picture. The pure text sharing includes full text sharing and partial text sharing.

[0117] Optionally, in this embodiment, the method of this embodiment further includes:

[0118] When a sharing trigger operation for the target image is detected, displaying a text sharing control and an image sharing control;

[0119] When a trigger operation for the text sharing control is detected, sharing the text in the target image;

[0120] When a trigger operation on the image sharing control is detected, the target image is shared.

[0121] Among them, the sharing of text can be the sharing of part of the text or the sharing of the entire text. Optionally, the recognition result page also includes a sharing trigger control. The step of "when a sharing trigger operation for the target image is detected, displaying the text sharing control and the image sharing control" may include: when a trigger operation for the sharing trigger control is detected, displaying the text sharing control and the image sharing control.

[0122] For example, reference Figure 3c The recognition result page shown in 301 includes a sharing trigger control such as a control named "Forward". When a trigger operation for the sharing trigger control is detected, such as a click operation, the sharing selection page shown in 308 is displayed. The sharing selection page includes text sharing controls such as a "text" control and image sharing controls such as a "picture" control. When a trigger operation for the "text" control is detected, the text content in the editable text in the target image is shared, for example, Figure 3c In step 309 , the text content in the target image is shared with friend D. It is understandable that sharing text is not limited to sharing the text content with friends, but can also be shared with user groups or friends circles, etc.

[0123] For example, reference Figure 3d The recognition result page shown in 301 includes a sharing trigger control such as a control named "forward". When a trigger operation for the sharing trigger control is detected, such as a click operation, the sharing selection page shown in 310 is displayed. The sharing selection page includes text sharing controls such as a "text" control and image sharing controls such as a "picture" control. When a trigger operation for the "picture" control is detected, the target image itself is shared, for example, Figure 3d When a trigger operation for the "Picture" control is detected, the target image is shared with the selected sharing object, such as friend D, based on the sharing object selection operation for the target image. However, it can be understood that the sharing object is not limited to users, but can also be the message integration page of the instant messaging client, such as the Moments page, etc.

[0124] Optionally, in this embodiment, the method of this embodiment further includes: when a text translation operation for the target image is detected, displaying a translation result page corresponding to the target image, the translation result page including a translation image corresponding to the target image, wherein the translation image includes: the translation result corresponding to the text in the target image, and the background content corresponding to the text in the target image.

[0125] Among them, the text translation operation can be some special touch operations, such as long press, double click, triple click and other touch operations, and the text translation operation can also be achieved by triggering the control. Optionally, in one embodiment, the recognition result page includes a translation control, such as Figure 3e The recognition result page 301 displays a translation control such as a control named "Translate".

[0126] The step of “when a text translation operation for the target image is detected, displaying a translation result page corresponding to the target image” may include:

[0127] When a triggering operation on a translation control in a target image is detected, a translation result page corresponding to the target image is displayed.

[0128] For example, reference Figure 3e When a trigger operation such as a click operation is detected for the "Translate" control in the recognition result page shown in 301, a translation result page shown in 311 is displayed. In the translation result page, the editable text in the target image is replaced by the corresponding translation result.

[0129] In this embodiment, a target image sharing solution is also provided. Optionally, the method of this embodiment further includes:

[0130] When an image sharing operation for the target image is detected, displaying a sharing setting page for the target image;

[0131] Determining a target sharing style for a target image based on a sharing style selection operation on the sharing setting page;

[0132] determining an image to be shared based on the target sharing style and the target image;

[0133] The image to be shared is shared.

[0134] The recognition result page may include a sharing trigger control, and the step of “when an image sharing operation for the target image is detected, displaying a sharing setting page for the target image” may include:

[0135] When a trigger operation for the sharing trigger control is detected, displaying a second text sharing control and a second image sharing control;

[0136] When a trigger operation for the second image sharing control is detected, the sharing settings page is displayed. In this embodiment, the sharing settings page can be used not only to select the sharing style of the target image, but also to select the sharing object of the target image. The process of selecting the sharing object can refer to the previous description and will not be repeated here.

[0137] In one embodiment, an operation on the translation result page may trigger the display of the sharing settings page. Optionally, the translation result page may also include a translation control and other functional controls, such as a forwarding control, etc.

[0138] Optionally, “when an image sharing operation for the target image is detected, displaying a sharing setting page for the target image” may include:

[0139] When a trigger operation on the sharing trigger control on the translation result page is detected, displaying a second text sharing control and a second image sharing control;

[0140] When a trigger operation for the second image sharing control is detected, the sharing settings page is displayed. In this embodiment, the sharing settings page can be used not only to select the sharing style of the target image, but also to select the sharing object of the target image. The process of selecting the sharing object can refer to the previous description and will not be repeated here.

[0141] For example, or refer to Figure 3e , when a trigger operation such as a click operation is detected for the "Translate" control in the recognition result page shown in 301, the translation result page shown in 311 is displayed. In the translation result page, the editable text in the target image is replaced with the corresponding translation result. The translation result page displays a sharing trigger control such as a control named "Forward". When a trigger operation is detected for the "Forward" control, a second text sharing control such as a control named "Text" and a second image sharing control such as a control named "Picture" are displayed, wherein the second text sharing control and the second image sharing control can be displayed on the translation result page or on the recognition result page (refer to Figure 3e ), this embodiment has no limitation on this. When a trigger operation for the "Picture" control is detected, the sharing setting page shown in 313 is displayed, and based on the sharing style selection operation on the sharing setting page, the target sharing style of the target image is determined, and the image to be shared is determined based on the target sharing style and the target image; the image to be shared is shared

[0142] The sharing modes in this embodiment include three types: sharing the recognition result of the original image, sharing the translation result, and sharing the translation comparison result.

[0143] Optionally, if the target sharing style is a sharing recognition result, determining the image to be shared based on the target sharing style and the target image includes: determining the target image as the image to be shared.

[0144] For example, in Figure 3e In the sharing setting page 313 shown, if the selected target sharing style is “recognition result”, the image to be shared is the target image.

[0145] Optionally, the sharing settings page includes preview images of the images to be shared under each sharing style. Determining the target sharing style of the target image based on the sharing style selection operation on the sharing settings page may include: determining the target sharing style of the image to be shared based on the selection operation on the preview images in the sharing settings page.

[0146] Optionally, if the target sharing style is to share a translation result, determining the image to be shared based on the target sharing style and the target image includes: determining a translation image corresponding to the target image as the image to be shared.

[0147] In this embodiment, if no text translation operation is detected for the target image before determining the image to be shared based on the target sharing style and the target image, the target image may be first translated to obtain a translated image of the target image. Optionally, determining the image to be shared based on the target sharing style and the target image includes:

[0148] A translation image of the target image is obtained, and the translation image corresponding to the target image is determined as the image to be shared.

[0149] For example, in Figure 3f In the sharing settings page shown, if the selected target sharing style is "Translation Result", the image to be shared is the translated image of the target image.

[0150] Optionally, if the target sharing style is to share the translation comparison result, determining the image to be shared based on the target sharing style and the target image includes:

[0151] A translation comparison image of the target image is obtained, where the translation comparison image includes the content of the target image and the content of the translation image of the target image.

[0152] In this example, the translation comparison image can be obtained by splicing the target image and the translation image of the target image. The splicing can be completed by the terminal, or the terminal can send a splicing instruction to the server, and the server completes the splicing of the target image and the translation image.

[0153] For example, in Figure 3gIn the sharing settings page shown, if the selected target sharing style is "Translation Comparison", the image to be shared will be a translation comparison image.

[0154] In this embodiment, editable text may be extracted from the target image for display and editing. Optionally, the image processing method of this embodiment may further include:

[0155] When a text extraction operation is detected for the target image in the recognition result page, a text extraction result page of the target image is displayed, wherein the text extraction result page includes editable text in the target image.

[0156] That is, the text in the text extraction result page comes from the text recognized from the original image.

[0157] The text extraction operation may be a specific touch operation, such as double-clicking, long pressing, or the like. In addition, the text extraction operation may also be implemented by triggering a control.

[0158] For example, reference Figure 4a The recognition result page 401 includes a text extraction control such as a control named "Extract Part". When a trigger operation for the control is detected, a text extraction result page 402 of the target image is displayed.

[0159] The text in the recognition result page 401 is editable. If the user selects a portion of the text box in the recognition result page 401, the text in the selected text box becomes the extracted text corresponding to the text extraction control. Optionally, the step of "when a text extraction operation is detected for the target image in the recognition result page, displaying the text extraction result page for the target image" may include:

[0160] determining selected text in the target image based on a selection operation on the text in the target image;

[0161] When a trigger operation on a text extraction control in a recognition result page is detected, a text extraction result page of a target image is displayed, wherein the text extraction result page includes the selected text.

[0162] In this way, partial extraction of text in the target image can be achieved.

[0163] Optionally, in this embodiment, after the step of “displaying the text extraction result page of the target image”, the following steps may also be included:

[0164] When a comparison display operation for the text extraction result page is detected, a comparison page is displayed, wherein the comparison page includes a first display area and a second display area, wherein the first display area is used to display the target image, and the second display area is used to display the text extraction result of the target image.

[0165] Optionally, the control display operation may be a specific touch operation or may be implemented by operating a control.

[0166] Optionally, the text extraction result page further includes a comparison display control. When a comparison display operation for the text extraction result page is detected, the comparison page is displayed. This may include: when a trigger operation for the comparison display control is detected, the comparison page is displayed.

[0167] For example, reference Figure 4b The text extraction result page shown in 402 includes a comparison display control such as a control named "comparison control". When a trigger operation for the "comparison control" is detected, the comparison page 403 is displayed. The comparison page 403 includes two display areas, a first display area 4031 and a second display area 4032. The first display area is used to display the target image, and the second display area is used to display the text extraction result of the target image.

[0168] In one embodiment, the first display area may display not the target image but the original image. The text extraction results of the original image and the target image are displayed in comparison, which can provide an original image comparison function for the text extraction results, making it easier to check whether the text recognized from the original image is incorrect.

[0169] Optionally, when the first display area displays a target image, the method of this embodiment further includes:

[0170] When a text selection operation is detected for the target image in the first display area, determining selected text corresponding to the text selection operation in the target image;

[0171] The text extraction result displayed in the second display area is adjusted based on the selected text, wherein after the adjustment, the text extraction result displayed in the second display area includes the text extraction result corresponding to the selected text.

[0172] For example, reference Figure 4bIn the comparison page shown in 403, when a text selection operation is detected for the line of text "He puts a big snowball on top. He adds a" in the first display area, the line of text is treated as selected text, and the text extraction result displayed in the second display area is adjusted based on the selected text. The adjusted second display area can be referred to in 404. Compared with 403, in the second display area in 404, the display position of "He puts a big snowball on top. He adds a" is located at the upper part of the second display area, which is more obvious.

[0173] In this embodiment, the text extraction result in the second display area is editable. When an input trigger operation is detected in the second display area, a second text input control is displayed in the second display area. The input trigger operation can be a click operation. At the location where the user clicks, a text input control can be displayed. Figure 4b The cursor indicated by A in the middle is convenient for prompting the user the text input position. When displaying the text input control, this embodiment can increase the area of ​​the second display area, such as raising the position of the upper boundary line of the second display area.

[0174] Optionally, when the first display area displays an original image, the method of this embodiment further includes:

[0175] When a text selection operation on the original image in the first display area is detected, determining selected text corresponding to the text selection operation in the original image;

[0176] The text extraction result displayed in the second display area is adjusted based on the selected text, wherein after the adjustment, the text extraction result displayed in the second display area includes the text extraction result corresponding to the selected text.

[0177] In this embodiment, the text in the original image may have position information. For example, the text in the original image may be marked by a text box, and the position information of the text box may be used as the position information of the text marked by the text box. The position information of the text box may be determined based on the position information of the corresponding text box in the target image.

[0178] Optionally, in this embodiment, the step of “displaying a recognition result page of the original image based on the image text recognition operation on the original image” may include:

[0179] triggering acquisition of a text recognition result of the original image based on an image text recognition operation on the original image, wherein the text recognition result includes text recognized from the original image and a text position of the text in the original image;

[0180] Replacing the original text at the corresponding text position in the original image with the recognized text in the form of editable text to obtain a target image corresponding to the original image;

[0181] A recognition result page of the original image is displayed, wherein the recognition result page includes the target image.

[0182] In which, the text recognition result and the target image can be generated independently by the terminal, or the text recognition result can be obtained by the server based on the recognition of the original image, and the target image can be generated by the terminal based on the original image and the text recognition result, or the text recognition result and the target image can be both generated by the server. This embodiment has no restrictions on this.

[0183] Optionally, in this embodiment, the original image may be recognized using OCR technology to obtain a text recognition result.

[0184] Optionally, the step of “replacing the text at the corresponding text position in the original image with the recognized text in the form of editable text to obtain a target image corresponding to the original image” includes:

[0185] Analyzing the recognized text based on the text position of the text in the text recognition result to obtain at least one text block;

[0186] Sort text blocks and format text within text blocks;

[0187] The corresponding text content in the original image is replaced with the typeset text block to obtain a target image.

[0188] Among them, the text in the original image can be removed based on the position information of the text in the text recognition result, and the original image can be modified to fill the removed text with the background content near the removed text to obtain a background image, and the typeset text block can be drawn into the background image in the form of editable text to obtain the target image.

[0189] By using the image processing method of this embodiment, a chat session page of an instant messaging client can be displayed, wherein the chat session page includes an original image sent by a chat session user; based on an image text recognition operation on the original image, a recognition result page of the original image is displayed, and the recognition result page includes a target image, and the target image includes: text recognized from the original image, and background content corresponding to the text, the text is editable text, and the background content is content other than the text in the original image; when an editing operation on the text in the target image is detected, the editing result of the text is displayed, thereby, the original image sent by the chat session user in the chat session can be identified as a target image with editable text, and the user can directly edit the text in the target image, obtaining an editing experience similar to editing text in the original image.

[0190] The method described in the above embodiment will be further described in detail below with examples.

[0191] In this embodiment, description will be made by taking an example where the first image processing apparatus is specifically integrated into a terminal and the second image processing apparatus is specifically integrated into a server.

[0192] like Figure 5a As shown, an image processing method, the specific process is as follows:

[0193] 501. The terminal displays a chat session page of an instant messaging client, wherein the chat session page includes an original image sent by a user of the chat session.

[0194] 502. The terminal sends an image recognition request to the server based on the image text recognition operation on the original image, wherein the image recognition request may carry the original image;

[0195] Among them, reference Figure 5b As shown in the optional timing diagram of the image processing method, the user can long press the original image in the chat session page of the instant messaging client to send the original image to the server to trigger the recognition of the original image.

[0196] In this embodiment, the server can be composed of many components, such as Figure 5b , including but not limited to: cloud recognition backend component, OCR recognition service component, cloud recognition typesetting component, drawing component and image generation component. These components can be integrated into one server or into different servers, and this embodiment has no limitation on this.

[0197] Optionally, the terminal may send an image recognition request to the cloud recognition backend component of the server through a big data channel.

[0198] The Cloud Recognition backend integrates numerous classification and recognition services, and supports configuring different recognition types. This includes image classification, and recognition types include number recognition, text recognition, and barcode image recognition, among others.

[0199] For the image text recognition operation of the original image, it can be considered that the terminal side actively selects the text recognition type, or, when sending an image recognition request, the terminal can directly write the text recognition type into the image recognition request. After receiving the image recognition request, the cloud recognition background will start the cloud recognition OCR service to extract the text in the original image.

[0200] 503. The server receives the image recognition request sent by the terminal, and obtains the original image based on the image recognition request;

[0201] 504. The server performs text recognition on the original image to obtain an original text recognition result, wherein the text recognition result includes the text recognized from the original image and the text position of the text in the original image;

[0202] After the cloud recognition backend of the server receives the image recognition request, it finds that the recognition type for the original image is text recognition type, calls the OCR service component, performs OCR recognition on the original image, and receives the recognition result of the OCR service component. In this embodiment, the OCR service component can segment the text in the image and identify each segmented text separately. The obtained text recognition result can include the recognized text and the text position and confidence of each text, where the text position of each text can be represented by coordinates.

[0203] 505. The server replaces the text at the corresponding text position in the original image with the recognized text in the form of editable text, thereby obtaining a target image corresponding to the original image.

[0204] The server may determine the original content corresponding to the text position in the original image based on the text position of the recognized text, and replace the original content at the text position with editable text using the recognized text to obtain the target image.

[0205] However, this direct replacement method may result in irregular text layout after replacement, which is not conducive to reading. In this embodiment, the OCR recognized text can be typeset first, and then the above-mentioned replacement can be performed to obtain the target image.

[0206] Optionally, the step of “the server replacing the text at the corresponding text position in the original image with the recognized text in the form of editable text to obtain a target image corresponding to the original image” may include:

[0207] The server analyzes the recognized text based on the position information of the text in the text recognition result to obtain at least one text block;

[0208] The server sorts the text blocks and typesets the text within the text blocks;

[0209] The server replaces the corresponding text content in the original image with the typeset text block to obtain a target image.

[0210] The above-mentioned text recognition result may be an OCR recognition result. After the cloud recognition backend of the server receives the OCR recognition result, it may call the cloud recognition typesetting component to typeset the OCR recognition result.

[0211] For example, the cloud recognition backend of the server calls the cloud recognition typesetting component to analyze the text in the OCR recognition result based on the text position of the text in the OCR recognition result to obtain at least one text block, sort the text blocks, and typeset the text in the text block.

[0212] Among them, the cloud recognition and typesetting component can first determine whether the original image contains a preset document through a classification algorithm. If the original image does not contain the preset document, the original image is simply typeset. For example, if only a small amount of text such as a line of text is recognized in the original image, it is considered that the original image does not contain the preset document, and the original image can be simply typeset.

[0213] If the original image contains a preset document, for example, a large number of texts are recognized in the original image, then the original image is considered to contain a preset document, and the cloud recognition typesetting component can use a layout analysis algorithm for typesetting.

[0214] In this embodiment, the layout analysis algorithm used by the cloud recognition typesetting component can be an optimized Docstrum algorithm, which uses the text position of the text extracted by OCR (such as the coordinates of the four corners of the text box) as input, solving the problems of the traditional Docstrum algorithm such as time consumption and difficult threshold control, and finally merges the text boxes extracted by OCR into text blocks.

[0215] The cloud recognition layout component can first determine the text line based on the text position of the text extracted by OCR, and then divide the text line into at least one text block based on the centroid of the text line. After dividing the text blocks, the cloud recognition layout component can sort the text blocks. For example, it can recursively cut the text blocks vertically and horizontally to construct a binary tree. Based on the binary tree, the order of the text blocks is determined to match the user's reading order and reading logic.

[0216] Afterwards, the cloud recognition typesetting component can typeset the text in the text block so that the text in the text block conforms to the user's reading logic.

[0217] Among them, the cloud recognition typesetting component can also segment the original image and obtain the location information of other content in the original image besides the text. For example, the image position of the illustration in the original image can be obtained. The server can obtain the location information of the sorted text block based on the text position of the text in the text block, as well as the sorting and in-block typesetting of the text block. Based on the location information and the location information of the background image in the original image, such as the illustration, the text block is drawn in the original image to obtain the target image.

[0218] In one embodiment, the server obtains the position information of the sorted text blocks and the information of the original image, such as the position information of the background content of the original image, and then sends the position information of the text blocks and the information of the original image to the terminal.

[0219] After receiving the information, the terminal can send the text block, the position information of the text block, and the information of the original image to the drawing component of the server. The drawing component draws the text block in the original image based on the information of the text block, the position information of the text block, and the information of the original image to obtain the target image, wherein the text in the text block drawn in the target image is editable text. Optionally, while drawing the text, the drawing component can also draw a text box for the text, wherein each line of text can correspond to a text box, and the text box can be used to respond to the user's touch operation. For example, the text in the text box clicked by the user is regarded as the selected text, and the selected text can be subjected to editing operations such as copying and forwarding.

[0220] In one embodiment, the process of drawing the target image may be executed by the terminal.

[0221] 506. The server sends the target image to the terminal.

[0222] 507. The terminal receives the target image and displays a recognition result page, wherein the recognition result page includes the target image.

[0223] 508. When the terminal detects a text translation operation for a target image, it displays a translation result page corresponding to the target image, where the translation result page includes a translation image corresponding to the target image, wherein the translation image includes: a translation result corresponding to the text in the target image, and background content corresponding to the text in the target image.

[0224] The translation of the editable text of the target image and the generation of the translated image may be performed by the terminal or the server.

[0225] The translation result page may include a sharing control.

[0226] The method of this embodiment may further include:

[0227] When the terminal detects a sharing operation on the sharing control on the translation result page, displaying a sharing setting page for the target image;

[0228] The terminal determines a target sharing style of the target image based on a sharing style selection operation on the sharing setting page;

[0229] The terminal determines an image to be shared based on the target sharing style and the target image;

[0230] The terminal shares the image to be shared.

[0231] For example, reference Figure 5b When the user clicks to share an image, the terminal may send an image sharing request to the image generation component, triggering the image generation component to generate the image to be shared.

[0232] When the target sharing style is to share the translation result, the image to be shared is determined to be the translation image corresponding to the target image. The terminal may request the image generation component to translate the image.

[0233] When the target sharing style is to share the translation comparison result, the shared image is the translation comparison image, which includes the content in the target image and the content in the translation image of the target image.

[0234] Optionally, the terminal may send an image sharing request for the translation comparison image to the image generation component of the server, triggering the image generation component to synthesize the target image and the translation image, for example, by left-right splicing to obtain the translation comparison image.

[0235] In this embodiment, quick operations are also provided for images. For the flowchart of implementing quick operations, refer to Figure 5c As shown, when the terminal detects that the user is viewing an image, for example, when a display operation such as a click operation is detected on the original image, the terminal sends the original image to the server, triggering the server to start the cloud recognition background, and calling the coarse classification service under the cloud recognition service through the big data channel. The cloud recognition background configures different recognition types for images in different scenarios. For example, for ID card photos, a number recognition type is configured, and for images with QR codes, a code image recognition type is configured, and so on.

[0236] After the coarse classification service identifies the identification type corresponding to the image, the identification type is returned to the client, and the client displays the image shortcut control corresponding to the identification type.

[0237] Therefore, this embodiment can provide users with a target image with editable text, giving users an experience similar to directly modifying the text in the original image. Moreover, the text processing for the target image can be local text processing. Users can freely select the text to be processed and perform operations such as translation, highlighting, copying, etc., which is conducive to improving the user's image processing experience.

[0238] In order to better implement the above method, an image processing device is also provided accordingly, wherein the image processing device can be integrated into a terminal, or integrated into a server, or integrated into a terminal and a server.

[0239] For example, Figure 6 , as shown, the image processing device may include

[0240] The conversation page display unit 601 is used to display the chat conversation page of the instant messaging client, wherein the chat conversation page includes the original image sent by the chat conversation user;

[0241] A recognition result display unit 602 is configured to display a recognition result page of the original image based on the image text recognition operation on the original image, wherein the recognition result page includes a target image, and the target image includes: text recognized from the original image and background content corresponding to the text, wherein the text is editable text and the background content is content in the original image other than the text;

[0242] The editing result display unit 603 is configured to display the editing result of the text when an editing operation on the text in the target image is detected.

[0243] Optionally, the recognition result display unit is used to display an image text recognition control based on a control display operation for the original image; when a trigger operation for the image text recognition control is detected, display a recognition result page of the original image.

[0244] Optionally, the recognition result display unit is used to display an image magnification page of the original image when a display operation for the original image is detected, and the image magnification page includes the original image in a full-screen display state; when a control display trigger operation for the original image is detected on the image magnification page, the image text recognition control is displayed.

[0245] Optionally, the image magnification page further includes: a quick operation control corresponding to the target content in the original image, wherein the quick operation control is used to perform an operation indicated by the quick operation control on the target content when triggered.

[0246] Optionally, an editing result display unit is used to display a text input control when a modification trigger operation is detected for a target text in the text; determine the modified text corresponding to the target text based on the text input operation for the text input control; and display a modified target image when a text input end operation is detected for the text input control, in which the target text is replaced by the modified text.

[0247] Optionally, the device further comprises:

[0248] a first sharing triggering unit, configured to display a text sharing control and an image sharing control when a sharing triggering operation for the target image is detected;

[0249] A first sharing unit, configured to share the text in the target image when a triggering operation on the text sharing control is detected;

[0250] The second sharing unit is configured to share the target image when a triggering operation on the image sharing control is detected.

[0251] Optionally, the device further includes: a translation result display unit, configured to display a translation result page corresponding to the target image when a text translation operation for the target image is detected, wherein the translation result page includes a translation image corresponding to the target image, wherein the translation image includes: a translation result corresponding to the text in the target image, and background content corresponding to the text in the target image.

[0252] Optionally, the device further comprises:

[0253] a second sharing triggering unit, configured to display a sharing setting page of the target image when an image sharing operation for the target image is detected;

[0254] a sharing setting unit, configured to determine a target sharing style of a target image based on a sharing style selection operation on the sharing setting page;

[0255] a determining unit, configured to determine an image to be shared based on the target sharing style and the target image;

[0256] The third sharing unit is configured to share the image to be shared.

[0257] Optionally, if the target sharing style is a sharing recognition result, the determining unit is configured to determine the target image as an image to be shared;

[0258] Optionally, if the target sharing style is to share a translation result, the determining unit is configured to determine the translated image corresponding to the target image as the image to be shared;

[0259] Optionally, if the target sharing style is to share a translation comparison result, the determination unit is configured to obtain a translation comparison image of the target image, where the translation comparison image includes content in the target image and content in the translation image of the target image.

[0260] Optionally, the device further includes: an extraction unit, configured to display a text extraction result page of the target image when a text extraction operation is detected for the target image in the recognition result page, wherein the text extraction result page includes editable text in the target image.

[0261] Optionally, the device also includes a comparison display unit for displaying a comparison page after the extraction unit displays the text extraction result page of the target image when a comparison display operation for the text extraction result page is detected, the comparison page including a first display area and a second display area, the first display area being used to display the target image, and the second display area being used to display the text extraction result of the target image.

[0262] Optionally, the device further includes: a text selection unit, configured to, when a text selection operation on the target image in the first display area is detected, determine selected text corresponding to the text selection operation in the target image;

[0263] A positioning unit is used to adjust the text extraction results displayed in the second display area based on the selected text, wherein after the adjustment, the text extraction results displayed in the second display area include the text extraction results corresponding to the selected text.

[0264] Optionally, the recognition result display unit includes:

[0265] a triggering subunit, configured to trigger acquisition of a text recognition result of the original image based on an image text recognition operation on the original image, wherein the text recognition result includes text recognized from the original image and a text position of the text in the original image;

[0266] a replacement subunit, configured to replace the text at the corresponding text position in the original image with the recognized text in the form of editable text, thereby obtaining a target image corresponding to the original image;

[0267] The display subunit is configured to display a recognition result page of the original image, wherein the recognition result page includes the target image.

[0268] In addition, an embodiment of the present invention further provides a computer device, which may be a terminal or a server. Figure 7, which shows a schematic diagram of the structure of a computer device involved in an embodiment of the present invention, specifically:

[0269] The computer device may include one or more processing core processors 701, one or more computer readable storage media memories 702, a power supply 703, an input unit 704 and other components. Those skilled in the art will understand that Figure 7 The computer device structure shown in the figure does not constitute a limitation on the computer device, and may include more or fewer components than shown in the figure, or combine certain components, or arrange components differently.

[0270] Processor 701 is the control center of the computer device. It connects all components of the computer device using various interfaces and circuits. It executes the various functions of the computer device and processes data by running or executing software programs and / or modules stored in memory 702 and accessing data stored in memory 702. Optionally, processor 701 may include one or more processing cores. Preferably, processor 701 may integrate an application processor and a modem processor. The application processor primarily processes the operating system, user interface, and application programs, while the modem processor primarily handles wireless communications. It is understood that the modem processor may not be integrated into processor 701.

[0271] The memory 702 can be used to store software programs and modules. The processor 701 executes various functional applications and data processing by running the software programs and modules stored in the memory 702. The memory 702 may mainly include a program storage area and a data storage area, wherein the program storage area may store an operating system, an application required for at least one function (such as a sound playback function, an image playback function, etc.), etc.; the data storage area may store data created according to the use of the computer device, etc. In addition, the memory 702 may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one disk storage device, a flash memory device, or other volatile solid-state storage device. Accordingly, the memory 702 may also include a memory controller to provide the processor 701 with access to the memory 702.

[0272] The computer device also includes a power supply 703 for supplying power to various components. Preferably, the power supply 703 can be logically connected to the processor 701 via a power management system, thereby enabling the power management system to manage charging, discharging, and power consumption. The power supply 703 can also include one or more DC or AC power supplies, a recharging system, a power failure detection circuit, a power converter or inverter, a power status indicator, and other arbitrary components.

[0273] The computer device may further include an input unit 704, which may be configured to receive input digital or character information and generate keyboard, mouse, joystick, optical or trackball signal input related to user settings and function control.

[0274] Although not shown, the computer device may further include a display unit, etc., which will not be described in detail here. Specifically, in this embodiment, the processor 701 in the computer device will load the executable files corresponding to one or more application processes into the memory 702 according to the following instructions, and the processor 701 will run the application stored in the memory 702 to implement various functions as follows:

[0275] Displaying a chat session page of an instant messaging client, wherein the chat session page includes an original image sent by a user of the chat session;

[0276] Based on the image text recognition operation on the original image, displaying a recognition result page of the original image, the recognition result page including a target image, the target image including: text recognized from the original image, and background content corresponding to the text, the text being editable text, and the background content being content in the original image other than the text;

[0277] When an editing operation on the text in the target image is detected, the editing result of the text is displayed.

[0278] The specific implementation of the above operations can be found in the previous embodiments and will not be repeated here.

[0279] Those skilled in the art will appreciate that all or part of the steps in the various methods of the above embodiments may be accomplished by instructions, or by controlling related hardware through instructions. The instructions may be stored in a computer-readable storage medium and loaded and executed by a processor.

[0280] To this end, an embodiment of the present invention further provides a storage medium, in which a plurality of instructions are stored. The instructions can be loaded by a processor to execute the image processing method provided by the embodiment of the present invention.

[0281] The system involved in the embodiment of the present invention may be a distributed system formed by connecting a client and multiple nodes (any form of computer equipment in an access network, such as a server and a terminal) through network communication.

[0282] Taking the distributed system as the blockchain system as an example, see Figure 8 , Figure 8This is a schematic diagram of an optional architecture for a distributed system 800 provided in an embodiment of the present invention, applied to a blockchain system. The system consists of multiple nodes 801 (any type of computing device connected to a network, such as a server or user terminal) and clients 802. The nodes form a peer-to-peer (P2P) network. The P2P protocol is an application layer protocol that runs on top of the Transmission Control Protocol (TCP). In a distributed system, any machine, such as a server or terminal, can join and become a node. Nodes include hardware, middleware, operating system, and application layers. The original image, target image, and translated versions of the target image can be stored in the blockchain system's shared ledger.

[0283] See also Figure 8 The functions of each node in the blockchain system shown include:

[0284] 1) Routing: A basic function of a node, used to support communication between nodes.

[0285] In addition to the routing function, nodes can also have the following functions:

[0286] 2) Applications, deployed in the blockchain, implement specific services based on actual business needs, record data related to the implementation of functions to form record data, carry digital signatures in the record data to indicate the source of the task data, and send the record data to other nodes in the blockchain system for other nodes to add the record data to a temporary block when they successfully verify the source and integrity of the record data.

[0287] For example, the services implemented by the application include:

[0288] 2.1) Wallet: This provides the functionality for conducting electronic currency transactions, including initiating transactions (i.e., sending the current transaction record to other nodes in the blockchain system. Upon successful verification by other nodes, the transaction record data is stored in a temporary block of the blockchain as a response to acknowledge the transaction's validity). The wallet also supports querying the remaining electronic currency in an electronic currency address.

[0289] 2.2) Shared ledgers are used to store, query, and modify account data. Records of operations on account data are sent to other nodes in the blockchain system. After verification, other nodes acknowledge the validity of the account data by storing the recorded data in a temporary block. They can also send a confirmation to the node that initiated the operation.

[0290] 2.3) Smart contracts are computerized protocols that can enforce the terms of a contract. They are implemented through code deployed on a shared ledger that is executed when certain conditions are met. Based on actual business needs, the code is used to complete automated transactions, such as querying the logistics status of a buyer's purchased goods and transferring the buyer's electronic currency to the merchant's address after the buyer signs for the goods. Of course, smart contracts are not limited to executing contracts for transactions, but can also execute contracts that process received information.

[0291] 3) Blockchain, including a series of blocks that are connected to each other in the order of their generation. Once a new block is added to the blockchain, it will not be removed. The block records the record data submitted by the nodes in the blockchain system.

[0292] In this embodiment, the content browsed by the current user and / or associated users, and / or the record data of the content browsed by the current user and / or associated users (such as content description information and link information, etc.) can be stored in the shared ledger of the regional chain through the node, and the computer device (such as a terminal or server) can obtain the content browsed by the current user and / or associated users based on the data stored in the shared ledger.

[0293] See also Figure 9 , Figure 9 This is an optional schematic diagram of the block structure provided by an embodiment of the present invention. Each block includes the hash value of the transaction records stored in the block (the hash value of the current block) and the hash value of the previous block. The blocks are connected by hash values ​​to form a blockchain. In addition, the block may also include information such as the timestamp when the block was generated. Blockchain is essentially a decentralized database, a series of data blocks generated using cryptographic methods. Each data block contains relevant information used to verify the validity of the information (anti-counterfeiting) and generate the next block.

[0294] The specific implementation of the above operations can be found in the previous embodiments and will not be repeated here.

[0295] The storage medium may include a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, etc.

[0296] Since the instructions stored in the storage medium can execute the steps in the image processing method provided in the embodiment of the present invention, the beneficial effects that can be achieved by the image processing method provided in the embodiment of the present invention can be achieved. Please refer to the previous embodiment for details and will not be repeated here.

[0297] The above is a detailed introduction to an image processing method, device, computer equipment and storage medium provided by an embodiment of the present invention. Specific examples are used herein to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only used to help understand the method of the present invention and its core ideas. At the same time, for those skilled in the art, according to the ideas of the present invention, there may be changes in the specific implementation methods and application scopes. In summary, the contents of this specification should not be understood as limiting the present invention.

Claims

1. An image processing method, characterized in that: include: Displaying a chat session page of an instant messaging client, wherein the chat session page includes an original image sent by a user of the chat session; Based on an image text recognition operation on the original image, a recognition result page of the original image is displayed, the recognition result page including: text recognized from the original image and background content corresponding to the text, the text being editable text, and the background content being content other than the text in the original image, the position of the editable text in the recognition result page corresponding to the position of the text in the original image, the position of the background content in the recognition result page corresponding to the position of content other than the text in the original image, the editable text being selectable by a user and operations being performed on the selected text, and the recognition result page being obtained by replacing the original text at the corresponding text position in the original image with the editable text; When an editing operation on the text in the recognition result page is detected, the editing result of the text is displayed.

2. The image processing method according to claim 1, wherein: The step of displaying a recognition result page of the original image based on the image text recognition operation on the original image includes: Displaying an image text recognition control based on a control display operation for the original image; When a trigger operation for the image text recognition control is detected, a recognition result page of the original image is displayed.

3. The image processing method according to claim 2, wherein: The displaying of the image text recognition control based on the control display operation for the original image includes: When a display operation for the original image is detected, displaying an image magnification page of the original image, the image magnification page including the original image in a full-screen display state; When a control display triggering operation for the original image is detected on the image magnification page, an image text recognition control is displayed.

4. The image processing method according to claim 3, wherein: The image magnification page further includes: a quick operation control corresponding to the target content in the original image, wherein the quick operation control is used to execute the operation indicated by the quick operation control on the target content when triggered.

5. The image processing method according to claim 1, wherein: When an editing operation on the text in the recognition result page is detected, displaying the editing result of the text includes: When a modification trigger operation for target text in the text on the recognition result page is detected, displaying a text input control; Determining a modified text corresponding to the target text based on a text input operation on the text input control; When a text input completion operation for the text input control is detected, a text-modified recognition result page is displayed, in which the target text is replaced by the modified text.

6. The image processing method according to claim 1, wherein: Also includes: When a sharing trigger operation for the recognition result page is detected, displaying a text sharing control and an image sharing control; When a trigger operation for the text sharing control is detected, sharing the text in the recognition result page; When a trigger operation for the image sharing control is detected, the recognition result page is shared.

7. The image processing method according to claim 1, wherein: Also includes: When a text translation operation for the recognition result page is detected, a translation result page corresponding to the recognition result page is displayed, wherein the translation result page includes a translation image corresponding to the recognition result page, and the translation image includes: a translation result corresponding to the text in the recognition result page, and background content corresponding to the text in the recognition result page.

8. The image processing method according to claim 7, characterized in that: Also includes: When an image sharing operation for the recognition result page is detected, displaying a sharing setting page; Determining a target sharing style for the content in the recognition result page based on a sharing style selection operation on the sharing setting page; Determining an image to be shared based on the target sharing style and content in the recognition result page; The image to be shared is shared.

9. The image processing method according to claim 8, characterized in that: If the target sharing style is to share the recognition result, determining the image to be shared based on the target sharing style and the content on the recognition result page includes: Determining an image to be shared based on the content in the recognition result page; If the target sharing style is to share the translation result, determining the image to be shared based on the target sharing style and the content in the recognition result page includes: Determining the translated image corresponding to the recognition result page as the image to be shared; If the target sharing style is to share the translation comparison result, determining the image to be shared based on the target sharing style and the content in the recognition result page includes: A translation comparison image of the content in the recognition result page is obtained, where the translation comparison image includes the content in the recognition result page and the content in the translation image.

10. The image processing method according to claim 1, wherein: Also includes: When a text extraction operation on the recognition result page is detected, the text extraction result page is displayed, wherein the text extraction result page includes the editable text.

11. The image processing method according to claim 10, wherein: After displaying the text extraction result page, the method further includes: When a comparison display operation for the text extraction result page is detected, a comparison page is displayed, wherein the comparison page includes a first display area and a second display area, wherein the first display area is used to display the content in the recognition result page, and the second display area is used to display the text extraction result of the recognition result page.

12. The image processing method according to claim 10, wherein: After displaying the text extraction result page, the method further includes: In response to an editing operation on editable text in a text extraction result page, the editable text in the text extraction result page is updated based on the editing operation.

13. The image processing method according to claim 11, wherein: Also includes: When a text selection operation on the content in the first display area is detected, determining selected text corresponding to the text selection operation in the first display area; The text extraction results displayed in the second display area are adjusted based on the selected text. After the adjustment, the text extraction results displayed in the second display area include the text extraction results corresponding to the selected text.

14. The image processing method according to claim 1, wherein: Based on the image text recognition operation on the original image, displaying a recognition result page of the original image includes: triggering acquisition of a text recognition result of the original image based on an image text recognition operation on the original image, wherein the text recognition result includes text recognized from the original image and a text position of the text in the original image; Replacing the original text at the corresponding text position in the original image with the recognized text in the form of editable text, thereby obtaining the content of the recognition result page corresponding to the original image; The recognition result page of the original image is displayed.

15. The image processing method according to any one of claims 1 to 14, characterized in that: A text mark is also displayed corresponding to the text area where the editable text is located in the recognition result page.

16. An image processing device, characterized in that: include: a conversation page display unit, configured to display a chat conversation page of an instant messaging client, wherein the chat conversation page includes an original image sent by a chat conversation user; a recognition result display unit, configured to display a recognition result page of the original image based on an image text recognition operation performed on the original image, the recognition result page including: text recognized from the original image and background content corresponding to the text, the text being editable text, and the background content being content other than the text in the original image; a position of the editable text in the recognition result page corresponding to a position of the text in the original image in the original image; a position of the background content in the recognition result page corresponding to a position of content other than the text in the original image in the original image; the editable text being available for selection by a user, and operations performed on the selected text; and the recognition result page being obtained by replacing the original text at the corresponding text position in the original image with the editable text; The editing result display unit is configured to display the editing result of the text when an editing operation on the text in the recognition result page is detected.

17. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 15 are implemented.

18. A storage medium having a computer program stored thereon, wherein: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 15 are implemented.

19. An image processing method, characterized in that: include: Displaying a chat session page of an instant messaging client, wherein the chat session page includes an original image sent by a user of the chat session; Based on an image text recognition operation on the original image, a recognition result page of the original image is displayed, the recognition result page including: text recognized from the original image and background content corresponding to the text, the text being editable text, and the background content being content other than the text in the original image, the editable text being selectable by a user and allowing operations to be performed on the selected text, the recognition result page being obtained by replacing original text at a corresponding text position in the original image with the editable text; In response to a text translation operation on the recognition result page, a translation result corresponding to the editable text is displayed on the recognition result page, wherein the editable text is replaced by the corresponding translation result.

20. An image processing method, characterized in that: include: Displaying a chat session page of an instant messaging client, wherein the chat session page includes an original image sent by a user of the chat session; Based on an image text recognition operation on the original image, a recognition result page of the original image is displayed, the recognition result page including a sharing trigger control, text recognized from the original image, and background content corresponding to the text, the text being editable text, and the background content being content in the original image other than the text, the editable text being selectable by a user and allowing operations to be performed on the selected text, the recognition result page being obtained by replacing the original text at a corresponding text position in the original image with the editable text; When a trigger operation for the sharing trigger control is detected, displaying a text sharing control and an image sharing control; When a trigger operation for the text sharing control is detected, sharing at least part of the text in the recognition result page; When a trigger operation for the image sharing control is detected, the content in the recognition result page is shared.

Citation Information

Patent Citations

  • Information processing method and electronic equipment

    CN105739832A

  • Chat data input method and device and communication terminal

    CN106909270A

  • Text recognition method and device, mobile terminal, and storage medium

    CN109002759A