Methods, Apparatus, Devices, and Storage Media for Document Processing

By adding images and merging and identifying texts by receiving image data units, the problem of inability to integrate multiple image texts in the prior art is solved, automatic merging and editing of texts is realized, and operation convenience is improved.

CN118658173BActive Publication Date: 2025-07-04BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410804849.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-06-20
Publication Date
2025-07-04
Estimated Expiration
2044-06-20

AI Technical Summary

Technical Problem

The prior art cannot effectively integrate the text content in multiple images, which limits the convenience of text recognition and editing operations of multiple images.

Method used

By adding at least one image by receiving the image data unit, identifying texts corresponding to each image are acquired, and the texts are merged according to the order of image addition, and finally the merged text is added to the corresponding data unit of the target document.

Benefits of technology

Automatic merging and editing of text in different images is realized, improving the convenience of multiple image text recognition and editing, and eliminating the limitation on the number of images.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118658173B_ABST
    Figure CN118658173B_ABST
Patent Text Reader

Abstract

Embodiments of the present disclosure provide a method, apparatus, device, and storage medium for document processing. The method includes: receiving an operation of adding at least one first image to an image data unit in a target document; obtaining at least one first recognition text respectively corresponding to the at least one first image, where each first recognition text is obtained by text recognition of the corresponding first image; merging the at least one first recognition text into a target text according to the addition order of the at least one first image; and adding the target text to a full text unit corresponding to the image data unit in the target document. In this way, the texts recognized from different images can be automatically merged together, thus eliminating the limitation on the number of images to be recognized. This provides a convenient solution for text recognition and text editing of multiple images.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Example embodiments of the present disclosure generally relate to the field of computers, and particularly to methods, apparatuses, devices, and computer-readable storage media for document processing. Background Art

[0002] In many usage scenarios, it is necessary to edit or store the text content in an image. To this end, image recognition technologies such as optical character recognition (OCR) can be used to recognize the text in the image. OCR refers to the process of analyzing an image to obtain the text information in the image. For example, by scanning a form or a receipt, an image of the scanned form or receipt can be obtained. Usually, the text in such an image cannot be edited, searched, or counted. In this case, OCR can be used to convert the image into text. However, there are still many inconveniences in the scenario of using OCR to convert an image into text. Summary of the Invention

[0003] In a first aspect of the present disclosure, a method for document processing is provided. The method includes: receiving an operation of adding at least one first image to an image data unit in a target document; obtaining at least one first recognition text respectively corresponding to the at least one first image, each first recognition text being obtained by text recognition of the corresponding first image; merging the at least one first recognition text into a target text according to the addition order of the at least one first image; and adding the target text to a full text unit corresponding to the image data unit in the target document.

[0004] In a second aspect of the present disclosure, an apparatus for document processing is provided. The apparatus includes: a receiving module configured to receive an operation of adding at least one first image to an image data unit in a target document; an obtaining module configured to obtain at least one first recognition text respectively corresponding to the at least one first image, each first recognition text being obtained by text recognition of the corresponding first image; a merging module configured to merge the at least one first recognition text into a target text according to the addition order of the at least one first image; and an adding module configured to add the target text to a full text unit corresponding to the image data unit in the target document.

[0005] In a third aspect of the present disclosure, an electronic device is provided. The device includes at least one processing unit; and at least one memory coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit. The instructions, when executed by the at least one processing unit, cause the electronic device to execute the method of the first aspect.

[0006] In a fourth aspect of the present disclosure, a computer-readable storage medium is provided. A computer program is stored on the medium, and when the computer program is executed by a processor, the method of the first aspect is implemented.

[0007] In a fifth aspect of the present disclosure, a computer program product is provided. The product includes a computer program, and when the computer program is executed by a processor, the method according to the first aspect of the present disclosure is implemented.

[0008] It should be understood that the content described in this part is not intended to define the key features or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] In combination with the accompanying drawings and with reference to the following detailed description, the above and other features, advantages, and aspects of the embodiments of the present disclosure will become more apparent. In the drawings, the same or similar reference numerals denote the same or similar elements, where:

[0010] Figure 1 A schematic diagram showing an example environment in which the embodiments of the present disclosure can be implemented;

[0011] Figures 2A through 2G Schematic diagrams respectively showing examples of an editing page according to some embodiments of the present disclosure;

[0012] Figure 3 A flowchart showing a process for document processing according to some embodiments of the present disclosure;

[0013] Figure 4 A schematic block diagram showing an exemplary structure of a device for document processing according to some embodiments of the present disclosure; and

[0014] Figure 5 A block diagram showing an electronic device that can implement one or more embodiments of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0015] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although some embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. On the contrary, these embodiments are provided to more thoroughly and completely understand the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes and are not used to limit the protection scope of the present disclosure.

[0016] In the description of the embodiments of the present disclosure, the term "including" and its similar terms should be understood as open inclusion, that is, "including but not limited to". The term "based on" should be understood as "at least partially based on". The term "one embodiment" or "the embodiment" should be understood as "at least one embodiment". The term "some embodiments" should be understood as "at least some embodiments". There may also be other explicit and implicit definitions hereinafter.

[0017] In this article, unless explicitly stated, performing a step "in response to A" does not mean that the step is immediately performed after "A", but may include one or more intermediate steps.

[0018] It can be understood that the data involved in the technical solution (including but not limited to the data itself, the acquisition, use, storage, or deletion of the data) should comply with the requirements of the corresponding laws, regulations, and related provisions.

[0019] It can be understood that before using the technical solutions disclosed in the embodiments of the present disclosure, the type, scope of use, usage scenarios, etc. of the information involved in the present disclosure should be informed to the relevant users in an appropriate manner and the authorization of the relevant users should be obtained according to the relevant laws and regulations. Among them, the relevant users may include any type of right subject, such as individuals, enterprises, and groups.

[0020] For example, when receiving an active request from a user, a prompt message is sent to the relevant user to clearly prompt the relevant user that the operation requested by them will require obtaining and using the information of the relevant user, so that the relevant user can autonomously choose whether to provide information to software or hardware such as an electronic device, application program, server, or storage medium that performs the operation of the technical solution of the present disclosure according to the prompt message.

[0021] As an optional but non-limiting implementation manner, the way of sending a prompt message to the relevant user in response to receiving an active request from the relevant user may be, for example, in the form of a pop-up window. The prompt message can be presented in text in the pop-up window. In addition, the pop-up window can also carry a selection control for the user to choose "agree" or "disagree" to provide information to the electronic device.

[0022] It can be understood that the above process of notifying and obtaining user authorization is only illustrative and does not limit the implementation manner of the present disclosure. Other ways that meet the relevant laws and regulations can also be applied to the implementation manner of the present disclosure.

[0023] As used herein, the term "model" can learn the correlation between corresponding inputs and outputs from training data, so that after training, for a given input, the corresponding output can be generated. The generation of the model can be based on machine learning techniques. Deep learning is a machine learning algorithm that processes inputs and provides corresponding outputs by using multiple layers of processing units. A neural network model is an example of a model based on deep learning. In this document, "model" can also be referred to as "machine learning model", "learning model", "machine learning network" or "learning network", and these terms are used interchangeably in this document.

[0024] As discussed above, in the scenario of image text recognition, text in an image can be recognized through image recognition technology. However, conventional image recognition technology only allows users to upload one image at a time and generate the text corresponding to the currently uploaded image, and cannot integrate the text content in multiple images. In addition, conventional image recognition technology usually can only accurately parse and extract images in a specific format, and is usually only for specific application scenarios. When developing a new scenario, a large amount of development work is required. This is very difficult for users without relevant skills and is usually difficult to achieve autonomous adjustment.

[0025] In view of this, in the embodiments of the present disclosure, an improved solution for document processing is provided. According to the solution of the embodiments of the present disclosure, if a user adds at least one image to an image data unit in a target document, at least one recognition text corresponding to the at least one image is obtained, and each recognition text is obtained through text recognition of the corresponding image. According to the addition order of the at least one image, the at least one recognition text is merged into a target text, and the target text is added to the full text unit corresponding to the image data unit in the target document.

[0026] In this way, the text recognized from different images can be automatically merged together and added to the data unit of the target document. This eliminates the limitation on the number of images to be recognized and provides a convenient solution for text recognition and text editing of multiple images.

[0027] Example environment

[0028] Figure 1 FIG. shows a schematic diagram of an example environment 100 in which the embodiments of the present disclosure can be implemented. In this example environment 100, one or more applications 130-1, 130-2,..., 130-N are installed in the terminal device 120. For ease of discussion, the applications 130-1, 130-2,..., 130-N can be collectively referred to or individually referred to as the application 130. The user 140 can interact with the application 130 via the terminal device 120 and / or an attached device of the terminal device 120.

[0029] In some embodiments, the terminal device 120 may include an image acquisition component (such as a camera). The terminal device 120 may capture an image 110 through the image acquisition component. In some embodiments, the terminal device 120 may also receive the image 110 sent by other electronic devices, or obtain the image 110 stored by itself. The image 110 may contain a text sequence, and the text sequence may include text units in any font and any language.

[0030] In some embodiments, the application 130 may be downloaded and installed on the terminal device 120. In some embodiments, the application 130 may also be accessed in other ways, such as through a web page. In Figure 1 In the environment 100, in response to the application 130 being launched, the terminal device 120 may present a page 150 of the application 130.

[0031] The application 130 includes, but is not limited to, one or more of the following: a chat application (also referred to as an instant messaging application), a document application, an audio and video conferencing application, an email application, a task application, a calendar application, an Objectives and Key Results (OKR) application, and so on. Although Figure 1 a single application is shown, in fact, multiple applications may be installed on the terminal device 120. In some embodiments, the application 130 may include a multi-functional collaboration platform, such as an office collaboration platform (also referred to as an office suite). Such a multi-functional collaboration platform can provide the integration of multiple types of applications or components to facilitate activities such as office work and communication for people. In the multi-functional collaboration platform, people can start different applications or components according to their needs to complete the processing, sharing, communication, etc. of corresponding content entities.

[0032] The application 130 may provide a content entity 135. The content entity 135 may be a content instance created by the user 140 or other users on the application 130. For example, depending on the type of the application 130, the content entity 135 may be a document (such as a word document, a pdf document, a presentation, a spreadsheet document, etc.), an email, a message (such as a session message on an instant messaging application), a calendar, a schedule, a task, a project entity, an audio, a video, an image, and so on.

[0033] In some embodiments, the environment 100 may include a text recognition model 160. The text recognition model 160 may perform text recognition on the image 110 to obtain the text included in the image 110. The text recognition model 160 may be deployed on the body of the terminal device 120 or may be deployed at a remote device. Alternatively or additionally, the text recognition model 160 may be a model based on OCR technology, and the text recognition model 160 may also be a machine learning model based on deep learning technology. It can be understood that the text recognition model 160 may also be a machine learning model or a non-machine learning model constructed based on other technologies. The type of the text recognition model 160 is not limited herein as long as it can implement the function of recognizing text from the image 110.

[0034] In some embodiments, the environment 100 may include a target model 180. The target model 180 may include any suitable machine learning model. The target model 180 may be deployed on the side of the terminal device 120, may also be deployed on the side of the server 170, or may also be deployed at a suitable remote device. In some embodiments, one or more target models 180 may be constructed based on a language model (LM). The target model used may be a content generation model that can generate a corresponding output based on the model input. In some embodiments, the machine learning model based on the language model can process model inputs in text modalities (such as natural language and / or machine language) and / or non-text modalities (such as images, voices, videos, etc.), and can generate a desired output according to the model input and the prompt words. The prompt words here are used to guide the machine learning model to generate an output that can solve the user needs indicated by the model input.

[0035] In some embodiments, the terminal device 120 communicates with the server 170 to implement the supply of the services of the application 130. The terminal device 120 may be any type of mobile terminal, fixed terminal or portable terminal, including mobile phones, desktop computers, laptop computers, notebook computers, netbook computers, tablet computers, media computers, multimedia tablets, personal communication system (PCS) devices, personal navigation devices, personal digital assistants (PDAs), audio / video players, digital cameras / cameras, television receivers, radio broadcast receivers, e-book devices, game devices, or any combination of the foregoing, including accessories and peripherals of these devices or any combination thereof. In some embodiments, the terminal device 120 can also support any type of user interface (such as a "wearable" circuit, etc.). The server 170 may be various types of computing systems / servers capable of providing computing power, including but not limited to mainframes, edge computing nodes, computing devices in a cloud environment, and the like.

[0036] It should be understood that the structures and functions of the various elements in the environment 100 are described only for exemplary purposes, without implying any limitation on the scope of the present disclosure.

[0037] Example interaction

[0038] Some example embodiments of the present disclosure will be further described below with reference to the accompanying drawings.

[0039] Figures 2A through 2G A schematic diagram showing example pages 200A to 200G (which may also be simply referred to as examples 200A to 200G) according to some embodiments of the present disclosure is shown. It should be understood that the pages shown in the drawings are only examples, and various page designs may actually exist. The various graphical elements in the pages may have different arrangements and different visual representations, one or more of which may be omitted or replaced, and one or more other elements may also exist. The embodiments of the present disclosure are not limited in this regard.

[0040] The pages shown in examples 200A to 200G may be presented at the terminal device 120. For ease of discussion, examples 200A to 200G will be described with reference to Figure 1 the environment 100. It should be noted that the operations performed by the aforementioned terminal device 120 and the subsequent operations described as being performed by the terminal device 120 may specifically be performed by a relevant application program (such as application 130) installed on the terminal device 120. In some embodiments, the operations performed on the terminal device 120 may be completed with the assistance of the server 170.

[0041] In an embodiment of the present disclosure, the terminal device 120 receives an operation of adding at least one first image to an image data unit in a target document. Here, the target document includes, but is not limited to, Word documents, PDF documents, presentation documents, table documents, and the like. Here, the image data unit may be any field or any data item in the target document. For example, in the case where the target document is a table document, the image data unit may be any cell in the table document. The at least one first image may include one or more images, and the at least one first image may contain a text sequence. Of course, the at least one first image may also not contain a text sequence, and the terminal device 120 may present a prompt message in response to not extracting text from the at least one first image. In the case of adding multiple images, these images may be added in a batch at one time or gradually added.

[0042] In some embodiments, the terminal device 120 may present a target document on an editing page in response to an editing request for the target document. The editing page includes an image addition control. The terminal device 120 may add at least one first image to the image data unit through the image addition control in response to receiving a triggering operation on the image addition control. The editing page here is not limited to including an image addition control, but may also include, but is not limited to, a document editing control, a document saving control, a text addition control, a text deletion control, and so on. It can be understood that the types of controls included in the editing page may vary according to the type of the target document or the functions of the application 130, and can be specifically selected according to actual needs.

[0043] Exemplarily, as Figure 2A shown, Example 200A shows an example of an editing page. Example 200A shows the editing page of a table (i.e., in this example, the target document is a table). Example A includes a title row 201 of the table (such as attachment, full text, incremental text, field 3, field 4, etc.), an image data unit (such as cell 202a), and image addition controls (such as attachment addition control 203 and attachment addition control 204). Here, the image addition control may be presented within the image data unit in response to a selection operation on the image data unit. For example, as Figure 2A shown, when cell 202a is selected, the terminal device 120 presents attachment addition control 203 and attachment addition control 204 within cell 202a. It can be understood that the attachment addition control 203 and the attachment addition control 204 may also be presented in other areas of the editing page, or each cell may separately display the attachment addition control 203 and the attachment addition control 204.

[0044] There are various implementation manners for the terminal device 120 to add at least one first image to the image data unit. In one example, the terminal device 120 may add an image to cell 202 through an addition manner corresponding to the attachment addition control 202 or the attachment addition control 203 in response to a triggering operation on the attachment addition control 202 or the attachment addition control 203.

[0045] In another example, the terminal device 120 may present an attachment addition pop-up window 205 on the editing page in response to a triggering operation on the attachment addition control 202 or the attachment addition control 203, as Figure 2BAs shown, an attachment addition pop-up window 205 may present an attachment addition control 206 and an attachment addition control 207. The terminal device 120 may, in response to dragging the icons of the images 208a and 208b to the area of the attachment addition control 206, add the images 208a and 208b to the cell 201a. The terminal device 120 may also, in response to a paste operation on the images 208a and 208b in the area of the attachment addition control 206, add the images 208a and 208b to the cell 201a. The terminal device 120 may further, in response to a trigger operation of the attachment addition control 207, present a file selection window. The terminal device 120 may, in response to selecting the images 208a and 208b through the file selection window, add the images 208a and 208b to the cell 201a.

[0046] In yet another example, the terminal device 120 may also, in response to dragging the image icons of the images 208a and 208b to the area of the cell 201a, add the images 208a and 208b to the cell 201a.

[0047] In some embodiments, in a case where at least one first image has been added to the image data unit, the terminal device 120 may present the at least one first image or a thumbnail of the at least one first image in the image data unit. Exemplarily, as Figure 2C shown, in a case where the images 208a and 208b have been added to the cell 202a, thumbnails of the images 208a and 208b may be displayed in the cell 202a.

[0048] In embodiments of the present disclosure, the terminal device 120 obtains at least one first recognition text respectively corresponding to at least one first image. Each first recognition text is obtained through text recognition of the corresponding first image.

[0049] In some embodiments, the terminal device 120 may provide at least one first image to the text recognition model 160 to perform text recognition on the at least one first image by using the text recognition model 160. The terminal device 120 may receive at least one first recognition text fed back by the text recognition model 160. Exemplarily, the terminal device 120 may, in response to the images 208a and 208b being added to the cell 202a, provide the images 208a and 208b to the text recognition model 160, perform text recognition on the images 208a and 208b by using the text recognition model 160, and receive the text 209a (e.g., may include the characters "XXXXXX number: 12345") and 209b (e.g., may include the characters "YYYYYYYYYY") fed back by the text recognition model 160.

[0050] Alternatively or additionally, the terminal device 120 may provide the at least one first image to the text recognition model 160 in the order of addition, and sequentially obtain at least one first recognized text fed back by the text recognition model 160. For example, the terminal device 120 first adds the image 208a to the cell 202a, and then adds the image 208b to the cell 202a. The terminal device 120 may provide the image 208a to the text recognition model 160 and receive the text 209a fed back by the text recognition model 160. Thereafter, the terminal device 120 may provide the image 208b to the text recognition model 160 and receive the text 209b fed back by the text recognition model 160. Of course, the terminal device 120 may also provide the at least one first image to the text recognition model 160 in other orders or synchronously. Here, the order and manner in which the terminal device 120 provides the first image to the text recognition model 160 are not limited.

[0051] In an embodiment of the present disclosure, the terminal device 120 may merge at least one first recognized text into a target text according to the addition order of the at least one first image. Exemplarily, the terminal device 120 may receive the text 209a and the text 209b fed back by the text recognition model 160 in the order of addition. The terminal device 120 may merge the text 209a and the text 209b into the text 209c (for example, including the characters "XXXXXX Serial number: 12345YYYYYYYYYY").

[0052] In an embodiment of the present disclosure, the terminal device 120 may add the target text to the full-text unit corresponding to the image data unit in the target document. Exemplarily, as Figure 2C shown, the cell 202a is an example of an image data unit, and the cell 202b is an example of a full-text unit corresponding to the cell 202a. The terminal device 120 may add the text 209c formed by merging the text 209a and the text 209b to the cell 202b.

[0053] It should be understood that the full-text unit may be any data unit in the target document that has a corresponding relationship with the image data unit. For example, as Figure 2C shown, the cell in the column corresponding to the title "Attachment" in the table may be determined as the image data unit, and the cell in the column corresponding to the title "Field 1" in the table may be determined as the full-text unit. The image data unit and the full-text unit in the same row correspond to each other. Thus, in the case where the cell 202a is the image data unit, the cell 202b is the full-text unit corresponding to the cell 202a. Another example is that, as Figure 2CThe cell in the first row of the shown table is determined as an image data unit. The cell in the second row of the table can be determined as a full-text unit. It can be determined that the image data unit and the full-text unit in the same column have a corresponding relationship. Also, for example, Figure 2C the cell at the intersection of the first row and the first column of the shown table can be determined as an image data unit, and the cell at the intersection of the second row and the second column can be determined as the corresponding full-text unit.

[0054] Alternatively or additionally, the terminal device 120 can obtain a configuration instruction for the image data unit or a configuration instruction for the full-text unit, and determine the corresponding relationship between the image data unit and the full-text unit. Exemplarily, as Figure 2D shown, the terminal device 120 can, in response to the selection operation of the "Field 1" title 201b, present an instruction input pop-up window 210 on the editing page. The instruction input pop-up window 210 can include an instruction input box 211, an instruction cancel control 212, and an instruction confirmation control 213. The user 140 can input a configuration instruction 214 into the instruction input box 211 through the terminal device 120 or an attached device of the terminal device 120. The terminal device 120 can, in response to the triggering operation for the instruction confirmation control 213, determine the corresponding relationship between the cell 202a and the cell 202b based on the configuration instruction 214.

[0055] It can be understood that the user 140 can also trigger the "Attachment" title 201a to cause the editing page to present an instruction input pop-up window similar to the instruction input pop-up window 210. The terminal device 120 can, in response to the configuration instruction obtained through the instruction input pop-up window corresponding to the "Attachment" title 201a, determine the corresponding relationship between the cell 202a and the cell 202b. In actual applications, the corresponding relationship between the image data unit and the full-text unit can also be constructed in other ways.

[0056] In some embodiments, the terminal device 120 can add the first recognition text in at least one first recognition text corresponding to the latest added first image to the incremental text unit corresponding to the image data unit. For example, as Figure 2C shown, if the cell 202a is an image data unit, the cell 202c is the incremental text unit corresponding to the cell 202a, the image 208b is the latest added first image in the cell 202a, and the text 209b is the text recognized from the image 208, then the text 209b can be added to the cell 202c. In this way, it is beneficial for the user 140 to determine the text content recognized from the latest added first image.

[0057] It can be understood that the incremental text unit can be any data unit in the target document that has a corresponding relationship with the image data unit. Exemplarily, as Figure 2CAs shown, the cells in the column corresponding to the "Attachment" header in the table (i.e., the first column) can be determined as image data units, the cells in the column corresponding to the "Field 2" header (i.e., the third column) can be determined as incremental text units, and it can be determined that the image data units and the incremental text units in the same row have a corresponding relationship. Also, for example, as Figure 2C shown, the cells in the first row of the table can be determined as image data units, the cells in the third row can be determined as incremental text units, and it can be determined that the image data units and the incremental text units in the same column have a corresponding relationship.

[0058] In some embodiments, the terminal device 120 can receive an operation to add a second image to the image data unit, add the second recognition text corresponding to the second image to the incremental text unit to replace the first recognition text in the incremental text unit. Here, the second recognition text is obtained by text recognition of the second image. The terminal device 120 can add the second recognition text to the full text unit to update the target text in the full text unit. That is, the text recognized in the newly added image and the existing text in the full text unit are combined into a new target text. Alternatively or additionally, the terminal device 120 can provide the second image to the text recognition model 160, use the text recognition model 160 to perform text recognition on the second image, and obtain the second recognition text corresponding to the second image (i.e., the text included in the second image). The terminal device 120 receives the second recognition text fed back by the text recognition model 160.

[0059] Exemplarily, as Figure 2C and Figure 2E shown, the image 208c can be added to the cell 202a. The terminal device 120 can provide the image 208c to the text recognition model 160, perform text recognition on the image 208c through the text recognition model 160, and obtain the text 209d (e.g., may include the characters "ZZZZZZZZZZZZ"). The terminal device 120 can add the text 209 to the cell 202c to replace the text 209b in the cell 202c. The terminal device 120 can also combine the text 209b with the text 209c in the cell 202b to form the text 209e, add the text 209e to the cell 202b to replace the text 209c in the cell 202b.

[0060] In some embodiments, the terminal device 120 may add the field value of the target field extracted from the target text to the extraction data unit of the target document. In this way, the editability and convenience of the target document can be improved. Here, the target field can be any field in the target text, and the field value corresponds to the target field. The field value may include, but is not limited to, text, numbers, symbols, etc. For example, the target field may include, but is not limited to, name, gender, address, date, number, product, category, customer, company name, question, answer, etc. When the target field is an address, the field value may include a combination of text and numbers. When the target field is a date, the field value may include numbers (e.g., XXXX-XX-XX), or may also include a combination of numbers and text (e.g., XXXX year XX month XX day). Here, the extraction data unit is similar to the full-text unit and the incremental text unit, and can be any data unit in the target document. Taking the target document as a table as an example, the extraction data unit can be any cell in the table. Exemplarily, as Figure 2F shown, when the target field is "number", the terminal device may extract the field value 209f (e.g., 12345) of "number" from the text 209e and add the field value 209f to the cell 202d.

[0061] In some embodiments, the terminal device 120 may receive a text extraction instruction for the target field. The text extraction instruction at least includes the identifier of the full-text unit and the field feature for characterizing the target field. The terminal device 120 may provide the target text and the text extraction instruction to the target model 180, so that the target model 180 extracts the field value of the target field from the target text based on the text extraction instruction. The terminal device 120 may add the field value fed back by the target model 180 to the extraction data unit.

[0062] Alternatively or additionally, the terminal device 120 may send the target text and the text extraction instruction to the server 170, and the server 170 sends the target text and the text extraction instruction to the target model 180.

[0063] The identifier of the full-text unit here may include various information capable of identifying the full-text unit. Exemplarily, taking the target document as a table as an example, the full-text unit can be a cell, and the identifier of the full-text unit may include, but is not limited to, row title, column title, cell coordinates, etc. It can be understood that when the target document is other documents, the full-text unit can also be identified by other information.

[0064] The field feature here is used to indicate the target field. It should be understood that the field feature may include various characteristic information capable of indicating the target field. Alternatively or additionally, the field feature includes, but is not limited to, the target field itself, the text sequence for describing the target field, or the identifier capable of indicating the target field, etc.

[0065] In some embodiments, the terminal device 120 may present an instruction input box and an instruction confirmation control in response to an instruction input request. The terminal device 120 may receive user input to the instruction input box. The terminal device 120 may also determine the user input received in the instruction input box as a text extraction instruction in response to receiving a trigger operation on the instruction confirmation control. Alternatively or additionally, the text extraction instruction may be written in natural language.

[0066] Exemplarily, as Figure 2F and Figure 2G shown, in Example 200G, the user 140 may trigger the "Field 3" caption 201d through the terminal device 120 or an attached device of the terminal device 120 to generate an instruction input request. The terminal device 120 may present an instruction input pop-up window 220 on the editing page in response to the instruction input request. The instruction input pop-up window 220 may include an instruction input box 221, an instruction confirmation control 222, and an instruction cancellation control 223. The user 140 may input text extraction instruction 224 into the instruction input box 221 through the terminal device 120 or an attached device of the terminal device 120. For example, please extract [Number] from [Field 1], and only give [Number] in your answer without providing any other additional content. The text extraction instruction 224 includes an identifier 225 of the full text unit (e.g., [Field 1]). The text extraction instruction also includes a field feature 226 for characterizing the target field (e.g., [Number]). The terminal device 120 may send the text extraction instruction 224 to the server 170 in response to a trigger operation on the instruction confirmation control 222, and provide the text extraction instruction 224 to the target model 180 through the server 170 to trigger the target model 180 to determine the target value of the target field. Thereafter, the server 170 may feedback the field value 209f (e.g., 12345) determined by the target model 180 to the terminal device 120. The terminal device 120 may add the field value 209f to the cell 202d, as Figure 2F shown.

[0067] In summary, according to the embodiments of the present disclosure, it is possible to recognize text in at least one image, merge the recognized at least one text, and add the merged text to a data unit of a target document, providing a convenient solution for text recognition and text editing of multiple images.

[0068] Example processes, devices, and apparatuses

[0069] Figure 3 FIG. shows a flowchart of a process 300 for document processing according to some embodiments of the present disclosure. The process 300 may be implemented at the terminal device 120 or may be jointly implemented by the terminal device 120 and the server 170.

[0070] At block 310, the terminal device 120 receives an operation of adding at least one first image to an image data unit in a target document.

[0071] At block 320, the terminal device 120 obtains at least one first recognition text respectively corresponding to the at least one first image, and each first recognition text is obtained by text recognition of the corresponding first image.

[0072] At block 330, the terminal device 120 combines the at least one first recognition text into a target text according to the addition order of the at least one first image.

[0073] At block 340, the terminal device 120 may add the target text to a full - volume text unit corresponding to the image data unit in the target document.

[0074] In some embodiments, process 300 further includes: adding the first recognition text corresponding to the most recently added first image in the at least one first recognition text to an incremental text unit corresponding to the image data unit.

[0075] In some embodiments, process 300 further includes: receiving an operation of adding a second image to the image data unit, where the second image is different from the at least one first image; replacing the first recognition text in the incremental text unit with a second recognition text corresponding to the second image, where the second recognition text is obtained by text recognition of the second image; and adding the second recognition text to the full - volume text unit to update the target text in the full - volume text unit.

[0076] In some embodiments, receiving an operation of adding at least one first image to an image data unit in a target document includes: in response to an edit request for the target document, presenting the target document on an edit page, where the edit page includes an image addition control; and in response to receiving a trigger operation on the image addition control, adding at least one first image to the image data unit through the image addition control.

[0077] In some embodiments, obtaining at least one first recognition text includes: providing the at least one first image to a text recognition model so that the text recognition model performs text recognition on the at least one first image; and receiving at least one first recognition text fed back by the text recognition model.

[0078] In some embodiments, process 300 further includes: adding the field value of a target field extracted from the target text to an extraction data unit of the target document.

[0079] In some embodiments, adding the field value of the target field extracted from the target text to the extraction data unit of the target document includes: receiving a text extraction instruction for the target field, where the text extraction instruction at least includes the identifier of the full text unit and the field feature used to characterize the target field; providing the target text and the text extraction instruction to the target model so that the target model extracts the field value of the target field from the target text based on the text extraction instruction; and adding the field value fed back by the target model to the extraction data unit.

[0080] In some embodiments, receiving the text extraction instruction includes: in response to an instruction input request, presenting an instruction input box and an instruction confirmation control; receiving user input to the instruction input box; and in response to receiving a trigger operation on the instruction confirmation control, determining the user input received in the instruction input box as the text extraction instruction.

[0081] In some embodiments, the text extraction instruction is written in natural language.

[0082] Embodiments of the present disclosure also provide corresponding apparatuses for implementing the above methods or processes. Figure 4 An exemplary structural block diagram of an apparatus 400 for document processing according to some embodiments of the present disclosure is shown. The apparatus 400 may be implemented as or included in a terminal device 120. Each module / component in the apparatus 400 may be implemented by hardware, software, firmware, or any combination thereof.

[0083] As Figure 4 shown, the apparatus 400 may include a receiving module 410, an obtaining module 420, a merging module 430, and an adding module 440. The receiving module 410 is configured to receive an operation of adding at least one first image to the image data unit in the target document. The obtaining module 420 is configured to obtain at least one first recognition text respectively corresponding to at least one first image, and each first recognition text is obtained by text recognition of the corresponding first image. The merging module 430 is configured to merge at least one first recognition text into a target text according to the adding order of at least one first image. The adding module 440 is configured to add the target text to the full text unit corresponding to the image data unit in the target document.

[0084] In some embodiments, the adding module 440 is further configured to: add the first recognition text corresponding to the most recently added first image in at least one first recognition text to the incremental text unit corresponding to the image data unit.

[0085] In some embodiments, the adding module 440 is further configured to: receive an operation of adding a second image to the image data unit, where the second image is different from at least one first image; replace the first recognition text in the incremental text unit with a second recognition text corresponding to the second image, where the second recognition text is obtained by performing text recognition on the second image; and add the second recognition text to the full text unit to update the target text in the full text unit.

[0086] In some embodiments, the adding module 440 is further configured to: in response to an editing request for a target document, present the target document on an editing page, where the editing page includes an image adding control; and in response to receiving a triggering operation on the image adding control, add at least one first image to the image data unit through the image adding control.

[0087] In some embodiments, the obtaining module 420 is further configured to: provide at least one first image to a text recognition model, so that the text recognition model performs text recognition on at least one first image; and receive at least one first recognition text fed back by the text recognition model.

[0088] In some embodiments, the adding module 440 is further configured to: add the field value of the target field extracted from the target text to the extraction data unit of the target document.

[0089] In some embodiments, the adding module 440 is further configured to: receive a text extraction instruction for a target field, where the text extraction instruction at least includes an identifier of the full text unit and a field feature for characterizing the target field; provide the target text and the text extraction instruction to a target model, so that the target model extracts the field value of the target field from the target text based on the text extraction instruction; and add the field value fed back by the target model to the extraction data unit.

[0090] In some embodiments, the adding module 440 is further configured to: in response to an instruction input request, present an instruction input box and an instruction confirmation control; receive user input in the instruction input box; and in response to receiving a triggering operation on the instruction confirmation control, determine the user input received in the instruction input box as the text extraction instruction.

[0091] In some embodiments, the text extraction instruction is written in natural language.

[0092] The units and / or modules included in apparatus 400 may be implemented in various ways, including software, hardware, firmware, or any combination thereof. In some embodiments, one or more units and / or modules may be implemented using software and / or firmware, such as machine-executable instructions stored on a storage medium. In addition to or as an alternative to the machine-executable instructions, some or all of the units and / or modules in apparatus 400 may be implemented at least in part by one or more hardware logic components. By way of example and not limitation, exemplary types of hardware logic components that may be used include field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system on a chip (SOCs), complex programmable logic devices (CPLDs), and the like.

[0093] It should be understood that one or more steps in the above methods may be performed by a suitable electronic device or a combination of electronic devices. Such an electronic device or combination of electronic devices may include, for example, Figure 1 the terminal device 120 in

[0094] Figure 5 A block diagram of an electronic device 500 in which one or more embodiments of the present disclosure may be implemented is shown. It should be understood that Figure 5 the electronic device 500 shown is merely exemplary and should not constitute any limitation on the functions and scope of the embodiments described herein. Figure 5 The electronic device 500 shown may be used to implement Figure 1 the terminal device 120 of Figure 4 or to implement

[0095] As Figure 5 shown, the electronic device 500 is in the form of a general-purpose electronic device. The components of the electronic device 500 may include, but are not limited to, one or more processors or processing units 510, processing unit 520, storage device 530, one or more communication units 540, one or more input devices 550, and one or more output devices 560. The processing unit 510 may be an actual or virtual processor and be capable of performing various processes according to the programs stored in the processing unit 520. In a multi-processor system, multiple processing units execute computer-executable instructions in parallel to improve the parallel processing ability of the electronic device 500.

[0096] The electronic device 500 generally includes multiple computer storage media. Such media can be any accessible media that the electronic device 500 can access, including but not limited to volatile and non-volatile media, removable and non-removable media. The processing unit 520 can be volatile memory (such as registers, caches, random access memory (RAM)), non-volatile memory (such as read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory), or some combination thereof. The storage device 530 can be removable or non-removable media and can include machine-readable media, such as flash drives, magnetic disks, or any other media that can be capable of storing information and / or data and can be accessed within the electronic device 500.

[0097] The electronic device 500 can further include additional removable / non-removable, volatile / non-volatile storage media. Although not shown in Figure 5 , a disk drive for reading from or writing to a removable, non-volatile magnetic disk (such as a "floppy disk") and an optical disk drive for reading from or writing to a removable, non-volatile optical disk can be provided. In these cases, each drive can be connected to a bus (not shown) by one or more data media interfaces. The processing unit 520 can include a computer program product 525 having one or more program modules that are configured to perform the various methods or actions of the various embodiments of the present disclosure.

[0098] The communication unit 540 enables communication with other electronic devices via a communication medium. Additionally, the functions of the components of the electronic device 500 can be implemented by a single computing cluster or multiple computer machines that can communicate via a communication connection. Thus, the electronic device 500 can operate in a networked environment using a logical connection with one or more other servers, network personal computers (PCs), or another network node.

[0099] The input device 550 can be one or more input devices, such as a mouse, keyboard, trackball, etc. The output device 560 can be one or more output devices, such as a display, speaker, printer, etc. The electronic device 500 can also communicate with one or more external devices (not shown) as needed via the communication unit 540, such as storage devices, display devices, etc., communicate with one or more devices that enable a user to interact with the electronic device 500, or communicate with any device that enables the electronic device 500 to communicate with one or more other electronic devices (such as a network card, modem, etc.). Such communication can be performed via an input / output (I / O) interface (not shown).

[0100] According to an exemplary implementation of the present disclosure, a computer-readable storage medium is provided, on which computer-executable instructions are stored, and the computer-executable instructions are executed by a processor to implement the method described above. According to an exemplary implementation of the present disclosure, a computer program product is also provided, the computer program product being tangibly stored on a non-transitory computer-readable medium and including computer-executable instructions, and the computer-executable instructions being executed by a processor to implement the method described above.

[0101] Aspects of the present disclosure are described herein with reference to the flowcharts and / or block diagrams of methods, apparatuses, devices, and computer program products according to the present disclosure. It should be understood that each block of the flowcharts and / or block diagrams, and the combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer-readable program instructions.

[0102] These computer-readable program instructions can be provided to a processing unit of a general-purpose computer, a special-purpose computer, or other programmable data processing device, thereby producing a machine such that when these instructions are executed by the processing unit of the computer or other programmable data processing device, a device is produced that implements the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium, and these instructions cause a computer, a programmable data processing device, and / or other devices to work in a specific manner. Thus, the computer-readable medium storing the instructions includes a manufactured article that includes instructions for implementing various aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.

[0103] The computer-readable program instructions can be loaded onto a computer, other programmable data processing device, or other device, such that a series of operation steps are executed on the computer, other programmable data processing device, or other device to produce a computer-implemented process, so that the instructions executed on the computer, other programmable data processing device, or other device implement the functions / actions specified in one or more blocks of the flowchart and / or block diagram.

[0104] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various implementations of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a portion of an instruction, which contains one or more executable instructions for implementing a specified logical function. In some alternative implementations, the functions noted in the blocks may occur out of the order noted in the figures. For example, two consecutive blocks may in fact be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending upon the functionality involved. It should also be noted that each block of the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented by special purpose hardware-based systems that perform the specified functions or acts, or by combinations of special purpose hardware and computer instructions.

[0105] The various implementations of the present disclosure have been described above. The above description is exemplary, not exhaustive, and is not limited to the disclosed implementations. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described implementations. The choice of terms used herein is intended to best explain the principles of the implementations, the practical application, or improvements made to the technology in the marketplace, or to enable other ordinary skill in the art to understand the various implementations disclosed herein.

Claims

1. A method for document processing, comprising: Receiving an operation of adding at least one first image to an image data unit in a target document; Obtaining at least one first recognition text respectively corresponding to the at least one first image, where the first recognition text is obtained through text recognition of the corresponding first image; Merging the at least one first recognition text into a target text according to the addition order of the at least one first image; And Adding the target text to a full - volume text unit corresponding to the image data unit in the target document, where the full - volume text unit is a data unit in the target document that has a corresponding relationship with the image data unit.

2. The method according to claim 1, further comprising: Adding the first recognition text corresponding to the latest - added first image in the at least one first recognition text to an incremental text unit corresponding to the image data unit, where the incremental text unit is another data unit in the target document that has a corresponding relationship with the image data unit.

3. The method according to claim 2, further comprising: Receiving an operation of adding a second image to the image data unit, where the second image is different from the at least one first image; Replacing the first recognition text in the incremental text unit with a second recognition text corresponding to the second image, where the second recognition text is obtained through text recognition of the second image; And Adding the second recognition text to the full - volume text unit to update the target text in the full - volume text unit.

4. The method according to claim 1, wherein receiving an operation of adding at least one first image to an image data unit in a target document comprises: In response to an edit request for the target document, presenting the target document on an edit page, where the edit page includes an image - adding control; And In response to receiving a trigger operation on the image - adding control, adding the at least one first image to the image data unit through the image - adding control.

5. The method according to claim 1, wherein obtaining the at least one first recognition text comprises: Providing the at least one first image to a text - recognition model so that the text - recognition model performs text recognition on the at least one first image; And Receiving the at least one first recognition text fed back by the text - recognition model.

6. The method according to claim 1, further comprising: Adding the field value of a target field extracted from the target text to an extraction data unit in the target document.

7. The method according to claim 6, wherein adding the field value of the target field to the extraction data unit comprises: Receiving a text - extraction instruction for the target field, where the text - extraction instruction at least includes an identifier of the full - volume text unit and a field feature for characterizing the target field; Providing the target text and the text - extraction instruction to a target model so that the target model extracts the field value of the target field from the target text based on the text - extraction instruction; And Add the field value fed back by the target model to the extraction data unit.

8. The method according to claim 7, wherein receiving a text extraction instruction comprises: In response to an instruction input request, presenting an instruction input box and an instruction confirmation control; Receiving user input to the instruction input box; And In response to receiving a trigger operation on the instruction confirmation control, determining the user input received in the instruction input box as the text extraction instruction.

9. The method according to claim 7, wherein the text extraction instruction is written in natural language.

10. An apparatus for document processing, comprising: A receiving module configured to receive an operation of adding at least one first image to an image data unit in a target document; An obtaining module configured to obtain at least one first recognition text respectively corresponding to the at least one first image, the first recognition text being obtained by text recognition of the corresponding first image; A merging module configured to merge the at least one first recognition text into a target text according to the adding order of the at least one first image; And An adding module configured to add the target text to a full - volume text unit corresponding to the image data unit in the target document, the full - volume text unit being a data unit in the target document having a corresponding relationship with the image data unit.

11. An electronic device, comprising: At least one processing unit; And At least one memory, the at least one memory being coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, the instructions when executed by the at least one processing unit causing the electronic device to perform the method according to any one of claims 1 to 9.

12. A computer - readable storage medium having stored thereon a computer program, the computer program being executable by a processor to implement the method according to any one of claims 1 to 9.

13. A computer program product comprising a computer program, wherein the computer program when executed by a processor implements the method according to any one of claims 1 to 9.

Citation Information

Patent Citations

  • PDF (Portable Document Format) document processing method and device, equipment and medium

    CN117173729A

  • Method and system for extracting information from a document image

    US20220156490A1