Image-text proofreading method, device and equipment and readable storage medium

By obtaining external supplementary information of images and text, and using multimodal large language model for graphic and text proofreading, the problem of insufficient accuracy of graphic and text proofreading in the existing technology is solved, and a more efficient and accurate graphic and text proofreading effect is achieved.

CN120495854APending Publication Date: 2025-08-15SHANGHAI MIDU INFORMATION TECH CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510977894.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-16
Publication Date
2025-08-15

AI Technical Summary

Technical Problem

The existing graphic and text proofreading methods lack accuracy, making it difficult to effectively utilize the correlation between pictures and text, and lack of knowledge in specific fields, resulting in complex models and low accuracy.

Method used

By obtaining the text content and external supplementary information of the image to be proofed, the multimodal large language model is used to proofread the graphic and text, and combining the information of the pictures and text for more accurate proofreading.

Benefits of technology

It improves the information richness and accuracy of graphic and text proofreading, enhances the model's knowledge base, and generates more accurate and specific text content.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120495854A_ABST
    Figure CN120495854A_ABST
Patent Text Reader

Abstract

The invention provides an image-text proofreading method and device, equipment and a readable storage medium. The method comprises the steps of obtaining a to-be-corrected image; determining text content corresponding to the to-be-corrected image; acquiring external supplementary information of the text content; determining to-be-corrected information corresponding to the to-be-corrected image, the text content and the external supplementary information; inputting the information to be proofread into a multi-modal large language model to obtain an image-text proofread result; the to-be-proofread image and the text content corresponding to the to-be-proofread image can be obtained at the same time by obtaining the to-be-proofread image and the text content corresponding to the to-be-proofread image, the information richness of subsequent image-text proofread is improved, and a knowledge base of the model can be supplemented by utilizing a large amount of external knowledge by obtaining the external supplementary information corresponding to the text content, so that the proofread efficiency is improved Therefore, the more accurate and specific text is generated, and the to-be-corrected image and the text content are combined, so that the accuracy of image-text correction is greatly improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application belongs to the field of image processing technology, and relates to a method for proofreading images and texts, and in particular to a method, device, equipment and readable storage medium for proofreading images and texts. Background Art

[0002] Existing image-text proofreading methods have some limitations when extracting image content. First, these methods tend to focus too much on the details of the image, while the types of objects in the image are too numerous to be enumerated one by one. Second, in order to process an image, multiple visual models are usually required to work together. For example, when identifying an architectural image, it is necessary to first perform image classification, and then perform target detection or semantic segmentation, which makes the model complex and inaccurate. In addition, these methods lack the ability to logically analyze the correlation between image content and text content, have poor robustness, and lack knowledge in specific fields. On the other hand, these models also lack timeliness and cannot effectively utilize the latest information for knowledge analysis, which affects the accuracy of image-text proofreading. Therefore, how to improve the accuracy of image-text proofreading has become a technical problem that needs to be solved urgently. Summary of the Invention

[0003] The present application provides a method, device, equipment and readable storage medium for image and text proofreading, which are used to solve the technical problem of the lack of accurate image and text proofreading methods in the prior art.

[0004] In a first aspect, an embodiment of the present application provides a method for image-text proofreading, the method comprising: obtaining an image to be proofread; determining the text content corresponding to the image to be proofread; obtaining external supplementary information of the text content; determining the information to be proofread corresponding to the image to be proofread, the text content and the external supplementary information; inputting the information to be proofread into a multimodal large language model to obtain an image-text proofreading result.

[0005] In an implementation of the first aspect, determining the text content corresponding to the image to be proofread includes: extracting text information in the image to be proofread based on an OCR model, and determining the text information in the image to be proofread as the text content corresponding to the image to be proofread.

[0006] In an implementation of the first aspect, determining the text content corresponding to the image to be proofread includes: detecting external description text of the image to be proofread based on a text detector, and determining the external description text as the text content corresponding to the proofreading image.

[0007] In an implementation of the first aspect, determining the text content corresponding to the image to be proofread includes: extracting text information in the image to be proofread based on an OCR model; detecting external description text of the image to be proofread based on a text detector; combining the text information in the image to be proofread and the external description text of the image to be proofread, and determining the combined text information and the external description text as the text content corresponding to the image to be proofread.

[0008] In an implementation of the first aspect, obtaining the external supplementary information of the text content includes: extracting an entity structure corresponding to the text information; and obtaining the external supplementary information corresponding to the entity structure based on an external database.

[0009] In an implementation of the first aspect, determining the information to be proofread corresponding to the image to be proofread, the text content, and the external supplementary information includes: determining first prompt information corresponding to the image to be proofread; determining second prompt information corresponding to the text content; determining third prompt information corresponding to the external supplementary information; and performing a combination operation on the first prompt information, the second prompt information, and the third prompt information to obtain the information to be proofread.

[0010] In an implementation of the first aspect, the method further includes: performing a post-processing operation on the image and text proofreading result.

[0011] The embodiment of the present application provides a method for image-text proofreading, including: obtaining an image to be proofread; determining the text content corresponding to the image to be proofread; obtaining external supplementary information of the text content; determining the information to be proofread corresponding to the image to be proofread, the text content, and the external supplementary information; and inputting the information to be proofread into a multimodal large language model to obtain an image-text proofreading result. In the above method, by obtaining the image to be proofread and the text content corresponding to the image to be proofread, the image to be proofread and the text content can be obtained at the same time, thereby improving the information richness of subsequent image-text proofreading. In addition, by obtaining the external supplementary information corresponding to the text content, a large amount of external knowledge can be used to supplement the knowledge base of the model itself, thereby generating more accurate and specific text. In combination with the image to be proofread and the text content, the accuracy of image-text proofreading is greatly improved.

[0012] In a second aspect, an embodiment of the present application provides a graphic proofreading device, which includes: an image acquisition module for acquiring an image to be proofread; a text content determination module for determining the text content corresponding to the image to be proofread; an external supplementary information acquisition module for acquiring external supplementary information of the text content; an information to be proofreading determination module for determining the information to be proofread corresponding to the image to be proofread, the text content and the external supplementary information; and a graphic proofreading result acquisition module for inputting the information to be proofread into a multimodal large language model to obtain a graphic proofreading result.

[0013] In a third aspect, an embodiment of the present application provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the image-text proofreading method described in any one of the first aspects of the embodiment of the present application is implemented.

[0014] In a fourth aspect, an embodiment of the present application provides an electronic device comprising a memory and a processor; the memory is used to store a computer program; the processor is used to execute the computer program and implement the image and text proofreading method as described in any one of the first aspects of the embodiment of the present application when executing the computer program. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] Figure 1A Shown is a structural diagram of an application scenario of a method for image-text proofreading in one embodiment of the present application.

[0016] Figure 1B Shown is a flowchart of a method for proofreading images and texts provided in one embodiment of the present application.

[0017] Figure 2 Shown is a flowchart of determining the text content corresponding to the image to be proofread in one embodiment of the present application.

[0018] Figure 3 Shown is a flowchart of obtaining external supplementary information of the text content in one embodiment of the present application.

[0019] Figure 4 Shown is a flowchart of determining the information to be proofread corresponding to the image to be proofread, the text content, and the external supplementary information in one embodiment of the present application.

[0020] Figure 5 Shown is a flowchart of another image-text proofreading method provided by an embodiment of the present application.

[0021] Figure 6 Shown is a schematic diagram of an image-text proofreading device provided by an embodiment of the present application.

[0022] Figure 7 Shown is a structural schematic diagram of an electronic device provided by an embodiment of the present application.

[0023] Component number description DETAILED DESCRIPTION

[0024] The following describes the embodiments of the present application through specific examples. Those skilled in the art can easily understand the other advantages and effects of the present application from the content disclosed in this specification. The present application can also be implemented or applied through other different specific embodiments. The details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present application. It should be noted that the following embodiments and features in the embodiments can be combined with each other unless they conflict.

[0025] It should be noted that the illustrations provided in the following embodiments are only schematic illustrations of the basic concept of the present application. Therefore, the drawings only show components related to the present application and are not drawn according to the number, shape and size of components in actual implementation. In actual implementation, the shape, quantity and proportion of each component can be changed at will, and the component layout type may also be more complicated.

[0026] The prior art lacks an accurate method for proofreading images and texts.

[0027] At least in response to the above-mentioned problems, an embodiment of the present application provides a method for image-text proofreading. The method can obtain an image to be proofread and temporarily store the image to be proofread in a first prompt box; extract text information from the image to be proofread and temporarily store the image to be proofread and the text information in a second prompt box; obtain external supplementary information corresponding to the text information and temporarily store the image to be proofread and the external supplementary information in a third prompt box; perform prompt processing on the image to be proofread in the first prompt box, the image to be proofread and the text information in the second prompt box, and the image to be proofread and the external supplementary information in the third prompt box to obtain the information to be proofread after prompt processing, and input the information to be proofread into a multimodal large model to obtain the image-text proofreading result output by the multimodal large model, which can solve the technical problem of the lack of accurate image-text proofreading methods in the existing technology.

[0028] Figure 1A Shown is a schematic diagram of an application scenario of the image and text proofreading method provided in one embodiment of the present application. Figure 1A As shown, this application scenario includes an image acquisition device and an electronic device, which are in communication with each other. The image acquisition device is used to capture the image to be proofread and transmit the captured image to the electronic device. In the electronic device, the image to be proofread undergoes a series of image and text proofreading processes to obtain an image-text proofreading result.

[0029] The technical solutions in the embodiments of the present application will be described in detail below with reference to the accompanying drawings in the embodiments of the present application.

[0030] Figure 1B The flowchart of the image and text proofreading method provided by one embodiment of the present application is shown. Figure 1B As shown, the image and text proofreading method provided in the embodiment of the present application includes the following steps S11 to S15.

[0031] S11, obtaining the image to be proofread.

[0032] For example, the image to be proofread may be obtained by directly photographing with a digital camera, a smart phone, or the like, or may be obtained from other services or databases through an application programming interface (API).

[0033] It should be noted that the above-listed methods for obtaining the image to be proofread are only for illustrative purposes. In actual applications, any other appropriate method can be selected to obtain the image to be proofread, and this application does not impose any limitation on this.

[0034] Illustratively, the graphics to be proofread include publication-related images, business-related images, product display pictures, icons, etc.

[0035] It should be noted that the types of images to be proofread listed above are only for illustrative purposes. In the actual image and text proofreading process, the images to be proofread can be any suitable images, and this application does not impose any restrictions on the types of images to be proofread.

[0036] The image to be proofread is transmitted to the first prompt box to generate first prompt information.

[0037] Exemplarily, the first prompt box includes a dialog box / pop-up window, a prompt bar / notification bar, etc.

[0038] It should be noted that the types of the first prompt boxes listed above are only for illustrative purposes. In actual applications, any appropriate first prompt box can be selected according to specific application requirements, and this application does not impose any restrictions on this.

[0039] S12: Determine the text content corresponding to the image to be proofread.

[0040] In some embodiments, determining the text content corresponding to the image to be proofread includes: extracting text information in the image to be proofread based on an Optical Character Recognition (OCR) model, and determining the text information in the image to be proofread as the text content corresponding to the image to be proofread.

[0041] Exemplarily, text recognition models include: convolutional neural networks, recurrent neural networks, long short-term memory networks, etc.

[0042] It should be noted that the types of text recognition models listed above are only for illustrative purposes. In actual applications, any other suitable text recognition model can be selected according to actual needs, and this application does not impose any restrictions on this.

[0043] This embodiment makes up for the situation where the other image-text proofreading methods mentioned above are not applicable when the image to be proofread is pure text, and expands the application scenarios of image-text proofreading; it can promote the development of image-text proofreading technology in a more efficient and intelligent direction to meet the needs of multiple fields and multiple scenarios.

[0044] In some embodiments, determining the text content corresponding to the image to be proofread includes: detecting external description text of the image to be proofread based on a text detector, and determining the external description text as the text content corresponding to the proofread image.

[0045] The external description text refers to text that is independent of the image to be proofread, and the description object of the text is the image to be proofread.

[0046] Exemplarily, the text detector includes a RAFT text detector, an optical character recognition tool, and the like.

[0047] It should be noted that the text detectors listed above are only for illustrative purposes. In actual applications, any other suitable text detector can be selected according to specific application requirements, and this application will not go into details.

[0048] It should be noted that in actual applications, any appropriate text detector can be selected according to specific application requirements, and this application does not impose any restrictions on this.

[0049] After obtaining the text content corresponding to the image to be proofread, the text content corresponding to the image to be proofread is transmitted to the second prompt box to generate second prompt information.

[0050] It should be noted that the type of the second prompt box is similar to that of the first prompt box mentioned above, and this application will not go into details.

[0051] In actual applications, any appropriate second prompt box can be selected to generate the second prompt information according to specific application requirements, and this application does not impose any limitation on this.

[0052] It should be noted that the second prompt box and the first prompt box are independent of each other, and the types of the second prompt box and the first prompt box can be reasonably selected according to actual conditions, and this application does not impose any restrictions on this.

[0053] S13: Obtain external supplementary information of the text content.

[0054] For example, relevant information can be retrieved from external knowledge sources (such as document collections, web pages, databases, etc.) based on information retrieval enhancement technology (RAG), and then this information can be integrated into the generation process to obtain external supplementary information, and the external supplementary information can be transmitted to the third prompt box to generate third prompt information.

[0055] It should be noted that the type of the third prompt box is similar to that of the first and second prompt boxes described above, and this application will not go into details. In actual applications, any appropriate third prompt box can be selected to generate the second prompt information according to specific application requirements, and this application does not impose any restrictions on this. In addition, the first prompt box, the second prompt box, and the third prompt box are all independent of each other.

[0056] S14, determining the information to be proofread corresponding to the image to be proofread, the text content and the external supplementary information.

[0057] S15: Input the information to be proofread into a multimodal large language model to obtain a picture-text proofreading result.

[0058] In some embodiments, the method further includes: performing a structured processing operation on the image and text proofreading result.

[0059] Specifically, the image and text proofreading results output by the multimodal large model may be relatively mechanical, and the coherence of the sentences is relatively poor. In order to facilitate the extraction of text information in the image to be proofread, and for the technicians to understand or analyze the image and text proofreading results, the image and text proofreading results are visually structured to obtain structured image and text proofreading results. For example, the image and text proofreading results output by the multimodal large model are a sentence description including the image content, text content, and whether the proofreading results of the image and text are consistent. The entire sentence description is relatively redundant, and it is time-consuming for the relevant technicians to view it directly. They need to select the required answers from the large sentence description. After the image and text proofreading results are structured, the large and relatively redundant sentence descriptions will be in the form of three paragraphs, such as image content, text content, and proofreading results of the image and text, for the relevant technicians to directly view the image and text proofreading results without manual classification and screening. This improves the work efficiency of the relevant technicians and improves the user experience.

[0060] An embodiment of the present application provides a method for image-text proofreading, in which the information to be proofread corresponding to the image to be proofread, the text content and the external supplementary information is determined, and the information to be proofread is input into a multimodal large language model to obtain an image-text proofreading result. The information to be proofread corresponding to the image to be proofread, the text content and the external supplementary information obtained based on the above three methods greatly improves the richness of the text information or image, and by determining the external supplementary information of the text content, a large amount of external knowledge is used to supplement the knowledge base of the model itself, thereby generating more accurate and specific external supplementary information. This method greatly improves the accuracy of image-text proofreading.

[0061] Figure 2 Shown is a flowchart of determining the text content corresponding to the image to be proofread in one embodiment of the present application. Figure 2 As shown, the process of determining the text content corresponding to the image to be proofread in the embodiment of the present application includes the following steps S21 to S23.

[0062] S21 . Extracting text information from the image to be proofread based on an OCR model.

[0063] S22: Detecting the external description text of the image to be proofread based on a text detector.

[0064] S23: combining the text information in the image to be proofread and the external description text of the image to be proofread, and determining the combined text information and the external description text as the text content corresponding to the image to be proofread.

[0065] For example, when combining the text information in the image to be proofread and the external description text of the image to be proofread, the combination methods include simple splicing, segmented splicing, feature-level fusion, text enhancement, and the like.

[0066] It should be noted that, in actual applications, a specific text combination method may be selected according to specific text features, and details thereof will not be given here.

[0067] It should be noted that the application scenario corresponding to the embodiment of the present application is that text information can be extracted from the image to be proofread, and external description text of the image to be proofread can be detected outside the image to be proofread, and the text information in the image to be proofread and the external description text of the image to be proofread can be combined.

[0068] An embodiment of the present application provides a method for determining the text content corresponding to an image to be proofread. In this method, text information in the image to be proofread is extracted, and external description text of the image to be proofread is detected; and the text information in the image to be proofread and the external description text of the image to be proofread are combined to obtain the text content corresponding to the image to be proofread. This can better understand the environment and background of the text to be proofread and improve the accuracy of text comprehension. In the process of combining the text information of the text to be proofread and the image to be proofread, inconsistent, erroneous or omitted information will be corrected in a timely manner to ensure the accuracy, completeness and consistency of the combined text. It can overcome the limitations of single information and generate richer, more accurate and more contextually meaningful text representations, thereby significantly improving the accuracy, efficiency and in-depth understanding ability of image-text proofreading; and provide accurate text information for image-text proofreading of multimodal large models.

[0069] Figure 3 Shown is a flow chart of obtaining the external supplementary information of the text content in one embodiment of the present application. Figure 3 As shown, the process of obtaining the external supplementary information of the text content in the embodiment of the present application includes the following steps S31 to S32.

[0070] S31, extracting the entity structure corresponding to the text information.

[0071] Exemplarily, the entity structure corresponding to the text information may be extracted based on an entity recognition model.

[0072] Exemplary entity recognition models include: hidden Markov model, bidirectional long short-term memory network and conditional random field, BERT model, etc.

[0073] It should be noted that the entity recognition models listed above are only for illustrative purposes. In actual applications, any other suitable entity recognition model can be selected according to specific application requirements, and this application does not impose any restrictions on this.

[0074] S32: Acquire external supplementary information corresponding to the entity structure based on an external database.

[0075] Exemplarily, the external database includes a vector database (ctor database), a search engine (search engine), etc.

[0076] In some embodiments, the external supplemental information includes an image URL.

[0077] For example, an entity structure is input into an external database, which then searches for relevant textual information based on the input entity structure. After the search is complete, the external database outputs the image URL and textual description corresponding to the entity structure, which is the external supplementary information. The image URL refers to the unique address on the internet for the image corresponding to the entity structure.

[0078] An embodiment of the present application provides a method for obtaining external supplementary information corresponding to the text information. In this method, by extracting the entity structure corresponding to the text to be proofread, the entity structure is input into the external database, and the external supplementary information output by the external database is obtained, thereby achieving the purpose of using a large amount of external knowledge to supplement the text information, expanding the dimension of the text information and the coverage of the text content, improving the integrity of the information, and generating more accurate and specific text information for multimodal large models to perform image and text proofreading.

[0079] Figure 4 Shown is a flow chart of determining the information to be proofread corresponding to the image to be proofread, the text content and the external supplementary information in one embodiment of the present application. Figure 4 As shown, another image-text proofreading method in one embodiment of the present application includes the following steps S41 to S44.

[0080] S41, determining first prompt information corresponding to the image to be proofread.

[0081] Specifically, the image to be proofread is transmitted to the first prompt box to generate the first prompt information.

[0082] S42: Determine second prompt information corresponding to the text content.

[0083] Specifically, the text content is transmitted to the second prompt box to generate the second prompt information.

[0084] S43: Determine third prompt information corresponding to the external supplementary information.

[0085] Specifically, the external supplementary information is transmitted to the third prompt box to generate the third prompt information.

[0086] S44: performing a combination operation on the first prompt information, the second prompt information, and the third prompt information to obtain the information to be proofread.

[0087] Specifically, before combining the first prompt information, the second prompt information, and the third prompt information, the first prompt information, the second prompt information, and the third prompt information are deleted and necessary text modifications are made to the text content to make the modified prompt information more concise.

[0088] For example, if it is just simple text proofreading, the first prompt information, the second prompt information and the third prompt information are transmitted to the information combination box, and the combination operation in the information combination box can be simple splicing or segmented splicing.

[0089] For example, if more complex context understanding or information association is required, the first prompt information, the second prompt information, and the third prompt information are transmitted to an information combination box, and the combination operation in the information combination box may be feature-level fusion or weighted combination.

[0090] Exemplarily, if interaction with the user or other interactions are required, the first prompt information, the second prompt information, and the third prompt information are transmitted to the information combination box, and the combination operation in the information combination box can be an interactive combination.

[0091] In actual applications, any appropriate combination method can be selected to combine the first prompt information, the second prompt information and the third prompt information according to specific application requirements, and this application does not impose any limitation on this.

[0092] In the method for determining the information to be proofread corresponding to the image to be proofread, the text content and the external supplementary information provided in an embodiment of the present application, by determining the first prompt information corresponding to the image to be proofread, determining the second prompt information corresponding to the text content, determining the third prompt information corresponding to the external supplementary information, and combining the first prompt information, the second prompt information and the third prompt information, the combined information to be proofread is obtained, which helps to understand and recognize the multimodal large model and improves the image and text proofreading efficiency and accuracy of the multimodal large model.

[0093] See also Figure 5 , Figure 5 The flowchart of another image and text proofreading method in one embodiment of the present application is shown. The specific implementation process includes the various steps in the above embodiment, which will not be described in detail in this application.

[0094] The protection scope of the image and text proofreading method of the embodiment of the present application is not limited to the execution order of the steps listed in this embodiment. All solutions implemented by adding, reducing, or replacing steps in the existing technology based on the principles of the present application are included in the protection scope of the present application.

[0095] The embodiment of the present application also provides a graphic proofreading device, which can implement the graphic proofreading method of the present application. However, the implementation device of the graphic proofreading method of the present application includes but is not limited to the structure of the graphic proofreading device listed in this embodiment. All structural deformations and replacements of the existing technology made according to the principles of the present application are included in the protection scope of the present application.

[0096] like Figure 6As shown, in one embodiment, the image and text proofreading device 60 of the present application includes an image acquisition module 61, a text content determination module 62, an external supplementary information acquisition module 63, a to-be-proofread information determination module 64, and an image and text proofreading result acquisition module 65.

[0097] The image acquisition module 61 is used to acquire the image to be proofread; the text content determination module 62 is used to determine the text content corresponding to the image to be proofread; the external supplementary information acquisition module 63 is used to acquire the external supplementary information of the text content; the information to be proofread determination module 64 is used to determine the information to be proofread corresponding to the image to be proofread, the text content and the external supplementary information; the image-text proofreading result acquisition module 65 is used to input the information to be proofread into the multimodal large language model to obtain the image-text proofreading result.

[0098] Among them, the structures and principles of the image acquisition module 61, text content determination module 62, external supplementary information acquisition module 63, information to be proofread determination module 64, and image and text proofreading result acquisition module 65 correspond one to one to the steps in the above-mentioned image and text proofreading method, so they will not be repeated here.

[0099] In the several embodiments provided in this application, it should be understood that the disclosed devices or methods can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of modules / units is only a logical function division. There may be other division methods in actual implementation. For example, multiple modules or units can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or modules or units, which can be electrical, mechanical or other forms.

[0100] The modules / units described as separate components may or may not be physically separate, and the components displayed as modules / units may or may not be physical modules, that is, they may be located in one place or distributed across multiple network elements. Some or all of the modules / units may be selected according to actual needs to achieve the purpose of the embodiments of the present application. For example, the functional modules / units in the various embodiments of the present application may be integrated into a processing module, or each module / unit may exist physically separately, or two or more modules / units may be integrated into a single module / unit.

[0101] Those skilled in the art should further appreciate that the units and steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the above description has generally described the composition and steps of each example according to function. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.

[0102] The embodiment of the present application also provides a computer-readable storage medium. Those skilled in the art will understand that all or part of the steps in the method for implementing the above embodiment can be completed by instructing the processor through a program, and the program can be stored in a computer-readable storage medium, and the storage medium is a non-transitory medium, such as a random access memory, a read-only memory, a flash memory, a hard disk, a solid-state hard disk, a magnetic tape, a floppy disk, an optical disc, and any combination thereof. The above storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or a data center that includes one or more available media. The available medium can be a magnetic medium (for example, a floppy disk, a hard disk, a tape), an optical medium (for example, a digital video disc (DVD)), or a semiconductor medium (for example, a solid-state disk (SSD)), etc.

[0103] An embodiment of the present application also provides an electronic device. Figure 7 The diagram shows the structure of an electronic device 70 in one embodiment of the present application. The image and text proofreading method provided in the embodiment of the present application can be applied to Figure 7 The electronic device 70 shown is, but not limited to, Figure 7 As shown, the electronic device 70 includes a processor 71 , a memory, a system bus 73 , and a network interface 75 , wherein the memory may include a non-volatile storage medium 72 and an internal memory 74 .

[0104] The non-volatile storage medium 72 can store an operating system and a computer program. The computer program includes program instructions, which, when executed, can cause the processor to execute any one of the image and text proofreading methods provided in the embodiments of the present application.

[0105] The processor is used to provide computing and control capabilities and support the operation of the entire computer equipment.

[0106] The internal memory 74 provides an environment for the operation of the computer program in the non-volatile storage medium. When the computer program is executed by the processor, the processor can execute any one of the image and text proofreading methods provided in the embodiments of the present application.

[0107] The network interface 75 is used for network communication, such as sending assigned tasks, etc. It will be understood by those skilled in the art that Figure 7 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.

[0108] It should be understood that the processor 71 may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor.

[0109] The electronic device 70 in the embodiment of the present application may include terminal devices such as tablet computers, laptop computers, mobile phones, supercomputers, smart wearable devices, etc., and can also be applied to databases, servers, and service response systems based on terminal artificial intelligence. The embodiment of the present application does not impose any restrictions on the specific type of electronic device.

[0110] For example, the electronic device can be a station (STAION, ST) in a WLAN, a cellular phone, a cordless phone, a Session Initiation Protocol (SIP) phone, a Wireless Local Loop (WLL) station, a handheld device with wireless communication capabilities, a computing device or other processing device connected to a wireless modem, a computer, a laptop computer, a handheld communication device, a handheld computing device, and / or other devices for communicating on a wireless system and a next-generation communication system, such as a mobile terminal in a 5G network, a mobile terminal in a future-evolved Public Land Mobile Network (PLMN), or a mobile terminal in a future-evolved Non-terrestrial Network (NTN).

[0111] As an example and not a limitation, when the electronic device is a wearable device, the wearable device can also be a general term for wearable devices that are intelligently designed and developed using wearable technology for daily wear, such as gloves and watches equipped with near-field communication modules. A wearable device is a portable device that is worn directly on the body or integrated into the user's clothes or accessories. It is attached to the user and performs payment, authentication and other operations through a pre-bound electronic card. Wearable devices are not just hardware devices, but also achieve powerful functions through software support, data interaction, and cloud interaction. Broadly speaking, wearable smart devices include those that are fully functional, large in size, and can achieve complete or partial functions without relying on smartphones, such as smart watches or smart glasses, as well as those that only focus on a certain type of application function and need to be used in conjunction with other devices such as smartphones, such as various types of smart watches and smart bracelets with display screens.

[0112] The descriptions of the processes or structures corresponding to the above figures have different emphases. For parts that are not described in detail in a certain process or structure, please refer to the relevant descriptions of other processes or structures.

[0113] The above embodiments are merely illustrative of the principles and effects of this application and are not intended to limit this application. Anyone skilled in the art may modify or alter the above embodiments without departing from the spirit and scope of this application. Therefore, all equivalent modifications or alterations made by one of ordinary skill in the art without departing from the spirit and technical concepts disclosed in this application shall be covered by the claims of this application.

Claims

1. A method for proofreading images and texts, characterized in that: The method comprises: Obtain the image to be proofread; Determining text content corresponding to the image to be proofread; Obtaining external supplementary information of the text content; Determining information to be proofread corresponding to the image to be proofread, the text content, and the external supplementary information; The information to be proofread is input into a multimodal large language model to obtain a picture-text proofreading result.

2. The image-text proofreading method according to claim 1, characterized in that: Determining the text content corresponding to the image to be proofread includes: The text information in the image to be proofread is extracted based on an OCR model, and the text information in the image to be proofread is determined as text content corresponding to the image to be proofread.

3. The image-text proofreading method according to claim 1, wherein: Determining the text content corresponding to the image to be proofread includes: The external description text of the image to be proofread is detected based on a text detector, and the external description text is determined as the text content corresponding to the proofread image.

4. The image-text proofreading method according to claim 1, characterized in that: Determining the text content corresponding to the image to be proofread includes: Extracting text information from the image to be proofread based on an OCR model; Detecting the external description text of the image to be proofread based on a text detector; The text information in the image to be proofread and the external description text of the image to be proofread are combined, and the combined text information and the external description text are determined as the text content corresponding to the image to be proofread.

5. The image-text proofreading method according to claim 1, characterized in that: The obtaining of the external supplementary information of the text content includes: Extracting the entity structure corresponding to the text content; External supplementary information corresponding to the entity structure is obtained based on an external database.

6. The image-text proofreading method according to claim 1, characterized in that: The determining the information to be proofread corresponding to the image to be proofread, the text content, and the external supplementary information includes: Determining first prompt information corresponding to the image to be proofread; Determining second prompt information corresponding to the text content; Determining third prompt information corresponding to the external supplementary information; The first prompt information, the second prompt information and the third prompt information are combined to obtain the information to be proofread.

7. The image-text proofreading method according to claim 1, characterized in that: The method further includes: performing post-processing operations on the image and text proofreading results.

8. A graphic and text proofreading device, characterized in that: include: An image acquisition module, used to acquire the image to be proofread; A text content determination module, configured to determine the text content corresponding to the image to be proofread; An external supplementary information acquisition module, used to acquire external supplementary information of the text content; a module for determining information to be proofread, configured to determine information to be proofread corresponding to the image to be proofread, the text content, and the external supplementary information; The image-text proofreading result acquisition module is used to input the information to be proofread into the multimodal large language model to obtain the image-text proofreading result.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the image-text proofreading method according to any one of claims 1 to 7 is implemented.

10. An electronic device, characterized in that: The electronic device comprises: a memory storing a computer program; A processor is communicatively connected to the memory, and executes the image-text proofreading method according to any one of claims 1 to 7 when calling the computer program.

Citation Information

Patent Citations

  • Picture processing method and system based on image-text combination and readable storage medium

    CN115186119A

  • Image content structured information extraction method and device, equipment and storage medium

    CN117437643A

  • Picture and text joint editing and correcting method and system for publications

    CN119250063A

  • Image-text quality evaluation method and device, electronic equipment and readable medium

    CN120217041A