Document checking method and device, equipment, medium and program product
By using a multimodal data extraction and comprehensive review module, the problems of low efficiency and insufficient comprehensiveness in document review in existing technologies are solved, achieving comprehensive document review and improving efficiency and accuracy.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- BEIJING FOUNDER ELECTRONICS CO LTD
- Filing Date
- 2025-12-11
- Publication Date
- 2026-05-01
AI Technical Summary
In existing technologies, document review is inefficient, prone to overlooking problems, cannot fully cover the inspection of text, images, and style, and lacks a collaborative processing mechanism.
By extracting multimodal data, including text, images, and style, and using the corresponding review modules for comprehensive review, the review results are obtained and output.
It enables multi-dimensional and comprehensive document review, improving review efficiency and accuracy, reducing labor costs, and ensuring manuscript quality.
Smart Images

Figure CN121963234A_ABST
Abstract
Description
Document review methods, devices, equipment, media and procedures products Technical Field
[0001] This application relates to the field of document review, and in particular to a document review method, apparatus, equipment, medium, and program product. Background Technology
[0002] In the manuscript editing process of publishing houses, journals, and research institutions, the documents submitted by authors typically contain text content, images, and complex formatting styles. These documents must undergo a rigorous review process before formal publication or academic dissemination to ensure the standardization, accuracy, and consistency of the content and formatting.
[0003] Currently, manuscript review mainly relies on manual operation. For example, editors need to check for typos, grammatical errors, sensitive words and other textual problems word by word. They also need to manually check whether the images contain illegal content (such as terrorism, explosives or advertising). They also need to manually compare the document's title hierarchy, font size, line spacing, indentation and other style formats.
[0004] However, this process has significant drawbacks: manual review is inefficient, prone to overlooking issues, and cannot comprehensively check images and style within the document, lacking a collaborative processing mechanism. Therefore, a document review method is urgently needed to improve the comprehensiveness of the review process. Summary of the Invention
[0005] This application provides a document review method, apparatus, device, medium, and program product to improve the comprehensiveness of document review.
[0006] In a first aspect, embodiments of this application provide a document review method, the method comprising:
[0007] Obtain the target document to be reviewed;
[0008] Multimodal data extraction is performed on the target document to obtain the data to be reviewed in the target document; the multimodal data includes: text, images, and at least one of the style of the target document;
[0009] The review module corresponding to the data to be reviewed is used to review the data to be reviewed, and the review result is obtained.
[0010] Output the review results.
[0011] In one possible implementation, the multimodal data includes: the image; the step of extracting multimodal data from the target document to obtain the data to be reviewed for the target document includes:
[0012] From the target document, extract images with pixel values greater than or equal to preset pixel values to obtain the target image in the target document;
[0013] Obtain the location description text of the target image; the location description text is used to describe the location of the target image in the target document;
[0014] Based on the location description text, the name of the target image is obtained;
[0015] Based on the target image and the naming, the data to be reviewed is obtained.
[0016] In one possible implementation, when the data to be reviewed includes image data, the review module corresponding to the data to be reviewed includes: an image content review module and a text review module. The step of reviewing the data to be reviewed through the review module corresponding to the data to be reviewed to obtain a review result includes:
[0017] In response to extracting text content from the target image, determining that the target image contains a first character, the text review module reviews the first character to obtain the text review result of the target image;
[0018] The image content review module obtains the image content review result of the target image.
[0019] The review result is obtained based on the text review result and the image content review result.
[0020] In one possible implementation, the multimodal data includes: the style and format; the step of extracting multimodal data from the target document to obtain the data to be reviewed for the target document includes:
[0021] From the target document, extract the text with the preset style to obtain the second text in the target document;
[0022] Based on the style of the second character, the style description text of the second character is obtained;
[0023] Based on the second text and the style description text, the data to be reviewed is obtained.
[0024] In one possible implementation, when the data to be reviewed includes style description text, the review module corresponding to the data to be reviewed includes: a style review module. The step of reviewing the data to be reviewed through the review module to obtain the review result includes:
[0025] The style review module compares the style description text with the preset style template information to obtain the style review result.
[0026] The review result is obtained based on the paragraph number of the second text in the target document and the style review result.
[0027] In one possible implementation, the multimodal data includes: the text; the step of extracting multimodal data from the target document to obtain the data to be reviewed for the target document includes:
[0028] Extract the third text from the target document, and the paragraph number in which the third text is located in the target document;
[0029] In response to the presence of a character of a preset type in the third text, the text with the preset type character is replaced with the preset character to obtain the updated third text; the preset type character includes at least one of the following: formula type character and inline image type character;
[0030] Based on the updated third text, the data to be reviewed is obtained.
[0031] Secondly, embodiments of this application provide a document review apparatus, the apparatus comprising:
[0032] The acquisition module is used to acquire the target document to be reviewed;
[0033] The processing module is used to extract multimodal data from the target document to obtain the data to be reviewed in the target document; the multimodal data includes: text, images, and at least one of the style of the target document;
[0034] The review module is used to review the data to be reviewed through the review module corresponding to the data to be reviewed, and obtain the review result;
[0035] The output module is used to output the review results.
[0036] Thirdly, embodiments of this application provide an electronic device, including: a memory and a processor;
[0037] The memory stores computer-executed instructions;
[0038] The processor executes computer execution instructions stored in the memory, causing the processor to perform the method described in any of the first aspects above.
[0039] Fourthly, embodiments of this application provide a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, are used to implement the method described in any of the first aspects above.
[0040] Fifthly, embodiments of this application provide a computer program product, including a computer program that, when executed by a processor, implements the method described in any of the first aspects above.
[0041] This application provides a document review method, apparatus, device, medium, and program product. Through multimodal data extraction, it can comprehensively obtain target document information and review it using the corresponding review module. This enables multi-dimensional and comprehensive document review. Compared with the one-sidedness of single review in the prior art, this application can improve the comprehensiveness of document review. Attached Figure Description
[0042] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.
[0043] Figure 1 is a schematic diagram of a document review process provided in an embodiment of this application;
[0044] Figure 2 is a schematic diagram of a process for obtaining image data to be reviewed according to an embodiment of this application;
[0045] Figure 3 is a schematic diagram of a process for obtaining image review results provided in an embodiment of this application;
[0046] Figure 4 is a schematic diagram of a process for obtaining image data to be reviewed according to an embodiment of this application;
[0047] Figure 5 is a schematic diagram of a process for obtaining image review results provided in an embodiment of this application;
[0048] Figure 6 is a schematic diagram of a process for obtaining image data to be reviewed according to an embodiment of this application;
[0049] Figure 7 is a schematic diagram of a text proofreading errata provided in an embodiment of this application;
[0050] Figure 8 is a schematic diagram of an image proofreading errata provided in an embodiment of this application;
[0051] Figure 9 is a schematic diagram of a proofreading errata table provided in an embodiment of this application;
[0052] Figure 10 is a structural schematic diagram of a document review device provided in this application;
[0053] Figure 11 is a schematic diagram of the structure of an electronic device provided in this application.
[0054] The accompanying drawings illustrate specific embodiments of this application, which will be described in more detail below. These drawings and descriptions are not intended to limit the scope of the concept in any way, but rather to illustrate the concept of this application to those skilled in the art through reference to particular embodiments. Detailed Implementation
[0055] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application.
[0056] In this application, the term "comprising" and its variations can refer to non-limiting inclusion; the term "or" and its variations can refer to "and / or". The terms "first", "second", etc., in this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. In this application, "multiple" refers to two or more. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, and B existing alone. The character " / " generally indicates that the preceding and following related objects have an "or" relationship.
[0057] The following is an explanation of the proper nouns used in the embodiments of this application.
[0058] Text proofreading: Performs a series of text content standardization checks on the document text content, including checking for common errors, punctuation marks, variant characters, misuse of simplified and traditional characters, grammar, sensitive words, key words, dates, and place names, and displays them through an errata table for easy and quick location.
[0059] Image review: This involves checking the text and content of images in the document. Text review uses Optical Character Recognition (OCR) to extract text from images, and then reviews the extracted text. Image content review identifies content related to terrorism, explosives, and advertising, ensuring the documents are properly formatted and that any suspicious issues are promptly identified and displayed in an errata sheet.
[0060] Style check: Check the font, font size, indentation, alignment, and line spacing of the first-level headings, second-level headings, third-level headings, body text, figure captions, and table footnotes in the document. Promptly identify any content that does not meet the style requirements and display it in the errata for quick location.
[0061] Multimodal document review: During document review, the text content and images within the document are reviewed simultaneously. Additionally, the document's style is checked, providing a comprehensive review and inspection from multiple dimensions, including text content, images, and style, to promptly identify any issues.
[0062] In the manuscript editing process of publishing houses, journals, and research institutions, the documents submitted by authors typically contain text content, images, and complex formatting styles. These documents must undergo a rigorous review process before formal publication or academic dissemination to ensure the standardization, accuracy, and consistency of the content and formatting.
[0063] In existing technologies, text content review in documents can be automated using natural language processing techniques (such as misspelling detection, grammar checking, and sensitive word filtering). These techniques are typically based on rule bases or machine learning models and can identify spelling errors, punctuation misuse, grammatical problems, etc., but they are limited to checking plain text content and cannot handle images or formatting in the document.
[0064] The checking of image content and format still relies on manual operation. For example, it is necessary to manually zoom in on the image to identify the text content (such as extracting it through an OCR tool and then checking it word by word), or check whether the image contains sensitive elements; the proofreading of the format requires manual comparison of the document style with the standard template paragraph by paragraph, and adjustment of parameters such as font, font size, and indentation item by item.
[0065] The aforementioned existing technologies have significant drawbacks: manual review is inefficient, prone to overlooking issues, and cannot comprehensively check images and formatting within the document. For example, authors may embed sensitive text or prohibited images in pictures, which are difficult for humans to identify one by one; formatting checks require editors to repeatedly switch between document style settings, which is time-consuming and labor-intensive.
[0066] Furthermore, the review of text, images, and style are typically handled by different systems or tools, lacking a collaborative processing mechanism, resulting in a fragmented review process and scattered results. Therefore, there is an urgent need for a document review method that can automate and integrate the processing of multimodal content (text, images, style) to improve review efficiency, reduce labor costs, and ensure comprehensive quality assurance of manuscripts.
[0067] The technical solution of this application and how the technical solution of this application solves the above-mentioned technical problems are described in detail below with specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments. The embodiments of this application will now be described with reference to the accompanying drawings.
[0068] It should be noted that the executing entity of this application can be any electronic device with processing capabilities, such as a user terminal or a server, for example, a computer.
[0069] Figure 1 is a schematic diagram of a document review process provided in an embodiment of this application. As shown in Figure 1, the method includes:
[0070] S101. Obtain the target document to be reviewed.
[0071] Optionally, the target document can be a document that requires review and editing, and it can be any type of file, such as academic papers, business reports, press releases, book manuscripts, etc. Optionally, the target document can be in the format of Word (a document format), Portable Document Format (PDF), etc.
[0072] Alternatively, electronic devices can retrieve the target document from a storage location (such as a local disk, a network server, or cloud storage) by calling an interface. For example, the target document to be reviewed can be retrieved by selecting the target document from a local folder on the computer.
[0073] S102. Perform multimodal data extraction on the target document to obtain the data to be reviewed in the target document; the multimodal data includes: text, images, and at least one of the style of the target document.
[0074] Optionally, multimodal data can contain data of various different types. The text in the target document can include body text, titles, footers, and other textual content. Images can include visual elements such as pictures, charts, and screenshots of formulas contained in the document. Style can be any one or more of the document's layout format, such as font size, paragraph spacing, header and footer settings, and heading level styles.
[0075] S103. The review module corresponding to the data to be reviewed is used to review the data to be reviewed and the review results are obtained.
[0076] Optionally, the review module can be a functional module or program component designed for different types of data to be reviewed. Different data types (such as text, images, and style) can correspond to different review modules. Each module has specific review rules and algorithms used to check whether the corresponding data has errors, conforms to specifications, etc. For example, the review module may include any one or more of the following: text review module, image content review module, style review module, etc.
[0077] The review results can be the conclusive information obtained after reviewing the data to be reviewed. It usually includes a detailed description of the problems in the data, such as the location of the error, the type of error, and any one or more possible correction suggestions.
[0078] Optionally, the electronic device can utilize a pre-designed review module that matches the type of data to be reviewed, and perform a comprehensive and detailed inspection and analysis of the data to be reviewed according to the rules and algorithms set within the module, in order to discover any potential problems.
[0079] S104. Output the review results.
[0080] Optionally, the electronic device can display or transmit the review results in a specific form, such as displaying a list of review results on the interactive interface of the electronic device, generating a report file containing the review results, sending the review results to a designated email address, or storing them in a specific database, or any one or more of these methods.
[0081] This application embodiment extracts multimodal data to comprehensively obtain target document information. By using the corresponding review module for review, it can achieve multi-dimensional and comprehensive document review. Compared with the one-sidedness of single review in the prior art, this application embodiment can improve the comprehensiveness of document review.
[0082] Taking multimodal data, including images, as an example, the following details how electronic devices extract images from target documents and how they review the images in the documents.
[0083] Figure 2 is a schematic diagram of a process for obtaining image data to be reviewed according to an embodiment of this application. As shown in Figure 2, the above-mentioned S102 includes:
[0084] S201. Extract images with pixel values greater than or equal to preset pixel values from the target document to obtain the target image in the target document.
[0085] Optionally, the preset pixel value can be a pre-defined pixel standard value used for filtering images. For example, the pixel value can be set to 80*80. By extracting images with pixel values greater than or equal to the preset pixel value, decorative images in the document that do not require review can be filtered out.
[0086] Optionally, the electronic device can use existing image recognition and processing technologies to scan and analyze the images in the target document, filter out images that meet the condition of having pixel values greater than or equal to preset pixel values, separate them from the document, form individual image data, and obtain the target image in the target document.
[0087] S202. Obtain the location description text of the target image; the location description text is used to describe the location of the target image in the target document.
[0088] Optionally, the location description text of the target image can be information describing the location of the target image in the target document in text form. For example, it may include one or more of the following: the page number where the image is located, its relative position on the page (such as the upper left corner of the page, the middle right, etc.), and its relative relationship with surrounding text or other elements.
[0089] Alternatively, the electronic device can analyze the document structure using existing document parsing algorithms, determine the page where the image is located and its coordinate position on the page, and convert it into a text description to obtain the location description text of the target image.
[0090] S203. Based on the location description text, obtain the name of the target image.
[0091] Optionally, the target image can be named with a specific meaning and identification function, facilitating image recognition, classification, and management. Electronic devices can use location description text as a basis, extracting key information from the location description text or generating a name for the target image according to certain rules.
[0092] S204. Based on the target image and its name, the data to be reviewed is obtained.
[0093] Optionally, the electronic device can combine the names of the target image in S201 and the target image in S203 to form the data to be reviewed.
[0094] This application embodiment can extract images that meet pixel requirements, accurately focus on key images, and locate the image's position in the document by naming it with location description text, which facilitates subsequent review and problem tracing, and improves the efficiency and accuracy of image review.
[0095] When the data to be reviewed includes image data, the review module corresponding to the data to be reviewed may include an image content review module and a text review module. Figure 3 is a flowchart illustrating an image review result provided in an embodiment of this application. As shown in Figure 3, the above-mentioned S103 may include:
[0096] S301. In response to extracting text content from the target image, determining that the target image contains the first text, and reviewing the first text through the text review module to obtain the text review result of the target image.
[0097] Optionally, the first text can be text content existing in the target image. For example, it can include any one or more of the title, body text, and annotation information in the image. The text review module can be a functional module or program component used to review text content. It can have various review rules and algorithms built-in to check for spelling errors, grammatical errors, semantic logic problems, formatting issues, etc., to ensure the accuracy and standardization of the text content.
[0098] The text review result of the target image can be the conclusive information obtained after the text review module reviews the first text. It can include a detailed description of the problems in the text, such as the location of the error, the type of error, and possible correction suggestions.
[0099] Optionally, the electronic device can analyze and process the target image using image recognition technology (such as OCR technology) to identify the text content contained in the image and identify it as the first text. The electronic device can use the preset rules and algorithms in the text review module to conduct a comprehensive and detailed inspection and analysis of the first text to find any possible errors or non-compliance with standards, and finally form a text review result for the first text in the target image.
[0100] S302. Obtain the image content review result of the target image through the image content review module.
[0101] Optionally, the image content review module can be a functional module or program component for reviewing image content. It can check issues such as image clarity, color accuracy, content integrity, and whether there are any illegal or inappropriate elements, such as whether the image belongs to the political, terrorist, explosive, or advertising categories, in order to ensure the quality of the image content.
[0102] The image content review results of the target image can be conclusive information obtained after reviewing the target image. It can include a detailed description of the problems existing in the image, such as blurred areas, color deviation, the type and location of illegal elements, and possible improvement suggestions.
[0103] S303. Based on the text review results and the image content review results, the review results are obtained.
[0104] Optionally, the electronic device can integrate and analyze the text review results in S301 and the image content review results in S302 to ultimately form a comprehensive review result for the target image.
[0105] This application embodiment combines a text review module and an image content review module to review images from both text and image content perspectives. This allows for a more comprehensive and accurate identification of problems in images, improving the quality of image review and ensuring the standardization of document images.
[0106] Taking multimodal data, including style and format, as an example, the following provides a detailed explanation of how electronic devices can extract style and format from target documents, and how to review and verify the style and format in documents.
[0107] Figure 4 is a schematic diagram of a process for obtaining image data to be reviewed according to an embodiment of this application. As shown in Figure 4, S102 above includes:
[0108] S401. Extract the text with the preset style from the target document to obtain the second text in the target document.
[0109] Optionally, the text in the preset style can be one or more pre-defined text formatting standards, covering aspects such as font, font size, color, line spacing, paragraph indentation, and heading level, used to filter text in the target document that meets specific formatting requirements. For example, the preset style may specify that the body text uses SimSun, size 12, and 1.5 line spacing; and the headings use Heiti, size 3, and center alignment.
[0110] The second text in the target document can be text content filtered from the target document according to a preset style, where the text conforms to the preset style. Optionally, the electronic device can use existing document parsing technology to scan and analyze the target document, and identify and separate the text content that meets the conditions according to the rules of the preset style. For example, by parsing the document's formatting marks, text with fonts, font sizes, etc., that meet preset values can be extracted to obtain the second text in the target document.
[0111] S402. Based on the style of the second character, obtain the style description text of the second character.
[0112] Optionally, the style description text can be information that describes the style of the second text in text form. It can clearly describe the specific characteristics of the second text in terms of font, font size, color, line spacing, paragraph format, etc., such as "the font is SimSun, the font size is 12pt, the line spacing is 1.5, and the first line of the paragraph is indented by 2 characters".
[0113] Optionally, the electronic device can analyze and organize the second text style to convert it into an accurate and clear text description, forming a style description text.
[0114] S403. Based on the second text and style description text, obtain the data to be reviewed.
[0115] Optionally, the electronic device can combine or encapsulate any one or more of the second text in S401 and the style description text in S402 to form data to be reviewed.
[0116] This application embodiment can extract text that conforms to a preset style and generate descriptive text, which can quickly and accurately obtain document style information, and combine the two to obtain the data to be reviewed, providing an accurate basis for subsequent style review and improving the efficiency and accuracy of style review.
[0117] When the data to be reviewed includes style description text, the review module corresponding to the data to be reviewed may include a style review module. Figure 5 is a flowchart illustrating an image review result provided by an embodiment of this application. As shown in Figure 5, the above-mentioned S103 may include:
[0118] S501. The style description text is compared with the preset style template information through the style review module to obtain the style review result.
[0119] Optionally, the style review module can be a functional module or program component used to review and proofread the style of a document. It can have built-in preset style template information, i.e., template data referencing style standards, capable of analyzing and judging the input style-related information to detect whether the document style conforms to the specifications. For example, the preset style template information can define the style specifications that the document should follow in different parts (such as headings, body text, and comments), such as the font, font size, and alignment of headings, and the line spacing and indentation of body text, or any one or more of these.
[0120] The style review result can be a conclusive information obtained by comparing the style description text with the preset style template information. It can include a judgment on whether the style conforms to the preset template, and a description of the specific differences when it does not conform, such as font mismatch, font size too large, or any one or more other issues.
[0121] Optionally, the electronic device can use the algorithms and rules in the style review module to compare and analyze each style feature in the style description text with the corresponding standard in the preset style template information to check whether the two are consistent.
[0122] S502. Based on the paragraph number of the second text in the target document and the style review results, the review results are obtained.
[0123] Optionally, the paragraph number can be used to identify the position of the second text in the target document, and the specific position of the second text in the document can be accurately located by the paragraph number.
[0124] Optionally, the electronic device can combine the paragraph number of the second text in the target document with the style review result in S501 to obtain the review result. This review result not only includes a judgment on whether the style review is passed, but also incorporates the position information of the second text in the document.
[0125] This application's embodiment can automatically compare styles using a style review module, quickly and accurately identifying differences between the style and the preset template. Combined with paragraph numbering, it can precisely pinpoint the problem location, facilitating modifications, improving the efficiency and accuracy of style review, and ensuring the standardization of document style.
[0126] Taking multimodal data, including text, as an example, the following is a detailed explanation of how electronic devices can extract text from target documents.
[0127] Figure 6 is a schematic diagram of a process for obtaining image data to be reviewed according to an embodiment of this application. As shown in Figure 6, the above-mentioned S102 includes:
[0128] S601. Extract the third text from the target document, and the paragraph number in which the third text is located in the target document.
[0129] Optionally, the third text can be text in the target document that needs to be reviewed, and the paragraph number in the target document where the third text is located is used to identify the paragraph position of the third text in the target document.
[0130] Optionally, electronic devices can employ document parsing technology to scan and analyze the target document, identify and extract third-party text content from the document, and determine the paragraph positions of these texts and assign them corresponding paragraph numbers. For example, by parsing the document's formatting marks or text structure, text can be extracted and its paragraph information recorded.
[0131] S602. In response to the presence of a character of a preset type in the third text, the text of the preset type character is replaced with the preset character to obtain the updated third text; the preset type character includes at least one of the following: formula type character and inline figure type character.
[0132] Optionally, formula type characters can represent formulas in subjects such as mathematics and physics, while inline figure type characters can be small pictures or graphic symbols that are mixed with text in the document. Preset characters can be simple text characters or placeholders with specific meanings, which can convert complex formulas and / or inline figures into a uniform and easy-to-process format.
[0133] The updated third text can be the new text content obtained by replacing the preset type characters existing in the third text with preset characters. After the replacement operation, the format and content of the third text may change, but the main information and semantics of the original text are still preserved.
[0134] Optionally, electronic devices can use character recognition technology or rules to determine whether the third text contains characters of a preset type. For example, regular expressions or specific parsing algorithms can be used to check whether formulas or inline image markers appear in the text.
[0135] Optionally, the electronic device can use string processing technology to replace preset type characters in the third text with preset characters according to predetermined rules. For example, the formula character can be replaced with "formula", and the inline image character can be replaced with "image".
[0136] S603. Based on the updated third character, obtain the data to be reviewed.
[0137] Optionally, the electronic device can organize and package the updated third-party text to obtain the data to be reviewed. Optionally, the data to be reviewed may also include other information related to the third-party text, such as paragraph number, document source, or any one or more of these.
[0138] This application embodiment can replace text with preset type characters with preset characters, unify complex content, simplify the review process, and improve the efficiency and accuracy of text review.
[0139] For example, the document review method of this application embodiment can be divided into three parts: content extraction, content review, and errata presentation. These three parts will be described in detail below.
[0140] Content extraction section: Extracting text content, images, and style elements from the document.
[0141] 1. Text Content Extraction: Text content is extracted from the document using methods such as the Software Development Kit (SDK), preserving complete paragraphs and their order during extraction. Since formulas and inline figures in the text are unrecognizable during proofreading, they need to be replaced with special characters. The extracted content is formatted as JSON, and the type of proofreading performed is indicated after extraction, such as: common word check, sensitive word check, and key word check. An example is shown below:
[0142] [
[0143] {
[0144] "PageIndex": 0, can represent the page number.
[0145] "ParagraphIndex": 0, can represent a paragraph number, and this paragraph number is unique within the same page.
[0146] "Text": "
Example
[0147] }
[0148] ]
[0149] 2. Image Content Extraction: Extract the image content from the document and name it with the text describing the image's location in the document. In addition, decorative images with a pixel value less than 80*80 should be filtered out during image extraction.
[0150] 3. Style Information Extraction: Since style information is contained in the corresponding document, style checking in the document is based on the original document. Therefore, style information needs to be extracted from the corresponding document file, and the automatic numbering and other field information in the document needs to be converted into normal text information. Inline images and formulas are replaced with special characters to enable more accurate checking.
[0151] Content review section: After collecting information from the document, the server reviews the text, images, and style, and returns the corresponding review results to the client for display.
[0152] 1. Text Proofreading: Based on the extracted JSON (one format) file content and the type of text to be proofread, a text proofreading engine is invoked. This engine tracks the selected proofreading type and calls error-prone word checkers, grammar checkers, place name checkers, and sensitive word checkers, among others. The proofreading results undergo a series of post-processing operations, primarily filtering the whitelist and merging the results to form a unified proofreading JSON (one format) result. It should be noted that the whitelist can include proper nouns. An example is shown below:
[0153] {
[0154] "status": 0, used to indicate the review status.
[0155] "message": "Call successful!", used to indicate the review status.
[0156] "data": {
[0157] "charCount": "49", used to represent the total number of words reviewed.
[0158] "detail":
[0159] {
[0160] "inspectType": "Error-prone word check", used to represent the result type
[0161] "content": "Modeling clay", used to represent the suspicious word
[0162] "lookup": "Simulation", used to represent the proofreading suggestion word
[0163] "offset": "13", used to represent the paragraph offset position of the current result
[0164] "paragraphIndex": "1", used to represent the paragraph position where the current result is located.
[0165] "detail": "Artificial intelligence is a new technical science that studies and develops theories, methods, technologies, and application systems for <em>Clay< / em> , extending, and expanding human intelligence", used to represent the context of the proofreading result.
[0166] "pageIndex": "0, used to represent the page number where the proofreading suspicious word is located
[0167] }
[0168]
[0169] }
[0170] }
[0171] 2. Picture proofreading: When proofreading pictures in a document, the picture engine is called to send the document identifier (Identifier, ID), picture name, and picture to the picture proofreading engine. The picture proofreading engine first extracts the OCR text content from the picture, and the extracted text content forms the format of text proofreading. Then, the text proofreading engine is called to proofread for spelling mistakes, sensitive words, etc.
[0172] At the same time, the picture content proofreading engine is called to analyze the picture to check whether it belongs to pictures of political, terrorist, explosive, or advertising types. After obtaining the OCR text proofreading result and the picture content proofreading result, the picture proofreading engine merges the results and returns a unified picture proofreading result. The example is as follows:
[0173] {
[0174] "status": 0,
[0175] "message": "Call successful!",
[0176] "data": {
[0177] "charCount": "49",
[0178] "detail": [
[0179] {
[0180] "docId": "aa1ed8a66e764df5ad08455e566c4e11", is used to represent the document ID.
[0181] "fileName": "1.jpg", is used to represent the document name.
[0182] "inspectType": "Error-prone word check",
[0183] "content": "mold clay",
[0184] "lookup": "simulation",
[0185] "offset": "13",
[0186] "paragraphIndex": "1",
[0187] "detail": "Artificial intelligence is the research and development of applications..." <em> Clay< / em> It is a new technical science that extends and expands the theories, methods, technologies, and application systems of human intelligence.
[0188] "pageIndex": "0",
[0189] },
[0190] {
[0191] "docId": "aa1ed8a66e764df5ad08455e566c4e11",
[0192] "fileName": "1.jpg",
[0193] "inspectType": "Image content inspection",
[0194] "content": "advertisement"
[0195] "lookup": "Please check",
[0196] "pageIndex": "0",
[0197] }
[0198] ]
[0199] }
[0200] }
[0201] 3. Style Review: Before style review, you need to define the correct style template information. Only when style information such as title, body text, figure captions, and table captions is defined can the corresponding comparison and check be performed in the review template. The defined template is in Extensible Markup Language (xml) format. During the review, the template xml file will be submitted to the style review engine along with the document.
[0202] The style review engine iterates through the paragraphs in the document, extracting structural information (such as headings and body text, and using regular expressions and intelligent recognition to determine if the extracted content contains figure captions or table notes) and style information (such as font, font size, reduction distance, alignment, and line spacing). This information is compared to the style specified in the template. For each discrepancy, the engine records the corresponding paragraph number (starting from 1, based on the paragraph in the document) and the reason for the discrepancy. Reduction and line spacing have a 5% tolerance margin; exceeding this margin results in a non-compliance. The engine records the corresponding paragraph number and error type information. An example is shown below:
[0203] <root>
[0204] <!-- Heading style check parameters (associated with Word style ID) -->
[0205] <param name="heading" value="1">
[0206] <!-- Main heading (associated with Heading1 style) -->
[0207] <param name="heading-level-1" value="1" style-id="Heading1">
[0208] <param name="font-name" value="1" expected="宋体">
[0209] <param name="font-size" value="1" expected="16">
[0210] <param name="alignment" value="1" expected="Center"> <舍]
[0211] <param name="space-before" value="1" expected="12">
[0212] <param name="space-after" value="1" expected="12">
[0213] <param name="line-spacing" value="1" expected="1.5">
[0214] <param name="numbering-format" value="1" expected="1.">
[0215] <00004’47>
[0216]
[0217] <!-- Body style check parameters (associated with Normal style) -->
[0218] <param name="body-text" value="1">
[0219] <param name="normal-style" value="1" style-id="Normal">
[0220] <param name="font-name" value="1" expected="宋体">
[0221] <param name="font-size" value="1" expected="12">
[0222] <param name="alignment" value="1" expected="Justify">
[0223] <param name="space-before" value="1" expected="0">
[0224] <param name="space-after" value="1" expected="0">
[0225] <param name="line-spacing" value="1" expected="1.5">
[0226] <param name="first-line-indent" value="1" expected="20">
[0227]
[0228] It should be noted that there may be some inaccuracies in the translation due to the lack of clear semantic context for some of the original text. You can provide more detailed information if needed for a more accurate translation.
[0229] <!-- Image style check parameters (related to image styles) -->
[0230] <param name="images" value="1">
[0231] <!-- Figure caption -->
[0232] <param name="image-legend" value="1" style-id="Legend">
[0233] <param name="font-name" value="1" expected="宋体">
[0234] <param name="font-size" value="1" expected="10">
[0235] <param name="alignment" value="1" expected="Justify">
[0236] <param name="line-spacing" value="1" expected="1.2">
[0237] <param name="space-before" value="1" expected="6">
[0238] <param name="space-after" value="1" expected="12">
[0239]
[0240]
[0241] <!-- Table style check parameters -->
[0242] <param name="tables" value="1">
[0243] <!-- Table caption -->
[0244] <param name="table-note" value="1" style-id="TableNote">
[0245] <param name="font-name" value="1" expected="宋体">
[0246] <param name="font-size" value="1" expected="10">
[0247] <param name="alignment" value="1" expected="Left">
[0248] <param name="line-spacing" value="1" expected="1.2">
[0249] <param name="space-before" value="1" expected="6">
[0250] <param name="space-after" value="1" expected="12">
[0251] <param name="indent" value="1" expected="20">
[0252]
[0253] <!-- Table cell text -->
[0254] <param name="table-cell-text" value="1">
[0255] <param name="font-name" value="1" expected="宋体">
[0256] <param name="font-size" value="1" expected="10">
[0257] <param name="line-spacing" value="1" expected="1.2">
[0258]
[0259]
[0260] < / root>
[0261] The document review results presentation section allows users to select the review type during the review process. Users can choose one or more of the following: text review, image review, and style review. After the review is complete, the corresponding review results will be displayed separately. Specifically, this includes:
[0262] 1. Obtaining Results: The review client obtains the review results according to the submitted review type, and then parses and displays the results in sequence.
[0263] 2. Text Proofreading Errata: The text proofreading results are displayed in the form of an errata. Clicking on the errata will locate the corresponding suspicious document.
[0264] 3. Image Review and Errata: The errata list shows the problematic images in an album format. Clicking on an image will take you to its location in the document, and double-clicking an image will open the image review results page to view the specific error details.
[0265] 4. Style Review and Errata: The style review results are presented in the Errata in the form of a list. Click to locate the corresponding paragraph and the style details of the corresponding paragraph will be displayed in the list.
[0266] For example, Figure 7 is a schematic diagram of a text errata sheet provided in an embodiment of this application. As shown in Figure 7, it can display the errata type, errata result, and suggested content. Figure 8 is a schematic diagram of an image errata sheet provided in an embodiment of this application. As shown in Figure 8, double-clicking the errata sheet image opens a details page. Figure 9 is a schematic diagram of a style errata sheet provided in an embodiment of this application. As shown in Figure 9, the style errata sheet can include the errata type and errata result.
[0267] The purpose of this application is to address the limitations of current manuscript format verification processes used by publishing houses, journals, and research institutions. These methods can only perform textual proofreading and cannot handle multimodal verification simultaneously. This application aims to provide an efficient and comprehensive multimodal document verification method. Given that existing verification methods can only detect textual content (such as typos, punctuation errors, grammatical problems, and sensitive words), they cannot cover equally important image content and style / formatting. Furthermore, the latter two types of checks rely on manual labor, leading to low efficiency and the risk of omissions. This application aims to automate the collaborative verification of bimodal content (text, images, titles, and body text), including line spacing, indentation, font, and font size. This reduces manual intervention, improves overall manuscript verification efficiency, lowers the probability of missing issues, and helps relevant institutions identify problems earlier and more comprehensively during the manuscript receiving and editing process, ensuring manuscript quality and optimizing manuscript processing.
[0268] The above are the method embodiments provided in this application. The apparatus provided in this application will be described below.
[0269] Figure 10 is a structural schematic diagram of a document review device provided in this application. As shown in Figure 10, the document review device 400 provided in this embodiment includes:
[0270] Module 701 is used to acquire the target document to be reviewed;
[0271] The processing module 702 is used to extract multimodal data from the target document to obtain the data to be reviewed in the target document; the multimodal data includes: text, images, and at least one of the style of the target document;
[0272] The review module 703 is used to review the data to be reviewed by the review module corresponding to the data to be reviewed, and obtain the review result.
[0273] Output module 704 is used to output the review results.
[0274] Optionally, the multimodal data includes: images; processing module 702, specifically used to: extract images with pixel values greater than or equal to preset pixel values from the target document to obtain the target image in the target document; obtain location description text for the target image; the location description text is used to describe the location of the target image in the target document; obtain a name for the target image based on the location description text; and obtain the data to be reviewed based on the target image and the name.
[0275] For example, when the data to be reviewed includes image data, the review modules corresponding to the data to be reviewed include: an image content review module and a text review module. Review module 703 is specifically used to, in response to extracting text content from the target image, determine that the target image contains a first character, review the first character through the text review module, and obtain the text review result of the target image. The image content review module then obtains the image content review result of the target image. Based on the text review result and the image content review result, a final review result is obtained.
[0276] Optionally, the multimodal data includes: style patterns. Processing module 702 is specifically used to: extract text with a preset style pattern from the target document to obtain the second text in the target document; obtain a style pattern description text for the second text based on the style pattern of the second text; and obtain the data to be reviewed based on the second text and the style pattern description text.
[0277] For example, when the data to be reviewed includes style description text, the review modules corresponding to the data to be reviewed include: a style review module, and review module 703, which is specifically used to compare the style description text with the preset style template information through the style review module to obtain the style review result. Based on the paragraph number of the second text in the target document and the style review result, the review result is obtained.
[0278] Optionally, the multimodal data includes: text. Processing module 702 is specifically used to extract the third text from the target document, and the paragraph number where the third text is located in the target document. In response to the presence of text with a preset type of character within the third text, the text with the preset type of character is replaced with the preset character to obtain the updated third text; the preset type of character includes at least one of: formula type character and inline figure type character. Based on the updated third text, the data to be reviewed is obtained.
[0279] The document review device provided in this embodiment can execute the methods provided in any of the above method embodiments. Its implementation principle and technical effect are similar, and will not be described in detail here.
[0280] Figure 11 is a schematic diagram of the structure of an electronic device provided in this application. As shown in Figure 11, the electronic device 500 provided in this embodiment includes at least one processor 501 and a memory 502. Optionally, the device 500 further includes a communication component 503. The processor 501, the memory 502, and the communication component 503 are connected via a bus 504.
[0281] In a specific implementation, at least one processor 501 executes computer execution instructions stored in memory 502, causing at least one processor 501 to perform the above-described method.
[0282] The specific implementation process of processor 501 can be found in the above method embodiments, and its implementation principle and technical effect are similar. It will not be repeated here.
[0283] In the above embodiments, it should be understood that the processor can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), etc. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the method disclosed in this invention can be directly implemented by a hardware processor, or implemented by a combination of hardware and software modules within the processor.
[0284] The memory may include random access memory (RAM) and may also include non-volatile memory (NVM), such as at least one disk storage device.
[0285] The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. Buses can be categorized as address buses, data buses, control buses, etc. For ease of illustration, the buses shown in the accompanying drawings are not limited to a single bus or a single type of bus.
[0286] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the above-described method.
[0287] This application also provides a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, implement the above-described method.
[0288] The aforementioned readable storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk. The readable storage medium can be any available medium accessible to a general-purpose or special-purpose computer.
[0289] An exemplary readable storage medium is coupled to a processor, enabling the processor to read information from and write information to the readable storage medium. Of course, the readable storage medium can also be a component of the processor. The processor and the readable storage medium can reside in an Application Specific Integrated Circuit (ASIC). Alternatively, the processor and the readable storage medium can exist as discrete components in the device.
[0290] The division of units is merely a logical functional division; in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be indirect coupling or communication connection through some interfaces, devices, or units, and may be electrical, mechanical, or other forms.
[0291] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0292] In addition, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.
[0293] If a function is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0294] Those skilled in the art will understand that all or part of the steps of the above-described method embodiments can be implemented by hardware related to program instructions. The aforementioned program can be stored in a computer-readable storage medium. When executed, the program performs the steps of the above-described method embodiments; and the aforementioned storage medium includes various media capable of storing program code, such as ROM, RAM, magnetic disks, or optical disks.
[0295] Finally, it should be noted that other embodiments of this application will readily conceive of by those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of this application that follow the general principles of this application and include common knowledge or customary techniques in the art not disclosed herein, and is not limited to the precise structures described above and shown in the accompanying drawings, and various modifications and changes may be made without departing from its scope.
Claims
1. A document review method, characterized in that, The method includes: acquiring a target document to be reviewed; extracting multimodal data from the target document to obtain review data of the target document; the multimodal data includes: text, images, and at least one of the style of the target document; reviewing the review data through the review module corresponding to the review data to obtain a review result; and outputting the review result.
2. The method according to claim 1, characterized in that, The multimodal data includes the image. The process of extracting multimodal data from the target document to obtain the data to be reviewed for the target document includes: extracting images with pixel values greater than or equal to preset pixel values from the target document to obtain target images in the target document; obtaining location description text for the target image; the location description text describing the location of the target image in the target document; obtaining a name for the target image based on the location description text; and obtaining the data to be reviewed based on the target image and the name.
3. The method according to claim 2, characterized in that, When the data to be reviewed includes image data, the review module corresponding to the data to be reviewed includes: an image content review module and a text review module. The step of reviewing the data to be reviewed through the review module corresponding to the data to be reviewed to obtain a review result includes: in response to extracting text content from the target image, determining that the target image includes a first character, reviewing the first character through the text review module to obtain a text review result for the target image; obtaining an image content review result for the target image through the image content review module; and obtaining the review result based on the text review result and the image content review result.
4. The method according to any one of claims 1-3, characterized in that, The multimodal data includes: the style; the multimodal data extraction of the target document to obtain the data to be reviewed of the target document includes: extracting text with a preset style from the target document to obtain the second text in the target document; obtaining the style description text of the second text based on the style of the second text; and obtaining the data to be reviewed based on the second text and the style description text.
5. The method according to claim 4, characterized in that, When the data to be reviewed includes style description text, the review module corresponding to the data to be reviewed includes: a style review module. The step of reviewing the data to be reviewed through the review module to obtain a review result includes: comparing the style description text with preset style template information through the style review module to obtain a style review result; and obtaining the review result based on the paragraph number of the second text in the target document and the style review result.
6. The method according to any one of claims 1-3, characterized in that, The multimodal data includes: the text; the multimodal data extraction of the target document to obtain the data to be reviewed of the target document includes: extracting the third text in the target document, and the paragraph number in which the third text is located in the target document; in response to the presence of a character of a preset type in the third text, replacing the character of the preset type with a preset character to obtain the updated third text; the preset type character includes at least one of: formula type character and inline image type character; and obtaining the data to be reviewed based on the updated third text.
7. A document review device, characterized in that, The apparatus includes: an acquisition module for acquiring a target document to be reviewed; a processing module for performing multimodal data extraction on the target document to obtain the review data of the target document; the multimodal data includes: text, images, and at least one of the style of the target document; a review module for reviewing the review data through the review module corresponding to the review data to obtain a review result; and an output module for outputting the review result.
8. An electronic device, characterized in that, include: Memory, processor; The memory stores computer-executable instructions; the processor executes the computer-executable instructions stored in the memory, causing the processor to perform the method as described in any one of claims 1-6.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions, which, when executed by a processor, are used to implement the method as described in any one of claims 1-6.
10. A computer program product, characterized in that, Includes a computer program that, when executed by a processor, implements the method described in any one of claims 1-6.