Document detection method and apparatus, electronic device, and medium

CN122549408APending Publication Date: 2026-08-11EMPYREAN TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-29
Publication Date
2026-08-11

AI Technical Summary

Technical Problem

然而,相关技术中的检测策略单一且固定,在检测时,很容易产生大量错误误报,或者遗漏错误,从而导致检测准确度差

Benefits of technology

[0008] The technical solutions provided by the embodiments of this disclosure can include the following beneficial effects: This disclosure can determine the target detection method corresponding to the document based on the document's business type and document format, and obtain the document's detection results. Thus, this disclosure can perform more comprehensive and accurate orthogonal detection on documents from multiple dimensions, thereby improving the accuracy of document detection, accurately locating errors in the document, facilitating error repair by technicians, and improving the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122549408A_ABST
    Figure CN122549408A_ABST
Patent Text Reader

Abstract

This disclosure relates to a document detection method, apparatus, electronic device, and storage medium. The method includes identifying at least one target document; determining the target business type and target document format for each target document; determining a target detection method corresponding to each target document based on the target business type and target document format, wherein the target detection method includes at least one of a first detection method, a second detection method, and a third detection method; and detecting at least one target document based on the target detection method corresponding to each target document to obtain a detection result for each target document. This disclosure performs more comprehensive and accurate orthogonal detection of documents from multiple dimensions, thereby improving the accuracy of document detection, accurately locating errors in documents, facilitating error repair by technicians, and improving the user experience.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of document detection, specifically to a document detection method, apparatus, electronic device, and medium. Background Technology

[0002] In related technologies, documents can be inspected to identify errors and facilitate document management. However, the inspection strategies in these technologies are singular and fixed, easily leading to numerous false alarms or missed errors, resulting in poor inspection accuracy. Summary of the Invention

[0003] To overcome the problems existing in related technologies, this disclosure provides a document detection method, apparatus, electronic device, and medium.

[0004] According to a first aspect of the present disclosure, a document detection method is provided, comprising: Identify at least one target document, and determine the target business type and target document format for each target document; Based on the target business type and target document format of each target document, a target detection method is determined for each target document. The target detection method includes at least one of a first detection method, a second detection method, and a third detection method. The target detection method is used to detect documents of the target business type and the target document format. The first detection method is used to detect language rule-related errors in the target document. The second detection method is used to detect errors related to the target business type in the target document. The third detection method is used to detect errors related to the target document format in the target document. Based on the target detection method corresponding to each target document, the at least one target document is detected, and the detection result of each target document is obtained.

[0005] According to a second aspect of the present disclosure, a document detection apparatus is provided, comprising: The scheduling agent module is configured to identify at least one target document and determine the target business type and target document format for each target document. The detection skills module is configured to determine a target detection method for each target document based on the target business type and the target document format. The target detection method includes at least one of a first detection method, a second detection method, and a third detection method. The target detection method is used to detect documents of the target business type and the target document format. The first detection method is used to detect language rule-related errors in the target document, the second detection method is used to detect target business-related errors in the target document, and the third detection method is used to detect document format-related errors in the target document. The processing module is configured to detect the at least one target document based on the target detection method corresponding to each target document, and obtain the detection result of each target document.

[0006] According to a third aspect of the present disclosure, an electronic device is provided, including a memory and a processor, the memory storing a computer program, the processor executing the computer program to implement the steps of the method described in the first aspect.

[0007] According to a fourth aspect of the present disclosure, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements the steps of the method described in the first aspect.

[0008] The technical solutions provided by the embodiments of this disclosure can include the following beneficial effects: This disclosure can determine the target detection method corresponding to the document based on the document's business type and document format, and obtain the document's detection results. Thus, this disclosure can perform more comprehensive and accurate orthogonal detection on documents from multiple dimensions, thereby improving the accuracy of document detection, accurately locating errors in the document, facilitating error repair by technicians, and improving the user experience.

[0009] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this disclosure. Attached Figure Description

[0010] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with the invention and, together with the description, serve to explain the principles of the invention.

[0011] Figure 1 This is a flowchart illustrating a document detection method according to an exemplary embodiment.

[0012] Figure 2 This is a flowchart illustrating a document detection method according to an exemplary embodiment.

[0013] Figure 3 This is a schematic diagram illustrating a document detection method according to an exemplary embodiment.

[0014] Figure 4 This is a block diagram illustrating a document detection device according to an exemplary embodiment.

[0015] Figure 5 This is a block diagram of an electronic device according to an exemplary embodiment. Detailed Implementation

[0016] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numerals in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present invention. Rather, they are merely examples of apparatuses and methods consistent with some aspects of the invention as detailed in the appended claims.

[0017] In related technologies, documents can be inspected to identify errors and facilitate document management. However, the inspection strategies in these technologies are singular and fixed, easily leading to numerous false alarms or missed errors, resulting in poor inspection accuracy.

[0018] To address the aforementioned issues, this disclosure provides a document detection method. Based on the document's business type and format, this disclosure determines the corresponding target detection method and obtains the document's detection results. Thus, this disclosure allows for more comprehensive and accurate orthogonal detection of documents from multiple dimensions, thereby improving the accuracy of document detection, precisely locating errors within the document, facilitating error repair by technical personnel, and enhancing the user experience.

[0019] This disclosure provides an exemplary embodiment of a document detection method that can be applied to electronic devices, specifically smart devices such as mobile phones, tablets, laptops, smart robots, and smart wearable devices. Furthermore, the electronic device also includes various hardware resources and energy storage devices that provide power for the operation of these hardware resources.

[0020] like Figure 1 As shown in this embodiment, the document detection method includes: Step S101: Identify at least one target document and determine the target business type and target document format for each target document.

[0021] In some embodiments, electronic devices can automatically identify the target business type and target document format of a document.

[0022] In some embodiments, the target business type can be used to indicate the business scenario to which the target document belongs, the purpose of the target document, and the business corresponding to the target document, and the target document format can be used to indicate the file format in which the target document is stored and presented.

[0023] Optionally, the target business type can indicate the business corresponding to the target document, such as user manual, release instructions, auxiliary tutorials, etc. The target document format can include Portable Document Format (PDF), doc, docx, HyperText Markup Language (HTML), Extensible Markup Language (XML), LaTeX.

[0024] In some embodiments, each target document may correspond to a different target business type and a different target document format. For example, the business type of one target document may be a user manual, and the target document format may be PDF. The business type of another target document may be a release notice, and the target document format may be HTML.

[0025] Step S102: Based on the target business type and target document format of each target document, determine the target detection method corresponding to each target document. The target detection method includes at least one of the first detection method, the second detection method, and the third detection method.

[0026] Among them, the target detection method is used to detect documents of target business type and target document format. The first detection method is used to detect errors related to language rules in the target document. The second detection method is used to detect errors related to target business type in the target document. The third detection method is used to detect errors related to target document format in the target document.

[0027] In one example, the object detection method can be expressed by the following formula: RuleSet=GeneralRules∪TypeRules(document_type)∪FormatRules(file_format).

[0028] In this framework, RuleSet represents the target detection method, GeneralRules represents the first detection method, TypeRules represents the second detection method, document_type represents the business type, FormatRules represents the third detection method, and file_format represents the document format. It's understandable that the second detection method is related to the business type, and the third detection method is related to the document format.

[0029] In some embodiments, a second detection method can be determined based on the target business type, and a third detection method can be determined based on the target document format. It should be noted that errors related to different target business types may differ, therefore the second detection methods corresponding to different target business types will differ; similarly, errors related to different target document formats may differ, therefore the third detection methods corresponding to different target document formats will differ.

[0030] In some embodiments, a first association is established between the target business type and the second detection method, and a second association is established between the target document format and the third detection method. The second detection method and the third detection method can be determined based on the above association. The specific process will not be described in detail here.

[0031] The three detection methods are explained below.

[0032] In some embodiments, language rules may refer to writing standards, and errors related to language rules in the target document may include typos, grammatical errors, punctuation marks, and other basic formatting errors in the target document.

[0033] In some embodiments, the second detection method is used to detect errors in at least one of the following in the target document: terminology, product name, function name, version identifier, parameter name, completeness of operation steps, logic of document content, and correlation of operation steps.

[0034] In other words, errors related to the target business type can include at least one of the following: terminology, product name, function name, version identifier, parameter name, completeness of operation steps, logical structure of document content, and relevance of operation steps in the target document. It should be noted that the content of errors related to different target business types may be the same or different. For example, release notes focus on detecting version change information and compatibility descriptions; for release notes, the version identifier and logical structure of the document content can be checked. User manuals focus on detecting operation steps and technical terminology; for user manuals, the completeness, relevance, and terminology of the operation steps can be checked.

[0035] Among these measures, the detection of terminology in the target document can include checking whether the terminology in the target document is standardized, thereby unifying the professional expression of the document, avoiding the misuse, miswriting, and confusion of alternative names, ensuring that the document is expressed in a standardized and professional manner, and reducing reading comprehension bias and communication ambiguity.

[0036] Detecting product names in a target document can include checking for spelling errors and / or ensuring consistency between product names within the document's context. This eliminates issues such as misspelled or inconsistent product names, guaranteeing that the document matches the actual product information and improving document accuracy and information recognizability.

[0037] Detecting function names in the target document can include checking for spelling errors and / or ensuring consistency between function names within the target document's context. This ensures that function naming aligns with the product definition, preventing misaligned or incorrect function descriptions, allowing users to accurately identify product functions, and avoiding operational and misunderstanding errors.

[0038] Detecting version identifiers in a target document can include checking for spelling errors and / or ensuring consistency between version identifiers within the target document's context. This allows for verification of correct version number spelling, labeling, and matching relationships, distinguishing between different document versions, preventing version confusion and mislabeling, and ensuring the effectiveness of document version traceability and management.

[0039] The detection of parameter names in the target document can include checking for spelling errors and / or ensuring consistency between parameter names within the target document's context. This guarantees that parameter names are standardized, consistent, and error-free, avoiding ambiguous or incorrect parameter designations, ensuring accurate transmission of technical parameters, and supporting the normal operation of equipment debugging, use, and maintenance.

[0040] Checking the completeness of operation steps in a target document can include detecting whether any steps are missing or duplicated. This helps identify issues such as missing or duplicated steps, ensuring a complete and closed-loop workflow. This allows users to smoothly complete the entire process based on the document, reducing operational delays and failures.

[0041] The logic for detecting document content in a target document can include checking for logical conflicts, such as discrepancies in descriptions before and after changes (the document content can record added, modified, and adjusted features in this version update). This can correct inconsistencies, logical inconsistencies, and disjointed expressions in the document content, making the target document clearer and improving its readability and rigor.

[0042] Detecting the relevance of operation steps in a target document can include checking whether the order and connection between these steps are reasonable. This avoids errors in the order of operation steps (such as reversed steps or disjointed connections), ensures that the operation steps conform to actual usage logic, and improves the smoothness and security of the operation.

[0043] Of course, it is understood that the content detected by the second detection method described above is only an example. The content detected by the second detection method can be added or deleted based on the actual situation in order to detect the target document from more dimensions.

[0044] In some embodiments, the third detection method is used to detect errors in at least one of the following in the target document: instruction syntax, formula syntax, tags, links, anchors, table fields, and reference relationships.

[0045] It should be noted that errors related to different target document formats may be the same or different. For example, LaTeX documents focus on detecting tags and formula syntax, while XML documents focus on detecting tags and reference relationships.

[0046] Among them, detecting the syntax of instructions in the target document can include checking whether the syntax of instructions in the target document is compliant and whether the calls between instructions are compliant. In this way, it can avoid document parsing, rendering and opening failures caused by abnormal instruction syntax, and ensure that the document can be read and formatted normally.

[0047] Detecting the formula syntax in a target document can include checking whether the formulas in the target document are written in a standardized way and formatted correctly. This can prevent formula parsing errors, display anomalies, and calculation failures, and ensure that the formula content in the document can be correctly identified and used.

[0048] Tag detection in a target document can include checking for typos, whether tags are closed, and whether tags are nested incorrectly, thereby eliminating problems such as missing tags, miswritten tags, and overlapping tags, ensuring the document structure is complete and can be parsed normally.

[0049] Detecting links in a target document can include checking whether the links are valid, incorrect, or broken, thereby ensuring normal access to links and improving document interactivity.

[0050] Detecting anchor points in a target document can include checking whether the anchor point names, associated positions, and jump relationships are correct. This can prevent anchor point jump failures and positioning errors, ensuring that the fixed-point jump function within the document works properly.

[0051] Detecting table fields in a target document can include checking for missing, incorrect, or structurally abnormal table fields, thereby ensuring that both the table data and the table structure are well-organized.

[0052] Detecting the reference relationships in a target document can include checking whether the reference paths in the target document are incorrect, whether the referenced objects are missing, and whether the reference associations are incorrect. This ensures that the reference relationships are valid, can be invoked normally, and guarantees that the target document runs normally and the associations are matched correctly.

[0053] Of course, it is understood that the content detected by the third detection method mentioned above is only an example. The content detected by the third detection method can be added or deleted based on the actual situation in order to detect the target document from more dimensions.

[0054] In this embodiment of the disclosure, different target business types and different target document formats can correspond to different second detection methods and third detection methods, thereby corresponding to different target detection methods. In this way, the target detection method can be determined based on the actual business type and the actual document format, so as to perform targeted detection on the target document and obtain more accurate detection results of the target document.

[0055] In some embodiments, the detection results may include the error location, error type, error severity level, and error cause, and may also generate corresponding modification suggestions based on the detection results.

[0056] Step S103: Based on the target detection method corresponding to each target document, detect at least one target document and obtain the detection result for each target document.

[0057] In some embodiments, each target document can be detected based on the target detection method corresponding to each target document, and target data in at least one target document can be detected based on the target detection method corresponding to each target document to obtain the detection result of each target document; the target data is data that is related in at least one target document. In this way, on the basis of single-document internal detection, the detection of target data related to multiple documents can be added, thereby identifying cross-document conflict problems that cannot be detected by single-document detection.

[0058] In some embodiments, each target document can be detected based on a target detection method corresponding to each target document to obtain a first detection result, and target data in at least one target document can be detected to obtain a second detection result. Based on the first and second detection results, a detection result for each target document can be obtained. The target data may include parameters such as version identifier, function name, parameter name, and reference relationships.

[0059] In this way, on the one hand, various errors in a single document can be detected and eliminated independently, ensuring the accuracy of the content, format, and business information of a single document; on the other hand, cross-document related data detection can identify interconnected defects that cannot be detected by single-document detection, such as information conflicts, abnormal references, and logical mismatches between multiple documents; the final document detection result is output by combining the results of the two types of detection, expanding the error detection coverage, reducing missed detection scenarios, and comprehensively improving the completeness and accuracy of document detection, ensuring the overall information unity and logical consistency of the entire target document cluster.

[0060] It should be noted that, based on the target detection method corresponding to each target document, the process of detecting each target document can refer to the content of the first detection method, the second detection method, and the third detection method mentioned above, and will not be repeated here.

[0061] In some embodiments, the electronic device can perform parallel detection on at least one target document, thereby enabling batch detection of documents and improving the efficiency of document detection.

[0062] In some embodiments, a detection report can be generated based on the detection results for each target document. The detection report may include a summary of the detection results for at least one target document. In one example, the results may be categorized and integrated according to document path, detection time, and error type, allowing users to easily view all errors present in at least one target document. Alternatively, the detection results for each target document can be displayed separately, facilitating users to view and analyze the errors present in each document. Accordingly, referring to the content included in the detection results, the detection report may include error location, error type, error severity level, error cause, and modification suggestions.

[0063] Optionally, the detection report can be a visual report, and the detection report can be exported. When generating the detection report, a timestamped running log can be generated at the same time. The running log is used to record the detection duration, errors, and model call records, so as to facilitate technicians to troubleshoot errors.

[0064] In some embodiments, for each target document, the space occupied by the target document can be determined. When the space occupied by the target document is large, the electronic device may not be able to support the recognition and detection of the target document. Therefore, for such target documents, the target document can be split to obtain multiple text fragments after the target document is split.

[0065] In some embodiments, multiple text segments after the target document has been split can be detected to obtain a detection result for each text segment, and the detection result of the target document can be obtained based on the detection results of each text segment. In some embodiments, each text segment in the multiple text segments after the target document has been split can be detected to obtain a third detection result, and the associated data in the multiple text segments can be detected to obtain a fourth detection result, and the detection result of each text segment can be obtained based on the third detection result and the fourth detection result.

[0066] In some embodiments, the target detection method corresponding to the target document can be used as the target detection method for each text segment among the multiple text fragments obtained from splitting the target document. Based on the target detection method of each text segment, each text segment is detected to obtain the detection result for each text segment. In one example, the target detection method corresponding to the target document can be bound to each text segment among the multiple text fragments obtained from splitting the target document, and the target detection method corresponding to the target document can be used as metadata for each text segment, thereby facilitating the subsequent detection process.

[0067] It should be noted that the related data in multiple text fragments can be referenced from the target data, which will not be elaborated here.

[0068] In some embodiments, the space occupied by the target document can be compared with a preset threshold, and the corresponding processing method can be determined based on the comparison result.

[0069] In some embodiments, the comparison results are categorized into the following two cases: In the first scenario, in some embodiments, for each target document, in response to the target document's space occupancy being greater than or equal to a preset threshold, the target document is split based on a target splitting method to obtain multiple text fragments after the target document is split. The target splitting method is used to split semantic units in the target document based on target splitting boundaries, and the target splitting boundaries are used to indicate the boundaries of semantic units in the target document that cannot be split into different text fragments. Then, each text fragment in the multiple text fragments after the target document is split can be detected based on the target detection method corresponding to the target document to obtain the detection result of each text fragment. Based on the detection result of each text fragment, the detection result of the target document is determined.

[0070] The preset threshold value can be set and selected based on actual conditions. It should be noted that the target document needs to be input into the preset model; therefore, the preset threshold can be related to the maximum space occupied by the preset model, which can refer to the maximum space occupied by the preset model's context window.

[0071] Optionally, the space usage can be expressed as the storage capacity. For example, if the maximum space usage supported by the preset model is 384KB, then the preset threshold can be 384KB.

[0072] Optionally, the space usage can be expressed in bytes. For example, if the maximum space usage supported by the preset model is 128K tokens, and each token can include 3~4 bytes, then the preset threshold can be 512 bytes.

[0073] It's important to note that some content in the target document needs to be detected from a holistic perspective. For example, a function in the code might require detection of the entire function. If this function is split into two different text segments during the splitting process, it will affect the detection results for that function. The target splitting boundary indicates the boundaries of semantic units in the target document that cannot be split into different text segments. Therefore, splitting the target document based on the target splitting boundary ensures that semantic units that cannot be split into different text segments are grouped into the same text segment, avoiding splitting them into different text segments and thus affecting the detection results of the target document.

[0074] In some embodiments, the target splitting method includes a first splitting method, a second splitting method, and a third splitting method. The first splitting method is used to split semantic units in the target document based on a first splitting boundary, where the first splitting boundary represents the boundary of semantic units in the target document that cannot be split into different text fragments corresponding to the outline level. The second splitting method is used to split semantic units in the target document based on a second splitting boundary, where the second splitting boundary represents the boundary of semantic units in the target document that cannot be split into different text fragments corresponding to the target business type. The third splitting method is used to split semantic units in the target document based on a third splitting boundary, where the third splitting boundary represents the boundary of semantic units in the target document that cannot be split into different text fragments corresponding to the target document format.

[0075] In this way, three types of splitting boundaries can be set to correspond to outline level, business content, and file format, respectively, so as to match the splitting strategy with the document and flexibly select the splitting method according to the characteristics of the document.

[0076] The following is a brief explanation of the above splitting methods.

[0077] For the first splitting method: In some embodiments, the first splitting method is used to split the target document based on at least one of the chapters, headings and paragraphs in the target document. That is, the outline level in the target document may include at least one of the chapters, headings and paragraphs in the target document.

[0078] For the second splitting method: In some embodiments, the second splitting method is used to split semantic units in the target document based on at least one of function information, interface information, command line examples, parameter configurations, version change information, operation steps, and tables in the target document. That is, the second splitting boundary may include at least one of function information, interface information, command line examples, parameter configurations, version change information, operation steps, and tables in the target document.

[0079] In some embodiments, similar to the second detection method, the boundaries of semantic units that cannot be divided into different text fragments corresponding to different target business types are different. For example, for R&D documents such as user manuals, release instructions, and tutorials, function / interface description blocks, command line example blocks, parameter configuration blocks, version change blocks, operation step blocks, and table blocks can be used as the second splitting boundaries.

[0080] Among them, splitting semantic units in the target document based on function information in the target document can refer to using functions as semantic splitting boundaries, thereby accurately segmenting independent functional logic blocks in the document, distinguishing the explanatory text corresponding to different functions, facilitating subsequent detection of the consistency of terms, parameters, and descriptions by function dimension, realizing independent quality inspection of function-related content, and improving the accuracy of semantic recognition and error location.

[0081] Separating semantic units in a target document based on interface information can refer to using interface information as the semantic decomposition boundary. This can isolate the input parameters, output parameters, and call logic descriptions of each interface, avoid the interference of multiple interface contents with semantic parsing, and facilitate subsequent compliance testing and correlation comparison for a single interface description, ensuring the independence and integrity of interface information.

[0082] Separating semantic units in a target document based on command-line examples can refer to using command-line example blocks as semantic separation boundaries. Each command-line example corresponds to a text fragment, so that the format, parameters, operation descriptions, and other parameters of the command-line example can be detected separately, and errors in each command-line example can be accurately identified.

[0083] Semantic units in a target document can be split based on parameter configurations in the target document. This means using parameters such as parameter names and configuration descriptions as semantic splitting boundaries, which facilitates subsequent detection of whether the parameters themselves have errors and detection of consistency of parameters across documents.

[0084] Separating semantic units in a target document based on version change information can refer to using version change information as the semantic decomposition boundary, thereby isolating changes, additions, and deletions corresponding to different versions, distinguishing version-specific semantic information, and effectively detecting cross-content logic problems such as conflicting descriptions of different versions and mismatched version information.

[0085] Separating semantic units in a target document based on operation steps can refer to using complete operation steps as semantic segmentation boundaries. This avoids splitting operation steps in an operation process into different text fragments, which would affect the subsequent detection of operation steps.

[0086] Semantic units in a target document can be split by using the entire table as the semantic split boundary. This allows for the separation of fields and main text within the table, enabling separate detection of the table and avoiding interference between the main text and the table during detection. This ensures the accuracy of both the main text and the table detection.

[0087] In this way, at least one of the following can be used as the split boundary: function, interface, command line, parameter configuration, version, operation steps, or table. This allows highly related and functionally independent content to be grouped into the same text segment. Subsequent detection can then be performed on individual text segments, reducing missed detections and false positives caused by interference between different text segments, improving the accuracy of semantic parsing and the efficiency of document error localization, and providing standardized and structured foundational data for single-document detection and multi-document correlation comparison.

[0088] For the third splitting method: In some embodiments, the third segmentation method is used to segment semantic units in the target document based on at least one of tags, anchors, links, chapter commands, contextual information, formulas, and tables in the target document. That is, at least one of tags, anchors, links, chapter commands, contextual information, formulas, and tables in the target document.

[0089] In some embodiments, similar to the third detection method, the boundaries of semantic units that cannot be divided into different text fragments differ for different target document formats. For example, for XML / HTML format documents, parent-child tag nodes, anchor marks, link targets, and tag closure relationships can be used as the third splitting boundaries. For LaTeX documents, chapter commands, environment blocks, formula blocks, reference tags, and table environments can be used as the third splitting boundaries.

[0090] Among them, splitting semantic units in the target document based on tags in the target document can refer to using parent-child tag nodes and tag closure relationships as semantic splitting boundaries, thereby avoiding the destruction of parent-child tag node relationships and tag closure relationships in the target document.

[0091] Semantic units in a target document can be split based on anchor points. This means that anchor points can be used as the boundaries for semantic splitting, allowing for individual verification of anchor naming and jump correspondence, and quickly locating issues such as anchor failure and mismatch between anchor points and text descriptions.

[0092] Semantic units in a target document can be decomposed based on links in the target document. This means using the link carrier text, link address, link description, etc., as semantic decomposition boundaries, which facilitates subsequent detection of link validity and consistency between link description and jump target.

[0093] Semantic units in a target document can be split based on the chapter commands in the target document. This means using the LaTeX chapter commands of the document as the semantic split boundary, thereby distinguishing first-level, second-level, and multi-level chapters in a LaTeX document, realizing chapter-by-chapter detection, and facilitating the detection of chapter logical order, chapter content completeness, and conflicting descriptions between chapters.

[0094] Separating semantic units from target documents based on environmental information can refer to using the document's runtime environment, deployment environment, and hardware / software environment descriptions as semantic decomposition boundaries. This isolates various environment configuration descriptions into independent semantic units, allowing for the separate detection of environment parameters and version compatibility requirements, and accurately detecting inconsistencies and parameter mismatches in multiple environmental descriptions.

[0095] Decomposing semantic units in a target document based on environmental information can mean dividing formulas and their accompanying explanatory text into independent semantic units, distinguishing formulas from ordinary text. This allows for separate verification of formula syntax, symbols, and parameter definitions, preventing interference from the main text with formula parsing, and quickly identifying errors in formula writing and inconsistencies in symbols.

[0096] The content of the table can be found in the second splitting method mentioned above, and will not be repeated here.

[0097] In some embodiments, attribute information can be set for each text fragment. This attribute information may include a chunk identifier (chunk_id), a document identifier (file_id), the target document format (file_format), the target business type (document_type), the section path (section_path), the start and end line numbers (start_line and end_line), the parent tag path (parent_tag_path), the function or command line name (function_or_command_name), the table caption (table_caption), the anchor identifier (anchor_id), the referenced target (referring to the target resource pointed to by links, references, jumps, etc.), and the matched detection method identifier (matched_rules). In this way, each text fragment can be targeted for detection based on the attribute information of each document, and the relevant parameters of document detection can be recorded for convenient subsequent traceability and management.

[0098] Thus, by splitting the target document in the above manner, the integrity of text fragments can be detected. For example, compared to previous methods involving cross-text fragment references, table continuation, cross-text fragment environment information, cross-text fragment tags, cross-text fragment interface information, and cross-text fragment function information, this embodiment of the disclosure can split related information into the same text fragment, thereby enabling the integrity detection of text fragments and obtaining more accurate detection results.

[0099] In some embodiments, similar to the detection process for the target document, associated data between multiple text segments can be detected, such as verifying the consistency of associated data within multiple text segments. Additionally, for two adjacent text segments, secondary detection can be performed at their boundary; the boundary can be set and selected based on the actual situation.

[0100] In some embodiments, detection results for each text segment can be obtained, resulting in multiple detection results. Then, these multiple detection results can be deduplicated and merged to obtain the detection results for the target document. Deduplication removes duplicate errors from multiple detection results, while merging combines different errors from multiple detection results. For example, if multiple detection results include error 1, error 2, and error 1, then the detection results for the target document can include error 1 and error 2.

[0101] Thus, limiting the second splitting method to business elements such as functions, interfaces, command lines, versions, and operation steps can divide independent business function descriptions into complete text fragments, facilitating subsequent consistency verification of business dimensions; limiting the third splitting method to formatted structured elements such as tags, anchors, formulas, and chapter commands can completely preserve the document format structure units, ensuring that format syntax detection and link anchor verification are not destroyed by the splitting.

[0102] In the second scenario, for each target document, in response to the target document's space occupancy being less than a preset threshold, the target document is detected using the target document detection method, and the detection result of the target document is obtained.

[0103] In some embodiments, if the space occupied by the target document is less than a preset threshold, the target document is not split, and the entire target document can be detected directly to obtain the detection result of the target document.

[0104] In some embodiments, the electronic device can perform parallel detection on at least one target document, thereby enabling batch document detection and improving document detection efficiency. Similarly, the electronic device can perform parallel detection on multiple text fragments.

[0105] The method provided in this disclosure can set a file space usage threshold to differentiate between large and small documents. Only documents with large space usage are split, while documents with small space usage are directly detected completely, thus balancing detection efficiency and resource consumption. Furthermore, by determining semantic splitting boundaries from multiple dimensions, the semantic integrity of the split text fragments can be ensured, thereby reducing missed detections or false positives due to semantic incompleteness and improving the stability and processing speed of large document detection.

[0106] According to embodiments of this disclosure Figure 2 A flowchart of a detection method is shown, see reference. Figure 2 The method includes: S1. For each target document, determine whether the space occupied by the target document is greater than or equal to a preset threshold. If yes, proceed to step S2; otherwise, proceed to step S5.

[0107] S2. Determine the target splitting method based on the target business type and target document format.

[0108] S3. Based on the target splitting method, the target document is split to obtain multiple text fragments corresponding to the target document.

[0109] S4. Detect each text segment based on the object detection method to obtain the detection results of the target document.

[0110] S5. Detect the target document based on the target detection method and obtain the detection results of the target document.

[0111] In some embodiments, such as Figure 3 The diagram illustrates a document detection method. In some embodiments, each target document from at least one target document can be input into a document detection device to obtain a detection result for each target document. The document detection device includes a scheduling agent module, a detection skill module, and a processing module. The scheduling agent module can perform document recognition, document splitting, deduplication and merging of detection results, contextual correlation detection, and full-process scheduling and control of the document detection method. The detection skill module is independent of the scheduling agent module and can determine the corresponding detection method based on the document format and business type determined by the scheduling agent module. It should be noted that the detection skill module can pre-build multiple orthogonal detection methods to facilitate broader, multi-dimensional detection of documents, enabling comprehensive detection of potential document problems. The detection skill module can determine the corresponding detection method based on the actual document format and business type, thereby achieving targeted document detection. The processing module includes a preset model, which can be used to perform document detection operations.

[0112] Among these features, document recognition can traverse a specified file directory, collect at least one target document from that directory, and statistically analyze the attribute information (quantity, size, and storage path, etc.) of each target document. Each target document is then labeled to distinguish between files with large and small file sizes, and the document's business type and format are identified. Document splitting can handle scenarios with large or numerous documents by splitting them into multiple text fragments, thus avoiding the context window limitations of preset models and ensuring smooth document detection. Deduplication and merging of detection results can summarize the results for easy user review. Context association detection can perform correlation detection between documents and text fragments, thereby increasing the detection scope and accuracy. Full-process scheduling and control can manage the document detection process. In one example, it can manage and monitor the entire lifecycle, such as task distribution, status monitoring, exception retries, and process termination. For corrupted documents, documents with abnormal formats, and tasks that time out via interfaces, it can automatically mark abnormal documents, skip invalid files, and retry abnormal tasks.

[0113] like Figure 3 As shown, for each target document, the scheduling agent module can determine the target document format and target business type. Furthermore, the scheduling agent module can determine whether to split the target document. If so, it can split the document into multiple text fragments. If not, the target document is used for subsequent detection processes.

[0114] In one example, for each target document, after determining the target business type and target document format, the scheduling agent module can call the detection skill module so that the detection skill module can determine the target detection method for the target document based on the target business type and target document format.

[0115] Similarly, in one example, after the scheduling agent module determines the target business type and target document format of the target document, and after the scheduling agent module splits the target document into multiple text fragments, for each text fragment, the scheduling agent module can call the detection skill module so that the detection skill module determines the target detection method of the target document based on the target business type and target document format, and uses the detection method of the target document as the detection method of each text fragment obtained from the split target document.

[0116] In some embodiments, the scheduling agent module may invoke a preset model in the processing module to enable the preset model to detect the target document and / or text fragments.

[0117] In some embodiments, after invoking a preset model, for each target document, the target document and the target detection method determined by the detection skill module can be input into the preset model of the processing module. The preset model can then detect the target document based on the target detection method to obtain the detection result of the target document. Similarly, in some embodiments, for each text segment, the text segment determined by the scheduling agent module and the target detection method determined by the detection skill module can be input into the preset model of the processing module to obtain the detection result of the text segment. The scheduling agent module can then obtain the detection result of the target document based on the detection result of each text segment among multiple text segments of the same target document.

[0118] In some embodiments, the preset model can be a large language model. It should be noted that prompt words for the preset model can be determined based on object detection methods, and the target document and prompt words, and / or text fragments and prompt words, are respectively input into the preset model to obtain the detection results for the target document and / or the detection results for the text fragments. Thus, by inputting prompt words, false errors and irrelevant prompts generated by model illusions can be avoided, reducing the probability of false errors and misleading information.

[0119] In some embodiments, the scheduling agent module can merge the detection results of each target document into an initial detection report, and call a preset model to perform cross-document detection on each target document to generate a visual detection report in HTML format. This allows technicians to view the problems and modification suggestions in the document and make corrections to the document.

[0120] In some embodiments, a specific example is given below: The environment of the electronic device is configured, and then a startup command is entered. This startup command is used to start a preset model. The preset model can be input with a folder path and a detection command for detecting documents within the folder. Correspondingly, a scheduling agent module can be activated. The scheduling agent module can identify all documents in the folder and recognize the business type and document format of each document. For example, it can identify that document 1 is a user manual and a LaTeX document, and document 2 is a version update description and an HTML document. Then, the scheduling agent module can split the documents that need to be split, obtaining multiple text fragments corresponding to each document. For example, if document 1 needs to be split, multiple text fragments corresponding to document 1 can be obtained. Next, the scheduling agent module can call the detection skill module. The detection skill module can determine the detection method corresponding to each document and use this detection method as the detection method for each text fragment obtained after splitting the document. Finally, the detection skill module can call the preset model. The preset model can perform parallel detection on the document or its text fragments based on the detection method corresponding to each document, obtaining the detection result for each document.

[0121] Exemplary embodiments of this disclosure provide a document detection device, such as... Figure 4 As shown in the figure, a block diagram of a document detection device disclosed herein includes: The scheduling agent module 401 is configured to identify at least one target document and determine the target business type and target document format for each target document. The detection skill module 402 is also configured to determine the target detection method corresponding to each target document based on the target business type and target document format of each target document. The target detection method is used to detect documents of the target business type and target document format. The target detection method includes at least one of the first detection method, the second detection method, and the third detection method. The first detection method is used to detect errors related to language rules in the target document, the second detection method is used to detect errors related to the target business in the target document, and the third detection method is used to detect errors related to document format in the target document. The processing module 403 is configured to detect at least one target document based on the target detection method corresponding to each target document, and obtain the detection result of each target document.

[0122] In some embodiments, the second detection method is used to detect errors in at least one of the following in the target document: terminology, product name, function name, version identifier, parameter name, completeness of operation steps, logic of document content, and correlation of operation steps. The third detection method is used to detect errors in at least one of the following in the target document: instruction syntax, formula syntax, tags, links, anchors, table fields, and reference relationships.

[0123] In some embodiments, the scheduling agent module 401 is configured as follows: For each target document, in response to the target document's space occupancy being greater than or equal to a preset threshold, the target document is split based on the target splitting method to obtain multiple text fragments after the target document is split. The target splitting method is used to split semantic units in the target document based on the target splitting boundary. The target splitting boundary is used to indicate the boundary of semantic units in the target document that cannot be split into different text fragments. Processing module 403 is configured as follows: Based on the target detection method corresponding to the target document, detect each text segment in the multiple text segments after the target document is split, and obtain the detection result of each text segment; Based on the detection results of each text segment, the detection results of the target document are determined.

[0124] In some embodiments, the target splitting method includes a first splitting method, a second splitting method, and a third splitting method; The first splitting method is used to split semantic units in the target document based on the first splitting boundary, where the first splitting boundary is used to characterize the boundary of semantic units in the target document that cannot be split into different text fragments at the outline level. The second splitting method is used to split semantic units in the target document based on the second splitting boundary. The second splitting boundary is used to characterize the boundary of semantic units in the target document that cannot be split into different text fragments corresponding to the target business type. The third splitting method is used to split semantic units in the target document based on the third splitting boundary, which is used to characterize the boundary of semantic units in the target document that cannot be split into different text segments according to the target document format.

[0125] In some embodiments, the second splitting method is used to split semantic units in the target document based on at least one of function information, interface information, command line examples, parameter configuration, version change information, operation steps, and tables in the target document, and the third splitting method is used to split semantic units in the target document based on at least one of tags, anchors, links, chapter commands, environment information, formulas, and tables in the target document.

[0126] In some embodiments, the processing module 403 is configured to: For each target document, in response to the target document's space occupancy being less than a preset threshold, the target document is detected based on the target document's target detection method, and the detection result of the target document is obtained.

[0127] In some embodiments, the processing module 403 is configured to: Based on the target detection method corresponding to each target document, target data in at least one target document is detected to obtain the detection result of each target document; the target data is the data that is related in at least one target document.

[0128] Each module in the aforementioned document inspection device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in the processor of the electronic device in hardware form or independent of it, or stored in the memory of the electronic device in software form, so that the processor can call and execute the operations corresponding to each module.

[0129] In one exemplary embodiment, an electronic device is provided, including a processor and a memory, the memory storing a computer program, the processor executing the computer program to implement the steps of any of the document detection methods described above.

[0130] In one exemplary embodiment, a computer-readable storage medium is provided, on which a computer program is stored, which, when executed by a processor, implements the steps of any of the document detection methods described above. The computer-readable storage medium may be a ROM, random access memory (RAM), CD-ROM, magnetic tape, floppy disk, or optical document storage device, etc.

[0131] refer to Figure 5 The following description serves as a structural block diagram of the electronic device 500 disclosed herein. The electronic device 500 includes a computing unit 501, which can perform various appropriate actions and processes based on a computer program stored in a read-only memory (ROM) 502 or a computer program loaded from a storage unit 508 into a random access memory (RAM) 503. The RAM 503 may also store various programs and documents required for the operation of the electronic device 500. The computing unit 501, ROM 502, and RAM 503 are interconnected via a bus 504. An input / output (I / O) interface 505 is also connected to the bus 504.

[0132] Multiple components in electronic device 500 are connected to I / O interface 505, including: input unit 506, output unit 507, storage unit 508, and communication unit 509. Input unit 506 can be any type of device capable of inputting information to electronic device 500. Input unit 506 can receive input digital or character information and generate key signal inputs related to user settings and / or function control of electronic device 500, and may include, but is not limited to, a mouse, keyboard, touchscreen, trackpad, trackball, joystick, microphone, and / or remote control. Output unit 507 can be any type of device capable of presenting information, and may include, but is not limited to, a monitor, speaker, video / audio output terminal, vibrator, and / or printer. Storage unit 508 may include, but is not limited to, a hard disk and an optical disk. Communication unit 509 allows electronic device 500 to exchange information / documents with other devices through computer networks such as the Internet and / or various telecommunications networks, and may include, but is not limited to, modems, network cards, infrared communication devices, wireless communication transceivers, and / or chipsets, such as Bluetooth™ devices, WiFi devices, WiMax devices, cellular communication devices, and / or the like.

[0133] The computing unit 501 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 501 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 501 performs the various methods and processes described above, such as document detection methods. For example, in some embodiments, the document detection method may be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 508. In some embodiments, part or all of the computer program may be loaded and / or installed on the electronic device 500 via ROM 502 and / or communication unit 509. When the computer program is loaded into RAM 503 and executed by the computing unit 501, one or more steps of the document detection method described above may be performed. Alternatively, in other embodiments, the computing unit 501 may be configured to perform the document detection method by any other suitable means (e.g., by means of firmware).

[0134] The electronic device 500 may be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components to perform the document detection method described above.

[0135] Other embodiments of the invention will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of the invention that follow the general principles of the invention and include common knowledge or customary techniques in the art not disclosed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of the invention are indicated by the following claims.

[0136] It should be understood that the present invention is not limited to the precise structure described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of the invention is limited only by the appended claims.

Claims

1. A document detection method characterized by, include: Identify at least one target document, and determine the target business type and target document format for each target document; Based on the target business type and target document format of each target document, a target detection method corresponding to each target document is determined, wherein the target detection method includes at least one of a first detection method, a second detection method, and a third detection method; The target detection method is used to detect documents of the target business type and the target document format. The first detection method is used to detect language rule-related errors in the target document, the second detection method is used to detect errors related to the target business type in the target document, and the third detection method is used to detect errors related to the target document format in the target document. Based on the target detection method corresponding to each target document, the at least one target document is detected, and the detection result of each target document is obtained.

2. The document detection method of claim 1, wherein, The second detection method is used to detect errors in at least one of the following in the target document: terminology, product name, function name, version identifier, parameter name, completeness of operation steps, logic of document content, and relevance of operation steps. The third detection method is used to detect errors in at least one of the following in the target document: instruction syntax, formula syntax, tags, links, anchors, table fields, and reference relationships.

3. The document detection method according to claim 1, characterized in that, The method further includes: For each target document, in response to the space occupancy of the target document being greater than or equal to a preset threshold, the target document is split based on the target splitting method to obtain multiple text segments after the target document is split. The target splitting method is used to split semantic units in the target document based on the target splitting boundary. The target splitting boundary is used to indicate the boundary of semantic units in the target document that cannot be split into different text segments. The step of detecting the at least one target document based on the target detection method corresponding to each target document, and obtaining the detection result for each target document, includes: Based on the target detection method corresponding to the target document, each text segment in the multiple text segments after the target document is split is detected, and the detection result of each text segment is obtained; Based on the detection results of each text fragment, the detection result of the target document is determined.

4. The document detection method according to claim 3, characterized in that, The target splitting method includes a first splitting method, a second splitting method, and a third splitting method; The first splitting method is used to split semantic units in the target document based on a first splitting boundary, wherein the first splitting boundary is used to characterize the boundary of semantic units in the target document that cannot be split into different text segments corresponding to the outline level; The second splitting method is used to split semantic units in the target document based on a second splitting boundary, whereby the second splitting boundary is used to characterize the boundary of semantic units in the target document that cannot be split into different text segments corresponding to the target business type; The third splitting method is used to split semantic units in the target document based on the third splitting boundary, whereby the third splitting boundary is used to characterize the boundary of semantic units in the target document that cannot be split into different text segments corresponding to the target document format.

5. The document detection method according to claim 4, characterized in that, The second splitting method is used to split semantic units in the target document based on at least one of the following: function information, interface information, command line examples, parameter configuration, version change information, operation steps, and tables. The third splitting method is used to split semantic units in the target document based on at least one of the following: tags, anchors, links, chapter commands, environment information, formulas, and tables.

6. The document detection method according to claim 1, characterized in that, The step of detecting the at least one target document based on the target detection method corresponding to each target document, and obtaining the detection result for each target document, includes: For each target document, in response to the target document's space occupancy being less than a preset threshold, the target document is detected based on the target document's target detection method, and the detection result of the target document is obtained.

7. The document detection method according to claim 1, characterized in that, The step of detecting the at least one target document based on the target detection method corresponding to each target document, and obtaining the detection result for each target document, includes: Based on the target detection method corresponding to each target document, target data in the at least one target document is detected to obtain the detection result of each target document; the target data is data that is related in the at least one target document.

8. A document detection device, characterized in that, include: The scheduling agent module is configured to identify at least one target document and determine the target business type and target document format for each target document. The detection skills module is configured to determine the target detection method corresponding to each target document based on the target business type and the target document format of each target document, wherein the target detection method includes at least one of a first detection method, a second detection method, and a third detection method; The target detection method is used to detect documents of the target business type and the target document format. The first detection method is used to detect language rule-related errors in the target document, the second detection method is used to detect target business-related errors in the target document, and the third detection method is used to detect document format-related errors in the target document. The processing module is configured to detect the at least one target document based on the target detection method corresponding to each target document, and obtain the detection result of each target document.

9. An electronic device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 7.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 7.