Text detection method and system and server

By using automated methods for detecting garbled characters, signatures, and content, the detection challenges in batch processing of judicial electronic documents have been solved, improving processing efficiency and accuracy, reducing the workload of manual detection, and minimizing judicial risks.

CN120874815APending Publication Date: 2025-10-31SHENZHEN HAIGUI NETWORK TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511030371.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-25
Publication Date
2025-10-31

AI Technical Summary

Technical Problem

Existing technologies lack automated detection methods when processing judicial electronic documents in batches, resulting in problems such as garbled characters, missing signatures, and incomplete content, making the processing time-consuming, labor-intensive, and posing judicial risks.

Method used

This paper provides a text detection method that detects garbled characters through character parsing and encoding rules, detects the position and clarity of signatures using signature templates, and detects document content by combining content templates, thereby achieving automated integrity detection.

Benefits of technology

It improves the efficiency and accuracy of electronic document processing, reduces the workload of manual inspection, and reduces judicial risks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120874815A_ABST
    Figure CN120874815A_ABST
Patent Text Reader

Abstract

The invention provides a text detection method and system and a server, and relates to the technical field of text detection, the method can automatically and comprehensively detect the integrity of judicial electronic documents, and can realize messy code detection, signature completion detection and content integrity detection at the same time, so that the processing efficiency and accuracy of the electronic documents are improved, and the user experience is improved. And the workload of manual detection is greatly reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of text detection technology, and in particular to a text detection method, system and server. Background Technology

[0002] For judicial documents, the corresponding electronic documents contain a considerable amount of content. Besides the text related to the judicial case, they also involve sensitive personal data, judgment clauses, case number information, signature information, and many other types. The generation of electronic documents is usually based on scanning paper documents. This process can produce garbled characters due to technical limitations, and the scanning process does not check the integrity of signatures and text content. Therefore, when processing electronic documents in batches, manual integrity checks are necessary, which is time-consuming and labor-intensive. Summary of the Invention

[0003] To address the technical problem of the lack of automated means for batch processing of judicial electronic documents in existing technologies, the present invention aims to provide a text detection method, system, and server that can automatically and comprehensively detect the integrity of judicial electronic documents. It can simultaneously detect garbled characters, complete signatures, and content integrity, thereby improving the processing efficiency and accuracy of electronic documents and significantly reducing the workload of manual detection.

[0004] In a first aspect, embodiments of the present invention provide a text detection method, the method comprising: The steps for garbled text detection are as follows: 1. Obtain the electronic document and its format parameters corresponding to the text to be detected. 2. Determine the character parsing rules and character encoding rules corresponding to the electronic document using the format parameters. 3. Use the character parsing rules to parse the characters contained in the electronic document. 4. Use the character encoding rules to determine the character encoding result corresponding to the characters. 5. Determine the garbled text detection result of the electronic document based on the character encoding result. Signature detection steps: After converting the electronic document into a digital image using the image format parameters contained in the format parameters, the signature image contained in the digital image is obtained using the preset signature template, and the signature detection result of the electronic document is determined by the signature position data and signature clarity data corresponding to the signature image. Content detection steps: Determine the content template corresponding to the electronic document using the document type data contained in the format parameters; obtain the corresponding field content in the electronic document based on the key fields contained in the content template and their corresponding position parameters; and determine the content detection result of the electronic document based on the field content. The detection summary steps are as follows: Based on the garbled character detection results, signature detection results, and content detection results, determine the text detection results corresponding to the text to be detected.

[0005] Optional garbled character detection steps include: Obtain the electronic document corresponding to the portable file format parameters of the text to be detected, and determine the character parsing rules and character encoding rules corresponding to the electronic document based on the character set corresponding to the portable file format parameters; After parsing the text content of the electronic document using character parsing rules, the characters contained in the text content are obtained, and the baseline format result corresponding to the characters under the character parsing rules is determined. After encoding the characters using character encoding rules, the character encoding result is obtained. The garbled character detection result corresponding to the electronic document is determined based on the comparison between the character encoding result and the baseline format result.

[0006] Optionally, after retrieving the characters contained in the text content, the method also includes: Obtain the semantic analysis results corresponding to the text content, and determine the grammar rules corresponding to the characters based on the semantic analysis results; The semantic detection result corresponding to the character is determined according to the grammar rules, and the semantic detection result is used to update the garbled character detection result.

[0007] Optional signature verification steps include: Obtain the portable file format parameters contained in the format parameters, determine the image format parameters corresponding to the portable file format parameters, and use the corresponding conversion parameters between the portable file format parameters and the image format parameters to convert the electronic document into a digital image. Obtain a preset signature template, determine the sub-region in the digital image that matches the signature template based on the signature template, and determine the signature image contained in the digital image based on the boundary parameters of the sub-region; Obtain the preset signature area corresponding to the signature template in the digital image, determine the area boundary and the first center point of the preset signature area, and determine the second center point of the signature image. Then, use the offset between the first center point and the second center point and the relative distance between the area boundary and the second center point to determine the signature position data corresponding to the signature. After performing edge detection calculations on the signature in the signature image, the edge information corresponding to the signature is obtained, and the signature clarity data corresponding to the signature image is determined based on the continuous parameters, intensity parameters and contrast parameters corresponding to the edge information. The signature detection results for electronic documents are determined based on signature location data and signature clarity data.

[0008] Optionally, the process of determining signature clarity data may also include: After performing texture detection calculations on the signature image, the texture information corresponding to the signature in the signature image is obtained, and the signature clarity data is determined based on the texture density corresponding to the texture information.

[0009] Optional content detection steps include: Obtain the document type data contained in the format parameters, use the document type data to determine the format type of the electronic document, and after obtaining the field parameters corresponding to the format type based on the format parameters, use the field parameters to determine the content template corresponding to the electronic document. After obtaining the key fields corresponding to the field parameters based on the content template and the corresponding position parameters of the key fields, the field display area corresponding to the key fields is obtained using the position parameters, and the corresponding field content in the electronic document is obtained using the field display area. The standard length of the field corresponding to the field display area is obtained using the content template, and the content detection result of the electronic document is determined based on the length difference between the field length of the field content and the standard length of the field.

[0010] Optionally, after retrieving the corresponding field content from the electronic document using the field display area, the method further includes: After performing semantic analysis on the field content and key fields, the first semantic result corresponding to the field content and the second semantic result corresponding to the key fields are obtained. The content detection result of the electronic document is determined based on the semantic comparison between the first semantic result and the second semantic result.

[0011] Optional, the detection summary step includes: Obtain the first moment corresponding to the garbled character detection result, and associate the first moment with the garbled character detection result to obtain the first detection report corresponding to the garbled character detection step; Obtain the second moment corresponding to the signature detection result, and associate the second moment with the signature detection result to obtain the second detection report corresponding to the signature detection step; Obtain the third moment corresponding to the content detection result, and then associate the third moment with the content detection result to obtain the third detection report corresponding to the content detection step; The first, second, and third test reports are combined according to the first, second, and third time points to obtain the test report, which is then determined as the text test result.

[0012] Secondly, the present invention provides a text detection system, the system comprising: The garbled text detection unit is used to obtain the electronic document and its format parameters corresponding to the text to be detected, determine the character parsing rules and character encoding rules corresponding to the electronic document through the format parameters, parse the characters contained in the electronic document using the character parsing rules, determine the character encoding result corresponding to the characters using the character encoding rules, and then determine the garbled text detection result of the electronic document based on the character encoding result. The signature detection unit is used to convert electronic documents into digital images using image format parameters included in the format parameters, obtain the signature image contained in the digital image using a preset signature template, and determine the signature detection result of the electronic document through the signature position data and signature clarity data corresponding to the signature image. The content detection unit is used to determine the content template corresponding to the electronic document using the document type data contained in the format parameters, obtain the corresponding field content in the electronic document based on the key fields contained in the content template and their corresponding position parameters, and determine the content detection result of the electronic document based on the field content. The detection summary unit is used to determine the text detection result corresponding to the text to be detected based on the garbled character detection result, signature detection result, and content detection result.

[0013] Thirdly, embodiments of the present invention also provide a server, including a processor and a memory, the memory storing computer-executable instructions that can be executed by the processor, the processor executing the computer-executable instructions to implement the steps of the text detection method provided in the first aspect.

[0014] Fourthly, embodiments of the present invention also provide a storage medium storing computer-executable instructions, which, when invoked and executed by a processor, cause the processor to implement the steps of the text detection method provided in the first aspect.

[0015] This invention provides a text detection method, system, and server. In the process of detecting judicial text, the method first acquires the electronic document and its format parameters corresponding to the text to be detected. The format parameters determine the character parsing rules and character encoding rules corresponding to the electronic document. The character parsing rules are used to parse the characters contained in the electronic document, and the character encoding rules are used to determine the character encoding results. Based on the character encoding results, the garbled text detection result of the electronic document is determined. Then, the image format parameters contained in the format parameters are used to convert the electronic document into a digital image. A preset signature template is used to acquire the signature image contained in the digital image, and the signature detection result of the electronic document is determined based on the signature position data and signature clarity data corresponding to the signature image. Subsequently, the document type data contained in the format parameters is used to determine the content template corresponding to the electronic document. Based on the key fields contained in the content template and their corresponding position parameters, the corresponding field content in the electronic document is acquired, and the content detection result of the electronic document is determined based on the field content. Finally, the text detection result corresponding to the text to be detected is determined based on the garbled text detection result, the signature detection result, and the content detection result. This method can automatically and comprehensively detect the integrity of judicial electronic documents, and can simultaneously detect garbled characters, complete signatures, and content integrity, thereby improving the processing efficiency and accuracy of electronic documents and significantly reducing the workload of manual detection.

[0016] Other features and advantages of the invention will be set forth in the description which follows, and will be apparent in part from the description, or may be learned by practicing the invention. The objects and other advantages of the invention are realized and obtained in accordance with the structures particularly pointed out in the description, claims and drawings.

[0017] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, preferred embodiments are described below in detail with reference to the accompanying drawings. Attached Figure Description

[0018] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0019] Figure 1 A flowchart of a text detection method provided in an embodiment of the present invention; Figure 2 This is a flowchart of the garbled character detection step in a text detection method provided by an embodiment of the present invention; Figure 3A flowchart illustrating the process after obtaining characters contained in the text content in a text detection method provided in this embodiment of the invention; Figure 4 This is a flowchart of the signature detection step in a text detection method provided by an embodiment of the present invention; Figure 5 This is a flowchart of the content detection step in a text detection method provided in an embodiment of the present invention; Figure 6 This is a flowchart of the detection and summarization step in a text detection method provided by an embodiment of the present invention; Figure 7 This is a schematic diagram of the structure of a text detection system provided in an embodiment of the present invention; Figure 8 This is a schematic diagram of the structure of a server provided in an embodiment of the present invention.

[0020] icon: 710 - Garbled Character Detection Unit; 720 - Signature Detection Unit; 730 - Content Detection Unit; 740 - Detection Summary Unit; 101 - Processor; 102 - Memory; 103 - Bus; 104 - Communication interface. Detailed Implementation

[0021] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the present invention will be clearly and completely described below in conjunction with the embodiments. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0022] For judicial documents, the corresponding electronic documents contain a wealth of content. Besides the text related to the judicial case, they also involve sensitive personal data, judgment clauses, case numbers, signatures, and many other types of information. The generation of electronic documents is usually based on scanning paper documents. This process can result in garbled characters due to technical limitations, and the scanning process does not check the integrity of signatures and text content. Therefore, when processing electronic documents in batches, manual inspection of their integrity is necessary, which is time-consuming and labor-intensive. In specific scenarios, when processing hundreds of similar cases in batches, ensuring the integrity of each electronic document (PDF format) requires dedicated personnel for manual inspection, which is extremely time-consuming.

[0023] Furthermore, because the inspection is done manually, omissions are prone to occur during the inspection of key elements, specifically involving: There is a risk of garbled text; unrecognizable garbled text may be generated during document generation or transmission. Documents lacking signatures or without electronic signatures are not legally valid. Content is incomplete, with key fields (such as case number, party information, and judgment clauses) missing or misplaced.

[0024] The aforementioned actions will bring serious legal risks. The service of incomplete documents will result in procedural violations and serious consequences such as retrial of the case.

[0025] In the process of examining judicial electronic documents using image recognition technology, general verification methods cannot recognize the unique signatures and seals of judicial documents, and traditional OCR technology is difficult to adapt to the specific typesetting format of judicial electronic documents. This results in a lack of automated detection methods for batch processing of judicial electronic documents in existing technologies. Therefore, this invention provides a text detection method, system, and server that can automatically and comprehensively detect the integrity of judicial electronic documents. It can simultaneously detect garbled characters, complete signatures, and content integrity, thereby improving the processing efficiency and accuracy of electronic documents and significantly reducing the workload of manual inspection.

[0026] To facilitate understanding of this embodiment, a text detection method disclosed in this invention will first be described in detail, such as... Figure 1 As shown, the method includes: S101: Obtain the electronic document and its format parameters corresponding to the text to be detected. Determine the character parsing rules and character encoding rules corresponding to the electronic document through the format parameters. Parse the characters contained in the electronic document using the character parsing rules. After determining the character encoding result corresponding to the character using the character encoding rules, determine the garbled text detection result of the electronic document based on the character encoding result. Signature detection step S102: After converting the electronic document into a digital image using the image format parameters contained in the format parameters, the signature image contained in the digital image is obtained using the preset signature template, and the signature detection result of the electronic document is determined by the signature position data and signature clarity data corresponding to the signature image. Content detection step S103: Determine the content template corresponding to the electronic document using the document type data contained in the format parameters, obtain the corresponding field content in the electronic document based on the key fields contained in the content template and their corresponding position parameters, and determine the content detection result of the electronic document based on the field content; Detection summary step S104: Determine the text detection result corresponding to the text to be detected based on the garbled character detection result, signature detection result, and content detection result.

[0027] Optional, garbled character detection step S101, such as Figure 2 As shown, it includes: Step S201: Obtain the electronic document with portable file format parameters corresponding to the text to be detected, and determine the character parsing rules and character encoding rules corresponding to the electronic document based on the character set corresponding to the portable file format parameters; Step S202: After parsing the text content corresponding to the electronic document using character parsing rules, obtain the characters contained in the text content and determine the baseline format result corresponding to the characters under the character parsing rules; Step S203: After encoding the characters using character encoding rules, the character encoding result is obtained. The garbled character detection result corresponding to the electronic document is determined based on the comparison between the character encoding result and the baseline format result.

[0028] Optionally, after obtaining the characters contained in the text content, such as Figure 3 As shown, the method also includes: Step S301: Obtain the semantic analysis results corresponding to the text content, and determine the grammar rules corresponding to the characters based on the semantic analysis results; Step S302: Determine the semantic detection result corresponding to the character according to the grammar rules, and update the garbled character detection result using the semantic detection result.

[0029] Optional, signature verification step S102, such as Figure 4 As shown, it includes: Step S401: Obtain the portable file format parameters contained in the format parameters, determine the image format parameters corresponding to the portable file format parameters, and convert the electronic document into a digital image using the conversion parameters corresponding to the portable file format parameters and the image format parameters. Step S402: Obtain a preset signature template, determine the sub-region in the digital image that matches the signature template based on the signature template, and determine the signature image contained in the digital image according to the boundary parameters of the sub-region. Step S403: Obtain the preset signature area corresponding to the signature template in the digital image, determine the area boundary and the first center point of the preset signature area, and determine the second center point of the signature image. Then, use the offset between the first center point and the second center point and the relative distance between the area boundary and the second center point to determine the signature position data corresponding to the signature. Step S404: After performing edge detection calculation on the signature in the signature image, obtain the edge information corresponding to the signature, and determine the signature clarity data corresponding to the signature image based on the continuous parameters, intensity parameters and contrast parameters corresponding to the edge information. Step S405: Determine the signature detection result corresponding to the electronic document based on the signature location data and signature clarity data.

[0030] Optionally, the process of determining signature clarity data may also include: After performing texture detection calculations on the signature image, the texture information corresponding to the signature in the signature image is obtained, and the signature clarity data is determined based on the texture density corresponding to the texture information.

[0031] Optionally, content detection step S103, such as Figure 5 As shown, it includes: Step S501: Obtain the document type data contained in the format parameters, determine the format type of the electronic document using the document type data, and after obtaining the field parameters corresponding to the format type based on the format parameters, determine the content template corresponding to the electronic document using the field parameters. Step S502: After obtaining the key fields corresponding to the field parameters based on the content template and obtaining the position parameters corresponding to the key fields, the field display area corresponding to the key fields is obtained using the position parameters, and the corresponding field content in the electronic document is obtained using the field display area. Step S503: Use the content template to obtain the standard length of the field corresponding to the field display area, and determine the content detection result of the electronic document based on the length difference between the field length of the field content and the standard length of the field.

[0032] Optionally, after retrieving the corresponding field content from the electronic document using the field display area, the method further includes: After performing semantic analysis on the field content and key fields, the first semantic result corresponding to the field content and the second semantic result corresponding to the key fields are obtained. The content detection result of the electronic document is determined based on the semantic comparison between the first semantic result and the second semantic result.

[0033] Optionally, the detection summary step S104, such as Figure 6 As shown, it includes: Step S601: Obtain the first moment corresponding to the garbled character detection result, and associate the first moment with the garbled character detection result to obtain the first detection report corresponding to the garbled character detection step; Step S602: Obtain the second moment corresponding to the signature detection result, and associate the second moment with the signature detection result to obtain the second detection report corresponding to the signature detection step; Step S603: Obtain the third moment corresponding to the content detection result, and associate the third moment with the content detection result to obtain the third detection report corresponding to the content detection step; Step S604: The first detection report, the second detection report, and the third detection report are summarized according to the first time point, the second time point, and the third time point to determine the text detection result.

[0034] As can be seen from the text detection method in the above embodiments, this method can automatically and comprehensively detect the integrity of judicial electronic documents, and can simultaneously detect garbled characters, complete signatures, and content integrity, thereby improving the processing efficiency and accuracy of electronic documents and significantly reducing the workload of manual detection.

[0035] Corresponding to the aforementioned text detection method, this embodiment of the invention also provides a text detection system, such as... Figure 7 As shown, the system includes: The garbled text detection unit 710 is used to obtain the electronic document and its format parameters corresponding to the text to be detected, determine the character parsing rules and character encoding rules corresponding to the electronic document through the format parameters, parse the characters contained in the electronic document using the character parsing rules, determine the character encoding result corresponding to the characters using the character encoding rules, and determine the garbled text detection result of the electronic document based on the character encoding result. The signature detection unit 720 is used to convert electronic documents into digital images using image format parameters contained in the format parameters, obtain the signature image contained in the digital image using a preset signature template, and determine the signature detection result of the electronic document through the signature position data and signature clarity data corresponding to the signature image. The content detection unit 730 is used to determine the content template corresponding to the electronic document using the document type data contained in the format parameters, obtain the corresponding field content in the electronic document based on the key fields contained in the content template and their corresponding position parameters, and determine the content detection result of the electronic document based on the field content. The detection and summarization unit 740 is used to determine the text detection result corresponding to the text to be detected based on the garbled character detection result, the signature detection result, and the content detection result.

[0036] As can be seen from the above text detection system, the system can automatically and comprehensively detect the integrity of judicial electronic documents. It can simultaneously detect garbled characters, complete signatures, and content integrity, thereby improving the processing efficiency and accuracy of electronic documents and significantly reducing the workload of manual detection.

[0037] The text detection system provided in this embodiment of the invention has the same implementation principle and technical effect as the aforementioned text detection method embodiment. For the sake of brevity, any parts not mentioned in the system embodiment can be referred to the corresponding content in the aforementioned text detection method embodiment.

[0038] This embodiment also provides a server, the structural diagram of which is shown below. Figure 8 As shown, the device includes a processor 101 and a memory 102; wherein, the memory 102 is used to store one or more computer instructions, which are executed by the processor to implement the steps of the above-described text detection method.

[0039] Figure 8 The server shown also includes a bus 103 and a communication interface 104. The processor 101, the communication interface 104, and the memory 102 are connected via the bus 103.

[0040] The memory 102 may include high-speed random access memory (RAM) and may also include non-volatile memory, such as at least one disk storage device. The bus 103 may be an ISA bus, PCI bus, or EISA bus, etc. The bus can be divided into address bus, data bus, control bus, etc. For ease of representation, Figure 8 The symbol is represented by a single double-headed arrow, but this does not mean that there is only one bus or one type of bus.

[0041] The communication interface 104 is used to connect to at least one user terminal and other network units through a network interface, and to send encapsulated IPv4 packets or IPv4 packets to the user terminal through the network interface.

[0042] Processor 101 may be an integrated circuit chip with signal processing capabilities. In implementation, each step of the above method can be completed by the integrated logic circuitry in the hardware of processor 101 or by instructions in software form. The processor 101 can be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), etc.; it can also be a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field-Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. It can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of this disclosure. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the methods disclosed in the embodiments of this disclosure can be directly manifested as execution by a hardware decoding processor, or execution by a combination of hardware and software modules in the decoding processor. The software module can reside in a mature storage medium in the art, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, or registers. This storage medium is located in memory 102. The processor 101 reads the information in memory 102 and, in conjunction with its hardware, completes the steps of the method described in the foregoing embodiments.

[0043] This invention also provides a storage medium storing a computer program, which, when run by a processor, executes the steps of the text detection method described in the foregoing embodiments.

[0044] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, devices, and methods can be implemented in other ways. The system embodiments described above are merely illustrative. For example, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. Furthermore, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Additionally, the coupling or direct coupling or communication connection shown or discussed may be through some communication interfaces; the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms.

[0045] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0046] In addition, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.

[0047] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a processor-executable, non-volatile, computer-readable storage medium. Based on this understanding, the technical solution of this invention, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0048] Finally, it should be noted that the above-described embodiments are merely specific implementations of the present invention, used to illustrate the technical solutions of the present invention, and not to limit it. The scope of protection of the present invention is not limited thereto. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that any person skilled in the art can still modify or easily conceive of changes to the technical solutions described in the foregoing embodiments within the technical scope disclosed in the present invention, or make equivalent substitutions for some of the technical features; and these modifications, changes, or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be covered within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.

Claims

1. A text detection method, characterized in that, The method includes: The steps for detecting garbled characters are as follows:

1. Obtain the electronic document and its format parameters corresponding to the text to be detected.

2. Determine the character parsing rules and character encoding rules corresponding to the electronic document using the format parameters.

3. Parse the characters contained in the electronic document using the character parsing rules.

4. Determine the character encoding result corresponding to the character using the character encoding rules.

5. Determine the garbled character detection result of the electronic document based on the character encoding result. Signature detection steps: After converting the electronic document into a digital image using the image format parameters included in the format parameters, the signature image contained in the digital image is obtained using a preset signature template, and the signature detection result of the electronic document is determined by the signature position data and signature clarity data corresponding to the signature image. Content detection steps: Determine the content template corresponding to the electronic document using the document type data contained in the format parameters; obtain the corresponding field content in the electronic document based on the key fields contained in the content template and their corresponding position parameters; and determine the content detection result of the electronic document based on the field content. Detection summary steps: Based on the garbled character detection results, the signature detection results, and the content detection results, determine the text detection results corresponding to the text to be detected.

2. The text detection method according to claim 1, characterized in that, The garbled text detection steps include: The electronic document is obtained by acquiring the portable file format parameters corresponding to the text to be detected, and the character parsing rules and character encoding rules corresponding to the electronic document are determined based on the character set corresponding to the portable file format parameters. After parsing the text content corresponding to the electronic document using the character parsing rules, the characters contained in the text content are obtained, and the baseline format result corresponding to the characters under the character parsing rules is determined; The character encoding result is obtained by encoding the character using the character encoding rules, and the garbled character detection result corresponding to the electronic document is determined based on the comparison result between the character encoding result and the baseline format result.

3. The text detection method according to claim 2, characterized in that, After obtaining the characters contained in the text content, the method further includes: Obtain the semantic analysis results corresponding to the text content, and determine the grammar rules corresponding to the characters based on the semantic analysis results; The semantic detection result corresponding to the character is determined according to the grammar rules, and the garbled character detection result is updated using the semantic detection result.

4. The text detection method according to claim 1, characterized in that, The signature detection steps include: Obtain the portable file format parameters contained in the format parameters, determine the image format parameters corresponding to the portable file format parameters, and convert the electronic document into the digital image using the conversion parameters corresponding to the portable file format parameters and the image format parameters; Obtain the preset signature template, determine the sub-region in the digital image that matches the signature template based on the signature template, and determine the signature image contained in the digital image according to the boundary parameters of the sub-region; Obtain the preset signature area corresponding to the signature template in the digital image, determine the area boundary and the first center point of the preset signature area, and determine the second center point of the signature image. Then, use the offset between the first center point and the second center point and the relative distance between the area boundary and the second center point to determine the signature position data corresponding to the signature. After performing edge detection calculation on the signature in the signature image, the edge information corresponding to the signature is obtained, and the signature clarity data corresponding to the signature image is determined based on the continuous parameters, intensity parameters and contrast parameters corresponding to the edge information. The signature detection result corresponding to the electronic document is determined based on the signature location data and the signature clarity data.

5. The text detection method according to claim 4, characterized in that, The process of determining the signature clarity data also includes: After performing texture detection calculation on the signature image, the texture information corresponding to the signature in the signature image is obtained, and the signature clarity data is determined based on the texture density corresponding to the texture information.

6. The text detection method according to claim 1, characterized in that, The content detection step includes: The document type data contained in the format parameters is obtained, the format type of the electronic document is determined using the document type data, and the field parameters corresponding to the format type are obtained based on the format parameters. Then, the content template corresponding to the electronic document is determined using the field parameters. After obtaining the key field corresponding to the field parameter based on the content template and obtaining the position parameter corresponding to the key field, the field display area corresponding to the key field is obtained using the position parameter, and the field content corresponding to the field in the electronic document is obtained using the field display area; The standard length of the field corresponding to the field display area is obtained using the content template, and the content detection result of the electronic document is determined based on the length difference between the field length of the field content and the standard length of the field.

7. The text detection method according to claim 6, characterized in that, After obtaining the corresponding field content in the electronic document using the field display area, the method further includes: After performing semantic analysis on the field content and the key field, a first semantic result corresponding to the field content and a second semantic result corresponding to the key field are obtained, and the content detection result of the electronic document is determined based on the semantic comparison result of the first semantic result and the second semantic result.

8. The text detection method according to claim 1, characterized in that, The detection and aggregation steps include: Obtain the first moment corresponding to the garbled character detection result, and associate the first moment with the garbled character detection result to obtain the first detection report corresponding to the garbled character detection step; Obtain the second moment corresponding to the signature detection result, and associate the second moment with the signature detection result to obtain the second detection report corresponding to the signature detection step; Obtain the third moment corresponding to the content detection result, and associate the third moment with the content detection result to obtain the third detection report corresponding to the content detection step; The text detection result is determined by summarizing the first detection report, the second detection report, and the third detection report according to the first time point, the second time point, and the third time point.

9. A text detection system, characterized in that, The system includes: The garbled text detection unit is used to acquire the electronic document and its format parameters corresponding to the text to be detected, determine the character parsing rules and character encoding rules corresponding to the electronic document through the format parameters, parse the characters contained in the electronic document using the character parsing rules, determine the character encoding result corresponding to the character using the character encoding rules, and determine the garbled text detection result of the electronic document based on the character encoding result. The signature detection unit is used to convert the electronic document into a digital image using the image format parameters included in the format parameters, obtain the signature image contained in the digital image using a preset signature template, and determine the signature detection result of the electronic document through the signature position data and signature clarity data corresponding to the signature image. The content detection unit is used to determine the content template corresponding to the electronic document using the document type data contained in the format parameters, obtain the corresponding field content in the electronic document according to the key fields contained in the content template and their corresponding position parameters, and determine the content detection result of the electronic document based on the field content. The detection and aggregation unit is used to determine the text detection result corresponding to the text to be detected based on the garbled character detection result, the signature detection result, and the content detection result.

10. A server, characterized in that, The method includes a processor and a memory, the memory storing computer-executable instructions that can be executed by the processor, the processor executing the computer-executable instructions to implement the steps of the text detection method according to any one of claims 1 to 8.