Material identification method based on checkerboard grid topological coding and intelligent pre-auditing system
Through the material recognition method based on chess grid topological encoding, intelligent identification and audit of materials are realized, and the problems of large amount of manual audits and difficult positioning of key fields in complex tables are solved, and service efficiency and audit accuracy are improved.
Patent Information
- Application Number
- CN202510723567.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-31
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2045-05-31
AI Technical Summary
In the prior art, the manual review volume during the material review process is too large, resulting in low service efficiency. It is difficult for traditional methods to accurately locate key fields in complex tables, and the time complexity is high, making it difficult to meet the real-time pre-examination needs.
The material recognition method based on chess grid topological encoding is adopted, and intelligent identification and audit of materials is realized through high-definition image acquisition, dynamic ambiguity adjustment, grayscale conversion, binary processing, checkerboard grid processing, information processing and feature encoding, intelligent comparison, material element segmentation and comprehensive quality evaluation model calculation.
It greatly reduces the amount of manual review, improves service efficiency, can accurately locate key fields in complex tables, meets real-time pre-examination needs, and improves the accuracy and reliability of material reviews.
Smart Images

Figure CN120218874A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of material identification systems, and particularly to a material identification method and an intelligent pre-trial system based on chessboard grid topology coding. Background Art
[0002] At present, the functions of traditional service terminals are relatively single, usually only able to achieve basic functions such as basic material submission and information query. In actual operation, its core review link highly depends on the manual window, that is, it is necessary for staff to manually receive materials, check information and conduct reviews. The manual review process takes a long time and is prone to omissions or judgment errors.
[0003] At present, there are also some material identification systems to assist in the review, but traditional methods rely on fixed templates or connected component analysis, and it is difficult to accurately locate key fields (such as official seals and signature columns) in complex forms (such as business licenses and application forms). Traditional methods need to perform pixel-by-pixel matching or line-by-line traversal, and the time complexity reaches The complexity of line-by-line matching is , which is difficult to meet the real-time pre-trial requirements. In addition, text, image, and structure information are processed separately, lacking a unified representation framework, resulting in low cross-modal verification efficiency. Summary of the Invention
[0004] The purpose of the present invention is to provide a material identification method and an intelligent pre-trial system based on chessboard grid topology coding, so as to solve the problem in the prior art that the manual review volume is too large during the material review process, reducing the service efficiency.
[0005] To solve the above technical problems, the first technical solution adopted by the present invention is:
[0006] A material identification method based on chessboard grid topology coding, comprising the following steps: Step S1, the user selects a matter to be handled through an operating device; Step S2, the user places the materials in the pre-audit area of the material collection device; Step S3, high-definition image collection is performed by the material collection device, and dynamic blur adjustment is performed during the collection process; Step S4, the material collection device sends the collected high-definition image to the operating device, and the operating device calculates the material tilt angle by identifying the high-definition image, corrects the high-definition image and performs affine transformation to obtain the first sample form image; Step S5, gray conversion and binarization processing are performed on the first sample form image to obtain the second sample form image; Step S6, 3*3 chessboard grid processing is performed on the second processed image to obtain the third sample form image; Step S7, informatization processing and feature coding are performed on all chessboard grids in the third sample form image to obtain the fourth sample form image; Step S8, the fourth sample form image is intelligently compared with the blank form, and the redundant blank forms are excluded to obtain the fifth sample form image; Step S9, material element segmentation is performed on the fifth sample form image; Step S10, a comprehensive quality evaluation model calculation is performed on the elements of the fifth sample form image; Step S11, intelligent verification is performed on all the collected materials through a business rule library; Step S12, the audit result is output.
[0007] A further technical solution is that in step S3, the dynamic blur adjustment during the collection process includes the following steps; Step K1, constructing a blur prediction model, , where represents the feature transformation function of the th detection index, is the defocus degree index, is a constant term, is the defocus sensitivity coefficient, is the predicted value of the target index; Step K2, constructing the hardware-image quality mapping relationship is defined as follows, the optimal focal plane distance , the real-time z-axis displacement , the defocus amount , the absolute deviation between the real-time displacement and the optimal focal plane distance, that is , which is directly used as the defocus degree index in the original model; Step K3, constructing a blur-displacement correlation model, substituting the hardware parameters into the original model to form a blur prediction model, , where is the comprehensive constant term of non-defocus factors; Step K4, constructing an optimal soft measurement model based on multi-modal information fusion, setting the multi-modal feature fusion term as the prediction model of the target index, and the basic framework is , is the th detection index, is the basic function transformation, is the weight coefficient to be obtained; Step K5, finally form the linearized eigenvector, , collect m groups of sample data, and construct the observation matrix and the target value vector , the detection index matrix, , the target value vector: , use the formula to calculate the weight coefficient, where is the weight coefficient group, is the detection index value after linearization processing, is the detection index value group after linearization processing, is the observed value of the target index, take the training set = , minimize the sum of squared prediction errors , let , , then obtain , represents the minimization operation, and the goal is to find the optimal defocus sensitivity coefficient k6; Step K6, for each model, substitute the detection index value to calculate the predicted value , is the model number, calculate the error , where is the error, is the observed value of the target index, standard deviation evaluation, , calculate the standard deviation set of all errors, where is the standard deviation set, is the number of errors, is the error, is the average error; select the model with the smallest .
[0008] A further technical solution is that the correction model for calculating the material tilt angle in the step S4 is , by detecting the boundary line of the target area, select two points and to calculate the slope, and then convert it to an angle. The corrected high-definition image can be realized through affine transformation .
[0009] A further technical solution is that in the step S5, the grayscale processing is By eliminating the color dimension RGB → single channel, the first sample table image is converted into a matrix containing only luminance information, significantly reducing the subsequent computational complexity; the binarization process is Threshold segmentation is adopted to divide the pixels of the first sample table image into white pixels and black pixels. The white pixels are represented by "0", and the black pixels are represented by "1", achieving a strict separation of the table lines, text, and blank areas. The mathematical expression is ; In step S6, to achieve efficient parsing and feature extraction of the structure of the third sample table image, for all empty tables after preprocessing, a checkerboard division operation with a 3×3 specification is uniformly performed, and finally each empty table is accurately split into 9 mutually independent and logically related checkerboards; In step S7, by calculating the proportion of black pixels , where is the total amount of black pixels, m×n is the size of the checkerboard, and it is compared with the threshold T to achieve information discrimination of the image region. When the threshold, the information entropy of this region reaches a significant level and is marked as "1", indicating that it carries sufficiently prominent structural information; otherwise, it is marked as "0" and regarded as an information-sparse region. For each checkerboard c, its local information entropy is calculated, , the larger the entropy value, the richer the information in this region, and vice versa, the information is sparse. Define the optimal threshold to maximize the global entropy, is the indicator function, the parameter value that makes the function reach the maximum value, find a threshold T to make the value of the subsequent summation term the largest, and thus determine the optimal threshold . When ; In the construction of the symbol system for row-column mode classification, coding theory is used to achieve hierarchical mapping of features. Each row of 3-bit binary combinations (000 - 111) constitutes 8 states, which are respectively mapped to different symbol sets, constructing a three-level coding system: the first row is used as the title area and uses capital letters (A - H); the second row is used as the transition area and uses lowercase letters (a - h), and the third row is used as the number area and uses numbers (0 - 7). Define the mapping function is the 3-bit binary space, .
[0010] A further technical solution is that in step S8, let the row-column codes of the sample table and the empty table be symbol sequences , where , , define the numerical mapping function as , , define the matching condition as, , the termination condition is if there exists .
[0011] A further technical solution is that in step S9, the material element segmentation step is as follows: in step T1, the difference feature map calculation is , where represents the th difference feature map; is an activation function used to enhance the non-linear expression of features, is a gradient operator, respectively represent the feature maps of different scales of the material, is a dynamic convolution kernel, where is a 3×3 Gaussian kernel used to smooth the feature map and reduce noise; is a 5×5 Laplacian kernel used to enhance the edge information of the feature map; in step T2, cross-scale feature fusion is ; represents the optimized fused feature map for subsequent material element segmentation; is the Softmax activation function; image clarity - pre-trial confidence dynamic adjustment model; is the dynamic weight and satisfies α + β = 1, is the difference feature map obtained after optimization; and are the multi-scale pooling features respectively, where is obtained by global average pooling through 1×1 convolution and can capture the global information of the material, is obtained by local maximum pooling through 7×7 convolution and can highlight the local features of the material; represents the feature concatenation operation.
[0012] A further technical solution is that in step S10, to comprehensively and accurately evaluate the accuracy and reliability of material review, the model is , the OCR recognition confidence (0 - 100%) is obtained by weighted averaging the character-level probabilities output by the OCR model after Softmax or CTC decoding, reflecting the recognition reliability of the text content, : seal / signature matching degree (%), calculated by hash similarity: = number of matching bits of total feature bits × 100%, where the feature bits come from the local features of the image or the feature vectors of the neural network, and the number of matching bits is the feature coincidence degree between the material to be reviewed and the standard template, is the accuracy of table structured parsing (%), weight assignment: , and the weight is dynamically adjusted according to the matter type.
[0013] A further technical solution is that in the step S11, the material integrity logical expression, , logical operator: represents "logical AND", and it is necessary to simultaneously satisfy "material existence" and "quantity compliance". The cross-material consistency verification formula, , function definition: is the k-th key information extraction function. When all key information matches, the consistency = 1, otherwise points are deducted according to the proportion of unmatched items.
[0014] The second technical solution adopted by the present invention is:
[0015] An intelligent pre-trial system, which is applied to the intelligent pre-trial system of a material recognition method based on checkerboard grid topology coding in the first technical solution, includes an operating device and a material collection device that are electrically connected. The main screen and the secondary screen are respectively arranged on the opposite sides of the operating device. The operating device is provided with a face recognition module, a fingerprint collection module, an ID card and RF card collection module at the edge of the main screen; the material collection device includes a material collection area, a telescopic component vertically arranged on one side of the material collection area, an installation plate horizontally arranged at the upper end of the telescopic component, and a fill light and an image collection module arranged on the lower side of the installation plate and aligned with the material collection area.
[0016] Compared with the prior art, the beneficial effects of the present invention are: the collected material images are intelligently processed, and image enhancement, noise reduction and other technologies are used to improve the image quality, providing a high-quality data basis for subsequent analysis. Through intelligent image recognition and analysis technology and content analysis technology, the content of the material images is deeply analyzed to extract key information, such as material name, specification, quantity, text content, etc. The analyzed material information is matched with the material review rules preset in the system. These rules cover review points such as material integrity, standardization, and accuracy, such as whether the application materials are complete, whether the format meets the requirements, and whether the content is true and valid. If the material information completely matches the rules, the review result is passed, and the system generates an acceptance receipt; if there is a situation that does not conform to the rules, the system automatically marks the problem points, such as material missing, filling in non-standard, etc. This greatly reduces the manual review volume and improves the service efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Figure 1 It is a flow chart of a material recognition method based on checkerboard grid topology coding of the present invention.
[0018] Figure 2 It is a checkerboard grid diagram of a material recognition method based on checkerboard grid topology coding of the present invention.
[0019] Figure 3Schematic diagram for constructing a symbol system for row and column pattern classification in a material recognition method based on chessboard grid topology coding according to the present invention.
[0020] Figure 4 Schematic diagram of Ce5 in a material recognition method based on chessboard grid topology coding according to the present invention.
[0021] Figure 5 Schematic diagram of the key chessboard squares in a material recognition method based on chessboard grid topology coding according to the present invention.
[0022] Figure 6 Schematic diagram of composite naming in a material recognition method based on chessboard grid topology coding according to the present invention.
[0023] Figure 7 Schematic diagram of the operating device of an intelligent preliminary examination system according to the present invention.
[0024] Figure 8 Schematic diagram of the material collection device of an intelligent preliminary examination system according to the present invention.
[0025] Icons: 1 - operating device, 2 - material collection device, 3 - main screen, 4 - secondary screen, 5 - face recognition module, 6 - fingerprint collection module, 7 - ID card and RF card collection module, 8 - material collection area, 9 - telescopic component, 10 - mounting plate, 11 - fill light, 12 - image collection module. Detailed implementation manners
[0026] In order to make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0027] As Figures 1 - 8 shown.
[0028] Embodiment 1:
[0029] As Figure 1As shown in the figure, a material recognition method based on chessboard grid topology coding includes the following steps: Step S1, the user selects a matter to handle through an operating device; Step S2, the user places the materials in the pre-trial area of the material collection device; Step S3, the material collection device performs high-definition image acquisition and dynamically adjusts the blur degree during the acquisition process; Step S4, the material collection device sends the acquired high-definition image to the operating device, and the operating device calculates the tilt angle of the material by identifying the high-definition image, corrects the high-definition image and performs affine transformation to obtain the first sample form image; Step S5, performs gray conversion and binarization processing on the first sample form image to obtain the second sample form image; Step S6, performs 3*3 chessboard grid processing on the second processed image to obtain the third sample form image; Step S7, performs information processing and feature coding on all chessboard grids in the third sample form image to obtain the fourth sample form image; Step S8, intelligently compares the fourth sample form image with the empty form, eliminates the redundant empty forms, and obtains the fifth sample form image; Step S9, performs material element segmentation on the fifth sample form image; Step S10, calculates the comprehensive quality evaluation model for the elements of the fifth sample form image; Step S11, performs intelligent verification on all the collected materials through the business rule library; Step S12, outputs the audit result.
[0030] In step S3, high-definition image acquisition is performed by the material collection device, and the blur degree is dynamically adjusted during the acquisition process, which specifically includes the following steps;
[0031] 1. Construct a blur prediction model:
[0032] 。
[0033] Among them, represents the feature transformation function of the th detection index. It includes the composition of basic functions such as linear transformation, power function, and logarithmic function.
[0034] is the defocus degree index. In the intelligent pre-trial system, the defocus degree is determined by the deviation between the moving distance of the telescopic device along the Z axis and the distance from the optimal focal plane.
[0035] Define the hardware-image quality mapping relationship as follows:
[0036] (1) Optimal focal plane distance : The ideal distance between the material plane and the optical center of the lens. At this time, the response value of the modulation transfer function (MTF) of the imaging system in the high-frequency region is the largest.
[0037] (2) Real-time z-axis displacement ( ) : The actual distance between the material plane and the lens in real-time feedback by the telescopic device.
[0038] (3) Defocus amount ( ): The absolute deviation between the real-time displacement and the distance of the optimal focal plane, that is , which is directly used as the defocus degree index in the original model .
[0039] 2. Construct the ambiguity-displacement correlation model:
[0040] Substitute the hardware parameters into the original model to form a special ambiguity prediction model for intelligent preliminary review: .
[0041] Among them, is the comprehensive constant term of non-defocus factors, including the influence of fixed detection indexes such as resolution, edge sharpness, and contrast; is the defocus sensitivity coefficient, which represents the contribution degree of unit defocus amount to the ambiguity (unit: ambiguity value / mm).
[0042] 3. Construct the optimal soft sensor model based on multi-modal information fusion:
[0043] (I) Construct the basic framework of the soft sensor model:
[0044] Set the multi-modal feature fusion term as the prediction model of the target index (such as ambiguity), and the basic framework is: .
[0045] : Predicted value of the target index;
[0046] : The th detection index (such as resolution , edge sharpness , contrast , etc.);
[0047] : Basic function transformation (such as power function, logarithmic function, exponential function, etc.);
[0048] : Weight coefficient to be solved.
[0049] (II) Mapping of multi-modal features and basic functions:
[0050] Define 5 detection indexes and their basic function transformations, as shown in Table 1:
[0051]
[0052] Finally, form a linearized feature vector: .
[0053] Collect m groups of sample data to construct the observation matrix and the target value vector .
[0054] Detection index matrix: .
[0055] Target value vector: .
[0056] Using the formula:
[0057] .
[0058] Calculate the weight coefficient, where K* is the weight coefficient group, is the detection index value after linearization processing, is the detection index value group after linearization processing, is the observed value of the target index.
[0059] Take the training set = .
[0060] Minimize the sum of squared prediction errors , represents the minimization operation, and the goal is to find the optimal defocus sensitivity coefficient . If is large, it indicates that the defocus amount has a significant impact on the blur. It is necessary to reduce Δd through hardware adjustment (such as the displacement of the telescopic device) to improve the image clarity. If is small, it indicates that the image blur is mainly caused by non-defocus factors (such as insufficient resolution). It is necessary to optimize from other parameters (such as fill light, resolution).
[0061] Let , = , then obtain .
[0062] 4. Select the optimal model by standard deviation:
[0063] (1) Prediction value calculation: For each model, substitute the detection index value to calculate the prediction value (j is the model number);
[0064] (2) Calculate the error .
[0065] Calculate the error matrix, where is the error, is the observed value of the target index, is the predicted value of the target index.
[0066] (3) Standard deviation evaluation: .
[0067] Calculate the standard deviation set of all errors, where is the standard deviation set, is the number of errors, is the error, is the average error; select the model with the smallest value and its weight coefficients as the final solution.
[0068] .
[0069] In step S4, the material acquisition device sends the acquired high-definition image to the operation device. The operation device calculates the material tilt angle by recognizing the high-definition image, corrects the high-definition image and performs an affine transformation to obtain the first sample form image;
[0070] The correction model for calculating the material tilt angle is
[0071]
[0072] Physical meaning:
[0073] The tilt angle of the material image (such as the rotation angle during the scanning of an ID card or business license), used for image correction.
[0074] Calculation logic:
[0075] By detecting the boundary line of the target area (such as the edge of an ID card), select two points and calculate the slope and then convert it to an angle.
[0076] The corrected image can be achieved through an affine transformation to obtain the first sample form image. cv2.getRotationMatrix2D() is a function in OpenCV used to create a rotation matrix.
[0077] In step S5, perform grayscale conversion and binarization processing on the first sample form image to obtain the second sample form image;
[0078] 1. Image preprocessing:
[0079] (1) Grayscale conversion and binarization processing:
[0080] Grayscale processing By eliminating the color dimension RGB→single channel, convert the table image into a matrix containing only luminance information, significantly reducing the subsequent computational complexity.
[0081] Binarization By using threshold segmentation, the pixels of the first sample form are divided into white pixels and black pixels. The white pixels are represented by "0", and the black pixels are represented by "1", so as to strictly separate the table lines, text and blank areas, and obtain the second sample form image. The mathematical expression is:
[0082] .
[0083] Among them, is the binarization threshold.
[0084] Step S6: Perform 3×3 chessboard grid processing on the second processed image to obtain the third sample form image;
[0085] Chessboard grid processing:
[0086] To achieve efficient parsing and feature extraction of the table structure, for all empty tables after preprocessing, a chessboard division operation of 3×3 specification is uniformly performed, and finally each empty table is accurately split into 9 mutually independent and logically related chessboard grids. Each chessboard grid corresponds to a logical area of the table (such as the title area, data area, dividing line area), forming a hierarchical representation system of "local feature global layout", forming a first-level topology, as Figure 2 shown, to obtain the third sample form image.
[0087] Step S7: Perform information processing and feature encoding on all chessboard grids in the third sample form image to obtain the fourth sample form image;
[0088] Chessboard information processing:
[0089] (1) Pixel ratio judgment:
[0090] In the theoretical system of chessboard information processing, pixel ratio judgment is a typical application of image binarization threshold segmentation. By calculating the ratio of black pixels is the total amount of black pixels, m×n is the size of the chessboard grid), and comparing it with the threshold T, the information discrimination of the image area is realized. From the perspective of information theory, when the threshold, the information entropy of this area reaches a significant level and is marked as "1", indicating that it carries sufficiently prominent structural information; otherwise, it is marked as "0" and regarded as an information-sparse area. Among them, for each chessboard grid c (such as a 3×3 sub-region), its local information entropy is calculated.
[0091] .
[0092] The larger the entropy value, the richer the information in this area (such as the text-dense area), and vice versa, the sparser the information (such as the blank area). Define the optimal threshold to maximize the global entropy:
[0093]
[0094] 。
[0095] (2) Construction of the symbol system for row-column pattern classification:
[0096] In the construction of the symbol system for row-column pattern classification, coding theory is used to achieve hierarchical mapping of features. Each row of 3-bit binary combinations (000∼111) constitutes 8 states, which are respectively mapped to different symbol sets to construct a three-level coding system: The first row serves as the title area, using capital letters (A−H), and leveraging their visual salience to strengthen the dominant position in the information structure; The second row serves as the transition area, using lowercase letters (a−h), forming a transition buffer at the visual level; The third row serves as the digital area, using numbers (0−7), facilitating the intuitive expression of numerical information. This symbol system follows the hierarchical semantic expression principle of semiotics. Through the hierarchical differentiation of visual symbols, it improves the readability and recognizability of information, realizes the cognitive conversion from binary coding to semantic symbols, and defines the mapping function as a 3-bit binary space, 。
[0097] Title area mapping (capital letters): , corresponding rule:
[0098] b 1 b 2 b 3→chr(65+bin2dec( b 1 b 2 b 3)) For example, 101→chr(65+5)=F.
[0099] Transition area mapping (lowercase letters): , corresponding rule:
[0100] b 1 b 2 b 3→chr(97+bin2dec( b 1 b 2 b 3)) For example, 011→chr(97+3)= g 。
[0101] Digital area mapping (numerical symbols): , corresponding rule:
[0102] b 1 b 2 b 3→bin2dec( b 1 b 2b 3) For example, 101 → 6. As Figure 3 shown.
[0103] In summary, as Figure 4 shown, it can be named Ce5, where bin2dec() is a function that converts a 3-bit binary number to a decimal number and is used to implement the mapping from binary encoding to the symbol set, and chr() is a character conversion function that converts a decimal ASCII code value to the corresponding character.
[0104] If there is one sample table corresponding to multiple empty tables, we screen out the key checkerboards in the sample table for further 3*3 checkerboard division to form a secondary topological structure. The key checkerboards are as Figure 5 shown:
[0105] By increasing the threshold, the checkerboard information is more finely distinguished.
[0106] Perform checkerboard informatization processing on the 3*3 checkerboards in the key checkerboards of the sample table, and add a separator for compound naming on the basis of the original naming, as Figure 6 shown:
[0107] In summary, the compound naming of this sample table is Ce5-Ec4-Db2-Cc1. Form a compound code , and obtain the fourth sample table image.
[0108] By visually differentiating between uppercase and lowercase letters and numbers, construct a semantic coding system corresponding to "hierarchy - function", so that the number itself has self-explanatory properties (e.g., in "Gc7", G → first row mode, c → middle row mode, 7 → last row mode).
[0109] Step S8, perform intelligent comparison between the fourth sample table image and the empty table, eliminate the redundant empty tables, and obtain the fifth sample table image;
[0110] Principle of high efficiency of Ascii coding comparison:
[0111] Let the row and column encodings of the sample table and the empty table be symbol sequences .
[0112] Among them , define the numerical mapping function as:
[0113] .
[0114] This mapping satisfies injectivity, ensuring a one-to-one correspondence between characters and numerical values. Among them, define: map the ASCII character to the corresponding positive integer value. Comparison rule: adopt the numerical value in lexicographical order to implement, and define the matching condition as
[0115] 。
[0116] Termination condition: If there exists to obtain the fifth sample table image, where It means: Compare the ASCII values of the characters at each position of S and E in sequence.
[0117] Step S9: Segment the material elements of the fifth sample table image;
[0118] 1. Difference feature map calculation:
[0119] 。
[0120] where represents the th difference feature map; is an activation function used to enhance the non - linear expression of features; is a gradient operator. By calculating the gradient of the feature map, it can highlight the edge contour information of elements such as official seals and handwritten characters in the material. Compared with the direct difference calculation of the original formula, it can capture detailed features more effectively; respectively represent the feature maps of different scales of the material. is a dynamic convolution kernel, where is a 3×3 Gaussian kernel used to smooth the feature map and reduce noise; is a 5×5 Laplacian kernel used to enhance the edge information of the feature map. The use of this dynamic convolution kernel can perform adaptive feature extraction according to the characteristics of different - scale features.
[0121] 2. Cross - scale feature fusion:
[0122] 。
[0123] represents the optimized fused feature map for subsequent material element segmentation; is the Softmax activation function, which can normalize the fused features, enhance the non - linear expression of features, and improve the segmentation accuracy for irregular elements (such as handwritten annotations); Image clarity - pre - review confidence dynamic adjustment model; is the dynamic weight, and α + β = 1. These two weights can be adaptively adjusted according to different types of materials (such as business licenses, medical bills, etc.) to better balance the contributions of different features; is the difference feature map obtained after the previous optimization; are the multi - scale pooling features respectively, where It is obtained through global average pooling with 1×1 convolution and can capture the global information of the material; It is obtained through local maximum pooling with 7×7 convolution and can highlight the local features of the material; Represents the feature concatenation operation, which concatenates different feature maps in the channel dimension. Compared with the multiplication operation in the original formula, it can retain more feature information.
[0124] Step S10, perform calculations using the comprehensive quality assessment model on the elements of the fifth sample form image;
[0125] Achieve a comprehensive and accurate assessment of the accuracy and reliability of material review. The model is:
[0126] .
[0127] : Comprehensive quality assessment of the elements of the fifth sample form image.
[0128] : OCR recognition confidence (0 - 100%), obtained by weighted averaging the character-level probabilities output by the OCR model after Softmax or CTC decoding, reflecting the recognition reliability of the text content.
[0129] : Seal / signature matching degree (%), calculated through hash similarity: = (Number of matching bits / Total number of feature bits) × 100%.
[0130] Among them, the number of feature bits comes from the local features of the image (such as SIFT / SURF key point descriptors) or the feature vectors of the neural network (such as the deep features extracted by CNN), and the number of matching bits is the feature coincidence degree between the material to be reviewed and the standard template.
[0131] : Table structure parsing accuracy rate (%), calculated by detecting indicators such as table row and column alignment and field integrity, for example: = (1 - (Number of missing fields + Number of misaligned fields) / Total number of fields) × 100%.
[0132] are all weight distribution coefficients, and , The weight values are dynamically adjusted according to the matter type.
[0133] Step S11, perform intelligent verification on all the collected materials through the business rule library;
[0134] 1. Material integrity logical expression:
[0135] .
[0136] Logical operator: Represents "logical AND", indicating that all items from i = 1 to n must satisfy the conditions simultaneously, and both "material existence" and "quantity compliance" need to be satisfied simultaneously;
[0137] : For each material item in all the materials submitted by the user, determine whether the material item submitted by the user belongs to the material items listed in the list.
[0138] : For the quantity of each material item to be submitted, determine whether the quantity of each material item meets the quantity requirements of the material items listed in the list and whether the minimum quantity requirement is satisfied.
[0139] Example: For enterprise registration, it is required to have "Business License (1 copy) + Articles of Association (3 copies)". If either is missing or the quantity is insufficient, the integrity = false.
[0140] 2. Cross - material consistency verification formula:
[0141]
[0142] Function definition: is the key information extraction function for the k - th item (such as "name", "ID number", "valid period"); is the product operator, calculating the product of all items from k = 1 to p.
[0143] When all key information matches, the consistency = 1; otherwise, points are deducted according to the proportion of unmatched items (for example, if "the name is consistent but the address is deviated", then the consistency = 0.8).
[0144] In step S12, the audit result is output to generate a receipt or problem annotation.
[0145] Embodiment 2:
[0146] Such as Figure 7 and Figure 8 shown, an intelligent pre - examination system includes an operating device 1 and a material collection device 2 that are electrically connected. On opposite sides of the operating device, a main screen 3 and a secondary screen 4 are respectively provided. At the edge of the main screen 3 of the operating device 1, a face recognition module 5, a fingerprint collection module 6, an ID card and RF card collection module 7 are provided; the material collection device 2 includes a material collection area 8, a telescopic component 9 vertically arranged on one side of the material collection area, an installation plate 10 horizontally arranged at the upper end of the telescopic component 9, and a supplementary light 11 and an image collection module 12 that are arranged on the lower side of the installation plate 10 and are aligned with the material collection area 8.
[0147] Although the present invention has been described herein with reference to a number of illustrative embodiments, it should be understood that those skilled in the art can devise many other modifications and embodiments that will fall within the scope and spirit of the principles disclosed in this application. More specifically, within the scope of the present application disclosure, the accompanying drawings, and the claims, various variations and improvements can be made to the components and / or layout of the subject combination layout. In addition to the variations and improvements made to the components and / or layout, other uses will also be apparent to those skilled in the art.
Claims
1. A material recognition method based on chessboard grid topology coding, characterized in that, It includes the following steps. Step S1, the user selects the matter to be handled through the operating device; Step S2, the user places the materials in the pre-audit area of the material collection device; Step S3, high-definition image collection is performed through the material collection device, and dynamic blur adjustment is performed during the collection process; Step S4, the material collection device sends the collected high-definition image to the operating device. The operating device calculates the material tilt angle by identifying the high-definition image, corrects the high-definition image and performs affine transformation to obtain the first sample form image; Step S5, performs gray conversion and binarization processing on the first sample form image to obtain the second sample form image; Step S6, performs 3*3 checkerboard grid processing on the second processed image to obtain the third sample form image; Step S7, performs informatization processing and feature encoding on all checkerboard grids in the third sample form image to obtain the fourth sample form image; Step S8, performs intelligent comparison between the fourth sample form image and the blank form, eliminates the redundant blank forms to obtain the fifth sample form image; Step S9, performs material element segmentation on the fifth sample form image; Step S10, calculates the comprehensive quality evaluation model for the elements of the fifth sample form image; Step S11, performs intelligent verification on all the collected materials through the business rule library; Step S12, outputs the audit result.
2. The material identification method based on chessboard grid topology coding according to claim 1, characterized in that: In step S3, the dynamic blur adjustment during the collection process includes the following steps; Step K1, construct a blur prediction model, where represents the feature transformation function of the th detection index, is the defocus degree index, is the constant term, is the defocus sensitivity coefficient, is the predicted value of the target index; Step K2, construct the hardware-image quality mapping relationship defined as follows, the optimal focal plane distance , the real-time z-axis displacement , the defocus amount , the absolute deviation between the real-time displacement and the optimal focal plane distance, i.e., , which is directly used as the defocus degree index in the original model ; Step K3, construct a blur-displacement correlation model, substitute the hardware parameters into the original model to form a blur prediction model, , where is the comprehensive constant term of non-defocus factors; Step K4, construct an optimal soft sensor model based on multi-modal information fusion, set the multi-modal feature fusion term as the prediction model of the target index, and the basic framework is , is the th detection index, is the basic function transformation, is the weight coefficient to be solved; Step K5, finally form a linearized feature vector, , collect m groups of sample data, construct the observation matrix and the target value vector , the detection index matrix, , the target value vector: , use the formula to calculate the weight coefficient, where is the weight coefficient group, is the detection index value after linearization, is the detection index value group after linearization, is the observed value of the target index, take the training set = , minimize the sum of squared prediction errors , let, , , then = , obtain , Indicates a minimization operation, with the goal of finding the optimal defocus sensitivity coefficient k6; Step K6, for each model, substitute the detection index value to calculate the predicted value , is the model number, calculate the error , where is the error, is the observed value of the target index, evaluated by standard deviation, , calculate the standard deviation set of all errors, where is the standard deviation set, is the number of errors, is the error, is the average error; Select the model with the smallest .
3. A material identification method based on chessboard grid topology coding according to claim 2, characterized in that: The calibration model for calculating the tilt angle of the material in step S4 is , by detecting the boundary line of the target area, two points are selected and to calculate the slope and then convert it into an angle. The calibrated high-definition image can be achieved through affine transformation to realize.
4. The material recognition method based on chessboard grid topology coding according to claim 3, characterized in that: In the step S5, the grayscale processing is By eliminating the color dimension RGB → single channel, the first sample table image is converted into a matrix containing only luminance information, significantly reducing the subsequent computational complexity: The binarization processing is Threshold segmentation is adopted to divide the pixels of the first sample table image into white pixels and black pixels. The white pixels are represented by "0", and the black pixels are represented by "1", realizing a strict separation of table lines, text, and blank areas. The mathematical expression is : In step S6, to achieve efficient parsing and feature extraction of the structure of the third sample table image, for all empty tables after preprocessing, a checkerboard division operation with a 3×3 specification is uniformly performed, and finally each empty table is accurately split into 9 independent and logically related checkerboards: In step S7, by calculating the proportion of black pixels , where is the total amount of black pixels, m×n is the size of the checkerboard, and it is compared with the threshold T to realize the information discrimination of the image area. When the threshold, the information entropy of this area reaches a significant level and is marked as "1", indicating that it carries sufficiently prominent structural information; otherwise, it is marked as "0" and regarded as an information-sparse area. For each checkerboard c, its local information entropy is calculated, , the larger the entropy value, the richer the information in this area, and vice versa, the information is sparse. Define the optimal threshold to maximize the global entropy, , is an indicator function. When ; In the construction of the symbol system for row-column mode classification, coding theory is used to realize the hierarchical mapping of features. Each row of 3-bit binary combinations (000 ∼ 111) constitutes 8 states, which are respectively mapped to different symbol sets to construct a three-level coding system: The first row is used as the title area and uses capital letters (A−H); the second row is used as the transition area and uses lowercase letters (a−h), and the third row is used as the digital area and uses numbers (0−7). Define the mapping function , where is a 3-bit binary space, is a composite symbol set.
5. A material recognition method based on chessboard grid topology coding according to claim 4, characterized in that: In the step S8, the row and column encodings of the sample table and the empty table are respectively symbol sequences and , where , , define the numerical mapping function as , , define the matching condition as , and the termination condition is that if there exists , it is determined as a mismatch and the comparison is terminated.
6. The material recognition method based on chessboard grid topology coding according to claim 5, characterized in that: In the step S9, the material element segmentation step is as follows: in step T1, the difference feature map is calculated. , where represents the -th difference feature map; is an activation function used to enhance the non-linear expression of features. is a gradient operator. and represent the feature maps of different scales of the material respectively. and are dynamic convolution kernels. Among them, is a 3×3 Gaussian kernel used to smooth the feature map and reduce noise; is a 5×5 Laplacian kernel used to enhance the edge information of the feature map. In step T2, cross-scale feature fusion is performed. ; represents the optimized fused feature map for subsequent material element segmentation; is the Softmax activation function; the image clarity - pre-trial confidence dynamic adjustment model. is the dynamic weight and satisfies , and are the difference feature maps obtained after optimization; and are the multi-scale pooling features respectively. Among them, is obtained by global average pooling through 1×1 convolution and can capture the global information of the material. is obtained by local maximum pooling through 7×7 convolution and can highlight the local features of the material; represents the feature concatenation operation.
7. A material recognition method based on chessboard grid topology coding according to claim 6, characterized in that: In the step S10, a comprehensive and accurate evaluation of the accuracy and reliability of material review is achieved. The model is , : OCR recognition confidence (0 - 100%), which is obtained by weighted averaging the character-level probabilities output by the OCR model after Softmax or CTC decoding, reflecting the recognition reliability of the text content. : Seal / signature matching degree (%), calculated by hash similarity: = (Number of matching feature bits / Total number of feature bits) × 100%, where the number of feature bits comes from local image features or neural network feature vectors, and the number of matching bits is the feature coincidence degree between the material to be reviewed and the standard template. is the accuracy rate of table structured parsing (%), and the weight assignment: , and the weight is dynamically adjusted according to the matter type.
8. A method for material identification based on chessboard grid topology coding according to claim 7, characterized in that: In the step S11, the material integrity logical expression, , logical operator: represents "logical AND", and both "material existence" and "quantity compliance" need to be satisfied simultaneously. The cross-material consistency verification formula, , function definition: is the k-th key information extraction function. When all key information matches, the consistency = 1; otherwise, points are deducted according to the proportion of unmatched items.
9. An intelligent pre-trial system, characterized in that, It includes the operating device and the material collection device that are electrically connected. A main screen and a secondary screen are respectively arranged on the opposite sides of the operating device. A face recognition module, a fingerprint collection module, an ID card and RF card collection module are arranged on the edge of the main screen of the operating device; The material collection device includes a material collection area, a telescopic component vertically arranged on one side of the material collection area. An installation plate is horizontally arranged at the upper end of the telescopic component. A fill light and an image collection module that are aligned with the material collection area are arranged on the lower side of the installation plate.
Citation Information
Patent Citations
Self-service all-in-one machine based on face detection and character recognition and using method thereof
CN105469513A
Intelligent government administration service working system and application thereof
CN110322643A
Robot camera calibration method based on edge scale adaptive defocusing fuzzy estimation
CN112950723A
Batch fuzzy identifier reconstruction method based on deformable convolution
CN116796773A
Image definition recognition method and device and storage medium
CN117409215A