A material identification method and intelligent pre-screening system based on chessboard grid topology coding
Through the material identification method based on chess grid topological encoding, the problem of time-consuming and low accuracy of traditional material audit systems is solved, intelligent material audits are realized, and service efficiency and accuracy are improved.
Patent Information
- Application Number
- CN202510723567.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-31
- Publication Date
- 2025-08-29
- Estimated Expiration
- 2045-05-31
AI Technical Summary
Traditional materials audit systems rely on manual audits, which are time-consuming and error-prone, making them difficult to meet the needs of real-time pre-examination. Traditional methods are difficult to accurately locate key fields when processing complex tables, and the cross-modal verification efficiency is low.
The material recognition method based on chess grid topological encoding is adopted, and material element segmentation and intelligent verification are achieved through high-definition image acquisition, ambiguity adjustment, image correction, chess grid processing and feature encoding, combined with intelligent comparison and comprehensive quality evaluation models.
Significantly reduce the amount of manual review, improve service efficiency, improve the accuracy and audit accuracy of material information analysis, and meet the needs of real-time pre-examination.
Smart Images

Figure CN120218874B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of material identification systems, and in particular to a material identification method and an intelligent pre-screening system based on chessboard grid topology coding. Background Art
[0002] At present, the functions of traditional service terminals are relatively simple, and they can usually only realize basic functions such as basic material submission and information query. In actual operation, its core review link is highly dependent on manual windows, that is, staff are required to manually receive materials, check information and conduct reviews. The manual review process is time-consuming and prone to omissions or misjudgments.
[0003] At present, some materials can only be identified by the system to assist in the review. However, the traditional method relies on fixed templates or connected domain analysis, which makes it difficult to accurately locate key fields (such as official seals and signature columns) in complex forms (such as business licenses and application forms). The traditional method requires pixel-by-pixel matching or line-by-line traversal, which takes a long time. , the row-by-row matching complexity is , it is difficult to meet the real-time pre-examination requirements. In addition, the text, image, and structural information are processed separately and lack a unified representation framework, resulting in low cross-modal verification efficiency. Summary of the Invention
[0004] The purpose of the present invention is to provide a material identification method and intelligent pre-examination system based on chessboard grid topology coding, so as to solve the problem in the prior art that the manual review process of materials is too large, which reduces service efficiency.
[0005] In order to solve the above technical problems, the first technical solution adopted by the present invention is:
[0006] A material identification method based on chessboard grid topology coding includes the following steps: step S1, a user selects an item to be handled through an operating device; step S2, the user places the material in a pre-examination area of a material collection device; step S3, high-definition image collection is performed through the material collection device, and the fuzziness is dynamically adjusted during the collection process; step S4, the material collection device sends the collected high-definition image to the operating device, and the operating device calculates the material inclination angle by identifying the high-definition image, corrects the high-definition image and performs affine transformation to obtain a first sample image; step S5, grayscale conversion and binarization processing are performed on the first sample image to obtain a second sample table image; step S6, performing 3*3 chessboard gridding processing on the second processed image to obtain a third sample table image; step S7, performing information processing and feature coding on all chessboard grids in the third sample table image to obtain a fourth sample table image; step S8, performing intelligent comparison between the fourth sample table image and the blank table, eliminating unnecessary blank tables, and obtaining a fifth sample table image; step S9, performing material element segmentation on the fifth sample table image; step S10, performing comprehensive quality assessment model calculation on the elements of the fifth sample table image; step S11, performing intelligent verification on all collected materials through the business rule library; step S12, outputting the audit results.
[0007] A further technical solution is that in step S3, the dynamic adjustment of ambiguity during the acquisition process includes the following steps: step K1, constructing an ambiguity prediction model, ,in, Indicates the The characteristic transformation function of the detection index, is the defocus degree indicator, is a constant term, is the defocus sensitivity coefficient, is the predicted value of the target indicator; Step K2, construct the hardware-image quality mapping relationship defined as follows, the optimal focal plane distance , real-time z-axis displacement , defocus amount , the absolute deviation between the real-time displacement and the optimal focal plane distance, that is , directly used as the defocus index in the original model Step K3, construct the ambiguity-displacement correlation model, substitute the hardware parameters into the original model to form the ambiguity prediction model, ,in, is the comprehensive constant term of non-defocusing factors; Step K4, construct the optimal soft measurement model based on multimodal information fusion, set the multimodal feature fusion term as the prediction model of the target indicator, the basic framework is , For the Detection indicators, is the basic function transformation, is the weight coefficient to be determined; step K5, finally forms the linearized eigenvector, , collect m groups of sample data and build the observation matrix and the target value vector , detection indicator matrix, , target value vector: , using the formula Calculate the weight coefficient, where is the weight coefficient group, is the detection index value after linearization processing, is the detection index value group after linearization processing, is the observed value of the target indicator, take the training set = , minimize the sum of squared prediction errors ,make , ,but Obtain , Represents a minimization operation, the goal is to find the optimal defocus sensitivity coefficient k6; Step K6, for each model, the detection index value is brought in to calculate the predicted value , is the model number, calculation error ,in is the error, is the observed value of the target indicator, the standard deviation is evaluated, , calculate the set of standard deviations of all errors, where is the set of standard deviations, is the number of errors, is the error, is the error mean; select The smallest model and its weight coefficient are taken as the final solution. .
[0008] A further technical solution is that the correction model for calculating the material tilt angle in step S4 is , by detecting the boundary line of the target area, select two points and Calculate the slope and convert it into an angle. The corrected high-definition image can be transformed by affine transformation. accomplish.
[0009] A further technical solution is that in step S5, the grayscale processing is By eliminating the color dimension RGB→single channel, the first sample image is converted into a matrix containing only brightness information, which significantly reduces the complexity of subsequent calculations; the binarization process is Threshold segmentation is used to divide the pixels of the first sample image into white pixels and black pixels. White pixels are represented by "0" and black pixels are represented by "1". The strict separation of table lines, text and blank areas is achieved. The mathematical expression is: In step S6, in order to achieve efficient analysis and feature extraction of the image structure of the third sample table, a 3×3 chessboard division operation is uniformly performed on all the pre-processed empty tables, and each empty table is finally accurately divided into 9 independent and logically related chessboard grids; in step S7, the black pixel ratio is calculated ,in is the total amount of black pixels, m×n is the chessboard size, and is compared with the threshold T to realize the information discrimination of the image area. When the threshold is reached, the information entropy of the region reaches a significant level and is marked as "1", indicating that it carries sufficiently prominent structural information; otherwise, it is marked as "0", which is considered an information sparse region. For each checkerboard cell c, its local information entropy is calculated. The larger the entropy value, the richer the information in the area, and vice versa, the information is sparse. The optimal threshold is defined as Maximize the global entropy, is the indicator function, The parameter value that makes the function reach the maximum value is used to find a threshold T so that the value of the subsequent summation item is maximized, thereby determining the optimal threshold value. .when In the construction of the symbol system for row and column pattern classification, the coding theory is used to realize the hierarchical mapping of features. Each row of 3-bit binary combinations (000-111) constitutes 8 states, which are mapped to different symbol sets respectively, to build a three-level coding system: the first row is used as the title area, using uppercase letters (AH); the second row is used as the transition area, using lowercase letters (ah); the third row is used as the number area, using numbers (0-7), and the mapping function is defined. is a 3-bit binary space, .
[0010] A further technical solution is that in step S8, the row and column codes of the sample table and the empty table are symbol sequences ,in , , define the numerical mapping function as , , define the matching conditions as, , the termination condition is if there is .
[0011] A further technical solution is that in step S9, the material element segmentation step is as follows: step T1, difference feature map calculation, ,in To indicate the A difference feature map; is an activation function used to enhance the nonlinear expression of features. is the gradient operator, Represent the characteristic maps of different scales of materials, is a dynamic convolution kernel, where is a 3×3 Gaussian kernel used to smooth the feature map and reduce noise; is a 5×5 Laplacian kernel, which is used to enhance the edge information of the feature map; step T2, cross-scale feature fusion, ; Represents the optimized fusion feature map, which is used for subsequent material element segmentation; Softmax activation function; image clarity-pre-screening confidence dynamic adjustment model; is a dynamic weight and satisfies α+β=1, It is the difference feature map obtained after optimization; and They are multi-scale pooling features, It is obtained by global average pooling of 1×1 convolution, which can capture the global information of the material. It is obtained by local maximum pooling of 7×7 convolution, which can highlight the local features of the material; Represents a feature concatenation operation.
[0012] A further technical solution is that in step S10, a comprehensive and accurate assessment of the accuracy and reliability of the material review is achieved. ,OCR recognition confidence (0-100%), obtained by weighted average of the character-level probability output by the OCR model after Softmax or CTC decoding, reflecting the recognition reliability of the text content, : Seal / signature matching degree (%), calculated by hash similarity: = Total number of feature bits matched × 100%, where the number of feature bits comes from local features of the image or neural network feature vectors, and the number of matching bits is the degree of feature overlap between the material to be reviewed and the standard template. The accuracy of table structured parsing (%), weight distribution: , the weight is dynamically adjusted according to the type of matter.
[0013] A further technical solution is that in step S11, the material integrity logical expression is: , logical operators: Represents "logical AND", which requires both "material existence" and "number of copies to be compliant" to be met. Cross-material consistency verification formula, , function definition: The kth key information extraction function is used. When all key information matches, consistency = 1. Otherwise, points are deducted according to the proportion of mismatched items.
[0014] The second technical solution adopted by the present invention is:
[0015] An intelligent pre-screening system is applied to an intelligent pre-screening system of a material identification method based on chessboard grid topology coding in the first technical solution, comprising an electrically connected operating device and a material collection device, wherein a main screen and a secondary screen are respectively provided on opposite sides of the operating device, and a facial recognition module, a fingerprint collection module, and an ID card and RF card collection module are provided on the edge of the main screen of the operating device; the material collection device comprises a material collection area, a telescopic component vertically arranged on one side of the material collection area, a mounting plate horizontally arranged on the upper end of the telescopic component, and a fill light and an image collection module aimed at the material collection area are provided on the lower side of the mounting plate.
[0016] Compared with the existing technology, the beneficial effects of the present invention are: intelligent processing of the collected material images, application of image enhancement, noise reduction and other technologies to improve image quality, providing a high-quality data foundation for subsequent analysis, and in-depth analysis of the material image content through intelligent image recognition and analysis technology and content analysis technology, extracting key information such as material name, specifications, quantity, text content, etc. The parsed material information is matched with the material review rules preset in the system. These rules cover the integrity, standardization, accuracy and other review points of the materials, such as whether the application materials are complete, whether the format meets the requirements, whether the content is true and valid, etc. If the material information fully matches the rules, the review result is passed and the system generates an acceptance receipt; if there is a situation that does not comply with the rules, the system automatically marks the problem points, such as missing materials, non-standard filling, etc., which greatly reduces the amount of manual review and improves service efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Figure 1 The figure is a flow chart of a material identification method based on chessboard grid topology coding according to the present invention.
[0018] Figure 2 A schematic diagram of chessboard gridding of a material identification method based on chessboard grid topology coding according to the present invention.
[0019] Figure 3A schematic diagram of the symbol system construction for row and column pattern classification of a material identification method based on chessboard grid topology coding according to the present invention.
[0020] Figure 4 Schematic diagram of Ce5 of a material identification method based on chessboard grid topology coding according to the present invention.
[0021] Figure 5 This is a schematic diagram of the key checkerboard grids of a material identification method based on checkerboard grid topology coding according to the present invention.
[0022] Figure 6 This is a composite naming schematic diagram of a material identification method based on chessboard grid topology coding according to the present invention.
[0023] Figure 7 The figure is a schematic diagram of an operating device of an intelligent pre-screening system of the present invention.
[0024] Figure 8 This is a schematic diagram of a material collection device for an intelligent pre-screening system of the present invention.
[0025] Icons: 1-operating device, 2-material collection device, 3-main screen, 4-secondary screen, 5-facial recognition module, 6-fingerprint collection module, 7-ID card and RF card collection module, 8-material collection area, 9-telescopic component, 10-mounting plate, 11-fill light, 12-image acquisition module. DETAILED DESCRIPTION
[0026] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.
[0027] like Figures 1-8 shown.
[0028] Example 1:
[0029] like Figure 1As shown, a material identification method based on chessboard grid topology coding includes the following steps: step S1, the user selects the task to be handled through the operating device; step S2, the user places the material in the pre-examination area of the material collection device; step S3, the material collection device collects high-definition images, and dynamically adjusts the fuzziness during the collection process; step S4, the material collection device sends the collected high-definition image to the operating device, and the operating device calculates the material inclination angle by identifying the high-definition image, corrects the high-definition image and performs affine transformation to obtain a first sample image; step S5, grayscale conversion and binarization processing are performed on the first sample image. Obtain a second sample table image; Step S6, perform 3*3 chessboard gridding processing on the second processed image to obtain a third sample table image; Step S7, perform information processing and feature encoding on all chessboard grids in the third sample table image to obtain a fourth sample table image; Step S8, perform intelligent comparison between the fourth sample table image and the blank table, eliminate unnecessary blank tables, and obtain a fifth sample table image; Step S9, perform material element segmentation on the fifth sample table image; Step S10, perform comprehensive quality assessment model calculation on the elements of the fifth sample table image; Step S11, perform intelligent verification on all collected materials through the business rule library; Step S12, output the audit results.
[0030] In step S3, high-definition image acquisition is performed by a material acquisition device, and the fuzziness is dynamically adjusted during the acquisition process, which specifically includes the following steps:
[0031] 1. Build a fuzzy prediction model:
[0032] .
[0033] in, Indicates the Characteristic transformation functions of the detection indicators, including linear transformation, power function, logarithmic function and other basic function combinations.
[0034] It is an indicator of the degree of defocus. In the intelligent pre-screening system, the degree of defocus is determined by the deviation between the Z-axis movement distance of the telescopic device and the distance from the optimal focal plane.
[0035] The hardware-image quality mapping relationship is defined as follows:
[0036] (1) Optimal focal plane distance : The ideal distance between the material plane and the optical center of the lens, at which the imaging system modulation transfer function (MTF) has the maximum response value in the high-frequency region.
[0037] (2) Real-time z-axis displacement ( ): The actual distance between the material plane and the lens is fed back in real time by the telescopic device.
[0038] (3) Defocus amount ( ): The absolute deviation between the real-time displacement and the optimal focal plane distance, that is, , directly used as the defocus index in the original model .
[0039] 2. Constructing the ambiguity-displacement correlation model:
[0040] Substitute the hardware parameters into the original model to form a fuzzy prediction model dedicated to intelligent pre-examination: .
[0041] in, It is a comprehensive constant term of non-defocus factors, including the influence of fixed detection indicators such as resolution, edge sharpness, and contrast; is the defocus sensitivity coefficient, which represents the contribution of unit defocus amount to blur (unit: blur value / mm).
[0042] 3. Construct an optimal soft-sensing model based on multimodal information fusion:
[0043] (1) Constructing the basic framework of soft measurement model:
[0044] Assuming that the multimodal feature fusion item is the prediction model of the target indicator (such as fuzziness), the basic framework is: .
[0045] : predicted value of target indicator;
[0046] : No. Detection indicators (such as resolution , edge sharpness , contrast wait);
[0047] : Basic function transformation (such as power function, logarithmic function, exponential function, etc.);
[0048] : Weight coefficient to be determined.
[0049] (2) Multimodal features and basic function mapping:
[0050] Define five detection indicators and their basic function transformations, as shown in Table 1:
[0051]
[0052] Finally, the linearized eigenvector is formed: .
[0053] Collect m groups of sample data and build an observation matrix and the target value vector .
[0054] Detection indicator matrix: .
[0055] Target value vector: .
[0056] Using the formula:
[0057] .
[0058] Calculate the weight coefficient, where K* is the weight coefficient group, is the detection index value after linearization processing, is the detection index value group after linearization processing, is the observed value of the target indicator.
[0059] Take the training set = .
[0060] Minimize the sum of squared prediction errors , Represents a minimization operation, the goal is to find the optimal defocus sensitivity coefficient .like If the defocus value is large, it means that the blur has a significant impact on the image quality, and hardware adjustments (such as telescopic device displacement) need to be made to reduce Δd to improve image clarity. If it is smaller, it means that the image blur is mainly caused by non-defocus factors (such as insufficient resolution) and other parameters (such as fill light and resolution) need to be optimized.
[0061] make , = ,but Obtain .
[0062] 4. Standard deviation screening optimal model:
[0063] (1) Prediction value calculation: For each model, the detection index value is used to calculate the prediction value. (j is the model number);
[0064] (2) Calculation error .
[0065] Compute the error matrix, where is the error, is the observed value of the target indicator, is the predicted value of the target indicator.
[0066] (3) Standard deviation assessment: .
[0067] Compute the set of standard deviations of all errors, where is the set of standard deviations, is the number of errors, is the error, is the error mean; select The smallest model and its weight coefficients are taken as the final solution.
[0068] .
[0069] Step S4: The material acquisition device sends the acquired high-definition image to the operating device. The operating device calculates the material tilt angle by recognizing the high-definition image, corrects the high-definition image, and performs affine transformation to obtain a first sample image.
[0070] The correction model for calculating the material tilt angle is,
[0071]
[0072] Physical meaning:
[0073] The tilt angle of the material image (such as the rotation angle when scanning an ID card or business license) is used for image correction.
[0074] Calculation logic:
[0075] By detecting the boundary line of the target area (such as the edge of an ID card), select two points and Calculate the slope and convert it to an angle.
[0076] The rectified image can be transformed by affine Implementation, get the first sample image, cv2.getRotationMatrix2D() is a function in OpenCV used to create a rotation matrix.
[0077] In step S5, grayscale conversion and binarization are performed on the first sample image to obtain a second sample image;
[0078] 1. Image preprocessing:
[0079] (1) Grayscale conversion and binarization processing:
[0080] Grayscale processing By eliminating the color dimension RGB → single channel, the table image is converted into a matrix containing only brightness information, which significantly reduces the complexity of subsequent calculations.
[0081] Binarization Threshold segmentation is used to divide the pixels of the first sample image into white pixels and black pixels. White pixels are represented by "0" and black pixels are represented by "1". This achieves strict separation of table lines, text and blank areas, and obtains the second sample image. The mathematical expression is:
[0082] .
[0083] in, is the binarization threshold.
[0084] Step S6, performing 3*3 chessboard gridding processing on the second processed image to obtain a third sample table image;
[0085] Checkerboard gridding:
[0086] In order to achieve efficient analysis and feature extraction of table structure, a 3×3 chessboard partitioning operation is uniformly performed on all pre-processed empty tables, and each empty table is finally accurately split into 9 independent and logically related chessboard grids. Each chessboard grid corresponds to a logical area of the table (such as the title area, data area, and dividing line area), forming a "local feature The hierarchical representation system of "global layout" forms a first-level topology, such as Figure 2 As shown, the third sample image is obtained.
[0087] Step S7, performing information processing and feature coding on all chessboard grids in the third sample table image to obtain a fourth sample table image;
[0088] Chessboard informationization:
[0089] (1) Pixel ratio judgment:
[0090] In the theoretical system of chessboard information processing, pixel ratio judgment is a typical application of image binarization threshold segmentation. is the total number of black pixels, m×n is the chessboard size), and is compared with the threshold T to achieve information discrimination of the image area. From the perspective of information theory, when When the information entropy of the region reaches a significant level at the threshold, it is marked as "1", indicating that it carries sufficient prominent structural information; otherwise, it is marked as "0", indicating that it is an information-sparse region. For each checkerboard cell c (such as a 3×3 sub-region), its local information entropy is calculated.
[0091] .
[0092] The larger the entropy value, the richer the information in the area (such as text-dense areas), and vice versa, the sparser the information (such as blank areas). Define the optimal threshold Maximize the global entropy:
[0093]
[0094] .
[0095] (2) Construction of symbol system for row and column pattern classification:
[0096] In the construction of the symbol system for row and column pattern classification, coding theory is used to achieve hierarchical mapping of features. Each row of three binary combinations (000∼111) constitutes eight states, which are mapped to different symbol sets to construct a three-level coding system: the first row is used as the title area, using uppercase letters (A-H), using their visual salience to strengthen their dominant position in the information structure; the second row is used as the transition area, using lowercase letters (a-h), forming a transition buffer in the visual hierarchy; the third row is used as the number area, using numbers (0-7) to facilitate the intuitive expression of numerical information. This symbol system follows the principle of hierarchical semantics in semiotics. By distinguishing the levels of visual symbols, it improves the readability and recognizability of information, realizes the cognitive transformation from binary coding to semantic symbols, and defines the mapping function is a 3-bit binary space, .
[0097] Title area mapping (uppercase letters): , corresponding rules:
[0098] b 1 b 2 b 3→chr(65+bin2dec( b 1 b 2 b 3))For example, 101→chr(65+5)=F.
[0099] Transition zone mapping (lowercase letters): , corresponding rules:
[0100] b 1 b 2 b 3→chr(97+bin2dec( b 1 b 2 b 3)) For example, 011→chr(97+3)= g .
[0101] Numeric area mapping (numeric symbols): , corresponding rules:
[0102] b 1 b 2 b 3→bin2dec( b 1 b 2b 3) For example, 101→6. Figure 3 shown.
[0103] In summary, if Figure 4 As shown, it can be named Ce5, where bin2dec() is a function that converts a 3-bit binary number into a decimal number, which is used to implement the mapping from binary encoding to symbol set, and chr() is a character conversion function that is used to convert the decimal ASCII code value into the corresponding character.
[0104] If there is a sample table corresponding to multiple empty tables, we filter out the key checkerboards in the sample table and further divide them into 3*3 checkerboards to form a secondary topological structure. Figure 5 As shown:
[0105] By increasing the threshold, the chessboard information can be distinguished more finely.
[0106] Perform chessboard information processing on the 3*3 chessboard in the key chessboard of the sample table, and add separators to the original naming to make a compound name, such as Figure 6 As shown:
[0107] In summary, the composite name of this sample table is Ce5-Ec4-Db2-Cc1. , and obtain the fourth sample image.
[0108] By visually distinguishing between uppercase and lowercase letters and numbers, a semantic coding system corresponding to "level-function" is constructed, making the numbering itself self-explanatory (such as "Gc7" in which G→first row mode, c→middle row mode, and 7→last row mode).
[0109] Step S8, performing intelligent comparison between the fourth sample table image and the blank table, eliminating redundant blank tables, and obtaining a fifth sample table image;
[0110] The efficiency principle of Ascii encoding comparison:
[0111] Assume that the row and column codes of the sample table and the empty table are symbol sequences .
[0112] in , define the numerical mapping function as:
[0113] .
[0114] This mapping satisfies the injectivity property, ensuring a one-to-one correspondence between characters and values, where Definition: Map ASCII characters to corresponding positive integer values. Comparison rules: Use lexicographical numerical values to implement the matching conditions.
[0115] .
[0116] Termination condition: If there is Get the fifth sample image, where Meaning: Compare the ASCII values of each position character of S and E in sequence.
[0117] Step S9: performing material element segmentation on the fifth sample image;
[0118] 1. Difference feature map calculation:
[0119] .
[0120] in To indicate the A difference feature map; is an activation function used to enhance the nonlinear expression of features; It is a gradient operator that can highlight the edge contour information of elements such as official seals and handwritten text in the material by calculating the gradient of the feature map. Compared with the direct difference calculation of the original formula, it can capture detailed features more effectively. They represent the characteristic maps of materials at different scales. is a dynamic convolution kernel, where is a 3×3 Gaussian kernel used to smooth the feature map and reduce noise; The 5×5 Laplacian kernel is used to enhance the edge information of the feature map. The use of this dynamic convolution kernel can perform adaptive feature extraction based on the characteristics of features at different scales.
[0121] 2. Cross-scale feature fusion:
[0122] .
[0123] Represents the optimized fusion feature map, which is used for subsequent material element segmentation; The Softmax activation function can normalize the fused features, enhance the nonlinear expression of features, and improve the segmentation accuracy of irregular elements (such as handwritten annotations); image clarity-preliminary confidence dynamic adjustment model; are dynamic weights, satisfying α + β = 1. These two weights can be adaptively adjusted based on different types of materials (such as business licenses, medical receipts, etc.) to better balance the contributions of different features; It is the difference feature map obtained after the previous optimization; They are multi-scale pooling features, It is obtained by global average pooling of 1×1 convolution, which can capture the global information of the material; It is obtained by local maximum pooling of 7×7 convolution, which can highlight the local features of the material; It represents the feature concatenation operation, which concatenates different feature maps in the channel dimension. Compared with the multiplication operation in the original formula, it can retain more feature information.
[0124] Step S10, performing comprehensive quality assessment model calculation on the elements of the fifth sample image;
[0125] Achieve a comprehensive and accurate assessment of the accuracy and reliability of material review. The model is:
[0126] .
[0127] : Comprehensive quality assessment of elements of the fifth sample table image.
[0128] : OCR recognition confidence (0-100%), obtained by weighted averaging the character-level probabilities output by the OCR model after Softmax or CTC decoding, reflecting the recognition reliability of the text content.
[0129] : Seal / signature matching degree (%), calculated by hash similarity: =Total number of matching digits × 100%.
[0130] Among them, the number of feature bits comes from local image features (such as SIFT / SURF key point descriptors) or neural network feature vectors (such as deep features extracted by CNN), and the number of matching bits is the degree of feature overlap between the material to be reviewed and the standard template.
[0131] : Table structured parsing accuracy (%), calculated by checking indicators such as table row and column alignment and field completeness, for example: =(1 − total number of fields, number of missing fields + number of misplaced fields) × 100%.
[0132] are weight distribution coefficients, and , The weight value is dynamically adjusted according to the type of matter.
[0133] Step S11, performing intelligent verification on all collected materials through the business rule library;
[0134] 1. Material integrity logical expression:
[0135] .
[0136] Logical operators: Represents "logical AND", meaning that for all items from i=1 to n, both "material existence" and "number of copies" must be met.
[0137] : For each material item in all materials submitted by the user, determine whether the material item submitted by the user belongs to the material items listed in the list.
[0138] : Used to determine the number of copies of each material item submitted and whether the number of copies of each material item meets the minimum quantity requirements listed on the list.
[0139] Example: Business registration requires "Business License (1 copy) + Articles of Association (3 copies)". If one copy is missing or the number of copies is insufficient, the completeness = false.
[0140] 2. Cross-material consistency verification formula:
[0141]
[0142] Function definition: Extract the function for the kth key information (such as "name", "ID number", "validity period"); is the product operator, which calculates the product of all terms with k ranging from 1 to p.
[0143] When all key information matches, consistency = 1, otherwise points are deducted according to the proportion of mismatches (for example, "name is the same but address is different", consistency = 0.8).
[0144] Step S12: Output the audit results and generate a receipt or problem mark.
[0145] Example 2:
[0146] like Figure 7 and Figure 8 As shown, an intelligent pre-screening system includes an electrically connected operating device 1 and a material collection device 2, wherein a main screen 3 and a sub-screen 4 are respectively provided on opposite sides of the operating device, and the operating device 1 is provided with a facial recognition module 5, a fingerprint collection module 6, and an ID card and RF card collection module 7 at the edge of the main screen 3; the material collection device 2 includes a material collection area 8, a telescopic component 9 vertically arranged on one side of the material collection area, a mounting plate 10 is horizontally provided at the upper end of the telescopic component 9, and a fill light 11 and an image collection module 12 aimed at the material collection area 8 are provided on the lower side of the mounting plate.
[0147] Although the present invention has been described herein with reference to a number of illustrative embodiments thereof, it will be understood that numerous other modifications and implementations may be devised by those skilled in the art that fall within the scope and spirit of the principles disclosed herein. More specifically, within the scope of the present disclosure, the drawings, and the claims, numerous variations and modifications may be made to the components and / or layout of the subject combination arrangement. In addition to variations and modifications to the components and / or layout, other uses will also be apparent to those skilled in the art.
Claims
1. A material identification method based on chessboard grid topology coding, characterized in that: The method comprises the following steps: Step S1, the user selects an item to be handled through the operating device; Step S2: The user places the material in the pre-screening area of the material collection device; Step S3: The material collection device collects high-definition images and dynamically adjusts the blur during the collection process; In step S4, the material acquisition device sends the collected high-definition image to the operating device, which calculates the material inclination angle by identifying the high-definition image, corrects the high-definition image and performs affine transformation to obtain a first sample image; in step S5, the first sample image is grayscale converted and binarized to obtain a second sample image; in step S6, the second processed image is gridded into a 3*3 chessboard to obtain a third sample image; in step S7, all chessboard grids in the third sample image are subjected to information processing and feature encoding to obtain a fourth sample image; in step S8, the fourth sample image is intelligently compared with the blank table, and redundant blank tables are eliminated to obtain a fifth sample image; in step S9, the fifth sample image is segmented into material elements; in step S10, a comprehensive quality assessment model is calculated for the elements of the fifth sample image; in step S11, all collected materials are intelligently verified through a business rule library; in step S12, the audit result is output; in step S3, the dynamic adjustment of fuzziness during the acquisition process includes the following steps; Step K1, constructing the fuzziness prediction model, in, Indicates the The characteristic transformation function of the detection index, is the defocus degree indicator, is a constant term, is the defocus sensitivity coefficient, is the predicted value of the target indicator; Step K2, construct the hardware-image quality mapping relationship defined as follows, the optimal focal plane distance , real-time z-axis displacement , defocus amount , the absolute deviation between the real-time displacement and the optimal focal plane distance, that is , directly used as the defocus index in the original model Step K3, construct the ambiguity-displacement correlation model, substitute the hardware parameters into the original model to form the ambiguity prediction model, ,in, is the comprehensive constant term of non-defocusing factors; Step K4, construct the optimal soft measurement model based on multimodal information fusion, set the multimodal feature fusion term as the prediction model of the target indicator, the basic framework is , For the Detection indicators, in =0,1,2,3,4, The corresponding detection index is resolution, The corresponding detection index is edge sharpness. The corresponding detection index is contrast, The corresponding detection index is MTF high frequency response, The corresponding detection index is noise intensity, is the basic function transformation, is the weight coefficient to be determined; step K5, finally forms the linearized eigenvector, , collect m groups of sample data and build the observation matrix and the target value vector , detection indicator matrix, , target value vector: , using the formula Calculate the weight coefficient, where is the weight coefficient group, is the detection index value after linearization processing, is the detection index value group after linearization processing, is the observed value of the target indicator, take the training set = , minimize the sum of squared prediction errors ,make, , ,but = , to obtain , Represents a minimization operation, the goal is to find the optimal defocus sensitivity coefficient k6; Step K6, for each model, the detection index value is brought in to calculate the predicted value , is the model number, calculation error ,in is the error, is the observed value of the target indicator, the standard deviation is evaluated, , calculate the set of standard deviations of all errors, where is the set of standard deviations, is the number of errors, is the error, is the error mean; select The smallest model and its weight coefficient are taken as the final solution. .
2. The material identification method based on chessboard grid topology coding according to claim 1, characterized in that: The correction model for calculating the material tilt angle in step S4 is: , by detecting the boundary line of the target area, select two points and Calculate the slope and convert it into an angle. The corrected high-definition image can be transformed by affine transformation. accomplish.
3. The material identification method based on chessboard grid topology coding according to claim 2, characterized in that: In step S5, the grayscale processing is By eliminating the color dimension RGB → single channel, the first sample image is converted into a matrix containing only brightness information, which significantly reduces the complexity of subsequent calculations: the binarization process is Threshold segmentation is used to divide the pixels of the first sample image into white pixels and black pixels. White pixels are represented by "0" and black pixels are represented by "1". The strict separation of table lines, text and blank areas is achieved. The mathematical expression is: : In step S6, in order to achieve efficient analysis and feature extraction of the image structure of the third sample table, a 3×3 chessboard division operation is uniformly performed on all the pre-processed empty tables, and each empty table is finally accurately divided into 9 independent and logically related chessboard grids: In step S7, the black pixel ratio is calculated ,in is the total amount of black pixels, m×n is the chessboard size, and is compared with the threshold T to realize the information discrimination of the image area. When the threshold is reached, the information entropy of the region reaches a significant level and is marked as "1", indicating that it carries sufficiently prominent structural information; otherwise, it is marked as "0", which is considered an information sparse region. For each checkerboard square c, its local information entropy is calculated. The larger the entropy value, the richer the information in the area, and vice versa, the information is sparse. The optimal threshold is defined as Maximize the global entropy, , is the indicator function, when In the construction of the symbol system for row and column pattern classification, the coding theory is used to realize the hierarchical mapping of features. Each row of 3-bit binary combinations 000-111 constitutes 8 states, which are mapped to different symbol sets respectively to build a three-level coding system: the first row is used as the title area, using uppercase letters AH; the second row is used as the transition area, using lowercase letters ah; the third row is used as the number area, using numbers 0-7, and the mapping function is defined. ,in is a 3-bit binary space, A composite symbol set.
4. The material identification method based on chessboard grid topology coding according to claim 3, characterized in that: In step S8, the row and column codes of the sample table and the empty table are symbol sequences and ,in , , define the numerical mapping function as , , define the matching conditions as, , the termination condition is if there is , it is judged as mismatch and the comparison is terminated.
5. The material identification method based on chessboard grid topology coding according to claim 4, characterized in that: In step S9, the material element segmentation step is as follows: step T1, difference feature map calculation, ,in To indicate the A difference feature map; is an activation function used to enhance the nonlinear expression of features. is the gradient operator, and Represent the characteristic maps of different scales of materials, and is a dynamic convolution kernel, where is a 3×3 Gaussian kernel used to smooth the feature map and reduce noise; is a 5×5 Laplacian kernel, which is used to enhance the edge information of the feature map; step T2, cross-scale feature fusion, ; Represents the optimized fusion feature map, which is used for subsequent material element segmentation; is the Softmax activation function; image clarity-preliminary confidence dynamic adjustment model; is a dynamic weight and satisfies , and It is the difference feature map obtained after optimization; and They are multi-scale pooling features, It is obtained by global average pooling of 1×1 convolution, which can capture the global information of the material. It is obtained through local maximum pooling of 7×7 convolution, which can highlight the local features of the material; Represents a feature concatenation operation.
6. The material identification method based on chessboard grid topology coding according to claim 5, characterized in that: In step S10, a comprehensive and accurate assessment of the accuracy and reliability of the material review is achieved. The model is , : OCR recognition confidence, obtained by weighted averaging the character-level probability output by the OCR model after Softmax or CTC decoding, reflecting the recognition reliability of the text content. : Seal or signature matching degree, calculated by hash similarity: = Total number of feature bits matched × 100%, where the number of feature bits comes from local features of the image or neural network feature vectors, and the number of matching bits is the degree of feature overlap between the material to be reviewed and the standard template. For the accuracy of table structured parsing, weight distribution: , the weight is dynamically adjusted according to the type of matter.
7. The material identification method based on chessboard grid topology coding according to claim 6, characterized in that: In step S11, the material integrity logical expression is: , logical operators: Indicates "logical AND", which requires both "material existence" and "number of copies to be compliant" to be met. The cross-material consistency verification formula is: , function definition: The kth key information extraction function is used. When all key information matches, consistency = 1. Otherwise, points are deducted according to the proportion of mismatched items.
Citation Information
Patent Citations
Self-service all-in-one machine based on face detection and character recognition and using method thereof
CN105469513A
Intelligent government administration service working system and application thereof
CN110322643A
Batch fuzzy identifier reconstruction method based on deformable convolution
CN116796773A
Image definition sorting method and device, electronic equipment and storage medium
CN118429336A