Digital employee intelligent report automatic generation method and system
By collecting data on document lighting and paper characteristics using a sensor array, adaptively optimizing scanning parameters and monitoring environmental noise, and constructing a data traceability chain, this technology solves the problems of high recognition error rate, difficulty in tracing the source, and low review efficiency in traditional OCR technology, achieving highly accurate and traceable document processing.
Patent Information
- Application Number
- CN202511097677.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-06
- Publication Date
- 2025-11-21
AI Technical Summary
Traditional OCR technology suffers from high error rates when processing complex or low-quality documents, making it difficult to guarantee data accuracy, trace errors, and conduct manual review inefficiently. Furthermore, it cannot intelligently identify high-risk data items.
By collecting data on document lighting and paper characteristics using a sensor array, a document quality index is calculated, scanning parameters are adaptively optimized, environmental noise is monitored, a data traceability chain is constructed, and a risk perception report is generated.
It improves OCR recognition accuracy, reduces recognition errors, enables data traceability and efficient manual review, provides more reliable data accuracy assessment, and optimizes the document digitization process.
Smart Images

Figure CN120997857A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of document digitization processing and intelligent report generation, and particularly relates to a digital employee intelligent report automatic generation method and system. BACKGROUND
[0002] Traditional OCR mainly relies on the quality of scanned images for recognition. However, the physical environment before scanning (such as uneven lighting, light source color temperature deviation) and the paper itself characteristics (such as wrinkles, stains, paper color, light transmission, light reflection) directly affect the quality of the scanned image (contrast, definition, color, shadow, etc.). Traditional methods usually use fixed or artificially experienced set scanning parameters, which cannot be optimized for specific physical conditions, resulting in large fluctuations in input image quality, and further causing high OCR recognition error rate, especially when dealing with complex or low-quality documents, data accuracy is difficult to guarantee.
[0003] When data errors are found in subsequent business processes, traditional methods often have difficulty determining the specific causes of errors. Due to the lack of systematic recording of physical environment, document itself state and key parameters and environmental factors in the processing process, the association between data and original physical conditions and processing process is broken, making it difficult to trace errors and locate the root cause of the problem for targeted improvement.
[0004] The recognition result of traditional OCR usually accompanies a confidence score, but this mainly reflects the recognition grasp of the algorithm itself, and does not comprehensively consider the error possibility introduced by external risk factors such as physical environment. This leads to the system being unable to intelligently and accurately identify data items that are truly "high risk" and require manual review. At the same time, environmental noise and other factors during manual review can also affect the judgment of the reviewer, but the existing system cannot perceive and record these risks, resulting in large manual review workload, low efficiency, and the review itself also introduces new errors.
[0005] In summary, the existing technology ignores physical environmental factors and external interference in the processing process, resulting in problems such as insufficient data accuracy, difficult error tracing, and low efficiency of manual review, which need to be solved. SUMMARY
[0006] Therefore, it is necessary to provide a digital employee intelligent report automatic generation method and system to solve at least one of the above technical problems.
[0007] To achieve the above purpose, a digital employee intelligent report automatic generation method comprises the following steps:
[0008] Step S1: Collecting a light characteristic dataset and a paper physical characteristic set of the document to be scanned by a sensor array; detecting environmental interference factors on the document to be scanned to obtain an environmental interference dataset; calculating a document physical quality index according to the light characteristic dataset, the paper physical characteristic set and the environmental interference dataset, and then generating a document quality distribution map;
[0009] Step S2: Determining a scanning parameter reference table according to the document physical quality index; correcting the scanning parameter reference table according to paper characteristic parameters to obtain a paper adaptive parameter table; calculating regional adaptive parameters on the document quality distribution map to obtain a regional parameter adjustment matrix; identifying and enhancing a key information region on the document quality distribution map according to the regional parameter adjustment matrix to obtain a key region enhancement matrix; collecting a scanned result image, and simultaneously integrating the paper adaptive parameter table and the key region enhancement matrix to perform scanning quality evaluation to obtain a scanning quality score map;
[0010] Step S3: Monitoring and recording environmental noise data by an environmental noise sensor during a document scanning operation to obtain environmental noise data; performing document quality partition OCR processing based on the scanning quality score map to obtain a partition recognition result set; performing artificial environmental adaptability analysis according to the environmental noise data to obtain an environmental influence time axis; calculating data reliability according to the environmental influence time axis on the partition recognition result set to obtain a data reliability matrix;
[0011] Step S4: Constructing a data traceability chain according to the document physical quality index, the scanning quality score map and the data reliability matrix; generating a digital employee intelligent report template according to the data traceability chain; optimizing a report flow on the digital employee intelligent report template to obtain a flow optimization suggestion report.
[0012] The present application realizes quantitative evaluation of the physical environment before document scanning and the characteristics of the document itself through multi-dimensional sensor perception and data acquisition. The calculated document physical quality index and document quality distribution map provide an objective and detailed report on the "health status" of the document, which overcomes the drawbacks of traditional methods that ignore physical conditions, provides key input data for subsequent processes, and lays the foundation for data quality management and error traceability. Based on the physical quality evaluation results of step S1, intelligent adaptive optimization of scanning parameters is realized, including global parameter correction and regional fine adjustment, especially enhanced processing for low-quality areas and key information areas. Through real-time parameter fine-tuning closed-loop control, high-quality scanned images can be obtained under different physical conditions. The generated scanning quality score map objectively quantifies the image quality after scanning, directly improving the input quality of OCR recognition and reducing recognition errors caused by image problems. Combined with the scanning quality score map of step S2, the optimal OCR recognition strategy is intelligently selected according to the regional quality and text type, improving the recognition accuracy of complex documents. At the same time, by monitoring environmental noise and evaluating its potential impact on manual judgment, environmental factors are taken into account. The most core benefit is to calculate the comprehensive data reliability matrix, which not only reflects the confidence of the OCR algorithm, but also integrates external risk factors such as scanning quality and environmental impact, providing more reliable data accuracy evaluation than traditional methods, and accurately identifying high-risk data that need attention. Based on comprehensive physical, scanning, and recognition link data, a complete data traceability chain is established, making any recognized data item traceable to its generation process and influencing factors, completely solving the problem of difficult error traceability in traditional methods. The generated intelligent report intuitively highlights key information through risk classification and environmental impact identification, and provides interactive traceability functions, greatly improving the efficiency and accuracy of manual review. Long-term data accumulation and analysis can further generate process optimization suggestions, driving continuous improvement and quality improvement of the entire document digitization process.
[0013] Therefore, the present application provides a digital employee intelligent report automatic generation method, which constructs an intelligent processing closed loop of physical environment perception and risk association. By perceiving and quantifying the physical environment and paper characteristics before scanning, scanning parameters can be adaptively optimized to obtain higher quality original images. In the OCR processing stage, scanning quality and real-time environmental noise are combined to intelligently select recognition strategies and calculate data reliability, which is not only based on recognition results, but also integrates physical and environmental risk factors. Finally, based on these comprehensive perception and evaluation data, a risk perception report is generated and a complete data traceability chain is established, accurately marking high-risk data, prompting environmental impact, and providing process optimization suggestions, thereby fundamentally improving data accuracy, achieving traceability of errors, and significantly improving the efficiency of manual review.
[0014] Preferably, the present application also provides a digital employee intelligent report automatic generation system for executing the digital employee intelligent report automatic generation method as described above, and the digital employee intelligent report automatic generation system comprises:
[0015] A physical environment perception module is configured to collect a light characteristic dataset and a paper physical characteristic set of a document to be scanned through a sensor array, detect environmental interference factors of the document to be scanned to obtain an environmental interference dataset, and calculate a document physical quality index according to the light characteristic dataset, the paper physical characteristic set and the environmental interference dataset, and further generate a document quality distribution atlas;
[0016] An intelligent scanning optimization module is configured to determine a scanning parameter benchmark table according to the document physical quality index, correct paper characteristic parameters of the scanning parameter benchmark table to obtain a paper adaptive parameter table, calculate regional adaptive parameters of the document quality distribution atlas to obtain a regional parameter adjustment matrix, identify and enhance a key information region of the document quality distribution atlas according to the regional parameter adjustment matrix to obtain a key region enhancement matrix, collect a scanning result image, and simultaneously integrate the paper adaptive parameter table and the key region enhancement matrix to perform scanning quality evaluation to obtain a scanning quality score atlas;
[0017] An environment perception identification module is configured to monitor and record environmental noise of a document scanning operation through an environmental noise sensor to obtain environmental noise data, perform partitioned OCR processing of the document quality based on the scanning quality score atlas to obtain a partitioned recognition result set, perform artificial environmental adaptability analysis according to the environmental noise data to obtain an environmental influence time axis, and calculate data reliability of the partitioned recognition result set according to the environmental influence time axis to obtain a data reliability matrix;
[0018] A risk report generation module is configured to construct a data traceability chain according to the document physical quality index, the scanning quality score atlas and the data reliability matrix, generate a digital employee intelligent report template according to the data traceability chain, and perform report flow optimization on the digital employee intelligent report template to obtain a flow optimization suggestion report.
[0019] The digital employee intelligent report automatic generation system, through the physical environment perception module, quantitatively evaluates the physical conditions before scanning and the document characteristics, and provides basic quality information for subsequent processing; the intelligent scanning optimization module uses the evaluation results to adaptively adjust the scanning parameters and perform key area enhancement, significantly improving the scanning image quality; the environment sensing identification module combines the scanning quality and real-time environmental noise to intelligently perform OCR, and calculates the data reliability including physical and environmental risk factors, accurately identifies high-risk data; the risk report generation module integrates all link data to build a complete traceability chain and generate a risk perception report, which intuitively prompts the risk and environmental impact, greatly improves the artificial review efficiency and supports the continuous optimization of the process. The system as a whole builds an intelligent closed loop from the physical environment to the final report, and fundamentally improves the data accuracy, traceability and processing efficiency. BRIEF DESCRIPTION OF DRAWINGS
[0020] Figure 1 It is a step flowchart of a digital employee intelligent report automatic generation method.
[0021] Figure 2 It is a detailed implementation step flowchart of step S1 in the application.
[0022] The application purposes, functional features and advantages will be further described with reference to the embodiments and the accompanying drawings. DETAILED DESCRIPTION
[0023] The technical method of the application will be described clearly and completely below in combination with the accompanying drawings. Obviously, the described embodiments are part of the embodiments of the application, rather than all the embodiments. Based on the embodiments in the application, all other embodiments obtained by those skilled in the art without creative labor fall within the protection scope of the application.
[0024] In addition, the accompanying drawings are only schematic illustrations of the application, and are not necessarily drawn to scale. The same reference signs in the drawings represent the same or similar parts, and thus repeated descriptions thereof will be omitted. Some block diagrams shown in the drawings are functional entities, which do not necessarily correspond to physically or logically independent entities. The functional entities can be implemented in the form of software, or in one or more hardware modules or integrated circuits, or in different network and / or processor methods and / or microcontroller methods.
[0025] It should be understood that, although the terms "first", "second", etc. can be used herein to describe various elements, these elements should not be limited by these terms. These terms are only used to distinguish one element from another. For example, a first element could be termed a second element, and, similarly, a second element could be termed a first element, without departing from the scope of the example embodiments. The term "and / or" as used herein includes any and all combinations of one or more of the associated listed items.
[0026] In the embodiments of the present application, referring to Figure 1 As shown in the figure, it is a schematic diagram of the step flow of the digital employee intelligent report automatic generation method. In the present example, the digital employee intelligent report automatic generation method comprises the following steps:
[0027] Step S1: Collecting the light characteristic data set and the paper physical characteristic set of the to-be-scanned document through the sensor array; detecting the environmental interference factors of the to-be-scanned document to obtain the environmental interference data set; calculating the document physical quality index according to the light characteristic data set, the paper physical characteristic set and the environmental interference data set, and then generating a document quality distribution map;
[0028] In the embodiments of the present application, by deploying the sensor array, the system collects the light characteristic data set of the to-be-scanned document, such as light intensity, uniformity and color temperature, and measures the paper physical characteristic set, such as paper thickness, flatness, light transmittance, light reflectance and texture complexity. At the same time, the environmental interference factors, such as environmental humidity, paper moisture content, environmental temperature and air particulate matter concentration, are detected to obtain the environmental interference data set. These data sets are comprehensively used to calculate the document physical quality index, which quantifies the overall physical quality of the document. According to the index and the local physical characteristic data, a document quality distribution map is generated to mark the quality level and high-risk area of each region of the document.
[0029] Step S2: Determining the scanning parameter reference table according to the document physical quality index; correcting the paper characteristic parameter of the scanning parameter reference table to obtain a paper adaptive parameter table; calculating the region adaptive parameter of the document quality distribution map to obtain a region parameter adjustment matrix; identifying and enhancing the key information region of the document quality distribution map according to the region parameter adjustment matrix to obtain a key region enhancement matrix; collecting the scanning result image, and simultaneously integrating the paper adaptive parameter table and the key region enhancement matrix to perform scanning quality evaluation to obtain a scanning quality score map;
[0030] In the embodiment of the present application, the system determines the scanning parameter reference value according to the document physical quality index, and corrects the reference value in combination with the paper physical characteristics to obtain a paper adaptive parameter table. The resolution, contrast, focal length and other adaptive adjustment parameters of each region (especially the high-risk region) are calculated by analyzing the document quality distribution map, a region parameter adjustment matrix is generated, and the parameter transition is ensured to be natural through smoothing processing. The pre-scanning is performed to identify the key information regions such as seals, signatures and table frames, the targeted image enhancement parameters are formulated, and are integrated into the region parameter adjustment matrix to form a key region enhancement matrix. The scanning execution scheme is generated by integrating the paper adaptive parameter table and the key region enhancement matrix, and the scanning is performed. The parameter fine-tuning closed-loop control is performed through real-time scanning feedback data during the process to obtain the scanning result image. Finally, the actual quality of each region of the scanning result image is evaluated, the deviation from the expected quality is calculated, and the scanning quality score map reflecting the quality distribution of the scanned image is generated.
[0031] Step S3: environmental noise monitoring and recording of the document scanning operation is performed through the environmental noise sensor to obtain environmental noise data; document quality partition OCR processing is performed based on the scanning quality score map to obtain a partition recognition result set; artificial environmental adaptability analysis is performed according to the environmental noise data to obtain an environmental influence time axis; data reliability calculation is performed on the partition recognition result set according to the environmental influence time axis to obtain a data reliability matrix;
[0032] In the embodiment of the present application, during the document scanning and OCR processing, the intensity, type and timing of the environmental noise are monitored and recorded through the environmental noise sensor to obtain environmental noise data. Based on the scanning quality score map, the scanning result image is divided into high-quality, medium-quality and low-quality regions, and the optimal OCR algorithm combination is matched according to the text type in the region. The standard, enhanced or multi-algorithm cross-verification OCR processing is performed on different quality regions, and the environmental noise data of the processing period is synchronously associated. The environmental noise data is analyzed to evaluate the interference risk of different periods on the attention of artificial review, and an environmental influence time axis is constructed. Finally, the OCR recognition results of each region are calculated according to the environmental influence time axis. The reliability takes into account the OCR algorithm confidence, the region scanning quality and the environmental influence factor, and a data reliability matrix containing the reliability score of each data item is generated.
[0033] Step S4: a data traceability chain is constructed according to the document physical quality index, the scanning quality score map and the data reliability matrix; a digital employee intelligent report template is generated according to the data traceability chain; a report flow optimization suggestion report is obtained by optimizing the digital employee intelligent report template.
[0034] In the embodiment of the present application, according to the data credibility matrix, a threshold is set to screen and classify (high, medium and low risk) the identified data items, and a comprehensive risk index is calculated. A unique identifier is generated for each data item, and it is associated with the document physical quality index, the quality score of the corresponding area in the scanning quality score map, the used OCR algorithm and the environmental impact factor in the data credibility matrix to construct a complete data traceability chain. Based on the data traceability chain and the risk data classification result, a basic report template is loaded, the identified content is filled in, and high-risk data is marked with different colors, environmental impact identification is added, interactive traceability function is embedded, and a digital employee intelligent report template instance is generated. According to the comprehensive risk index and the like, the report content is dynamically optimized and displayed to obtain a dynamically optimized report. Finally, historical processing data, risk patterns and environmental correlations are analyzed to automatically generate a process optimization suggestion report and propose specific measures to improve the scanning environment, equipment or process.
[0035] As an example of the present application, reference is made to Figure 2 In this example, the step S1 includes:
[0036] Step S11: Deploy a multi-point light illumination sensor array to sample the light illumination characteristics at the four corners and the center position of the document to be scanned to obtain a light illumination characteristic data set;
[0037] In the embodiment of the present application, an array composed of nine digital light illumination sensors is deployed above the scanning area of the document to be scanned, and the sensors are uniformly distributed at the four corners, the midpoints of the four sides and the center position of the document area. Before scanning the document, the system triggers the sensor array to simultaneously collect ambient light data. Each sensor measures and outputs the light intensity at its position, with the unit being lux. These nine light intensity measurement values form an array, which is recorded as a light intensity original value array. The system calculates the average (μ) and the standard deviation (σ) of the array, and then calculates the light uniformity coefficient, whose calculation formula is light uniformity coefficient = 1-(σ / μ), and the value of the coefficient ranges from 0 to 1, and the closer the value is to 1, the more uniform the light is. At the same time, the sensor in the sensor array that has the color temperature measurement function measures the color temperature of the ambient light, with the unit being Kelvin (K), and the value is recorded as the light source color temperature value. The light intensity original value array, the light uniformity coefficient and the light source color temperature value are integrated into the light illumination characteristic data set.
[0038] Step S12: Measure the paper physical characteristics of the document to be scanned according to the light illumination characteristic data set to obtain a paper physical characteristic set;
[0039] In the embodiment of the present application, the thickness of the document to be scanned at multiple sampling points is measured by using a non-contact ultrasonic sensor, and the average of these thickness measurements is calculated as the paper thickness value, with the unit of millimeters (mm). The surface of the document is scanned by using a laser profilometer to obtain an array of elevation data points, and the variance of the array of elevation data points is calculated and standardized to generate a flatness index, for example, flatness index = 1-(elevation variance / reference elevation variance), where the reference elevation variance is the elevation variance of a perfectly flat paper, and the index value ranges from 0 to 1, with a value closer to 1 indicating a flatter paper. In combination with the raw light intensity value collected in step S11, the document to be scanned is placed above a standard light source with a known brightness, and a light sensor located above the document is used to measure the light intensity transmitted through the paper to calculate the transmittance value, with the formula being transmittance value = (transmitted light intensity / standard light source intensity) x 100%, and the unit being percentage. A light sensor with spectral analysis capability is used to measure the intensity and spectral distribution of the light reflected from the surface of the paper under the illumination of the light source in step S11, and the reflected light is compared with the light reflected from a standard white board to calculate the reflectance coefficient, for example, the reflectance coefficient can be calculated as the average value of the reflected intensity in a specific waveband divided by the corresponding intensity of the standard white board, with the unit being percentage. A high-resolution camera is used to capture microscopic images of the surface of the paper, and image processing algorithms such as Local Binary Pattern (LBP) or Fourier transform are used to analyze the texture features to calculate the texture complexity value, for example, based on the quantification of the entropy or contrast of the texture features, which is mapped to the range of 1-10, with a larger value indicating more complex texture. The paper thickness value, flatness index, transmittance value, reflectance coefficient, and texture complexity value are integrated into the paper physical property set.
[0040] Step S13: detecting the environmental interference factors of the document to be scanned to obtain an environmental interference data set;
[0041] In the embodiment of the present application, a digital humidity sensor is used to collect the environmental humidity of the scanning area, which is recorded as the environmental humidity value with the unit of percentage (%). A non-destructive paper moisture content measuring instrument is used to measure the moisture content of the paper, with the unit of percentage (%), and in combination with the paper thickness value obtained in step S12, the paper moisture absorption coefficient is calculated, for example, paper moisture absorption coefficient = paper moisture content x paper thickness value, which reflects the moisture absorption degree of the paper under the current humidity and its impact on the thickness. A digital temperature sensor is used to collect the environmental temperature of the scanning area, which is recorded as the environmental temperature value with the unit of Celsius (℃). A laser scattering particulate matter sensor is used to measure the concentration of particulate matter in the air of the scanning area, with the unit of micrograms per cubic meter (μg / m 3Combining the reflectivity coefficient obtained in step S12, the dust interference factor is calculated. For example, the dust interference factor = particulate matter concentration × reflectivity coefficient, which represents the degree of reflection interference from dust particles in the air on the paper surface. The ambient humidity value, paper moisture absorption coefficient, ambient temperature value, and dust interference factor are integrated into an environmental interference dataset.
[0042] Step S14: Calculate the document physical quality index from the lighting characteristics dataset, paper physical characteristics dataset, and environmental interference dataset to obtain the document physical quality index;
[0043] In this embodiment of the invention, the collected illumination uniformity coefficient, light source color temperature value, flatness index, transmittance value, reflectance coefficient, paper thickness value, texture complexity value, ambient humidity value, paper moisture absorption coefficient, ambient temperature value, and dust interference factor are standardized and mapped to a unified scale of 0-1. Some indicators (such as illumination uniformity coefficient, flatness index, and transmittance value) are positively correlated with quality, while others (such as reflectance coefficient, paper thickness value, texture complexity value, ambient humidity value, paper moisture absorption coefficient, ambient temperature value, and dust interference factor) are negatively correlated with quality or exhibit non-linear relationships. Appropriate inverse or non-linear mapping processing is required to generate a standardized score. The illumination suitability score is calculated; for example, illumination suitability score = W4 × standardized (illumination uniformity coefficient) + W5 × standardized (light source color temperature value), where W4 and W5 are preset illumination factor weights. Calculate the paper scan suitability score, for example, Paper Scan Suitability Score = W6 × Standardized (Flatness Index) + W7 × Standardized (Transmittance Value) + W8 × Standardized (Reflectance Coefficient) + W9 × Standardized (Paper Thickness Value) + W 10 × Normalization (texture complexity value), W6 to W 10 This is the preset paper factor weight. Calculate the environmental impact weight, for example, environmental impact weight = W. 11 ×Standardization (paper moisture absorption coefficient) + W 12 ×Standardized (ambient temperature value) + W 13 ×Standardization (dust interference factor), W 11 To W 13 These are preset environmental factor weights. A weighted average algorithm is applied to calculate the document physical quality index. For example, the document physical quality index = W1 × lighting suitability score + W2 × paper scanning suitability score + W3 × environmental impact weight, where W1, W2, and W3 are preset main factor weights, and W1 + W2 + W3 = 1. The calculated index value is scaled to 0-100 points. Based on this score, the document physical quality index is divided into quality levels: 80-100 points is excellent, 60-79 points is good, 40-59 points is average, 20-39 points is poor, and 0-19 points is very poor. The document physical quality index score and quality level are output together.
[0044] Step S15: generating a document quality distribution map according to the document physical quality index, the light characteristic dataset, and the paper physical characteristic set.
[0045] In the embodiment of the present application, the scanning area of the document to be scanned is divided into a 9x9 grid, forming 81 regions. For each grid region, the local light uniformity coefficient and the local flatness index of the center of the region are estimated by an interpolation algorithm (such as bilinear interpolation or Kriging interpolation) using the multi-point light data collected in step S11 and the multi-point paper physical characteristic data collected in step S12. Combining the overall score of the document physical quality index calculated in step S14, and the local light uniformity coefficient and the local flatness index, the local quality score of each region is calculated, for example, local quality score = W 14 × document physical quality index + W 15 × local light uniformity coefficient + W 16 × local flatness index, W 14 , W 15 , W 16 are preset weights. The local quality score of each region is mapped to a color gradient (for example, high score corresponds to green color, and low score corresponds to red color), and a region quality heat map is generated. At the same time, the regions with local quality scores lower than a preset threshold (for example, 40) are identified and marked as high-risk regions that need special attention, for example, by superimposing a border or a special icon on the region quality heat map. The region quality heat map and the high-risk region marking together constitute the document quality distribution map.
[0046] Especially important is that step S14 includes:
[0047] light parameter standardization is performed on the light characteristic dataset to obtain a light suitability score;
[0048] light intensity correction is performed according to the light suitability score to obtain a corrected light suitability;
[0049] paper scanning adaptability evaluation is performed on the paper physical characteristic set to obtain a paper scanning adaptability score;
[0050] environmental impact factor analysis is performed on the environmental interference dataset to obtain an environmental comprehensive score;
[0051] an environmental impact weight is calculated according to the environmental comprehensive score and the paper scanning adaptability score;
[0052] the document physical quality index is calculated according to the environmental impact weight, the corrected light suitability, and the paper scanning adaptability score;
[0053] In the embodiment of the present application, the illumination uniformity coefficient and the light source color temperature value in the illumination characteristic data set obtained in step S11 are subjected to illumination parameter standardization processing to obtain an illumination suitability score. The illumination uniformity coefficient itself is between 0 and 1 and is directly used as a standardized value. The light source color temperature value (unit: K) is mapped to a high score according to an industry standard or an empirical value, and a color temperature deviating from the ideal range (for example, 5000K to 6500K) gradually decreases in score, for example, a Gaussian function or a piecewise linear function is used for mapping, and 7500K and above or 3500K and below are mapped to the lowest score 0, and 5750K is mapped to the highest score 100. Then, the illumination uniformity coefficient and the standardized light source color temperature value are weighted and averaged to calculate the illumination suitability score. The calculation formula is, for example: illumination suitability score = W1 x illumination uniformity coefficient + W2 x standardized (light source color temperature value), wherein W1 and W2 are preset weights and W1 + W2 = 1.
[0054] According to the average value (denoted as μ_illumination intensity) of the illumination intensity original value array collected in step S11 and the illumination suitability score calculated in the previous step, the illumination intensity is corrected to obtain a corrected illumination suitability. This correction takes into account that even if the illumination is uniform and the color temperature is suitable, the absolute illumination intensity that is too high or too low will still affect the scanning quality (for example, overexposure or underexposure). A range of ideal illumination intensity (for example, 300 lux to 500 lux) is set. A correction factor function f(μ_illumination intensity) is constructed, when μ_illumination intensity is within the ideal range, f(μ_illumination intensity) is close to 1; when μ_illumination intensity deviates from the range, f(μ_illumination intensity) is less than 1, the more it deviates, the smaller the factor. For example, a reverse U-shaped function can be used for mapping. The corrected illumination suitability = illumination suitability score x f(μ_illumination intensity).
[0055] The paper thickness value, flatness index, light transmittance value, light reflectance coefficient, and texture complexity value are standardized and mapped to a score of 0-100. For example, the flatness index (0-1) is directly multiplied by 100; the light transmittance value and the light reflectance coefficient (percentage) are mapped according to their influence on scanning, and a too high or too low light reflectance or light transmittance has a lower score; the paper thickness value is scored with reference to the standard document paper thickness, and a too thick or too thin paper has a lower score; the texture complexity value (1-10) is inversely mapped, and the higher the complexity, the lower the score. Then, the standardized paper feature scores are weighted and averaged to calculate the paper scanning adaptability score. The calculation formula is, for example: paper scanning adaptability score = W3 x standardized (paper thickness value) + W4 x standardized (flatness index) + W5 x standardized (light transmittance value) + W6 x standardized (light reflectance coefficient) + W7 x standardized (texture complexity value), wherein W3 to W7 are preset weights and W3 + W4 + W5 + W6 + W7 = 1.
[0056] The environmental humidity value, paper hygroscopicity coefficient, environmental temperature value, dust interference factor are standardized and mapped to a score of 0-100, wherein the higher the score of these factors indicates the greater the potential negative impact. For example, the score increases as the environmental humidity value exceeds the ideal range (e.g. 40%-60%); the score increases as the paper hygroscopicity coefficient is higher; the score increases as the environmental temperature value exceeds the ideal range (e.g. 20-25°C); the score increases as the dust interference factor is higher. The standardized environmental interference scores are then weighted and averaged to calculate an environmental comprehensive score. The calculation formula is, for example: Environmental comprehensive score = W8 x standardized (environmental humidity value) + W9 x standardized (paper hygroscopicity coefficient) + W 10 x standardized (environmental temperature value) + W 11 x standardized (dust interference factor), wherein W8 to W 11 are preset weights and W8 + W9 + W 10 + W 11 = 1.
[0057] The environmental comprehensive score and the paper scanning adaptability score calculated in the previous step are used to calculate an environmental impact weight. This weight reflects the degree of the combined negative impact on the scanning quality after the environmental interference factors and the characteristics of the paper itself are superimposed, and the value ranges from 0 to 1, with a higher value indicating a greater negative impact. The calculation formula is, for example: Environmental impact weight = W 12 x (environmental comprehensive score / 100) + W 13 x (1-paper scanning adaptability score / 100), wherein W 12 and W 13 are preset weights and W 12 + W 13 = 1. This weight value will be used to reduce the final physical quality index.
[0058] A basic quality score is first calculated, for example: Basic quality score = W 14 x corrected lighting suitability + W 15 x paper scanning adaptability score, wherein W 14 and W 15 are preset weights and W 14 + W 15 = 1. Then, the basic quality score is corrected using the environmental impact weight to obtain the final document physical quality index. The calculation formula is, for example: Document physical quality index = basic quality score x (1-environmental impact weight). The final index value is scaled to 0-100 points. According to this score, the document physical quality index is divided into quality levels: 80-100 points for excellent, 60-79 points for good, 40-59 points for general, 20-39 points for poor, and 0-19 points for very poor. The numerical score and quality level of the document physical quality index are output together.
[0059] Preferably, the environmental interference factor detection on the document to be scanned in step S13 comprises:
[0060] Multi-point humidity collection is performed on the document to be scanned to obtain a humidity uniformity coefficient;
[0061] Paper moisture content is determined according to the humidity uniformity coefficient to obtain an average paper moisture content;
[0062] Temperature distribution detection is performed on the document to be scanned to obtain a temperature fluctuation index;
[0063] Particle concentration measurement is performed according to the temperature fluctuation index to obtain a temperature-particle correlation factor;
[0064] Reflective light interference evaluation is performed according to the temperature-particle correlation factor to obtain a reflective light interference distribution map;
[0065] A dust interference factor is calculated according to the reflective light interference distribution map;
[0066] The average paper moisture content, the temperature-particle correlation factor and the dust interference factor are integrated to obtain an environmental interference data set.
[0067] In the embodiment of the application, a plurality of (for example, 9) digital humidity sensor arrays are arranged in a non-contact manner around the scanning area of the document to be scanned. Before scanning, these sensors simultaneously measure the ambient humidity at their locations, in percentage (%). These measurement values are recorded in a humidity value array. The average value (μ_humidity) and the standard deviation (σ_humidity) of the array are calculated. The humidity uniformity coefficient is calculated by the formula humidity uniformity coefficient = 1-(σ_humidity / μ_humidity), which reflects the spatial distribution uniformity of the humidity of the scanning area, and the value range is between 0 and 1, and 1 indicates complete uniformity.
[0068] A non-contact paper moisture content measuring instrument (for example, a sensor based on the principle of capacitance or microwave) is used to measure at a plurality of preset points on the surface of the document to be scanned. Paper moisture content measurement is easily affected by ambient humidity, especially when the humidity is not uniform. Using the humidity uniformity coefficient calculated in the previous step, the system can adjust the measurement strategy (for example, increase the sampling point density, preferentially sample in areas with small humidity change gradient) or correct the original measurement values to reduce the error caused by humidity unevenness. The average value of the corrected multi-point moisture content measurement values is calculated to obtain the average paper moisture content, in percentage (%).
[0069] The document surface and the surrounding area are scanned using an infrared temperature sensor array or a thermal imaging camera to obtain multiple temperature measurements. These temperature measurements are recorded in a temperature value array. The standard deviation of this array (σ_temp_distribution) is calculated. The temperature fluctuation index is obtained by comparing this standard deviation with a pre-set reference standard deviation and normalizing it, for example, temperature fluctuation index = σ_temp_distribution / reference temperature standard deviation. This index reflects the degree of unevenness of the temperature distribution in the scanned area, and the higher the value, the greater the fluctuation.
[0070] The concentration of particulate matter in the air of the scanned area is measured using a laser scattering particulate matter sensor in micrograms per cubic meter (μg / m 3 ). Fluctuations in air temperature can affect air flow, leading to uneven distribution of suspended particulate matter or greater susceptibility to being picked up by air currents during scanning. The temperature fluctuation index calculated in the previous step is used to weight or correct the measured particulate matter concentration to calculate a temperature-particulate matter correlation factor. For example, temperature-particulate matter correlation factor = particulate matter concentration × (1 + W × temperature fluctuation index), where W is a pre-set weight coefficient for the influence of temperature fluctuations on particulate matter. This factor reflects the enhanced potential for environmental temperature fluctuations to interfere with particulate matter.
[0071] In the on state of the scanning light source, a high-resolution camera is used to take an image of the document surface. Particulate matter in the air settles on the paper surface or is suspended in the scanning path, forming additional light reflection points or scattering light, which affects the quality of the scanned image. These additional light reflection points appear as high-light areas in the image. Image processing algorithms (such as threshold-based segmentation, edge detection combined with connected component analysis) are used to detect and quantify the high-light reflection areas in the image. The position, size and brightness information of the detected light reflection areas are plotted into a two-dimensional map, i.e. a light reflection interference distribution map, which can be a matrix corresponding to the size of the document image, with matrix element values representing the degree of light reflection interference in the corresponding area. The temperature-particulate matter correlation factor calculated in the previous step can be used to guide the sensitivity adjustment of the light reflection detection algorithm or to weight and correct the interference degree in the map, for example, when the correlation factor is high, the detection sensitivity of small or low-brightness light reflection points is appropriately increased.
[0072] The light reflection interference distribution map obtained in the previous step is quantitatively analyzed to calculate the overall dust interference degree. For example, the total area of non-zero elements, the average brightness, or the weighted sum of these values in the map can be calculated. This quantitative value is standardized into a single value, i.e. the dust interference factor, which usually ranges from 0 to 1, with a higher value indicating more serious light reflection interference caused by dust. For example, dust interference factor = standardization (total area of light reflection interference regions × average light reflection brightness).
[0073] The paper average moisture content, temperature-particle correlation factor and dust interference factor are integrated to obtain an environmental interference data set. The calculated paper average moisture content, temperature-particle correlation factor and dust interference factor are packaged into a data structure as an environmental interference data set output, which is used for subsequent document physical quality index calculation.
[0074] Preferably, the region adaptive parameter calculation of the document quality distribution map in step S2 includes:
[0075] The document quality distribution map is subjected to high-risk region identification to obtain a region risk assessment table.
[0076] The base resolution parameter value is extracted from the paper adaptive parameter table, and the resolution parameter of each high-risk region in the region risk assessment table is optimized to obtain a region resolution mapping table.
[0077] According to the document quality distribution map, uneven illumination compensation processing is performed to obtain a region contrast mapping table.
[0078] According to the document quality distribution map, flatness focal length adjustment is performed to obtain a region focal length mapping table.
[0079] The region resolution mapping table, the region contrast mapping table and the region focal length mapping table are used to generate a smooth transition mapping table.
[0080] The smooth transition mapping table and the paper adaptive parameter table are used to solve parameter conflicts to obtain a non-conflict parameter scheme.
[0081] The non-conflict parameter scheme and the paper adaptive parameter table are used to generate a region parameter adjustment matrix.
[0082] In the embodiment of the application, the document quality distribution map contains local quality scores and high-risk region labels of each region of the document. The system traverses all regions (for example, 9x9 grid) in the map, identifies regions with local quality scores lower than a preset threshold (for example, 40 points), and marks these regions as high-risk regions. For each region, its grid coordinates and corresponding local quality score are recorded. These information is arranged into a region risk assessment table, such as a list or matrix containing region coordinates and risk levels (divided according to scores).
[0083] Extract the global baseline resolution parameter value (e.g., 300 dpi) from the paper adaptation parameter table. For each high-risk area in the area risk assessment table, optimize the resolution parameter for that area to obtain an area resolution mapping table. For areas marked as high-risk, the system increases their scan resolution parameter by a fixed percentage (e.g., 20%), for example, from 300 dpi to 360 dpi. For non-high-risk areas, maintain the baseline resolution parameter value. Generate an area resolution mapping table corresponding to the document grid, where each element represents the scan resolution value that should be used for the corresponding area.
[0084] Based on the document quality distribution map, uneven illumination compensation is performed to obtain a regional contrast mapping table. The document quality distribution map reflects the illumination uniformity of each region. For regions with low illumination uniformity coefficients (e.g., below 0.8), the system calculates the required increase in contrast compensation value. The more uneven the illumination, the more contrast compensation is needed. For example, an inverse mapping function can be used to map the illumination uniformity coefficient to contrast compensation values (e.g., uniformity 0.5 maps to a 20% increase in contrast, uniformity 0.9 maps to a 5% increase in contrast). A regional contrast mapping table corresponding to the document grid is generated, where each element represents the contrast adjustment value (relative to the base contrast) that should be applied to the corresponding region.
[0085] Based on the document quality distribution map, flatness focal length adjustment is performed to obtain a region focal length mapping table. The document quality distribution map includes the flatness index of each region. For regions with a low flatness index (e.g., below 0.7), it indicates the presence of wrinkles or warping, requiring adjustment of the scanning device's focal length to ensure clear imaging. The system calculates the required focal length adjustment based on the flatness index. The worse the flatness, the greater the focal length adjustment. A region focal length mapping table corresponding to the document grid is generated, where each element represents the focal length adjustment value (relative to the base focal length) to be applied to the corresponding region.
[0086] A smooth transition mapping table is generated based on the region resolution mapping table, region contrast mapping table, and region focal length mapping table. Directly applying discrete region parameters causes abrupt changes at region boundaries in the scanned image, affecting aesthetics and subsequent processing. To avoid this, the three region mapping tables are smoothed. For example, Gaussian filtering or bilinear interpolation algorithms are used to create smooth transition regions between different region parameter values. In non-high-risk regions near high-risk regions, the resolution, contrast, and focal length parameters smoothly transition to the values of the high-risk regions, rather than changing abruptly. A smooth transition mapping table containing the smoothed resolution, contrast, and focal length parameter values is generated; this is actually a matrix containing multiple parameter layers.
[0087] Parameter conflicts are resolved using the smooth transition mapping table and the paper adaptation parameter table to obtain a conflict-free parameter scheme. The smooth transition mapping table provides local optimization parameters for each region, while the paper adaptation parameter table provides global basic parameters applicable to the entire sheet of paper. Conflicts or inconsistencies may exist between local optimization parameters and global parameters. For example, the global parameters may set a maximum contrast ratio, while the contrast compensation calculated by local optimization may exceed this maximum value. The system needs to check all parameter values in the mapping table and compare them with global constraints (e.g., minimum, maximum, and allowable variation range) in the paper adaptation parameter table. For cases exceeding constraints or where conflicts exist, preset conflict resolution rules (e.g., prioritizing stricter constraint values, taking a weighted average of global and local optimization values) are applied for adjustment to obtain a set of conflict-free parameter schemes that meet the constraints in all regions.
[0088] A region parameter adjustment matrix is generated based on the conflict-free parameter scheme and the paper adaptation parameter table. Local optimization parameters (adjusted or final values for resolution, contrast, focal length, etc.) from the conflict-free parameter scheme are integrated with other global parameters (e.g., brightness, color balance, sharpening, etc.) from the paper adaptation parameter table. The paper adaptation parameter table provides baseline parameters applicable to the entire sheet of paper, while the conflict-free parameter scheme provides adjustment or coverage values for different regions. A multi-layer matrix is generated, where each layer represents a scanning parameter (e.g., resolution, contrast, brightness, focal length, color balance, sharpening, etc.), and each element of the matrix corresponds to a region in the document grid, storing the final scanning parameter value to be used for that region. This multi-layer matrix is the region parameter adjustment matrix, which guides the scanning device to apply different scanning parameters to different regions of the document.
[0089] Preferably, step S2, which involves identifying and enhancing key information regions of the document quality distribution map, includes:
[0090] Pre-scanning feature extraction is performed on the region parameter adjustment matrix to obtain a document feature hierarchy map;
[0091] Stamp region detection is performed on the document feature hierarchy map based on the document quality distribution map to obtain stamp region data; stamp enhancement strategy is formulated based on the stamp region data to obtain stamp enhancement parameters;
[0092] The handwritten signature region is identified by analyzing the document feature hierarchy map to obtain signature region data; signature parameters are then enhanced by analyzing the signature region data to obtain signature enhancement parameters.
[0093] Perform table border detection on the document feature hierarchy map to obtain table border data; generate border enhancement parameters based on the table border data.
[0094] The key region priority table is obtained by performing key region priority sorting on the seal region data, signature region data and table border data;
[0095] The key region parameter layer is obtained by performing key region parameter integration on the seal enhancement parameter, signature enhancement parameter and border enhancement parameter according to the key region priority table;
[0096] The key region enhancement matrix is generated according to the key region parameter layer and the region parameter adjustment matrix.
[0097] In the embodiment of the present application, the region parameter adjustment matrix is pre-scanned to extract features, and a document feature hierarchical graph is obtained. The system performs a fast and low-resolution document pre-scanning by using the basic or low-resolution parameter configuration determined in the region parameter adjustment matrix, and obtains a pre-scanning image. Then, a series of image processing and analysis algorithms are applied to the pre-scanning image to extract features. These algorithms include but are not limited to edge detection (such as Canny operator), texture analysis (such as Gabor filter), color segmentation (such as threshold segmentation based on HSV or Lab color space), shape detection (such as circle detection, straight line detection) and layout analysis (such as projection-based or connected domain analysis). The extracted feature information is organized into a multi-layer structure to form a document feature hierarchical graph. The graph represents different types of features at different levels, such as pixel-level edge intensity or color value at the bottom layer, basic shapes and patterns such as lines, curves, texture blocks at the middle layer, and potential seal, signature, table and other complex structure candidate regions at the high layer. The graph is aligned with the grid division of the document, and the feature information of each region is stored independently or in association.
[0098] Then, the document feature hierarchy is subjected to seal region detection based on the document quality distribution map, and seal region data is obtained. Using the color (especially red or carmine), shape (circular, elliptical), and texture features extracted from the document feature hierarchy, the system analyzes the pre-scanned image using pattern recognition algorithms (such as support vector machine SVM, convolutional neural network CNN model) to identify potential seal regions. Meanwhile, combining the document quality distribution map obtained in step S15, the system prioritizes or focuses on analyzing regions with lower quality but red color or complex texture features, as seals often result in poor quality due to ink penetration or uneven stamping pressure. The position coordinates, approximate size, shape information, and preliminary recognition confidence of the detected seal regions are recorded, forming the seal region data. Based on the seal region data, seal enhancement strategies are developed, and seal enhancement parameters are obtained. For example, for the detected seal region, the system develops enhancement strategies based on its color features (judging whether it is red ink or blue seal) and the quality score of the region in the document quality distribution map. If the seal color is light or the region quality is low, the contrast or sensitivity of the red channel is increased; if there is ink penetration, a deblurring algorithm is applied; if the edges are not clear, edge sharpening is performed. These specific image processing parameters for the seal region are determined, forming the seal enhancement parameter set.
[0099] Subsequently, the document feature hierarchy is subjected to handwritten signature region recognition, and signature region data is obtained. Using the line, ink thickness variation, and connected-pen features extracted from the document feature hierarchy, the system searches for handwritten signature patterns in the pre-scanned image using algorithms in the field of handwriting recognition or signature verification (such as based on stroke sequence comparison, shape context descriptor). These algorithms usually require a training set to recognize different styles of handwriting. The position coordinates, bounding box, ink color, and preliminary recognition confidence of the detected handwritten signature region are recorded, forming the signature region data. The signature region data is subjected to signature parameter enhancement, and signature enhancement parameters are obtained. For example, for the detected signature region, the system develops enhancement strategies based on the ink color (black, blue) and the quality score of the region in the document quality distribution map. If the ink is thin or the region quality is low, line thickening or contrast enhancement algorithms are applied; if the paper texture interference is severe, noise removal or texture suppression algorithms are applied; if the signature is tilted or distorted, affine transformation correction is performed. These specific image processing parameters for the signature region are determined, forming the signature enhancement parameter set.
[0100] Meanwhile, table border detection is performed on the document feature hierarchy map to obtain table border data. Using the extracted straight line segments, intersection points and regular arrangement structure information in the document feature hierarchy map, the system uses table structure analysis algorithms (e.g. Hough transform based straight line detection, connected component analysis to identify cells) to identify table border and cell structure in the pre-scanned image. The detected table border line positions, thickness, color and cell division information are recorded to form table border data. Border enhancement parameters are generated based on the table border data. For example, for a detected table border, the system formulates an enhancement strategy based on the border's clarity (evaluated from edge strength in the feature hierarchy map) and the quality score of the region in the document quality distribution map. If the border is blurred or broken, edge enhancement algorithms (e.g. non-maximum suppression, hysteresis thresholding) or morphological closing operations are applied to connect broken lines, ensuring the integrity of the table structure. These specific image processing parameters for table border are determined to form a set of border enhancement parameters.
[0101] Next, the stamp region data, signature region data and table border data are sorted by key region priority to obtain a key region priority table. In some documents, stamps, signatures and tables overlap. In order to handle such overlapping cases during parameter enhancement, the priority of different types of key regions needs to be determined. The system assigns different processing priorities to stamps, signatures and table borders according to pre-set business rules or user configurations (e.g. stamp priority is highest, signature is second, and table border is lowest). When multiple key regions overlap, the enhancement parameters of the region with higher priority will be applied preferentially or superimposed on the overlapping region. Each detected key region and its corresponding priority information are sorted into a key region priority table.
[0102] According to the key region priority table, the stamp enhancement parameters, signature enhancement parameters and border enhancement parameters are integrated into a key region parameter layer. The system iterates through all detected key regions. For each region, the corresponding enhancement parameter set (stamp enhancement parameters, signature enhancement parameters, border enhancement parameters) is found according to its type (stamp, signature, table border). If there is overlap between regions, the final parameter combination applied to the overlapping part is determined according to the key region priority table, for example, the parameters of the region with the highest priority are used, or the parameters of different regions are weighted and averaged or logically combined. These final local enhancement parameters (e.g. local contrast adjustment, local sharpening intensity, local color enhancement factor) determined for each key region (including the region after handling the overlap) are integrated into a key region parameter layer corresponding to the document grid. This layer is a matrix, where non-zero elements represent additional or modified scanning parameters to be applied to the corresponding region.
[0103] Finally, the key region enhancement matrix is generated based on the key region parameter layer and the region parameter adjustment matrix. The region parameter adjustment matrix contains the basic region adaptive scanning parameters calculated based on the overall document quality and the region quality distribution. The key region parameter layer contains more refined local enhancement parameters for specific seal, signature, and table border regions. The system superimposes or replaces the parameter values in the key region parameter layer to the corresponding key region positions in the region parameter adjustment matrix. For example, if the region parameter adjustment matrix specifies the contrast of a region as C, and the key region parameter layer specifies that the contrast of this region (which happens to be a seal region) needs to be increased by AC, then the contrast parameter of this region in the final key region enhancement matrix will be C+AC. If the key region parameter layer specifies the absolute value of a parameter (such as the forced sharpening intensity S), then the sharpening parameter of the corresponding region in the region parameter adjustment matrix is directly replaced. The integrated multi-layer matrix is the key region enhancement matrix, which contains all the optimized, region-adaptive scanning parameters required for scanning all regions of the document, including key information regions.
[0104] Preferably, the step S2 acquires the scanned result image, and simultaneously integrates the paper adaptation parameter table and the key region enhancement matrix to perform scanning quality evaluation, which includes:
[0105] Integrating the paper adaptation parameter table and the key region enhancement matrix to generate scanning execution parameters, to obtain a scanning execution scheme;
[0106] Obtaining real-time scanning feedback data; performing real-time parameter fine-tuning control on the scanning execution scheme according to the real-time scanning feedback data, and synchronously acquiring image data to obtain parameter fine-tuning instructions and a scanned result image, and sending the parameter fine-tuning instructions to the scanning device in real time to complete the closed-loop control;
[0107] Calculating the deviation value of the expected quality of the scanned result image to obtain quality deviation data;
[0108] Integrating the scanned result image and the quality deviation data to obtain a scanning quality score atlas.
[0109] In the embodiments of the present application, the paper adaptation parameter table and the key area enhancement matrix are integrated to generate a scanning execution parameter set, which is used to guide the scanning device to scan the document. The system reads the paper adaptation parameter table, which contains global scanning parameter baseline values applicable to the whole document, such as basic resolution, brightness, contrast, color balance, etc. Meanwhile, the system reads the key area enhancement matrix, which provides local parameter adjustment or override values for each grid area or specific key information area (such as seal, signature, table border) in the document. A parameter integration module combines the parameters from the two sources. For each grid area of the document, the module first adopts the global parameters in the paper adaptation parameter table as the basis. Then, it checks whether there are specific parameter values for the area in the key area enhancement matrix. If there are, the module calculates the final scanning parameter set to be used for the area according to the preset merging rules (for example, the local adjustment value is superimposed on the global baseline value, or the local value completely overrides the global value). For example, if the global contrast is 50 and the key area enhancement matrix specifies that the contrast of a certain area needs to be increased by 10, then the final contrast parameter of the area is 60. If the matrix directly specifies that the resolution of a certain area is 400 dpi, then the resolution parameter of the area is 400 dpi, ignoring the global setting. This process generates a complete parameter configuration set covering all areas of the document, and formats it into a command sequence recognizable by the scanning device driver or firmware, forming a scanning execution scheme that guides the scanning device to apply different scanning parameters to different areas of the document during scanning.
[0110] Real-time scanning feedback data is acquired. Real-time parameter fine-tuning control is performed on the scanning execution plan according to the real-time scanning feedback data, and image data is synchronously collected, to obtain parameter fine-tuning instructions and scanning result images. The parameter fine-tuning instructions are sent to the scanning device in real time, to complete closed-loop control. The scanning device receives and starts to execute the scanning execution plan. In a scanning system supporting real-time feedback and parameter adjustment, a scanning head or a sensor continuously collects data during movement. For example, real-time scanning feedback data such as average light intensity, contrast, definition, and the like of a current scanning line can be acquired through an auxiliary light sensor in front of the scanning head or rapid analysis of a very narrow image strip that has been scanned. A real-time control module in the system receives the feedback data, compares the feedback data with expected parameter values or expected image quality indicators set in the scanning execution plan for a current scanning area, and calculates a real-time quality deviation. For example, if the current area should have high contrast according to the plan, but the real-time feedback shows that the contrast is low, a negative deviation value is calculated. A fast-response control algorithm (for example, based on a lookup table or simple proportional control) uses the deviation value to generate parameter fine-tuning instructions, for example, to increase scanning light source intensity, adjust exposure time, or slightly change focal length. The parameter fine-tuning instructions are sent back to the scanning device in real time through a high-speed interface, and the device immediately adjusts scanning parameters according to the instructions, thereby forming a closed-loop control loop to dynamically correct transient fluctuations in the scanning process. At the same time, the scanning device outputs a complete image data stream after the real-time adjustment, and the system assembles the image data stream into a final scanning result image.
[0111] Quality deviation data is calculated by calculating deviation values of expected quality of the scanning result image. After the complete scanning result image is acquired (or accumulated while being processed in real time), the system performs detailed quality analysis on the scanning result image. The scanning result image is divided into the same grid (for example, 9x9) as the document quality distribution map and the key area enhancement matrix. For each area, actual image quality indicators are calculated, for example, a sharpness score is calculated using a Laplacian operator, contrast is analyzed using a pixel histogram, average brightness is measured, and the like. At the same time, according to the parameters set for the area in the key area enhancement matrix and the original document physical quality index or area quality score, expected image quality indicator values that the area should reach in an ideal case are predicted or looked up. For example, if the original quality is low but a high enhancement parameter is applied, the expected quality should be significantly higher than the original prediction. The difference or ratio between the actual image quality indicator value and the expected image quality indicator value of each area is calculated as the quality deviation value of the area. The matrix or data structure composed of these deviation values is the quality deviation data. A positive deviation indicates that the actual quality is better than expected, and a negative deviation indicates that the actual quality is worse than expected.
[0112] The scan result image and the quality deviation data are integrated to obtain a scan quality score map. The system uses the scan result image as a basic visualization carrier. The region division information and key information region markers contained in the key region enhancement matrix are superimposed on the image layout. In combination with the quality deviation data calculated in the previous step, the system calculates the final scan quality score for each region. This score takes into account the original document quality (implicit in the parameter settings of the key region enhancement matrix), the effect of the application parameters (reflected by the quality deviation data), and the importance of the key regions. For example, the final scan quality score of a region can be calculated according to the formula: Scan quality score = W1 x (expected quality score + W2 x quality deviation value) + W3 x key region importance score, where W1, W2, and W3 are weights, the expected quality score is based on the target after the application of parameters, and the key region importance score reflects whether the region contains seals, signatures, etc. The scan quality score of each region is mapped to a color code to generate a color-coded scan quality heat map, which is superimposed on the scan result image. The identified key information regions are clearly marked on the map, and their specific scan quality scores or levels (e.g., excellent, good, etc.) are displayed. The scan quality heat map, key region markers, and quality level information are integrated to form the final scan quality score map, which visually displays the image quality status of each region of the scanned document, especially the actual effect after adaptive parameter optimization.
[0113] Preferably, the document quality partitioning OCR processing based on the scan quality score map in step S3 comprises:
[0114] Based on the scan quality score map, the document is regionally divided and the processing order is determined to obtain document region data, wherein the document region data includes high-quality regions, medium-quality regions, and low-quality regions.
[0115] Based on the scan result image, the text types in the document region data are identified, and OCR algorithm matching is performed for each text type to obtain a region algorithm matching table.
[0116] According to the region algorithm matching table and the environmental noise data, standard region OCR processing is performed on the high-quality regions to obtain standard region recognition results.
[0117] According to the region algorithm matching table and the environmental noise data, enhanced region OCR processing is performed on the medium-quality regions to obtain enhanced region recognition results.
[0118] According to the region algorithm matching table and the environmental noise data, difficult region multi-algorithm processing is performed on the low-quality regions to obtain difficult region recognition results.
[0119] In the embodiments of the present application, the document is divided into regions and the processing order is determined based on the scan quality score map, and the document region data is obtained, wherein the document region data includes high-quality regions, medium-quality regions and low-quality regions. The system reads the scan quality score of each grid region in the scan quality score map. A quality threshold is set, and regions with scores greater than or equal to 80 points are classified as high-quality regions; regions with scores greater than or equal to 40 points and less than 80 points are classified as medium-quality regions; and regions with scores less than 40 points are classified as low-quality regions. All grid regions in the document are grouped according to their quality categories, and the coordinate range of each region is recorded. At the same time, the system determines the processing order of the regions according to the preset strategy, for example, high-quality regions are processed first, medium-quality regions are processed second, and low-quality regions are processed last. The results of this division and sorting are integrated into the document region data, which records the quality category, position and order in the OCR processing queue of each region.
[0120] Based on the scan result image, the text type in the document region data is identified, and the OCR algorithm matching table is obtained for each text type. The system traverses each region defined in the document region data and performs content analysis in the corresponding scan result image region. Using image analysis and pattern recognition techniques, such as classifiers based on local features (such as stroke shape, character structure) and global features (such as line spacing, character density, color), the main text type contained in the region is identified, such as printed text, handwritten text, number string, table cell content, seal text, etc. According to the identified text type and the quality category (high quality, medium quality, low quality) to which the region belongs, the system matches the most suitable OCR algorithm or algorithm combination from the preconfigured OCR algorithm library. For example, high-quality printed text matches a general high-precision OCR engine; medium-quality handwritten text matches a specialized handwriting recognition engine; and low-quality numbers match multiple number recognition algorithms for cross-validation. The coordinates of each region, the identified text type, and the recommended OCR algorithm (or algorithm combination) are recorded to form the region algorithm matching table.
[0121] High-quality regions are processed by standard region OCR according to the region algorithm matching table and ambient noise data, resulting in standard region recognition results. The system processes high-quality regions first according to the processing order determined in the document region data. For each high-quality region, the system looks up the standard OCR algorithm matched for it from the region algorithm matching table. The standard OCR algorithm is invoked to perform text recognition on the corresponding region in the scanned result image. During the OCR processing, the system records the ambient noise data of the region processing period (e.g., the processing start and end time stamps, noise intensity, type, and other information associated in the ambient noise data) synchronously. The OCR algorithm outputs the recognized text content and the base confidence score of each character. These recognition results, location information, character confidence, and associated ambient noise data are integrated into the standard region recognition results.
[0122] Medium-quality regions are processed by enhanced region OCR according to the region algorithm matching table and ambient noise data, resulting in enhanced region recognition results. The system then processes medium-quality regions. For each medium-quality region, the system looks up the enhanced OCR algorithm combination matched for it from the region algorithm matching table. Before performing OCR, the system performs targeted pre-processing enhancement on the image according to the scanning quality score of the region and the recognized text type, such as moderate image sharpening, local contrast adjustment, or applying a specific filter to suppress paper texture or slight stains. Then, the enhanced OCR algorithm combination is invoked to perform text recognition on the pre-processed image region. Again, the ambient noise data of the processing period is recorded and associated with the recognition results. Enhanced OCR algorithms usually provide more detailed recognition information, such as multiple recognition candidates, character-level feature scores, etc. These enhanced recognition results, location information, more refined confidence scores (taking into account the pre-processing effect), and associated ambient noise data are integrated into the enhanced region recognition results.
[0123] According to the region algorithm matching table and the environmental noise data, the low-quality region is processed by a plurality of algorithms of difficult regions, and a difficult region identification result is obtained. The system finally processes the low-quality region. For each low-quality region, the system searches for a plurality of sets of OCR algorithms (for example, algorithm A, algorithm B, algorithm C) matched therewith from the region algorithm matching table. The system calls these different OCR algorithms in parallel or in series to independently identify the corresponding regions in the scanned result image. For the identification result of each algorithm, the environmental noise data during the processing period thereof is recorded. Then, the system compares and fuses the identification results of different algorithms. For example, for the same character position, the identification results given by algorithms A, B and C are compared. A voting mechanism is adopted. If two or more algorithms have the same identification result, the result is adopted. If the identification results of all algorithms are inconsistent, the character is marked as a high-risk character. The consistency score of each identification result (for example, how many algorithms give the same result) is calculated. The best identification result after fusion, position information, character consistency score, high-risk mark and associated environmental noise data are integrated into a difficult region identification result.
[0124] Preferably, the artificial environmental adaptability analysis according to the environmental noise data in step S3 comprises:
[0125] The noise timing segmentation processing is performed on the environmental noise data to obtain a noise feature table;
[0126] The attention influence table is calculated according to the noise feature table and a preset attention interference coefficient;
[0127] The artificial review risk assessment is performed according to the attention influence table to obtain an artificial review risk table;
[0128] The OCR time correlation is established according to the artificial review risk table to obtain a block environment correlation table;
[0129] The environmental conversion point detection is performed according to the block environment correlation table and the noise feature table to obtain an environmental conversion point list;
[0130] The environmental influence time axis is constructed according to the block environment correlation table, the environmental conversion point list and the artificial review risk table.
[0131] In the embodiments of the present application, the environmental noise data is subjected to noise timing segmentation processing to obtain a noise feature table. The environmental noise data contains the environmental noise intensity values (in decibels dB) continuously collected during the scanning operation, noise type identification (for example, environmental background sound, human voice conversation, equipment operation noise) and their corresponding time stamps. The system analyzes the curve of noise intensity change over time, and uses signal processing techniques (for example, sliding window averaging, peak detection, change point detection algorithm) to identify the moments when the noise level or the main noise type changes significantly. These moments are determined as timing segmentation points. The entire scanning duration is divided into a series of consecutive time periods according to these segmentation points. For each time period, the duration, the average noise intensity, the maximum noise intensity, the main noise type and the frequency or duration ratio of noise exceeding the preset threshold (for example, 65 dB) in the time period are calculated. These calculated feature values are recorded in a table or data structure together with the start and end time stamps of the time period, forming a noise feature table. Each row of the noise feature table represents a noise stable time period and its descriptive features.
[0132] According to the noise feature table and the preset attention interference coefficient, an attention impact table is calculated. The preset attention interference coefficient is a knowledge base or lookup table, which quantifies the interference degree of different noise types and intensities on human attention concentration and task execution based on the research results of psychology and human factors engineering. For example, continuous low-intensity white noise has a low interference coefficient, while intermittent high-intensity human voice conversation has a high interference coefficient. The system traverses each time period in the noise feature table. For the noise features (average intensity, peak value, type, etc.) of the time period, the system queries the attention interference coefficient library to calculate the attention interference score of the time period. The calculation method can combine multiple noise features and corresponding coefficients by weighting, for example, attention interference score = f (average noise intensity, basic interference coefficient corresponding to the main noise type, peak intensity enhancement on score, high-intensity noise duration ratio). The calculated attention interference score (for example, 0-100 points, the higher the score, the greater the interference) is recorded together with the corresponding time period information to form an attention impact table.
[0133] The risk assessment is manually reviewed according to the attention impact table, and a manual review risk table is obtained. The attention impact table provides attention distraction scores for different time periods during the scanning operation. The system maps these scores to a risk level or probability of introducing errors in the manual review task. This mapping relationship is based on historical data analysis (e.g., error rate statistics of manual review under different noise environments) or empirical models. For example, time periods with attention distraction scores below 30 points can be evaluated as low risk, 30-60 points as medium risk, and above 60 points as high risk. Alternatively, a function is used to convert the score to the expected error probability, for example, expected error probability = g(attention distraction score), where g is a monotonically increasing function. The evaluation results are recorded in a table or data structure, forming the manual review risk table, in which each time period is associated with its evaluated manual review risk level or expected error probability.
[0134] According to the manual review risk table, the OCR time association is established, and a block environment association table is obtained. The OCR processing result contains the location coordinates of each recognized data item (e.g., a word, a number, a field) in the original scanned image and its timestamp or time range processed by the OCR engine. The system traverses all the OCR recognition results. For each data item, the system looks up the time period to which the timestamp belongs and its corresponding manual review risk (risk level or probability) in the manual review risk table according to the processing timestamp. The location information, recognition result of the data item, and the manual review risk information corresponding to the processing period are associated to form the block environment association table. Each row of the block environment association table represents an OCR recognized data item and records the data item processed under what potential manual review risk environment.
[0135] According to the block environment association table and the noise feature table, the environment transition point detection is performed, and an environment transition point list is obtained. The block environment association table records the processing time of each data item. The noise feature table defines the time points at which the noise environment occurs. The system analyzes the timestamp distribution of the data items in the block environment association table and compares it with the time period boundaries in the noise feature table. Those time stamps that fall exactly on the time period boundaries of the noise feature table, or the time stamps of the data items processed near these boundaries, are identified. These time points mark the transition of the processing environment from one noise feature segment to another, and are therefore called environment transition points. These transition points are crucial for understanding the impact of environmental changes on continuous processing. The timestamps of all detected environment transition points are collected to form the environment transition point list.
[0136] According to the block environment association table, the environment transition point list and the artificial review risk table, an environment impact timeline is constructed. The environment impact timeline is a comprehensive data structure that visually or structurally presents the changes of the processing environment (especially the noise environment and its potential impact on artificial review) over time during the entire document scanning and OCR processing process, and is associated with specific processed data items. The system constructs a timeline starting from the scanning start time and ending at the scanning end time. Important environmental change moments are marked on the timeline using the timestamps in the environment transition point list. For each time period on the timeline (divided by the transition points), the system queries the artificial review risk table to mark the artificial review risk level or expected error probability corresponding to the period. At the same time, the system traverses the block environment association table to link each data item at its processing time point to the timeline and display its associated environmental risk information. This environment impact timeline intuitively shows which data items are processed during high-risk environmental periods, providing key evidence for subsequent risk-aware report generation and error tracing.
[0137] Preferably, step S4 comprises the following steps:
[0138] Step S41: Risk data screening and grading according to the data credibility matrix, to obtain a risk data grading result;
[0139] Step S42: Constructing a data traceability chain according to the risk data grading result, the document physical quality index, the scanning quality score atlas and the data credibility matrix;
[0140] Step S43: Generating a digital employee intelligence report template according to the data traceability chain and the risk data grading result;
[0141] Step S44: Dynamically optimizing the report content of the digital employee intelligence report template to obtain a dynamically optimized report;
[0142] Step S45: Generating a process optimization suggestion report according to the dynamically optimized report.
[0143] In the embodiments of the present application, risk data screening and classification are performed according to the data credibility matrix to obtain a risk data classification result. The system sets three credibility thresholds: a high-risk threshold (e.g., 40 points), a medium-risk threshold (e.g., 55 points), and a low-risk threshold (e.g., 70 points). The system iterates through each data item in the data credibility matrix. For each data item, its associated credibility score is read. If the credibility score is less than or equal to the high-risk threshold, the data item is marked as high-risk data. If the credibility score is greater than the high-risk threshold and less than or equal to the medium-risk threshold, it is marked as medium-risk data. If the credibility score is greater than the medium-risk threshold and less than or equal to the low-risk threshold, it is marked as low-risk data. Data items with credibility scores greater than the low-risk threshold are marked as low-risk or no-risk data (depending on specific requirements). All data items marked as risk (high, medium, and low) and their corresponding risk levels, location coordinates (obtained from the data credibility matrix), and credibility scores are organized into a risk data classification result list. The system also counts the number of data items at different risk levels and analyzes their location distribution on the original layout of the document to generate a risk hotspot map, such as a two-dimensional matrix, whose element values reflect the risk data density or average risk level in the corresponding document area. In addition, in combination with the environmental impact factor in the data credibility matrix, the system calculates the comprehensive risk index of each risk data item, such as comprehensive risk index = W1 x (100 - credibility score) + W2 x standardized (environmental impact factor), where W1 and W2 are preset weights, and standardized (environmental impact factor) is the mapping of the environmental impact factor to a scale of 0-100. The risk data classification result, risk hotspot map, and comprehensive risk index of each risk data item are output together.
[0144] According to the risk data grading results, the document physical quality index, the scanning quality score map and the data credibility matrix, a data provenance chain is constructed. For each data item in the data credibility matrix (whether marked as risk or not), the system assigns it a unique number or combination of letters as provenance code. This provenance code serves as the primary identifier of the data item. The system constructs an associated record structure for each data item. In this structure, in addition to the content, location and credibility score of the data item itself, the system links the following information as nodes on the chain: 1. Physical quality association: linked to the overall document physical quality index (including numerical score and quality level) when the document is scanned. 2. Scanning quality association: according to the location coordinates of the data item, linked to the scanning quality score and level of the corresponding area in the scanning quality score map. 3. Processing method record: linked to the specific OCR algorithm or algorithm combination used when performing OCR processing on the area where the data item is located (obtained from the data credibility matrix or extracted from the detailed processing log of step S3). 4. Environmental impact record: linked to the environmental impact factor or the environmental risk period information of the data item when it is processed by OCR (obtained from the data credibility matrix or queried from the environmental impact timeline). 5. Original parameter record: optionally, linked to the specific scanning parameters used when scanning the area where the data item is located (queried from the key area enhancement matrix). All these associated information is stored in a structured manner, forming a data provenance chain throughout the entire process of data item generation.
[0145] According to the data provenance chain obtained in step S42 and the risk data classification results obtained in step S41, a digital employee intelligence report template is generated. The system loads a basic template from a pre-set report template library according to the business type of the report to be generated (for example, an invoice recognition report, a contract information extraction report). The basic template defines the overall layout, field structure and basic display style of the report. The system traverses all data items contained in the data provenance chain, and fills the recognized text content into the corresponding fields of the report basic template according to their positions and types. During the filling process, the system refers to the risk data classification results. For data items marked as high risk, they are marked in a prominent way in the report, for example, using red font, red background highlighting or adding a risk icon next to them. Medium-risk data items are marked with yellow, and low-risk data items are marked with blue. For data items in the data provenance chain that record environmental impact information (for example, data processed in an evaluation period with high noise), an environmental risk identifier is added next to the data item in the report, such as an ear icon indicating noise interference or a thermometer icon indicating temperature impact. At the same time, the system embeds an interactive element (for example, a hyperlink or a clickable button) on each data item in the report. When the user clicks on the element, the system calls a function that retrieves and displays detailed provenance information for the data item from the data provenance chain, including physical quality, scanning quality, processing method, environmental impact, etc., according to the provenance code of the data item. The report structure filled with data, risk / environmental identifiers and interactive provenance functions is saved as an instance of the digital employee intelligence report template.
[0146] The digital employee intelligence report template obtained in step S43 is dynamically optimized for report content to obtain a dynamically optimized report. The system reads the comprehensive risk index of each risk data item calculated in step S41. According to these indices, the system adjusts the display priority of the report content. Items with higher comprehensive risk indices are prioritized or highlighted. For example, the system creates a "risk data summary" area at the top or side of the report, and arranges all high-risk and medium-risk data items in this area in descending order of comprehensive risk index, making it easy for users to quickly focus on key issues. The system can generate personalized report views according to different user roles or configurations. For example, the quality inspector view focuses on displaying all risk data and provenance information, while the business user view only marks risks in the main report and hides detailed provenance information behind interactive clicks. In addition, the system can generate a report thumbnail or overview, superimpose the risk hotspot map obtained in step S41 on the document layout thumbnail, and embed it on the first page of the report to provide a quick preview of the overall risk distribution of the document. These optimization measures are applied to the digital employee intelligence report template instance to generate the dynamically optimized report that is finally presented to the user.
[0147] The dynamic optimization report obtained according to step S44 generates a process optimization suggestion report. The system analyzes the risk data distribution, risk hotspot area in the dynamic optimization report of the current batch processing, and the risk correlation factors revealed through the data traceability chain. The system compares and analyzes these information with the historical processing data (assuming that the system accumulates the document physical quality index, scanning quality score map, environmental noise data, data reliability matrix, and correction records after manual review of historical batches). Through the statistics and pattern mining of a large amount of historical data, the system identifies common patterns and root causes that lead to data errors or low reliability, such as: a specific type of paper (identified by the paper physical property set) has a high recognition error rate under certain lighting conditions; the proportion of data items processed under environmental noise intensity exceeding a certain threshold is significantly increased by manual correction; there is a persistent low-quality area in the scanning quality score map, and the data in this area is often marked as high-risk. Based on these analysis results, the system automatically generates a process optimization suggestion report. The report content includes: identifying the main risk sources (for example, "highly reflective paper scanned under strong top light" is a common problem), quantifying the impact of these risk sources on data quality (for example, "causing the average reliability of related data items to drop by 15 points"), providing specific improvement suggestions (for example, "suggesting to add side light sources or use soft light covers for the scanning table", or "suggesting to suspend the scanning and OCR processing of sensitive documents during high-noise periods"), and preliminary evaluation of the investment and expected effect of implementing these suggestions.
[0148] Especially important is that step S42 includes:
[0149] generating a data item identification table according to the data reliability matrix;
[0150] performing physical environment traceability correlation on the data item identification table and the document physical quality index to obtain a physical quality correlation table;
[0151] performing scanning process tracking on the scanning quality score map according to the physical quality correlation table to obtain a scanning quality correlation table;
[0152] performing identification algorithm backtracking according to the scanning quality correlation table and the data reliability matrix to obtain a processing method record table;
[0153] performing risk factor integration on the processing method record table and the risk data classification result to obtain a risk traceability table;
[0154] performing traceability path construction and visualization according to the risk traceability table to obtain a data traceability chain;
[0155] In the embodiments of the present application, a data item identification table is generated according to the data confidence matrix. The system traverses all entries in the data confidence matrix. Each entry represents a data item recognized by OCR, containing its location coordinates in the scanned result image, recognized text content, confidence score and environmental impact factor. For each data item, the system generates a unique internal identifier, such as a string generated based on the UUID algorithm, or a serial number related to the document ID and the order of the data item in the matrix. This unique identifier will serve as the primary key of the data item in the entire provenance chain. The system creates a data item identification table, which contains at least two columns: data item unique identifier and the precise location coordinates of the data item in the original scanned result image (e.g. the pixel coordinates of the top-left corner and the bottom-right corner). This table provides a basic index for subsequent provenance association operations.
[0156] The data item identification table and the document physical quality index are associated with the physical environment to obtain a physical quality association table. The system creates a physical quality association table. The table takes the data item unique identifier as the primary key. The system traverses each data item unique identifier in the data item identification table. For each identifier, the system obtains its corresponding document ID (because the data item identifier is generated for a specific document). According to the document ID, the system looks up the overall document physical quality index (including numerical score and quality level) of the document. The system associates the document physical quality index of the document with the corresponding data item unique identifier, and stores the association relationship in the physical quality association table. This association indicates that the data item comes from a document with a specific overall physical quality.
[0157] The scanned quality scoring map is tracked according to the physical quality association table to obtain a scanned quality association table. The system creates a scanned quality association table. The table takes the data item unique identifier as the primary key. The system traverses each data item unique identifier in the physical quality association table. For each identifier, the system looks up the precise location coordinates of the data item in the original scanned result image from the data item identification table. Using the location coordinates, the system looks up the region to which the location belongs in the scanned quality scoring map (which is based on document grid division). The scanned quality score and quality level of the region are extracted from the scanned quality scoring map. The system associates the scanned quality information of the region with the corresponding data item unique identifier, and stores the association relationship in the scanned quality association table. This association indicates that the region where the data item is located has reached a specific image quality level in the scanning link.
[0158] The processing method record table is obtained by tracing the recognition algorithm according to the scan quality association table and the data confidence matrix. The system creates the processing method record table. The table takes the data item unique identifier as the primary key. The system traverses each data item unique identifier in the scan quality association table. For each identifier, the system looks up the detailed information of the data item from the data confidence matrix. The generation process of the data confidence matrix depends on the determined regional algorithm matching table and the regional quality division. According to the location coordinates or the region to which the data item belongs, the system traces which specific OCR algorithm or algorithm combination (for example, the standard algorithm A used in high-quality regions, the algorithm B, C, and D combination used in low-quality regions) is called when the region is processed by OCR in step S3. The system associates these specific OCR algorithm identifiers or descriptions with the corresponding data item unique identifiers, and stores the association relationship in the processing method record table. This association indicates that the data item is the result obtained after using a specific recognition algorithm or algorithm combination for processing.
[0159] The processing method record table is obtained by tracing the recognition algorithm according to the scan quality association table and the data confidence matrix. The system creates the processing method record table. The table takes the data item unique identifier as the primary key. The system traverses each data item unique identifier in the scan quality association table. For each identifier, the system looks up the detailed information of the data item from the data confidence matrix. The generation process of the data confidence matrix depends on the determined regional algorithm matching table and the regional quality division. According to the location coordinates or the region to which the data item belongs, the system traces which specific OCR algorithm or algorithm combination (for example, the standard algorithm A used in high-quality regions, the algorithm B, C, and D combination used in low-quality regions) is called when the region is processed by OCR in step S3. The system associates these specific OCR algorithm identifiers or descriptions with the corresponding data item unique identifiers, and stores the association relationship in the processing method record table. This association indicates that the data item is the result obtained after using a specific recognition algorithm or algorithm combination for processing.
[0160] The risk factor integration is performed on the processing method record table and the risk data grading result obtained in step S41 to obtain a risk traceability table. The system creates the risk traceability table. The table takes the data item unique identifier as the primary key. The system traverses each data item unique identifier in the processing method record table. For each identifier, the system looks up the risk level (high, medium, low) and the comprehensive risk index of the data item from the risk data grading result obtained in step S41. At the same time, the system looks up the environmental impact factor associated with the data item from the data credibility matrix. The system integrates the algorithm information in the processing method record table with the risk level / comprehensive index in the risk data grading result and the environmental impact factor in the data credibility matrix. The information (data item unique identifier, OCR algorithm used, risk level, comprehensive risk index, environmental impact factor) is stored in the risk traceability table. The table collects the identification method of the data item, the evaluated risk, and the environmental associated factors.
[0161] According to the risk traceability table, the traceability path construction and visualization are performed to obtain a data traceability chain. The risk traceability table contains the core information for constructing the traceability chain. The system takes each data item unique identifier in the risk traceability table as the starting or terminal node of the traceability chain. Centering on the node, the system links the relevant traceability information to the data item node by querying the physical quality association table, the scanning quality association table, and the risk traceability table itself. The system saves the structured association information in a database or memory model to form a complete data traceability chain data structure. To realize visualization, the system uses a graph drawing library or a visualization component to render the data traceability chain data structure into a graphical interface. In the graphical interface, the data items are represented by text or icons, different types of traceability information (physical quality, scanning quality, algorithm, risk, environment) are represented by different node types or colors, and the relationships between the nodes are represented by lines. Users can expand or highlight the complete traceability path of a data item by interactive operations (for example, clicking the data item node) to intuitively trace the generation process and influencing factors of the data item. The structured data model and its visualization representation jointly constitute the final data traceability chain.
[0162] Therefore, the embodiments should be considered in all respects as illustrative and not restrictive, the scope of the application being indicated by the appended claims rather than by the description given above, and all changes which come within the meaning and range of equivalency of the claims are therefore intended to be embraced therein.
[0163] The foregoing is considered as illustrative only of the principles of the application. Numerous modifications and changes will readily occur to those skilled in the art, and it is intended to embrace all such modifications and changes that fall within the scope of the application. Accordingly, the application is not to be restricted in scope to the specific embodiments disclosed herein but is to be accorded the full scope that the principles and novel features request appropriately granted.
Claims
1. A method for digital employee intelligence report auto-generation, characterized in that, The method comprises the following steps: Step S1: Collecting a light characteristic data set and a paper physical characteristic set of the document to be scanned by a sensor array; detecting environmental interference factors of the document to be scanned to obtain an environmental interference data set; calculating a document physical quality index according to the light characteristic data set, the paper physical characteristic set and the environmental interference data set, and then generating a document quality distribution map; Step S2: Determining a scanning parameter reference table according to the document physical quality index; Correcting the scanning parameter reference table according to paper characteristic parameters to obtain a paper adaptive parameter table; and calculating regional adaptive parameters of the document quality distribution map to obtain a regional parameter adjustment matrix; Identifying and enhancing a key information region of the document quality distribution map according to the regional parameter adjustment matrix to obtain a key region enhancement matrix; collecting a scanned result image, and simultaneously integrating the paper adaptive parameter table and the key region enhancement matrix to perform scanning quality evaluation to obtain a scanning quality score map; Step S3: Monitoring and recording environmental noise of a document scanning operation by an environmental noise sensor to obtain environmental noise data; performing document quality partition OCR processing based on the scanning quality score map to obtain a partition recognition result set; performing artificial environmental adaptability analysis according to the environmental noise data to obtain an environmental influence time axis; and performing data reliability calculation on the partition recognition result set according to the environmental influence time axis to obtain a data reliability matrix; Step S4: Constructing a data traceability chain according to the document physical quality index, the scanning quality score map and the data reliability matrix; and generating a digital employee intelligent report template according to the data traceability chain; Optimizing a report flow of the digital employee intelligent report template to obtain a flow optimization suggestion report.
2. The digital employee intelligence report auto-generation method of claim 1, wherein, Step S1 comprises the following steps: Step S11: Deploying a multi-point light sensor array to sample light characteristics at four corners and a center position of the document to be scanned to obtain a light characteristic data set; Step S12: Measuring paper physical characteristics of the document to be scanned according to the light characteristic data set to obtain a paper physical characteristic set; Step S13: Detecting environmental interference factors of the document to be scanned to obtain an environmental interference data set; Step S14: Calculating a document physical quality index according to the light characteristic data set, the paper physical characteristic set and the environmental interference data set to obtain the document physical quality index; Step S15: Generating a document quality distribution map according to the document physical quality index, the light characteristic data set and the paper physical characteristic set.
3. The digital employee intelligence report auto-generation method of claim 2, wherein, The detection of environmental interference factors of the document to be scanned in step S13 comprises: Collecting humidity of the document to be scanned at multiple points to obtain a humidity uniformity coefficient; Determining a paper moisture content according to the humidity uniformity coefficient to obtain an average paper moisture content; Detecting a temperature distribution of the document to be scanned to obtain a temperature fluctuation index; Measuring a particulate matter concentration according to the temperature fluctuation index to obtain a temperature-particulate matter correlation factor; Evaluating a light reflection interference according to the temperature-particulate matter correlation factor to obtain a light reflection interference distribution map; Calculating a dust interference factor according to the light reflection interference distribution map; Integrating the average paper moisture content, the temperature-particulate matter correlation factor and the dust interference factor to obtain the environmental interference data set.
4. The method of claim 1, wherein, The region adaptive parameter calculation on the document quality distribution map in step S2 includes: Performing high-risk region identification on the document quality distribution map to obtain a region risk assessment table; Extracting a basic resolution parameter value from the paper adaptive parameter table to perform resolution parameter optimization on each high-risk region in the region risk assessment table to obtain a region resolution mapping table; Performing uneven illumination compensation processing according to the document quality distribution map to obtain a region contrast mapping table; Performing flatness focal length adjustment according to the document quality distribution map to obtain a region focal length mapping table; Generating a smooth transition mapping table according to the region resolution mapping table, the region contrast mapping table, and the region focal length mapping table; Performing parameter conflict resolution according to the smooth transition mapping table and the paper adaptive parameter table to obtain a non-conflict parameter scheme; Generating a region parameter adjustment matrix according to the non-conflict parameter scheme and the paper adaptive parameter table.
5. The digital employee intelligence report auto-generation method of claim 1, wherein, The key information region identification and enhancement on the document quality distribution map in step S2 includes: Performing pre-scanning feature extraction on the region parameter adjustment matrix to obtain a document feature hierarchy map; Performing seal region detection on the document feature hierarchy map according to the document quality distribution map to obtain seal region data; and performing seal enhancement strategy formulation on the seal region data to obtain seal enhancement parameters; Performing handwritten signature region identification on the document feature hierarchy map to obtain signature region data; and performing signature parameter enhancement on the signature region data to obtain signature enhancement parameters; Performing table border detection on the document feature hierarchy map to obtain table border data; and generating border enhancement parameters according to the table border data; Performing key region priority sorting on the seal region data, the signature region data, and the table border data to obtain a key region priority table; Performing key region parameter integration on the seal enhancement parameters, the signature enhancement parameters, and the border enhancement parameters according to the key region priority table to obtain a key region parameter layer; Generating a key region enhancement matrix according to the key region parameter layer and the region parameter adjustment matrix.
6. The digital employee intelligence report auto-generation method of claim 1, wherein, The scanning result image acquisition and the integration of the paper adaptive parameter table and the key region enhancement matrix for scanning quality evaluation in step S2 include: Integrating the paper adaptive parameter table and the key region enhancement matrix to generate a scanning execution parameter, thereby obtaining a scanning execution scheme; Obtaining real-time scanning feedback data; performing real-time parameter fine-tuning control on the scanning execution scheme according to the real-time scanning feedback data, and synchronously acquiring image data, thereby obtaining parameter fine-tuning instructions and a scanning result image; and sending the parameter fine-tuning instructions to a scanning device in real time to complete closed-loop control; Performing expected quality deviation value calculation on the scanning result image to obtain quality deviation data; Integrating the scanning result image and the quality deviation data to obtain a scanning quality score map.
7. The digital employee intelligence report auto-generation method of claim 1, wherein, The document quality partition OCR processing based on the scanning quality score map in step S3 includes: Dividing the document into regions and determining a processing order based on the scanning quality score map to obtain document region data, wherein the document region data includes high-quality regions, medium-quality regions, and low-quality regions; Identifying text types in the document region data based on the scanning result image, and matching an OCR algorithm for each text type to obtain a region algorithm matching table; According to the region algorithm matching table and the environmental noise data, the high-quality region is subjected to standard region OCR processing to obtain a standard region recognition result; According to the region algorithm matching table and the environmental noise data, the medium-quality region is subjected to enhanced region OCR processing to obtain an enhanced region recognition result; According to the region algorithm matching table and the environmental noise data, the low-quality region is subjected to difficult region multi-algorithm processing to obtain a difficult region recognition result.
8. The digital employee intelligent report auto-generation method of claim 1, wherein, The artificial environmental adaptability analysis according to the environmental noise data in step S3 includes: The noise time sequence segmentation processing is performed on the environmental noise data to obtain a noise feature table; The attention influence table is calculated according to the noise feature table and a preset attention interference coefficient; The artificial review risk assessment is performed according to the attention influence table to obtain an artificial review risk table; The OCR time correlation is established according to the artificial review risk table to obtain a block environment correlation table; The environmental conversion point detection is performed according to the block environment correlation table and the noise feature table to obtain an environmental conversion point list; The environmental influence time axis is constructed according to the block environment correlation table, the environmental conversion point list and the artificial review risk table.
9. The digital employee intelligent report auto-generation method of claim 1, wherein, Step S4 includes the following steps: Step S41: performing risk data screening and grading according to the data credibility matrix to obtain a risk data grading result; Step S42: constructing a data traceability chain according to the risk data grading result, the document physical quality index, the scanning quality score graph and the data credibility matrix; Step S43: generating a digital employee intelligent report template according to the data traceability chain and the risk data grading result; Step S44: performing report content dynamic optimization on the digital employee intelligent report template to obtain a dynamic optimization report; Step S45: generating a process optimization suggestion report according to the dynamic optimization report.
10. A digital employee intelligence report auto-generation system, characterized in that, The digital employee intelligent report automatic generation system is used for performing the digital employee intelligent report automatic generation method as claimed in claim 1, and includes: The physical environment perception module is used for collecting an illumination characteristic data set and a paper physical characteristic set of the to-be-scanned document through a sensor array; detecting environmental interference factors of the to-be-scanned document to obtain an environmental interference data set; calculating a document physical quality index according to the illumination characteristic data set, the paper physical characteristic set and the environmental interference data set, and then generating a document quality distribution graph; The intelligent scanning optimization module is used for determining a scanning parameter benchmark table according to the document physical quality index; correcting the scanning parameter benchmark table to obtain a paper adaptation parameter table; calculating region self-adaptive parameters according to the document quality distribution graph to obtain a region parameter adjustment matrix; identifying and enhancing a key region according to the region parameter adjustment matrix to obtain a key region enhancement matrix; collecting a scanning result image, and simultaneously integrating the paper adaptation parameter table and the key region enhancement matrix to perform scanning quality evaluation, thereby obtaining a scanning quality score graph; The intelligent scanning optimization module is used for determining a scanning parameter benchmark table according to the document physical quality index; correcting the scanning parameter benchmark table to obtain a paper adaptation parameter table; calculating region self-adaptive parameters according to the document quality distribution graph to obtain a region parameter adjustment matrix; identifying and enhancing a key region according to the region parameter adjustment matrix to obtain a key region enhancement matrix; collecting a scanning result image, and simultaneously integrating the paper adaptation parameter table and the key region enhancement matrix to perform scanning quality evaluation, thereby obtaining a scanning quality score graph; The environment sensing and identifying module is used for environment noise monitoring and recording of the document scanning operation through an environment noise sensor to obtain environment noise data; performing document quality partition OCR processing based on a scanning quality score map to obtain a partition identification result set; performing artificial environment adaptability analysis according to the environment noise data to obtain an environment influence time axis; and performing data reliability calculation on the partition identification result set according to the environment influence time axis to obtain a data reliability matrix; The risk report generation module is used for constructing a data traceability chain according to the document physical quality index, the scanning quality score map and the data reliability matrix; generating a digital employee intelligent report template according to the data traceability chain; and performing report process optimization on the digital employee intelligent report template to obtain a process optimization suggestion report.