A method and system for generating an MRZ synthetic sample

By generating MRZ synthetic samples that conform to international standards, the problems of data scarcity and incompatibility of learning rate strategies in MRZ recognition models are solved, the recognition rate is improved and computing resources are saved, and it is suitable for document recognition.

CN120910561BActive Publication Date: 2026-02-03BEIJING NJA INFORMATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511017383.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-23
Publication Date
2026-02-03
Estimated Expiration
2045-07-23

AI Technical Summary

Technical Problem

In existing technologies, machine-readable zone (MRZ) recognition models suffer from poor recognition performance due to difficulties in data acquisition, lack of training data, insufficient model generalization ability, and inappropriate learning rate strategies. Furthermore, the convergence speed of pre-trained models is slow, which fails to meet the requirements for high accuracy.

Method used

By acquiring preprocessed corpora and combining synthesis parameters such as ambiguity, font size, character spacing, font thickness, and tilt angle, the image quality changes are simulated to generate MRZ synthetic samples that meet international standards. Based on an intelligent model, the samples are evaluated and optimized to generate samples that meet international standards.

Benefits of technology

It improved the MRZ recognition rate by 7%, saved 30% of computing resources, and the generated samples meet international standards and are suitable for document recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120910561B_ABST
    Figure CN120910561B_ABST
Patent Text Reader

Abstract

The application provides a method and system for generating MRZ synthetic samples, and is applied to the technical field of certificate data processing. The application combines the pre-processed machine-readable area corpus and the preset synthetic parameters, simulates the quality change, information truncation and physical defects of the machine-readable area image, generates parameter combination constraint conditions based on international standards and real MRZ image data; processes based on the synthetic parameter combination, corpus preprocessing, sample proportion control and parameter combination constraint conditions to generate machine-readable area synthetic samples; processes the generated machine-readable area synthetic samples to generate sample-related evaluation results; processes the evaluation results and parameter settings in the synthetic parameters according to the international standards to generate machine-readable area synthetic sample generation bases; processes the parameter combination, sample proportion control, sample-related evaluation results and parameter compliance feature vectors to generate machine-readable area synthetic sample generation optimization results.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of certificate data processing, and in particular to a method and system for generating MRZ synthetic samples. BACKGROUND

[0002] In the field of data processing, the training of a machine-readable zone (MRZ) recognition model faces many challenges. From the perspective of data acquisition, due to factors such as data privacy protection, it is difficult to obtain real samples, resulting in a lack of training data. This makes it difficult for the model to fully cover a variety of recognition scenarios, such as when facing complex situations such as character confusion and special font deformation, the recognition effect is poor.

[0003] In terms of training methods, the traditional approach often uses a fixed proportion of real samples and synthetic samples. This approach cannot be dynamically adjusted according to the training process, making it difficult to balance data expansion and precision optimization, and thus leading to the problem of insufficient generalization ability of the model.

[0004] When fine-tuning a pre-trained model, the conventional learning rate strategy also has obvious defects. For example, the fixed learning rate strategy cannot meet the needs of different stages of the model training process, and the simple linear decay strategy can easily cause gradient shock phenomenon, both of which can lead to slow model convergence speed and limited final accuracy, which cannot meet the requirements of high accuracy in practical applications.

[0005] It should be noted that the information disclosed in the above background section is only used to strengthen the understanding of the background of the present disclosure, and therefore can include information that does not constitute prior art known to those of ordinary skill in the art. SUMMARY

[0006] The purpose of the present application is to provide a method and system for generating MRZ synthetic samples, which at least partially overcomes the problems existing in the prior art. By obtaining preprocessed corpus, synthetic parameters containing ambiguity, and international standards, the image quality is simulated after combination, the parameter association rules are established according to the standards, the conditions containing hard constraints are generated, and then the samples are generated and the compliance is verified. Then extract features, evaluate quality deviation, calculate parameter compliance, and finally analyze risks based on intelligent models, combine historical data and standards to update and generate optimization instructions. The scheme makes the sample comply with international standards, improves the recognition rate by 7%, saves 30% of computing resources, and is suitable for certificate recognition.

[0007] Other characteristics and advantages of the present application will become apparent from the following detailed description, or will be learned by practice of the present application.

[0008] According to an aspect of the present application, a method for generating a machine-readable zone (MRZ) synthetic sample is provided, including: obtaining a MRZ synthetic sample and an international standard, the MRZ synthetic sample including using preprocessed MRZ corpus, specific synthetic parameters, and controlling a proportion of the synthetic sample, wherein the synthetic parameters include blurriness, font size, inter-character spacing, font thickness, and tilt angle, and the corpus preprocessing includes randomly cutting the MRZ corpus in a length range of [5, 14]; the international standard is used to regulate a format, a character set, and a printing quality of a machine-readable area of a travel document; combining the preprocessed MRZ corpus and the preset synthetic parameters, simulating quality variation, information truncation, and physical defects of a MRZ image, establishing a correlation rule between parameters based on the international standard and real MRZ image data, and generating parameter combination constraint conditions; processing based on the synthetic parameter combination, the corpus preprocessing, the sample proportion control, and the parameter combination constraint conditions to generate the MRZ synthetic sample; processing the generated MRZ synthetic sample to generate sample-related evaluation results; processing the evaluation results and parameter settings in the synthetic parameters based on the international standard to generate MRZ synthetic sample generation bases; and processing based on an intelligent model in combination with the parameter combination, the sample proportion control, the sample-related evaluation results, and a parameter compliance feature vector to generate a MRZ synthetic sample generation optimization result.

[0009] According to another aspect of the present application, a device for generating a machine-readable zone (MRZ) synthetic sample is provided, including: an obtaining module configured to obtain a MRZ synthetic sample and an international standard, the MRZ synthetic sample including using preprocessed MRZ corpus, specific synthetic parameters, and controlling a proportion of the synthetic sample, wherein the synthetic parameters include blurriness, font size, inter-character spacing, font thickness, and tilt angle, and the corpus preprocessing includes randomly cutting the MRZ corpus in a length range of [5, 14]; the international standard is used to regulate a format, a character set, and a printing quality of a machine-readable area of a travel document; a processing module configured to combine the preprocessed MRZ corpus and the preset synthetic parameters, simulate quality variation, information truncation, and physical defects of a MRZ image, establish a correlation rule between parameters based on the international standard and real MRZ image data, and generate parameter combination constraint conditions; process based on the synthetic parameter combination, the corpus preprocessing, the sample proportion control, and the parameter combination constraint conditions to generate the MRZ synthetic sample; process the generated MRZ synthetic sample to generate sample-related evaluation results; process the evaluation results and parameter settings in the synthetic parameters based on the international standard to generate MRZ synthetic sample generation bases; and process based on an intelligent model in combination with the parameter combination, the sample proportion control, the sample-related evaluation results, and a parameter compliance feature vector to generate a MRZ synthetic sample generation optimization result.

[0010] According to still another aspect of the present application, an electronic device includes a first processor and a memory storing executable instructions of the first processor, wherein the first processor is configured to perform the above-mentioned method for generating MRZ synthetic samples by executing the executable instructions.

[0011] The method and system for generating MRZ synthetic samples provided by the present application obtain pre-processed corpus (cut 5-14 length), synthetic parameters containing ambiguity, and international standards, combine them to simulate image quality changes, build parameter association rules according to the standards, generate conditions containing hard constraints, and then generate samples and verify compliance. Then, features are extracted, quality deviation is evaluated, parameter compliance is calculated, and finally, risks are analyzed based on an intelligent model, and optimization instructions are generated based on historical data and standards. The scheme makes the samples comply with international standards, improves the recognition rate by 7%, saves 30% of computing resources, and is suitable for certificate identification.

[0012] It should be understood that the foregoing general description and the following detailed description are only exemplary and explanatory, and are not limiting to the present disclosure. BRIEF DESCRIPTION OF DRAWINGS

[0013] Figure 1 A flowchart of a method for generating MRZ synthetic samples according to an embodiment of the present application is shown.

[0014] Figure 2 A structural schematic diagram of a device for generating MRZ synthetic samples according to an embodiment of the present application is shown. DETAILED DESCRIPTION

[0015] The preferred embodiments of the present application are described below in conjunction with the accompanying drawings, and it should be understood that the preferred embodiments described herein are only used to illustrate and explain the present application, and are not used to limit the present application.

[0016] The method for generating MRZ synthetic samples according to the exemplary embodiments of the present application is described below in conjunction with Figure 1 It should be noted that the following application scenarios are only shown for the purpose of facilitating understanding of the spirit and principles of the present application, and the embodiments of the present application are not limited in this respect. On the contrary, the embodiments of the present application are applicable to any applicable scenarios.

[0017] In one embodiment, the present application further provides a method and system for generating MRZ synthetic samples. Figure 1 A flowchart of a method for generating MRZ synthetic samples according to an embodiment of the present application is shown. As shown in Figure 1 the method is applied to a server and includes:

[0018] S101, obtain machine-readable area synthetic samples and international standards, the machine-readable area synthetic samples include using pre-processed machine-readable area corpus, specific synthetic parameters, and controlling the proportion of synthetic samples.

[0019] In an implementation, the machine-readable area corpus is randomly cut into a length between 5 and 14, while the field integrity is preserved through semantic unit identification technology. For example, the original corpus is P<CHEN<<HAO<<<<<<1234567>8901234 (including the name CHEN<<HAO and the ID number 12345678901234), and the field core position is avoided during cutting, such as being cut into P<CHEN<<H (length 7) and AO<<<<<<12345 (length 10), to ensure that the cut string still retains part of the valid semantics, simulating the information truncation scenario caused by real certificate wear.

[0020] The setting of the synthetic parameters (based on the ICAO Doc9303 standard) is as follows: blur: the value range is a Gaussian blur kernel ≤ 5px, for example, the blur is set to 3px, which is used to simulate the image quality degradation caused by the diffusion of printing ink. Font size: it needs to be in the range of 2.0-4.0mm, for example, 2.5mm is selected as the standard font size, which corresponds to the normal printing scenario. Character spacing: the specification is 0.3-0.8mm, and the conventional value is 0.5mm, which is used to simulate the typesetting deviation. Font thickness: it is divided into regular and bold, for example, it is set to bold to simulate the thickening of character edges caused by uneven printing pressure. Angle of inclination: it allows a deviation within ±15°, for example, 8° is set to simulate the inclination of the certificate during scanning.

[0021] Sample proportion control (based on real scene statistics): standard parameter combination (no defects): the proportion is 50%, which is set according to the statistical result that normal certificates account for about 50%-60% in real scenes, and is used as the basic data for model training. Single defect sample (such as blur = 3px): the proportion is 30%, which corresponds to the proportion of single defect certificates (such as slight wear) in reality, and is used to improve the recognition robustness of the model to single defects. Compound defect sample (blur + inclination): the proportion is 20%, which simulates complex defect scenarios such as folding and stains, accounts for about 20% of real certificate defects, and is used to train the comprehensive recognition ability of the model to multi-dimensional defects.

[0022] International standard application (taking ICAO Doc9303 as an example): font size compliance verification: if the parameter is set to 1.8mm (lower than the lower limit of 2.0mm), the system automatically warns and adjusts to 2.0mm to ensure compliance with international standards. Angle of inclination constraint: when the parameter is set to 20° (exceeding the ±15° range), the system forcibly corrects it to 15°, and records the deviation value for subsequent evaluation, to ensure the compliance of the sample generation.

[0023] Parameter association rules (based on printing physical laws), font weight is associated with character spacing: when the font weight is set to bold, the character spacing automatically increases by 10% (e.g., adjusted from 0.5mm to 0.55mm) to avoid characters sticking together due to bolding, which conforms to the logic of printing technology. Blur is associated with resolution: for every 1-level increase in blur (e.g., from 2px to 3px), the resolution decreases synchronously by 5dpi (from 300dpi to 295dpi), simulating the negative correlation between print quality and blur, and ensuring that the parameter combination conforms to physical laws.

[0024] The optimized solution ensures the authenticity of samples through semantic corpus segmentation, takes the ICAO international standard as the parameter setting benchmark, establishes parameter association rules in combination with printing physical laws, and sets the sample ratio based on real-scene statistics, forming a complete sample generation system. This system not only meets the requirements of industry norms but also can simulate various defects of real certificates through multi-dimensional parameter combinations, providing scientific and effective dataset support for the training of the machine-readable zone (MRZ) recognition model, ensuring that the generated samples have high authenticity and reliability in practical applications.

[0025] S102, Combine the preprocessed machine-readable zone corpus and preset synthesis parameters, simulate the quality changes, information truncation, and physical defects of the machine-readable zone image, and establish the association rules between parameters based on international standards and real MRZ image data to generate parameter combination constraint conditions.

[0026] In one implementation, perform a correlation degree combination analysis on the preprocessed machine-readable zone corpus and preset synthesis parameters to generate the simulation correlation degree analysis results of quality changes, information truncation, and physical defects. Among them, the machine-readable zone corpus includes the data after random segmentation based on semantic unit recognition, and the specific synthesis parameters include the blur, font size, character spacing, font weight, and tilt angle parameter data for simulating image quality changes. Perform an association analysis on the semantically segmented corpus and preset parameters to generate the simulation correlation degree results of quality changes, information truncation, and physical defects. The string P<CHEN<<H (name segment segmentation) and AO<<<<<<12345 (document number segment segmentation) after semantic recognition segmentation. The preset parameters are blur = 3px, font size = 2.5mm, character spacing = 0.5mm, font weight = bold, and tilt angle = 8°. The association between blur and font size: when the blur ≥ 3px, the font size needs to be ≥ 2.5mm to ensure readability (correlation coefficient 0.78); the association between tilt angle and character spacing: when the tilt angle > 5°, the character spacing needs to increase by 5% (0.5mm → 0.525mm) to reduce character overlap (correlation coefficient 0.65); the association between the corpus segmentation position and physical defects: when the segmentation point is at the starting position of the field, the probability of simulating document edge wear increases by 40%.

[0027] The correlation analysis results are filtered to generate a target simulation variable list, which includes ambiguity parameters, font size parameters, character spacing parameters, font weight parameters, tilt angle parameters, and corpus random segmentation parameters. Key variables are selected from the correlation analysis results to form the target simulation variable list. Variables with a correlation coefficient > 0.6 are included in the list. The target simulation variable list includes: ambiguity parameters (0-5px, step 1px); font size parameters (2.0-4.0mm, step 0.5mm); character spacing parameters (0.3-0.8mm, step 0.1mm); font weight parameters (regular / bold); tilt angle parameters (0-15°, step 1°); and corpus random segmentation parameters (segmentation position, field integrity markers).

[0028] The results of the correlation analysis are filtered to generate parameter combination constraints, including hard constraints, collaborative constraints, and semantic constraints that conform to international standards. Based on the correlation analysis and international standards, hard constraints, collaborative constraints, and semantic constraints are generated. Hard constraints (international standards) are: font size ∈ [2.0, 4.0] mm (ICAODoc9303); tilt angle ≤ 15°. Collaborative constraints (physical laws): font weight = bold → character spacing = base value × 1.1 (e.g., base character spacing 0.5mm → 0.55mm); for every 1px increase in ambiguity → resolution decreases by 5dpi (300dpi → 295dpi). Semantic constraints (corpus processing) are: the segmentation point must not destroy the core semantics of the field (e.g., the first letter of the name cannot be truncated); the segmented string must contain at least one valid character segment (e.g., the first 3 digits of the ID number).

[0029] Based on parameter combination constraints, multi-dimensional simulation input variables are generated. These input variables characterize key features associated with changes in image quality, information truncation, and physical defects in the machine-readable region. Input variables containing features of quality changes, information truncation, and physical defects are generated based on the constraints. Input variable 1 (single defect):

[0030] {"blur":3px,"font_size":2.5mm,"letter_spacing":0.5mm,"font_weight":"bold","slant_angle":0°,"cut_position":"Middle of field","field_integrity":"Name field incomplete"} (Simulation: blurry printing + bold font + worn document with middle segmentation).

[0031] Input variable 2 (compound defect):

[0032] {"blur":2px,"font_size":3.0mm,"letter_spacing":0.55mm,"font_weight":"bold","slant_angle":10°,"cut_position":"Field start","field_integrity":"ID number segment missing"} (Simulation: A composite defect of slight blur + bold font + 10° slant + start segmentation).

[0033] Through a four-layer logic of correlation analysis, variable selection, constraint generation, and variable combination, a systematic design of parameter combinations is achieved: quantitative correlations are established between parameters based on real MRZ data and international standards, avoiding physical contradictions caused by independent parameter settings; hard constraints directly reference ICAO standards to ensure samples conform to industry norms; semantic constraints and collaborative constraints are combined to make the simulated quality changes and information truncation more closely resemble real document defect scenarios. The multi-dimensional input variables generated by this process can effectively cover defect simulations in both single and composite scenarios, providing high-quality synthetic samples for MRZ recognition model training.

[0034] S103 generates machine-readable synthetic samples by processing based on the combination of synthesis parameters, corpus preprocessing, sample ratio control, and parameter combination constraints.

[0035] In one implementation, feature extraction is performed on the combination of synthesis parameters, corpus preprocessing, and sample ratio control to generate ambiguity features, font size features, character spacing features, font thickness features, tilt angle features, corpus random segmentation features, and sample ratio control features. The synthesis parameter combination features include specific combination parameter data conforming to international standards for ambiguity, font size, character spacing, font thickness, and tilt angle. The sample ratio control features include sample ratio setting data for different parameter combinations. Features from parameter combinations, corpus preprocessing, and sample ratios are extracted to generate structured feature data. For the synthesis parameter combination features, the input parameter combinations are: ambiguity = 3px, font size = 2.5mm, character spacing = 0.55mm (bold text automatically increases by 10%), and tilt angle = 8°.

[0036] The feature extraction results are as follows:

[0037] {"blur_feature":3,"font_size_feature":2.5,"letter_spacing_feature":0.55,"font_weight_feature":"bold","slant_angle_feature":8,"is_standard_compliant":true / / Comply with ICAO standards}.

[0038] The corpus preprocessing features include the original corpus: P<CHEN<<HAO<<<<<<1234567>8901234; the segmented corpus: P<CHEN<<H (length 7, segmented in the middle of the name segment, field integrity = partially missing).

[0039] The feature extraction results are as follows:

[0040] {"cut_length":7,"cut_position":"middle","field_integrity":"partial_loss","semantic_unit":"na me"}. The sample ratio control features include the preset ratio: standard sample 50%, single defect 30%, compound defect 20%; the current parameter combination belongs to "single defect (blur degree 3px)", and the corresponding ratio allocation: 30%.

[0041] Feature extraction processing is performed on the generation process of generating machine-readable area synthetic samples to generate parameter combination simulation features, corpus preprocessing simulation features, and sample ratio control features. Among them, the parameter combination simulation features include the simulation parameter combination data of image quality change and physical defects; the corpus preprocessing simulation features include the randomly segmented data of the truncated corpus; the sample ratio control features include the ratio allocation data of samples with different parameter combinations. Extract the parameter combination simulation, corpus preprocessing simulation, and ratio control features in the generation process. Specifically, the parameter combination simulation feature is the simulated defect: print blur + font boldening;

[0042] The feature extraction results are as follows:

[0043] {"quality_change":"blur","physical_defect":"bold_enlargement","parameter_combination":"blur=3+font_weight=bold"}.

[0044] The corpus preprocessing simulation feature is the simulated information truncation: segmented in the middle of the name segment, simulating character loss caused by document wear.

[0045] The feature extraction results are as follows:

[0046] {"truncation_type":"mechanical_wear","truncation_position":"name_field_middle","characte r_loss_count":4}.

[0047] The sample ratio control feature is the actual generated sample type: single defect sample, and the feature extraction result is:

[0048] {"current_ratio_type":"single_defect","remaining_ratio":{"standard":50%,"compound_defect":20%}}.

[0049] Based on parameter combination constraints, this study analyzes and processes synthetic parameter combination features, corpus preprocessing features, sample ratio control features, as well as parameter combination simulation features, corpus preprocessing simulation features, and sample ratio control features. During sample generation, the parameter combinations are verified to conform to international standards, generating machine-readable synthetic samples. These machine-readable synthetic samples are used to simulate sample features characterizing MRZ image quality changes, information truncation, and physical defects, forming a MRZ synthetic sample generation result that includes parameter combination, corpus preprocessing, and sample ratio control. The parameter combination compliance is verified based on feature extraction results before sample generation. Parameter combination verification is as follows: Input parameters: tilt angle = 12°, ambiguity = 5px; Constraint verification: tilt angle 12° ≤ 15° (compliant with hard constraints); ambiguity 5px ≤ 5px (compliant with hard constraints); Cooperative constraint: resolution is automatically adjusted to 275dpi (300dpi - 5×5dpi) when ambiguity is 5px. The verification result is compliant, and generation is allowed. The sample generation results are as follows. The synthetic sample features are: {"defect_types":["blur=5","slant=12°"],"compliance_grade":"A", / / fully compliant with ICAO standards"sample_type":"compound_defect","visualization":"blurred and tilted MRZ image, name segment split in the middle"}

[0050] Through a three-layer processing approach—feature extraction, process simulation, and compliance verification—a systematic generation of MRZ synthetic samples was achieved. Parameter combinations, corpus segmentation, and proportion control were transformed into structured features, ensuring traceability of the generation process. Parameter combinations were verified based on ICAO standards and physical laws to avoid generating non-compliant or physically contradictory samples. Combined with semantic segmentation and parameter association rules, the generated samples' quality variations and information truncation more closely resemble real-world document defect scenarios. The samples generated by this process not only meet international standard requirements but also cover defect simulations in both single and complex scenarios, providing MRZ recognition models with training data that combines compliance and realism.

[0051] S104, process the generated machine-readable synthesized sample to generate sample-related evaluation results.

[0052] In one implementation, based on the combination features of synthesis parameters, corpus preprocessing features, and sample ratio control features, the quality changes, information truncation, and physical defect simulation features of the generated machine-readable synthesized samples are extracted to generate a set of sample feature vectors. The input sample parameters are: ambiguity = 3px, font size = 2.5mm (standard range), character spacing = 0.55mm (bold + 10%), tilt angle = 8°, corpus segmentation = middle segmentation of the name segment (length 7), and sample type = single defect (ambiguity). The feature vector extraction results are as follows:

[0053] [{"blur":3,"font_size":2.5,"letter_spacing":0.55,"slant_angle":8,"cut_position":"middle","def ect_type":"single_blur","compliance":"compliant"}].

[0054] By comparing the feature vector sets of samples under different parameter combinations, the pattern of sample quality variation in the machine-readable area is determined, and sample quality change assessment information is generated. The feature vectors of different parameter combinations are compared to determine the pattern of quality variation with parameters. The comparison scenarios are as follows: Scenario 1: Blur = 2px, font size = 2.0mm → OCR recognition rate 92%; Scenario 2: Blur = 3px, font size = 2.5mm → OCR recognition rate 85%; Scenario 3: Blur = 5px, font size = 3.0mm → OCR recognition rate 70%.

[0055] For every 1px increase in blur, the font size needs to be increased by at least 0.5mm to maintain a recognition rate drop of ≤5%. Generate quality change assessment information: {"blur_impact":"For every 1px increase, the recognition rate drops by an average of 7%","font_size_compensation":"Needs to be adjusted at a ratio of 0.5mm / 1px"}.

[0056] The sample quality change assessment information is compared with a preset normal sample pattern library and an abnormal sample pattern library containing international standard samples to generate a sample quality deviation assessment result. The assessment information is compared with the preset pattern library to calculate the deviation. Normal sample pattern library standards: ambiguity ≤ 2px, font size 2.0-3.0mm, recognition rate ≥ 90%; Current sample parameters: ambiguity = 3px, recognition rate = 85%; Deviation calculation: Ambiguity deviation: (3-2) / 2 = 50%; Recognition rate deviation: (90-85) / 90 ≈ 5.56%; Overall deviation: (50% + 5.56%) / 2 = 27.78%. The assessment results are as follows:

[0057] {"deviation_level":"moderate","standard_reference":"Ambiguity exceeds the standard by 50%, recognition rate is close to the threshold"}.

[0058] By combining historical evaluation data of machine-readable synthesized samples and the rationality of parameter combinations, the sample quality deviation evaluation results are weighted to generate a sample evaluation risk feature vector. The deviation results are weighted based on historical data and parameter rationality. Historical data weight is calculated as follows: historical defect rate of samples exceeding ambiguity limits: 30% (weight 40%); current parameter combination rationality score: 70 points (weight 60%). The weight allocation is based on the analysis of the degree of impact on MRZ sample quality risk, specifically as follows: Historical data weight (40%): Referencing the frequency of similar parameter deviations (such as font size, exceeding ambiguity limits) leading to a decrease in recognition rate in historical evaluation data (e.g., in the past 3 evaluations, the risk caused by historical deviations accounted for 40%), reflecting the impact of historical data on current risk. Parameter rationality weight (60%): The rationality of the current parameter combination directly determines the compliance and effectiveness of sample generation, and has a greater real-time impact on risk, hence it is given a higher weight. Furthermore, the weights can be dynamically adjusted according to actual scenarios (e.g., when international standards are updated or the type of sample defect changes, the weight ratio can be reallocated).

[0059] The risk vector is calculated as follows. The formula for calculating the risk value is: Risk Value = (Deviation × Historical Data Weight) + ((1 - Reasonableness Coefficient) × Parameter Reasonableness Weight). Substituting the above values ​​into the formula, we get the risk value = 0.2778 × 40% + (1 - 70 / 100) × 60% = 0.111 + 0.18 = 0.2911. The percentage form of deviation is converted into a decimal in the range [0,1] (e.g., 27.78% → 0.2778); the reasonableness score of 100 points is converted into a reasonableness coefficient in the range [0,1] (e.g., 70 points → 70 / 100 = 0.7). This ensures that the values ​​of the two dimensions can be directly calculated under the same scale. The risk feature vector is {"risk_level":"medium","historical_correlation":0.3,"parameter_rationality":0.7}.

[0060] The sample quality deviation assessment results and sample assessment risk feature vectors are normalized and comprehensively calculated to generate a set of sample assessment feature parameters, including the quality deviation coefficient, sample validity index, and parameter compliance index. This set of parameter parameters consists of a sample feature vector set, sample quality change assessment information, sample quality deviation assessment results, sample assessment risk feature vector, quality deviation coefficient, and sample validity index. After normalization, comprehensive assessment parameters are generated. The normalization calculation is as follows: Quality deviation coefficient: 27.78% → 0.28 (normalized to [0,1]); Sample validity index: 1 - 29.11% = 0.71; Parameter compliance index: 1 (complies with standards such as font size and tilt angle). Comprehensive parameter set: {"feature_vectors":[first step feature vectors],"quality_change":[second step evaluation information],"deviation_result":[third step deviation],"risk_vector":[fourth step risk vector],"quality_bias":0.28,"validity_index":0.71,"compliance_index":1.0}.

[0061] A systematic evaluation of MRZ synthetic samples was achieved through a five-layer process: feature extraction, pattern analysis, deviation assessment, risk weighting, and comprehensive calculation. Unstructured information such as quality variations and defect types were transformed into computable feature vectors, ensuring the objectivity of the evaluation. Comparison with an international standard sample pattern library quantified the sample's compliance and quality deviation. Weighted processing using historical evaluation data improved the reliability of the risk assessment. Indicators such as the quality deviation coefficient and effectiveness index formed a multi-dimensional evaluation matrix, providing clear guidance for sample optimization. The feature parameter set generated by this evaluation process reflects both the sample's conformity to international standards and the quantification of the impact of defects on recognition performance, providing a scientific basis for subsequent intelligent optimization.

[0062] S105 processes the evaluation results and parameter setting basis in the synthesis parameters in accordance with international standards to generate the basis for generating machine-readable synthesized samples.

[0063] In one implementation, based on the parameter evaluation results and the international standard benchmarking method, the deviation values ​​of the actual values ​​of ambiguity, font size, character spacing, font thickness, and tilt angle in the synthesized parameters are calculated from the international standards, generating parameter deviation data. The parameter deviation calculation is based on the absolute deviation method, that is, the degree of deviation is quantified by the difference between the actual parameter values ​​and the international standard threshold. The formula is: Deviation value = Actual value - Standard threshold. For interval-type standards (e.g., font size 2.0-4.0mm), the lower limit is 2.0mm and the upper limit is 4.0mm; for unidirectional standards (e.g., tilt angle ≤15°), the threshold is 15°.

[0064] The international standard details (ICAO Doc 9303) are as follows: Font size standard range: 2.0-4.0mm, to ensure character recognition accuracy of OCR devices; fonts that are too small (<2.0mm) may cause blurry character edges during scanning, while fonts that are too large (>4.0mm) exceed the layout specifications for machine-readable areas on documents. Tilt angle, standard threshold: ≤15°, allowing normal offset of documents during scanning or photography; tilts exceeding 15° will significantly reduce OCR recognition rate and need to be controlled through parameter deviation. Character spacing, standard range: 0.3-0.8mm, to avoid character sticking or sparseness; character spacing <0.3mm may cause character overlap, while >0.8mm will affect the compactness of the layout.

[0065] The actual parameter values ​​and deviation calculation process are as follows: Font size deviation: Actual value: 1.8mm; Standard lower limit: 2.0mm; Deviation calculation: 1.8mm - 2.0mm = -0.2mm. The deviation is negative (below the standard lower limit), which may increase the difficulty of character recognition. Tilt angle deviation: Actual value: 12°; Standard threshold: 15°; Deviation calculation: 12° - 15° = -3° (Since the standard is a unidirectional threshold, the actual deviation is taken as 0°). Compliance judgment: 12° ≤ 15°, deviation value recorded as 0°, conforming to the standard. Character spacing deviation: Actual value: 0.2mm; Standard lower limit: 0.3mm; Deviation calculation: 0.2mm - 0.3mm = -0.1mm. The deviation is negative (below the standard lower limit), which may cause characters to stick together.

[0066] The impact of a font size reduction of -0.2mm: Experimental data shows that for every 0.1mm decrease in font size, the OCR recognition rate drops by an average of 3%; in this example, -0.2mm may lead to a recognition rate decrease of approximately 6%, which needs to be closely monitored in sample evaluation. The impact of a character spacing reduction of -0.1mm: A character spacing of 0.2mm is close to 40% of the character width (approximately 0.5mm), which may cause overlapping edges of adjacent characters; it may also simulate ink diffusion or layout errors during document printing.

[0067] By using standardized deviation calculation methods, parameter compliance can be transformed into quantitative indicators, providing data support for subsequent parameter optimization. In this example, the negative deviations in font size and character spacing need to be corrected during sample generation through parameter adjustments or constraints (such as automatically increasing the font size to 2.0mm and the character spacing to 0.3mm) to ensure that the synthesized samples meet international standards and improve the effectiveness of OCR model training.

[0068] The parameter deviation data undergoes validity filtering to generate filtered parameter deviation data. The deviation data filtering is based on a dual standard: assessing whether the deviation has a substantial impact on MRZ identification effectiveness or international standard compliance; and determining whether it is noise or measurement error based on the absolute value of the deviation.

[0069] The original deviation data is analyzed as follows:

[0070]

[0071]

[0072] The absolute value threshold for deviation is set as follows: core parameters (font size, fuzziness, character spacing): threshold 0.05mm / 0.1px, that is, deviation absolute value ≥0.05mm / 0.1px is considered valid; non-critical parameters (font thickness, tilt angle, etc.): the threshold can be relaxed to 0.1mm, but it needs to be judged in conjunction with the business impact.

[0073] The priority of business impact is as follows: High priority parameters include font size, ambiguity, and character spacing: these directly affect character readability, and deviations must be included in the evaluation 100%; Low priority parameters include the following: font thickness: this only affects the character edge outline, and minor deviations (such as 0.03mm) can be filtered out; tilt angle: in this example, 12°≤15°, the deviation is 0°, and no filtering is required.

[0074] The filtering process and result analysis are as follows: Font size deviation (-0.2mm), with an absolute value of 0.2mm > 0.05mm, and being a high-priority parameter, is retained. Blur deviation (+0.5px), with an absolute value of 0.5px > 0.1px (the ambiguity parameter threshold), and exceeding the standard upper limit, is retained. Character spacing deviation (-0.1mm), with an absolute value of 0.1mm > 0.05mm, and affecting character spacing, is retained. Font thickness deviation (+0.03mm), with an absolute value of 0.03mm < 0.05mm, and being a non-critical parameter, is filtered.

[0075] The retained deviations are all parameters that significantly affect MRZ recognition (font size, ambiguity, and character spacing), avoiding non-critical deviations from interfering with the evaluation focus. After filtering noisy data, the computational workload for parameter compliance is reduced by approximately 30% while maintaining evaluation accuracy. For a character spacing deviation of -0.1mm, the character spacing can be automatically adjusted to 0.3mm during sample generation, simultaneously verifying the combined effect of font size and ambiguity.

[0076] Validity filtering, through the combination of quantitative standards and business logic, achieves noise reduction of biased data, ensuring that subsequent compliance assessments focus on parameter anomalies that substantially affect the quality of MRZ samples. This mechanism avoids assessment distortion caused by "bias overload," provides more accurate input data for intelligent optimization, and effectively improves the standard compliance of synthetic samples and the training efficiency of the recognition model.

[0077] Based on the filtered parameter deviation data, the degree to which the synthesized parameters conform to international standards is calculated, generating parameter compliance data. The compliance calculation employs a relative deviation ratio method. The core logic is to convert the absolute deviation of the parameter into a ratio relative to the standard's allowable range. The formula is: Compliance = 1 - (|deviation value| / upper limit of the standard's allowable range). This model quantifies compliance as a value in the range [0,1], where 1 represents complete compliance and 0 represents complete non-compliance. The upper limit of the standard's allowable range is dynamically determined based on the parameter type. For interval-type parameters (such as font size): the upper limit is the difference between the standard's maximum and minimum values ​​(4.0mm - 2.0mm = 2.0mm); for unidirectional parameters (such as ambiguity): the upper limit is the standard threshold (5px).

[0078] The calculation process for compliance of each item is as follows: Font size compliance, deviation value: |-0.2mm|=0.2mm (absolute value of negative deviation); upper limit of the standard allowable range: 4.0mm-2.0mm=2.0mm. The calculation process is as follows: Compliance = 1-(0.2mm / 2.0mm)=0.9 (90%). The font size deviation accounts for 10% of the standard range, which is still within the acceptable range, but attention should be paid to the impact on the recognition rate.

[0079] Blur tolerance compliance, deviation value: |+0.5px|=0.5px (absolute value of positive deviation); upper limit of standard allowable range: 5px (one-way standard threshold). The calculation process is as follows: compliance = 1 - (0.5px / 5.0px) = 0.9 (90%); blur tolerance exceeding the standard by 10% may cause slight blurring of the image, but has not yet exceeded the threshold that seriously affects recognition.

[0080] Character spacing compliance, deviation value: |-0.1mm|=0.1mm (absolute value of negative deviation); upper limit of standard allowable range: 0.8mm-0.3mm=0.5mm; calculation process is as follows, compliance = 1-(0.1mm / 0.5mm)=0.8 (80%); character spacing deviation accounts for 20% of the standard range, which poses a risk of character adhesion and needs to be adjusted first when generating samples.

[0081] The comprehensive compliance data structure is analyzed as follows: {"individual_compliance":{"font_size":0.9, / / font size compliance 90%"blur":0.9, / / blur compliance 90%"letter_spacing":0.8 / / letter spacing compliance 80%},"average_compliance":0.87 / / average compliance 87%}. Individual compliance reflects the independent compliance status of each parameter, facilitating the identification of specific issues; average compliance provides an overall compliance overview through the arithmetic mean ((0.9+0.9+0.8) / 3=0.87), suitable for batch evaluation of samples.

[0082] Set a threshold (e.g., average compliance ≥ 0.85) to automatically filter low-compliance samples, ensuring the training dataset conforms to international standards. For cases where the character spacing compliance (0.8) is lower than the font size and blurriness, prioritize adjusting the character spacing parameter to improve overall compliance. When the compliance of a certain parameter is < 0.8 (e.g., character spacing 0.8), trigger an alert, indicating that this parameter combination may lead to a decrease in recognition rate.

[0083] A standardized compliance calculation model can transform parameter deviations into quantifiable compliance indicators, providing data support for the quality assessment and optimization of MRZ synthetic samples. This model balances computational efficiency with business logic, enabling rapid identification of parameter anomalies while providing standardized input for intelligent models. Further optimization using parameter weights and nonlinear mapping can enhance the accuracy of compliance assessments, ensuring that generated samples not only conform to international standards but also effectively simulate various defects in real-world scenarios.

[0084] By combining historical evaluation data of the synthesis parameters and the update frequency of international standards, the parameter compliance data is weighted to generate a parameter compliance feature vector. The historical data weight is the font size historical deviation impact coefficient: 0.4 (in the last 3 evaluations, font size deviation caused a 40% decrease in recognition rate). The international standard update frequency: the ICAO standard has not been updated in the last 5 years, with a weight coefficient of 0.9 (stable). The weighted calculation yields the following compliance feature vector: Compliance Feature Vector = Overall Compliance × Historical Impact Weight × Standard Stability Weight; Font Size: 0.9 × 0.4 × 0.9 = 0.324; Ambiguity: 0.9 × 0.2 × 0.9 = 0.162; Character Spacing: 0.8 × 0.3 × 0.9 = 0.216; The final feature vector is [0.324, 0.162, 0.216].

[0085] Through a four-layer process of deviation calculation, validity filtering, compliance scoring, and weighted processing, quantitative assessment of parameter compliance is achieved. This process transforms parameter deviations into calculable numerical indicators, ensuring the objectivity of the compliance assessment; eliminates noise bias, focusing on parameter anomalies that substantially impact recognition performance; combines historical assessment data with standard update frequency to make compliance features more closely aligned with real-world application scenarios; and integrates multi-dimensional compliance data into a unified vector, providing standardized input for intelligent optimization. The parameter compliance feature vector generated by this process reflects both the degree of standard compliance of the current parameter combination and dynamically adapts to historical data patterns and changes in industry standards, providing a scientific compliance basis for subsequent intelligent model optimization.

[0086] S106, based on the intelligent model, combines parameter combinations, sample ratio control, sample-related evaluation results, and parameter compliance feature vectors to generate optimized results for machine-readable synthetic samples.

[0087] In one implementation, risk level analysis is performed on parameter combinations, sample ratio control, and sample-related assessment results to generate a parameter combination risk level assessment result. The risk level is analyzed based on the parameter combinations, sample ratio, and assessment results. Input data includes: parameter combinations: ambiguity = 5px (standard ≤ 5px), font size = 1.8mm (standard 2.0-4.0mm); sample ratio: high ambiguity samples account for 30%; assessment results are: quality deviation coefficient = 0.28 (threshold 0.2), parameter compliance index = 0.87.

[0088] The risk level is calculated as follows: a negative deviation in font size (-0.2mm) leads to a 6% decrease in recognition rate, with a risk weight of 40%; the proportion of highly ambiguous samples exceeds the historical average (20%), with a risk weight of 30%; the quality deviation coefficient exceeds the threshold, with a risk weight of 30%; the overall risk level is medium risk (65%).

[0089] Anomaly detection results are generated by comparing quality risk characteristics from sample-related assessment results with preset quality thresholds. The quality anomaly detection method is based on a two-dimensional threshold comparison, which quantifies and compares actual sample characteristics with preset quality thresholds to achieve automated identification of abnormal states. Its core logical framework is as follows: Threshold system construction: establishing threshold standards for core indicators such as recognition rate and compliance index; Feature data collection: extracting key quality characteristics from sample assessment results; Threshold comparison calculation: determining whether features exceed allowable ranges through numerical comparison; Anomaly type classification: generating specific anomaly types and cause descriptions based on the comparison results.

[0090] For details on the preset threshold system, please see the table below:

[0091]

[0092] The impact of an 82% recognition rate is analyzed as follows. The data source is the actual recognition rate of the OCR model on the samples. The deviation is due to the font size being 1.8mm (lower than the standard lower limit of 2.0mm), which causes blurry character edges; the blurriness being 5px (close to the standard upper limit of 5px), which further reduces character clarity; a recognition rate below 85% may cause the false recognition rate in actual document recognition scenarios to rise to more than 5% (the false recognition rate in normal scenarios is ≤2%).

[0093] The compliance index of 0.87 is analyzed as follows: Font size compliance: 0.9 (deviation -0.2mm); Blur compliance: 0.9 (deviation +0.5px); Character spacing compliance: 0.8 (deviation -0.1mm); The weighted calculation is (0.9 + 0.9 + 0.8) / 3 = 0.87. The character spacing compliance (0.8) is close to the lower limit of the threshold, which may lead to character adhesion defects exceeding the allowable range of international standards. The recognition rate anomaly is determined by the following formula: 82% < 85% → result is true; the anomaly type is classified as recognition effect anomaly; the severity is moderate anomaly (deviation 3%, not exceeding the emergency threshold of 80%). The compliance index anomaly is determined by the following formula: 0.87 < 0.9 → result is true; the anomaly type is classified as parameter compliance anomaly; the severity is critical anomaly (deviation 0.03, close to but not completely exceeding the standard). The overall anomaly assessment result is as follows: Anomaly type: Composite anomaly (recognition rate not up to standard + compliance threshold); Risk level: Medium risk; It may lead to the model being overly adaptable to low-quality samples during training, reducing the robustness of recognition in real-world scenarios.

[0094] Automatically label anomalous samples to avoid including samples with a recognition rate <85% or a compliance index <0.9 in the core training set; add a "[Parameters need adjustment]" label to current samples to prevent them from entering the model fine-tuning stage. Parameter optimization is guided: anomalous recognition rates indicate the need to adjust font size and ambiguity parameters; anomalous compliance indicates the need to prioritize correcting character spacing parameters (e.g., from 0.2mm to 0.3mm). Trigger parameter combination risk warnings, indicating a synergistic risk associated with the combination of "font size + ambiguity + character spacing"; correlate with historical data to retrieve handling methods for similar anomalous samples within the past 3 months (e.g., increasing font size by 0.2mm to improve the recognition rate to 88%).

[0095] The threshold is adaptively adjusted, recalibrated based on quarterly historical data: if the average recognition rate of qualified samples increases to 88% in a quarter, the threshold can be raised to 86%; if international standards are updated, leading to higher compliance requirements, the compliance index threshold is adjusted to 0.95. A weighting mechanism is introduced, setting differentiated weights for different thresholds: recognition rate threshold weight 60% (directly affecting model performance); compliance index weight 40% (ensuring industry standard compliance); comprehensive anomaly score = recognition rate deviation × 60% + compliance deviation × 40% = 3% × 60% + 3.33% × 40% = 3.13%.

[0096] Anomaly grading response: Mild anomalies (e.g., recognition rate 84%): only record without intervention, observe subsequent sample trends; Moderate anomalies (e.g., current sample): trigger parameter fine-tuning suggestions; Severe anomalies (recognition rate < 80%): pause sample generation and initiate full-process diagnosis. Quality anomaly judgment achieves quantitative identification and classification of MRZ synthetic sample quality problems by constructing a standardized threshold comparison system. This mechanism can accurately locate potential risks to recognition effectiveness and compliance, and provide a clear basis for parameter optimization and sample selection. In practical applications, combining dynamic threshold adjustment and weighting mechanisms can further improve the accuracy and adaptability of anomaly judgment, ensuring that generated samples always meet international standard requirements and model training needs.

[0097] The results of the parameter combination risk level assessment and quality anomaly judgment are processed to generate sample risk warning information. Specifically, the warning content is as follows: the parameter combination (ambiguity = 5px + font size = 1.8mm) has a medium risk: the recognition rate is lower than the threshold of 3%, the compliance index is close to the critical value; the proportion of high ambiguity samples exceeds the historical average by 10%, which may lead to model overfitting.

[0098] Acquire historical generation information of machine-readable synthesized samples, intelligent model configuration information, and international standard update records. Historical data is from the past 3 months; samples with ambiguity > 4px and font size < 2.0mm showed an average decrease in OCR recognition rate of 8%. Regarding international standard updates, ICAO recently released a supplementary standard requiring the lower limit of font size to be adjusted to 2.1mm (previously 2.0mm).

[0099] Based on historical synthetic sample generation information, sample generation risk warning information, and international standard update records, the intelligent model configuration information is processed to generate machine-readable synthetic sample generation optimization instructions. The optimization instruction generation is based on a 3D-driven model, integrating historical generation data, real-time risk warnings, and international standard updates to form a systematic parameter adjustment and model configuration optimization scheme. Its logical framework is as follows: Data Integration Layer: Collects historical sample performance data, current risk warning information, and standard update clauses; Rule Reasoning Layer: Generates candidate schemes through preset optimization rules (such as standard thresholds triggering parameter adjustments) and historical experience (such as a parameter combination previously improving recognition rates); Instruction Generation Layer: Transforms the reasoning results into executable parameter adjustment instructions and model configuration rules.

[0100] The font size optimization logic is based on the updated international standards: the latest supplementary standard of ICAO raises the lower limit of font size from 2.0mm to 2.1mm; historical data verification: data from the past 6 months shows that the average recognition rate of samples with font size ≥2.1mm is 7% higher than that of samples with font size <2.1mm; risk warning correlation: the current sample font size = 1.8mm triggers the compliance threshold anomaly (compliance index 0.87 < 0.9).

[0101] The adjustment of ambiguity control is based on the following: risk warning: the proportion of high-ambiguity samples (ambiguity ≥ 5px) exceeds the historical reasonable range (20%), and the average recognition rate of such samples is below 80%; physical laws: 5px of ambiguity is close to the limit of printing equipment, and further increase may lead to unreadable characters; model training feedback: excessively high-ambiguity samples will lead to a decrease in the model's generalization ability to normal samples. The adjustment of the sample ratio is based on historical statistics: when the proportion of high-defect samples (such as ambiguity + tilt composite defects) is 20%, the standard deviation of the model recognition rate is the smallest (the stability is the best); risk balance: although the current proportion of high-defect samples of 30% improves the model's noise resistance, it leads to insufficient training of normal samples (a decrease in recognition rate of 3%).

[0102] The model configuration optimization is analyzed as follows: A new sample filtering rule has been added, with the following logic: Filtering condition: Font size < 2.1mm (violates the latest international standard); Execution timing: Automatically triggered after sample generation and before data entry; Historical effect: After filtering samples with font size < 2.0mm, the model's compliance-related indicators improved by 15%. The loss function weights have been adjusted, with the following weight allocation: Original weights: Quality deviation coefficient 60% + Compliance index 30% + Other 10%; New weights: Quality deviation coefficient 60% → 50%, Compliance index 30% → 40% (due to the increased importance of compliance after the standard update); On the test set, this adjustment increased the compliant sample recognition rate from 88% to 91%. The calculation formula is: New loss function = 0.5 × Quality deviation loss + 0.4 × Compliance loss + 0.1 × Other losses.

[0103] After adjustments, 1000 samples were generated, with 100% of them having a font size ≥ 2.1mm. The average recognition rate improved from 82% to 87%. Every 1000 samples generated, the execution effect of the optimization instructions was automatically calculated: if the compliance index did not reach 0.95, a second optimization was triggered (such as further increasing the recommended font size to 2.8mm); if the recognition rate improvement was < 5%, the weight allocation of the loss function was re-evaluated.

[0104] Ensure that sample generation always complies with the latest international standards (e.g., ICAO may update printing parameter requirements annually); avoid compliance risks caused by outdated standards (e.g., samples being rejected by customs due to non-compliance). The optimized parameter combination shortens the model training cycle by 20% (due to a reduction in invalid samples); the standard deviation of the recognition rate decreases from 8% to 5%, significantly improving the model's generalization ability. Reduce the cost of sample regeneration due to unreasonable parameters (estimated to save 30% of computing resources annually); reduce the cost of manual parameter adjustment and compliance review (automated filtering and weight adjustment replace 50% of manual work).

[0105] The optimized instruction generation mechanism achieves intelligent iteration in the MRZ synthetic sample generation process through deep integration of historical data, risk warnings, and standard updates. This mechanism not only ensures that samples strictly conform to international standards but also continuously improves model training performance through data-driven approaches. In practical applications, combined with dynamic verification and feedback mechanisms, a closed loop of "problem identification - solution generation - effect verification - continuous optimization" can be formed, ensuring that the sample generation system always maintains the industry's best practice level and effectively solving the problems of lag in parameter adjustment and rigid model configuration in traditional methods.

[0106] This application addresses the issues of insufficient real samples and distorted synthetic samples through multi-dimensional parameter simulation and intelligent optimization. First, preprocessed MRZ corpora, synthesis parameters, and international standards are acquired. The corpus is segmented into fragments of length 5-14, and parameters include ambiguity, font size, etc., while adhering to international standard formatting and printing quality. Next, the corpus and parameters are combined to simulate image quality changes, and parameter association rules are established based on international standards and real data to generate parameter combination constraints containing hard, collaborative, and semantic constraints. Then, synthetic samples are generated based on parameter combinations, corpus preprocessing, sample ratio control, and constraints, and parameter compliance is verified. After sample generation, features such as quality changes are extracted, and different parameter combinations are compared to determine the patterns of quality changes. A deviation assessment is generated by comparing with a standard pattern library, and weighted processing using historical data is applied to generate a set of evaluation feature parameters including a quality deviation coefficient.

[0107] The deviation of parameters from international standards is then calculated, invalid deviations are filtered out, and compliant data and feature vectors are generated. Finally, based on the intelligent model, combined with parameter combinations, sample ratios, evaluation results, and compliant feature vectors, the risk level is analyzed, anomaly detection and risk warnings are generated by comparing with quality thresholds, and optimization instructions are generated by combining historical data and standard updates, such as adjusting parameters like font size and ambiguity, to optimize the sample ratio and model configuration. The synthetic samples generated by this solution conform to international standards, improve the model recognition rate by 7%, save 30% of computing resources, provide high-quality data support for MRZ recognition model training, and are suitable for document recognition scenarios.

[0108] In one implementation, such as Figure 2 As shown, this application also provides an apparatus for generating MRZ synthetic samples, comprising:

[0109] The acquisition module 201 is used to acquire machine-readable area synthesized samples and international standards. The machine-readable area synthesized samples include pre-processed machine-readable area corpus, specific synthesis parameters, and control of the proportion of synthesized samples. Among them, the synthesis parameters include ambiguity, font size, character spacing, font thickness, and tilt angle. The corpus preprocessing involves randomly segmenting the machine-readable area corpus into a length range of [5,14]. The international standards are used to standardize the machine-readable area format, character set, and printing quality of travel documents.

[0110] The processing module 202 is used to combine the preprocessed machine-readable zone corpus with preset synthesis parameters to simulate quality changes, information truncation, and physical defects in machine-readable zone images. Based on international standards and real MRZ image data, it establishes association rules between parameters and generates parameter combination constraints. Based on the synthesis parameter combination, corpus preprocessing, sample ratio control, and parameter combination constraints, it generates machine-readable zone synthesized samples. It processes the generated machine-readable zone synthesized samples to generate sample-related evaluation results. It processes the evaluation results and parameter setting basis in the synthesis parameters in conjunction with international standards to generate the basis for generating machine-readable zone synthesized samples. Based on the intelligent model, it processes the parameter combination, sample ratio control, sample-related evaluation results, and parameter compliance feature vectors to generate optimized results for generating machine-readable zone synthesized samples.

[0111] The computer-readable storage medium provided in the above embodiments of this application and the method for generating MRZ synthetic samples provided in the embodiments of this application are based on the same inventive concept and have the same beneficial effects as the methods adopted, run or implemented by the applications stored therein.

[0112] The various embodiments in this application are described in a related manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the embodiments for the method of generating MRZ synthetic samples, electronic devices, electronic devices, and readable storage media are basically similar to the embodiments of the method of generating MRZ synthetic samples described above, and therefore the descriptions are relatively simple. Relevant parts can be referred to in the descriptions of the embodiments of the method of generating MRZ synthetic samples described above.

Claims

1. A method for generating MRZ synthetic samples, characterized in that, include: The process involves obtaining machine-readable area (MDI) synthesized samples and international standards. The MDI synthesized samples include pre-processed MDI corpus, specific synthesis parameters, and control over the proportion of synthesized samples. The synthesis parameters include ambiguity, font size, character spacing, font weight, and tilt angle. The corpus preprocessing involves randomly segmenting the MDI corpus into a length range of [5,14]. International standards are used to standardize the format, character set, and printing quality of machine-readable areas on travel documents. The preprocessed machine-readable zone corpus and preset synthesis parameters are combined to simulate the quality changes, information truncation and physical defects of machine-readable zone images. Based on international standards and real MRZ image data, the association rules between parameters are established and parameter combination constraints are generated. Based on the combination of synthesis parameters, corpus preprocessing, sample ratio control and parameter combination constraints, machine-readable synthesized samples are generated. The generated machine-readable synthetic samples are processed to generate sample-related evaluation results; The evaluation results and parameter settings in the synthesis parameters are processed in accordance with international standards to generate a basis for generating machine-readable synthesized samples. This includes calculating the deviations between the actual values ​​of ambiguity, font size, character spacing, font weight, and tilt angle in the synthesis parameters and international standards, based on the parameter evaluation results and international standard benchmarking methods, generating parameter deviation data; performing validity filtering on the parameter deviation data to generate filtered parameter deviation data; calculating the degree to which the synthesis parameters conform to international standards based on the filtered parameter deviation data to generate parameter compliance data; and weighting the parameter compliance data by combining historical evaluation data of the synthesis parameters and the update frequency of international standards to generate a parameter compliance feature vector. Based on the intelligent model, combined with parameter combination, sample ratio control, sample-related evaluation results and parameter compliance feature vectors, the generated optimized results of synthesized samples in the machine-readable area are produced.

2. The method as described in claim 1, characterized in that, The preprocessed machine-readable region (MRZ) corpus and preset synthesis parameters are combined to simulate quality changes, information truncation, and physical defects in machine-readable region images. Based on international standards and real MRZ image data, association rules between parameters are established, and parameter combination constraints are generated, including: The preprocessed machine-readable corpus and preset synthesis parameters are combined for correlation analysis to generate simulated correlation analysis results of quality changes, information truncation and physical defects. The machine-readable corpus includes data after random segmentation based on semantic unit recognition, and the specific synthesis parameters include data on blur, font size, character spacing, font thickness and tilt angle parameters that simulate image quality changes. The correlation analysis results are filtered to generate a list of target simulation variables, which includes ambiguity parameters, font size parameters, character spacing parameters, font thickness parameters, tilt angle parameters, and corpus random segmentation parameters. The correlation analysis results are filtered and processed to generate parameter combination constraints, including hard constraints, collaborative constraints and semantic constraints that conform to international standards; Based on parameter combination constraints, multi-dimensional simulation input variables are generated, which are used to characterize key features related to changes in image quality, information truncation, and physical defects in the machine-readable area.

3. The method as described in claim 1, characterized in that, Based on the combination of synthesis parameters, corpus preprocessing, sample ratio control, and parameter combination constraints, machine-readable synthesized samples are generated, including: Feature extraction is performed on the combination of synthesis parameters, corpus preprocessing, and sample ratio control to generate ambiguity features, font size features, character spacing features, font thickness features, tilt angle features, corpus random segmentation features, and sample ratio control features. Among them, the combination of synthesis parameters includes specific combination parameter data of ambiguity, font size, character spacing, font thickness, and tilt angle that conform to international standards; the sample ratio control features include sample ratio setting data for different parameter combinations. Feature extraction is performed on the process of generating machine-readable synthesized samples to generate parameter combination simulation features, corpus preprocessing simulation features, and sample ratio control features. Among them, the parameter combination simulation features include simulated parameter combination data of image quality changes and physical defects; the corpus preprocessing simulation features include randomly segmented corpus data with information truncation; and the sample ratio control features include proportional allocation data of samples with different parameter combinations. Based on parameter combination constraints, this study combines synthetic parameter combination features, corpus preprocessing features, sample ratio control features, as well as parameter combination simulation features, corpus preprocessing simulation features, and sample ratio control features for analysis and processing. When generating samples, the study verifies whether the parameter combination conforms to international standards and generates machine-readable synthetic samples. The machine-readable synthetic samples are used to characterize the quality changes, information truncation, and physical defects of MRZ images, forming an MRZ synthetic sample generation result that includes parameter combination, corpus preprocessing, and sample ratio control.

4. The method as described in claim 3, characterized in that, The generated machine-readable region synthetic samples are processed to generate sample-related evaluation results, including: Based on the features of the synthesis parameter combination, the corpus preprocessing features, and the sample ratio control features, the quality changes, information truncation, and physical defect simulation features of the generated machine-readable synthesized samples are extracted to generate a set of sample feature vectors. By comparing the sample feature vector sets under different parameter combinations, the pattern of changes in the quality of synthesized samples in the machine-readable area with parameters is determined, and sample quality change assessment information is generated. The sample quality change assessment information is compared with a preset normal sample pattern library and an abnormal sample pattern library containing international standard samples to generate a sample quality deviation assessment result. By combining historical evaluation data of the synthesized samples in the machine-readable area and the rationality of parameter combinations, the sample quality deviation evaluation results are weighted to generate a sample evaluation risk feature vector. The sample quality deviation assessment results and sample assessment risk feature vectors are normalized and comprehensively calculated to generate a sample assessment feature parameter set that includes the quality deviation coefficient, sample validity index, and parameter compliance index. The sample assessment feature parameter set consists of the sample feature vector set, sample quality change assessment information, sample quality deviation assessment results, sample assessment risk feature vector, quality deviation coefficient, and sample validity index.

5. The method as described in claim 4, characterized in that, Based on an intelligent model, combined with parameter combinations, sample ratio control, sample-related evaluation results, and parameter compliance feature vectors, optimized results for generating machine-readable synthetic samples are produced, including: Risk level analysis is performed on parameter combinations, sample ratio control, and sample-related assessment results to generate parameter combination risk level assessment results. Based on the quality risk characteristics in the sample-related assessment results and the preset quality threshold, anomaly characteristics are compared and processed to generate quality anomaly judgment results. The results of the parameter combination risk level assessment and the quality anomaly judgment are processed to generate sample risk warning information. Acquire historical generation information of synthesized samples in the machine-readable region, intelligent model configuration information, and international standard update records; Based on the historical generation information of synthetic samples, the risk warning information of sample generation, and the update records of international standards, the configuration information of the intelligent model is processed to generate machine-readable synthetic sample generation optimization instructions.

6. An apparatus for generating MRZ synthetic samples, characterized in that, The apparatus for implementing the method of claim 1 includes: The acquisition module is used to acquire machine-readable area synthesized samples and international standards. The machine-readable area synthesized samples include pre-processed machine-readable area corpus, specific synthesis parameters, and control of the proportion of synthesized samples. Among them, the synthesis parameters include ambiguity, font size, character spacing, font thickness, and tilt angle. The corpus preprocessing involves randomly segmenting the machine-readable area corpus into a length range of [5,14]. The international standards are used to standardize the machine-readable area format, character set, and printing quality of travel documents. The processing module combines preprocessed machine-readable zone (MRZ) corpus with preset synthesis parameters to simulate quality changes, information truncation, and physical defects in machine-readable zone images. Based on international standards and real MRZ image data, it establishes association rules between parameters and generates parameter combination constraints. Based on the synthesis parameter combination, corpus preprocessing, sample ratio control, and parameter combination constraints, it generates machine-readable zone synthesized samples. It then processes the generated machine-readable zone synthesized samples to generate sample-related evaluation results. Finally, it processes the evaluation results and parameter setting basis in the synthesis parameters in accordance with international standards to generate the basis for generating machine-readable zone synthesized samples. Finally, based on an intelligent model, it processes parameter combinations, sample ratio control, sample-related evaluation results, and parameter compliance feature vectors to generate optimized results for generating machine-readable zone synthesized samples.

7. An electronic device, characterized in that, include: First processor; and memory for storing executable instructions of the first processor; The first processor is configured to execute the method for generating MRZ synthetic samples according to any one of claims 1 to 5 by executing the executable instructions.

8. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the second processor, it implements the method for generating MRZ synthetic samples according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Method for recognizing machine-readable travel certificate

    CN101038686A

  • Image editing method and device, storage medium and electronic equipment

    CN119399327A