A dynamic data enhancement method in OCR recognition model training

By randomly selecting the training direction and enhancing the depth and width of the model area in OCR recognition model training, the problem of poor dynamic data enhancement effect in the prior art is solved, and stronger dynamic data recognition and output capabilities are achieved.

CN119206734BActive Publication Date: 2025-05-09JIANGSU ELECTRIC POWER INFORMATION TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411738937.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-29
Publication Date
2025-05-09
Estimated Expiration
2044-11-29

AI Technical Summary

Technical Problem

The lack of comprehensive and in-depth analysis of complex comprehensive situations in the training of existing OCR recognition models for dynamic images that may move and rotate at the same time, resulting in poor dynamic data enhancement effect.

Method used

In OCR recognition model training, the training direction is randomly selected based on the set training direction (including horizontal movement and rotation direction) to train the model. The movement speed and acceleration associated with each training are different. The direction of meeting or failure direction is determined based on the specific training results, and the depth and width of the model area are enhanced in the direction of failure until it is changed to the direction of meeting.

Benefits of technology

The comprehensive recognition capability of the OCR recognition model has been improved, and the recognition capability of dynamic data has been initially enhanced. Through the correlation training of the fusion direction and the allocation of CPU utilization, the model has the strongest output capability during dynamic analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119206734B_ABST
    Figure CN119206734B_ABST
Patent Text Reader

Abstract

The present invention discloses a dynamic data enhancement method in OCR recognition model training, which relates to the technical field of text training and solves the problem of no comprehensive and in-depth analysis of the complex comprehensive situation in which a dynamic image may move and rotate at the same time. The present invention determines a fusion direction from the target-reaching direction through a determined target-reaching direction, and performs fusion training on the model based on the determined fusion direction, and preferentially identifies whether the training stage reaches the target. If the target is reached, no associated enhancement is required. If the target is not reached, associated enhancement is performed by allocating the associated CPU utilization. Through the actual enhancement effect, the corresponding model can achieve the best dynamic training effect in the actual fusion training process, thereby improving the specific recognition ability of the corresponding model for dynamic images.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of text training, and in particular to a dynamic data enhancement method in OCR recognition model training. Background Art

[0002] The OCR recognition model is a model specifically used to convert text information in an image into an editable text format; it extracts text from the image background by analyzing and recognizing features such as the shape and structure of characters in the image.

[0003] When the model is trained for recognition, it generally gives priority to defining templates for various characters, and matches the characters in the input image with the templates during recognition. For example, there are corresponding standard templates for the handwritten or printed numbers 0-9 and the letters AZ. During the recognition process, the similarity between the input character and each template is calculated, and a measurement method such as the minimum distance method is used, and the character corresponding to the most similar template is used as the recognition result.

[0004] In the actual application scenarios of OCR recognition models, text extraction and association recognition tasks of dynamic images are often encountered. When performing association training for dynamic images, the traditional approach is usually to move or rotate the image to determine the features used for training, and then carry out enhancement work in the dynamic data recognition process; however, this original enhancement method has certain limitations. It only focuses on the consideration of two single situations of image movement or rotation, and does not conduct a comprehensive and in-depth analysis of the complex and comprehensive situation in which dynamic images may move and rotate at the same time; this one-sidedness results in unsatisfactory results for dynamic data enhancement during model training. Summary of the invention

[0005] In view of the deficiencies in the prior art, the present invention provides a dynamic data enhancement method in OCR recognition model training, which solves the problem of no comprehensive and in-depth analysis of the complex comprehensive situation in which dynamic images may move and rotate at the same time.

[0006] To achieve the above objectives, the present invention is implemented by the following technical scheme: a dynamic data enhancement method in OCR recognition model training, comprising the following steps:

[0007] Step 1: Based on the set training direction, which includes the horizontal movement direction and the rotation direction, a set of training directions are randomly selected to train the model. The associated movement speed and movement acceleration are different during each training process. According to the specific training results, the direction that meets the standard or does not meet the standard is determined. The training method is:

[0008] S11, recording the set horizontal moving direction as the associated direction, and moving the training image in the associated direction. The initial moving speed and moving acceleration corresponding to each different moving stage are different. The training image is a preset image, and the initial moving speed and moving acceleration are prepared in advance by the operator. There are one or more groups of horizontal moving directions, so each different horizontal moving direction needs to be recorded as an associated direction and related processing is performed;

[0009] S12, record the text data generated in the corresponding moving stage, the text data includes text and time value, the time value is the period from the start of the training image movement to the end of the model outputting text, record the text and time value generated in each moving stage one by one, and based on the large amount of text data generated in this association direction, confirm whether this association direction is the target direction, the specific sub-steps are:

[0010] S121. Accuracy of recognizing characters from a large amount of character data. Compare the characters in the character data with the standard characters. When they are completely consistent, determine the accurate characters. The standard characters are characters preset in advance. Mark the accurate characters as qualified characters. Record the ratio of qualified characters to the total number of characters. The ratio = the total number of qualified characters / the total number of characters. If the ratio is ≥ 97%, it means that the character recognition in this related direction has met the standard. Otherwise, it means that the character recognition in this related direction has not met the standard.

[0011] S122, performing average processing on the time values ​​corresponding to the recorded multiple different moving stages to determine the associated average value, if the associated average value ≥ Y1, where Y1 is a preset value, it means that the time interval of the associated direction does not meet the standard, otherwise, it means that the time interval of the associated direction meets the standard;

[0012] S123, marking the associated direction where both the time interval and the text recognition meet the standards as the standard direction, otherwise marking the corresponding associated direction as the non-standard direction, the standard direction is the standard horizontal moving direction, and the non-standard direction is the non-standard horizontal moving direction;

[0013] When the training direction is the rotation direction, the model training method is:

[0014] The set rotation direction is recorded as the associated direction, and the training image is moved in the associated direction. The initial rotation speed and the rotation acceleration corresponding to each different moving stage are different. There are one or more groups of rotation directions, so each different rotation direction needs to be recorded as an associated direction and related processing is performed. Then, the same method as steps S12 and S121-S123 is used to confirm whether the associated direction is a standard direction. The standard direction is the standard rotation direction, and the non-standard direction is the non-standard rotation direction.

[0015] Step 2: Based on the determined non-standard direction, determine the model area associated with the non-standard direction in the model, enhance the original depth and width of the model area, and perform real-time training during the enhancement process to stop the enhancement when the non-standard direction is transformed into the standard direction. The specific sub-steps are:

[0016] S21. Based on the determined non-conforming direction, the model area corresponding to the non-conforming direction is locked from the model, and the original depth and width of the model area are recorded to confirm the non-conforming reason of the non-conforming direction:

[0017] If the reason for failure to meet the standard is that the text recognition fails to meet the standard, the original depth of this model area is enhanced, and a convolution layer that does not exist in this model area is added, so that the number of convolution layers is increased on the original basis, and the addition of new convolution layers is stopped when the direction of failure to meet the standard turns into the direction of meeting the standard;

[0018] If the reason for failure to meet the standard is that the time interval does not meet the standard, the original width of this model area is enhanced, and the number of neurons in this model area is increased, so that the number of neurons is increased on the original basis, and the number of neurons is stopped when the direction of failure to meet the standard changes to the direction of meeting the standard;

[0019] If the reason for failure to meet the standard is not only that the text recognition fails to meet the standard, but also that the time interval fails to meet the standard, the depth and width of the corresponding model area are enhanced simultaneously, and the enhancement is stopped when the direction of failure to meet the standard turns into the direction of meeting the standard;

[0020] Step 3: Based on the determined target-reaching direction, determine the fusion direction from the target-reaching direction, and perform fusion training on the model based on the determined fusion direction, and first identify whether the training stage reaches the target. If it reaches the target, there is no need to perform association enhancement. If it does not reach the target, association enhancement is performed by allocating the associated CPU utilization. The specific sub-steps are:

[0021] S31, randomly selecting a set of target horizontal movement directions and a set of target rotation directions from the target directions, combining them and recording them as fusion directions, so that the training image moves according to the target horizontal movement direction and rotates according to the target rotation direction at the same time, and the movement speed and rotation speed in the movement and rotation process are set in advance, and the movement speed and rotation speed corresponding to each different movement stage are different, and the movement acceleration and rotation acceleration are not set in this movement stage;

[0022] S32, recording the time value generated by the corresponding moving stage, where the time value is the period from when the training image moves to when the model outputs text, and performing mean processing on the time value associated with each different moving stage, determining the relevant mean, and identifying whether the relevant mean satisfies: relevant mean ≥ Y2, where Y2>Y1, and Y2 is a preset value;

[0023] If it is satisfied, it means that this fusion direction does not meet the standard and needs to be enhanced;

[0024] If not, it means that this fusion direction has met the standard and no further processing is required;

[0025] The association enhancement method for when the fusion direction does not meet the standard is as follows:

[0026] S321, preferentially record the CPU utilization LY occupied by the corresponding fusion direction at the current training moment, and perform correlation extraction in the CPU buffer of this model. Part of the utilization in this CPU buffer belongs to unused resources, which are called during training to increase LY. The increased value is one unit of utilization, and the unit is a preset unit;

[0027] S322. Identify the change difference of the relevant mean after LY increases the utilization rate by one unit: this change difference = the relevant mean before the utilization rate is increased - the relevant mean after the utilization rate is increased. The determined change difference is calibrated as Bz. Then identify the total difference Zc between the relevant mean before the utilization rate is increased and the preset value Y2. Zc = Y2 - the relevant mean before the utilization rate is increased. Use Zc ÷ Bz = Ts to confirm the unit utilization rate to be increased. Then increase the utilization rate of LY by (Ts-1) units, and lock the relevant mean JJ after the increase:

[0028] If JJ<Y2, then gradually reduce the changed LY so that JJ satisfies: JJ<Y2 and the difference between Y2 and JJ is the smallest, and lock the LY value corresponding to the moment before the stop moment as the execution value, and then perform data identification on this fusion direction based on this execution value;

[0029] If JJ=Y2, the LY value increased at the current moment is directly used as the execution value, and the data recognition of this fusion direction is performed based on this execution value later;

[0030] If JJ>Y2, then gradually increase LY again, and identify whether the corresponding JJ after the increase is lower than the corresponding JJ before the increase:

[0031] If so, then gradually increase LY until JJ satisfies: JJ ≤ Y2 or JJ does not decrease;

[0032] If not, stop increasing LY, and use the increased LY value at the current moment as the execution value, and subsequently perform data identification on this fusion direction based on this execution value.

[0033] The present invention provides a dynamic data enhancement method in OCR recognition model training. Compared with the prior art, it has the following beneficial effects:

[0034] The present invention trains the set training directions in sequence, identifies the standard direction and the non-standard direction based on the specific training results, and proportionally enhances the associated depth and width of the specified model area based on the determined non-standard direction. Based on the specific enhancement process, the corresponding non-standard direction is transformed into the standard direction, thereby improving the comprehensive recognition ability of the model and preliminarily enhancing the recognition ability of dynamic data.

[0035] Subsequently, the corresponding fusion direction is determined based on the confirmed target-reaching direction, and then associated training is performed through the fusion direction. Based on the specific training process, the corresponding training data is determined, and whether the time interval meets the target is identified from the training data. For non-target situations, the associated enhancement is performed by allocating the corresponding CPU utilization association method, and based on the specific allocation process, the training ability of the corresponding model in the corresponding fusion direction reaches the optimal state. When the fusion direction is dynamically analyzed again, its corresponding analysis ability can be fully guaranteed, so that the output capacity of dynamic data reaches the strongest state. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] Figure 1 It is a schematic diagram of the process of the present invention;

[0037] Figure 2 It is a schematic diagram of generating the fusion direction of the present invention. DETAILED DESCRIPTION

[0038] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0039] See also Figure 1 , the present application provides a method for dynamic data enhancement in OCR recognition model training. When training the OCR recognition model, data should be collected from different types of documents, such as scanned books, magazines, invoices, reports, etc. in priority; these documents should cover various fonts, font sizes, typesetting methods, and languages, and based on a specific C4.5 decision tree algorithm, the text can be classified and recognized according to image features (such as character stroke features, character structure features, etc.), and the text recognition capability of the OCR recognition model can be realized through related training of large-volume data. However, this type of recognition is only for static recognition. For dynamic recognition, since different text image features exist at different moving positions, the OCR recognition model needs to undergo multiple related trainings to enhance the recognition capability of the corresponding dynamic data, wherein the method for enhancing the recognition capability of dynamic data specifically includes the following steps:

[0040] Step 1: Based on the set training direction, which includes the horizontal movement direction and the rotation direction, both are set in advance by the relevant operators, a group of training directions are randomly selected to train the model. The associated movement speed and movement acceleration are different during each training process. According to the specific training results, the direction that meets the standard or the direction that does not meet the standard is determined. Specifically, the direction that meets the standard includes the horizontal movement direction that meets the standard or the horizontal movement direction that does not meet the standard, and the rotation direction includes the rotation direction that meets the standard or the rotation direction that does not meet the standard. The specific sub-steps for determination are:

[0041] When the training direction is horizontal movement:

[0042] S11, record the set horizontal moving direction as the associated direction, and move the training image in the associated direction. The initial moving speed and moving acceleration corresponding to each different moving stage are different. The training image is a preset image. The initial moving speed and moving acceleration are both prepared in advance by the operator and are changed in association according to the set logic. For example, if the initial speed is 0-10 (excluding 0) and the acceleration is 0-5 (including 0), a large amount of moving feature data will be generated after combination. Assuming that the initial speed is 1 and the acceleration is 0, then the training image moves in the associated direction at a speed of 1 in this moving stage. If the acceleration is 1, then the training image moves in the associated direction at an initial speed of 1 and an acceleration of 1. There are one or more groups of horizontal moving directions, so each different horizontal moving direction needs to be recorded as an associated direction and related processing is performed, which is set in advance by the relevant operator.

[0043] S12, record the text data generated in the corresponding moving stage, the text data includes text and time value, the time value is the period from the time when the training image moves to the time when the model outputs text, record the text and time value generated in each moving stage one by one, and confirm whether this associated direction is the standard direction based on the large amount of text data generated in this associated direction. For example: the text data here is associated with: ABC, and its standard is ABC, then the text ABC recognized in this moving stage meets the standard. If the text recognized in the corresponding stage is: AB or AC, then it means that the recognized text does not meet the standard. That is, the text here can be regarded as the corresponding text column, and only when it is completely consistent, it is considered to be a standard text:

[0044] S121. Accuracy of recognizing characters from a large amount of character data. Compare the characters in the character data with the standard characters. When they are completely consistent, determine the accurate characters. The standard characters are characters preset in advance. Mark the accurate characters as qualified characters. Record the ratio of qualified characters to the total number of characters. The ratio = the total number of qualified characters / the total number of characters (the total number of characters is the sum of the total number of qualified characters and the total number of unqualified characters). If the ratio is ≥ 97%, it means that the character recognition in this associated direction has met the standard. Otherwise, it means that the character recognition in this associated direction has not met the standard.

[0045] S122, performing mean processing on the time values ​​corresponding to the recorded multiple different moving stages to determine the associated mean value, if the associated mean value ≥ Y1, where Y1 is a preset value, and its specific value is determined by the operator based on experience, it means that the time interval of the associated direction does not meet the standard, otherwise, it means that the time interval of the associated direction meets the standard;

[0046] S123, marking the associated direction where both the time interval and the text recognition meet the standards as the standard direction, otherwise marking the corresponding associated direction as the non-standard direction, the standard direction is the standard horizontal moving direction, and the non-standard direction is the non-standard horizontal moving direction;

[0047] When the training direction is the rotation direction:

[0048] The set rotation direction is recorded as the associated direction, and the training image is moved in the associated direction. The initial rotation speed and the rotation acceleration corresponding to each different movement stage are different. The initial rotation speed and the rotation acceleration are both prepared in advance by the operator, and are consistent with the change logic associated with the initial movement speed and the movement acceleration. There are one or more groups of rotation directions, so each different rotation direction needs to be recorded as an associated direction and related processing is performed, which is set in advance by the relevant operator, and then the same method as step S12 is used to confirm whether the associated direction is a standard direction (synchronously execute steps S121-S123). The standard direction is the standard rotation direction, and the non-standard direction is the non-standard rotation direction;

[0049] Because the processing content of the rotation direction is relatively consistent with that of the movement direction, an abbreviated description method is used here to confirm the rotation direction, so as to complete the specific confirmation of the corresponding standard direction and the non-standard direction. For the confirmed non-standard direction, the model area of ​​this model about this part of the direction is strengthened, so that this non-standard direction is gradually transformed into the standard direction. By strengthening the depth and width of the corresponding model area, the depth is the number of convolutional layers added to the corresponding model area to make it more recognizable text features, and the width is the increase in the number of corresponding neurons to improve the overall computing power of the corresponding model area;

[0050] Step 2: Based on the determined non-standard direction, determine the model area associated with the non-standard direction in the model, enhance the original depth and width of the model area, and perform real-time training during the enhancement process to stop the enhancement when the non-standard direction is transformed into the standard direction. The specific sub-steps of the enhancement are:

[0051] S21. Based on the determined non-conforming direction, the model area corresponding to the non-conforming direction is locked from the model, and the original depth and width of the model area are recorded to confirm the non-conforming reason of the non-conforming direction:

[0052] If the reason for failure to meet the standard is that the text recognition fails to meet the standard, the original depth of this model area is enhanced, and a convolution layer that does not exist in this model area is added, so that the number of convolution layers is increased on the original basis, and the addition is stopped when the failure direction turns into the standard direction. Specifically, a text recognition error means that the corresponding text feature does not exist in the corresponding model, and its text feature corresponds to the convolution layer feature set in the model area. Each convolution layer corresponds to a set of text features. By adding a corresponding convolution layer to ensure a text recognition capability of this model area, the error rate can be fully reduced;

[0053] If the reason for failure to meet the standard is that the time interval does not meet the standard, the original width of this model area is enhanced, and the number of neurons in this model area is increased, so that the number of neurons is increased on the original basis, and the number of neurons is stopped when the direction of failure to meet the standard turns to the direction of meeting the standard. Specifically, the failure of the time interval to meet the standard means that the original computing power is insufficient, resulting in insufficient analysis ability, so the analysis ability of the corresponding area needs to be enhanced, so the corresponding number of neurons needs to be added;

[0054] If the reason for failure to meet the standard is not only that the text recognition fails to meet the standard, but also that the time interval fails to meet the standard, the depth and width of the corresponding model area are enhanced simultaneously, and the enhancement is stopped when the direction of failure to meet the standard turns into the direction of meeting the standard;

[0055] Furthermore, when the number of convolutional layers in the corresponding model area increases to 150 groups or the total number of neurons reaches 200, the number of convolutional layers stops increasing. When the number of convolutional layers and neurons in the corresponding model area increases to the limit value and still cannot turn the non-compliant direction into the compliant direction, an error signal is directly generated for display, indicating that manual intervention is required. The limit value here is the corresponding 150 groups or 200.

[0056] Specifically, the reason why the corresponding model has errors and large delays is either that the internally stored features are insufficient or that the internal computing power and analysis capabilities are insufficient. In order to gradually make it meet the standards, it is necessary to manage the depth and width of the associated area so that it gradually tends to meet the standards.

[0057] Step 3: Based on the determined target-reaching direction, determine the fusion direction from the target-reaching direction, and perform fusion training on the model based on the determined fusion direction, and first identify whether the training stage meets the target. If it meets the target, there is no need to perform association enhancement. If it does not meet the target, perform association enhancement by allocating the associated CPU utilization. The specific sub-steps of performing fusion training are:

[0058] S31, randomly select a set of target horizontal movement directions and a set of target rotation directions from the target directions, combine them and record them as fusion directions, so that the training image moves according to the target horizontal movement direction and rotates according to the target rotation direction, and the movement speed and rotation speed during the movement and rotation process are set in advance, and the movement speed and rotation speed corresponding to each different movement stage are different, and the movement acceleration and rotation acceleration are not set in this movement stage, combined Figure 2 , which is a schematic diagram of the display of a two-dimensional plane. By combining the horizontal movement direction and the rotation direction in the two-dimensional plane, the corresponding training image can be rotated in the horizontal movement direction, thereby completing the associated movement of the corresponding training image according to the fusion direction;

[0059] S32, record the time value generated by the corresponding moving stage, the time value is the period from the start of the training image moving to the end of the model outputting text, and average the time values ​​associated with each different moving stage, determine the relevant mean, and identify whether the relevant mean satisfies: relevant mean ≥ Y2, where Y2> Y1, and Y2 is a preset value. Specifically, when performing association analysis training here, there is no need to identify the accuracy of the text, so that all are in the standard direction, so there is no problem with the accuracy of the text, but after the direction is combined, the corresponding model output time will change, and the adaptability will be reduced, so it is necessary to analyze the relevant mean generated to assess whether the corresponding fusion direction is up to standard;

[0060] If it is satisfied, it means that this fusion direction does not meet the standard and needs to be enhanced to ensure that the time value is minimized or meets the standard;

[0061] If not, it means that this fusion direction has met the standard and no further processing is required;

[0062] The related ways of association enhancement are:

[0063] S321, preferentially record the CPU utilization LY occupied by the corresponding fusion direction at the current training moment, and perform correlation extraction from the CPU buffer of this model. Part of the utilization in this CPU buffer belongs to unused resources, which are called during training to increase LY. The increased value is a unit of utilization, and the unit is a preset unit, which can be understood as a percentage or a thousandth;

[0064] S322. Identify the change difference of the relevant mean after LY increases the utilization rate by one unit: this change difference = the relevant mean before the utilization rate is increased - the relevant mean after the utilization rate is increased. The determined change difference is calibrated as Bz. Then identify the total difference Zc between the relevant mean before the utilization rate is increased and the preset value Y2. Zc = Y2 - the relevant mean before the utilization rate is increased. Use Zc ÷ Bz = Ts to confirm the unit utilization rate to be increased. Then increase the utilization rate of LY by (Ts-1) units, and lock the relevant mean JJ after the increase:

[0065] If JJ<Y2, then gradually reduce the changed LY to make JJ satisfy: JJ<Y2 and the difference between Y2 and JJ is the smallest (difference = Y2-JJ, synchronously maintain JJ<Y2, and JJ cannot be ≥Y2), and lock the LY value corresponding to the moment before the stop moment as the execution value. Subsequently, data identification is performed on this fusion direction based on this execution value. Specifically, JJ is proposed to be 5, and Y2 is set to 7. Then after analysis and processing, its JJ is much smaller than Y2, so there is an excessive waste of CPU utilization. In order to make it meet the standard and not cause CPU waste, the LY value after the improvement process is gradually reduced, and the corresponding JJ value is determined in real time. When JJ gradually approaches Y2 and is at the maximum state, it stops, thereby achieving better CPU utilization effect;

[0066] If JJ=Y2, the LY value increased at the current moment is directly used as the execution value, and the data recognition of this fusion direction is performed based on this execution value later;

[0067] If JJ>Y2, then gradually increase LY again, and identify whether the corresponding JJ after the increase is lower than the corresponding JJ before the increase:

[0068] If so, then gradually increase LY until JJ satisfies: JJ ≤ Y2 or JJ does not decrease;

[0069] If not, stop increasing LY, and use the increased LY value at the current moment as the execution value, and then perform data identification on this fusion direction based on this execution value;

[0070] Specifically, when JJ=Y2, it means that the standard is just met. In this case, it can be understood that the corresponding CPU utilization is fully utilized and no associated processing is required;

[0071] If JJ>Y2, it means that after the improvement, the corresponding time interval still does not meet the standard. In order to identify whether such time interval is affected by the actual analysis capability of the corresponding model, the method of further improving the utilization rate LY is adopted here to reduce the corresponding JJ. If JJ is in a reduced state (that is, the corresponding JJ after the improvement is lower than the corresponding JJ before the improvement), then the utilization rate LY can be continued to be improved to make JJ reach the minimum state. If JJ does not decrease, there is no need to improve its LY again, because no matter how the corresponding LY is improved in the future, its JJ will not decrease at all, because the analysis capability of the corresponding model is limited, no matter how much LY is shared, it cannot meet the corresponding analysis behavior.

[0072] Some of the data in the above formulas are dimensionless and numerically calculated. Meanwhile, the contents not described in detail in this specification belong to the prior art known to those skilled in the art.

[0073] The above embodiments are only used to illustrate the technical method of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical method of the present invention may be modified or replaced by equivalents without departing from the spirit and scope of the technical method of the present invention.

Claims

1. A dynamic data enhancement method in OCR recognition model training, characterized in that: The following steps are involved: Step 1: Based on the set training direction, which includes the horizontal movement direction and the rotation direction, a group of training directions are randomly selected to train the model. The associated movement speed and movement acceleration are different during each training process. According to the specific training results, the direction that meets the standard or does not meet the standard is determined; Step 2: Based on the determined non-standard direction, determine the model area associated with the non-standard direction in the model, enhance the original depth and width of the model area, and perform real-time training during the enhancement process to stop the enhancement when the non-standard direction is transformed into the standard direction; Step 3: Based on the determined target-reaching direction, determine the fusion direction from the target-reaching direction, and perform fusion training on the model based on the determined fusion direction, and first identify whether the training stage reaches the target. If it reaches the target, there is no need to perform association enhancement. If it does not reach the target, association enhancement is performed by allocating the associated CPU utilization. The specific sub-steps are: S31, randomly selecting a set of target horizontal movement directions and a set of target rotation directions from the target directions, combining them and recording them as fusion directions, so that the training image moves according to the target horizontal movement direction and rotates according to the target rotation direction at the same time, and the movement speed and rotation speed in the movement and rotation process are set in advance, and the movement speed and rotation speed corresponding to each different movement stage are different, and the movement acceleration and rotation acceleration are not set in this movement stage; S32, recording the time value generated by the corresponding moving stage, where the time value is the period from when the training image moves to when the model outputs text, and performing mean processing on the time value associated with each different moving stage, determining the relevant mean, and identifying whether the relevant mean satisfies: relevant mean ≥ Y2, where Y2>Y1, and Y2 is a preset value; If not, it means that this fusion direction has met the standard and no further processing is required; If it is satisfied, it means that this fusion direction does not meet the standard and needs to be enhanced. The enhancement method is: S321, preferentially record the CPU utilization LY occupied by the corresponding fusion direction at the current training moment, and perform correlation extraction in the CPU buffer of this model. Part of the utilization in this CPU buffer belongs to unused resources, which are called during training to increase LY. The increased value is one unit of utilization, and the unit is a preset unit; S322. Identify the change difference of the relevant mean after LY increases the utilization rate by one unit: this change difference = the relevant mean before the utilization rate is increased - the relevant mean after the utilization rate is increased. The determined change difference is calibrated as Bz. Then identify the total difference Zc between the relevant mean before the utilization rate is increased and the preset value Y2. Zc = Y2 - the relevant mean before the utilization rate is increased. Use Zc ÷ Bz = Ts to confirm the unit utilization rate to be increased. Then increase the utilization rate of LY by (Ts-1) units, and lock the relevant mean JJ after the increase: If JJ<Y2, then gradually reduce the changed LY so that JJ satisfies: JJ<Y2 and the difference between Y2 and JJ is the smallest, and lock the LY value corresponding to the moment before the stop moment as the execution value, and then perform data identification on this fusion direction based on this execution value; If JJ=Y2, the LY value increased at the current moment is directly used as the execution value, and the data recognition of this fusion direction is performed based on this execution value later; If JJ>Y2, then gradually increase LY again, and identify whether the corresponding JJ after the increase is lower than the corresponding JJ before the increase: If so, then gradually increase LY until JJ satisfies: JJ ≤ Y2 or JJ does not decrease; If not, stop increasing LY, and use the increased LY value at the current moment as the execution value, and subsequently perform data identification on this fusion direction based on this execution value.

2. The method for dynamic data enhancement in OCR recognition model training according to claim 1, characterized in that: In the step 1, the training method of the model when the training direction is the horizontal movement direction is: S11, recording the set horizontal moving direction as the associated direction, and moving the training image in the associated direction. The initial moving speed and moving acceleration corresponding to each different moving stage are different. The training image is a preset image, and the initial moving speed and moving acceleration are prepared in advance by the operator. There are one or more groups of horizontal moving directions, so each different horizontal moving direction needs to be recorded as an associated direction and related processing is performed; S12. Record the text data generated in the corresponding moving stage. The text data includes text and time value. The time value is the period from the time when the training image moves to the time when the model outputs text. The text and time value generated in each moving stage are recorded one by one, and based on the large amount of text data generated in this associated direction, confirm whether this associated direction is the target direction.

3. The method for dynamic data enhancement in OCR recognition model training according to claim 2, characterized in that: In step S12, the specific sub-steps for confirming whether the association direction is a target-reaching direction are: S121. Accuracy of recognizing characters from a large amount of character data. Compare the characters in the character data with the standard characters. When they are completely consistent, determine the accurate characters. The standard characters are characters preset in advance. Mark the accurate characters as qualified characters. Record the ratio of qualified characters to the total number of characters. The ratio = the total number of qualified characters / the total number of characters. If the ratio is ≥ 97%, it means that the character recognition in this related direction has met the standard. Otherwise, it means that the character recognition in this related direction has not met the standard. S122, performing average processing on the time values ​​corresponding to the recorded multiple different moving stages to determine the associated average value, if the associated average value ≥ Y1, where Y1 is a preset value, it means that the time interval of the associated direction does not meet the standard, otherwise, it means that the time interval of the associated direction meets the standard; S123, mark the associated direction where both the time interval and the text recognition meet the standards as the standard direction, otherwise mark the corresponding associated direction as the non-standard direction, the standard direction is the standard horizontal moving direction, and the non-standard direction is the non-standard horizontal moving direction.

4. The method for dynamic data enhancement in OCR recognition model training according to claim 3, characterized in that: In step 1, the training method of the model when the training direction is the rotation direction is: The set rotation direction is recorded as the associated direction, and the training image is moved in the associated direction. The initial rotation velocity and rotation acceleration corresponding to each different moving stage are different. There are one or more groups of rotation directions, so each different rotation direction needs to be recorded as an associated direction and related processing is performed. Then, the same method as steps S12 and S121-S123 is used to confirm whether this associated direction is a standard direction. This standard direction is a standard rotation direction, and this non-standard direction is a non-standard rotation direction.

5. The method for dynamic data enhancement in OCR recognition model training according to claim 1, characterized in that: In step 2, the specific sub-steps for enhancing the original depth and width of the model area are: S21. Based on the determined non-conforming direction, the model area corresponding to the non-conforming direction is locked from the model, and the original depth and width of the model area are recorded to confirm the non-conforming reason of the non-conforming direction: If the reason for failure to meet the standard is that the text recognition fails to meet the standard, the original depth of this model area is enhanced, and a convolution layer that does not exist in this model area is added, so that the number of convolution layers is increased on the original basis, and the addition of new convolution layers is stopped when the direction of failure to meet the standard turns into the direction of meeting the standard; If the reason for failure to meet the standard is that the time interval does not meet the standard, the original width of this model area is enhanced, and the number of neurons in this model area is increased, so that the number of neurons is increased on the original basis, and the number of neurons is stopped when the direction of failure to meet the standard changes to the direction of meeting the standard; If the reason for non-compliance is not only that the text recognition fails to meet the standard, but also that the time interval fails to meet the standard, the depth and width of the corresponding model area are enhanced simultaneously, and the enhancement is stopped when the non-compliance direction turns into the compliance direction.

6. The method for dynamic data enhancement in OCR recognition model training according to claim 5, characterized in that: In the step S21, when the number of convolutional layers and neurons in the corresponding model area increases to the limit value and still cannot convert the non-compliant direction into the compliant direction, an error signal is directly generated for display, and the limit value is: 150 groups of convolutional layers or 200 total neurons.

Citation Information

Patent Citations

  • Scene text end-to-end identification method based on boundary point detection

    CN110837835A

  • OCR (Optical Character Recognition) model training method, OCR method and related device

    CN114565751A