A Model Adaptive Text Recognition Method and System for Decentralized Scenarios

By designing pseudo-label filtering methods based on confidence and uncertainty and diversity metric judgments based on confidence and uncertainty in decentralized scenarios, combined with integrated selection strategies, the challenge of adaptive text recognition of multi-source models is solved, and a more efficient text image recognition effect is achieved.

CN116434216BActive Publication Date: 2025-07-22ZHEJIANG UNIV +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310320095.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-03-29
Publication Date
2025-07-22
Estimated Expiration
2043-03-29

AI Technical Summary

Technical Problem

There is a lack of effective methods in the prior art to deal with model adaptive text recognition tasks in decentralized scenarios, especially in the case of style differences and potential malicious model interference between multiple source models, making it difficult to achieve accurate text image prediction.

Method used

By designing a pseudo-label screening method based on confidence and uncertainty, combining diversity metrics to determine the availability of pseudo-label pairs, and using an integrated selection strategy in the prediction stage, filter out pseudo-label pairs that can be used for adaptive training, and finally text recognition is performed through the adaptive training model.

Benefits of technology

It improves the adaptability of the text recognition model in decentralized scenarios, improves the accuracy and reliability of text image prediction, and reduces the typo and malword rates.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116434216B_ABST
    Figure CN116434216B_ABST
Patent Text Reader

Abstract

The present invention discloses a model adaptive text recognition method and system for a decentralized scenario. The method of the present invention includes the steps of: using multiple text recognition source models to predict text images in a set of target domains to obtain character sequence labels predicted by all models; screening based on confidence and uncertainty, and forming character sequences from the qualified character sequence labels, and using the corresponding text images as pseudo-label pairs; judging whether the pseudo-label pairs can be used for the adaptive training of the model based on diversity metrics, and if not, removing them, and forming a training set from the remaining pseudo-label pairs; using the training set to perform adaptive training on the model; the trained model recognizes the text images to be measured, and uses an ensemble selection strategy to determine the final text recognition result. The present invention designs a new pseudo-label screening strategy in a decentralized scenario, and only uses multiple models and unlabeled target domain text images to achieve the effect of model adaptive text recognition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of text recognition, and particularly to a model adaptive text recognition method and system for a decentralized scenario. Background Art

[0002] In the model adaptive text recognition task for a decentralized scenario, there are multiple source models. These source models are trained based on different source domains, and they may come from different devices. There may also be some malicious models among these source models, which may even interfere with the result prediction. The purpose of this task is to adaptively train these source models using unlabeled target domain data, and then use these adaptively trained text recognition models to accurately predict the text images in the target domain.

[0003] Since this task involves multiple source models and there are significant style differences between the source domain and the target domain, this task is very challenging. Currently, there have been some studies on the model adaptive text recognition task for a single model. However, there is currently no publicly available method for a decentralized scenario.

[0004] In summary, there is currently no model adaptive text recognition method for a decentralized scenario, and this task is a challenging new task. Summary of the Invention

[0005] The purpose of the present invention is to propose a method and system for a model adaptive text recognition task for a decentralized scenario. The present invention screens pseudo-labels by designing a pseudo-label screening method based on confidence and uncertainty, and determines the pseudo-label pairs for model adaptive training by judging whether the pseudo-label pairs can be used for the adaptive training of the source models through diversity measurement. Finally, at the prediction stage, an ensemble selection strategy is designed to determine the final prediction result for the text images in the target domain.

[0006] The specific technical solution adopted by the present invention is as follows:

[0007] In a first aspect, the present invention proposes a model adaptive text recognition method for a decentralized scenario, comprising the following steps:

[0008] 1) Collect multiple pre-trained text recognition source models from different scenarios, as well as unlabeled text images in the target domain;

[0009] 2) Use all the text recognition source models in step 1) to predict a group of text images in the target domain. After a text image is predicted by all the text recognition source models, a group of character sequence labels is obtained;

[0010] 3) Screen the multiple groups of character sequence tags obtained in step 2) based on confidence and uncertainty, and form a character sequence by combining the qualified character sequence tags in the same group. The character sequence and its corresponding text image are used as a pseudo-label pair;

[0011] 4) Based on the diversity metric, determine whether the pseudo-label pairs screened in step 3) can be used for the adaptive training of the text recognition source model. If not, eliminate them, and the remaining pseudo-label pairs form the training set;

[0012] 5) Use the training set obtained in step 4) to perform adaptive training on the text recognition source model;

[0013] 6) When recognizing the text image to be measured, use the integrated selection strategy for the text recognition source model after the adaptive training in step 5) to determine the final text recognition result.

[0014] Further, in step 2), for a group of target domain text images Use To represent the prediction of the text recognition source model Net i For it, where x j Represents the j-th unlabeled target domain text image, n t Represents the number of text images in the target domain, Net i Represents the i-th text recognition source model, Y i t Represents the character sequence label predicted by the i-th text recognition source model, y i,j Represents the character sequence label of the j-th target domain text image predicted by the i-th text recognition source model.

[0015] Further, step 3) includes:

[0016] 3.1) For a text image from the target domain, a character sequence y = {y1,... y l ,... y L , EOS} with a length of L + 1 is predicted through all text recognition source models. y l Represents the l-th valid character in the text image, L represents the length of the valid characters in the text image, and EOS represents the sequence terminator;

[0017] Assume that the number of text recognition source models is n. Then, for each step of prediction, there will be n prediction results. First, select the K prediction results with the largest softmax confidence scores. Take the mean of these K softmax vectors as the current softmax vector. Obtain the predicted character for the current step according to the current softmax vector, and then take the confidence score corresponding to the current softmax vector as the confidence score of the predicted character for the current step. Take the standard deviation of the selected K softmax vectors as the uncertainty score of the predicted character for the current step. Therefore, each predicted character has a corresponding softmax confidence score p l and an uncertainty score u l , and after all prediction steps, a character sequence of length L + 1 is obtained;

[0018] 3.2) The confidence score and uncertainty score of the character sequence are the means of the confidence scores and uncertainty scores of each prediction step. When the confidence score of the character sequence is greater than the first threshold δ d and the uncertainty score is less than the second threshold δ u , the character sequence meets the confidence and uncertainty conditions, and the character sequence and its corresponding text image are used as a pseudo-label pair.

[0019] The first threshold δ d and the second threshold δ u both have a value range of (0, 1).

[0020] Furthermore, the diversity metric judgment formula in step 4) is:

[0021]

[0022] where div represents the character diversity of the character sequence in the pseudo-label pair, m t represents the number of character types included in the character sequence in the pseudo-label pair, represents the number of characters included in the character sequence in the pseudo-label pair, and γ is the threshold.

[0023] Furthermore, the text recognition source model described in step 1) includes:

[0024] A regularization conversion module, which is used to regularize the input text image;

[0025] A visual feature extraction module, which is used to extract the visual features of the regularized text image;

[0026] A sequence modeling module, which is used to model the visual feature sequence;

[0027] A predictor module, which is used to perform dimensionality conversion on the modeled features and classify them to obtain the corresponding character sequence labels.

[0028] Furthermore, when adaptively training the text recognition source model in step 5), the parameters of the sequence modeling module and the predictor module are frozen, and only the parameters of the regularization transformation module and the visual feature extraction module are updated.

[0029] Further, the loss function for adaptive training is as follows:

[0030]

[0031] where Net i represents the i-th text recognition source model, represents a set of pseudo-label pairs, represents the set of text images in the set of pseudo-label pairs, represents the set of character sequences in the set of pseudo-label pairs, are the text image and the character sequence in the pseudo-label pair respectively, and θ i represents the parameters of the i-th text recognition source model, represents the loss value of the i-th text recognition source model.

[0032] Further, step 6) includes:

[0033] Using all text recognition source models to independently predict character sequences for the text images to be recognized. When performing the prediction at the current step, if the confidence score corresponding to the softmax vector of the predicted character is greater than the threshold ζ, then select the predicted character, and take the mean of the softmax vectors corresponding to all selected characters to construct a new softmax vector; then select the predicted character corresponding to the current step according to the best-first strategy based on the new softmax vector; if there is no selected character, use the character corresponding to the maximum confidence score as the character predicted at the current step.

[0034] The value range of the threshold ζ is (0, 1).

[0035] In a second aspect, the present invention proposes a model adaptive text recognition system for a decentralized scenario, including:

[0036] A data acquisition module, which is used to collect multiple pre-trained text recognition source models from different scenarios, and unlabeled text images in the target domain;

[0037] A text recognition source model module, which is used to predict a group of text images in the target domain, and one text image corresponds to one predicted character sequence label; the number of text recognition source model modules is the same as the number of text recognition source models in the data acquisition module, and a group of character sequence labels are obtained after one text image is predicted by all text recognition source model modules;

[0038] The first screening module is used to screen multiple groups of character sequence labels obtained by the text recognition source model module based on confidence and uncertainty, and form character sequences with the qualified character sequence labels in the same group. The character sequences and their corresponding text images are used as pseudo-label pairs.

[0039] The second screening module is used to judge whether the pseudo-label pairs screened by the first screening module can be used for the adaptive training of the text recognition source model based on diversity measurement. If not, they are eliminated, and the remaining pseudo-label pairs form a training set.

[0040] The adaptive training module is used to adaptively train the text recognition source model by using the training set obtained by the second screening module.

[0041] The text recognition module is used to recognize the text image to be measured by using the trained text recognition source model and determine the final text recognition result by using the integrated selection strategy.

[0042] Furthermore, the text recognition source model module is composed of a regularization conversion module, a visual feature extraction module, a sequence modeling module, and a predictor module. The adaptive training module only performs adaptive training on the regularization conversion module and the visual feature extraction module.

[0043] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0044] (1) The present invention designs a pseudo-label screening method based on the combination of confidence and uncertainty, which can screen out relatively accurate pseudo-label pairs. Combining with the self-training method, the present invention can improve the adaptability of the text recognition source model.

[0045] (2) The present invention designs a diversity judgment strategy, which can solve the problem that appears in some character-level label sets during the pseudo-label screening process, and further improve the adaptability of the text recognition source model. Description of the Drawings

[0046] Figure 1 is a schematic diagram of the present invention for text image recognition by using multiple pre-trained text recognition source models from different scenarios;

[0047] Figure 2 is a schematic diagram of the model adaptive text recognition method of the present invention for a decentralized scenario;

[0048] Figure 3 is a schematic diagram of the present invention for screening character sequence labels based on confidence and uncertainty. Detailed Embodiments

[0049] The present invention will be further described and explained below in conjunction with the accompanying drawings of the specification and specific embodiments.

[0050] As Figure 2 shown, the model adaptive text recognition method for the decentralized scenario proposed by the present invention mainly includes the following steps:

[0051] S1. Collect a plurality of pre-trained text recognition source models from different scenarios and unannotated text images in the target domain;

[0052] S2. Use all the text recognition source models in step S1 to predict a group of text images in the target domain. After a text image is predicted by all the text recognition source models, a group of character sequence labels is obtained;

[0053] S3. Screen the multiple groups of character sequence labels obtained in step S2 based on confidence and uncertainty, and form a character sequence by combining the character sequence labels that meet the conditions in the same group. The character sequence and its corresponding text image are used as a pseudo-label pair;

[0054] S4. Based on the diversity metric, determine whether the pseudo-label pairs screened in step S3 can be used for the adaptive training of the text recognition source model. If not, then eliminate them, and the remaining pseudo-label pairs form the training set;

[0055] S5. Use the training set obtained in step S4 to perform adaptive training on the text recognition source model;

[0056] S6. When recognizing the text image to be measured, as Figure 1 shown, use the integrated selection strategy for the text recognition source model after the adaptive training in step S5 to determine the final text recognition result.

[0057] In a specific implementation of the present invention, the text recognition models collected in step S1 from different scenarios have the same model structure, including:

[0058] A regularization conversion module, which is used to regularize the input text image. In this embodiment, the spatial transformation network STN is selected;

[0059] A visual feature extraction module, which is used to extract the visual features of the regularized text image. In this embodiment, the residual convolutional neural network ResNet is selected;

[0060] A sequence modeling module, which is used to model the visual feature sequence. In this embodiment, the bidirectional long short-term memory model BiLSTM is selected;

[0061] A predictor module, which is used to perform dimensionality conversion on the modeled features and classify them to obtain the corresponding character sequence labels. In this embodiment, a sequence prediction method based on the attention mechanism is selected.

[0062] Suppose there are n text recognition source models, denoted as {Net1, …, Net n}; use to represent the unlabeled target domain text images, and here N t refers to the number of text images in the target domain.

[0063] In a specific implementation of the present invention, in step S2, multiple text recognition source models are used to predict a group of text images in the target domain. For a group of target domain text images use to represent the prediction of the source model Net i for it. Among them, x j represents the j-th unlabeled target domain text image, n t represents the number of text images in the target domain, Net i represents the i-th text recognition source model, Y i t represents the character sequence label predicted by the i-th text recognition source model, and y i,j represents the character sequence label of the j-th target domain text image predicted by the i-th text recognition source model.

[0064] In a specific implementation of the present invention, in step S3, in order to initially screen out pseudo-label pairs, as Figure 3 shown, its implementation steps are as follows:

[0065] 3.1) For a text image from the target domain, a character sequence y = {y1, … y l , … y L , EOS} of length L + 1 is obtained through the prediction of all text recognition source models. y l represents the l-th valid character in the text image, L represents the length of the valid characters in the text image, and EOS represents the sequence end symbol;

[0066] Suppose the number of text recognition source models is n. Then, for each step of prediction, there will be n prediction results. First, select the K prediction results with the largest softmax confidence scores. Take the mean of these K softmax vectors as the current softmax vector. According to the current softmax vector, obtain the predicted character at the current step, and then use the confidence score corresponding to the current softmax vector as the confidence score of the predicted character at the current step; take the standard deviation of the selected K softmax vectors as the uncertainty score of the predicted character at the current step. Therefore, each predicted character has a corresponding softmax confidence score p l and an uncertainty score u l , and a character sequence of length L + 1 is obtained after all prediction steps;

[0067] 3.2) The confidence score p and uncertainty score u of the character sequence are the means of the confidence scores and uncertainty scores for each prediction step, expressed as:

[0068]

[0069]

[0070] When the confidence score of the character sequence is greater than the first threshold δ d and the uncertainty score is less than the second threshold δ u the character sequence meets the confidence and uncertainty conditions, and the character sequence and its corresponding text image are used as a pseudo-label pair. In this embodiment, δ d is set to 0.55, and δ u is set to 0.5.

[0071] In a specific implementation of the present invention, in step S4, the pseudo-label pairs are further screened based on diversity metrics. The diversity metric judgment formula is:

[0072]

[0073] where div represents the character diversity of the character sequence in the pseudo-label pair, m t represents the number of character types included in the character sequence in the pseudo-label pair, represents the number of characters included in the character sequence in the pseudo-label pair, and γ is the threshold.

[0074] For the text image group the image set in the corresponding pseudo-label pair screened by step S3 is and the corresponding character sequence set is Suppose includes character sequences, and these character sequences include a total of characters and m t different character types. In this embodiment, γ is set to 0.6.

[0075] In a specific implementation of the present invention, when adaptively training the text recognition source model, the parameters of the sequence modeling module and the predictor module are frozen, and only the parameters of the regularization transformation module and the visual feature extraction module are updated. Figure 2 where VP represents the regularization transformation module and the visual feature extraction module, SE represents the sequence modeling module, and PD represents the predictor module.

[0076] The loss function for adaptive training is as follows:

[0077]

[0078] Among them, Net i represents the i-th text recognition source model, represents the pseudo-label pair group, represents the set of text images in the pseudo-label pair group, represents the set of character sequences in the pseudo-label pair group, are respectively the text image and the character sequence in the pseudo-label pair, θ i represents the parameters of the i-th text recognition source model, represents the loss value of the i-th text recognition source model.

[0079] In a specific implementation of the present invention, step S6 is specifically as follows:

[0080] Use all text recognition source models to independently predict character sequences for the text image to be recognized. When performing the prediction of the current step, if the confidence score corresponding to the softmax vector of the predicted character is greater than the threshold ζ (in this embodiment, ζ is set to 0.9), then select the predicted character, and take the mean of the softmax vectors corresponding to all selected characters to construct a new softmax vector; then select the predicted character corresponding to the current step according to the best-first strategy based on the new softmax vector. If there is no selected character, use the character corresponding to the maximum confidence score as the character predicted in the current step.

[0081] To test the effectiveness of the above method, the present invention conducted experiments on the MSDA dataset, which is the existing largest-scale dataset for multi-domain text recognition, and contains a total of 5,209,215 Chinese text images; in addition, according to the different styles of text images, the MSDA dataset is divided into five different domains (scenarios): synthetic text domain (Sy), handwritten text domain (H), document text domain (D), street view text domain (St), and license plate text domain (C). Each of these five domains is divided into a training set and a test set according to a ratio of 9:1. In the MSDA dataset, there are a total of 3,816 different characters, including 3,754 Chinese character characters, 62 English letter characters, and numerical characters.

[0082] The evaluation metrics used in this invention include the Character Error Rate (CER for short) and the Word Error Rate (WER for short). The Character Error Rate (CER) is defined as the Levenstein distance between the predicted character sequence and the actual character sequence. The Word Error Rate (WER) represents the percentage of the incorrect character sequences recognized by the model. Obviously, for CER and WER, the lower the score, the better the performance of the model. This invention was compared with the baseline method (non - adaptive training) and the KD3A method, and the results are shown in Table 1.

[0083] Table 1 Test results of this invention, the baseline method, and the KD3A method on the MSDA dataset

[0084]

[0085] The baseline method directly takes the average of the source model softmax distributions as the final softmax distribution; the KD3A method is the best - performing method among the current domain - adaptive methods for the decentralized scenario of image classification. The KD3A method achieves domain adaptation for the decentralized scenario of the image classification task through data distillation. Here, the KD3A method is directly applied to the text recognition task and compared with the method proposed in this invention.

[0086] Compared with the baseline method, the method proposed in this invention screens out pseudo - labels for self - training of the source model, so the method proposed in this invention is better; compared with the KD3A method, this invention designs a domain - adaptive method for the text recognition task, designs a pseudo - label screening method and a diversity judgment strategy based on confidence and uncertainty, which is more targeted at the text recognition task, and this invention has achieved better results.

[0087] In this embodiment, a model - adaptive text recognition system for the decentralized scenario is also provided. This system is used to implement the above - mentioned embodiment. The following terms such as "module" and "unit" can be a combination of software and / or hardware that can achieve a predetermined function. Although the system described in the following embodiments is preferably implemented in software, implementation in hardware, or a combination of software and hardware is also possible.

[0088] A model - adaptive text recognition system for the decentralized scenario provided in this embodiment includes:

[0089] A data acquisition module, which is used to collect multiple pre - trained text recognition source models from different scenarios and unlabeled text images in the target domain;

[0090] The text recognition source model module is used to predict text images in a set of target domains, and one text image corresponds to a predicted character sequence label; the number of text recognition source model modules is the number of text recognition source models described in the data acquisition module, and a set of character sequence labels is obtained after a text image is predicted by all text recognition source model modules;

[0091] The first screening module is used to screen multiple sets of character sequence labels obtained by the text recognition source model module based on confidence and uncertainty, and form a character sequence with the character sequence labels that meet the conditions in the same group. The character sequence and its corresponding text image are used as a pseudo-label pair;

[0092] The second screening module is used to judge whether the pseudo-label pairs screened by the first screening module can be used for the adaptive training of the text recognition source model based on the diversity metric. If not, they are eliminated, and the remaining pseudo-label pairs form a training set;

[0093] The adaptive training module is used to adaptively train the text recognition source model by using the training set obtained by the second screening module;

[0094] The text recognition module is used to recognize the text image to be measured by using the trained text recognition source model and determine the final text recognition result by using the integrated selection strategy.

[0095] For the system embodiment, since it basically corresponds to the method embodiment, the relevant parts can be referred to the partial description of the method embodiment, and the implementation methods of the remaining modules are not elaborated here. The system embodiment described above is only illustrative, and the units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of the present invention. Those of ordinary skill in the art can understand and implement it without creative efforts.

[0096] The embodiment of the system of the present invention can be applied to any device with data processing capabilities, and the any device with data processing capabilities can be a device or apparatus such as a computer. The system embodiment can be implemented by software, or by hardware or a combination of software and hardware. Taking software implementation as an example, as a logically meaningful device, it is formed by the processor of any device with data processing capabilities reading the corresponding computer program instructions in the non-volatile memory into the memory and running.

[0097] Please note that the technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered to be within the scope described in this specification. It should be noted that for those of ordinary skill in the art, without departing from the concept of this application, several modifications and improvements can still be made, and these all belong to the protection scope of this application.

Claims

1. A model adaptive text recognition method for decentralized scenarios, characterized in that, The following steps are involved: 1) Collect multiple pre-trained text recognition source models from different scenarios and unlabeled text images in the target domain; 2) All text recognition source models in step 1) are used to predict a set of text images in the target domain. After a text image is predicted by all text recognition source models, a set of character sequence labels is obtained; 3) Based on confidence and uncertainty, the multiple groups of character sequence labels obtained in step 2) are screened, and the character sequence labels that meet the conditions in the same group are combined into a character sequence, and the character sequence and its corresponding text image are used as a pseudo label pair; 4) Based on the diversity metric, determine whether the pseudo-label pairs screened in step 3) can be used for adaptive training of the text recognition source model. If not, remove them, and the remaining pseudo-label pairs constitute the training set; 5) Adopting the training set obtained in step 4) to perform adaptive training on the text recognition source model; 6) When recognizing the text image to be tested, the text recognition source model after the adaptive training in step 5) is used to determine the final text recognition result using an integrated selection strategy; The step 3) comprises: 3.1) For a text image from the target domain, a character sequence y = {y1, … y l , … y L , EOS} of length L + 1 is predicted by the full text recognition source model, where y l represents the l-th valid character in the text image, L represents the length of the valid characters in the text image, and EOS represents the sequence end symbol; 3.2) The confidence score and uncertainty score of the character sequence are the means of the confidence scores and uncertainty scores for each prediction step. When the confidence score of the character sequence is greater than the first threshold δ d and the uncertainty score is less than the second threshold δ u the character sequence meets the confidence and uncertainty conditions, and the character sequence and its corresponding text image are used as a pseudo-label pair.

2. The model adaptive text recognition method for a decentralized scenario according to claim 1, wherein In step 2), for a set of target domain text images use to represent the prediction of the text recognition source model Net i for it, where x j represents the j-th unlabeled target domain text image, n t represents the number of text images in the target domain, Net i represents the i-th text recognition source model, represents the character sequence label predicted by the i-th text recognition source model, y i,j represents the character sequence label of the j-th target domain text image predicted by the i-th text recognition source model.

3. The model adaptive text recognition method for a decentralized scenario according to claim 1, wherein In step 3.1), assuming the number of text recognition source models is n, there will be n prediction results for each step of prediction. First, select the K prediction results with the largest softmax confidence scores. Take the mean of these K softmax vectors as the current softmax vector. Obtain the predicted character for the current step according to the current softmax vector, and then take the confidence score corresponding to the current softmax vector as the confidence score of the predicted character for the current step. Take the standard deviation of the selected K softmax vectors as the uncertainty score of the predicted character for the current step. Therefore, each predicted character has a corresponding softmax confidence score p l and an uncertainty score u l , and a character sequence of length L + 1 is obtained after all prediction steps.

4. The model adaptive text recognition method for a decentralized scenario according to claim 1, wherein The diversity measurement judgment formula in step 4) is: Among them, div represents the character diversity of the character sequence in the pseudo-label pair, and m t represents the number of character types included in the character sequence in the pseudo-label pair, represents the number of characters included in the character sequence in the pseudo-label pair, and γ is the threshold value.

5. The model adaptive text recognition method for a decentralized scenario according to claim 1, characterized in that, The text recognition source model described in step 1) includes: A regularization transformation module, which is used to regularize the input text image; A visual feature extraction module, which is used to extract visual features of the regularized text image; A sequence modeling module, which is used to model visual feature sequences; The predictor module is used to transform the dimensions of the modeled features and classify them to obtain the corresponding character sequence labels.

6. The model adaptive text recognition method for a decentralized scenario according to claim 5, wherein When the text recognition source model is adaptively trained in step 5), the parameters of the sequence modeling module and the predictor module are frozen, and only the parameters of the regularization conversion module and the visual feature extraction module are updated.

7. The model adaptive text recognition method for a decentralized scenario according to claim 6, wherein The loss function for adaptive training is as follows: Among them, Net i represents the i-th text recognition source model, represents the pseudo-label pair group, represents the set of text images in the pseudo-label pair group, represents the set of character sequences in the pseudo-label pair group, are respectively the text image and the character sequence in the pseudo-label pair, θ i represents the parameter of the i-th text recognition source model, represents the loss value of the i-th text recognition source model.

8. The model adaptive text recognition method for a decentralized scenario according to claim 1, wherein The step 6) comprises: All text recognition source models are used to independently predict the character sequence for the text image to be recognized. When executing the current step prediction, if the confidence score corresponding to the softmax vector of the predicted character is greater than the threshold ζ, the predicted character is selected, and the softmax vectors corresponding to all selected characters are averaged to construct a new softmax vector; then based on the new softmax vector, the predicted character corresponding to the current step is selected according to the best-first strategy; if there is no selected character, the character corresponding to the maximum confidence score will be used as the character predicted for the current step.

9. A model adaptive text recognition system for a decentralized scenario, which is used to implement the model adaptive text recognition method according to any one of claims 1-8, characterized in that include: A data acquisition module, which is used to collect multiple pre-trained text recognition source models from different scenarios and unlabeled text images of the target domain; A text recognition source model module is used to predict a set of text images in a target domain. One text image corresponds to one predicted character sequence label. The number of text recognition source model modules is the number of text recognition source models described in the data acquisition module. After a text image is predicted by all text recognition source model modules, a set of character sequence labels is obtained. The first screening module is used to screen multiple groups of character sequence tags obtained by the text recognition source model module based on confidence and uncertainty, and form a character sequence by combining the qualified character sequence tags in the same group. The character sequence and its corresponding text image are used as a pseudo-label pair; The second screening module is used to judge whether the pseudo-label pairs screened by the first screening module can be used for the adaptive training of the text recognition source model based on the diversity metric. If not, they are eliminated, and the remaining pseudo-label pairs form a training set; The adaptive training module is used to adaptively train the text recognition source model by using the training set obtained by the second screening module; The text recognition module is used to recognize the text image to be measured by using the trained text recognition source model and determine the final text recognition result by using the integrated selection strategy.

10. A model adaptive text recognition system for a decentralized scenario according to claim 9, characterized in that, The text recognition source model module is composed of a regularization conversion module, a visual feature extraction module, a sequence modeling module, and a predictor module. The adaptive training module only adaptively trains the regularization conversion module and the visual feature extraction module.

Citation Information

Patent Citations

  • Model adaptive text recognition method and system from printed form to handwritten form

    CN113592045A

  • Aggregation cross-entropy loss function-based sequence recognition method

    WO2020248471A1