Training method of image processing model, image processing method and related equipment
Patent Information
- Application Number
- CN202211236737.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-10-10
- Publication Date
- 2026-09-18
- Estimated Expiration
- 2042-10-10
AI Technical Summary
[0003]目前,训练图像处理模型需要收集大量训练样本,训练样本的分布情况通常会影响图像处理模型的识别能力,若多个样本类别的训练样本分布不均,会使得图像处理模型在一些样本类别分布的识别能力表现较好,在另一些样本类别分布的识别较差,导致图像处理模型的整体识别精度不高
[0009] The above scheme obtains a sample dataset containing K subsets of sample data. Different subsets within the K subsets contain samples of different categories, and at least two subsets contain samples of different numbers of samples. The dataset is then processed to obtain M sample data sequences, which are composed of multiple subsets sorted according to the distribution characteristics of the target sample category. Based on these M sample data sequences, an image processing model is trained. This trained model learns the distribution patterns of various categories of sample data subsets to a certain extent, while maintaining the model's representational ability for each category. Therefore, it improves the model's ability to learn the distribution patterns of different categories of sample data subsets and enhances the overall recognition accuracy of each category, thereby improving the overall image processing accuracy.
Smart Images

Figure CN115761754B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and in particular to a training method for an image processing model, an image processing method, a computer device, and a storage device. Background Technology
[0002] With the development of internet technology, a lot of useful information can be extracted from images, such as text recognition. Text recognition technology not only reduces manual recognition costs but also improves recognition efficiency. Therefore, the accuracy of text recognition is a crucial factor in evaluating its effectiveness. Text recognition technology typically uses image processing models to identify text in images, and the training of these models is particularly important for the accuracy of text recognition.
[0003] Currently, training image processing models requires collecting a large number of training samples. The distribution of training samples usually affects the recognition ability of image processing models. If the training samples of multiple sample categories are unevenly distributed, the image processing model will perform well in some sample category distributions but poorly in other sample category distributions, resulting in low overall recognition accuracy of the image processing model. Summary of the Invention
[0004] The main technical problem addressed in this application is to provide a training method for an image processing model, an image processing method, and related equipment, which can improve the accuracy of the image processing model in image processing.
[0005] To address the aforementioned issues, the first aspect of this application provides a method for training an image processing model. This method includes: acquiring a sample dataset; the sample dataset contains K subsets of sample data, where different subsets contain different categories of sample data, and at least two subsets contain different numbers of samples, where K is an integer greater than 1; processing the sample dataset to obtain M sample data sequences, where each sample data sequence is composed of multiple subsets of sample data sorted according to the distribution characteristics of the target sample category, where M is an integer greater than 1; and training the image processing model based on the M sample data sequences to obtain the trained image processing model.
[0006] To address the aforementioned issues, a second aspect of this application provides an image processing method, comprising: acquiring an image to be processed, wherein the image to be processed is an image containing text information; and processing the image to be processed using an image processing model trained by the aforementioned image processing model training method to obtain a processing result of the image to be processed.
[0007] To address the aforementioned problems, a third aspect of this application provides a computer device comprising a memory and a processor coupled to each other, wherein the memory stores program data and the processor executes the program data to implement any step of the training method and / or image processing method of the aforementioned image processing model.
[0008] To address the aforementioned problems, a fourth aspect of this application provides a storage device storing program data that can be executed by a processor. The program data is used to implement any step of the training method and / or image processing method of the aforementioned image processing model.
[0009] The above scheme obtains a sample dataset containing K subsets of sample data. Different subsets within the K subsets contain samples of different categories, and at least two subsets contain samples of different numbers of samples. The dataset is then processed to obtain M sample data sequences, which are composed of multiple subsets sorted according to the distribution characteristics of the target sample category. Based on these M sample data sequences, an image processing model is trained. This trained model learns the distribution patterns of various categories of sample data subsets to a certain extent, while maintaining the model's representational ability for each category. Therefore, it improves the model's ability to learn the distribution patterns of different categories of sample data subsets and enhances the overall recognition accuracy of each category, thereby improving the overall image processing accuracy. Attached Figure Description
[0010] To more clearly illustrate the technical solutions in this application, the accompanying drawings required in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort. Among them:
[0011] Figure 1 This is a flowchart illustrating an embodiment of the image processing model training method of this application;
[0012] Figure 2 This application Figure 1 A flowchart illustrating an embodiment of step S12;
[0013] Figure 3 This application Figure 1 A flowchart illustrating an embodiment of step S13;
[0014] Figure 4This is a schematic diagram of the structure of an embodiment of the image processing model training method of this application;
[0015] Figure 5 This is a schematic diagram of the structure of another embodiment of the training method of the image processing model of this application;
[0016] Figure 6 This is a schematic flowchart of an embodiment of the image processing method of this application;
[0017] Figure 7 This is a schematic diagram of the structure of an embodiment of the training device for the image processing model of this application;
[0018] Figure 8 This is a schematic diagram of the structure of an embodiment of the image processing apparatus of this application;
[0019] Figure 9 This is a schematic diagram of the structure of an embodiment of the computer device of this application;
[0020] Figure 10 This is a schematic diagram of the structure of an embodiment of the storage device of this application. Detailed Implementation
[0021] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of the embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of this application.
[0022] The terms "first" and "second" in this application are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. In the description of this application, "multiple" means at least two, such as two, three, etc., unless otherwise explicitly specified. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or apparatus that includes a series of steps or units is not limited to the listed steps or units, but may optionally include steps or units not listed, or may optionally include other steps or units inherent to such processes, methods, products, or apparatus.
[0023] In this application, the reference to "embodiment" means that a specific feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of this application. The appearance of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.
[0024] This application provides the following embodiments, and each embodiment is described in detail below.
[0025] Please see Figure 1 , Figure 1 This is a flowchart illustrating an embodiment of the image processing model training method of this application. The method may include the following steps:
[0026] S11: Obtain the sample dataset; the sample dataset contains K sample data subsets, and the different sample data subsets in the K sample data subsets contain different categories of sample data, and at least two sample data subsets in the K sample data subsets contain different numbers of sample data, where K is an integer greater than 1.
[0027] In some implementations, the acquired sample dataset may include K sample data subsets, each of which further includes multiple sample data, wherein the sample data in the sample data subset is a sample image containing text information, and K is an integer greater than 1.
[0028] In some implementations, the image processing model described above can be used for text recognition processing of images containing text information. Before using the image processing model, it needs to be trained using the sample data described above.
[0029] In some implementations, different subsets of the K sample data sets contain sample data of different categories; for example, they may be subsets of sample data with K categories. At least two subsets of the K sample data sets contain sample data of different numbers of samples; for example, at least two subsets of sample data with different numbers of samples from different categories.
[0030] In some implementations, the aforementioned sample data categories may include text categories, number categories, letter categories, etc. The text categories may include multiple language categories, such as Chinese, Korean, Japanese, etc. In some application scenarios, the categories may also include classifications of the font style of the characters, such as italic characters, artistic characters, deformed characters, blurred characters, etc.
[0031] In some implementations, categories can also be classified according to the frequency of use of the sample data, such as sample categories for high-frequency text and sample categories for low-frequency text. This application does not impose any restrictions on the sample categories of the sample dataset.
[0032] In some implementations, the number of samples in the subset of sample data corresponding to each category is different, that is, the distribution of sample data for each category is unbalanced. For example, in the field of natural scene text recognition, when training an image processing model, a few text categories contain a large number of samples, while most text categories contain a small number of samples.
[0033] In some implementations, the K subsets of sample data contained in the sample dataset can be long-tailed distributed data, meaning that the number of samples in the K categories follows a long-tailed distribution trend. Long-tailed distribution is a skewed distribution where a few categories (also called head classes) contain a large number of samples, while most categories (also called tail classes) have very few samples.
[0034] In some implementations, the sample dataset contains K subsets of sample data with progressively decreasing sample counts. The subset containing the category with the largest sample count or a sample count greater than a first threshold is called the head category, and the subset containing the category with the smallest sample count, or a sample count less than or equal to the first threshold, or a sample count less than a second threshold, is called the tail category. The sample count distribution of the subsets from the head category to the tail category is a decreasing distribution.
[0035] S12: Process the sample dataset to obtain M sample data sequences, where each sample data sequence is a subset of sample data ordered according to the distribution characteristics of the target sample category, and M is an integer greater than 1.
[0036] In some implementations, the image processing module can be trained in multiple rounds. In each round of training of the image processing model, the sample dataset can be sampled at least twice to obtain at least M sample data sequences.
[0037] In some implementations, sampling can be performed on at least two subsets of sample data from K subsets of sample data, such that the resulting sequence of sample data includes sample data from at least two subsets of sample data from different classes.
[0038] In some implementations, sampling can be performed on K subsets of sample data, such that the sampled M data sequences contain sample data from subsets of K categories. This application does not limit the number of sample data subsets used for sampling; it can be set according to the specific application scenario of the image processing model.
[0039] In some implementations, the sample data sequence consists of multiple subsets of sample data ordered according to the distribution characteristics of the target sample category, where M is an integer greater than 1. The target sample category can be the category to which the subsets of sample data included in the sample data sequence belong. The sorting of the target sample categories according to their characteristics can be a long-tailed distribution, such as the sorting from the sample data subsets of the head category to the sample data subsets of the tail category described above.
[0040] In some implementations, for example, when M is 2, sampling can be performed from at least two sample data subsets of K sample data subsets to obtain two sample data sequences with different distribution characteristics. For example, it can include sample data sequences where the number of samples in multiple sample data subsets ordered according to the distribution characteristics of the target sample category gradually increases or gradually decreases. If M is an integer greater than 2, in addition to the sample data sequences included when M is 2, it can also include sample data sequences where the number of samples in multiple sample data subsets ordered according to the distribution characteristics of the target sample category follows a preset distribution pattern. For example, the preset distribution pattern can be that the number of samples in even-numbered sample data subsets increases or gradually decreases sequentially, or the number of samples in odd-numbered sample data subsets increases or gradually decreases sequentially, etc., to obtain M sample data sequences. This application does not limit the method of obtaining the M sample data sequences.
[0041] In some implementations, the M sample data sequences include a first sample sequence and a second sample sequence, wherein the number of samples in each category contained in the first sample sequence and the second sample sequence is different.
[0042] In this context, among the multiple sample data subsets included in the first sample sequence, for example, in the K sample data subsets included in the first sequence, the number of samples in the sample data subset ranked at position i+1 is no greater than the number of samples in the sample data subset ranked at position i; i is greater than 0 and less than the number of subsets of the multiple sample data subsets included in the first sample sequence.
[0043] In the second sample sequence, for example, in the K sample data subsets included in the second sequence, the number of samples in the sample data subset ranked at position j+1 is not less than the number of samples in the sample data subset ranked at position j; j is greater than 0 and less than the number of subsets in the first sample sequence.
[0044] In some implementations, the sample data of the multiple subsets of sample data contained in the M sample data sequences can be distributed in a long-tailed manner. That is, the first sample sequence and the second sample sequence can be long-tailed data.
[0045] For example, the number of samples in the K subsets of the first sample sequence decreases sequentially, meaning the number of samples in the subsets from the first category to the last category decreases sequentially. The number of samples in the K subsets of the second sample sequence increases sequentially, meaning the number of samples in the subsets from the first category to the last category increases sequentially.
[0046] S13: Train the image processing model based on M sample data sequences to obtain the trained image processing model.
[0047] After obtaining M sample data sequences, these sequences can be used as training samples for an image processing model. The model is trained by inputting these sequences and then adjusting its parameters based on the processing results of the M sample data sequences to obtain a well-trained model that meets the preset training requirements. This allows the image processing model to be used for text recognition on images containing text.
[0048] In some implementations, step S12 described above can be used to perform multiple rounds of sampling on the K subsets of sample data contained in the sample dataset. The subsets of sample data for each category can be the same in each round, or they can be different in each round. Furthermore, the number of subsets of sample data in each round can be different or the same. This application does not limit the method of obtaining the M sample data sequences. By iteratively training the image processing model using multiple rounds of sample data sequences, a more accurate image processing model can be obtained.
[0049] In this embodiment, a sample dataset is obtained. The sample dataset contains K subsets of sample data, and different subsets of the K subsets contain sample data of different categories. At least two subsets of the K subsets contain sample data of different numbers of samples. The sample dataset is processed to obtain M sample data sequences. Since the sample data sequences are composed of multiple subsets of sample data sorted according to the distribution characteristics of the target sample category, the image processing model is trained based on the M sample data sequences to obtain the trained image processing model. This allows the image processing model to learn the distribution patterns of sample data subsets of various categories to a certain extent, while maintaining the representational ability of each category of sample data subsets in the image processing model. Therefore, it can improve the image processing model's ability to learn the distribution patterns of sample data subsets of various categories, and improve the overall recognition accuracy of the image processing model for each category, thereby improving the accuracy of image processing.
[0050] Furthermore, since the sample data of each subset of the sample dataset is unevenly distributed, the above method of training the image processing model can learn the different rules under uneven distribution conditions, and maintain the representational ability of each category of sample data in the image processing model. Overall, it improves the recognition accuracy of the image processing model for each category of sample data, and can improve the accuracy of the image processing model in text recognition when the number of samples of each category is unevenly distributed.
[0051] In some embodiments, please refer to Figure 2 This embodiment can further extend step S12 of the above embodiment. Processing the sample dataset to obtain M sample data sequences may include the following steps:
[0052] S121: From the K subsets of sample data in the sample dataset, select multiple subsets of sample data to determine multiple subsets of sample data for the target sample category.
[0053] Multiple subsets of sample data can be selected from the K subsets of sample data included in the sample dataset as the subsets of sample data for the target sample category, so that M sample data sequences are obtained from the multiple subsets of sample data for the target category. The target sample category can be selected based on the actual use case of the image processing model. For example, if the image processing model is used to recognize numbers, letters, characters, etc., a subset of sample data containing categories such as numbers, letters, characters, etc. can be selected as the subset of sample data for the target sample category. This application does not impose any restrictions on this.
[0054] In some implementations, K subsets of sample data can be used as multiple subsets of sample data for the target sample category, so that the training of the image processing model can learn the distinct rules of each category. This application does not limit this.
[0055] S122: Perform at least two different sampling processes on multiple subsets of sample data of the target sample category to obtain M sample data sequences.
[0056] When obtaining the first sample sequence, multiple subsets of sample data from the target sample category can be uniformly sampled to obtain a first sample sequence containing multiple subsets of sample data. The sampling probability corresponding to each subset of sample data for each category is proportional to the number of samples contained in that subset.
[0057] A uniform sampler can be used for sampling. Specifically, multiple subsets of sample data for the target sample category are uniformly sampled to obtain the first sample data corresponding to each category's subset. The first sample data corresponding to each category's subset is then used to form a first sample sequence. The sampling probability corresponding to each category's subset is proportional to the number of samples contained in the subset. For example, if the multiple subsets of sample data for the target sample category have a long-tailed distribution, the first sample sequence obtained after uniform sampling will also have a long-tailed distribution.
[0058] When obtaining the second sample sequence, a sample preset sampler can be used to sample multiple sample data subsets of the target sample category using a preset sampling method to obtain the second sample data corresponding to each category's sample data subset, and the second sample sequence can be formed using the second sample data corresponding to each category's sample data subset.
[0059] Specifically, the preset sampling probability corresponding to each subset of sample data in the target sample category is obtained; wherein, the preset sampling probability corresponding to each subset of sample data in each category is inversely proportional to the number of samples contained in the subset. The preset probability can be obtained by referring to the following method.
[0060] Get the ratio of the number of samples in the subset of sample data for each category to the maximum number of samples. The maximum number of samples is the maximum number of samples in the multiple subsets of sample data contained in the target sample category.
[0061] The ratio of the number of samples in each category of sample data subset to the maximum number of samples can be expressed as:
[0062]
[0063] In the above formula (1), Num max Num represents the maximum number of samples in multiple subsets of sample data contained in the target sample category, i.e., the maximum number of samples. j v represents the number of samples contained in the subset of sample data of category j. j This represents the ratio of the number of samples in the subset of sample data of category j to the maximum number of samples.
[0064] Then, by using the ratio of the number of samples in each category's sample data subset to the maximum number of samples and the total ratio, the preset sampling probability corresponding to each category's sample data subset is obtained, where the total ratio is the sum of the ratios corresponding to multiple sample data subsets included in the target sample category.
[0065] The preset sampling probability corresponding to the subset of sample data for each category can be expressed as:
[0066]
[0067] In the above formula (2), That is, Num represents the number of samples in the subset of sample data of category i. i With the maximum number of samples Num max The ratio. C represents the total number of categories contained in the target sample category. It represents the total ratio, which is the sum of the ratios corresponding to multiple subsets of sample data included in the target sample category.
[0068] In some implementations, multiple sample data can be selected from each sample data subset according to a preset sampling probability corresponding to each sample data subset, to obtain a second sample sequence containing multiple sample data subsets.
[0069] In some implementations, the preset sampling probability corresponding to the sample data subset of each category is inversely proportional to the number of samples in the multiple sample data subsets included in the target sample category, and the sample data of the multiple sample data subsets included in the sampled second sample sequence gradually increases.
[0070] In some implementations, after obtaining the first sample sequence or the second sample sequence using a uniform sampling method in step S122, each sample data extracted can be put back into its respective sample data subset, so that step S122 can be repeated multiple times with replacement to obtain M sample data sequences in multiple rounds.
[0071] By obtaining M sample data sequences in the above manner, multiple sample data sequences with different distribution patterns can be obtained, so that different patterns can be learned when training the image processing model in the future.
[0072] In some embodiments, please refer to Figure 3 Step S13 of the above embodiment can be further extended. Training the image processing model based on M sample data sequences to obtain the trained image processing model can include the following steps:
[0073] S131: Extract the features of each sample data in each subset of the M sample data sequences to obtain the M image feature sequences corresponding to the M sample data sequences.
[0074] Please see Figure 4 The image processing model 100 includes a feature extraction layer 101, a first processing module 102, a second processing module 103, and a fusion module 104.
[0075] The image processing model 100 may include two branches: one branch (such as the first processing module 102) processes a first sample sequence comprising M sample data sequences, and the other branch (such as the second processing module 103) processes a second sample sequence comprising M sample data sequences. The two branches may share a feature extraction layer 101 for extracting features from the M sample data sequences. The shared feature extraction layer 101 may share parameters, which can be the network parameters of the feature extraction layer; this application does not impose any restrictions on this.
[0076] In some embodiments, the first processing module 102 and the second processing module 103 may be the same or different network models. Specifically, the first processing module 102 and the second processing module 103 may include a feature sequence layer and a preset neural network layer. At least one of the feature sequence layer and the preset neural network layer in the two branches may be different. For example, the preset neural network layers in the two branches may be different. Specifically, the difference could be that the network model structure of the preset neural network layers in the two branches is the same but the parameters are different, or the network model structure is different and the parameters are different. This application does not impose any limitations on this.
[0077] In some embodiments, the fusion module 104 can be used to fuse the processing results of the first processing module 102 and the second processing module 103. The fusion module 104 may include an adaptive updater and a fusion layer. The adaptive updater generates weighting coefficients for the processing results of the first processing module 102 and the second processing module 103. These weighting coefficients can be user-defined or pre-set, and are not limited here. The fusion layer is used to fuse and output the processing results after weighting coefficient processing.
[0078] When training the image processing model 100, in this step, M sample data sequences can be input into the image processing model 100. Then, the feature extraction layer 101 of the image processing model 100 can be used to extract the features of each sample data in each sample data subset contained in the M sample data sequences, and the M image feature sequences corresponding to the M sample data sequences can be obtained respectively.
[0079] Among them, the M image feature sequences obtained include the first image feature sequence corresponding to the first sample sequence and the second image feature sequence corresponding to the second sample sequence.
[0080] S132: Perform image processing on the M image feature sequences respectively to obtain M result sequences respectively.
[0081] Please see Figure 5The first image feature sequence corresponding to the first sample sequence can be obtained by the uniform sampler and input into the first processing module of the image processing model. The first processing module of the image processing model performs image processing on the first image feature sequence to obtain the first result sequence.
[0082] In some implementations, the first processing model includes a feature sequence layer and a preset neural network layer. The first image feature sequence is input into the feature sequence layer to obtain a sorted or sequenced first image feature sequence, which corresponds to each sample data of the first sample sequence.
[0083] In some embodiments, the first image feature sequence is input into a preset neural network layer for processing to identify a first result sequence corresponding to the first sample sequence. The first result sequence is, for example, a character recognition result. The preset neural network layer can be used for character recognition, and is, for example, an LSTM (Long Short-Term Memory) network. Of course, the preset neural network layer can also be other neural networks, and this application does not limit this.
[0084] The second processing module of the image processing model performs image processing on the second image feature sequence to obtain the second result sequence corresponding to the second sample sequence. It is understood that the processing procedure of the second processing model for the second sample sequence can refer to the processing procedure of the first processing model for the first sample sequence, and will not be elaborated here.
[0085] S133: Fuse the M result sequences to obtain the fusion result and obtain the loss value corresponding to the fusion result.
[0086] Please see Figure 5 The weighted coefficients corresponding to the M result sequences can be obtained using the adaptive updater included in the fusion module. The weighted coefficients can include a first weighted coefficient corresponding to the first result sequence and a second weighted coefficient corresponding to the second result sequence, with the sum of the first and second weighted coefficients being a preset value. The first weighted coefficient is obtained using the number of iterations trained on the image processing model and the maximum and minimum sample sizes of the subsets of the M sample data sequences.
[0087] For example, if the preset value is 1, the first weighting coefficient can be expressed as:
[0088]
[0089] In the above formula (3), α represents the first weighting coefficient, and T represents the number of iterations for training the image processing model. It can also be expressed as the number of iterations for M sample data sequences. For example, inputting M sample data sequences into the image processing model for training in one round can be considered as one iteration.max This represents the maximum number of iterations, which is the iteration threshold set for training the image processing model. β represents a hyperparameter, which is the ratio of the maximum to the minimum number of samples in a subset of M sample data sequences.
[0090] In some implementations, a second weighting coefficient can be obtained using a first weighting coefficient and a preset value. For example, when the preset value is 1, the second weighting coefficient can be expressed as: 1-α.
[0091] After obtaining the weighting coefficients, the fusion layer of the fusion module can be used to weight each result sequence using the weighting coefficients corresponding to each result sequence, resulting in M weighted results. Specifically, the first weighted result can be obtained by weighting the first result sequence using the first weighting coefficients corresponding to the first result sequence; and the second weighted result can be obtained by weighting the second result sequence using the second weighting coefficients corresponding to the second result sequence.
[0092] Therefore, the fusion layer of the fusion module can be used to fuse the weighted results corresponding to the M result sequences. Specifically, the first weighted result and the second weighted result can be summed proportionally, or the first weighted result and the second weighted result can be summed to obtain the fusion result.
[0093] In some implementations, after obtaining the fusion result, the corresponding loss value is acquired. A loss function can be used to obtain the loss between the fusion result and a preset result, where the loss function can be expressed as:
[0094] L(S)=-ln∏ (x,z)∈S p(z|x) Formula (4)
[0095] In formula (4) above, L(S) represents the loss function, and S represents the set of sample data in the M sample data sequences (which can be the set of sample data containing text labels, i.e., the set of sample data containing prediction results). x represents the sample data in the input M sample data sequences. z represents the preset result corresponding to the sample data in the M sample data sequences, i.e., the text labels contained therein.
[0096] In some implementations, the loss function L(S) may include the CTC (Connectionist Temporal Classification loss) loss function. The CTC loss function can solve the problem of input and output alignment, avoiding the need to label each character in the sample data sequence, and only needing to label the sample data line by line.
[0097] S134: Adjust the parameters of the image processing model using the loss value to obtain the trained image processing model.
[0098] During the training of the image processing model, special characters, such as whitespace, can be inserted between repeated characters when encoding the fused results of the sample data sequence. The Adam algorithm (Adaptive Momentum) is used to continuously adjust the parameters in the image processing model, such as the weights and biases, so that the accuracy of the image processing model reaches a preset accuracy, and / or the loss function value is less than a preset loss value. Specifically, the smaller the loss value of the CTC loss function and the higher the probability of all sequences matching the target, the closer the text recognition result of the image processing model is to the preset result, that is, the closer the predicted text sequence is to the real text sequence, and the higher the accuracy. During decoding, the most likely character is selected at each time step to calculate the optimal path, repeated characters are removed, and then all special characters are removed from the path. The remaining characters are taken as the recognized text, i.e., the predicted result.
[0099] In some implementations, the parameters of the image processing model are adjusted based on the loss value. That is, the parameters of the first processing model, the second processing module, etc., can be adjusted until an image processing model that meets the preset training requirements is obtained. It is understood that the parameters of the first processing module and the second processing module after parameter adjustment, i.e., after training, can be different.
[0100] In this embodiment, by training the image processing model using M sample data sequences, the image processing model can focus on the sample data of each category's sample data subset. That is, it can focus on both high-frequency characters in the sample data and low-frequency characters in the sample data, thereby improving the overall accuracy of the image processing model in recognizing characters in images.
[0101] Please see Figure 6 , Figure 6 This is a schematic flowchart of an embodiment of the image processing method of this application. The method may include the following steps:
[0102] S21: Obtain the image to be processed, which is an image containing text information.
[0103] The image to be processed can be an image that requires text recognition, containing text information such as letters, numbers, and characters.
[0104] S22: The image processing model trained using the above image processing model training method processes the image to be processed and obtains the processing result of the image to be processed.
[0105] The image processing model trained using the above-described image processing model training method includes a feature extraction layer, a first processing module, a second processing module, and a fusion module. The specific training process of the image processing model can be referred to the specific implementation process of the above embodiments, which will not be elaborated here.
[0106] The trained image processing model is used to perform text recognition on the image to be processed, and the processing result of the text information contained in the image to be processed is obtained, that is, the recognition result of the text information.
[0107] Since the image processing model trained using the above-mentioned image processing model training method can improve the accuracy of text recognition in the image to be processed, the image processing model can perform text recognition on the image to be processed.
[0108] In accordance with the above embodiments, this application provides a training apparatus for an image processing model. Please refer to... Figure 7 , Figure 7 This is a schematic diagram of the structure of an embodiment of the training device for the image processing model of this application. The training device 30 for the image processing model may include a sample module 31, a processing module 32, and a training module 33.
[0109] The sample module 31 is used to obtain the sample dataset; the sample dataset contains K sample data subsets, and the different sample data subsets in the K sample data subsets contain different categories of sample data, and at least two sample data subsets in the K sample data subsets contain different numbers of sample data, where K is an integer greater than 1.
[0110] The processing module 32 is used to process the sample dataset to obtain M sample data sequences, wherein each sample data sequence is composed of multiple sample data subsets sorted according to the distribution characteristics of the target sample category, and M is an integer greater than 1.
[0111] The training module 33 is used to train the image processing model based on M sample data sequences to obtain the trained image processing model.
[0112] The specific implementation of this embodiment can be referred to the implementation process of the above embodiments, and will not be repeated here.
[0113] In accordance with the above embodiments, this application provides an image processing apparatus. Please refer to... Figure 8 , Figure 8 This is a schematic diagram of an embodiment of the image processing apparatus of this application. The character recognition apparatus 40 may include an image module 41 and an image processing module 42.
[0114] The image module 41 is used to acquire the image to be processed, which is an image containing text information.
[0115] The image processing module 42 is used to process the image to be processed by the image processing model trained by the above image processing model training method, and obtain the processing result of the image to be processed.
[0116] The specific implementation of this embodiment can be referred to the implementation process of the above embodiments, and will not be repeated here.
[0117] Regarding the above embodiments, this application provides a computer device; please refer to [link / reference]. Figure 9 , Figure 9 This is a schematic diagram of the structure of a computer device according to an embodiment of the present application. The computer device 50 includes a memory 51 and a processor 52, wherein the memory 51 and the processor 52 are coupled to each other. The memory 51 stores program data, and the processor 52 is used to execute the program data to implement the steps of the above-described image processing model training method and any embodiment of the image processing method.
[0118] In this embodiment, processor 52 can also be referred to as CPU (Central Processing Unit). Processor 52 may be an integrated circuit chip with signal processing capabilities. Processor 52 can also be a general-purpose processor, digital signal processor (DSP), application-specific integrated circuit (ASIC), field-programmable gate array (FPGA), or other programmable logic device, discrete gate or transistor logic device, or discrete hardware component. The general-purpose processor can be a microprocessor, or processor 52 can be any conventional processor.
[0119] The specific implementation of this embodiment can be referred to the implementation process of the above embodiments, and will not be repeated here.
[0120] The methods described in the above embodiments can be implemented as computer programs; therefore, this application proposes a storage device. Please refer to [link to relevant documentation]. Figure 10 , Figure 10 This is a schematic diagram of a storage device according to an embodiment of the present application. The storage device 60 stores program data 61 that can be executed by a processor. The program data 61 can be executed by the processor to implement the steps of any embodiment of the above-described image processing model training method and image processing method.
[0121] In this embodiment, the storage device 60 can be a medium that can store program data 61, such as a USB flash drive, a portable hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk. Alternatively, it can be a server that stores the program data 61. The server can send the stored program data 61 to other devices for execution, or it can run the stored program data 61 itself.
[0122] The specific implementation of this embodiment can be referred to the implementation process of the above embodiments, and will not be repeated here.
[0123] In the several embodiments provided in this application, it should be understood that the disclosed methods and apparatus can be implemented in other ways. For example, the apparatus implementations described above are merely illustrative. For instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms.
[0124] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment, depending on actual needs.
[0125] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0126] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a storage device, which is a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause an electronic device (which may be a personal computer, server, or network device, etc.) or processor to execute all or part of the steps of the methods of the various embodiments of this application.
[0127] Obviously, those skilled in the art should understand that the modules or steps of this application described above can be implemented using general-purpose computing devices. They can be centralized on a single computing device or distributed across a network of multiple computing devices. Optionally, they can be implemented using computer-executable program code, thereby storing them in a storage device for execution by a computing device, or fabricating them separately as individual integrated circuit modules, or fabricating multiple modules or steps as a single integrated circuit module. Thus, this application is not limited to any particular combination of hardware and software.
[0128] The above description is merely an embodiment of this application and does not limit the patent scope of this application. Any equivalent structural or procedural transformations made using the content of this application's specification and drawings, or direct or indirect applications in other related technical fields, are similarly included within the patent protection scope of this application.
Claims
1. A method for training an image processing model, characterized in that, The method includes: Obtain a sample dataset; the sample dataset contains K sample data subsets, and the different sample data subsets in the K sample data subsets contain sample data of different categories, and at least two sample data subsets in the K sample data subsets contain sample data of different numbers, where K is an integer greater than 1; the K sample data subsets in the sample dataset are long-tailed distributed data; wherein, the sample data in the sample dataset are sample images containing text information; the sample data categories are classified according to the usage frequency of the sample data; The sample dataset is processed to obtain M sample data sequences, wherein each sample data sequence consists of multiple sample data subsets sorted according to the target sample category distribution characteristics, where M is an integer greater than 1. The M sample data sequences include sample data sequences where the number of samples in the multiple sample data subsets sorted according to the target sample category distribution characteristics gradually increases and gradually decreases, including: From the K subsets of the sample dataset, multiple subsets of sample data are selected to determine multiple subsets of sample data for the target sample category; wherein, the target sample category is selected based on the actual use scenario of the image processing model; At least two different sampling processes are performed on multiple subsets of sample data of the target sample category to obtain the M sample data sequences; The image processing model is trained based on the M sample data sequences to obtain the trained image processing model. This image processing model is used to perform text recognition processing on images containing text information, including: Extract the features of each sample data in each subset of the M sample data sequences to obtain the M image feature sequences corresponding to the M sample data sequences; Image processing is performed on the M image feature sequences respectively to obtain M result sequences; The weighting coefficients corresponding to the M result sequences are obtained using an adaptive updater; Each result sequence is weighted using a weighting coefficient corresponding to each result sequence to obtain a weighted result for the M result sequences. The M result sequences include a first result sequence corresponding to a first sample sequence and a second result sequence corresponding to a second sample sequence. The weighting coefficients include a first weighting coefficient corresponding to the first result sequence and a second weighting coefficient corresponding to the second result sequence. The first weighting coefficient is obtained using the number of iterations trained by the image processing model, the maximum number of samples in the subset of the M sample data sequences, and a hyperparameter, where the hyperparameter is the ratio of the maximum number of samples to the minimum number of samples. The sum of the first weighting coefficient and the second weighting coefficient is a preset value. The weighted results corresponding to the M result sequences are fused to obtain the fused result, and the loss value corresponding to the fused result is obtained; The parameters of the image processing model are adjusted using the loss value to obtain the trained image processing model.
2. The method of claim 1, wherein, The M sample data sequences include a first sample sequence and a second sample sequence, wherein: In the first sample sequence, the number of samples in the sample data subset ranked at position i+1 is no greater than the number of samples in the sample data subset ranked at position i; where i is greater than 0 and less than the number of subsets in the first sample sequence. In the second sample sequence, the number of samples in the sample data subset ranked at position j+1 is not less than the number of samples in the sample data subset ranked at position j; where j is greater than 0 and less than the number of subsets in the first sample sequence.
3. The method of claim 1, wherein, The M sample data sequences include a first sample sequence, and the at least two different sampling processes include: Uniform sampling is performed on multiple subsets of sample data of the target sample category to obtain a first sample sequence containing the multiple subsets of sample data; The sampling probability corresponding to each category of sample data subset is proportional to the number of samples contained in the sample data subset.
4. The method of claim 1, wherein, The M sample data sequences include a second sample sequence, and the at least two different sampling processes include: Obtain the preset sampling probability corresponding to each subset of sample data in the target sample category; wherein, the preset sampling probability corresponding to each subset of sample data in the category is inversely proportional to the number of samples in the subset of sample data; According to the preset sampling probability corresponding to each subset of sample data, multiple sample data are selected from each subset of sample data to obtain a second sample sequence containing the multiple subsets of sample data.
5. The method of claim 4, wherein, The step of obtaining the preset sampling probability corresponding to each subset of sample data in the target sample category includes: Obtain the ratio of the number of samples in the subset of sample data for each category to the maximum number of samples, wherein the maximum number of samples is the maximum number of samples in multiple subsets of sample data included in the target sample category; By using the ratio of the number of samples in each subset of sample data for each category to the maximum number of samples and the total ratio, a preset sampling probability corresponding to each subset of sample data for each category is obtained, wherein the total ratio is the sum of the ratios corresponding to multiple subsets of sample data included in the target sample category.
6. The method of claim 1, wherein, The step of weighting each result sequence using the weighting coefficients corresponding to each result sequence includes: The first result sequence is weighted using the first weighting coefficient corresponding to the first result sequence; and The second result sequence is weighted using the second weighting coefficient corresponding to the second result sequence.
7. The method of claim 1, wherein, The M image feature sequences include a first image feature sequence corresponding to a first sample sequence and a second image feature sequence corresponding to a second sample sequence; the M result sequences include a first result sequence corresponding to a first sample sequence and a second result sequence corresponding to a second sample sequence. The step of performing image processing on the M image feature sequences respectively to obtain M result sequences includes: The first image feature sequence is processed using the first processing module of the image processing model to obtain the first result sequence. as well as The second image feature sequence is processed using the second processing module of the image processing model to obtain the second result sequence.
8. An image processing method characterized by, The method includes: Obtain the image to be processed, which is an image containing text information; The image to be processed is processed using the image processing model included in any one of claims 1 to 7 to obtain the processing result of the image to be processed.
9. A computer device, comprising: The method includes a memory and a processor coupled to each other, the memory storing program data and the processor executing the program data to implement the steps of the method according to any one of claims 1 to 8.
10. A memory device, comprising: The system stores program data that can be executed by a processor, the program data being used to implement the steps of the method according to any one of claims 1 to 8.
Citation Information
Patent Citations
Continuous learning method and device, terminal and storage medium
CN112990318A