Method and apparatus for selecting sample of textual image, and computer device

By performing different initialization training on multiple models and calculating the difficulty scores of text image prediction results, selecting text images with high difficulty scores as samples, solving the problem of insufficient data representation in the prior art and improving the performance and accuracy of the character recognition model.

WO2025129681A1PCT designated stage expired Publication Date: 2025-06-26SIEMENS AG +1
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2023/141216
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-12-22
Publication Date
2025-06-26

AI Technical Summary

Technical Problem

In the prior art, when establishing a high-precision character recognition model, a large amount of data is required to adjust, and random sampling may lead to insufficient data representation and affect model performance.

Method used

By performing different initialization training on multiple models and inputting text pictures to multiple models, the prediction results of text pictures of multiple models are obtained. Calculate the difficulty scores of each text position and combine these scores to obtain the difficulty scores of the text picture. Select the text pictures with high difficulty scores as samples.

Benefits of technology

Mining out representative sample subsets reduces annotation costs and improves the model's identification accuracy of text images.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2023141216_26062025_PF_FP_ABST
    Figure CN2023141216_26062025_PF_FP_ABST
Patent Text Reader

Abstract

The present application discloses a method and apparatus for selecting a sample of a textual image, a computer device, and a storage medium. Specifically, the present application discloses a method for selecting a sample of a textual image, comprising performing different initialization training on a plurality of models; inputting textual images into the plurality of models, and obtaining prediction results of the plurality of models for a plurality of text positions of the textual images; for the prediction results for the plurality of text positions, obtaining a difficulty score of each text position; integrating the difficulty scores of the plurality of text positions, and obtaining difficulty scores of the textual images; and, on the basis of the difficulty scores, select as a sample a textual image that meets a preset condition. By means of the described means, a sample of a textual image that is difficult for a model to determine can be identified, thus specific manual labeling can be carried out, and thus the working process of large-scale manual labeling can be reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Text and image sample selection method, device, and computer equipment Technical Field

[0001] The present application relates to the field of images, and in particular, to a sample selection method, apparatus, computer equipment, and storage medium for text images. Background Art

[0002] In existing technologies, to build a highly accurate character recognition model for a specific use case, it is crucial and time-consuming to tune the target domain using large amounts of data. As deployment cases expand, it becomes impossible to provide annotations for tons of data points.

[0003] Random sampling is the simplest way to reduce the workload of data annotation. However, if the representativeness of the data is not taken into account and only a reduced dataset is used for training, the performance of the data-driven model will undoubtedly be poor. On the other hand, using the entropy output of the probability distribution function (such as the entropy output of the softmax) or the convolutional layer is not suitable for character recognition because these methods cannot handle continuous prediction tasks.

[0004] Summary of the Invention

[0005] This summary is provided to introduce some selected concepts in a simplified form, which will be further described in the detailed description below. This summary is not intended to identify any key features or essential features of the claimed subject matter, nor is it intended to be used to help determine the scope of the claimed subject matter.

[0006] Based on this, the present application discloses a method for selecting sample text images, wherein:

[0007] Train multiple models with different initializations;

[0008] Inputting the text image into multiple models to obtain prediction results of the multiple models for multiple text positions in the text image;

[0009] For the prediction results of the multiple text positions, obtain a difficulty score for each of the text positions; and combine the difficulty scores of the multiple text positions to obtain a difficulty score for the text image;

[0010] According to the difficulty score, text images with high difficulty scores are selected as samples.

[0011] Through the above method, a representative sample subset can be mined, which is conducive to later annotation and training of other models, as well as labeling of other images.

[0012] Furthermore, the text image is input into multiple models to obtain prediction results of the multiple models for the text image, including:

[0013] Calculate the sequence features of the image through convolution and downsampling of each model;

[0014] According to the sequence features, a prediction vector is obtained through a recurrent neural network layer;

[0015] According to the prediction vector, a maximum independent variable point set function is used to obtain a prediction result of each model on each character position in the character image.

[0016] In this way, prediction results of different models for the same image can be obtained, which can be used for subsequent selection of sample subsets.

[0017] Furthermore, different initialization training is performed on multiple models, including:

[0018] Initialize weights differently for multiple models;

[0019] Set random numbers with different Gaussian distributions for multiple models;

[0020] Train multiple models using different training samples;

[0021] Use different neural network structures for multiple models.

[0022] By using the above method, different models can be trained. These different models can then be used to generate corresponding prediction results for the same text image. The prediction results can be understood as the results obtained from different models, and the corresponding prediction results have reference and comparison significance.

[0023] Furthermore, according to the difficulty score, selecting text images with high difficulty scores as samples includes:

[0024] According to a preset difficulty score threshold, text images with difficulty scores exceeding the threshold are selected as samples.

[0025] Through the above method, when the difficulty score exceeds a certain threshold, it means that multiple models have different prediction results for the image. In this case, this text image is more meaningful for manual labeling, so that the results can be further fed back to the model for training, thereby improving the model's recognition accuracy for text images.

[0026] Furthermore, the present application discloses a device for selecting text and image samples, wherein:

[0027] The initialization module is used to perform different initialization training on multiple models;

[0028] A prediction module, configured to input a text image into a plurality of models and obtain prediction results of the plurality of models for a plurality of text positions in the text image;

[0029] a score module, configured to obtain a difficulty score for each of the plurality of text positions based on the prediction results; and to obtain a difficulty score for the text image by combining the difficulty scores of the plurality of text positions;

[0030] The selection module is used to select text images that meet preset conditions as samples based on the difficulty score.

[0031] The present application also provides a computer device, comprising a memory and a processor, wherein the memory stores a computer program, and wherein the processor implements the above method when executing the computer program.

[0032] The present application also provides a computer-readable storage medium having a computer program stored thereon, and the computer program implements the above method when executed by a processor.

[0033] The present application also provides a computer program product, which is tangibly stored on a computer-readable medium and includes computer-executable instructions. When the computer-executable instructions are executed, the computer-executable instructions cause at least one processor to perform the method described above. BRIEF DESCRIPTION OF THE DRAWINGS

[0034] Implementations of the present disclosure are illustrated by way of example and not limitation in the figures of the accompanying drawings in which like references indicate the same or similar parts.

[0035] FIG1 is a schematic diagram of a process of a sample selection method for text images according to an embodiment of the present application.

[0036] FIG2 is a schematic diagram of a sample selection device for text images according to an embodiment of the present application.

[0037] FIG3 is a schematic diagram of a computer device for selecting text image samples according to an embodiment of the present application.

[0038] FIG4 is a schematic diagram of an example of a sample selection method for text images according to an embodiment of the present application.

[0039] The reference numerals are as follows: S101-S104 Step 200: Apparatus 201: Module 202: Module 203: Module 204: Module 300: Computer device 302: Processor 304: Memory DETAILED DESCRIPTION

[0040] In the following description, for the purpose of explanation, a large number of specific details are set forth. However, it is understood that the present invention can be implemented without these specific details. In other examples, well-known circuits, structures, and technologies are not shown in detail so as not to affect the understanding of the description.

[0041] References throughout this specification to "an implementation," "an implementation," "an exemplary implementation," "some implementations," "various implementations," etc., indicate that the implementations of the invention being described may include particular features, structures, or characteristics. However, it does not imply that every implementation must include those particular features, structures, or characteristics. Furthermore, some implementations may have some, all, or none of the features described for other implementations.

[0042] Therefore, this application proposes a solution to mine representative subsets instead of using the entire dataset, which is better than random sampling.

[0043] This application discloses a method for selecting sample text images, wherein:

[0044] S101, perform different initialization training on multiple models.

[0045] Specifically, multiple different models can be used for initialization training, and the multiple models can be multiple neural networks; the different initialization training of multiple models is mainly to enable different models to produce independent prediction results for the processing results of the same text image when recognizing the same text image later, so as to comprehensively judge whether the text content of the text image is easy to be recognized or labeled by the model.

[0046] Furthermore, different initialization training is performed on multiple models, including:

[0047] Initialize weights differently for multiple models;

[0048] Set random numbers with different Gaussian distributions for multiple models;

[0049] Train multiple models using different training samples;

[0050] Use different neural network structures for multiple models.

[0051] By using the above method, different models can be trained. These different models can then be used to generate corresponding prediction results for the same text image. The prediction results can be understood as the results obtained from different models, and the corresponding prediction results have reference and comparison significance.

[0052] S102: Input the text image into multiple models to obtain prediction results of the multiple models for multiple text positions in the text image.

[0053] Specifically, a text image is input into the above-mentioned multiple differently initialized models, in the hope of obtaining prediction results of each model for different text positions of the same text image. Specifically, there may be multiple text positions on the text image. Through the recognition of the text image by the model, on the one hand, the text position must be recognized, and on the other hand, the text at the text position must be recognized, so as to obtain the corresponding text recognition result. The prediction result is the result of each model predicting the text on the text image. In some embodiments, each model can recognize the same text position for the text image. In some embodiments, each model recognizes a different number of text positions for the same text image. Furthermore, at the determined text position, specific text is further identified, which may be numbers, Chinese characters, English letters, symbols, etc. The prediction result can be used subsequently to analyze and determine whether the text image is difficult or challenging for model recognition and labeling.

[0054] Furthermore, the text image is input into multiple models to obtain prediction results of the multiple models for multiple text positions in the text image, including:

[0055] Calculate the sequence features of the image through convolution and downsampling of each model;

[0056] According to the sequence features, a prediction vector is obtained through a recurrent neural network layer;

[0057] According to the prediction vector, a maximum independent variable point set function is used to obtain a prediction result of each model on each character position in the character image.

[0058] Specifically, for example, a 32*200 pixel image, after convolution and downsampling by the model, obtains the sequence features of the image, and the sequence features can be a 1*1024*200 matrix.

[0059] Secondly, for the sequence features, a prediction vector is obtained through a recursive neural network layer, and the prediction vector is a 1*1000*200 matrix.

[0060] Again, the prediction result of the text image is obtained by using the maximum independent variable point set function for the prediction vector. The prediction result can be a 1*1*200 matrix.

[0061] 200 can be represented by L, where L represents the number of text positions. In some embodiments, the number of text positions recognized by multiple models for the same text image is 200.

[0062] S103 , for the prediction results of the multiple text positions, obtaining a difficulty score for each of the text positions; and combining the difficulty scores of the multiple text positions to obtain a difficulty score for the text image.

[0063] Specifically, the prediction results of the multiple models are compared, and the difficulty score of the text image is obtained by aggregation. The difficulty score reflects the consistency of the prediction results between the multiple models when recognizing the same text image. If the prediction results between the models are more consistent, it means that the different models are more accurate in recognizing the text image. Conversely, the success rate of recognition of the text image for different models is relatively high, and the corresponding difficulty score will be lower. If the prediction results between the models are more inconsistent, it means that different models have different recognition results for the text image. Conversely, it can be said that the probability of successful recognition of the text image for different models is not high, and the difficulty score will be higher.

[0064] Specifically, use the formula to express:

[0065] in This represents the sum of the number of times the maximum index number appears at the i-th character position across all models. Alternatively, it can be understood as the sum of the number of times multiple models produce the same prediction for character position i. The maximum index value represents the model's prediction for that character position.

[0066] In some embodiments, The maximum value of is N, where N represents the number of models. In some embodiments, The minimum value is 0.

[0067] Then Indicates the average number of times each model predicts the result. Can be expressed as a difficulty score.

[0068] As mentioned above, if is N, which means that multiple models have produced the same prediction results for the text position. If the corresponding value is 0, it means that the difficulty score of the text position is 0, because all models produce the same prediction result for the text position.

[0069] On the contrary, if is 0, indicating that multiple models have produced different prediction results for the text position. If the corresponding value is 1, it means that the difficulty score of the position is 1 (which can also be understood as the maximum difficulty score).

[0070] In some embodiments, i is summed, where i ranges from 0 to L. That is, the difficulty score of each text position is summed to obtain the sum of the difficulty scores of all text positions in the entire text image, which can be represented by dp.

[0071] Furthermore, the formula used is as follows:

[0072] In the above formula, dp divided by L represents the average difficulty score of each text position in the text image, represented by d.

[0073] In some embodiments, for the above prediction results, a 1*1*200 matrix is ​​calculated to obtain a 1*200 matrix, which is then aggregated to obtain a final one-dimensional difficulty score.

[0074] S104: Select text images that meet preset conditions as samples based on the difficulty score.

[0075] Specifically, the preset condition may be that pictures with higher difficulty scores need to be screened out as samples, and more accurate annotations are given later through manual or other methods.

[0076] Furthermore, according to the difficulty score, selecting text images that meet preset conditions as samples includes: according to a preset difficulty score threshold, selecting text images whose difficulty scores exceed the threshold as samples.

[0077] Through the above method, when the difficulty score exceeds a certain threshold, it means that multiple models have different prediction results for the image. In this case, this text image is more meaningful for manual labeling, so that the results can be further fed back to the model for training, thereby improving the model's recognition accuracy for text images.

[0078] Furthermore, in order to solve the above-mentioned problem, in this embodiment of the present application, as shown in FIG4 , a hard instance selection method based on query-by-query is adopted to perform sequence recognition.

[0079] First, this application trains N basic convolutional recurrent neural network (CRNN) models with different performance on a basic dataset, from #1 to #n, as shown in Figure 4. In order to achieve the differences between multiple basic models, different training strategies can be adopted, such as starting from different initialization weights, using subsets of training samples, using different backbones, etc.

[0080] Then, this application uses the basic CRNN model as a group to measure the differences in their prediction results for the same text image. For each image sample in the sample pool, the sequence features of the text image are extracted through convolution and downsampling operations. The flattened sequence features of length L are then sent to the neural network recursive layer. The processing process in the recursive layer can be regarded as using a sliding window to scan from the first element to the last element in sequence, and is based on a comprehensive judgment and prediction of context relevance. Therefore, the output of the RNN layer at each position is the prediction vector of the element corresponding to the input feature. Here, this application compares the difference in the output probability of each text position (obtained using the softmax function) and uses it as the acquisition function. Specifically, this application measures the most common frequency of each time period in the sequence. The greater the frequency, the more consistent the prediction results between the models, which means that the difficulty information of the position is less. For sequential recognition tasks such as character recognition, this application will obtain sequential outputs from the RNN, so the final difficulty information or difficulty score d is calculated by performing an aggregation operation on the entire sequence. For the sake of intuitiveness, colored rectangles are used in Figure 4 to represent the difficulty score of each position. The lighter the color, such as red, the higher the difficulty score. The text recognition at this location is relatively inconsistent for each model.

[0081] Because our method uses a pool-based selection criterion, we extract features from each sample in the entire pool and calculate an aggregated difficulty score. We then select a subset from the pool based on the ranking results. The size of the selected sample depends on requirements or annotation costs.

[0082] The technical features and beneficial effects of this application are as follows:

[0083] Compared with random selection or other hard example mining methods designed for general classification tasks, the character recognition model trained using the subset selected by the method of this application will have better performance.

[0084] The proposed method can be used to reduce the annotation cost in model tuning work.

[0085] Furthermore, the use of a sample selection method to process the upgrade of a character recognition model can be seen as a possible implementation of the present application solution.

[0086] It should be understood that although the various steps in the flowchart of FIG1 are displayed in sequence as indicated by the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order restriction on the execution of these steps, and these steps can be executed in other orders. Moreover, at least a portion of the steps in FIG1 may include multiple steps or multiple stages, and these steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be executed in turn or alternately with other steps or at least a portion of the steps or stages in other steps.

[0087] FIG2 provides a device 200 for selecting sample text images. The device 200 includes:

[0088] Initialization module 201, used to perform different initialization training on multiple models;

[0089] Prediction module 202, inputs the text image into multiple models, and obtains prediction results of the multiple models for multiple text positions in the text image;

[0090] The score module 203 is configured to obtain a difficulty score for each of the plurality of text positions based on the prediction results; and to obtain a difficulty score for the text image by combining the difficulty scores of the plurality of text positions;

[0091] The selection module 204 is configured to select text images that meet preset conditions as samples based on the difficulty score.

[0092] Furthermore, the prediction module 202 is further configured to:

[0093] Calculate the sequence features of the image through convolution and downsampling of each model;

[0094] According to the sequence features, a prediction vector is obtained through a recurrent neural network layer;

[0095] According to the prediction vector, a maximum independent variable point set function is used to obtain a prediction result of each model on each character position in the character image.

[0096] Furthermore, the initial module 201 is further configured to:

[0097] Initialize weights differently for multiple models;

[0098] Set random numbers with different Gaussian distributions for multiple models;

[0099] Train multiple models using different training samples;

[0100] Use different neural network structures for multiple models.

[0101] Furthermore, the selection module 204 includes:

[0102] According to a preset difficulty score threshold, text images with difficulty scores exceeding the threshold are selected as samples.

[0103] It should be noted that the apparatus may include more or fewer modules to implement the described functionality. For example, at least one module in FIG. 2 may be further divided into a plurality of different submodules, each of which is configured to perform at least a portion of the operations described herein in conjunction with the corresponding module. Furthermore, in some examples, the apparatus 200 may further include additional modules for performing other operations already described in the specification. Furthermore, those skilled in the art will appreciate that the exemplary apparatus 200 may be implemented using software, hardware, firmware, or any combination thereof.

[0104] Figure 3 provides a computer device. According to one embodiment, the computer device 300 may include a processor 302, which executes a computer program stored in a memory 304. When the computer program is executed by the processor, the above method is implemented.

[0105] Those skilled in the art will understand that the structure shown in Figure 3 is merely a block diagram of a portion of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the present application is applied. The specific computer device may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.

[0106] Those skilled in the art will appreciate that all or part of the processes in the methods for implementing the above-mentioned embodiments can be accomplished by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include processes for the implementation of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the various embodiments provided in this application may include at least one of non-volatile and volatile memory. Non-volatile memory may include read-only memory (ROM), magnetic tape, floppy disk, flash memory or optical memory, etc. Volatile memory may include random access memory (RAM) or external cache memory. As an illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM).

[0107] The present application also provides a computer-readable storage medium having a computer program stored thereon, which implements the above steps when the computer program is executed by a processor.

[0108] The present application also provides a computer program product, which is tangibly stored on a computer-readable medium and includes computer-executable instructions. When the computer-executable instructions are executed, the computer-executable instructions cause at least one processor to perform the above method.

[0109] Furthermore, the computer program can be stored and run in the cloud to perform the method. Furthermore, the components of the program can be deployed on multiple devices and the cloud. For example, the corresponding steps can be deployed and run on a local or local computer, or run on different cloud devices, and transmit signals through a communication connection, or can also be deployed and run on a local or local computer. This application does not limit the manner or method, and the corresponding technology can be flexibly deployed to make full use of equipment and technologies such as the cloud, big data, and supercomputing capabilities to execute and complete the method.

[0110] Some implementations of the present disclosure may include articles of manufacture. Articles of manufacture may include storage media for storing logic. Examples of storage media may include one or more types of computer-readable storage media capable of storing electronic data, including volatile or non-volatile memory, removable or non-removable memory, erasable or non-erasable memory, writable or rewritable memory, and the like. Examples of logic may include various software units, such as software components, programs, applications, computer programs, application programs, system programs, machine programs, operating system software, middleware, firmware, software modules, routines, subroutines, functions, methods, processes, software interfaces, application program interfaces (APIs), instruction sets, computing codes, computer codes, code segments, computer code segments, words, values, symbols, or any combination thereof. In some implementations, for example, articles of manufacture may store executable computer program instructions that, when executed by a processor, cause the processor to perform the methods and / or operations described herein. Executable computer program instructions may include any suitable type of code, such as source code, compiled code, interpreted code, executable code, static code, dynamic code, and the like. Executable computer program instructions can be implemented according to a predefined computer language, method or syntax for commanding a computer to perform a specific function. The instructions can be implemented using any appropriate high-level, low-level, object-oriented, visual, compiled and / or interpreted programming language.

[0111] What has been described above includes examples of the disclosed architecture. It is, of course, not possible to describe every conceivable combination of components and / or methodologies, but those skilled in the art will appreciate that many other combinations and permutations are possible. Therefore, the novel architecture is intended to embrace all such alternatives, modifications, and variations that fall within the spirit and scope of the appended claims.

Claims

1. A method for sample selection of text images, wherein, Perform different initial training on multiple models; Input the text images into multiple models to obtain prediction results of multiple models for multiple text positions of the text images; For the prediction results of the multiple text positions, obtain the difficulty scores of each text position; And synthesize the difficulty scores of the multiple text positions to obtain the difficulty score of the text image; According to the difficulty score, select text images meeting preset conditions as samples.

2. The method according to claim 1, wherein Inputting the text images into multiple models to obtain prediction results of multiple models for multiple text positions of the text images includes: Calculate the sequence features of the text images through convolution and downsampling of each model; According to the sequence features, obtain prediction vectors through a recurrent neural network layer; According to the prediction vectors, use the maximum independent variable point set function to obtain the prediction results of each model for each text position of the text image.

3. The method according to claim 1, wherein Performing different initial training on multiple models includes: Perform different initial weights on multiple models; Set random numbers with different Gaussian distributions for multiple models; Train multiple models using different training samples; Adopt different neural network structures for multiple models.

4. The method according to claim 1, wherein According to the difficulty score, selecting text images meeting preset conditions as samples includes: According to a preset difficulty score threshold, select text images whose difficulty scores exceed the threshold as samples.

5. An apparatus (200) for sample selection of text images, wherein, An initial module (201) for performing different initial training on multiple models; A prediction module (202) that inputs text images into multiple models to obtain prediction results of multiple models for multiple text positions of the text images; A score module (203) for obtaining the Difficulty scores; And synthesize the difficulty scores of the multiple text positions to obtain the difficulty score of the text image; A selection module (204) for selecting text images meeting preset conditions as samples according to the difficulty score.

6. The device according to claim 5, wherein The prediction module (202) is further configured to: Calculate the sequence features of the text images through convolution and downsampling of each model; According to the sequence features, obtain prediction vectors through a recurrent neural network layer; According to the prediction vectors, use the maximum independent variable point set function to obtain the prediction results of each model for each text position of the text image.

7. The apparatus according to claim 5, wherein, The initial module (201) is further configured to: Perform different initial weights on multiple models; Set random numbers with different Gaussian distributions for multiple models; Train multiple models using different training samples; Adopt different neural network structures for multiple models.

8. The apparatus according to claim 5, wherein The selection module (204) includes: According to a preset difficulty score threshold, select text images whose difficulty scores exceed the threshold as samples.

9. A computer device, comprising a memory and a processor, the memory storing a computer program, wherein, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 4.

10. A computer-readable storage medium having a computer program stored thereon, wherein, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 4.

11. A computer program product, the computer program product being tangibly stored on a computer-readable medium and comprising computer-executable instructions that, when executed, cause at least one processor to perform the method according to any one of claims 1 to 4.

Citation Information

Patent Citations

  • Difficult sample mining and model training method, device and electronic equipment

    CN110610197A

  • Method and device for mining data

    CN111768007A

  • Model training method, device and equipment

    CN114677680A

  • Natural scene text detection model training method and device, server and storage medium

    CN115761752A

  • Image processing method, electronic device, and storage medium

    WO2023280229A1