Model training method and device, electronic equipment and storage medium
By performing image clustering and training models, and utilizing LayoutLMv3 and CRF models, the problem of slow recognition in image recognition schemes under conditions of varied styles was solved, achieving fast and accurate recognition of unknown styles.
Patent Information
- Application Number
- CN202510966966.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-14
- Publication Date
- 2025-11-18
AI Technical Summary
Existing image recognition solutions suffer from slow response and poor image recognition performance when faced with diverse styles, requiring manual adjustment of templates to adapt to new styles.
By acquiring the image to be processed, performing clustering, and using LayoutLMv3 and CRF models to train an image recognition model, the model can quickly identify unknown patterns, including feature extraction, dimensionality reduction and clustering of text data and location information. The model is then trained in conjunction with the OCR results.
It enables rapid recognition of unknown patterns, improves image recognition performance, reduces slow response to new patterns, and enhances the model's learning ability and recognition accuracy.
Smart Images

Figure CN120976934A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of computer, and particularly relates to a model training method and device, electronic equipment and storage medium. BACKGROUND
[0002] In actual application, the recognition scheme for the image with text content can be implemented in combination with an optical character recognition (OCR) technology and a machine learning algorithm. For example, taking a bill image as an example, the text in the image can be converted into machine recognizable text by an OCR tool, and then the text is analyzed by means of a machine learning algorithm, so as to adapt to different bill styles.
[0003] However, the existing image recognition scheme mainly relies on a fixed template or a manually configured template to implement a recognition process. Therefore, in the case of a large number of existing bills, cards, forms and other styles, if the text in the image is arranged in a new style, the template in the scheme often needs to be manually adjusted to adapt to the recognition requirement of the new style. Therefore, in the face of such a variable style, the whole process recognition response is slow, and the image recognition effect is poor. SUMMARY
[0004] The embodiments of the present application provide a model training method, device, electronic equipment and storage medium to improve the image recognition effect.
[0005] In a first aspect, the embodiments of the present application provide a model training method, comprising:
[0006] obtaining a to-be-processed image, text data in the to-be-processed image being arranged in a target style;
[0007] performing clustering processing on the to-be-processed image to obtain a clustering result, the clustering result being used to indicate that the target style is a new style;
[0008] performing model training on a to-be-trained model based on the to-be-processed image and an optical character recognition (OCR) result of the to-be-processed image to obtain an image recognition model, the to-be-trained model comprising a document base model (LayoutLMv3) and a conditional random field (CRF) model connected in sequence, the LayoutLMv3 being used to generate a graphic-text feature vector corresponding to the to-be-processed image based on the to-be-processed image and the OCR result, and the CRF model being used to process the graphic-text feature vector to obtain a recognition result of the text data.
[0009] Optionally, the performing clustering processing on the to-be-processed image to obtain a clustering result comprises:
[0010] processing the to-be-processed image through an OCR tool to obtain an OCR result, the OCR result including the text data and position information of the text data in the to-be-processed image;
[0011] processing the text data and the position information to obtain text features corresponding to the text data and position features corresponding to the position information;
[0012] performing feature extraction on the to-be-processed image to obtain image features of the to-be-processed image;
[0013] generating the image-text feature vector based on the image features, the text features, and the position features through the LayoutLMv3;
[0014] performing dimension reduction processing on the image-text feature vector to obtain a dimension reduction processing result;
[0015] performing clustering processing on the dimension reduction processing result to obtain the clustering result.
[0016] Optionally, the performing clustering processing on the dimension reduction processing result to obtain the clustering result includes:
[0017] setting a value of a minimum cluster for clustering processing, the value being a number of historical styles;
[0018] performing clustering processing on the dimension reduction processing result to obtain an initial clustering result, the initial clustering result being used to indicate whether the target style belongs to the historical styles;
[0019] generating the clustering result in a case where the initial clustering result indicates that the target style does not belong to the historical styles.
[0020] Optionally, the OCR result includes position information of the text data in the to-be-processed image, the text data including a plurality of words, and the LayoutLMv3 includes an attention mechanism;
[0021] the attention mechanism being used to learn a position relationship between the plurality of words and adjust the position information based on the position relationship in a case where the position information is incorrect.
[0022] Optionally, before the model training is performed on the to-be-processed image and an optical character recognition (OCR) result of the to-be-processed image to obtain an image recognition model, the method further includes:
[0023] creating a labeling task for the to-be-processed image, the labeling task being used to label the OCR result to obtain a labeling result;
[0024] The model training is performed on a to-be-trained model based on the to-be-processed image and an optical character recognition (OCR) result of the to-be-processed image, to obtain an image recognition model, including:
[0025] The model training is performed on a to-be-trained model based on the to-be-processed image, the OCR result and the annotation result, to obtain the image recognition model.
[0026] Optionally, the method further includes:
[0027] A model test task is created, the model test task is configured to receive a to-be-tested image, and model inference is performed on the to-be-tested image based on the image recognition model to obtain a model inference result of the to-be-tested image.
[0028] The model inference result is output.
[0029] Optionally, the model training method is applied to a model training system, and the method further includes:
[0030] Operation permission information of a plurality of different user roles in the model training system is set;
[0031] The model training system is managed based on the operation permission information.
[0032] In a second aspect, an embodiment of the present application provides a model training apparatus, including:
[0033] An image acquisition module is configured to acquire a to-be-processed image, and text data in the to-be-processed image is arranged in a target style;
[0034] An image clustering module is configured to perform clustering processing on the to-be-processed image to obtain a clustering result, and the clustering result is configured to indicate that the target style is a new style;
[0035] A model training module is configured to perform model training on a to-be-trained model based on the to-be-processed image and an optical character recognition (OCR) result of the to-be-processed image, to obtain an image recognition model, and the to-be-trained model includes a document base model (LayoutLMv3) and a conditional random field (CRF) model connected in sequence, the LayoutLMv3 is configured to generate a picture-text feature vector corresponding to the to-be-processed image based on the to-be-processed image and the OCR result, and the CRF model is configured to process the picture-text feature vector to obtain a recognition result of the text data.
[0036] In a third aspect, an embodiment of the present application provides an electronic device, including a processor, a memory and a system bus;
[0037] The processor and the memory are connected through the system bus.
[0038] The memory is configured to store a program, the program comprising instructions which, when executed by the processor, cause the processor to perform any implementation step of the above model training method.
[0039] In a fourth aspect, an embodiment of the present application provides a computer readable storage medium, and the computer readable storage medium stores a computer program. When the computer program is executed by a processor, any implementation step of the above model training method is implemented.
[0040] From the above technical solutions, the embodiments of the present application have the following advantages:
[0041] In the embodiments of the present application, first, a to-be-processed image can be obtained, and the text data in the to-be-processed image is arranged according to a target style. Then, the to-be-processed image is subjected to clustering processing, and a clustering result can be obtained, which is used to indicate that the target style is a new style. In this way, model training is performed on a to-be-trained model based on the to-be-processed image and the OCR result of the to-be-processed image, and an image recognition model is obtained, wherein the to-be-trained model comprises a LayoutLMv3 and a CRF model connected in sequence, the LayoutLMv3 is used to obtain a graphic-text feature vector corresponding to the to-be-processed image based on the to-be-processed image and the OCR result, and the CRF model is used to process the graphic-text feature vector to obtain a recognition result of the text data. It can be seen that, since the clustering result corresponding to the to-be-processed image can indicate whether the target style of the text data of the to-be-processed image is a new style, the fast recognition of unknown styles can be realized. In this way, when a new style appears (i.e., the target style is a new style), the model training process can be directly triggered by using the to-be-processed image and the OCR result corresponding to the to-be-processed image. In this way, the characteristics of the new style can be quickly learned through model training, so that the model trained subsequently can perform image recognition, thereby solving the problem of slow response to the recognition of new styles and improving the image recognition effect. BRIEF DESCRIPTION OF DRAWINGS
[0042] Figure 1 A flowchart of a model training method provided by an embodiment of the present application;
[0043] Figure 2 A schematic diagram of an exemplary electronic railway ticket provided by an embodiment of the present application;
[0044] Figure 3 A flowchart of a clustering processing process provided by an embodiment of the present application;
[0045] Figure 4 A flowchart of a model training process provided by an embodiment of the present application;
[0046] Figure 5a A schematic diagram for creating a labeling task provided by an embodiment of the present application;
[0047] Figure 5b A schematic diagram for user labeling provided by an embodiment of the present application;
[0048] Figure 6 A schematic diagram for performing a model test task provided by an embodiment of the present application;
[0049] Figure 7 A schematic diagram for user role management provided by an embodiment of the present application;
[0050] Figure 8 A structural schematic diagram of a model training system provided by an embodiment of the present application;
[0051] Figure 9 A structural schematic diagram of a model training device provided by an embodiment of the present application. DETAILED DESCRIPTION
[0052] As described above, the existing image recognition scheme mainly relies on fixed templates or manually configured templates to implement the recognition process. Therefore, in the case of a large number of styles of bills, cards, or forms, if the text in the image is arranged in a new style, the template in the scheme often needs to be manually adjusted to adapt to the recognition requirements of the new style. Therefore, in the face of such variable styles, the entire process recognition response is slow, and the image recognition effect is poor.
[0053] Based on this, in order to solve the above problems, an embodiment of the present application provides a model training method, first, a to-be-processed image can be obtained, the text data in the to-be-processed image is arranged in a target style. Then, the to-be-processed image is clustered to obtain a clustering result, which indicates that the target style is a new style. In this way, the to-be-trained model is trained based on the to-be-processed image and the OCR result of the to-be-processed image, that is, an image recognition model is obtained, wherein the to-be-trained model includes LayoutLMv3 and CRF model connected in sequence, LayoutLMv3 is used to obtain the corresponding graph-text feature vector of the to-be-processed image based on the to-be-processed image and the OCR result, and the CRF model is used to process the graph-text feature vector to obtain the recognition result of the text data.
[0054] It can be seen that, since the clustering result corresponding to the to-be-processed image can represent whether the target style of the text data of the to-be-processed image is a new style, the fast recognition of the unknown style can be realized. In this way, when a new style appears (i.e., the target style is a new style), the model training process can be directly triggered by using the to-be-processed image and the OCR result corresponding to the to-be-processed image. In this way, the characteristics of the new style can be quickly learned through model training, so that the model obtained through subsequent training can perform image recognition, thereby solving the problem of slow response to new style recognition and improving the image recognition effect.
[0055] It should be noted that the model training method of the embodiments of the present application can not be limited to the execution subject, for example, the model training method of the embodiments of the present application can be applied to a terminal device or a server and the like data processing device. The terminal device can be a smart phone, a computer, a personal digital assistant (PDA), a tablet computer and the like electronic device. The server can be a stand-alone server, a cluster server or a cloud server.
[0056] In order to make the purpose, technical solutions and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be described clearly and completely below in combination with the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative labor are within the scope of protection of the present application.
[0057] Figure 1 A flowchart of a model training method provided by the embodiments of the present application is shown. In combination with FIG. 1, the model training method provided by the embodiments of the present application can include the following steps S101-S103. Figure 1
[0058] S101: Obtain a to-be-processed image, and text data in the to-be-processed image is arranged according to a target style.
[0059] The to-be-processed image refers to an image that has not been used to train an image recognition model. For example, the to-be-processed image can be an image of various electronic invoices, such as a railway electronic ticket, an air transportation electronic ticket itinerary, a motor vehicle sales unified invoice and a second-hand vehicle sales unified invoice, wherein the text data is arranged according to the corresponding target style. Figure 2 A schematic diagram of a railway electronic ticket is shown.
[0060] Correspondingly, for the acquisition process of the to-be-processed image, in specific implementation, the to-be-processed image can be uploaded by the user himself / herself, or can be obtained by crawling from a related business system through data crawling.
[0061] For the user self-uploaded mode, the format of the to-be-processed image needs to be predefined, such as a standard picture format like jpg or png, and a compressed file format like zip or rar. For the compressed file format, the compressed file can be decompressed after being obtained, so as to restore the compressed file to the standard picture format and obtain the to-be-processed image. For the data crawling mode, the initial image can also be converted to the standard picture format after being obtained, so as to obtain the to-be-processed image.
[0062] S102: performing clustering processing on the to-be-processed image to obtain a clustering result, where the clustering result is used to indicate that the target style is a new style.
[0063] In the embodiment of the present application, the to-be-processed image is subjected to clustering processing, which aims to quickly distinguish whether the text data in the to-be-processed image is arranged in a new style. For example, if the to-be-processed image is a bill image, it is determined whether it is a new style of a certain bill. In this way, if it belongs to a new style, it can be marked for subsequent training process; if it does not belong to a new style, it can be integrated into the data set of the belonging style.
[0064] Corresponding to this, for the convenience of understanding, the clustering processing process is exemplarily described below in combination with the accompanying drawings. As shown in FIG. 3, the process of clustering processing the to-be-processed image can include the following steps 31-35. Figure 3
[0065] Step 31: processing the to-be-processed image by using an OCR tool to obtain an OCR result, where the OCR result includes text data and position information of the text data in the to-be-processed image.
[0066] Here, the implementation of the OCR tool and the corresponding processing process are not limited in the embodiment of the present application, and any existing or future OCR tool can be used for implementation, for example, some open source OCR tools with fast recognition speed and high recognition accuracy.
[0067] Step 32: processing the text data and the position information to obtain text features corresponding to the text data and position features corresponding to the position information.
[0068] For the process of obtaining the text features, first, the text data can be segmented to obtain a segmentation result, which can be a word or a token, that is, the text data can include multiple words. Here, the segmentation process is not limited in the embodiment of the present application, and any existing or future segmentation algorithm can be used for implementation.
[0069] Then, specific tokens, i.e., [CLS] and [SEP], can be added to the word segmentation result to classify and segment the text data. After adding the specific tokens, the text feature corresponding to the text data can be obtained.
[0070] For the position feature acquisition process, the boundary box coordinates (denoted as X0, Y0, X1, Y1) corresponding to each word segmentation result of the text data can be obtained as its position information. Then, the coordinates are normalized to the range of 0-1000, thereby obtaining the corresponding position feature.
[0071] Step 33: Feature extraction is performed on the to-be-processed image to obtain the image feature of the to-be-processed image.
[0072] In specific implementation, the size of the to-be-processed image can be adjusted to a size suitable for model input to obtain a first image, for example, the adjusted size is 224*224. Then, normalization processing is performed on the first image to obtain a second image. Finally, the second image can be subjected to feature extraction by a pre-trained image model, thereby obtaining the image feature of the to-be-processed image. For example, ResNet50, a deep convolutional neural network model suitable for feature extraction, can be used to implement the feature extraction process of the second image.
[0073] It should be noted that the execution order of steps 31 and 33 is not limited in the embodiments of the present application. Step 31 can be executed first and then step 33 can be executed, or step 33 can be executed first and then step 31 can be executed, or steps 31 and 33 can be executed in parallel, as long as step 32 is executed after step 31.
[0074] Step 34: An image-text feature vector is generated by LayoutLMv3 based on the image feature, the text feature, and the position feature.
[0075] Here, the above three features can be added to obtain an added feature. Then, the added feature is input to LayoutLMv3, thereby obtaining the image-text feature vector of the to-be-processed image through the learning processing of the LayoutLMv3. Here, the learning processing process of the LayoutLMv3 is not limited in the embodiments of the present application, and any algorithm corresponding to the architecture of the LayoutLMv3 that exists at present or may appear in the future can be implemented.
[0076] Step 35: Dimensionality reduction processing is performed on the image-text feature vector to obtain a dimensionality reduction processing result.
[0077] In the embodiments of the present application, the dimension of the image-text feature vector is too high, which can cause the clustering boundary to be blurred, resulting in clustering failure and high computational cost. Therefore, by performing dimension reduction processing on the image-text feature vector, the above problems can be effectively avoided, and the clustering effect can be improved.
[0078] In actual applications, the Uniform Manifold Approximation and Projection (UMAP) algorithm can be used for dimension reduction processing. The UMAP algorithm can maintain the proximity structure and style distribution in the original space, and is still suitable for clustering after dimension reduction, thereby helping to improve the clustering effect.
[0079] Correspondingly, in specific implementation, the umap library can be imported from the python third-party library, and the corresponding parameters such as the number of neighbors, the minimum distance, or the number of dimensions after dimension reduction are adjusted, and the umap.UMAP method is run to obtain the vector after dimension reduction. For example, the image-text feature vector can be reduced to two dimensions to facilitate subsequent visualization.
[0080] Step 36: performing clustering processing on the dimension reduction processing result to obtain a clustering result.
[0081] In actual applications, the Hierarchical Density-Based Spatial Clustering of Applications with Noise (HDBSCAN) algorithm can be used for clustering processing. It supports high-dimensional dense embedding and is particularly suitable for processing vectors output by LayoutLMv3, such as the image-text feature vector described above. At the same time, it can also automatically identify outliers and is sensitive to abnormal samples, so it can be used as a basis for triggering few-sample training.
[0082] Correspondingly, in specific implementation, the minimum cluster value for clustering processing can be set first, which is the number of historical styles. Here, the historical style refers to the style arranged by the historical image and text data. For example, if the number of historical styles saved in the database is 4, the minimum cluster value can be set to 4.
[0083] Then, the dimension reduction processing result can be clustered to obtain an initial clustering result, which can indicate whether the target style belongs to the historical style. Specifically, the hdbscan library can be imported from the python third-party library, and the parameters such as the distance measurement method or the minimum number of samples of the clustering core point are adjusted, and the hdbscan.HDBSCAN method is run to obtain the initial clustering result.
[0084] In this way, the clustering result is generated in a case that the initial clustering result indicates that the target pattern does not belong to the historical patterns, that is, the target pattern is a new pattern.
[0085] In addition, an index corresponding to the target pattern can be established in the database, so as to facilitate subsequent quick search of the target pattern.
[0086] S103: Model training is performed on the to-be-trained model based on the to-be-processed image and the OCR result of the to-be-processed image, to obtain an image recognition model, the to-be-trained model comprising a LayoutLMv3 and a CRF model connected in sequence, the LayoutLMv3 being configured to generate a graphic-text feature vector corresponding to the to-be-processed image based on the to-be-processed image and the OCR result, and the CRF model being configured to process the graphic-text feature vector to obtain a recognition result of the text data.
[0087] The to-be-processed OCR result can be directly obtained in the clustering process of the above step S102. Alternatively, in combination with Figure 4 As shown in FIG. 6, before the model training, the to-be-processed image can be processed again by using the OCR tool used in the above step S102 to obtain a corresponding OCR result. In this way, the same OCR tool is used for processing for the purpose of realizing a predetermined function, which can reduce the waste of GPU resources and avoid a too complicated training framework.
[0088] Further, the to-be-trained model can be implemented by joint modeling of the LayoutLMv3 and the CRF model, so as to strengthen the learning of the context by the image recognition model obtained through training. As can be seen, the LayoutLMv3 is used in both the model training and the image clustering stage, so that repeated calculation by using the same model algorithm can further reduce the waste of GPU resources and avoid a too complicated training framework.
[0089] The CRF model is a discriminative model for sequence labeling, which models the conditional probability of the input sequence (such as the token of the text data) and the output label sequence (such as the field category). Moreover, the CRF can not only consider what label the current word is, but also consider whether the label of the previous word has an impact on the judgment of the current word.
[0090] Based on this, the to-be-trained model is jointly modeled by the LayoutLMv3 and the CRF model, which can enhance the decoder layer of the LayoutLMv3 by using the CRF. In a specific implementation, a fully connected layer and the CRF model can be connected after the output layer of the LayoutLMv3. In this way, during the model training, the CRF model learns the correct label scores of each token and the transition probability matrix between labels, so as to learn the mutual relationship between the previous labels and the current label, thereby enhancing the reasoning accuracy.
[0091] Further, the CRF model can fully consider the dependency between labels, for example, for the text in the bill image, the amount usually follows the currency unit, so that the precision of field boundary recognition can be enhanced by means of the CRF model.
[0092] Further, the LayoutLMv3 can include an attention mechanism. Based on this, the position relationship between multiple words in the text data can be learned by means of the attention mechanism, so that in the case that the position information of the text data obtained by the OCR tool in the image to be processed is incorrect, the position information is adjusted based on the position relationship. For example, the "tax" should appear in the "amount" field, if the OCR tool identifies the "tax" as "tax solid", then the LayoutLMv3 can correct it through the attention mechanism, thereby reducing the probability of subsequent NLP extraction failure.
[0093] In this way, in the embodiment of the present application, the attention mechanism in the LayoutLMv3 and the context semantics provided by the CRF model are combined for joint error correction, which can significantly reduce the structure extraction failure rate caused by character recognition errors.
[0094] In addition, for the OCR result recognized by the OCR tool, the OCR result and the message for indicating the recognition error of the OCR result can also be fed back to the OCR tool, and periodically optimized and trained, so that the OCR tool can continuously learn and continuously improve the recognition accuracy.
[0095] Further, before the above-mentioned model training process is performed, the image to be processed of the new style (i.e. the target style) can also be labeled to obtain a labeling result, so as to further perform model training by using the labeling result and improve the training accuracy. For ease of understanding, the process is specifically described below.
[0096] In specific implementation, first, a labeling task for the image to be processed can be created, and the labeling task is used to label the OCR result to obtain a labeling result. In this way, in the process of model training, the model to be trained can be trained based on the image to be processed, the OCR result and the labeling result, to obtain an image recognition model.
[0097] In actual application, the labeling task can be realized by means of collaborative labeling, that is, multiple people are allowed to collaborate to label the image to be processed or other images, so as to improve the speed of process and improve the efficiency of the model.
[0098] As an example, in combination with Figure 5aAs shown, when creating a labeling task, multiple labelers can be added to label online synchronously or asynchronously, and an auditor can also be set to audit the labeling results. Accordingly, multiple auditors can also be added, and the purpose of auditing is to judge the labeling results to reduce the error rate of manual labeling.
[0099] In addition, in Figure 5a , each labeling task has a unique name to distinguish it from other labeling tasks. The "labeling data set" can select a pre-uploaded data set, which includes images to be processed. The "pre-labeling base model" can select the corresponding base pre-training model of LayoutLMv3, which is a model that has not learned any images, or it can also select a pre-training model that has learned specific images. The "training model" is a pre-set option, and other models can be added to facilitate joint learning with the pre-labeling base model described above. After the page is set, the labeling task of the image to be processed is created, so that each user can label the image to be processed to obtain the labeling results.
[0100] Further, in practical applications, in combination with Figure 5b As shown, in the labeling page, the user can label the key-value pair on the right side according to the text block (orange covered part) pre-recognized by the OCR tool on the left side. Accordingly, the background will automatically record the position and content of the text data to facilitate subsequent model training and learning.
[0101] In addition, after labeling the data and training the model, the model effect also needs to be verified.
[0102] After obtaining the image recognition model, the model effect can be further verified to understand the model inference effect and better arrange the subsequent model optimization process. For ease of understanding, the process is described in detail below.
[0103] As an example, first, a model test task can be created, which is used to receive a test image and perform model inference on the test image based on the image recognition model to obtain a model inference result of the test image. Then, based on the model test task, the model inference result can be output.
[0104] In practical applications, each style of image can create a corresponding model test task to distinguish the inference effect of the trained image recognition model for different styles. Accordingly, in combination with Figure 6 As shown, after uploading the test image, the inference can be performed by clicking the "execute" control to utilize the image recognition model. The inference result and inference time information are displayed at the same time.
[0105] Furthermore, in practical applications, a model training system can be provided to implement any of the above-described model training methods. Based on this, in this embodiment, to facilitate the management of the model training system and the model training method, various different user roles can be set with different access permissions within the model training system. For example, combined with... Figure 7 As shown, various user roles can include system administrators, regular users, annotators, and annotation reviewers. Correspondingly, the system administrator's access permissions can specify that they are allowed access to all pages provided by the model training system; regular users' access permissions can specify that they are only allowed access to the pages corresponding to the model testing tasks; annotators' access permissions can specify that they are only allowed access to the pages corresponding to the annotation tasks; and annotation reviewers' access permissions can specify that they are allowed access to the pages corresponding to the annotation tasks and the pages displaying each annotator's annotation results. This facilitates the management of the model training system based on access permissions, improving its security.
[0106] As can be seen from the above steps S101-S103, in this embodiment, firstly, an image to be processed is acquired, and the text data in the image is arranged according to the target style. Then, the image to be processed is clustered to obtain a clustering result, which indicates that the target style is a new style. Thus, based on the image to be processed and its OCR result, a training model is performed to obtain an image recognition model. This training model includes a LayoutLMv3 and a CRF model connected sequentially. LayoutLMv3 is used to obtain the image-text feature vector corresponding to the image to be processed based on the image to be processed and the OCR result. The CRF model is used to process the image-text feature vector to obtain the recognition result of the text data. It is evident that since the clustering result corresponding to the image to be processed can indicate whether the target style of the text data in the image to be processed is a new style, rapid recognition of unknown styles can be achieved. Therefore, when a new style appears (i.e., the target style is a new style), the image to be processed and its corresponding OCR result can be used to directly trigger the model training process. In this way, the features of new styles can be quickly learned through model training, so that the subsequently trained model can perform image recognition, thereby solving the problem of slow response to new style recognition and improving the image recognition effect.
[0107] Furthermore, based on the model training method provided in the above embodiments, this application embodiment can also provide a model training system. The model training system will now be described in conjunction with the embodiments and accompanying drawings.
[0108] Figure 8 A structural schematic diagram of a model training system provided by an embodiment of the present application is shown. The model training system provided by the embodiment of the present application can include a data acquisition module, a data clustering module, a data labeling module, a model training module, a model testing module, and a background management module. Figure 8
[0109] The data acquisition module can be configured to acquire the to-be-processed image. The specific implementation of the data acquisition module can refer to the description of step S101 in the above embodiment, and will not be repeated here.
[0110] The data clustering module can be configured to perform clustering processing on the to-be-processed image to obtain a clustering result, and the clustering result is used to indicate that the target style is a new style. The specific implementation of the data clustering module can refer to the description of step S102 in the above embodiment, and will not be repeated here.
[0111] The data labeling module can be configured to label the to-be-processed image. The specific implementation of the data labeling module can refer to the description of the related content for the labeling task in the above embodiment, and will not be repeated here.
[0112] The model training module can be configured to label the to-be-processed image. The specific implementation of the data training module can refer to the description of step S103 in the above embodiment, and will not be repeated here.
[0113] The model testing module can be configured to test the to-be-processed image. The specific implementation of the data testing module can refer to the description of the related content for the model testing task in the above embodiment, and will not be repeated here.
[0114] The background management module can be configured to manage the model training system as a whole. The overall management can include management of user roles, management of user information, and management of log information. The management of log information can include recording, querying, and analyzing various operations and events in the system running process. The management of user information can include maintaining and managing user login information. The specific implementation of the management of user roles can refer to the description of the related content such as setting operation permission information in the above embodiment, and will not be repeated here.
[0115] In this way, in the embodiment of the present application, based on the cooperation between the above-mentioned multiple modules, the rapid identification of unknown styles can be realized. When a new style appears, the model training process can be directly triggered by using the to-be-processed image and the OCR result corresponding to the to-be-processed image. In this way, the characteristics of the new style can be quickly learned through model training, so that the model obtained through subsequent training can perform image recognition, thereby solving the problem of slow response to new style recognition and improving the image recognition effect. In addition, the model training system also supports separate disassembly and update of each algorithm or flexible combination and use of different functional modules, so that the modules are highly decoupled and do not affect each other, thereby improving the application effect of the model training system.
[0116] Further, based on the model training method provided in the above embodiment, the embodiment of the present application can also provide a model training device. The model training device will be described below in combination with the embodiments and the drawings.
[0117] Figure 9 FIG. 1 is a structural schematic diagram of a model training device provided in an embodiment of the present application. As shown in FIG. 1, the model training device 900 provided in the embodiment of the present application can include: Figure 9
[0118] An image acquisition module 901, configured to acquire a to-be-processed image, wherein text data in the to-be-processed image is arranged in a target style;
[0119] An image clustering module 902, configured to perform clustering processing on the to-be-processed image to obtain a clustering result, wherein the clustering result is used to indicate that the target style is a new style;
[0120] A model training module 903, configured to perform model training on a to-be-trained model based on the to-be-processed image and an optical character recognition (OCR) result of the to-be-processed image to obtain an image recognition model, wherein the to-be-trained model includes a document base model (LayoutLMv3) and a conditional random field (CRF) model connected in sequence, the LayoutLMv3 is used to generate a graphic-text feature vector corresponding to the to-be-processed image based on the to-be-processed image and the OCR result, and the CRF model is used to process the graphic-text feature vector to obtain an identification result of the text data.
[0121] Optionally, the image clustering module 902 is specifically configured to:
[0122] A first processing module, configured to process the to-be-processed image through an OCR tool to obtain the OCR result, wherein the OCR result includes the text data and position information of the text data in the to-be-processed image;
[0123] The second processing module is configured to process the text data and the position information to obtain text features corresponding to the text data and position features corresponding to the position information.
[0124] The feature extraction module is configured to perform feature extraction on the to-be-processed image to obtain image features of the to-be-processed image.
[0125] The feature generation module is configured to generate the image-text feature vector based on the image features, the text features and the position features by using the LayoutLMv3.
[0126] The dimension reduction processing module is configured to perform dimension reduction processing on the image-text feature vector to obtain a dimension reduction processing result.
[0127] The image clustering submodule is configured to perform clustering processing on the dimension reduction processing result to obtain the clustering result.
[0128] Optionally, the image clustering submodule is specifically configured to:
[0129] set a value of a minimum cluster for clustering processing, the value being a number of historical styles;
[0130] perform clustering processing on the dimension reduction processing result to obtain an initial clustering result, the initial clustering result being used to indicate whether the target style belongs to the historical styles;
[0131] generate the clustering result in a case where the initial clustering result indicates that the target style does not belong to the historical styles.
[0132] Optionally, the OCR result includes position information of the text data in the to-be-processed image, the text data includes a plurality of words, and the LayoutLMv3 includes an attention mechanism.
[0133] The attention mechanism is configured to learn a position relationship between the plurality of words, and adjust the position information based on the position relationship in a case where the position information is incorrect.
[0134] Optionally, the model training apparatus 900 further includes:
[0135] The first task creation module is configured to create a labeling task for the to-be-processed image, the labeling task being used to label the OCR result to obtain a labeling result.
[0136] The model training module 903 is specifically configured to:
[0137] perform model training on a to-be-trained model based on the to-be-processed image, the OCR result and the labeling result to obtain the image recognition model.
[0138] Optionally, the model training apparatus 900 further comprises:
[0139] a second task creation module configured to create a model test task, the model test task being configured to receive a to-be-tested image and perform model inference on the to-be-tested image based on the image recognition model to obtain a model inference result of the to-be-tested image;
[0140] a data output module configured to output the model inference result.
[0141] Optionally, the model training method is applied to a model training system, and the model training apparatus 900 further comprises:
[0142] an information setting module configured to set operation permission information of a plurality of different user roles in the model training system;
[0143] a system management module configured to manage the model training system based on the operation permission information.
[0144] Further, an electronic device is provided in the embodiment of the present application, which comprises a processor, a memory and a system bus.
[0145] The processor and the memory are connected through the system bus.
[0146] The memory is configured to store one or more programs, the one or more programs comprising instructions which, when executed by the processor, cause the processor to perform any implementation step of the model training method described above.
[0147] Further, a computer readable storage medium is provided in the embodiment of the present application, the computer readable storage medium storing instructions, when the instructions run on an electronic device, causing any implementation step of the model training method described above.
[0148] Those skilled in the art can clearly understand that all or part of the steps of the above-mentioned method in the embodiments can be implemented by means of software and necessary universal hardware platforms based on the description of the embodiments. Based on such understanding, the technical solutions of the present application can be embodied in the form of a software product, which can be stored in a storage medium, such as a ROM / RAM, a magnetic disk, an optical disk, etc., and includes a number of instructions to make a computer device (which can be a personal computer, a server, or a network communication device such as a media gateway, etc.) execute the methods described in the various embodiments or some parts of the embodiments. It should be noted that the various embodiments in the specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the various embodiments can be referred to each other.
[0149] For the system disclosed in the embodiments, since it corresponds to the method disclosed in the embodiments, the description is relatively simple, and the relevant parts are referred to the method part.
[0150] It should also be noted that the relationship terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply that there is any such actual relationship or order between the entities or operations. Moreover, the terms "include", "contain" or any other variants thereof are intended to cover non-exclusive inclusion, so that the process, method, article or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed or inherent to such process, method, article or device. Without more limitations, the element defined by the statement "including a" does not exclude the presence of another identical element in the process, method, article or device including the element.
[0151] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present application. Various modifications to the embodiments will be apparent to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the application. Therefore, the present application will not be limited to the embodiments shown herein, but will conform to the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A model training method, characterized in that, include: Acquire an image to be processed, wherein the text data in the image to be processed is arranged according to a target style; The image to be processed is subjected to clustering to obtain clustering results, which are used to indicate that the target style is a new style; The image to be processed and the optical character recognition (OCR) result of the image to be processed are used to train the model to obtain an image recognition model. The model to be trained includes a document base model LayoutLMv3 and a conditional random field (CRF) model connected in sequence. The LayoutLMv3 is used to generate the image and text feature vector corresponding to the image to be processed based on the image to be processed and the OCR result. The CRF model is used to process the image and text feature vector to obtain the recognition result of the text data.
2. The model training method according to claim 1, characterized in that, The clustering process of the image to be processed to obtain the clustering result includes: The image to be processed is processed using an OCR tool to obtain the OCR result, which includes the text data and the position information of the text data in the image to be processed. The text data and the location information are processed to obtain the text features corresponding to the text data and the location features corresponding to the location information. Feature extraction is performed on the image to be processed to obtain the image features of the image to be processed; The image-text feature vector is generated using the LayoutLMv3 based on the image features, the text features, and the location features. The image and text feature vectors are subjected to dimensionality reduction processing to obtain the dimensionality reduction result; The dimensionality reduction result is then subjected to clustering to obtain the clustering result.
3. The model training method according to claim 2, characterized in that, The clustering process performed on the dimensionality reduction result to obtain the clustering result includes: Set the minimum cluster value used for clustering processing, where the value is the number of historical patterns; The dimensionality reduction result is subjected to clustering to obtain an initial clustering result, which is used to indicate whether the target style belongs to the historical style; If the initial clustering result indicates that the target style does not belong to the historical style, the clustering result is generated.
4. The model training method according to any one of claims 1 to 3, characterized in that, The OCR result includes the location information of the text data in the image to be processed, the text data includes multiple words, and the LayoutLMv3 includes an attention mechanism; The attention mechanism is used to learn the positional relationships between the multiple words; Furthermore, if the location information is incorrect, the location information is adjusted based on the location relationship.
5. The model training method according to any one of claims 1 to 3, characterized in that, Before training the image recognition model based on the image to be processed and the optical character recognition (OCR) results of the image to be processed, the method further includes: A labeling task is created for the image to be processed, the labeling task being used to label the OCR results to obtain the labeling results; The process of training the image recognition model based on the image to be processed and the optical character recognition (OCR) results of the image to be processed to obtain the image recognition model includes: The image recognition model is obtained by training the model based on the image to be processed, the OCR result, and the annotation result.
6. The model training method according to any one of claims 1 to 3, characterized in that, The method further includes: A model testing task is created, which is used to receive the image to be tested and perform model inference on the image to be tested based on the image recognition model to obtain the model inference result of the image to be tested. Output the inference results of the model.
7. The model training method according to any one of claims 1 to 3, characterized in that, The model training method is applied to a model training system, and the method further includes: Configure operation permission information for various user roles in the model training system; The model training system is managed based on the aforementioned operation permission information.
8. A model training device, characterized in that, include: An image acquisition module is used to acquire an image to be processed, wherein the text data in the image to be processed is arranged according to a target style; An image clustering module is used to perform clustering processing on the image to be processed to obtain clustering results, wherein the clustering results are used to indicate that the target style is a new style; The model training module is used to train the model to be trained based on the image to be processed and the optical character recognition (OCR) result of the image to be processed, so as to obtain an image recognition model. The model to be trained includes a document base model LayoutLMv3 and a conditional random field (CRF) model connected in sequence. The LayoutLMv3 is used to generate the image and text feature vector corresponding to the image to be processed based on the image to be processed and the OCR result. The CRF model is used to process the image and text feature vector to obtain the recognition result of the text data.
9. An electronic device, characterized in that, The device includes: a processor, a memory, and a system bus; The processor and the memory are connected via the system bus; The memory is used to store a program, the program including instructions that, when executed by the processor, cause the processor to perform the steps of the model training method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the steps of the model training method as described in any one of claims 1 to 7.