Image character reconstruction model training method and device, medium and electronic equipment

By jointly training the codec and character recognition model in the image character reconstruction model, obtaining and fusing the adapter output features and character recognition intermediate features, the problem of blurred and difficult to recognize the license plate text in image reconstruction is solved, and the text information quality of the reconstructed image is improved.

CN119942559APending Publication Date: 2025-05-06CHINA TELECOM CORP LTD TECHNOLOGY INNOVATION CENTER +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411998187.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-31
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

During the training process, the image codec based on intelligent tasks will filter out the background license plate text information of the image, resulting in the decoded features no longer containing the license plate text information, which will cause blurred and difficult to recognize in the pictures generated by image reconstruction.

Method used

By obtaining the preprocessed image of the image character reconstruction model, input it into the codec model to be jointly trained, obtaining the adapter output features and character recognition intermediate features. The codec model is trained based on the joint loss function of these features, obtaining the decoded features of the trained codec, and image reconstruction training is carried out on the image reconstruction module based on the decoded features and character recognition intermediate features.

Benefits of technology

The character recognition intermediate features with complete text information are fused on the reconstructed image to complete the background text information of the image filtered out during the decoding process, and improve the text information quality of the reconstructed image.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119942559A_ABST
    Figure CN119942559A_ABST
Patent Text Reader

Abstract

The invention discloses a training method and device of an image character reconstruction model, a medium and electronic equipment. The method comprises the following steps: acquiring a preprocessed image of the image character reconstruction model; inputting the preprocessed image into a codec model to be jointly trained to obtain adapter output features and character recognition intermediate features; the codec model to be jointly trained can realize improvement of image compression; training a codec model to be subjected to joint training based on a first joint loss function of the adapter output feature and the character recognition intermediate feature and an original training loss function of the codec; obtaining decoding characteristics of the trained codec; and image reconstruction training is performed on the image reconstruction module based on the decoding features and the character recognition intermediate features, so that the character recognition intermediate features with complete text information can be fused on the reconstructed image, image background text information filtered in the decoding process on the reconstructed image can be complemented, and the text information quality of the reconstructed image can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of image intelligent processing, and in particular, relates to a training method, device, medium and electronic equipment for an image character reconstruction model. Background Art

[0002] The current goal of image coding based on machine learning is to further improve the compression rate of images. Its compressed feature representation can be decoded into specific pixel values ​​for human viewing, or directly used as input for image processing or computer vision models based on neural networks. However, image codecs based on intelligent tasks may cause the loss of features that the human eye pays attention to. For example, image codecs based on target detection tasks do not pay attention to license plate target information. During the training process, they will filter out the background license plate text information of the image, resulting in the decoded features no longer containing license plate text information, which in turn causes the license plate text in the subsequent image reconstruction to be blurred and difficult to identify.

[0003] It should be noted that the information disclosed in the above background technology section is only used to enhance the understanding of the background of the present application, and therefore may include information that does not constitute the prior art known to ordinary technicians in the field. Summary of the invention

[0004] The purpose of the present invention is to provide a training method, device, medium and electronic equipment for an image character reconstruction model, so as to solve the problem that characters in reconstructed images are difficult to recognize in the related art.

[0005] According to one aspect of an embodiment of the present application, a method for training an image character reconstruction model is provided, the method comprising:

[0006] Obtaining a preprocessed image of an image character reconstruction model;

[0007] Inputting the preprocessed image into the codec model to be jointly trained to obtain adapter output features and character recognition intermediate features; the codec model to be jointly trained includes a codec, an adapter network and an optical character recognition model;

[0008] Training the jointly trained codec model based on a first joint loss function of the adapter output features and the character recognition intermediate features, and an original training loss function of the codec;

[0009] Get the decoded features of the trained codec;

[0010] The image reconstruction module is trained for image reconstruction based on the decoding features and the intermediate features of character recognition.

[0011] In some embodiments, a preprocessed image is input into a codec model to be jointly trained to obtain adapter output features and character recognition intermediate features, including: inputting the preprocessed image into the codec to obtain the decoding features of the codec; inputting the preprocessed image into an optical character recognition model to obtain the character recognition intermediate features of the optical character recognition model; inputting the decoding features into an adapter network to obtain adapter output features of the adapter network; the dimension of the adapter output features is the same as the dimension of the character recognition intermediate features.

[0012] In some embodiments, based on a first joint loss function of adapter output features and character recognition intermediate features, and an original training loss function of the codec, a codec model to be jointly trained is trained, including: determining the first loss function of the codec model to be jointly trained based on the first joint loss function and its first weight, the original training loss function of the codec and its second weight; adjusting the first weight and the second weight until the first loss function meets the first target value.

[0013] In some embodiments, an image reconstruction module is trained for image reconstruction based on decoding features and character recognition intermediate features, including: inputting the decoding features into an image reconstruction model to obtain a reconstructed image; inputting the reconstructed image into an optical character recognition model to obtain reconstructed image character features of the reconstructed image; and training the image reconstruction model based on a second joint loss function of the reconstructed image character features and the character recognition intermediate features, and the original training loss function of the image reconstruction model.

[0014] In some embodiments, the image reconstruction model is trained based on a second joint loss function of reconstructed image character features and character recognition intermediate features, and an original training loss function of the image reconstruction model, including: determining the second loss function of the image reconstruction model based on the second joint loss function and its third weight, the original training loss function of the image reconstruction model and its fourth weight; adjusting the third weight and the fourth weight until the second loss function meets the second target value.

[0015] In some embodiments, obtaining a preprocessed image of an image-character reconstruction model includes: obtaining a training image of the image-character reconstruction model; inputting the training image into a codec in a codec model to be jointly trained to perform image preprocessing to obtain a preprocessed image; the image preprocessing includes at least one of image normalization and image padding.

[0016] In some embodiments, the method provided by the present application also includes: building a network structure of a codec based on a business scenario of image character reconstruction; selecting an optical character recognition model that matches the codec structure; finding a segmentation point of the network layer of the optical character recognition model according to the multiple of the codec feature downsampling; dividing the network layer of the optical character recognition model based on the segmentation point to obtain a first network module before the segmentation point and a second network module after the segmentation point; the first network module is a network module in the optical character recognition model that processes preprocessed images.

[0017] According to one aspect of an embodiment of the present application, a training device for an image character reconstruction model is provided, the device comprising:

[0018] An image acquisition module, used for acquiring a preprocessed image of an image character reconstruction model;

[0019] A first feature output module is used to input the preprocessed image into the codec model to be jointly trained to obtain the adapter output feature and the character recognition intermediate feature; the codec model to be jointly trained includes a codec, an adapter network and an optical character recognition model;

[0020] A first joint training module, configured to train the codec model to be jointly trained based on a first joint loss function of the adapter output features and the character recognition intermediate features, and an original training loss function of the codec;

[0021] A feature acquisition module is used to obtain the decoding features of the trained codec;

[0022] The second joint training module is used to perform image reconstruction training on the image reconstruction module based on the decoding features and the character recognition intermediate features.

[0023] According to one aspect of an embodiment of the present application, a computer medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the training method of the image character reconstruction model provided by any embodiment of the present application is implemented.

[0024] According to one aspect of an embodiment of the present application, an electronic device is provided, comprising: a processor; a memory for storing executable instructions of the processor; the processor executes the executable instructions to enable the electronic device to implement the training method of the image character reconstruction model provided by any embodiment of the present application.

[0025] In the technical solution of the present application, a preprocessed image of an image-character reconstruction model is obtained; the preprocessed image is input into a codec model to be jointly trained to obtain adapter output features and character recognition intermediate features; the codec model to be jointly trained can achieve improved image compression; the codec model to be jointly trained is trained based on a first joint loss function of the adapter output features and the character recognition intermediate features, and the original training loss function of the codec; the decoding features of the trained codec are obtained to reduce the error of the image reconstruction result caused by the codec; finally, the image reconstruction module is trained based on the decoding features and the character recognition intermediate features, so that the character recognition intermediate features with complete text information can be integrated on the reconstructed image, the image background text information filtered out during the decoding process on the reconstructed image is completed, and the text information quality of the reconstructed image is improved.

[0026] It should be understood that the above general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] The drawings herein are incorporated into the specification and constitute a part of the specification, illustrate embodiments consistent with the present application, and together with the specification are used to explain the principles of the present application. Obviously, the drawings described below are only some embodiments of the present application, and for ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0028] Figure 1 The flowchart of the training method of the image character reconstruction model provided in one embodiment of the present application is schematically shown.

[0029] Figure 2 The schematic diagram shows the structure of a codec provided by an embodiment of the present application.

[0030] Figure 3 A schematic diagram of the training process of a codec model to be jointly trained provided in an embodiment of the present application is shown schematically.

[0031] Figure 4 The following is a schematic diagram schematically showing a basic unit structure of an adapter network provided by an embodiment of the present application.

[0032] Figure 5 The following is a schematic diagram showing the training process of the image reconstruction model provided in one embodiment of the present application.

[0033] Figure 6 The structural diagram of the training device for the image character reconstruction model provided by one embodiment of the present application is schematically shown.

[0034] Figure 7 The structural block diagram of an electronic device provided by an embodiment of the present application is schematically shown.

[0035] Figure 8 The structure block diagram of a computer system for implementing an electronic device according to an embodiment of the present application is schematically shown. DETAILED DESCRIPTION

[0036] Example embodiments will now be described more fully with reference to the accompanying drawings. However, example embodiments can be implemented in a variety of forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided so that this application will be more comprehensive and complete and fully convey the concept of the example embodiments to those skilled in the art.

[0037] In addition, the features, structures or characteristics described in the present application may be combined in one or more embodiments in any suitable manner. In the following description, many specific details are provided to provide a full understanding of the embodiments of the present application. However, those skilled in the art will appreciate that the technical scheme of the present application can be put into practice without one or more of the specific details, or other methods, components, devices, steps, etc. may be adopted. In other cases, known methods, devices, realizations or operations are not shown or described in detail to avoid blurring the various aspects of the present application.

[0038] The block diagrams shown in the accompanying drawings are merely functional entities and do not necessarily correspond to physically independent entities. That is, these functional entities may be implemented in software form, or in one or more hardware modules or integrated circuits, or in different networks and / or processor devices and / or microcontroller devices.

[0039] The flowcharts shown in the accompanying drawings are only exemplary and do not necessarily include all the contents and operations / steps, nor must they be executed in the order described. For example, some operations / steps can be decomposed, and some operations / steps can be combined or partially combined, so the actual execution order may change according to actual conditions.

[0040] like Figure 1 As shown, the present application provides a training method for an image character reconstruction model, the method includes S110 to S150, and the specific process is as follows.

[0041] S110, obtaining a preprocessed image of an image character reconstruction model.

[0042] Specifically, the image character reconstruction model is used to encode the image to be reconstructed, and decode and reconstruct the encoded image, and finally output the reconstructed image; the method provided in this application is applied to the image character reconstruction model to train the image encoding and decoding and image reconstruction process of the image character reconstruction model, so that the reconstructed image not only meets the requirements of high image compression rate, but also does not lose the information that the human eye pays attention to in the image. Figure 2 As shown, after receiving the input image, the image character reconstruction model starts to perform image encoding and decoding processing, and then reconstructs the image based on the image decoding features, and finally obtains the reconstructed image. The preprocessed image is the image obtained after preprocessing the training sample image.

[0043] In some embodiments, obtaining a preprocessed image of an image-character reconstruction model includes: obtaining a training image of the image-character reconstruction model; inputting the training image into a codec in a codec model to be jointly trained to perform image preprocessing to obtain a preprocessed image.

[0044] Specifically, the training sample images can be images prepared in advance in various scenes, and these images have character information, and these character information can be license plate numbers, vehicle logos, road sign text, etc. It should be understood that the real character information in the training image is used as the training teaching value of the image character reconstruction model, and can constrain the training results of the image character reconstruction model in the subsequent training process. Image preprocessing includes at least one of image normalization and image padding. Usually, image normalization refers to a process of performing a series of standard processing transformations on an image to transform it into a fixed standard form. The standard image is called a normalized image. Image padding refers to padding 0 in the right and bottom directions of the image until the image side length is aligned to an integer multiple of 2. Suppose the dimension of an image is [3, h, w], where 3 means that the image is a three-channel color image. The codec usually normalizes the original image and fills h (the height of the image) and w (the width of the image) to an integer multiple of 2 for preprocessing. Suppose the dimension of the preprocessed image is [3, new_h, new_w]. Image preprocessing can also include scaling and normalization, where scaling refers to scaling the short side of the training sample image to x min , the long side is scaled according to the original image ratio. If the long side exceeds the maximum side length threshold x after scaling max , then the long side is scaled to x max , the short side is scaled according to the original image ratio. min and x max is the image short side scaling length and the image scaling side length threshold. Standardization refers to performing a standardization operation on the image. For example, a standardization method uses the following formula 1:

[0045] Value pixel =(Value pixel -pixelmean ) / pixel std Formula 1

[0046] Among them, pixel mean and pixel std is a given constant, pixel in BGR channel order mean =[103.530, 116.280, 123.675], pixel std =[57.375, 57.120, 58.395].

[0047] In some embodiments, the method provided by the present application also includes: building a network structure of a codec based on a business scenario of image character reconstruction; selecting an optical character recognition model that matches the codec structure; finding a segmentation point of the network layer of the optical character recognition model according to the multiple of the codec feature downsampling; dividing the network layer of the optical character recognition model based on the segmentation point to obtain a first network module before the segmentation point and a second network module after the segmentation point; the first network module is a network module in the optical character recognition model that processes preprocessed images.

[0048] Specifically, the network structure of the codec will affect the results of image character reconstruction. Therefore, according to different business scenarios of image character reconstruction, the corresponding codec network results are built to encode and decode the image and reconstruct the image in a targeted manner. Common machine learning image codecs based on intelligent tasks include DCM codecs. Optical character recognition (OCR) is a technology that converts text in an image into a machine-readable format. This technology scans documents or images, and then uses complex algorithms to analyze and identify the characters in them, and converts them into machine-readable formats such as ASCII or Unicode, thereby achieving the purpose of electronic storage, editing and searching of text. Representative models of optical character recognition models include easyocr and PaddleOCR. The model structure can specifically include elements such as the various layers, nodes, weights and connections of the model, as well as the layout and organization of these elements. The structure of the optical character recognition model matches the codec structure, that is, the layout and organization of the above elements match. Assuming the downsampling multiple is n, find the segmentation point based on the OCR model, that is, locate the layer where the image is downsampled n times in the OCR network model, and divide the network into two. The network layer before the segmentation point is the first network module ocr head, and the network layer after the segmentation point is the second network module ocr_body. The dimension of the output feature feature2 of the ocr_head module is [channel1, height / n of the input image, width / n of the input image].

[0049] S120, inputting the preprocessed image into the codec model to be jointly trained to obtain adapter output features and character recognition intermediate features.

[0050] Specifically, Figure 3 As shown, the codec model to be jointly trained includes a codec, an adapter network and an optical character recognition model. In some embodiments, S120 specifically includes: inputting the preprocessed image into the codec to obtain the decoding features of the codec; inputting the preprocessed image into the optical character recognition model to obtain the character recognition intermediate features of the optical character recognition model; inputting the decoding features into the adapter network to obtain the adapter output features of the adapter network; the dimension of the adapter output features is the same as the dimension of the character recognition intermediate features. Specifically, after the preprocessed image is subjected to the codec to extract the decoding features, the h and w of the decoding features are integer multiples of 2. The downsampling multiple is set to n, that is, the dimension of the decoding features obtained after the original image is preprocessed, compressed and then decoded is [channel, new_h / n, new_w / n], and channel is the number of channels of the decoding features, which is generally an integer multiple of 2. The adapter network is used to adjust the dimension of the decoding features obtained by decoding the image to a dimension consistent with the OCR intermediate features. As Figure 4 The figure shows the basic unit structure of the adapter network, which includes a convolution layer (Conv) and a relu activation layer. Inputting the preprocessed image into the optical character recognition model actually inputs the preprocessed image into the first network module ocr_head in the optical character recognition model to obtain the intermediate features of character recognition. The dimension of the output feature feature2 (i.e., the intermediate feature of character recognition) of the first network module ocr_head is [channel1, height / n of the input image, width / n of the input image]. After the adapter network processes the decoded features, the dimension of the adapter output feature is the same as the dimension of the above-mentioned intermediate feature of character recognition.

[0051] S130, training the codec model to be jointly trained based on the first joint loss function of the adapter output features and the character recognition intermediate features, and the original training loss function of the codec.

[0052] Specifically, in the related art, the decoded features output by the codec may have lost most of the character information. By fusing the intermediate features of character recognition into the decoded features, the decoded features can be provided with text-related information. The decoded features output by the codec only change the feature dimensions through the adapter, and do not change or optimize the character information in the decoded features. Therefore, the adapter output features and the decoded features can be understood as different structural forms of the same features. When the codec model to be jointly trained is trained, the intermediate features of character recognition are fused on the basis of the decoded features, which is actually the adapter output features fused with the intermediate features of character recognition. The loss function is used to measure the degree of deviation between the predicted value (i.e., the output result of the model) made by the model and the ground truth. In simple terms, the loss function represents the error between the predicted value and the ground truth. When the predicted value is infinitely close to the ground truth, the difference between the two can be minimized, i.e., the loss function is the minimum. There are many types of loss functions, such as mean squared error (MSE), mean absolute error (MAE), or other types of loss functions. In this embodiment, the first joint loss function is used to represent the error between the predicted value after the decoding feature is integrated with the character recognition intermediate feature and the true value after image preprocessing. The original training loss function of the codec represents the error caused by the network structure of the codec itself. When the sum of the first joint loss function and the original training loss function of the codec is minimized, a codec model with optimal training effect can be obtained.

[0053] In some embodiments, based on a first joint loss function of adapter output features and character recognition intermediate features, and an original training loss function of the codec, a codec model to be jointly trained is trained, including: determining the first loss function of the codec model to be jointly trained based on the first joint loss function and its first weight, the original training loss function of the codec and its second weight; adjusting the first weight and the second weight until the first loss function meets the first target value.

[0054] Specifically, the proportion of errors caused by different reasons in the total error of the codec model is not the same. This embodiment introduces weights to adjust the influence of different errors on the codec model, so as to make the trained codec model more accurate. MSE1 represents the first joint loss function, w1 represents the first weight, Loss 编解 represents the original training loss function of the codec, w2 represents the second weight, then the first loss function Loss1 = w1loss MSE1 +w2Loss 编解The first target value is used to represent the minimum value of the first loss function. It should be understood that the minimum value here refers to the smallest possible value in the business scenario where the codec model to be jointly trained is located, not an absolute minimum value. When the first loss function value reaches the first target value, the codec model to be jointly trained completes training.

[0055] S140: Obtain decoding features of the trained codec.

[0056] Specifically, the trained codec refers to the codec model training step to be jointly trained in S130. By obtaining the decoding features output by the trained codec, the problem of large image reconstruction errors caused by changes in codec parameters can be improved.

[0057] S150, performing image reconstruction training on an image reconstruction module based on the decoding features and the character recognition intermediate features.

[0058] Specifically, after data processing by the codec, character information in the decoded features may have been missing. By combining the decoded features with character recognition intermediate features, the missing character information in the decoded features can be supplemented, so that the final reconstructed image has both a high compression rate and character information that the human eye pays attention to.

[0059] In the technical solution of the present application, a preprocessed image of an image-character reconstruction model is obtained; the preprocessed image is input into a codec model to be jointly trained to obtain adapter output features and character recognition intermediate features; the codec model to be jointly trained can achieve improved image compression; the codec model to be jointly trained is trained based on a first joint loss function of the adapter output features and the character recognition intermediate features, and the original training loss function of the codec; the decoding features of the trained codec are obtained to reduce the error of the image reconstruction result caused by the codec; finally, the image reconstruction module is trained based on the decoding features and the character recognition intermediate features, so that the character recognition intermediate features with complete text information can be integrated on the reconstructed image, the image background text information filtered out during the decoding process on the reconstructed image is completed, and the text information quality of the reconstructed image is improved.

[0060] In some embodiments, an image reconstruction module is trained for image reconstruction based on decoding features and character recognition intermediate features, including: inputting the decoding features into an image reconstruction model to obtain a reconstructed image; inputting the reconstructed image into an optical character recognition model to obtain reconstructed image character features of the reconstructed image; and training the image reconstruction model based on a second joint loss function of the reconstructed image character features and the character recognition intermediate features, and the original training loss function of the image reconstruction model.

[0061] Specifically, Figure 5As shown, the decoded features output by the trained codec are input into the image reconstruction model to obtain the reconstructed image. Although the decoded features output by the trained codec have text information, a lot of text information is lost in the compression of the image during the encoding and decoding process, and the text information in the reconstructed image is still unclear and unrecognizable. Therefore, the reconstructed image is input into the optical character recognition model, and the text information in the reconstructed image is recognized by the optical character recognition model to obtain the character features of the reconstructed image. Then the image preprocessed by the original image is also input into the optical character recognition model separately to obtain the text features of the preprocessed image, because the image preprocessing only processes the structure or size of the image, etc., and does not lose the text information. In other words, the text features of the preprocessed image can be taken as the true value, and the character features of the reconstructed image can be taken as the predicted value, then the second joint loss function represents the degree of deviation between the text features of the preprocessed image and the character features of the reconstructed image. When the second joint loss function reaches the minimum value, it means that the error between the character features of the reconstructed image and the text features of the preprocessed image is the smallest. In addition, limited by the structure of the image reconstruction model itself, the image reconstruction model will also cause an error between the reconstructed image and the true image. Then, by combining the second joint loss function and the original training loss function of the image reconstruction model, we can get the second loss function of the image reconstruction model. When the second loss function reaches the minimum value, it means that the image reconstruction model has completed training.

[0062] In some embodiments, the image reconstruction model is trained based on a second joint loss function of reconstructed image character features and character recognition intermediate features, and an original training loss function of the image reconstruction model, including: determining the second loss function of the image reconstruction model based on the second joint loss function and its third weight, the original training loss function of the image reconstruction model and its fourth weight; adjusting the third weight and the fourth weight until the second loss function meets the second target value.

[0063] Specifically, the proportion of errors due to different reasons in the total error of the image reconstruction model is not the same. This embodiment introduces weights to adjust the influence of different errors on the training results of the image reconstruction model, so as to make the trained image reconstruction model more accurate. MSE2 represents the second joint loss function, w3 represents the third weight, Loss 重建 represents the original training loss function of the codec, w4 represents the fourth weight, then the second loss function Loss2 = w3loss MSE2 +w4Loss 重建 The second target value is used to represent the minimum value of the second loss function. It should be understood that the minimum value here refers to the smallest possible value in the business scenario of the image reconstruction model, not the absolute minimum value. When the second loss function value reaches the second target value, the image reconstruction model completes training.

[0064] The following describes the device embodiments of the present application. Figure 6 As shown, the training device for the image character reconstruction model provided by the present application includes the following modules.

[0065] The image acquisition module 610 is used to acquire a pre-processed image of the image character reconstruction model.

[0066] The first feature output module 620 is used to input the preprocessed image into the codec model to be jointly trained to obtain the adapter output feature and the character recognition intermediate feature; the codec model to be jointly trained includes the codec, the adapter network and the optical character recognition model

[0067] The first joint training module 630 is used to train the codec model to be jointly trained based on the first joint loss function of the adapter output features and the character recognition intermediate features, and the original training loss function of the codec.

[0068] The feature acquisition module 640 is used to acquire the decoding features of the trained codec.

[0069] The second joint training module 650 is used to perform image reconstruction training on the image reconstruction module based on the decoding features and the character recognition intermediate features.

[0070] In some embodiments, the first feature output module 620 includes: a decoding unit, used to input the preprocessed image into the codec to obtain the decoding features of the codec; a first character recognition unit, used to input the preprocessed image into the optical character recognition model to obtain the character recognition intermediate features of the optical character recognition model; an adapter unit, used to input the decoding features into the adapter network to obtain the adapter output features of the adapter network; the dimension of the adapter output features is the same as the dimension of the character recognition intermediate features.

[0071] In some embodiments, the first joint training module 630 includes: a first error determination unit, used to determine the first loss function of the codec model to be jointly trained based on the first joint loss function and its first weight, the original training loss function of the codec and its second weight; a first weight adjustment unit, used to adjust the first weight and the second weight until the first loss function meets the first target value.

[0072] In some embodiments, the second joint training module 650 includes: a reconstruction result unit, used to input the decoded features into the image reconstruction model to obtain a reconstructed image; a reconstruction character recognition unit, used to input the reconstructed image into the optical character recognition model to obtain the reconstructed image character features of the reconstructed image; a second joint training unit, used to train the image reconstruction model based on a second joint loss function of the reconstructed image character features and the character recognition intermediate features, and the original training loss function of the image reconstruction model.

[0073] In some embodiments, the second joint training unit is also used to determine the second loss function of the image reconstruction model based on the second joint loss function and its third weight, the original training loss function of the image reconstruction model and its fourth weight; and adjust the third weight and the fourth weight until the second loss function meets the second target value.

[0074] In some embodiments, the image acquisition module 610 includes: a training image acquisition unit, used to acquire a training image of an image character reconstruction model; a preprocessing unit, used to input the training image into the codec in the codec model to be jointly trained to perform image preprocessing to obtain a preprocessed image; image preprocessing includes at least one of image normalization and image padding.

[0075] In some embodiments, the device provided by the present application also includes: a codec model building module, which is used to build a codec network structure based on a business scenario of image character reconstruction; select an optical character recognition model that matches the codec structure; find the segmentation point of the optical character recognition model network layer according to the multiple of codec feature downsampling; divide the network layer of the optical character recognition model based on the segmentation point to obtain a first network module before the segmentation point and a second network module after the segmentation point; the first network module is a network module in the optical character recognition model that processes preprocessed images.

[0076] It should be noted that the specific implementation contents of the device embodiments in the present application have been explained in detail in the corresponding method embodiments and will not be repeated here.

[0077] The electronic equipment of this application is introduced below, such as Figure 7 As shown, the present application provides an electronic device 700, which includes: a processor 710 and a memory 720, the memory 720 is used to store executable instructions of the processor; the processor 710 executes the executable instructions to enable the electronic device to implement the training method of the image character reconstruction model provided by any embodiment of the present application.

[0078] Specifically, the training method of the image character reconstruction model provided by the present application is stored in the memory 720 of the electronic device, and the training method of the image character reconstruction model is executed by the processor 710 to integrate the character recognition intermediate features with complete text information on the reconstructed image, to complete the image background text information on the reconstructed image that is filtered out during the decoding process, and to improve the text information quality of the reconstructed image.

[0079] It should be known that the specific implementation content of the electronic device in the present application has been explained in detail in the corresponding method embodiment and will not be repeated here.

[0080] Figure 8The structure block diagram of a computer system for implementing an electronic device according to an embodiment of the present application is schematically shown.

[0081] It should be noted that Figure 8 The computer system 800 of the electronic device shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present application.

[0082] like Figure 8 As shown, the computer system 800 includes a processor 801, which can be a CPU (Central Processing Unit) or an MCU (Microcontroller Unit). The processor 801 can perform various appropriate actions and processes according to the program stored in the read-only memory 802 (Read-Only Memory, ROM) or the program loaded from the storage part 808 to the random access memory 803 (Random Access Memory, RAM). In the random access memory 803, various programs and data required for system operation are also stored. The processor 801, the read-only memory 802 and the random access memory 803 are connected to each other through a bus 804. The input / output interface 805 (Input / Output interface, i.e., I / O interface) is also connected to the bus 804.

[0083] The following components are connected to the input / output interface 805: an input section 806 including a keyboard, a mouse, etc.; an output section 807 including a cathode ray tube (CRT), a liquid crystal display (LCD), etc., and a speaker; a storage section 808 including a hard disk, etc.; and a communication section 809 including a network interface card such as a local area network card, a modem, etc. The communication section 809 performs communication processing via a network such as the Internet. A drive 810 is also connected to the input / output interface 805 as needed. A removable medium 811, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 810 as needed so that a computer program read therefrom is installed into the storage section 808 as needed.

[0084] In particular, according to an embodiment of the present application, the processes described in the various method flow charts can be implemented as computer software programs. For example, an embodiment of the present application includes a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program contains program code for executing the method shown in the flow chart. In such an embodiment, the computer program can be downloaded and installed from a network through a communication section 809, and / or installed from a removable medium 811. When the computer program is executed by the processor 801, various functions defined in the system of the present application are executed.

[0085] It should be noted that the computer-readable medium shown in the embodiment of the present application may be a computer-readable signal medium or a computer-readable storage medium or any combination of the above two. The computer-readable storage medium may be, for example, but not limited to, a system, device or device of electricity, magnetism, light, electromagnetic, infrared, or semiconductor, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM), a flash memory, an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present application, a computer-readable storage medium may be any tangible medium containing or storing a program, which may be used by an instruction execution system, device or device or used in combination with it. In the present application, a computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, wherein a computer-readable program code is carried. Such propagated data signals may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. Computer readable signal media may also be any computer readable medium other than computer readable storage media, which may send, propagate, or transmit programs for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer readable medium may be transmitted using any suitable medium, including but not limited to: wireless, wired, etc., or any suitable combination of the above.

[0086] The flow chart and block diagram in the accompanying drawings illustrate the possible architecture, function and operation of the system, method and computer program product according to various embodiments of the present application. In this regard, each box in the flow chart or block diagram can represent a module, a program segment or a part of a code, and the above-mentioned module, program segment or a part of a code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order from the order marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram or flow chart, and the combination of the boxes in the block diagram or flow chart can be implemented with a dedicated hardware-based system that performs a specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.

[0087] It should be noted that, although several modules or units of the equipment for action execution are mentioned in the above detailed description, this division is not mandatory. In fact, according to the embodiments of the present application, the features and functions of two or more modules or units described above can be embodied in one module or unit. On the contrary, the features and functions of one module or unit described above can be further divided into being embodied by multiple modules or units.

[0088] Through the description of the above implementation methods, it is easy for those skilled in the art to understand that the example implementation methods described here can be implemented by software, or by software combined with necessary hardware. Therefore, the technical solution according to the implementation method of the present application can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, including several instructions to enable a computing device (which can be a personal computer, a server, a touch terminal, or a network device, etc.) to execute the implementation method according to the present application.

[0089] Those skilled in the art will readily appreciate other embodiments of the present application after considering the specification and practicing the invention disclosed herein. The present application is intended to cover any modification, use or adaptation of the present application, which follows the general principles of the present application and includes common knowledge or customary technical means in the art that are not disclosed in the present application.

[0090] It should be understood that the present application is not limited to the precise structures that have been described above and shown in the drawings, and that various modifications and changes may be made without departing from the scope thereof. The scope of the present application is limited only by the appended claims.

Claims

1. A training method for an image character reconstruction model, characterized in that: The method comprises: Obtaining a preprocessed image of an image character reconstruction model; Inputting the preprocessed image into a codec model to be jointly trained to obtain adapter output features and character recognition intermediate features; the codec model to be jointly trained includes a codec, an adapter network and an optical character recognition model; Training the codec model to be jointly trained based on a first joint loss function of the adapter output feature and the character recognition intermediate feature, and an original training loss function of the codec; Obtaining the trained decoding features of the codec; The image reconstruction module is trained for image reconstruction based on the decoding features and the character recognition intermediate features.

2. The training method for image character reconstruction model according to claim 1, characterized in that: The step of inputting the pre-processed image into the codec model to be jointly trained to obtain adapter output features and character recognition intermediate features includes: Inputting the preprocessed image into the codec to obtain a decoding feature of the codec; Inputting the preprocessed image into the optical character recognition model to obtain character recognition intermediate features of the optical character recognition model; The decoded features are input into the adapter network to obtain adapter output features of the adapter network; the dimension of the adapter output features is the same as the dimension of the character recognition intermediate features.

3. The training method for image character reconstruction model according to claim 1, characterized in that: The first joint loss function based on the adapter output feature and the character recognition intermediate feature, and the original training loss function of the codec, training the codec model to be jointly trained, comprises: Determine a first loss function of the codec model to be jointly trained based on the first joint loss function and its first weight, the original training loss function of the codec and its second weight; The first weight and the second weight are adjusted until the first loss function meets a first target value.

4. The training method for image character reconstruction model according to claim 1, characterized in that: The performing image reconstruction training on the image reconstruction module based on the decoding feature and the character recognition intermediate feature comprises: Inputting the decoded features into the image reconstruction model to obtain a reconstructed image; Inputting the reconstructed image into the optical character recognition model to obtain reconstructed image character features of the reconstructed image; The image reconstruction model is trained based on a second joint loss function of the reconstructed image character features and the character recognition intermediate features, and an original training loss function of the image reconstruction model.

5. The training method for image character reconstruction model according to claim 4, characterized in that: The second joint loss function based on the reconstructed image character features and the character recognition intermediate features, and the original training loss function of the image reconstruction model, is used to train the image reconstruction model, including: Determine a second loss function of the image reconstruction model based on the second joint loss function and its third weight, the original training loss function of the image reconstruction model and its fourth weight; The third weight and the fourth weight are adjusted until the second loss function meets the second target value.

6. The training method for image character reconstruction model according to claim 1, characterized in that: The step of obtaining a preprocessed image of an image character reconstruction model comprises: Obtaining a training image for an image-character reconstruction model; The training image is input into the codec in the codec model to be jointly trained for image preprocessing to obtain the preprocessed image; the image preprocessing includes at least one of image normalization and image padding.

7. The training method for image character reconstruction model according to claim 1, characterized in that: The method further comprises: Build the codec network structure based on the business scenario of image character reconstruction; Selecting an optical character recognition model that matches the codec structure; Finding the segmentation point of the network layer of the optical character recognition model according to the multiple of the downsampling of the codec feature; The network layer of the optical character recognition model is divided based on the segmentation point to obtain a first network module before the segmentation point and a second network module after the segmentation point; the first network module is a network module in the optical character recognition model that processes the preprocessed image.

8. A training device for an image character reconstruction model, characterized in that: The device comprises: An image acquisition module, used for acquiring a preprocessed image of an image character reconstruction model; A first feature output module, used for inputting the pre-processed image into a codec model to be jointly trained to obtain adapter output features and character recognition intermediate features; the codec model to be jointly trained includes a codec, an adapter network and an optical character recognition model; A first joint training module, configured to train the codec model to be jointly trained based on a first joint loss function of the adapter output feature and the character recognition intermediate feature, and an original training loss function of the codec; A feature acquisition module, used to acquire the decoding features of the trained codec; The second joint training module is used to perform image reconstruction training on the image reconstruction module based on the decoding features and the character recognition intermediate features.

9. A computer readable medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the training method for the image character reconstruction model described in any one of claims 1 to 7 is implemented.

10. An electronic device, characterized in that: include: processor; A memory, configured to store executable instructions of the processor; The processor executes the executable instructions to enable the electronic device to implement the training method for the image character reconstruction model as described in any one of claims 1 to 7.