Training Method, Device, Equipment and Storage Medium of Digital Recognition Model

By cropping and data augmenting the sample images, generating multiple training images, and iteratively training combined with loss function values ​​and similarity, the problems of high training costs and slow speed in the existing technology are solved, and more efficient digital recognition model training is achieved.

CN114417992BActive Publication Date: 2025-05-27PING AN TECH (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210044201.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-01-14
Publication Date
2025-05-27
Estimated Expiration
2042-01-14

AI Technical Summary

Technical Problem

Training of existing digital recognition models requires a large amount of training data, resulting in higher training costs and slower training speed.

Method used

By cropping and data augmenting the sample images, multiple training images are generated, and these images are input into the neural network separately, the loss function value and similarity are calculated, and iterative training is performed until the neural network converges.

Benefits of technology

The expansion of training samples is achieved, the convergence speed of the neural network is accelerated, and the training speed of the model is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114417992B_ABST
    Figure CN114417992B_ABST
Patent Text Reader

Abstract

This application relates to the fields of artificial intelligence and image recognition, and specifically discloses a training method, device, equipment and storage medium for a digital recognition model. The method includes: obtaining a sample image and a digital label corresponding to the sample image; performing image cropping on the sample image, and using the remaining image after image cropping as a first training image; performing data augmentation on the sample image to obtain a second training image; respectively inputting the first training image and the second training image into a neural network, obtaining a first output value corresponding to the first training image and a second output value corresponding to the second training image, and calculating a loss function value of the neural network and a similarity between the first output value and the second output value according to the digital label; performing iterative training on the neural network according to the loss function value and the similarity, and when the neural network converges, using the neural network as a digital recognition model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence, and particularly to a training method, device, equipment and storage medium for a digital recognition model. Background Art

[0002] Currently, when performing digital recognition, most often a deep neural network is trained to obtain a neural network model, and then the obtained neural network model is used to implement digital recognition. However, in order to ensure the accuracy of the trained classification model, a large amount of training data often needs to be obtained to participate in the model training, which makes the training cost relatively high. Summary of the Invention

[0003] This application provides a training method, device, equipment and storage medium for a digital recognition model to expand training samples and accelerate the training speed.

[0004] In a first aspect, this application provides a training method for a digital recognition model, the method including:

[0005] Obtain a sample image and the digital label corresponding to the sample image;

[0006] Perform image cropping on the sample image, and use the remaining image after image cropping as a first training image;

[0007] Perform data augmentation on the sample image to obtain a second training image;

[0008] Input the first training image and the second training image into a neural network respectively, obtain a first output value corresponding to the first training image and a second output value corresponding to the second training image, and calculate the loss function value of the neural network and the similarity between the first output value and the second output value according to the digital label;

[0009] Perform iterative training on the neural network according to the loss function value and the similarity, and when the neural network converges, use the neural network as a digital recognition model.

[0010] In a second aspect, this application also provides a training device for a digital recognition model, the device including:

[0011] A sample acquisition module, configured to obtain a sample image and the digital label corresponding to the sample image;

[0012] An image cropping module, configured to perform image cropping on the sample image, and use the remaining image after image cropping as a first training image;

[0013] A data augmentation module, configured to perform data augmentation on the sample image to obtain a second training image;

[0014] A loss calculation module, configured to input the first training image and the second training image into a neural network respectively, obtain a first output value corresponding to the first training image and a second output value corresponding to the second training image, and calculate a loss function value of the neural network and a similarity between the first output value and the second output value according to the digital label;

[0015] A model training module, configured to iteratively train the neural network according to the loss function value and the similarity, and when the neural network converges, use the neural network as a digital recognition model.

[0016] In a third aspect, the present application further provides a computer device, which includes a memory and a processor; the memory is used to store a computer program; the processor is configured to execute the computer program and implement the training method of the digital recognition model as described above when executing the computer program.

[0017] In a fourth aspect, the present application further provides a computer-readable storage medium, which stores a computer program, and when the computer program is executed by a processor, the processor is enabled to implement the training method of the digital recognition model as described above.

[0018] The present application discloses a training method, device, equipment and storage medium for a digital recognition model. By obtaining a sample image and a digital label corresponding to the sample image, then respectively performing image cropping and data augmentation on the sample image to obtain a first training image and a second training image, inputting the first training image and the second training image into a neural network respectively, and calculating a loss function value of the neural network and a similarity between the first training image and the second training image according to the digital label, finally training the neural network according to the loss function value and the similarity until the neural network converges to obtain a digital recognition model. Different methods are used to process the sample image to generate different training images to participate in the training of the neural network, realizing the expansion of training samples. In addition, the similarity between different training images is also added to the training of the neural network, accelerating the convergence speed of the neural network and improving the training speed of the model. Description of the Drawings

[0019] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings required for the description of the embodiments will be briefly introduced below. Obviously, the drawings in the following description are some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0020] Figure 1It is a schematic flowchart of a method for training a digital recognition model provided by an embodiment of the present application;

[0021] Figure 2 It is a schematic diagram of the steps for image cropping of a sample image provided by an embodiment of the present application;

[0022] Figure 3a It is a schematic diagram of a sample image in which the digital type is the first type provided by an embodiment of the present application;

[0023] Figure 3b It is a schematic diagram of a sample image in which the digital type is the second type provided by an embodiment of the present application;

[0024] Figure 4a It is a schematic diagram of image cropping of a sample image from both left and right ends provided by an embodiment of the present application;

[0025] Figure 4b It is a schematic diagram of image cropping of a sample image from both upper and lower ends provided by an embodiment of the present application;

[0026] Figure 5 It is a schematic flowchart of another method for training a digital recognition model provided by an embodiment of the present application;

[0027] Figure 6 It is a schematic block diagram of a device for training a digital recognition model provided by an embodiment of the present application;

[0028] Figure 7 It is a schematic block diagram of another device for training a digital recognition model provided by an embodiment of the present application;

[0029] Figure 8 It is a schematic block diagram of the structure of a computer device provided by an embodiment of the present application. Detailed implementation manners

[0030] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.

[0031] The flowcharts shown in the accompanying drawings are only illustrative examples, and do not necessarily include all the contents and operations / steps, nor do they necessarily need to be executed in the described order. For example, some operations / steps can also be decomposed, combined, or partially merged. Therefore, the actual execution order may be changed according to the actual situation.

[0032] It should be understood that the terms used in the specification of this application are only for the purpose of describing specific embodiments and are not intended to limit this application. As used in the specification of this application and the appended claims, unless the context clearly indicates otherwise, the singular forms "a", "an" and "the" are intended to include the plural forms.

[0033] It should also be understood that the term " / and" used in the specification of this application and the appended claims refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.

[0034] Embodiments of this application provide a training method, device, computer device and storage medium for a digital recognition model. The training method of the digital recognition model can be used for fraud insurance behaviors of patients and / or doctors, and provides an important reference for quickly identifying patients or doctors who commit fraud insurance.

[0035] The following will describe in detail some embodiments of this application with reference to the accompanying drawings. Without conflict, the following embodiments and the features in the embodiments can be combined with each other.

[0036] Please refer to Figure 1 , Figure 1 which is a schematic flowchart of a training method for a digital recognition model provided by an embodiment of this application. The training method of the digital recognition model achieves the purpose of expanding samples by performing different processes on the sample images.

[0037] As Figure 1 shown, the training method of the digital recognition model specifically includes: step S101 to step S105.

[0038] S101. Obtain a sample image and the digital label corresponding to the sample image.

[0039] Obtain a sample image for training the digital recognition model and the digital label corresponding to the sample image. The content of the sample image includes handwritten Roman numerals, and the digital label corresponding to the sample image represents the actual handwritten Roman numerals in the sample image. If the obtained sample image has no corresponding digital label, the sample image is labeled.

[0040] In some embodiments, before performing image cropping and data augmentation on the sample image, the sample image can be preprocessed first, and the preprocessing includes one or more processing methods such as binarization, denoising, normalization, and image thinning.

[0041] S102. Crop the sample image, and use the remaining image after the image cropping as the first training image.

[0042] After obtaining the sample image, the sample image can be cropped, that is, a specific area cutout is performed according to prior knowledge, and the remaining image after image cropping is used as the first training image. By randomly cropping the sample image, the neural network is guided to focus on more features and learn the information in the sample image more fully.

[0043] In one embodiment, please refer to Figure 2 , which is a schematic diagram of the steps for cropping the sample image. Step S102 may include step S1021 and step S1022.

[0044] S1021. Perform Hough transform and Sobel operator processing on the sample image to determine the digital type of the sample image.

[0045] Since during the process of randomly cropping the image, it is easy to change the digital category in the sample image. For example, cutting off the character "Ⅰ" on the right half of "Ⅵ" makes the picture become "Ⅴ". Therefore, in order to avoid randomly cropping from changing the numbers in the sample image, before performing image cropping, the sample image can be first processed by Hough Transform and Sobel operator.

[0046] Use the Hough transform to obtain the line feature map in the sample image, and use the Sobel operator to process to obtain the contour feature maps of the sample image in the horizontal and vertical directions. Based on the line feature map and the contour feature maps, the digital type of the sample image can be determined. Among them, the digital type of the sample image includes a first type and a second type. The first type can be a short and wide digital type that occupies more positions in the horizontal direction, such as Figure 3a shown, and the second type can be a tall and narrow digital type that occupies more positions in the vertical direction, such as Figure 3b shown.

[0047] S1022. Determine the image cropping method according to the digital type, and crop the sample image according to the image cropping method.

[0048] After determining the digital type, the corresponding image cropping method can be determined according to the digital type, so as to avoid changing the digital category in the sample image when cropping the sample image.

[0049] In one embodiment, determining the image cropping method according to the digital type includes: when the digital type is the first type, determining the image cropping method of the sample image as cropping at the left and right ends of the sample image; when the digital type is the second type, determining the image cropping method of the sample image as cropping at the upper and lower ends of the sample image.

[0050] If it is determined that the number type in the sample image is the first type, that is, a short and wide number type, then image cropping is performed from the left and right ends of the sample image, as Figure 4a shown. If it is determined that the number type in the sample image is the second type, that is, a tall and narrow number type, then image cropping is performed from the top and bottom ends of the sample image, as Figure 4b shown.

[0051] During the process of image cropping, the size of the rectangular frame for image cropping can be determined according to the length of the longest straight line in the sample image. The length of the longest straight line in the sample image can be calculated according to the Hough transform. When determining the size of the rectangular frame for image cropping, any multiple greater than 0 and not greater than 1 of the length of the longest straight line can be selected. For example, the size of the rectangular frame for image cropping can be 0.25 times the length of the longest straight line.

[0052] In addition, in one embodiment, the image cropping of the sample image and using the remaining image after image cropping as the first training image includes: performing image cropping on the sample image and performing data augmentation on the remaining image after image cropping to obtain a first image.

[0053] After cropping the sample image, data augmentation is performed on the remaining image after cropping to obtain a first image. Data augmentation can include at least one of transformation, rotation, and changing the color tone. In a specific implementation process, Augmix augmentation with a width of 1 and a depth of 2 can be used to perform data augmentation on the remaining image after cropping.

[0054] S103. Perform data augmentation on the sample image to obtain a second training image.

[0055] Among them, data augmentation can include various methods such as transformation, rotation, and changing the color tone. For example, Augmix augmentation with a width of 1 and a depth of 3 can be used to perform data augmentation on the sample image, and the image after data augmentation is used as the second training image.

[0056] S104. Input the first training image and the second training image into the neural network respectively to obtain a first output value corresponding to the first training image and a second output value corresponding to the second training image, and calculate the loss function value of the neural network and the similarity between the first output value and the second output value according to the digital label.

[0057] Input the first training image into the neural network to obtain a first output value of the neural network for the first training image, denoted as P M1 . Input the second training image into the neural network to obtain a second output value of the neural network for the second training image, denoted as P M2 .

[0058] The loss function of the neural network can adopt cross-entropy, based on the digital label corresponding to the sample image and the first output value P of the neural network for the first training image M1 Calculate a loss function value of the neural network; similarly, based on the digital label corresponding to the sample image and the second output value P of the neural network for the second training image M2 Calculate another loss function value of the neural network.

[0059] In addition, it is also necessary to calculate based on the first output value P of the neural network for the first training image M1 And the second output value P of the neural network for the second training image M2 Calculate P M1 And P M2 Calculate the similarity between them. The more similar P M1 And P M2 are, the better the prediction effect of the neural network indicates.

[0060] In the specific implementation process, the JS divergence loss can be used to calculate the similarity between P M1 And P M2 The smaller the calculated JS divergence loss value, the closer P M1 And P M2 are, and the better the prediction effect of the neural network.

[0061]

[0062] Among them, KL is the KL divergence.

[0063] S105. Iteratively train the neural network according to the loss function value and the similarity, and when the neural network converges, use the neural network as a digital recognition model.

[0064] Iteratively train the neural network according to the loss function value and the similarity. During the training process, the loss function value and the similarity can be given the same weight to participate in the iterative training of the neural network. That is, the loss function value and the similarity can be multiplied by their respective weights and then added together, and the obtained value is used as the final actual loss value. Based on this loss value, the parameters of the neural network are adjusted. When the loss value is the smallest, it is considered that the neural network converges at this time, and this converged neural network is used as the trained digital recognition model for recognizing handwritten Roman numerals.

[0065] Please refer to Figure 5 , Figure 5 which is a schematic flowchart of another method for training a digital recognition model provided by an embodiment of the present application.

[0066] As Figure 5As shown, the training method of the digital recognition model specifically includes: steps S201 to S207.

[0067] S201. Obtain a sample image and the digital label corresponding to the sample image.

[0068] Obtain a sample image for training the digital recognition model and the digital label corresponding to the sample image. The content of the sample image includes handwritten Roman numerals, and the digital label corresponding to the sample image represents the actual handwritten Roman numerals in the sample image. If the obtained sample image has no corresponding digital label, the sample image is labeled.

[0069] S202. Perform image cropping on the sample image, and use the remaining image after image cropping as the first training image.

[0070] After obtaining the sample image, image cropping can be performed on the sample image, that is, according to a specific region cutout of prior knowledge, and the remaining image after image cropping is used as the first training image. By randomly cropping the sample image, the neural network is guided to focus on more features and learn the information in the sample image more fully.

[0071] S203. Perform data augmentation on the sample image to obtain a second training image.

[0072] Among them, data augmentation can include at least one of transformation, rotation, and hue change. For example, Augmix augmentation with a width of 1 and a depth of 3 can be used to perform data augmentation on the sample image, and the image after data augmentation is used as the second training image.

[0073] S204. Determine the digital position in the sample image, determine a shear region at the digital position, and perform shearing on the shear region to obtain a shear region image and the remaining image after shearing.

[0074] Determine the digital position where the number is located in the sample image, and then determine the shear region according to the digital position, so that at least a part of the number is included in the sheared shear region image.

[0075] In the specific implementation process, the digital position in the sample image can be determined according to the pixel values of each pixel point in the sample image. According to the relationship between the pixel value of the pixel point in the sample image and the threshold, the position of the number in the sample image can be determined. For example, if the pixel value of the pixel point in the sample image is less than the threshold, it can be considered that the pixel point is a part of the number.

[0076] After determining the digital position, a shear region can be arbitrarily selected within the digital position, and the shear region is sheared to obtain the sheared image, that is, the shear region image, and the remaining image after shearing.

[0077] S205. Paste the image of the sheared area on the remaining image after shearing to obtain a third training image.

[0078] Then, randomly select any position on the remaining image after shearing to paste the sheared image of the sheared area, thereby obtaining a third training image.

[0079] During the pasting process, it is necessary to control that the characters in the cropped area image do not exceed the image range of the sample image. Therefore, in the specific implementation process, verification can be performed according to the sizes of the cropped area image and the sample image to ensure that the characters in the cropped area image do not exceed the picture range during pasting.

[0080] In one embodiment, the step of pasting the image of the sheared area on the remaining image after shearing to obtain a third training image includes: performing hole filling on the sheared area on the remaining image after shearing to obtain a filled image; pasting the image of the sheared area on the filled image to obtain a third training image.

[0081] After shearing the sample image, holes will appear at the digital positions in the remaining image after shearing. Therefore, it is necessary to perform hole filling on the holes generated by shearing. In the specific implementation process, the inpainting method can be used for hole filling.

[0082] After completing the hole filling, a completely filled image is obtained. Then, randomly select any position on the filled image to paste the sheared image of the sheared area, and the pasted image is the third training image.

[0083] In one embodiment, the step of pasting the image of the sheared area on the remaining image after shearing to obtain a third training image includes: obtaining the pasting position of the image of the sheared area; determining whether the pasting position of the image of the sheared area is within the remaining image after shearing; if the pasting position of the image of the sheared area is not within the remaining image after shearing, then adjust the pasting position of the image of the sheared area.

[0084] Obtain the pasting position of the image of the sheared area. The pasting position includes the four peripheral boundary positions of the image of the sheared area. Then, determine whether the entire image of the sheared area is within the range of the remaining image after shearing according to the four peripheral boundary positions. If it is not within the range of the remaining image after shearing, it is considered that the image of the sheared area exceeds the image range at this time, and the pasting position needs to be adjusted until the image of the sheared area is completely within the range of the remaining image after shearing.

[0085] In the specific implementation process, a coordinate system can be constructed according to the sample image, the boundary coordinates of the pasting position of the sheared area image are obtained, and by judging the relationship between the boundary coordinates of the sheared area image and the boundary coordinates of the sample image, it is determined whether the pasting position of the sheared area image is within the remaining image after shearing.

[0086] S206. Respectively input the first training image, the second training image, and the third training image into the neural network to obtain a first output value corresponding to the first training image, a second output value corresponding to the second training image, and a third output value corresponding to the third training image, and calculate the loss function value of the neural network and the similarity between the first output value, the second output value, and the third output value according to the digital label.

[0087] Input the first training image into the neural network to obtain the first output value of the neural network for the first training image, denoted as P M1 . Input the second training image into the neural network to obtain the second output value of the neural network for the second training image, denoted as P M2 . Input the third training image into the neural network to obtain the third output value of the neural network for the second training image, denoted as P M3 .

[0088] The loss function of the neural network can adopt cross entropy. Based on the digital label corresponding to the sample image and the first output value P of the neural network for the first training image M1 calculate a loss function value of the neural network; similarly, based on the digital label corresponding to the sample image and the second output value P of the neural network for the second training image M2 calculate another loss function value of the neural network; and, based on the digital label corresponding to the sample image and the third output value P of the neural network for the third training image M3 calculate another loss function value of the neural network.

[0089] In addition, it is also necessary to calculate according to the first output value P of the neural network for the first training image M1 , the second output value P of the neural network for the second training image M2 and the third output value P of the neural network for the third training image M3 calculate the similarity between P M1 , P M2 and P M3 . The more similar P M1 , P M2 and P M3 are, the better the prediction effect of the neural network indicates.

[0090] In the specific implementation process, the JS divergence loss can be used to calculate PM1 , P M2 and P M3 The smaller the calculated JS divergence loss value is, the closer P M1 , P M2 and P M3 are, and the better the prediction effect of the neural network is.

[0091]

[0092] Among them, KL is the KL divergence.

[0093] S207. Iteratively train the neural network according to the loss function value and the similarity, and when the neural network converges, use the neural network as a digital recognition model.

[0094] Iteratively train the neural network according to the loss function value and the similarity. During the training process, the loss function value and the similarity can be given the same weight to participate in the iterative training of the neural network. That is, the loss function value and the similarity can be multiplied by their respective weights and then added together, and the obtained value is used as the final actual loss value. Based on this loss value, the parameters of the neural network are adjusted. When the loss value is the smallest, it is considered that the neural network converges at this time, and the converged neural network is used as the trained digital recognition model for recognizing handwritten Roman numerals.

[0095] The training method of the digital recognition model provided in the above embodiment obtains the sample image and the digital label corresponding to the sample image, then respectively performs image cropping and data augmentation on the sample image to obtain the first training image and the second training image, inputs the first training image and the second training image into the neural network respectively, and calculates the loss function value of the neural network and the similarity between the first training image and the second training image according to the digital label. Finally, the neural network is trained according to the loss function value and the similarity until the neural network converges to obtain the digital recognition model. Different methods are used to process the sample image to generate different training images to participate in the training of the neural network, realizing the expansion of the training samples. In addition, the similarity between different training images is also added to the training of the neural network, which speeds up the convergence speed of the neural network and improves the training speed of the model.

[0096] Please refer to Figure 6 , Figure 6 which is a schematic block diagram of a training device for a digital recognition model provided in an embodiment of the present application. The training device for the digital recognition model is used to execute the foregoing training method of the digital recognition model. Among them, the training device for the digital recognition model can be configured in a server or a terminal.

[0097] Among them, the server can be an independent server or a server cluster. The terminal can be an electronic device such as a mobile phone, a tablet computer, a notebook computer, a desktop computer, a personal digital assistant, and a wearable device.

[0098] As Figure 6 shown, the training device 300 of the digital recognition model includes: a sample acquisition module 301, an image cropping module 302, a data augmentation module 303, a loss calculation module 304, and a model training module 305.

[0099] The sample acquisition module 301 is configured to acquire a sample image and a digital label corresponding to the sample image.

[0100] The image cropping module 302 is configured to crop the sample image, and use the remaining image after the image cropping as the first training image.

[0101] In one embodiment, the image cropping module 302 includes a type determination sub-module 3021 and a method determination sub-module 3022.

[0102] Among them, the type determination sub-module 3021 is configured to perform Hough transform and Sobel operator processing on the sample image to determine the digital type of the sample image. The method determination sub-module 3022 is configured to determine an image cropping method according to the digital type, and crop the sample image according to the image cropping method.

[0103] The data augmentation module 303 is configured to perform data augmentation on the sample image to obtain a second training image.

[0104] The loss calculation module 304 is configured to input the first training image and the second training image into a neural network respectively, obtain a first output value corresponding to the first training image and a second output value corresponding to the second training image, and calculate a loss function value of the neural network and a similarity between the first output value and the second output value according to the digital label.

[0105] The model training module 305 is configured to perform iterative training on the neural network according to the loss function value and the similarity, and use the neural network as the digital recognition model when the neural network converges.

[0106] Please refer to Figure 7 , Figure 7 which is a schematic block diagram of another training device of the digital recognition model provided by an embodiment of the present application. The training device of the digital recognition model is configured to execute the foregoing training method of the digital recognition model.

[0107] As Figure 7As shown in the figure, the training device 400 of the digital recognition model includes: a sample acquisition module 401, an image cropping module 402, a data augmentation module 403, an image shearing module 404, an image pasting module 405, a loss calculation module 406, and a model training module 407.

[0108] The sample acquisition module 401 is configured to acquire a sample image and a digital label corresponding to the sample image.

[0109] The image cropping module 402 is configured to crop the sample image, and use the remaining image after cropping as the first training image.

[0110] The data augmentation module 403 is configured to perform data augmentation on the sample image to obtain a second training image.

[0111] The image shearing module 404 is configured to determine the digital position in the sample image, determine a shearing area at the digital position, shear the shearing area, and obtain a shearing area image and the remaining image after shearing.

[0112] The image pasting module 405 is configured to paste the shearing area image on the remaining image after shearing to obtain a third training image.

[0113] The loss calculation module 406 is configured to input the first training image, the second training image, and the third training image into a neural network respectively, obtain a first output value corresponding to the first training image, a second output value corresponding to the second training image, and a third output value corresponding to the third training image, and calculate the loss function value of the neural network and the similarity between the first output value, the second output value, and the third output value according to the digital label.

[0114] The model training module 407 is configured to perform iterative training on the neural network according to the loss function value and the similarity, and when the neural network converges, use the neural network as the digital recognition model.

[0115] It should be noted that those skilled in the art can clearly understand that for the convenience and conciseness of description, the specific working processes of the above-described training device of the digital recognition model and each module can refer to the corresponding processes in the foregoing embodiments of the training method of the digital recognition model, and will not be elaborated herein.

[0116] The above-described training device of the digital recognition model can be implemented in the form of a computer program, and this computer program can run on a computer device as shown in Figure 8 the figure.

[0117] Please refer to Figure 8 , Figure 8It is a schematic block diagram of the structure of a computer device provided by an embodiment of the present application. The computer device can be a server or a terminal.

[0118] Referring to Figure 8 , the computer device includes a processor, a memory, and a network interface connected through a system bus. Among them, the memory can include a non-volatile storage medium and an internal memory.

[0119] The non-volatile storage medium can store an operating system and a computer program. The computer program includes program instructions, and when the program instructions are executed, the processor can be caused to execute any training method of a digital recognition model.

[0120] The processor is used to provide computing and control capabilities to support the operation of the entire computer device.

[0121] The internal memory provides an environment for the operation of the computer program in the non-volatile storage medium. When the computer program is executed by the processor, the processor can be caused to execute any training method of a digital recognition model.

[0122] The network interface is used for network communication, such as sending assigned tasks, etc. Those skilled in the art can understand that Figure 8 the structure shown in

[0123] is only a block diagram of some structures related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.

[0124] Among them, in one embodiment, the processor is used to run the computer program stored in the memory to implement the following steps:

[0125] Obtain a sample image and the digital label corresponding to the sample image;

[0126] Crop the sample image, and use the remaining image after cropping as the first training image;

[0127] Perform data augmentation on the sample image to obtain a second training image;

[0128] Input the first training image and the second training image into the neural network respectively to obtain a first output value corresponding to the first training image and a second output value corresponding to the second training image, and calculate the loss function value of the neural network and the similarity between the first output value and the second output value according to the digital label;

[0129] Iteratively train the neural network according to the loss function value and the similarity, and when the neural network converges, use the neural network as a digital recognition model.

[0130] In one embodiment, when the processor implements the image cropping of the sample image, it is used to implement:

[0131] Perform Hough transform and Sobel operator processing on the sample image to determine the digital type of the sample image;

[0132] Determine the image cropping method according to the digital type, and crop the sample image according to the image cropping method.

[0133] In one embodiment, the digital type of the sample image includes a first type and a second type; when the processor implements the determination of the image cropping method according to the digital type, it is used to implement:

[0134] When the digital type is the first type, determine the image cropping method of the sample image as cropping at the left and right ends of the sample image;

[0135] When the digital type is the second type, determine the image cropping method of the sample image as cropping at the upper and lower ends of the sample image.

[0136] In one embodiment, when the processor implements the image cropping of the sample image and uses the remaining image after cropping as the first training image, it is used to implement:

[0137] Crop the sample image, and perform data augmentation on the remaining image after cropping to obtain a first image, where the data augmentation includes at least one of transformation, rotation, and hue change.

[0138] In one embodiment, before the processor implements inputting the first training image and the second training image into a neural network respectively to obtain a first output value corresponding to the first training image and a second output value corresponding to the second training image, and calculating a loss function value of the neural network and a similarity between the first output value and the second output value according to the digital label, the processor is used to implement:

[0139] Determine the digital position in the sample image, determine a shearing area at the digital position, shear the shearing area to obtain a shearing area image and a remaining image after shearing;

[0140] Paste the shearing area image on the remaining image after shearing to obtain a third training image;

[0141] When the processor implements inputting the first training image and the second training image into a neural network respectively to obtain a first output value corresponding to the first training image and a second output value corresponding to the second training image, and calculating a loss function value of the neural network and a similarity between the first output value and the second output value according to the digital label, the processor is used to implement:

[0142] Input the first training image, the second training image, and the third training image into the neural network respectively to obtain a first output value corresponding to the first training image, a second output value corresponding to the second training image, and a third output value corresponding to the third training image, and calculate a loss function value of the neural network and a similarity between the first output value, the second output value, and the third output value according to the digital label.

[0143] In one embodiment, when the processor implements pasting the shearing area image on the remaining image after shearing to obtain a third training image, the processor is used to implement:

[0144] Perform hole filling on the shearing area on the remaining image after shearing to obtain a filled image;

[0145] Paste the shearing area image on the filled image to obtain a third training image.

[0146] In one embodiment, when the processor implements pasting the shearing area image on the remaining image after shearing to obtain a third training image, the processor is used to implement:

[0147] Obtain the pasting position of the shearing area image;

[0148] Determine whether the pasting position of the shearing area image is within the remaining image after shearing;

[0149] If the pasting position of the clipped area image is not within the remaining image after clipping, adjust the pasting position of the clipped area image.

[0150] An embodiment of the present application also provides a computer-readable storage medium. The computer-readable storage medium stores a computer program, and the computer program includes program instructions. The processor executes the program instructions to implement any one of the digital recognition model training methods provided by the embodiments of the present application.

[0151] Among them, the computer-readable storage medium may be an internal storage unit of the computer device described in the foregoing embodiments, such as the hard disk or memory of the computer device. The computer-readable storage medium may also be an external storage device of the computer device, such as a plug-in hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. equipped on the computer device.

[0152] As described above, the above is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily think of various equivalent modifications or replacements, and these modifications or replacements should all be covered within the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the protection scope of the claims.

Claims

1. A training method for a digital recognition model, characterized in that, the method comprises: obtaining a sample image and a digital label corresponding to the sample image; performing image cropping on the sample image, and taking the remaining image after the image cropping as a first training image; performing data augmentation on the sample image to obtain a second training image, and determining the digital position in the sample image, and determining a shearing region at the digital position, performing shearing on the shearing region to obtain a shearing region image and a remaining image after shearing; pasting the shearing region image on the remaining image after shearing to obtain a third training image; inputting the first training image, the second training image and the third training image into a neural network respectively, obtaining a first output value corresponding to the first training image, a second output value corresponding to the second training image and a third output value corresponding to the third training image, and calculating a loss function value of the neural network and a similarity between the first output value, the second output value and the third output value according to the digital label; performing iterative training on the neural network according to the loss function value and the similarity, and when the neural network converges, taking the neural network as a digital recognition model.

2. The training method for a digital recognition model according to claim 1, characterized in that, the performing image cropping on the sample image comprises: performing Hough transform and Sobel operator processing on the sample image to determine the digital type of the sample image; determining an image cropping method according to the digital type, and performing image cropping on the sample image according to the image cropping method.

3. The training method for a digital recognition model according to claim 2, characterized in that, the digital types of the sample image include a first type and a second type; the determining an image cropping method according to the digital type comprises: when the digital type is the first type, determining the image cropping method of the sample image as cropping at the left and right ends of the sample image; when the digital type is the second type, determining the image cropping method of the sample image as cropping at the upper and lower ends of the sample image.

4. The training method for a digital recognition model according to claim 1, characterized in that, the performing image cropping on the sample image, and taking the remaining image after the image cropping as a first training image comprises: performing image cropping on the sample image, and performing data augmentation on the remaining image after the image cropping to obtain a first image, and the data augmentation includes at least one of transformation, rotation and hue change.

5. The training method for a digital recognition model according to claim 1, characterized in that, the pasting the shearing region image on the remaining image after shearing to obtain a third training image comprises: performing hole filling on the shearing region on the remaining image after shearing to obtain a filled image; pasting the shearing region image on the filled image to obtain a third training image.

6. The training method for a digital recognition model according to claim 1, characterized in that, Said pasting the image of the shearing area on the remaining image after shearing to obtain a third training image includes: Obtaining the pasting position of the image of the shearing area; Determining whether the pasting position of the image of the shearing area is within the remaining image after shearing; If the pasting position of the image of the shearing area is not within the remaining image after shearing, adjusting the pasting position of the image of the shearing area.

7. A training device for a digital recognition model, Characterized in that, It includes: A sample acquisition module, configured to acquire a sample image and a digital label corresponding to the sample image; An image cropping module, configured to crop the sample image, and use the remaining image after image cropping as a first training image; A data augmentation module, configured to perform data augmentation on the sample image to obtain a second training image, determine the digital position in the sample image, determine a shearing area at the digital position, shear the shearing area to obtain a shearing area image and a remaining image after shearing; paste the shearing area image on the remaining image after shearing to obtain a third training image; A loss calculation module, configured to input the first training image, the second training image, and the third training image into a neural network respectively, to obtain a first output value corresponding to the first training image, a second output value corresponding to the second training image, and a third output value corresponding to the third training image, and calculate a loss function value of the neural network and the similarity between the first output value, the second output value, and the third output value according to the digital label; A model training module, configured to perform iterative training on the neural network according to the loss function value and the similarity, and when the neural network converges, use the neural network as a digital recognition model.

8. A computer device, Characterized in that, The computer device includes a memory and a processor; The memory is used to store a computer program; The processor is configured to execute the computer program and, when executing the computer program, implement the training method of the digital recognition model according to any one of claims 1 to 6.

9. A computer-readable storage medium, Characterized in that, The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the processor is caused to implement the training method of the digital recognition model according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Model generation method and device, electronic equipment and medium

    CN112529040A

  • Neural network training method, object detection method, device and equipment

    CN113052295A