Semantic segmentation-based casting image number identification method and system, storage medium and equipment

By combining semantic segmentation and OCR technology, the misjudgment and misjudgment problems in casting number identification are solved, the recognition accuracy and efficiency are improved, the cost is reduced, and the complex industrial environment is adapted.

CN120340049APending Publication Date: 2025-07-18NORTHEAST FORESTRY UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510483688.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-17
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The existing casting number identification methods have problems such as misjudgment and misjudgment, low efficiency, high requirements for image quality, requiring a large amount of labeling data and long training time, and high cost.

Method used

Combining semantic segmentation and OCR technology, the casting numbering area is accurately identified through the semantic segmentation model, and character recognition is combined with the deep learning model, reducing the amount and time of training data and reducing costs.

Benefits of technology

It improves the accuracy and efficiency of casting number identification, reduces the requirements for image quality, adapts to complex industrial environments, and reduces costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120340049A_ABST
    Figure CN120340049A_ABST
Patent Text Reader

Abstract

The invention relates to a casting image number identification method and system, a storage medium and equipment, in particular to a casting image number identification method and system based on semantic segmentation, and a storage medium and equipment. The method aims to solve the problems that an existing method is prone to misjudgment and missed judgment, low in efficiency, high in image quality requirement, long in training time, high in cost and likely to be affected by the metal environment, and a large amount of annotation data is needed. The method comprises the steps of 1, obtaining a casting image training set; 2, obtaining a trained semantic segmentation model; 3, inputting a to-be-processed image into the trained semantic segmentation model, and outputting a semantic segmentation result; 4, obtaining a final rectangular region; 5, cutting a corresponding area from the to-be-processed image for the final rectangular area according to the coordinates of the square frame to obtain a cut image; and 6, processing the cut image by adopting optical character recognition (OCR) equipment in combination with a deep learning model to obtain segmented characters. The method is used in the field of casting image number identification.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a method, system, storage medium and device for identifying the number of a casting image. Background Art

[0002] The existing methods for identifying the casting number mainly include the following:

[0003] Manual visual identification: directly observe the number on the surface of the casting and identify it by the naked eye or with the help of simple tools such as a magnifying glass. This method is suitable for the cases where the number is clear and simple, but for complex or tiny numbers, misjudgment and omission are likely to occur, and the efficiency is relatively low.

[0004] Optical Character Recognition (OCR) technology: use an optical device to collect the image of the casting number, and then use OCR software to identify and analyze the characters in the image. This method can process a large number of images quickly and has a relatively high recognition accuracy, but it has high requirements for the image quality, and stains, abrasions, etc. on the surface of the casting may affect the recognition effect.

[0005] Recognition method based on deep learning: represented by the Convolutional Neural Network (CNN), train with a large amount of casting number image data to enable the model to automatically learn the characteristics of the number. This method has good recognition ability for complex backgrounds, blurred or deformed numbers, and the recognition accuracy will continuously improve with the increase of the data volume and the optimization of the model, but a large amount of labeled data is required for training, and the model training time is relatively long.

[0006] Laser marking recognition: pre-mark a specific number on the casting with a laser, and then read the number information through a laser scanning device. The number marked by laser is clear and has good permanence, and the scanning and recognition speed is fast and the accuracy is high, but special laser marking and scanning devices are required, and the cost is relatively high.

[0007] Radio Frequency Identification (RFID) technology: implant or paste an RFID tag with number information on the casting, and identify the number in the tag through an RFID reader. This method has the advantages of non-contact, long-distance recognition, and the ability to identify multiple tags at the same time, but the cost of RFID tags is relatively high, and it may be affected by the metal environment. Summary of the Invention

[0008] The object of the present invention is to solve the problems that the existing manual visual recognition method is prone to misjudgment and missed judgment for complex or tiny numbers, and has low efficiency; the existing optical character recognition technology has high requirements for image quality, and stains, wear, etc. on the surface of castings may affect the recognition effect; the existing deep learning recognition method requires a large amount of labeled data for training, and the model training time is relatively long; the existing laser marking recognition requires special laser marking and scanning equipment, and the cost is relatively high; and the existing radio frequency identification technology has a relatively high tag cost and may be affected by the metal environment. Therefore, a casting image number recognition method, system, storage medium and device based on semantic segmentation are proposed.

[0009] The specific process of the casting image number recognition method based on semantic segmentation is as follows:

[0010] Step 1: Obtain a training set of casting images;

[0011] Step 2: Input the training set of casting images into the semantic segmentation model to train the semantic segmentation model, obtain a trained semantic segmentation model, and retain the parameters of the trained semantic segmentation model;

[0012] Step 3: Input the image to be processed into the trained semantic segmentation model. The trained semantic segmentation model outputs a semantic segmentation result, and the semantic segmentation result includes the range of key information of the casting. The key information of the casting is the area position coordinates where the casting model and the casting specification text are located;

[0013] Step 4: Process the semantic segmentation result to obtain a final rectangular area;

[0014] Step 5: According to the coordinates of the box of the final rectangular area obtained in Step 4, crop the corresponding area from the image to be processed in Step 3 to obtain a cropped image;

[0015] Step 6: Use an optical character recognition (OCR) device combined with a deep learning model to process the cropped image obtained in Step 5 to obtain segmented characters.

[0016] Preferably, in Step 1, obtaining a training set of casting images; the specific process is as follows:

[0017] Collect labeled casting images as a training set of casting images;

[0018] The label is the area where the casting model and specification text are located and the corresponding coordinates.

[0019] Preferably, in Step 4, processing the semantic segmentation result to obtain a final rectangular area;

[0020] The specific process is as follows:

[0021] Step 4-1: Use connected component analysis to process the located range containing the key information of the casting to obtain the regional contour;

[0022] Step 4-2: Use the contour detection algorithm to calculate the bounding rectangle of the regional contour determined in Step 4-1, and filter out the rectangular area;

[0023] Step 4-3: Set the standard size values of the areas where the casting model and casting specifications are located;

[0024] Based on the set standard size values of the areas where the casting model and casting specifications are located, adjust the length and width of the rectangle in the rectangular area filtered out in Step 4-2 proportionally to obtain the final rectangular area.

[0025] Preferably, in Step 5, the final rectangular area obtained in Step 4 is cropped from the image to be processed in Step 3 according to the coordinates of the bounding box to obtain the cropped image;

[0026] The specific process is as follows:

[0027] Save the cropped image in the standard image format.

[0028] Preferably, in Step 6, an optical character recognition (OCR) device combined with a deep learning model is used to process the cropped image obtained in Step 5 to obtain the segmented characters;

[0029] The specific process is as follows:

[0030] Step 6-1: Use an optical character recognition (OCR) device to preprocess the cropped image obtained in Step 5;

[0031] Step 6-2: Use a deep learning model to segment the characters in the image preprocessed in Step 6-1 to obtain the segmented characters.

[0032] Preferably, in Step 6-1, an optical character recognition (OCR) device is used to preprocess the cropped image obtained in Step 5; the specific process is as follows:

[0033] 1) Convert the cropped image obtained in Step 5 into a grayscale image;

[0034] 2) Use median filtering to remove noise from the grayscale image obtained in 1) to obtain the denoised image;

[0035] 3) Perform binarization processing on the denoised image to obtain the binarized image;

[0036] 4) Detect the tilt angle of the text line or character in the binarized image and perform rotation correction on the image.

[0037] Preferably, in step 6.2, a deep learning model is used to perform character segmentation on the pre-processed image in step 6.1 to obtain segmented characters. The specific process is as follows:

[0038] 1). Perform segmentation and positioning on the pre-processed image in step 6.1. The process is as follows:

[0039] Input the pre-processed image in step 6.1 into the PGNet model, and the PGNet model outputs the position of the character region and the center line of the text line.

[0040] 2). Based on TPS, perform geometric correction on the curved text and the curved center line in the character region output by the PGNet model to obtain the corrected text and center line.

[0041] Input the corrected text and center line into the PG-CTC model, and the PG-CTC model outputs a character sequence.

[0042] The casting image number recognition system based on semantic segmentation is used to execute the casting image number recognition method based on semantic segmentation.

[0043] A storage medium stores at least one instruction, and the at least one instruction is loaded and executed by a processor to implement the casting image number recognition method based on semantic segmentation.

[0044] A casting image number recognition device based on semantic segmentation includes a processor and a memory. At least one instruction is stored in the memory, and the at least one instruction is loaded and executed by the processor to implement the casting image number recognition method based on semantic segmentation.

[0045] The beneficial effects of the present invention are as follows:

[0046] In the field of casting number recognition, the present invention combines semantic segmentation with OCR and combines OCR with deep learning, which has significant advantages compared with traditional methods:

[0047] 1. Combination of semantic segmentation and OCR: When traditional OCR technology directly processes casting images, factors such as complex surrounding backgrounds, stains, and wear will interfere with recognition. By combining semantic segmentation, the area where the casting number is located can be accurately recognized first, and irrelevant parts can be filtered out. For example, if there are other patterns or stains on the casting surface, semantic segmentation can circle the number area, enabling OCR to focus on processing, reducing interference, improving the recognition accuracy, and being able to work effectively in complex industrial environments, reducing costs.

[0048] 2. Combining OCR and deep learning: Traditional deep learning for identifying casting numbers requires a large amount of labeled data and long training time. By combining with OCR technology, the advantage of OCR in recognizing common standard characters can be utilized to extract basic features and narrow the training scope of the deep learning model. For example, the preliminary recognition results of numbers and letters can provide initial information for the deep learning model, and the model only needs to train for special cases where OCR misjudges or is difficult to recognize, reducing the amount of training data and time, improving efficiency, and adapting to new number styles faster. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] Figure 1 This is the flowchart of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0050] DETAILED DESCRIPTION OF THE EMBODIMENT 1: The specific process of the method for identifying casting image numbers based on semantic segmentation in this embodiment is as follows:

[0051] Step 1: Obtain a training set of casting images;

[0052] Step 2: Input the training set of casting images into the semantic segmentation model (U-Net) to train the semantic segmentation model, obtain a trained semantic segmentation model, and retain the parameters of the trained semantic segmentation model;

[0053] Step 3: Input the image to be processed into the trained semantic segmentation model, and the trained semantic segmentation model outputs a semantic segmentation result, which includes the range of key information of the casting; the key information of the casting is the range of the coordinates of the area where the casting model and the casting specification text are located;

[0054] Step 4: Process the semantic segmentation result to obtain the final rectangular area;

[0055] Step 5: According to the coordinates of the box of the final rectangular area obtained in Step 4, crop the corresponding area from the image to be processed in Step 3 to obtain a cropped image;

[0056] Step 6: Use an optical character recognition (OCR) device combined with a deep learning model to process the cropped image obtained in Step 5 to obtain segmented characters.

[0057] DETAILED DESCRIPTION OF THE EMBODIMENT 2: The difference between this embodiment and Embodiment 1 is that in Step 1 of obtaining a training set of casting images; the specific process is as follows:

[0058] Collect labeled casting images as a training set of casting images;

[0059] The label is the area where the casting model, specification, etc. are located and the corresponding coordinates.

[0060] Other steps and parameters are the same as those in Embodiment 1.

[0061] Embodiment 3: Different from Embodiment 1 or 2, in step 4, the semantic segmentation result is processed to obtain the final rectangular area;

[0062] The specific process is as follows:

[0063] Step 4-1: Use connected component analysis to process the located range containing the key information of the casting to obtain the region contour;

[0064] Step 4-2: Use the contour detection algorithm to calculate the circumscribed rectangle of the region contour determined in step 4-1, and filter out the rectangle area most likely to contain the text for OCR recognition;

[0065] Step 4-3: Set the standard size values for the areas where the text such as the casting model and casting specifications is located;

[0066] Based on the set standard size values for the areas where the text such as the casting model and casting specifications is located, adjust the length and width of the rectangle obtained by filtering out the rectangle area most likely to contain the text for OCR recognition in step 2 in proportion to obtain the final rectangular area;

[0067] Check the filtered circumscribed rectangle. If the size is too large or too small, adjust the length and width of the rectangle in proportion according to the actual size of the text area; if there is a deviation in the position, fine-tune the position of the rectangle in the image according to the overall layout of the image to ensure that the box completely frames the useful information and does not contain too much redundant background.

[0068] Other steps and parameters are the same as those in Embodiment 1 or 2.

[0069] Embodiment 4: Different from one of Embodiments 1 to 3, in step 5, the final rectangular area obtained in step 4 is cropped from the image to be processed in step 3 according to the coordinates of the box to obtain the cropped image;

[0070] The specific process is as follows:

[0071] Save the cropped image in the standard image format (such as PNG, JPEG) to provide a clear and accurate input image for subsequent OCR recognition.

[0072] Other steps and parameters are the same as those in one of Embodiments 1 to 3.

[0073] Embodiment 5: Different from one of Embodiments 1 to 4, in step 6, an optical character recognition (OCR) device combined with a deep learning model is used to process the cropped image obtained in step 5 to obtain the segmented characters;

[0074] The specific process is as follows:

[0075] Step Six One: Use an optical character recognition (OCR) device to preprocess the cropped image obtained in Step Five;

[0076] Step Six Two: Use a deep learning model to perform character segmentation on the image preprocessed in Step Six One to obtain segmented characters.

[0077] Other steps and parameters are the same as those in any one of Embodiments One to Four.

[0078] Specific Embodiment Six: The difference between this embodiment and any one of Embodiments One to Five is that in Step Six One, an optical character recognition (OCR) device is used to preprocess the cropped image obtained in Step Five; the specific process is as follows:

[0079] 1) Convert the cropped image obtained in Step Five into a grayscale image to reduce the amount of data and computational complexity;

[0080] 2) Use median filtering to remove noise from the grayscale image obtained in 1) to obtain a denoised image;

[0081] Noise reduction: Use methods such as median filtering and Gaussian filtering to remove noise in the image and improve the image quality.

[0082] 3) Perform binarization processing on the denoised image to obtain a binarized image;

[0083] Binarization: Convert the grayscale image into a black-and-white binary image to separate the characters from the background. Common methods include global thresholding, adaptive thresholding, etc.

[0084] 4) Detect the tilt angle of the text line or characters in the binarized image, and perform rotation correction on the image to make the text horizontal or vertical.

[0085] Other steps and parameters are the same as those in any one of Embodiments One to Five.

[0086] Specific Embodiment Seven: The difference between this embodiment and any one of Embodiments One to Six is that in Step Six Two, a deep learning model is used to perform character segmentation on the image preprocessed in Step Six One to obtain segmented characters; the specific process is as follows:

[0087] 1) Perform segmentation and localization on the image preprocessed in Step Six One; the process is as follows:

[0088] Input the image preprocessed in Step Six One into the PGNet model. The PGNet model outputs the positions of the character regions and the centerlines of the text lines (read along the lines);

[0089] 2) Based on TPS, perform geometric correction on the curved text and curved centerlines in the character regions output by the PGNet model to obtain corrected text and centerlines;

[0090] Input the corrected text and the center line into the PG-CTC model, and the PG-CTC model outputs a character sequence;

[0091] Realize the integration of detection and recognition;

[0092] Other steps and parameters are the same as those in any one of the specific embodiments one to six.

[0093] Specific embodiment eight: This embodiment is a casting image number recognition system based on semantic segmentation, characterized in that the system is used to execute the casting image number recognition method based on semantic segmentation described in any one of claims 1 to 7.

[0094] Specific embodiment nine: This embodiment is a storage medium, in which at least one instruction is stored, and the at least one instruction is loaded and executed by a processor to implement the casting image number recognition method based on semantic segmentation.

[0095] It should be understood that any method corresponding to the present invention can be provided as a computer program product, software or a computerized method, which may include a non-transitory machine-readable medium storing instructions thereon, and the instructions can be used to program a computer system or other electronic devices. The storage medium may include, but is not limited to, magnetic storage media, optical storage media; magneto-optical storage media include: read-only memory ROM, random access memory RAM, erasable programmable memory (e.g., EPROM and EEPROM), and flash memory layers; or other types of media suitable for storing electronic instructions.

[0096] Specific embodiment ten: This embodiment is a casting image number recognition device based on semantic segmentation. The device includes a processor and a memory. It should be understood that any device including a processor and a memory described in the present invention, and the device may further include other units and modules for display, interaction, processing, control, etc. through signals or instructions, as well as other functions;

[0097] At least one instruction is stored in the memory, and the at least one instruction is loaded and executed by the processor to implement the casting image number recognition method based on semantic segmentation.

[0098] Model selection and loading: Use the semantic segmentation model to complete the construction of the model and load the trained weights. These models can accurately identify different categories in the image and distinguish the casting area from the complex background.

[0099] Image Input and Segmentation: Input the image to be processed into the semantic segmentation model with pre-loaded weights. The model analyzes the image features through forward propagation, outputs the probability of each pixel belonging to different classes, and thus obtains the semantic segmentation result to clarify the boundary between the casting and the background.

[0100] Region Location and Screening: Based on the semantic segmentation result, locate the approximate range that contains the key information of the casting (such as the area where text like model number and specifications are located). Use connected component analysis and contour detection algorithms to find the corresponding region contour, calculate the bounding rectangle of the contour, and screen out the rectangle region that is most likely to contain text for OCR recognition.

[0101] Bounding Box Adjustment and Determination: Check the screened bounding rectangle. If the size is too large or too small, adjust the length and width of the rectangle proportionally according to the actual size of the text region; if there is a deviation in the position, fine-tune the position of the rectangle in the image according to the overall layout of the image to ensure that the bounding box completely encloses the useful information without including too much redundant background.

[0102] Region Extraction and Saving: After determining the bounding box area, crop the corresponding region from the original image according to the coordinates of the bounding box. Save the cropped image in a standard image format (such as PNG, JPEG) to provide a clear and accurate input image for subsequent OCR recognition.

[0103] Combination of Django and Deep Learning

[0104] 1. Environment Setup: Install Python, Django, deep learning frameworks (such as TensorFlow, PyTorch) and common libraries (numpy, pandas, etc.), and isolate project dependencies with a virtual environment.

[0105] 2. Model Development: Build and train a model using a deep learning framework, and save the model weights after completion.

[0106] 3. Django Configuration: Create a Django project and application, configure the database and static file paths, define the data model in models.py and migrate the database.

[0107] 4. View and URL Settings: Write view functions in views.py to handle requests and call the model, and configure the mapping of URLs and views in urls.py.

[0108] 5. Front-End and Back-End Interaction: Develop the front-end page with HTML, CSS and JavaScript, interact with the back-end through AJAX, and the back-end receives requests, processes data and returns results.

[0109] Data Management and Permissions: Implement data addition, deletion, modification and query in the view function, and set user permissions using the Django authentication system.

[0110] Combining deep learning with Django: Traditional methods for identifying casting numbers are mostly local operations, which are inconvenient for management and data sharing. Connecting the deep learning recognition results to a web page made with Django can build a centralized management platform. Staff can view and manage casting number information anytime and anywhere through the web page, realizing real-time data sharing. For example, in a large factory, different workshops can access the system simultaneously to understand the production progress and numbering situation of castings, and can also conveniently perform data backup, statistical analysis, etc., which is convenient for management and decision-making.

[0111] The present invention may also have many other embodiments. Without departing from the spirit and essence of the present invention, those skilled in the art can make various corresponding changes and deformations according to the present invention, but these corresponding changes and deformations should all fall within the protection scope of the appended claims of the present invention.

Claims

1. A method for identifying the number of a casting image based on semantic segmentation, characterized in that: The specific process of the method is as follows: Step 1: Obtain a training set of casting images; Step 2: Input the training set of casting images into a semantic segmentation model to train the semantic segmentation model, obtain a trained semantic segmentation model, and retain the parameters of the trained semantic segmentation model; Step 3: Input the image to be processed into the trained semantic segmentation model. The trained semantic segmentation model outputs a semantic segmentation result, and the semantic segmentation result includes the range of key information of the casting; the key information of the casting is the casting model number and the position coordinates of the area where the casting specification text is located; Step 4: Process the semantic segmentation result to obtain the final rectangular area; Step 5: Crop the corresponding area from the image to be processed in Step 3 according to the coordinates of the box based on the final rectangular area obtained in Step 4 to obtain a cropped image; Step 6: Use an optical character recognition OCR device in combination with a deep learning model to process the cropped image obtained in Step 5 to obtain segmented characters.

2. The method for identifying the casting image number based on semantic segmentation according to claim 1, wherein: In Step 1, a training set of casting images is obtained; The specific process is as follows: Collect labeled casting images as the training set of casting images; The label is the casting model number, the area where the specification text is located, and the corresponding coordinates.

3. The method for identifying the casting image number based on semantic segmentation according to claim 2, characterized in that: In Step 4, the semantic segmentation result is processed to obtain the final rectangular area; The specific process is as follows: Step 4-1: Use connected component analysis to process the located range containing the key information of the casting to obtain a regional contour; Step 4-2: Use a contour detection algorithm to calculate the circumscribed rectangle of the regional contour determined in Step 4-1 and screen out the rectangular area; Step 4-3: Set the size standard values of the areas where the casting model number and the casting specification text are located; Based on the set size standard values of the areas where the casting model number and the casting specification text are located, adjust the length and width of the rectangle in the rectangular area screened out in Step 4-2 proportionally to obtain the final rectangular area.

4. The method for identifying the casting image number based on semantic segmentation according to claim 3, characterized in that: In Step 5, the corresponding area is cropped from the image to be processed in Step 3 according to the coordinates of the box based on the final rectangular area obtained in Step 4 to obtain a cropped image; The specific process is as follows: Save the cropped image in a standard image format.

5. The method for identifying the casting image number based on semantic segmentation according to claim 4, characterized in that: In Step 6, an optical character recognition OCR device in combination with a deep learning model is used to process the cropped image obtained in Step 5 to obtain segmented characters; The specific process is as follows: Step 6-1: Use an optical character recognition OCR device to preprocess the cropped image obtained in Step 5; Step 6-2: Use a deep learning model to perform character segmentation on the image preprocessed in Step 6-1 to obtain segmented characters.

6. The method for identifying the casting image number based on semantic segmentation according to claim 5, characterized in that: In Step 6-1, an optical character recognition OCR device is used to preprocess the cropped image obtained in Step 5; The specific process is as follows: 1) Convert the cropped image obtained in Step 5 into a grayscale image; 2) Use median filtering to remove noise from the grayscale image obtained in 1) to obtain a denoised image; 3) Perform binarization processing on the denoised image to obtain a binarized image; 4) Detect the tilt angle of the text line or character in the binarized image and perform rotation correction on the image.

7. The method for identifying the casting image number based on semantic segmentation according to claim 6, wherein: In Step 6-2, a deep learning model is used to perform character segmentation on the image preprocessed in Step 6-1 to obtain segmented characters; the specific process is as follows: 1), Segment and locate the image preprocessed in Step Six; the process is as follows: Input the image preprocessed in Step Six into the PGNet model, and the PGNet model outputs the positions of character regions and the centerlines of text lines; 2), Based on TPS, perform geometric correction on the curved text and curved centerlines in the character regions output by the PGNet model to obtain the corrected text and centerlines; Input the corrected text and centerlines into the PG-CTC model, and the PG-CTC model outputs a character sequence.

8. A casting image number recognition system based on semantic segmentation, characterized in that The system is used to execute the method for identifying the casting image number based on semantic segmentation described in any one of claims 1 to 7.

9. A storage medium, characterized in that, At least one instruction is stored in the storage medium, and the at least one instruction is loaded and executed by a processor to implement the method for identifying the casting image number based on semantic segmentation described in any one of claims 1 to 7.

10. A casting image number recognition device based on semantic segmentation, characterized in that, The device includes a processor and a memory, and at least one instruction is stored in the memory, and the at least one instruction is loaded and executed by the processor to implement the method for identifying the casting image number based on semantic segmentation described in any one of claims 1 to 7.