Material detection method and apparatus, electronic device, storage medium, and program product

WO2026199114A1PCT designated stage Publication Date: 2026-10-01BOE TECHNOLOGY GROUP CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2025/084412
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-03-24
Publication Date
2026-10-01

Smart Images

  • Figure CN2025084412_01102026_PF_FP_ABST
    Figure CN2025084412_01102026_PF_FP_ABST
Patent Text Reader

Abstract

The present disclosure specifically relates to the field of image processing, and provides a material detection method, applied to the technical field of detection. The method comprises: performing orientation recognition on a target text image in an image of a material to be detected, to obtain a text orientation recognition result; in response to the text orientation recognition result not satisfying a first preset condition, performing text recognition on the target text image and a supplementary text image, to obtain a text content recognition result, wherein the supplementary text image includes an image obtained by processing the target text image on the basis of a preset processing manner; and generating a detection result for said material on the basis of the text content recognition result. The present disclosure further provides a material detection apparatus, a device, a storage medium, and a program product.
Need to check novelty before this filing date? Find Prior Art

Description

Material testing methods, apparatus, electronic equipment, storage media and program products Technical Field

[0001] This disclosure relates to the field of detection technology, particularly to the field of image processing, and more specifically, to a material detection method, apparatus, electronic device, storage medium, and program product. Background Technology

[0002] The efficient operation of the production line is crucial for product quality and production efficiency. However, material misfeeding frequently occurs during actual production. Misfeeding not only leads to product defects but can also cause equipment malfunctions. For example, in display panel production, different models and applications require specific polarizers (POLs) from different manufacturers. If a POL is mistakenly fed into the production line, it will result in significant economic losses. Therefore, it is necessary to inspect the materials fed into the production line. Summary of the Invention

[0003] In view of the above problems, this disclosure provides a material testing method, apparatus, electronic device, storage medium, and program product.

[0004] According to one aspect of this disclosure, a method for detecting materials is provided, comprising:

[0005] Orientation recognition is performed on the target text image in the image of the material to be detected to obtain the text orientation recognition result;

[0006] In response to a text orientation recognition result that does not meet the first preset condition, text recognition is performed on the target text image and the supplementary text image to obtain a text content recognition result. The supplementary text image includes an image processed from the target text image according to a preset processing method; and

[0007] The detection results are generated based on the text content recognition results for the material to be detected.

[0008] According to another aspect of this disclosure, a material detection device is provided, comprising:

[0009] The text orientation recognition module is used to recognize the orientation of target text images in the image of the material to be detected, and to obtain the text orientation recognition result.

[0010] The text content recognition module is used to perform text recognition on the target text image and the supplementary text image in response to a text orientation recognition result that does not meet a first preset condition, thereby obtaining a text content recognition result. The supplementary text image includes an image processed from the target text image according to a preset processing method.

[0011] The first generation module is used to generate detection results for the material to be detected based on the text content recognition results.

[0012] According to another aspect of this disclosure, an electronic device is provided, comprising:

[0013] One or more processors;

[0014] Memory, used to store one or more computer programs.

[0015] In this process, one or more processors execute one or more computer programs to implement the steps of the above method.

[0016] According to another aspect of this disclosure, a computer-readable storage medium is provided that stores a computer program or instructions thereon, wherein the computer program or instructions, when executed by a processor, implement the steps of the above-described method.

[0017] According to another aspect of this disclosure, a computer program product is provided, including a computer program or instructions that, when executed by a processor, implement the steps of the above-described method.

[0018] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description

[0019] The above and other objects, features and advantages of this disclosure will become clearer from the following description of embodiments with reference to the accompanying drawings, in which:

[0020] Figure 1 schematically shows a structural diagram of a polarizer;

[0021] Figure 2 schematically illustrates a flowchart of a material detection method according to an embodiment of the present disclosure;

[0022] Figure 3 schematically illustrates a diagram of text orientation according to an embodiment of the present disclosure;

[0023] Figure 4 schematically illustrates a diagram of determining target-identified text according to an embodiment of the present disclosure;

[0024] Figure 5 schematically illustrates a diagram of determining target-identified text according to another embodiment of the present disclosure;

[0025] Figure 6 schematically illustrates a flowchart of a method for determining a target text image according to an embodiment of the present disclosure;

[0026] Figure 7 schematically illustrates a diagram of determining a matching pair according to an embodiment of the present disclosure;

[0027] Figure 8 schematically illustrates a target text detection box according to an embodiment of the present disclosure;

[0028] Figure 9 schematically illustrates a template configuration interface according to an embodiment of the present disclosure;

[0029] Figure 10 schematically illustrates an interface diagram of the fixed value option according to an embodiment of the present disclosure;

[0030] Figure 11 schematically illustrates an interface diagram facing a visualization window according to an embodiment of the present disclosure.

[0031] Figure 12 schematically illustrates a diagram of determining the detection result according to an embodiment of the present disclosure;

[0032] Figure 13 schematically illustrates a material detection method according to another embodiment of the present disclosure;

[0033] Figure 14 schematically illustrates a structural block diagram of a material detection device according to an embodiment of the present disclosure; and

[0034] Figure 15 schematically illustrates a block diagram of an electronic device suitable for implementing an information interaction method according to an embodiment of the present disclosure. Detailed Implementation

[0035] To make the objectives, technical solutions, and advantages of the embodiments of this disclosure clearer, the technical solutions of the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this disclosure. Based on the described embodiments of this disclosure, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this disclosure. It should be noted that throughout the accompanying drawings, the same elements are represented by the same or similar reference numerals. In the following description, some specific embodiments are used for descriptive purposes only and should not be construed as limiting this disclosure in any way, but are merely examples of embodiments of this disclosure. Conventional structures or configurations will be omitted where they may cause confusion in understanding this disclosure. It should be noted that the shapes and dimensions of the components in the figures do not reflect actual size and proportion, but are only schematic representations of the embodiments of this disclosure.

[0036] Unless otherwise defined, the technical or scientific terms used in the embodiments of this disclosure shall have the ordinary meaning as understood by those skilled in the art. The terms "first," "second," and similar words used in the embodiments of this disclosure do not indicate any order, quantity, or importance, but are merely used to distinguish different components.

[0037] In the embodiments disclosed herein, the collection, updating, analysis, processing, use, transmission, provision, disclosure, and storage of data (e.g., including but not limited to user personal information) comply with relevant laws and regulations, are used for legitimate purposes, and do not violate public order and good morals. In particular, necessary measures have been taken to prevent unauthorized access to user personal information data and to safeguard user personal information security and network security.

[0038] In the embodiments disclosed herein, user authorization or consent is obtained before acquiring or collecting user personal information.

[0039] Figure 1 schematically shows a structural diagram of a polarizer.

[0040] The polarizer (POL) is a key component of a display panel, primarily used to control and adjust the direction of light propagation, directly affecting the display's optical performance such as contrast, brightness, and viewing angle. The POL has a protective film and a release film. The protective film is a transparent film covering the surface of the polarizer, protecting the front optical surface from damage during transportation and installation. The release film covers the adhesive layer on the back of the polarizer, protecting this adhesive layer.

[0041] As shown in Figure 1, the protective film or release film of the polarizer 110 is printed with a material code encoded according to certain rules, representing different models and products from different manufacturers. It consists of a fixed value 112 related to the material number and a serial number 111. The material number includes numbers and letters. The serial number varies for different polarizers, while other values ​​remain constant. The serial number is matched to characters smaller than the fixed value, thus distinguishing it from the fixed value. Additionally, some POL protective films also have barcodes affixed, such as the barcode 113 on the polarizer 110 shown in Figure 1.

[0042] In actual production, different types of display panels require specific POLs from different manufacturers for their intended use. If confusion arises during the POL deployment process, it can lead to significant economic losses. Therefore, it is necessary to test the POL materials.

[0043] Based on the above-mentioned technical problems, this disclosure provides a material detection method, including: performing orientation recognition on a target text image in an image of the material to be detected to obtain a text orientation recognition result; responding to the text orientation recognition result not meeting a first preset condition, performing text recognition on the target text image and a supplementary text image to obtain a text content recognition result, wherein the supplementary text image includes an image after the target text image has been processed according to a preset processing method; and generating a detection result for the material to be detected based on the text content recognition result.

[0044] By identifying target text images in the image of the material to be detected, such as the fixed values ​​mentioned above, different POL materials can be distinguished, thereby reducing the possibility of material mixing. Furthermore, this embodiment of the present disclosure performs orientation recognition on the target text image. When the text orientation recognition result does not meet the first preset condition, text recognition is performed on the target text image in four directions, and the detection result is determined based on the text content recognition result. In addition, this embodiment of the present disclosure can also perform orientation recognition on the target text image and the rotated text image obtained by rotating the target text image by a preset angle when the aspect ratio of the target text image meets the second preset condition. When the text orientation recognition result does not meet the first preset condition, text recognition is performed on the target text image and the rotated text image in four directions respectively, and the detection result is determined based on the text content recognition result. This allows for effective recognition of approximately square images and symmetrical characters, contributing to improved detection accuracy.

[0045] Figure 2 schematically illustrates a flowchart of a material detection method according to an embodiment of the present disclosure.

[0046] As shown in Figure 2, the material detection method according to embodiment 200 of this disclosure includes operations S210 to S230.

[0047] In operation S210, the orientation of the target text image in the image of the material to be detected is recognized, and the text orientation recognition result is obtained.

[0048] The material to be inspected can be a polarizing film placed on the production line. The image of the material to be inspected can be an image taken of the material. For example, the image of the material to be inspected can be an image taken of the material by the AOI (Automated Optical Inspection) system on the production line.

[0049] The target text image can be an image of the text detection box region on the image of the material to be detected. For example, the target text image is an image obtained by cropping the text detection box region on the image of the material to be detected.

[0050] For example, the target text image can be a screenshot of the area corresponding to the fixed value 112 shown in Figure 1. Specifically, after acquiring the image of the material to be detected, text box detection and target text box filtering can be performed on the image of the material to be detected to obtain the target text image. For example, after acquiring the image of the material to be detected for the polarizer 110, text box detection processing can be performed on the image of the material to be detected to obtain a text detection box A for the serial number 111 and a text detection box B for the fixed value 112. Then, the target text detection box, i.e., the text detection box B corresponding to the fixed value 112, is determined from text detection boxes A and B. After that, a screenshot of the area corresponding to text detection box B is taken to obtain the target text image for processing.

[0051] It should be noted that the target text image may include one or more. The material to be detected may include one or more material numbers, therefore the target text image may also contain one or more.

[0052] The text orientation recognition result disclosed herein includes text orientation representations of relative orientation, such as the orientation relative to a horizontal target text image. Text orientation can include at least one of the following: 0° rotation, 180° rotation, horizontal flip, and vertical flip. 0° rotation indicates that the target text image is not rotated relative to the horizontal target text image. 180° rotation indicates that the target text image is rotated 180° relative to the horizontal target text image. Horizontal flip indicates that the target text image is symmetrical in the vertical direction relative to the horizontal target text image. Vertical flip indicates that the target text image is symmetrical in the horizontal direction relative to the horizontal target text image.

[0053] Figure 3 schematically illustrates a diagram of text orientation according to an embodiment of the present disclosure.

[0054] As shown in Figure 3, the target text image in this embodiment can be image 310. In this case, image 310 is rotated 0°, meaning it is a horizontal target text image. Image 320 is the image obtained by rotating image 310 180°, and in this case, image 320 is rotated 180°. Image 330 is the image obtained by horizontally flipping image 310, and in this case, image 330 is horizontally flipped. Image 340 is the image obtained by vertically flipping image 310, and in this case, image 340 is vertically flipped.

[0055] In operation S220, in response to the text orientation recognition result not meeting the first preset condition, text recognition is performed on the target text image and the supplementary text image to obtain the text content recognition result.

[0056] The text orientation recognition result can include the text orientation and the corresponding orientation confidence score. For example, the text orientation recognition result could be that the text orientation is rotated 180°, and the orientation confidence score is 0.9.

[0057] The first preset condition can be that the direction confidence score does not meet a preset threshold. For example, the first preset condition can be that the maximum direction confidence score in the text orientation recognition result is ≥0.995.

[0058] If the text orientation recognition result does not meet the first preset condition, it indicates that the orientation recognition of the target text image is inaccurate. In this case, text recognition is performed on both the target text image and the supplementary text image. The supplementary text image includes the image after the target text image has been processed according to a preset processing method.

[0059] The preset processing methods include rotation and flipping. For example, the supplementary image can be at least one of the following: an image obtained by rotating the target text image by 180°, an image obtained by flipping the target text image horizontally, or an image obtained by flipping the target text image vertically.

[0060] For example, if image 310 in Figure 3 is the target text image, then at least one of images 320, 330, and 340 can be a supplementary image to image 310.

[0061] In one example, the supplementary text images include image 1, which is obtained by rotating the target text image by 180°; image 2, which is obtained by horizontally flipping the target text image; and image 3, which is obtained by vertically flipping the target text image. At this time, text recognition is performed on the target text image and the supplementary text images, namely images 1, 2, and 3. That is, text content recognition is performed on the target text image in four directions to obtain the text content recognition results, thereby improving the accuracy of text content recognition.

[0062] The text content recognition result can include the recognized text of both the target text image and the supplementary text images. For example, the text recognition result can include the recognized text of the target text image mentioned above, as well as the recognized text of each of images 1, 2, and 3.

[0063] Text recognition of the target text image and supplementary text images can be performed using a text content recognition model. ParseQ can be used as the text content recognition model. The training dataset for the text content recognition model can include a character set, specifically numbers, uppercase and lowercase letters, and spaces. Considering the recognition of symmetrical text, images of the original image and an image rotated 90 degrees in four directions are added to the training dataset to enable the recognition of horizontally and vertically flipped text content. The four directions can be rotation of 0°, rotation of 180°, horizontal flip, and vertical flip.

[0064] In operation S230, detection results for the material to be detected are generated based on the text content recognition results.

[0065] Generating detection results for the material to be detected based on text content recognition results may include matching the text content recognition results with preset text in a preset matching template to obtain the detection results. The detection results may include the similarity between the recognized text in the text content recognition results and the preset text.

[0066] This embodiment of the disclosure performs orientation recognition on the target text image. When the text orientation recognition result does not meet the first preset condition, text recognition is performed on the target text image in four directions, and then the detection result is determined based on the text content recognition result. It can effectively recognize images that are approximately square and text with symmetrical characters, which helps to improve detection accuracy. In addition, by recognizing the target text image in the image of the material to be detected, such as the fixed value mentioned above, different POL materials can be distinguished, thereby reducing the possibility of material mixing.

[0067] In one example, the text content recognition result may also include the recognized text obtained after filtering the recognized text from both the target text image and the supplementary text image. For example, after obtaining the recognized text from both the target text image and the supplementary text image, the recognized text with a recognition confidence score ≥ 0.8 is selected as the text content recognition result.

[0068] Recognition confidence can indicate the accuracy of text content recognition results to a certain extent; low recognition confidence indicates low reliability of the recognition results. For example, for certain character sets, such as characters like GRA after being rotated 180°, the probability of misrecognition is high, and the recognition confidence is generally low. Based on this, it is considered that the text in the input image corresponding to the recognition result below a certain recognition confidence threshold is unreadable.

[0069] According to embodiments of this disclosure, the text content recognition result includes multiple recognized texts; the method may further include:

[0070] Based on a preset recognition threshold determined according to the character length of the preset text, the target recognition text for the target text image is determined from multiple recognition texts; the detection result for the material to be detected is generated based on the text content recognition result, including: generating the detection result for the material to be detected based on the target recognition text.

[0071] The preset recognition threshold is determined based on the character length; different character lengths allow for different recognition thresholds. For shorter preset text, a higher preset recognition threshold can be set. For example, when the character length is 1, the preset recognition threshold can be set to 0.99; when the character length is 3, the preset recognition threshold can be set to 0.95.

[0072] By using a preset recognition threshold to perform secondary filtering on the text content recognition results, false detections can be effectively eliminated, which helps to improve recognition accuracy.

[0073] According to embodiments of this disclosure, orientation recognition of a target text image in an image of a material to be detected may include: using a text orientation classification model to identify the orientation of the target text image. Since the text orientation classification model does not support text regions with a vertical orientation (height greater than width), the text region needs to be pre-rotated to a horizontal orientation (90°) before inputting into the model. Because the input to the text orientation classification model is a pre-rotated horizontal text region image, the output text orientation refers to the orientation relative to a readable text state (0°). A readable text state means that every character is readable from left to right. For example, if image 310 in Figure 3 represents a readable text state (0°), then the text orientation of image 320 is a 180° rotation, meaning that image 320 is rotated 180° relative to image 310.

[0074] Since the input to the text orientation classification model is a text region image pre-rotated to horizontal, it is difficult to accurately determine whether a text region image that is approximately square is a horizontal or vertical image, resulting in low orientation recognition accuracy. Therefore, this disclosure provides an orientation recognition method for approximately square text region images.

[0075] The text orientation classification model can use MobileNet_Large. The training dataset for the text orientation classification model can include a character set, specifically numbers, uppercase and lowercase letters, and spaces. Considering the recognition of symmetrical text, images of the original image in four directions—rotated 0°, rotated 180°, horizontally flipped, and vertically flipped—are added to the training dataset to enable the recognition of the orientation of horizontally and vertically flipped text.

[0076] According to embodiments of this disclosure, orientation recognition of a target text image in an image of a material to be detected includes: in response to the target text image satisfying a second preset condition, orientation recognition is performed on the target text image and a rotated text image, wherein the rotated text image is an image obtained by rotating the target text image by a preset angle; text recognition of the target text image and a supplementary text image includes: text recognition is performed on the target text image, the rotated text image, and the supplementary text image, wherein the supplementary text image includes an image of at least one of the target text image and the rotated text image processed according to a preset processing method.

[0077] The second preset condition can be that the target text graphic is an image that is close to a square. For example, the second preset condition can be that the aspect ratio of the target text image is within the range of [0.7, 1.5). If the aspect ratio of the target text image is within the range of [0.7, 1.5), it indicates that the icon text image is close to a square.

[0078] Orientation recognition for a near-square target text image may include: rotating the target text image by a preset angle, such as 90°, to obtain a rotated text image. Then, orientation recognition is performed on both the target text image and the rotated text image. Finally, the orientation with higher confidence is selected for the next step of text recognition.

[0079] The default processing method is the same as the default processing method described above, and will not be repeated here.

[0080] Figure 4 schematically illustrates a diagram of determining target-identified text according to an embodiment of the present disclosure.

[0081] As shown in Figure 4, the target text recognition method of this embodiment 400 is used to determine the target text when the target text image is close to a square. Specifically, it may include: first, acquiring a target text image 410; when the aspect ratio of the target text image 410 is within the range of [0.7, 1.5), rotating the target text image 410 by 90° to obtain a rotated text image 420; then, performing orientation recognition on the target text image 410 and the rotated text image 420 respectively to obtain a text orientation recognition result 430, which may include a first orientation and a second orientation. When the confidence scores for both the first and second orientations are less than 0.995, the text orientation classification is considered inaccurate. In this case, the target text image 410 and the rotated text image 420 are rotated by 180°, horizontally flipped, and vertically flipped, respectively, to obtain supplementary images 441, 442, 443, 444, 445, and 446. Then, text recognition is performed on the target text image 410, the rotated text image 420, and the supplementary images 441, 442, 443, 444, 445, and 446, respectively, to obtain recognized text 451, 452, and 456. 53. Recognize texts 454, 455, 456, 457, and 458; then, based on a preset recognition threshold of 0.8, select the texts 451, 452, 453, 454, 455, 456, 457, and 458 with a recognition confidence greater than 0.8, and obtain text content recognition result 460, which includes texts 451 and 453; then, based on a preset recognition threshold of 0.99, determine the texts with a recognition confidence greater than 0.99 from the text content recognition result 460 as target recognition texts 470.

[0082] By performing orientation recognition on approximately square target text images and on rotated text images (90° rotated), orientations with a confidence score ≥ 0.995 are taken as the text orientation of the target text image for the next step of text recognition, thus improving recognition accuracy. In addition, when the orientation confidence scores in the text orientation recognition results are all less than 0.995, indicating inaccurate classification, text recognition is performed on the target text image and the rotated text image in four directions to further improve recognition accuracy.

[0083] According to embodiments of this disclosure, the method further includes: in response to the target text image not meeting a second preset condition, performing orientation recognition on the target text image to obtain a text orientation recognition result; in response to the text orientation recognition result not meeting a third preset condition, performing text recognition on the target text image and a supplementary text image for the target text image to obtain a text content recognition result.

[0084] If the target text image does not meet the second preset condition, it indicates that the aspect ratio of the target text image is less than 0.7 or greater than 1.5, meaning the target text image is not an approximately square image. This allows for accurate determination of whether the target text image is horizontal or vertical. In this case, if the target text image is horizontal, orientation recognition is performed directly to obtain the text orientation recognition result; if the target text image is vertical, it is rotated 90° before orientation recognition to obtain the text orientation recognition result.

[0085] The third preset condition is configured for target text images with an aspect ratio less than 0.7 or greater than or equal to 1.5. The third preset condition can be a direction confidence score greater than or equal to 0.95 or less than 0.3. When the direction confidence score is greater than or equal to 0.3 and less than 0.95, it indicates that the direction recognition is inaccurate, and text recognition of the target text image needs to be performed in all four directions.

[0086] It should be noted that when the confidence score for the direction in which the text is oriented is less than 0.3, it indicates that the target text image is not text and should be directly removed. When the confidence score for the direction in which the text is oriented is greater than or equal to 0.95, it indicates that the text orientation is accurately identified, and text recognition can be performed directly on the target text image without the need for recognition in all four directions.

[0087] Figure 5 schematically illustrates a diagram of determining target-identified text according to another embodiment of the present disclosure.

[0088] As shown in Figure 5, the method for determining the target text in this embodiment 500 is a method for determining the target text when the target text image is horizontal text. Specifically, it may include: first, acquiring the target text image 510; when the aspect ratio of the target text image 510 is less than 0.7, directly performing orientation recognition on the target text image 510 to obtain the text orientation recognition result 520; when the direction confidence of the text orientation in the text orientation recognition result 520 is greater than or equal to 0.95, it indicates that the orientation classification is accurate, and directly recognizing the sandwich text in the target text image to obtain the target text 540. When the confidence score of the text orientation in the text orientation recognition result 520 is within the range of [0.3, 0.95), it indicates that the orientation classification is inaccurate. In this case, the target text image 510 is rotated 180°, horizontally flipped, and vertically flipped to obtain supplementary images 531, 532, and 533. Then, text recognition is performed on the target text image 510, supplementary images 531, 532, and 533 respectively to obtain recognized text 551 and recognized text 552. Text 553 and text 554 are identified; then, based on a preset identification threshold of 0.8, texts with a confidence level greater than 0.8 among texts 551, 552, 553, and 554 are selected to obtain text content identification result 560, which may include texts 551 and 552; then, based on a preset identification threshold of 0.99, texts with a confidence level greater than 0.99 are determined from text identification result 560 as target identification text 570.

[0089] In some examples, before executing the material detection method shown in Figure 2, a black image determination may also be included. Specifically, this may include performing grayscale processing on the image of the material to be detected to obtain a grayscale processed material image; and sending a black image anomaly alarm in response to the average pixel value of the grayscale processed material image being less than a second pixel threshold.

[0090] If the average pixel value of the material image after grayscale processing is less than the second pixel threshold, it indicates that the product is not lit up, the overall screen is black, and text cannot be seen or material detection can be performed. At this time, there is no need to execute the material detection direction, and the alarm of black image abnormality is returned directly, pending manual confirmation.

[0091] The second pixel threshold can be determined according to actual needs, for example, it can be 10.

[0092] In some of these examples, the material detection image to be detected needs to be preprocessed before the material detection method shown in Figure 2 is executed. This may include the following steps.

[0093] Black border cropping: In the images captured by the AOI system camera, some points of interest (POLs) have a certain range of black background areas around them. To improve text detection accuracy and facilitate the separate detection of fixed values ​​(Mark values) and serial numbers, the black borders need to be pre-cropped. Specifically, this can include: first, compressing the image of the material to be detected to 512×512 to reduce preprocessing time; then performing grayscale processing and Gaussian blur denoising; next, padding, for example, extending the black border by about 10 pixels around the image to prevent accidental cropping of areas without black background in the original image; then, using Hough transform for line detection to calculate the horizontal bounding rectangle of all line segments; finally, scaling the compressed image back to the original size of the image of the material to be detected, and cropping the black borders on the original image of the material to be detected based on the bounding rectangles.

[0094] Contrast Adjustment: Some Mark values ​​located on the release film on the back of the POL (Positioning Object) are too faint to be easily discernible in photographs. To address this, adaptive contrast adjustment is performed on the image of the material to be inspected after cropping the black edges, making the mirrored Mark values ​​on the back clearer. Specifically, since the Marks are multi-dot matrix fonts and the AOI camera captures images with a high resolution (up to 10,000 x 10,000 pixels), direct compression can easily lead to unclear text. Therefore, the image of the material to be inspected is first Gaussian blurred, and then image enhancement techniques are used to enhance contrast and suppress image noise. For example, the CLAHE algorithm can be used for adaptive histogram equalization to enhance contrast.

[0095] In some of these examples, before performing the material detection method shown in Figure 2, it is necessary to preprocess the image of the material to be detected and then determine the target text image in the material to be detected.

[0096] According to embodiments of this disclosure, the method further includes: performing text detection on the image of the material to be detected to obtain at least one text detection box; determining a target text detection box within the at least one text detection box based on the height of the text detection box and the distance between the at least one text detection box, and determining the area corresponding to the target text box as the target text image.

[0097] Text detection on images of materials to be inspected can include using word-level text detection models to detect bounding boxes. The text detection module can employ DBNet, where the backbone network can use the lightweight PPLCNetV3.

[0098] To enable text detection models to detect text boxes in symmetrical text, enhancements can be made to the training data during model training. For example, mirrored text or double-column bitmap text can be added.

[0099] At least one text detection box includes a text detection box with a fixed value on the polarizer, a text detection box for a serial number, and a text detection box for a barcode. Therefore, it is necessary to remove the text detection boxes for serial numbers and barcodes. The specific process for removing serial numbers is described below.

[0100] Figure 6 schematically illustrates a flowchart of a method for determining a target text image according to an embodiment of the present disclosure.

[0101] As shown in Figure 6, the method for determining the target text image in this embodiment 600 includes operations SS610 to S640.

[0102] In operation S610, a distance threshold is determined based on the height of at least one text detection box.

[0103] In operation S620, text detection box matching is performed based on the distance threshold and the distance between at least one text detection box, to obtain at least one matching pair.

[0104] In operation S630, the target text detection box in the matching pair is determined based on the height of the text detection box.

[0105] In operation S640, the area corresponding to the target text box is determined to be the target text image.

[0106] In one example, determining the distance threshold based on the height of at least one text detection box may include averaging the heights of the at least one text detection box to obtain the distance threshold.

[0107] In another example, determining the distance threshold based on the height of at least one text detection box may include: clustering the text detection boxes based on their heights to obtain a preset number of cluster centers; and determining the target cluster center among the cluster centers as the distance threshold.

[0108] The preset number can be determined based on actual data; for example, the preset number can be 2. The target cluster center is the cluster center with the smallest value among the preset number of cluster centers, which is used as the distance threshold.

[0109] By clustering at least one text detection box based on its height and determining the cluster center with the smallest value as the distance threshold, the determined distance threshold is more accurate, which helps to improve the accuracy of material detection.

[0110] Figure 7 schematically illustrates a diagram of determining a matching pair according to an embodiment of the present disclosure.

[0111] As shown in Figure 7, this embodiment 700 includes text detection boxes 710, 720, and 730. Then, distance analysis is performed on text detection boxes 710 and 720 to obtain a first distance 740; distance analysis is performed on text detection boxes 720 and 730 to obtain a second distance 750; distance analysis is performed on text detection boxes 710 and 730 to obtain a third distance 760. Then, the first distance 740, second distance 750, and third distance 760 are compared with distance thresholds, and two text detection boxes with distances less than the distance thresholds are generated as matching pairs. For example, if the first distance 740 is less than the distance threshold, text detection box 710 and text detection box 720 are matched to obtain a matching pair. In this case, the first text detection box can be text detection box 710, and the second text detection box can be text detection box 720.

[0112] In some of these examples, determining the target text detection box in a matching pair based on the height of the text detection box may include: determining the text detection box with the higher height in the matching pair as a fixed value, i.e., the target text detection box.

[0113] In other examples, determining the target text detection box in the matching pair based on the height of the text detection boxes may include: enlarging the first text detection box in the matching pair to obtain an enlarged first text detection box; determining the intersection area between the second text detection box in the matching pair and the enlarged first text detection box; wherein, determining the target text detection box in the matching pair based on the height of the text detection boxes includes: determining the target text detection box in the matching pair based on the intersection area, the height of the first text detection box, and the height of the second text detection box.

[0114] Enlarging the first text detection box may include enlarging the first text detection box by a preset factor, such as 3.5 times.

[0115] The intersecting region can be represented by the intersection-union ratio of the second text detection box and the enlarged first text detection box, where the denominator is the area of ​​the second text detection box.

[0116] Determining the target text detection box in the matching pair based on the intersection region, the height of the first text detection box, and the height of the second text detection box may include: determining that the sum of the regions is greater than the intersection-union ratio threshold, for example, 0.4, and that the height ratio of the second text detection box to the first text detection box is less than the height ratio threshold, for example, 0.9.

[0117] Figure 8 schematically illustrates a target text detection box determined according to an embodiment of the present disclosure.

[0118] As shown in Figure 8, this embodiment 800 can be executed after the matching pair is determined as shown in Figure 7. Specifically, it can include: the matching pair 810 includes a first text detection box and a second text detection box 812; the first text detection box 811 is enlarged, for example, by a factor of 3.5, to obtain an enlarged first text detection box 820; then the intersection-union ratio 830 of the enlarged first text detection box 820 and the second text detection box 812 is determined, and the height ratio 840 is determined based on the height of the first text detection box 811 and the height of the second text detection box 812, wherein the height ratio is the height ratio of the second text detection box to the first text detection box; then, if the intersection-union ratio 830 is greater than the intersection-union ratio threshold and the height ratio 840 is less than the height ratio threshold, the first text detection box 811 is determined as the target text detection box 850.

[0119] In one example, the method for removing serial numbers from text detection boxes can specifically include: First, sorting all text detection boxes (bboxes) for the image of the material to be detected by height, calculating two cluster center values ​​for the box height, and taking the smaller of the cluster center values ​​as the distance threshold th2. Then, sequentially traversing the bboxes, determining whether the current box should be removed. The removal criteria are as follows: if the distance d1 between box[i] and box[i+1] is greater than th3 (indicating that they are far apart), th3 is taken as th2*10, and the next traversal is performed; otherwise, box[i] is enlarged by 3.5 times, and the IoU with box[i+1] is calculated (the denominator is the area of ​​box[i+1]). If the IoU is greater than 0.4, and the height of box[i] * 0.9 is greater than the height of box[i+1], then box[i] and box[i+1] are considered to be paired, where box[i+1] is the serial number to be removed; otherwise, the next traversal is performed.

[0120] Some POL protective films have barcodes affixed to them. During image preprocessing, the contrast is adjusted, making the text on the barcodes visible. During text detection, some barcodes will be detected. Therefore, when there are multiple target text images, barcode removal can be performed. Specifically, this can include: selecting target text images whose average pixel value meets a first pixel threshold as the filtered target text images; and performing orientation recognition on the target text images in the material image to be detected, including orientation recognition on the filtered target text images.

[0121] The first pixel threshold can be 0.5 times the average value of the image of the material to be detected. Since the text brightness of the barcode is low, when the average pixel value of the image is less than the first pixel threshold, it can be considered as text in the barcode area and removed.

[0122] Before implementing the material detection method shown in Figure 2, a target production line configuration matching template can also be included, i.e., setting a whitelist of Mark values ​​for the target production line. The matching template is divided into a single Mark template and a double Mark template. A single Mark means that there is only one Mark value on each POL, and this Mark value can have multiple values; a double Mark means that there are two Mark values, one positive and one negative, on each POL, and each Mark value can have multiple values.

[0123] According to embodiments of this disclosure, the method for configuring a matching template may further include: displaying a template configuration interface in response to a user's configuration operation for a matching template of a target production line; displaying a template editing interface for the target template mode in response to a user selecting a target template mode from preset template modes; and generating a matching template for the target production line based on the editing information entered by the user in the template editing interface.

[0124] The preset template modes include single-mark templates and double-mark templates. Only one template mode can be selected for a single production line.

[0125] Figure 9 schematically illustrates a template configuration interface according to an embodiment of the present disclosure.

[0126] As shown in Figure 9, the template configuration interface 910 of this embodiment includes a product name for selecting the target production line; and a Mark quantity, which represents the preset template mode, allowing selection of a single-Mark template mode or a double-Mark template mode. After the user selects the target template mode, a template editing interface is displayed. For example, the template editing interface 920 shown in Figure 9 is the template editing interface corresponding to the double-Mark template mode. For the double-Mark template mode, the relative positions of the two Marks can be moved in the visual editing window 921 of the template editing interface 920, such as horizontally, vertically, or diagonally.

[0127] Users can enter text information in the template editing interface 920 to generate preset text. For example, if "ACD857" and "869" are entered in the input box corresponding to the Mark1 whitelist, and "EFD866" and "742" are entered in the input box corresponding to the Mark2 whitelist, then "ACD857", "869", "EFD866", and "742" will be the preset text.

[0128] Each Mark value requires an orientation setting. Users can perform at least one orientation adjustment operation on the preset text, such as rotating 0°, rotating 180°, horizontally flipping, or vertically flipping, to generate the text orientation for the preset text. For example, after entering a Mark value and confirming, the text is surrounded by a red bubble, indicating that the orientation needs to be set. Moving the mouse to the bubble area will bring up a fixed value option as shown in Figure 10. The fixed value option 1010 can include 0°, 45°, 90°, 180°, horizontally flipping, and vertically flipping. If the fixed value does not meet the requirements, right-clicking on the bubble area will bring up a visual window to open the editing state, as shown in Figure 11. Users can manually rotate, horizontally flip, and vertically flip the Mark value, such as Mark value "869". Users can perform multiple operations, and the algorithm accumulates and calculates the final state to obtain the preset text orientation. In Figure 11, the orientation of the Mark value is read from left to right, horizontally 0°, with no flipping.

[0129] According to embodiments of this disclosure, a user can perform multiple operations, and the algorithm accumulates and calculates the final state to obtain the preset text orientation. This can include: based on at least one orientation adjustment operation performed by the user on the template configuration interface for the preset text, converting the at least one orientation adjustment operation into a preset operation according to preset rules, wherein the preset operation includes a flip operation and a rotation operation; determining the text orientation of the preset text based on the flip type of the flip operation and the rotation angle of the rotation operation. Specifically, the user's orientation adjustment operation on the preset text on the front end includes rotation and flip, where the rotation angle can be any value; however, the text orientation classification model only determines four text orientation directions for horizontal text images: 0°, 180°, horizontal flip (fliplr), and vertical flip (flipud); therefore, it is necessary to convert the user-defined value into a value supported by the text orientation classification model before comparison.

[0130] The specific conversion method may be: when a user performs matching template setting, record the user's multiple flipping and rotation operations, and finally accumulate the operations into one operation of flipping first and then rotating [flip type (flipflag), rotation angle (rotate_angle)], wherein flipflag has three values: 'noflip', 'fliplr' and 'flipud', which represent no flipping, horizontal flipping and vertical flipping respectively; rotate_angle is an integer ranging from -180 to 180, where a positive number represents clockwise rotation and a negative number represents counterclockwise rotation. It should be noted that different sequences of rotation and flipping result in different final orientations of the preset text, two identical flips can cancel each other out, one fliplr and one flipud are equivalent to a 180° rotation, rotation first followed by flipping can be converted into flipping first followed by rotation by simply taking the negative of the angle, accumulation is performed step by step according to this logic, and finally rotate_angle is normalized to (-180, 180]). When flipflag is noflip, if rotate_angle∈(-67.5, 112.5], the corresponding text orientation is 0°, otherwise it is 180°; when flipflag is fliplr, if rotate_angle∈(-67.5, 112.5], the corresponding text orientation is fliplr, otherwise it is flipud; when flipflag is flipud, if rotate_angle∈(-67.5, 112.5], the corresponding text orientation is flipud, otherwise it is fliplr; the text orientation determined according to this logic is then compared with the orientation of the recognized text output by the text orientation classification model.

[0131] A suspected alarm digit count needs to be set for each Mark value; after a Mark value is input and confirmed, a digit input box is automatically expanded, and a user can input an alarm digit count to generate an edit distance threshold for preset text. If the number of mismatched digits in Mark matching is less than the alarm digit count, no alarm is triggered; if the alarm digit count ≤ the number of mismatched digits in Mark matching < 0.8 times the total Mark digit count, a suspected alarm is triggered; if the number of mismatched digits in Mark matching ≥ 0.8 times the total Mark digit count, an alarm is triggered. Setting an alarm position allows fault tolerance for algorithm errors and printing defects, which helps improve the accuracy of material detection.

[0132] After the setting of the preset text, the text orientation of the preset text and the edit distance threshold is completed, the user clicks confirm, and then a matching template is generated.

[0133] According to embodiments of this disclosure, the text content recognition result includes at least one recognized text; generating a detection result for the material to be detected based on the text content recognition result includes: performing quantity matching on at least one recognized text and a preset text in a preset matching template; in response to a successful quantity matching, performing text matching on at least one recognized text and the preset text respectively to obtain a detection result.

[0134] According to embodiments of this disclosure, the method further includes: in response to a quantity matching failure, sending a second alarm message and mapping address information of the image of the material to be detected.

[0135] A successful count match indicates that the number of recognized texts is the same as the preset number of texts. A failed count match indicates that the number of recognized texts is different from the preset number of texts. For example, multiple texts are recognized under a single Mark template, or only one text is recognized under a double Mark template.

[0136] The mapping address information of the material image to be detected can be displayed as a link on the front-end interface. Users can view the material image by clicking the link.

[0137] According to embodiments of this disclosure, the identified text includes multiple texts, and the identified texts have a text orientation; the method further includes: performing relative position matching on the multiple identified texts and a preset text in a preset matching template; in response to successful relative position matching, performing orientation matching on the identified texts and the preset texts at the corresponding positions; in response to successful orientation matching, performing text matching on the identified texts and the preset texts to obtain a detection result.

[0138] Relative position matching can include vertical, horizontal, or diagonal distribution. For templates with two Mark values, the template keywords are first read to distinguish whether the Mark values ​​are distributed vertically or horizontally. For templates with vertical distribution, the recognized text is sorted according to the Y-axis coordinate; for templates with horizontal distribution, the recognized text is sorted according to the X-axis coordinate, and then matched with preset text.

[0139] According to embodiments of this disclosure, text matching between identified text and preset text may include: determining the edit distance between the identified text and preset text in a preset matching template, thereby obtaining at least one edit distance; and sending a first alarm message in response to a target edit distance in at least one edit distance not meeting an edit distance threshold.

[0140] The target edit distance represents the minimum edit distance among at least one edit distance. The edit distance threshold is the number of alarm bits configured when matching the template.

[0141] Edit distance represents the minimum number of single-character editing operations required to convert a string (e.g., recognized text) into another string (e.g., preset text). These operations include insertion, deletion, and replacement, and correspond to the missed, over-recognized, and misrecognized characters in the text recognition model. For example, if the recognized text is "123EFG" and the corresponding preset text is "23EFG", the error length here is 1, not 6 bits for a one-to-one comparison.

[0142] Determining the edit distance between the identified text and the preset text in the preset matching template, and obtaining at least one edit distance, may include: for each identified text, traversing the list of preset texts in the matching template and calculating the minimum edit distance; if the identified text returned by the algorithm has multiple orientation results, for each preset text in the matching template list, calculating the edit distance Dist with the identified text in multiple orientations, and taking the minimum Dist value as the final edit distance.

[0143] In actual detection, there are cases where serial numbers (small characters) cannot be removed through post-processing. In such cases, they will be identified together with fixed values. To address this, a unified matching template is used to find the recognition text with the smallest edit distance. Then, the recognition text is truncated according to the character length of the matching template, and the edit distance is recalculated for the truncated recognition text.

[0144] The first alarm information can include alarms or suspected alarms. Specific alarm conditions are described in the matching template configuration section and will not be repeated here.

[0145] It should be noted that if every character in the preset text of the matching template belongs to the vertically symmetrical character set '038BCDEHIKOXo', then 0° is equivalent to vertical flipping. Since vertical symmetry does not affect the reading order, no alarm will be triggered as long as one of the recognized texts in the two orientations is the same as the preset text.

[0146] Figure 12 schematically illustrates a diagram of determining the detection result according to an embodiment of the present disclosure.

[0147] As shown in Figure 12, the text content recognition result of this embodiment 1200 includes multiple recognized texts, such as recognized text 1211 and recognized text 1212. First, the recognized texts in the text content recognition result 1210 are matched in quantity 1230 with the preset texts in the preset matching template 1220. If the quantities are inconsistent, an alarm message 1270 is sent directly. If the quantities are consistent, the recognized texts in the text content recognition result are matched in relative position 1240 with the preset texts in the preset matching template 1220. If the relative positions do not match, an alarm message 1270 is sent. If the relative positions match successfully, for example, the recognized texts are distributed horizontally, and the preset texts are also distributed horizontally, then the recognized texts in the text content recognition result 1210 are matched in relative position. The text is matched with the preset text in the preset matching template 1220 in terms of orientation 1250. If the orientation matching fails, an alarm message is sent 1270. If the matching is successful, the identified text is matched with the corresponding preset text to determine the edit distance between the identified text and the preset text, thereby obtaining the detection result 1260. The detection result may include edit distance 1261 and edit distance 1262. For example, edit distance 1261 is the edit distance between the identified text 1211 and the preset text 1221, and edit distance 1262 is the edit distance between the identified text 1212 and the preset text 1222.

[0148] Figure 13 schematically illustrates a material detection method according to another embodiment of the present disclosure.

[0149] As shown in Figure 13, the material detection method of this embodiment 1300 includes acquiring an image of the material to be detected 1310, performing black image judgment on the image of the material to be detected 1320, and if the black image judgment is normal, then preprocessing the image of the material to be detected 1310 1330, which may include black border cropping and contrast adjustment to obtain a preprocessed image of the material to be detected; then, performing text box detection on the preprocessed image of the material to be detected 1340 to obtain at least one text detection box; then, performing orientation recognition on the text screenshots at the positions corresponding to the at least one text detection box 1350 to obtain a text orientation recognition result; then, performing text content recognition on the text screenshots 1360 to obtain at least one recognized text; then, performing Mark matching 1370 on the recognized text and the preset text in the matching template to obtain a detection result; if the detection result is abnormal, then issuing an alarm 1380.

[0150] In embodiments of this disclosure, multiple orientation recognitions are performed on symmetrical characters, and a matching template with preset text orientations is used to verify the Mark on the polarizer. Furthermore, specific preprocessing is applied to the POL image captured by the AOI system to improve the Mark recognition accuracy.

[0151] Based on the above-described material detection method, this disclosure also provides a material detection device. The device will be described in detail below with reference to Figure 14.

[0152] Figure 14 schematically illustrates a structural block diagram of a material detection apparatus according to an embodiment of the present disclosure.

[0153] As shown in Figure 14, the material detection device 1400 includes a text orientation recognition module 1410, a text content recognition module 1420, and a first generation module 1430.

[0154] The text orientation recognition module 1410 is used to perform orientation recognition on the target text image in the image of the material to be detected, and obtain the text orientation recognition result.

[0155] The text content recognition module 1420 is used to perform text recognition on the target text image and the supplementary text image in response to the text orientation recognition result not meeting the first preset condition, and to obtain the text content recognition result. The supplementary text image includes the image after the target text image has been processed according to the preset processing method.

[0156] The first generation module 1430 is used to generate detection results for the material to be detected based on the text content recognition results.

[0157] According to embodiments of this disclosure, the above-described apparatus further includes a text detection module and a first determination module.

[0158] The text detection module is used to perform text detection on the image of the material to be detected, and obtain at least one text detection box.

[0159] The first determining module is used to determine the target text detection box in at least one text detection box based on the height of the text detection box and the distance between at least one text detection box, and to determine the area corresponding to the target text box as the target text image.

[0160] According to embodiments of this disclosure, the first determining module includes: a first determining submodule, a matching submodule, and a second determining submodule.

[0161] The first determination submodule is used to determine a distance threshold based on the height of at least one text detection box.

[0162] The matching submodule is used to perform text detection box matching based on a distance threshold and the distance between at least one text detection box, so as to obtain at least one matching pair.

[0163] The second determination submodule is used to determine the target text detection box in the matching pair based on the height of the text detection box.

[0164] According to embodiments of this disclosure, the first determining submodule includes: a clustering processing unit and a first determining unit.

[0165] The clustering processing unit is used to cluster text detection boxes based on the height of at least one text detection box to obtain a preset number of cluster centers.

[0166] The first determining unit is used to determine the target cluster center in the cluster center as the distance threshold.

[0167] According to embodiments of this disclosure, the above-described apparatus further includes: an amplification processing module and a second determination module.

[0168] The enlargement processing module is used to enlarge the first text detection box in the matching pair to obtain the enlarged first text detection box.

[0169] The second determining module is used to determine the intersection area between the second text detection box and the enlarged first text box in the matching pair.

[0170] According to embodiments of this disclosure, the second determining submodule is further configured to determine the target text detection box in the matching pair based on the intersecting region, the height of the first text detection box, and the height of the second text detection box.

[0171] According to embodiments of this disclosure, the text content recognition result includes at least one recognized text.

[0172] According to embodiments of this disclosure, the first generation module includes: a quantity matching submodule and a first text matching submodule.

[0173] The quantity matching submodule is used to perform quantity matching between at least one identified text and a preset text in a preset matching template.

[0174] The first text matching submodule is used to perform text matching between at least one identified text and a preset text in response to a successful quantity matching, so as to obtain the detection result.

[0175] According to embodiments of this disclosure, the identified text includes a plurality of texts, and the identified texts have a text orientation.

[0176] According to embodiments of this disclosure, the first generation module further includes: a relative position matching submodule, an orientation matching submodule, and a second text matching submodule.

[0177] The relative position matching submodule is used to perform relative position matching between multiple recognized texts and preset texts in preset matching templates.

[0178] The orientation matching submodule is used to perform orientation matching between the identified text and the preset text at the corresponding position in response to a successful relative position match.

[0179] The second text matching submodule is used to perform text matching between the identified text and the preset text in response to a successful orientation match, and obtain the detection result.

[0180] According to embodiments of this disclosure, the second text matching submodule includes: a second determination unit and an alarm unit.

[0181] The second determining unit is used to determine the editing distance between the identified text and the preset text in the preset matching template, and obtain at least one editing distance.

[0182] An alarm unit is used to send a first alarm message in response to a target edit distance in at least one edit distance not meeting an edit distance threshold.

[0183] According to embodiments of this disclosure, the above-described apparatus further includes a conversion module and a third determination module.

[0184] The conversion module is used to convert at least one orientation adjustment operation performed by the user on the template configuration interface into a preset operation according to preset rules. The preset operations include flipping and rotating operations.

[0185] The third module determines the text orientation of the preset text based on the flip type of the flip operation and the rotation angle of the rotation operation.

[0186] According to embodiments of this disclosure, the above-described apparatus further includes: a first alarm module.

[0187] The first alarm module is used to send a second alarm message and the mapping address information of the image of the material to be detected in response to a quantity matching failure.

[0188] According to embodiments of this disclosure, the above-described apparatus further includes: a first display module, a second display module, and a second generation module.

[0189] The first display module is used to respond to the user's configuration operation for the matching template of the target production line and display the template configuration interface.

[0190] The second display module is used to display the template editing interface for the target template mode in response to the user's selection of the target template mode.

[0191] The second generation module is used to generate a matching template for the target production line based on the editing information entered by the user in the template editing interface.

[0192] According to embodiments of this disclosure, the second generation module includes: a first generation submodule, a second generation submodule, a third generation submodule, and a fourth generation submodule.

[0193] The first generation submodule is used to generate preset text in response to user-inputted text information.

[0194] The second generation submodule is used to generate text orientation for the preset text in response to at least one orientation adjustment operation performed by the user for the preset text.

[0195] The third generation submodule is used to generate an edit distance threshold for a preset text in response to the number of alarm bits entered by the user.

[0196] The fourth generation submodule is used to generate a matching template based on the preset text, the text orientation of the preset text, and the editing distance threshold.

[0197] According to embodiments of this disclosure, the text content recognition result includes multiple recognized texts.

[0198] According to embodiments of this disclosure, the above-described apparatus further includes a fourth determining module.

[0199] The fourth determination module is used to determine the target recognition text for the target text image from multiple recognition texts based on a preset recognition threshold determined according to the character length of the preset text.

[0200] According to embodiments of this disclosure, the first generation module is further configured to generate detection results for the material to be detected based on the target recognition text.

[0201] According to an embodiment of this disclosure, the text orientation recognition module includes: a first text orientation recognition submodule.

[0202] The first text orientation recognition submodule is used to perform orientation recognition on the target text image and the rotated text image in response to the target text image meeting the second preset condition, wherein the rotated text image is the image obtained by rotating the target text image by a preset angle.

[0203] According to embodiments of this disclosure, the text content recognition module is further configured to perform text recognition on the target text image, the rotated text image, and the supplementary text image, wherein the supplementary text image includes an image after processing at least one of the target text image and the rotated text image according to a preset processing method.

[0204] According to embodiments of this disclosure, the text orientation recognition module further includes: a second text orientation recognition submodule and a text content recognition submodule.

[0205] The second text orientation recognition submodule is used to perform orientation recognition on the target text image in response to the target text image not meeting the second preset condition, and obtain the text orientation recognition result.

[0206] The text content recognition submodule is used to perform text recognition on the target text image and the supplementary text image for the target text image in response to the text orientation recognition result not meeting the third preset condition, so as to obtain the text content recognition result.

[0207] According to embodiments of this disclosure, the target text image includes multiple images.

[0208] According to embodiments of this disclosure, the above-described apparatus further includes a screening module.

[0209] The filtering module is used to select target text images from multiple target text images whose average pixel value meets the first pixel threshold as the filtered target text images.

[0210] According to embodiments of this disclosure, the text orientation recognition module is also used to perform orientation recognition on the filtered target text image and the rotated text image.

[0211] According to embodiments of this disclosure, the above-described apparatus further includes: a grayscale processing module and a second alarm module.

[0212] The grayscale processing module is used to perform grayscale processing on the image of the material to be detected, so as to obtain the grayscale processed image of the material.

[0213] The second alarm module is used to send a black image abnormality alarm in response to the fact that the average pixel value of the material image after grayscale processing is less than the second pixel threshold.

[0214] According to embodiments of this disclosure, any plurality of modules among the text orientation recognition module 1410, text content recognition module 1420, and first generation module 1430 can be combined into one module, or any one of these modules can be split into multiple modules. Alternatively, at least part of the functionality of one or more of these modules can be combined with at least part of the functionality of other modules and implemented in one module. According to embodiments of this disclosure, at least one of the text orientation recognition module 1410, text content recognition module 1420, and first generation module 1430 can be at least partially implemented as hardware circuitry, such as a field-programmable gate array (FPGA), a programmable logic array (PLA), a system-on-a-chip, a system-on-a-substrate, a system-on-package, an application-specific integrated circuit (ASIC), or implemented in hardware or firmware by any other reasonable means of integrating or packaging the circuitry, or implemented in software, hardware, or firmware, or in any suitable combination of any of these three implementation methods. Alternatively, at least one of the text orientation recognition module 1410, text content recognition module 1420, and first generation module 1430 may be implemented at least partially as a computer program module, which can perform corresponding functions when the computer program module is run.

[0215] It should be noted that the material detection device part in the embodiments of this disclosure corresponds to the material detection method part in the embodiments of this disclosure. For a detailed description of the material detection device part, please refer to the material detection method part, which will not be repeated here.

[0216] Figure 15 schematically illustrates a block diagram of an electronic device suitable for implementing an information interaction method according to an embodiment of the present disclosure.

[0217] As shown in FIG. 15, an electronic device 1500 according to an embodiment of the present disclosure includes a processor 1501, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 1502 or a program loaded from a storage portion 1508 into a random access memory (RAM) 1503. The processor 1501 may include, for example, a general-purpose microprocessor (e.g., a CPU), an instruction set processor and / or an associated chipset and / or a special-purpose microprocessor (e.g., an application-specific integrated circuit (ASIC)), etc. The processor 1501 may also include onboard memory for caching purposes. The processor 1501 may include a single processing unit or multiple processing units for performing different actions of the method flow according to an embodiment of the present disclosure.

[0218] RAM 1503 stores various programs and data required for the operation of electronic device 1500. Processor 1501, ROM 1502, and RAM 1503 are interconnected via bus 1504. Processor 1501 performs various operations of the method flow according to embodiments of the present disclosure by executing programs in ROM 1502 and / or RAM 1503. It should be noted that the programs may also be stored in one or more memories other than ROM 1502 and RAM 1503. Processor 1501 may also perform various operations of the method flow according to embodiments of the present disclosure by executing programs stored in said one or more memories.

[0219] According to embodiments of this disclosure, the electronic device 1500 may further include an input / output (I / O) interface 1505, which is also connected to a bus 1504. The electronic device 1500 may also include one or more of the following components connected to the input / output (I / O) interface 1505: an input section 1506 including a keyboard, mouse, etc.; an output section 1507 including a cathode ray tube (CRT), liquid crystal display (LCD), etc., and a speaker, etc.; a storage section 1508 including a hard disk, etc.; and a communication section 1509 including a network interface card such as a LAN card, modem, etc. The communication section 1509 performs communication processing via a network such as the Internet. A drive 1510 is also connected to the input / output (I / O) interface 1505 as needed. A removable medium 1511, such as a disk, optical disk, magneto-optical disk, semiconductor memory, etc., is installed on the drive 1510 as needed so that computer programs read from it can be installed into the storage section 1508 as needed.

[0220] According to embodiments of this disclosure, the method flow according to embodiments of this disclosure can be implemented as a computer software program. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a computer-readable storage medium, the computer program containing program code for performing the methods shown in the flowchart. In such embodiments, the computer program can be downloaded and installed from a network via communication section 1509, and / or installed from removable medium 1511. When the computer program is executed by processor 1501, it performs the functions defined in the system of embodiments of this disclosure. According to embodiments of this disclosure, the systems, devices, apparatuses, modules, units, etc., described above can be implemented by computer program modules.

[0221] This disclosure also provides a computer-readable storage medium, which may be included in the device / apparatus / system described in the above embodiments; or it may exist independently and not assembled into the device / apparatus / system. The computer-readable storage medium carries one or more programs that, when executed, implement the method according to the embodiments of this disclosure.

[0222] According to embodiments of this disclosure, the computer-readable storage medium can be a non-volatile computer-readable storage medium. Examples include, but are not limited to: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this disclosure, the computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.

[0223] For example, according to embodiments of this disclosure, a computer-readable storage medium may include one or more memories other than the ROM 1502 and / or RAM 1503 described above.

[0224] Embodiments of this disclosure also include a computer program product comprising a computer program containing program code for performing the methods provided in the embodiments of this disclosure. When the computer program product is run on an electronic device, the program code enables the electronic device to perform the material detection provided in the embodiments of this disclosure.

[0225] When the computer program is executed by the processor 1501, it performs the functions defined in the system / apparatus of this disclosure embodiments. According to embodiments of this disclosure, the systems, apparatuses, modules, units, etc., described above can be implemented by computer program modules.

[0226] In one embodiment, the computer program may rely on a tangible storage medium such as an optical storage device or a magnetic storage device. In another embodiment, the computer program may also be transmitted and distributed in the form of signals over a network medium, and may be downloaded and installed via the communication section 1509, and / or installed from the removable medium 1511. The program code contained in the computer program can be transmitted using any suitable network medium, including but not limited to: wireless, wired, etc., or any suitable combination thereof.

[0227] According to embodiments of this disclosure, program code for executing the computer programs provided in embodiments of this disclosure can be written in any combination of one or more programming languages. Specifically, these computational programs can be implemented using high-level procedural and / or object-oriented programming languages, and / or assembly / machine languages. Programming languages ​​include, but are not limited to, languages ​​such as Java, C++, Python, "C", or similar programming languages. The program code can execute entirely on the user's computing device, partially on the user's device, partially on a remote computing device, or entirely on a remote computing device or server. In cases involving remote computing devices, the remote computing device can be connected to the user's computing device via any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computing device (e.g., via the Internet using an Internet service provider).

[0228] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in a block diagram or flowchart, and combinations of blocks in a block diagram or flowchart, may be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions. Those skilled in the art will understand that the features described in the various embodiments of the present disclosure can be combined and / or combined in various ways, even if such combinations are not explicitly described in the present disclosure. In particular, the features described in the various embodiments of this disclosure may be combined and / or combined in various ways without departing from the spirit and teachings of this disclosure. All such combinations and / or combinations fall within the scope of this disclosure.

[0229] The embodiments of this disclosure have been described above. However, these embodiments are for illustrative purposes only and are not intended to limit the scope of this disclosure. Although various embodiments have been described above, this does not mean that the measures in the various embodiments cannot be used advantageously in combination. Various substitutions and modifications can be made by those skilled in the art without departing from the scope of this disclosure, and all such substitutions and modifications should fall within the scope of this disclosure.

Claims

1. A method for detecting materials, comprising: Orientation recognition is performed on the target text image in the image of the material to be detected to obtain the text orientation recognition result; In response to the text orientation recognition result not meeting the first preset condition, text recognition is performed on the target text image and the supplementary text image to obtain the text content recognition result, wherein the supplementary text image includes an image after the target text image has been processed according to a preset processing method; as well as The detection results for the material to be detected are generated based on the text content recognition results.

2. The method according to claim 1, further comprising: Text detection is performed on the image of the material to be detected to obtain at least one text detection box; Based on the height of the text detection boxes and the distance between at least one text detection box, a target text detection box is determined within at least one text detection box, and the region corresponding to the target text box is determined as the target text image.

3. The method according to claim 2, wherein, The step of determining the target text detection box in at least one text detection box based on the height of the text detection box and the distance between at least one text detection box includes: A distance threshold is determined based on the height of at least one of the text detection boxes; Based on the distance threshold and the distance between at least one of the text detection boxes, text detection box matching is performed to obtain at least one matching pair; The target text detection box in the matching pair is determined based on the height of the text detection box.

4. The method according to claim 3, wherein, Determining the distance threshold based on the height of at least one of the text detection boxes includes: Based on the height of at least one of the text detection boxes, the text detection boxes are clustered to obtain a preset number of cluster centers; The target cluster center among the cluster centers is determined as the distance threshold.

5. The method according to claim 3, further comprising: The first text detection box in the matching pair is enlarged to obtain an enlarged first text detection box; Determine the intersection area between the second text detection box and the enlarged first text box in the matching pair; The step of determining the target text detection box in the matching pair based on the height of the text detection box includes: The target text detection box in the matching pair is determined based on the intersecting region, the height of the first text detection box, and the height of the second text detection box.

6. The method according to claim 1, wherein, The text content recognition result includes at least one recognized text; The step of generating detection results for the material to be detected based on the text content recognition results includes: Perform quantity matching on at least one of the identified texts and preset texts in the preset matching template; In response to a successful quantity match, at least one of the identified texts is matched with the preset text to obtain the detection result.

7. The method according to claim 6, wherein, The identified text includes multiple texts, and the identified text has a text orientation; The method further includes: Relative position matching is performed on multiple identified texts and preset texts in preset matching templates; In response to a successful relative position match, orientation matching is performed between the identified text and the preset text at the corresponding position; In response to a successful orientation match, text matching is performed between the identified text and the preset text to obtain the detection result.

8. The method according to claim 7, wherein, The detection result obtained by performing text matching between the identified text and the preset text includes: Determine the edit distances between the identified text and preset texts in the preset matching template to obtain at least one edit distance; In response to the fact that the target edit distance of at least one of the edit distances does not meet the edit distance threshold, a first alarm message is sent.

9. The method according to claim 7, further comprising: Based on at least one orientation adjustment operation performed by the user on the template configuration interface for the preset text, the at least one orientation adjustment operation is converted into a preset operation according to a preset rule, wherein the preset operation includes a flip operation and a rotation operation; The text orientation of the preset text is determined based on the flip type of the flip operation and the rotation angle of the rotation operation.

10. The method of claim 6, further comprising: In response to a quantity matching failure, a second alarm message and the mapping address information of the image of the material to be detected are sent.

11. The method according to claim 1, further comprising: In response to the user's configuration operation for the matching template of the target production line, the template configuration interface is displayed; In response to the user selecting a target template mode from the preset template modes, a template editing interface for the target template mode is displayed; Based on the editing information entered by the user in the template editing interface, a matching template for the target production line is generated.

12. The method according to claim 11, wherein, The step of generating a matching template for the target production line based on the editing information entered by the user in the template editing interface includes: Responding to user-inputted text, generate preset text; In response to at least one orientation adjustment operation performed by the user on the preset text, generate a text orientation for the preset text; In response to the number of alarm bits entered by the user, an edit distance threshold is generated for the preset text; The matching template is generated based on the preset text, the text orientation of the preset text, and the edit distance threshold.

13. The method according to claim 12, wherein, The text content recognition result includes multiple recognized texts; The method further includes: Based on a preset recognition threshold determined according to the character length of the preset text, the target recognition text for the target text image is determined from a plurality of recognition texts; The step of generating detection results for the material to be detected based on the text content recognition results includes: The detection results for the material to be detected are generated based on the target recognition text.

14. The method according to claim 1, wherein, The orientation recognition of the target text image in the image of the material to be detected includes: In response to the target text image satisfying a second preset condition, orientation recognition is performed on the target text image and the rotated text image, wherein the rotated text image is the image obtained by rotating the target text image by a preset angle; The text recognition process for the target text image and the supplementary text image includes: Text recognition is performed on the target text image, the rotated text image, and the supplementary text image. The supplementary text image includes an image processed by a preset processing method on at least one of the target text image and the rotated text image.

15. The method of claim 14, further comprising: In response to the target text image not meeting the second preset condition, orientation recognition is performed on the target text image to obtain the text orientation recognition result; In response to the text orientation recognition result not meeting the third preset condition, text recognition is performed on the target text image and the supplementary text image for the target text image to obtain the text content recognition result.

16. The method according to claim 1, wherein, The target text image includes multiple images; The method further includes: Among the multiple target text images, the target text image whose average pixel value satisfies the first pixel threshold is selected as the filtered target text image; The orientation recognition of the target text image in the image of the material to be detected includes: Orientation recognition is performed on the filtered target text image.

17. The method according to claim 1, further comprising: The image of the material to be detected is processed into grayscale to obtain a grayscale processed image of the material. In response to the fact that the average pixel value of the material image after grayscale processing is less than the second pixel threshold, a black image abnormality alarm is sent.

18. A material detection device, comprising: The text orientation recognition module is used to identify the orientation of target text images in the image of the material to be detected, and to obtain the text orientation recognition result. A text content recognition module is used to perform text recognition on the target text image and the supplementary text image in response to the text orientation recognition result not meeting a first preset condition, and to obtain a text content recognition result, wherein the supplementary text image includes an image after the target text image has been processed according to a preset processing method; as well as The first generation module is used to generate detection results for the material to be detected based on the text content recognition results.

19. An electronic device comprising a memory and a processor, the memory storing instructions executable by the processor, the instructions, when executed by the processor, causing the processor to perform the method as described in any one of claims 1 to 16.

20. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to perform the method according to any one of claims 1 to 16.

21. A computer program product comprising a computer program that, when executed by a processor, implements the method according to any one of claims 1 to 16.