A method and device for detecting irregularly deformed characters of a product based on deep learning
The deep learning-based method addresses irregular surface detection challenges by optimizing image capture and using advanced character recognition techniques, improving accuracy and efficiency in character detection on complex product surfaces.
Patent Information
- Application Number
- CN202411248664.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-06
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2044-09-06
AI Technical Summary
The existing visual detection methods have problems such as high difficulty in modeling, high interference in imaging noise, and poor adaptability of detection equipment in the detection of irregular deformation characters, making it difficult to achieve efficient and accurate detection of special-shaped curved surface products.
The product irregular deformation character detection method based on deep learning is adopted to collect irregular deformation characters through hand-eye cooperative robot arms, and combine normalized cross-correlation and deep learning character recognition models to achieve character correction and recognition.
It improves the accuracy and efficiency of character detection of irregular surfaces, reduces the dependence of manual detection, and realizes automated high-precision detection of special-shaped surface products, which has strong robustness and adaptability.
Smart Images

Figure CN119049054B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a method and device for detecting deformed characters of products based on deep learning, and particularly to a general detection technology for deformed characters such as laser engraving, printing, and pad printing on different special-shaped curved surface products. Background Art
[0002] In the process of industrial product manufacturing, technologies such as laser engraving, printing, and pad printing are often used to present product information in the form of characters on the outer surface of the product, and then in the automatic production process, the characters on the product surface are identified through visual imaging technology to determine the corresponding information of the product. The characters presented on the surface of products with cylindrical surfaces, spherical surfaces, or irregular curved surfaces have phenomena such as deformity and distortion. At the same time, products with such irregular surfaces are affected by multiple factors such as reflection, material, and curvature during the imaging process. As Figure 9 shown, it makes the detection of product information characters difficult. Currently, it is difficult to use a unified standard to determine character defects such as character breakpoints, ghosting, offset, and different colors, and only manual word-by-word visual inspection can be relied on. The manual inspection method has disadvantages such as low efficiency, high missed detection rate, and strong subjectivity in detection results.
[0003] Traditional vision detection methods based on fixed template matching can only accurately identify relatively regular characters and are difficult to apply to the detection of products with deformed characters on irregular surfaces, which brings great troubles to the quality inspection of product production links. Therefore, there is an urgent need for a method and device for detecting irregular characters of products based on deep learning to solve the above problems.
[0004] The following technical problems are faced in the visual detection of deformed characters on the surface of complex special-shaped products:
[0005] (1) The imaging of characters on irregular curved surfaces is irregular and the modeling is difficult. In the process of pinhole imaging of cylindrical surfaces, spherical surfaces, or irregular bends, the curved surface characters are projected from three-dimensional space to a two-dimensional imaging plane, and due to the loss of depth information, the character images have features such as stretching, compression, or distortion, making it difficult to reversely judge the character quality by establishing a model.
[0006] (2) The product material differences are large and the imaging noise interference is large. The surface materials of different categories of products vary greatly, and the appearances have different differences such as smooth, rough, and different colors. When lighting and imaging, phenomena such as reflection, light absorption, and uneven brightness occur, resulting in strong noise interference such as overexposure, uneven brightness, and poor contrast during the imaging process, affecting the discrimination scale between the background and character information.
[0007] (3) The adaptability of existing surface general detection methods is poor. Most of the existing optical detection devices are developed and designed for a single product. The optical solutions are single, and the detection algorithms are simple. It is difficult to meet the requirements of the existing production line for compatible small varieties, diversification, and rapid model change. There is an urgent need to develop a surface character detection device with strong versatility.
[0008] In addition, Chinese Patent Application with Publication No. CN116168395A discloses a character detection method and device, an electronic device, and a storage medium. The technical solution of this patent is as follows: obtaining a text image, extracting features from the text image, and obtaining the central position and character region information corresponding to the characters in the text image, so as to detect the characters in the text image. However, it does not consider that the surface of the curved surface is prone to reflection, overexposure, etc. under light conditions, resulting in blurring, missing, etc. of the surface characters, affecting the extraction of subsequent character edge, color, or texture features, and thus affecting the detection performance.
[0009] Chinese Patent Application with Publication No. CN117727052A discloses a character defect detection method and device, an electronic device, and a storage medium. The technical solution of this patent is as follows: obtaining the information of the to-be-detected characters corresponding to the to-be-detected area in the image, generating a template character image based on the character information, and using the template character image to perform matching detection on at least one to-be-detected character, which can avoid the situation where the to-be-detected characters cannot be detected. However, new characters will continuously appear during the character detection process, which will frequently generate template character images. The to-be-detected characters need to be accurately matched from a large number of template character images, resulting in a large amount of time consumption during the detection process, thus affecting the production capacity.
[0010] Therefore, it is of great significance and market demand to develop a product irregular deformation character detection technology that combines the optimal perception posture of robot hand-eye coordination with deep learning algorithms to solve the deficiencies of existing detection methods. Summary of the Invention
[0011] The technical problem solved by the present invention is: aiming at the deficiencies of existing visual detection methods in the detection of irregularly deformed characters, to provide a deep learning-based product irregular deformation character detection method and device.
[0012] The present invention is implemented by adopting the following technical solutions:
[0013] The present invention first discloses a deep learning-based product irregular deformation character detection method, specifically including the following steps:
[0014] S1. Use a robotic arm equipped with a camera with hand-eye coordination to perform image sampling on the irregularly deformed characters on the product. Obtain the acquisition pose deviation of the robotic arm by comparing the position features and structural features between the sampled image of the irregularly deformed characters and the standard image, and optimize and adjust the optimal acquisition pose of the robotic arm. The robotic arm performs image acquisition on the characters of the product to be detected through the optimal acquisition pose;
[0015] S2. After denoising and preprocessing the image of the characters of the product to be detected collected in step S1, confirm the region of interest where the characters of the product to be detected are located in the image through normalized cross-correlation, and perform affine transformation on the region of interest and the standard image to correct the character distortion and tilt in the image of the characters of the product to be detected, and obtain the corrected character image of the region of interest;
[0016] S3. Send the corrected character image of the region of interest into a deep learning character recognition model based on image edge segmentation to separate the character text region and the background region in the image, and generate a text box for the character text region;
[0017] S4. Input the text box into a CRNN text recognition model, recognize the character line in the text box as a unit, and return the text sequence of the character line in sequence form after character recognition;
[0018] S5. Match the text sequence recognized in step S4 with the standard character text sequence of the product, determine whether the recognized characters of the product to be detected are unqualified or qualified, and output the product information indicating whether the product to be detected is qualified according to the matching result.
[0019] In a method for detecting irregularly deformed characters of a product based on deep learning according to the present invention, specifically, in step S1, the camera is fixed at the execution end of the robotic arm and moves with the robotic arm, and hand-eye calibration is performed on the camera and the robotic arm through the checkerboard calibration method.
[0020] In a method for detecting irregularly deformed characters of a product based on deep learning according to the present invention, specifically, in step S1, the sampled image of the irregularly deformed characters of the product and the standard image are input into an image matching model for matching, and the position features and structural features of the two images are extracted.
[0021] First, according to the position characteristics, select the feature points F with common characteristics in the trial acquisition image and the standard image, and calculate the Euclidean distance d between the feature points F of the two images. Compare the Euclidean distance d with the set threshold D. If d > D, convert the coordinate of the feature point of the trial acquisition image into the world coordinate where the robotic arm is located through the internal and external parameters of the camera, obtain the first movement vector of the feature points F in the trial acquisition image and the standard image in the world coordinate, calculate the second movement vector of the robotic arm in the world coordinate through the conversion of the first movement vector, solve the position and attitude parameters for the robotic arm to perform the trial acquisition movement, and adjust the position and attitude parameters of the robotic arm to move for the trial acquisition to ensure that d ≤ D;
[0022] Then, for the structural characteristics, obtain the SSIM index between the two images by the following formula for the brightness, contrast, and structural parameters of the trial acquisition image and the standard image:
[0023] SSIM(x, y) = l(x, y)α·c(x, y) β ·s(x, y)γ
[0024] l(x, y) represents the brightness parameter, obtained by comparing the average brightness of the two images. c(x, y) represents the contrast parameter, obtained by comparing the standard deviations of the two images. s(x, y) represents the structural parameter, obtained by comparing the covariance between the two images. (x, y) represents the position coordinates for extracting the brightness, contrast, and structural parameters from the trial acquisition image and the standard image. α, β, γ are weight parameters. The value range of the SSIM index is [-1, 1]. The SSIM index approaching 1 indicates that the two images are exactly the same, and the SSIM index approaching -1 indicates a poor similarity between the two images;
[0025] Then, optimize the SSIM index between the trial acquisition image and the standard image through the following loss function Loss:
[0026] Loss = (SSIM current - SSIM target ) 2 + λ·Penalty(x, T)
[0027] Among them, SSIM current is the current SSIM index between the trial acquisition image and the standard image, SSIMt arget is the target SSIM index between the trial acquisition image and the standard image. SSIM tarhet takes 0.8 - 0.9. λ is the weight for adjusting the penalty term. Penalty(x, T) is the penalty function that restricts the current position x of the robotic arm within the spherical region T. If the current position x of the robotic arm is within the spherical region with T P as the radius, the penalty function is 0, otherwise the penalty function is the distance from the robotic arm position x to the spherical region TP The Euclidean distance is used to adjust the position of the robotic arm x through the gradient descent optimization algorithm, so that the current SSIM index between the test-acquired image and the standard image reaches the set target SSIM index, or the optimization stops when the loss function Loss is minimized. At this time, the robotic arm stops moving, and the current position x of the robotic arm is the optimal acquisition pose for the irregularly deformed characters of the product.
[0028] In a method for detecting irregularly deformed characters of a product based on deep learning according to the present invention, specifically, in the step S2, the non-local means denoising method is used to denoise the acquired character image of the product to be detected.
[0029] In a method for detecting irregularly deformed characters of a product based on deep learning according to the present invention, specifically, in the step S2, the template matching method based on normalized cross-correlation is used to locate the core area of the characters in the character image of the product to be detected, and the template target image T of the character area is intercepted from the standard image in the step S1. m×n , using the template target image T m×n to perform sliding matching in the character image I of the product to be detected. During the sliding process, the normalized cross-correlation similarity is used to measure the similarity between the template target image T M×N and the character image I of the product to be detected. Specifically, it includes: m×n and the character image I of the product to be detected. M×N The similarity is specifically as follows:
[0030] The cross-correlation calculation is performed on the average gray value m×n of the template target image T and the average gray value M×N of the character image I of the product to be detected to obtain the average gray cross-correlation term corr(x, y). The average gray cross-correlation term corr(x, y) is normalized with the standard deviation term σ of the template target image T m×n and the standard deviation term σ T of the character image I of the product to be detected M×N through the following formula to obtain the NCC similarity between the template target image T I and the character image I of the product to be detected: m×n and the character image I of the product to be detected. M×N The NCC similarity is as follows:
[0031]
[0032] In the character image of the product to be detected, a sliding window with a step size of 1 and a sliding size of Q is selected. The NCC similarity of the template target image in each sliding window is calculated. The area with an NCC similarity higher than the set similarity threshold is regarded as the matching area, and the area with the largest NCC similarity is found as the matching region of interest. A bounding box is generated according to the position and size of the region of interest to identify the region of interest.
[0033] In a method for detecting irregularly deformed characters of products based on deep learning according to the present invention, specifically, in step S3, the deep learning character recognition model adopts the Canny edge segmentation algorithm, which specifically includes the following sub-steps:
[0034] Perform Gaussian blur on the corrected character image of the region of interest to reduce noise;
[0035] Use the Sobel operator to calculate the horizontal gradient G x and the vertical gradient G y of the blurred and corrected character image to obtain the gradient magnitude M. Perform non-maximum suppression on the gradient magnitude M, and compare the gradient magnitude of each pixel in the corrected character image with the gradient magnitudes of two adjacent pixels along the gradient direction to retain the local maximum;
[0036] Use the double-threshold edge connection method to threshold the corrected character image after non-maximum suppression, set the high threshold T high , select the strong edge pixels greater than the high threshold T high for edge tracking to determine the connected edge lines;
[0037] Finally, perform thinning processing on the connected edge lines. The internal area of the edge lines is the character text area of the region of interest.
[0038] In a method for detecting irregularly deformed characters of products based on deep learning according to the present invention, specifically, cluster the character text area in the region of interest, identify and merge the local areas belonging to the same text, screen the clustered character text areas, and retain the character text areas where the text characters are continuous and larger than the set region size threshold S a of the character text areas, and use an external rectangle to generate a text box for the screened and retained character text areas.
[0039] In a method for detecting irregularly deformed characters of products based on deep learning according to the present invention, specifically, in step S4, the CRNN text recognition model outputs the text sequence of the character line in the text box through the following sub-steps:
[0040] First step: Convert the character image in the text box into the CNN model format, adjust the character image in the text box to a fixed size and perform normalization;
[0041] Second step: Use the VGG16 model to perform convolution and pooling operations on the input character image in the text box to extract the feature map;
[0042] Third step: Serialize the extracted feature map to obtain the feature sequence;
[0043] Step 4: Input the serialized feature sequence into the RNN model. Combine the feature sequence by recursively integrating the text character context information. Output a vector at each time step, representing the prediction result at that time step.
[0044] Step 5: Perform a decoding operation on the prediction result sequence output by the RNN model. Use the CTC loss function to convert the output sequence into the final text sequence and output it. The finally obtained text sequence is the text sequence of the characters in the text box recognized by the CRNN text recognition model.
[0045] The present invention also discloses a device for detecting irregularly deformed characters of products based on deep learning, including a character detection robotic arm. The execution end of the character detection robotic arm is equipped with a camera component for image acquisition of the irregularly deformed characters of the product. The character detection robotic arm performs image acquisition on the irregularly deformed characters of the product to be detected and transmits the acquired image to the device control host. The device control host determines whether the irregularly deformed characters of the product are qualified through the above detection method of the present invention.
[0046] In a device for detecting irregularly deformed characters of products based on deep learning according to the present invention, it further includes a transport conveyor belt, a product loading module, and a product sorting module.
[0047] The product loading module is located upstream of the transport conveyor belt and includes a loading robotic arm for grasping the product to be detected onto the transport conveyor belt.
[0048] The product sorting module is located downstream of the transport conveyor belt and includes a sorting robotic arm for removing the detected product from the transport conveyor belt, and a material tray for classifying and placing qualified and unqualified products. The sorting robotic arm sorts the detected product into the corresponding material tray according to the detection result of the device control host on whether the product to be detected is qualified.
[0049] The character detection robotic arm is located between the loading robotic arm and the sorting robotic arm, and performs image acquisition on the deformed characters of the product to be detected grasped by the loading robotic arm onto the transport conveyor belt.
[0050] The present invention adopts the above technical solution and has the following beneficial effects:
[0051] 1. The present invention designs an optimal perception attitude adjustment method for special-shaped curved surfaces based on robot hand-eye coordination to adjust the optimal acquisition of regular characters on products, innovatively solving problems such as irregular imaging of characters on irregular curved surfaces, large imaging noise interference, and single detection of irregular products. This method uses a robotic arm to flexibly drive a camera module to perform trial acquisitions of special-shaped curved surface characters from multiple directions and angles, matches the real-time images obtained from the trial acquisitions with standard images, calculates the pose deviation parameters of the robotic arm, and selects the pose of the robotic arm with the highest matching degree between the real-time image and the standard image as the optimal acquisition pose. To improve the search speed of the optimal acquisition pose, a gradient descent algorithm is used to accelerate convergence, solving the problem of long initial pose adjustment time of the robotic arm. The camera module of the present invention uses a double telecentric lens for equal-proportion non-distorted imaging, eliminating problems such as inconsistent magnification ratios of different depths of field caused by pinhole imaging of ordinary lenses, solving the problem of irregular deformation of characters under non-optimal viewing angles on special-shaped curved surfaces, and mapping the imaging within a local curved surface range to planar imaging through the imaging scheme of the double telecentric lens under the optimal viewing angle, realizing the detection feasibility of traditional planar character detection algorithms and enhancing generality and flexibility.
[0052] 2. The present invention designs a method for detecting slightly deformed characters based on deep learning to solve the problem of poor adaptability of slightly deformed characters such as stretching, compression, inflation, and thinning, and has generality. In the detection process, template matching based on normalized cross-correlation (NCC) is introduced. NCC has a certain anti-interference ability against image contrast changes and character deformation. Then, affine transformation and feature extraction are used to effectively separate the character image from the background image. The separated character image is sent into the RCNN text recognition model of deep learning to obtain a character sequence. The character sequence is sent to a character database, and the character database matches the received characters and outputs whether the product character information is qualified according to the matching result. This method solves the problem of detecting characters with uncertain deformation and has strong robustness and adaptability compared with the problems of poor anti-interference ability and low accuracy of traditional fixed template matching.
[0053] 3. The present invention designs a device for automatically and highly accurately detecting irregularly deformed characters of products, solving the problem that current special-shaped curved surface characters mainly rely on a large amount of manual detection, and improving both efficiency and detection accuracy. Two four-degree-of-freedom robotic arms are arranged in both the loading module and the sorting module of this device, completely replacing manual loading and unloading, improving the working efficiency of the entire detection process; to avoid the subjective influence on the detection accuracy of products during the manual detection process, this device combines a vision detection device with a character detection robotic arm, and based on a deep learning character recognition model, realizes automatic and highly accurate detection of special-shaped product characters. In addition, the detection system composed of the character detection robotic arm and the machine vision camera module can adaptively adjust the vision detection angle to quickly respond to product model changes during detection.
[0054] In summary, the method and device for detecting irregularly deformed characters of products based on deep learning provided by the present invention can replace the existing manual detection, eliminate the influence of irregular characters on the product surface easily affected by light and curved surface materials, improve the detection accuracy and speed, and can meet the industrial production of general detection methods for deformed characters such as laser engraving, printing, and pad printing on different products with special-shaped curved surfaces.
[0055] The following further describes the present invention in conjunction with the accompanying drawings and specific embodiments. Description of the Drawings
[0056] Figure 1 It is a schematic diagram of the overall device for detecting irregularly deformed characters of products in Embodiment 1.
[0057] Figure 2 It is a schematic diagram of the product loading module of the device for detecting irregularly deformed characters of products in Embodiment 1.
[0058] Figure 3 It is a schematic diagram of the transport conveyor belt of the device for detecting irregularly deformed characters of products in Embodiment 1.
[0059] Figure 4 It is a schematic diagram of the character detection module of the device for detecting irregularly deformed characters of products in Embodiment 1.
[0060] Figure 5 It is a schematic diagram of the process flow of the method for detecting irregularly deformed characters of products in Embodiment 2.
[0061] Figure 6 It is a schematic diagram of the hand-eye calibration of the character detection robotic arm and camera assembly in Embodiment 2.
[0062] Figure 7 It is a schematic diagram of the process flow for optimizing and adjusting the optimal acquisition pose of the robotic arm in Embodiment 2.
[0063] Figure 8 It is a schematic diagram of the CRNN text recognition model in Embodiment 2.
[0064] Figure 9 It is a schematic diagram of the analysis of the product with irregularly deformed characters detected in Embodiment 2.
[0065] Figure 10 It is a detection effect diagram of the product with irregularly deformed characters in Embodiment 2.
[0066] Reference numerals in the figure:
[0067] 1 - Product loading module, 11 - Loading table, 12 - Material tray to be tested, 13 - Loading robotic arm, 14 - Pneumatic gripper;
[0068] 2 - Conveyor belt, 21 - Belt, 22 - Roller, 23 - Belt support guard plate, 24 - Support block, 25 - Motor, 26 - Frame;
[0069] 3 - Character detection module, 31 - Character detection support, 32 - Character detection robotic arm, 33 - Area array camera, 34 - Double telecentric lens, 35 - Light source, 36 - Vision acquisition fixture;
[0070] 4 - Product sorting module, 41 - Sorting table, 42 - Qualified material tray, 43 - Defective material tray, 44 - Sorting robotic arm, 45 - Sorting gripper. Detailed implementation mode
[0071] Embodiment 1
[0072] See Figure 1 , the figure shows a specific implementation of the product irregular deformation character detection device in the present invention. The device includes a product loading module 1, a conveyor belt 2, a character detection module 3, and a product sorting module 4. Among them, the product loading module 1 is located upstream of the conveyor belt 2 and grabs the product to be detected onto the conveyor belt 2 and conveys it downstream. The product sorting module 4 is located downstream of the conveyor belt 2 and takes the detected product off the conveyor belt 2. The character detection module 3 is located on one side of the conveyor belt 2 between the product loading module 1 and the product sorting module 4, and performs image acquisition on the irregular deformation characters of the product to be detected grabbed by the product loading module 1 onto the conveyor belt 2.
[0073] The following combines Figure 2 , Figure 3 and Figure 4 to further specifically describe this device.
[0074] The product loading module 1 of this embodiment is used to transfer the character product to be measured to the conveyor belt 2, and is composed of a loading table 11, a material tray 12 to be measured, a loading robotic arm 13, and a pneumatic gripper 14. Among them, the main body of the loading table 11 is built with aluminum alloy profiles and can place multiple material trays 12 of products to be measured; the loading robotic arm 13 is installed on the loading table 11 and is fixedly connected by bolts to grab the product to be detected onto the conveyor belt. The loading robotic arm 13 in this embodiment is an industrial robot with four degrees of freedom. The pneumatic gripper 14, as the end effector of the loading robotic arm 13, is connected to the front end of the loading robotic arm 13 by bolts and is used to pick up the product to be detected in the material tray 12 and place it on the conveyor belt 2.
[0075] The transport conveyor belt 2 is used to receive the products to be detected with characters transferred by the product loading module 1 and transport the products to be detected to the detection area of the character detection module 3. The transport conveyor belt 2 includes a belt 21, roller shafts 22, a belt support guard plate 23, support blocks 24, a motor 25, and a frame 26. Among them, the frame 26 is a box structure composed of aluminum and aluminum plates. The inside of the frame 26 is used to store the equipment driver and the equipment control host. A plurality of support blocks 24 are arranged above the frame 26 along the conveyor belt direction. The front and rear rows of the support blocks 24 are connected to the frame 26 by screws. The two rows of support blocks are connected to two belt support guard plates 23. The roller shafts 22 are installed on the two belt support guard plates 23, one in the front and one in the back. The two roller shafts 22 are wound around by the belt 21. The shaft of one of the roller shafts 22 passes through the belt support guard plate 23 and is connected to the motor 25. The motor 25 is fixedly installed on the side of the belt support guard plate 23 by screws.
[0076] The character detection module 3 is used to detect whether there are character defects on the surface of the products to be detected, and includes a character detection support 31, a character detection robotic arm 32, an area array camera 33, a double telecentric lens 34, a light source 35, and a vision acquisition fixture 36. Among them, in this embodiment, the character detection robotic arm 32 is a six-degree-of-freedom industrial robotic arm, which is connected to the character detection support 31 by screws. The execution end of the character detection robotic arm 32 is equipped with a camera assembly for image acquisition of the irregularly deformed characters on the products to be detected. The camera assembly is a combination of an area array camera 33, a double telecentric lens 34, and a light source 35. After the area array camera 33 and the double telecentric lens 34 are connected, they are installed on the vision acquisition fixture 36 by screws. The light source 35 is a circular hollow light source. After passing through the double telecentric lens 34, it is fixedly installed on the vision acquisition fixture 36. The camera assembly for vision acquisition is combined into the end effector of the robotic arm 32 through the vision acquisition fixture 36 and combined with the robotic arm to form a hand-eye detection system.
[0077] The product sorting module 4 is used to sort out unqualified character products and transfer the qualified character products to the corresponding material racks, and includes a sorting table 41, a qualified material tray 42, a defective material tray 43, a sorting robotic arm 44, and a sorting gripper 45. Among them, the main body of the sorting table 41 is built with aluminum alloy profiles. The qualified material tray 42 and the defective material tray 43 are respectively placed on the sorting table 41. The sorting robotic arm 44 is installed on the sorting table 41 and fixed by bolt connection. The sorting robotic arm 44 in this embodiment uses an industrial robot with four degrees of freedom. The sorting gripper 45, as the end effector of the sorting robotic arm 44, is connected to the front end of the sorting robotic arm by bolts and is used to pick up the products that have been detected on the transport conveyor belt and place them in the designated material trays.
[0078] When using the above-mentioned irregular deformation character detection device for products, first place the product to be detected in the to-be-detected material tray 12 of the product loading module 1. The to-be-detected material tray 12 is placed on the loading table 11. The loading robotic arm 13 drives the end pneumatic gripper 14 to the position where the to-be-detected product is located. The pneumatic gripper 14 is ventilated to grip the to-be-detected product and transfer the product above the belt 21 of the transport conveyor belt 2. Place the to-be-detected product, and successively and continuously transfer the to-be-detected product from the material tray 12 to the belt 21 through the loading robotic arm 13 and the pneumatic gripper 14. The belt 21 successively transports the to-be-detected product to the detection area where the character detection module 3 is located. The camera assembly in the character detection module 3 moves and adjusts through the character detection robotic arm 32 to collect images of the character products on the material tray. The collected images are transmitted to the device control host, and the device control host judges whether the irregular deformation characters of the product are qualified through the detection method of the present invention. The sorting robotic arm 44 of the product sorting module 4 sorts the detected products into the corresponding material trays according to the detection result of whether the to-be-detected product is qualified by the device control host. If it receives the sorting signal of the qualified product, the sorting robotic arm 44 transfers it to the qualified material tray 42. If it receives the sorting signal of the unqualified product, the sorting robotic arm 44 transfers it to the defective material tray 43.
[0079] Embodiment 2
[0080] This embodiment details the method for detecting irregular deformation characters of products by the device in Embodiment 1. It should be noted that this embodiment is only a specific implementation scheme of the method for detecting irregular deformation characters of products based on deep learning provided by the present invention applied to the detection device in Embodiment 1, and does not limit that Embodiment 1 is the only detection device applicable to the detection method of the present invention.
[0081] See Figure 5 , the method for detecting irregular deformation characters of products in this embodiment includes image acquisition of deformed characters, image processing of deformed characters, image processing of deformed characters, and post-processing of deformed character detection, specifically including the following steps:
[0082] S1. Use a character detection robotic arm with hand-eye coordination to perform trial image acquisition on the irregular deformation characters on the product. Obtain the acquisition pose deviation of the robotic arm by comparing the position features and structural features between the trial acquisition image of the irregular deformation characters and the standard image, and optimize and adjust the optimal acquisition pose of the robotic arm. The robotic arm performs image acquisition on the characters of the to-be-detected product through the optimal acquisition pose.
[0083] The camera component for collecting images is fixed at the execution end of the character detection robotic arm and moves with the robotic arm. The hand-eye calibration of the camera component and the character detection robotic arm is performed by the checkerboard calibration method. The transformation matrix between the camera coordinate system and the robotic arm coordinate system is calibrated, and the internal and external parameter matrices of the camera are calculated to establish the coordinate transformation relationship between the camera and the character detection robotic arm. The corresponding point of a certain point in the image in the real world is obtained through the coordinate transformation relationship. The specific process of the checkerboard calibration method is as follows.
[0084] As Figure 6 shown, a 10*7 checkerboard calibration board is selected. The side length of each small square on the checkerboard is 21.5 mm. The camera remains stationary, and the checkerboard is rotated to collect checkerboard images with different poses and tilt angles. Through the mapping from the two-dimensional calibration board to the two-dimensional imaging plane, the internal parameter matrix of the camera is solved using the calibration board information at different angles and positions. In this embodiment, the eye-in-hand calibration method is selected, that is, the camera is fixed at the end of the character detection robotic arm and moves with the robotic arm, and the transformation matrix between the camera coordinate system and the robotic arm coordinate system is calibrated.
[0085] The camera calibration process is mainly divided into two parts: First, according to the constraints of geometric optics, the optical and physical parameters of the imaging model are solved; Second, through establishing constraint relationships, the solved parameters are iteratively optimized. Assume that the checkerboard is located on the plane where Z W in the world coordinate system O W =0. The conversion relationship between the pixel coordinate system and the world coordinate system is known as:
[0086]
[0087] In the formula, t WC is the transformation matrix between the world coordinate system (base coordinate system) and the camera coordinate system, that is, the external parameter matrix of the camera. t CI is the mapping matrix between the two-dimensional image coordinate system and the three-dimensional camera coordinate system, that is, the internal parameter matrix. R 3×3 and T 1×3 are the rotation matrices and translation matrices around the X W , Y W and Z W axes between the camera coordinate system and the robotic arm coordinate system respectively. u and v represent a certain point in the pixel coordinate system, v0 and u0 are the coordinates of the center origin of the pixel coordinate system, f x , f y are set to represent the pixel values per millimeter in the x and y axis directions in the pixel coordinate system, and Z c is the depth information parameter of a certain point in the world coordinate system.
[0088] The result of the camera internal parameter calibration is the three-dimensional space (X C , Y C , Z C) The conversion relationship to the two-dimensional imaging plane (u, v). Each marked point on the checkerboard calibration board is denoted as [X i , Y i , 0] T in the world coordinate system, and the corresponding point in the pixel coordinate system is denoted as [u i , v i , 1] T . Substituting into the above conversion relationship, we can obtain:
[0089]
[0090] The A[r1 r2 t1] in the formula is denoted as the homography matrix, and H = [H1, H2, H3] is the homography matrix used to describe the corresponding relationship between two image planes. Then, for a point M(X i , Y i ) on the checkerboard calibration board, the relationship between it and its pixel coordinates (u i , v i ) is as follows:
[0091]
[0092] In the formula, Z i represents the depth information of point M. Eliminating the depth information Z i in the above formula, we can obtain:
[0093]
[0094] According to the relationship that r1 and r2 in the rotation matrix are unit orthogonal, and the homography matrix H = [H1, H2, H3] = A[r1 r2 t1], we know that:
[0095] r1 = A -1 H1, r2 = A -1 H2,
[0096]
[0097] Among them, H1 = Ar1 represents the mapping relationship between the image plane and the checkerboard plane in the first direction, H2 = Ar2 represents the mapping relationship between the image plane and the checkerboard plane in the second direction, H2 = At1 represents the translation relationship between the image plane and the checkerboard plane, and B = A -T A -1 ).
[0098] Let B = A -T A -1 , where A 3×3 is the internal parameter matrix of the camera. We can obtain:
[0099]
[0100] The internal parameters of the camera are as follows:
[0101]
[0102] α and β are the scale factors of the image axis (focal length), and γ is the angle between the image axes (sensor axis and optical axis). According to the homography matrix H=A[r1 r2 t1], we get [R1, R2, T] corresponding to each chessboard calibration plate image. Since R1, R2 and R3 are orthogonal to each other, we can get the rotation vector R3 and translation vector T of the camera extrinsic parameter:
[0103] R3=R1×R2,T=A -1 H3.
[0104] The “eyes on hands” calibration algorithm is as follows Figure 6 As shown, the camera is fixed at the end of the character detection robot arm. At this time, the rotation and translation matrix between the camera and the end of the character detection robot arm is The transformation matrix between the chessboard calibration plate and the robot base coordinate system remains unchanged. unchanged, according to these two unchanged transformation matrices, the following relationship exists:
[0105]
[0106] In the formula, is the position transformation relationship between the robot base coordinate system and the end of the robot. It is the conversion relationship between the camera coordinate system and the calibration plate coordinate system. Let the character detection robot be in two different positions, and the camera obtains the chessboard calibration plate image. According to the transformation relationship of the coordinate system, we can get:
[0107]
[0108] in It is the transformation relationship between two different robot end positions and the robot base coordinate system, which can be obtained by 5-point calibration of the robot. is the conversion relationship between two different pose cameras and the chessboard calibration plate, which can be obtained by the chessboard calibration method, that is, the camera external parameter. Combining the above two equations, we can get:
[0109]
[0110] Then we get:
[0111]
[0112] Then the above formula can be simplified to CX=XD, where is an unknown. Convert CX = XD into a homogeneous linear equation and solve for X using the least squares method. In the matrix, C, X, and D are all 4×4 homogeneous matrices composed of the rotation vector R3 and the translation vector T. The matrix form of the equation system CX = XD is as follows:
[0113]
[0114] Then CX = XD can be split into the following equations:
[0115]
[0116] Among them, R C represents the 3×3 rotation matrix of the robotic arm between two different poses, R D represents the 3×3 rotation matrix of the camera between different poses, R X represents the 3×3 rotation matrix of the end of the robotic arm between different poses, T C represents the 3×1 translation matrix of the robotic arm between two different poses, T D represents the 3×1 translation matrix of the camera between two different poses, T X represents the 3×1 translation matrix of the end of the robotic arm between two different poses. First, solve for R X , and then solve for T X , then the hand-eye calibration of the robotic arm vision system can be completed.
[0117] In step S1, an optimal perception pose adjustment method for special-shaped curved surfaces based on robot hand-eye coordination is used to perform a trial acquisition of the deformed characters of irregular products. This method performs an image trial acquisition of the deformed characters of the irregular products, matches the trial acquisition image of the deformed characters of the products and the standard image input to the image matching model, extracts the position features and structural features of the two images, calculates the position offset and structural difference vector between the two images, further calculates the pose deviation parameters of the character detection robotic arm, and determines the optimal acquisition pose of the robotic arm through the fast gradient descent optimization algorithm.
[0118] The camera component carried by the character detection robotic arm in this embodiment includes a area array CCD camera, a double telecentric lens, and a square shadowless light source, which is used to perform trial acquisition on the characters on the surface of irregular products. The acquired trial acquisition image is sent into an image matching model for matching. There are two inputs in the image matching model. One is the standard image of the deformed character, which is acquired by the same optical device. During the acquisition of the standard image, the light is parallel to the character mapping, ensuring that the character is located in the center of the image and the light interference is small. The other input is the trial acquisition image to be matched acquired by the character detection robotic arm. The trial acquisition image first acquired by the character detection robotic arm may have phenomena such as character offset, incompleteness, and character distortion. The image matching model is used to match the image to be matched with the standard image, calculate the pose deviation parameters during the trial acquisition of the character detection robotic arm, and select the robotic arm pose with the highest matching degree between the trial acquisition real-time image and the standard image as the optimal acquisition pose.
[0119] Specifically, as Figure 7 shown, first process the position features between the trial acquisition image and the standard image, and perform a difference analysis on the Euclidean distance in the position features. Select the feature point F with common features in the trial acquisition image and the standard image and calculate the Euclidean distance d between the feature points F of the two images. Compare the Euclidean distance d with the set threshold D. If d > D, convert the coordinate of the feature point of the trial acquisition image into the world coordinate where the robotic arm is located through the internal and external parameters of the camera, obtain the first movement vector of the feature points F in the trial acquisition image and the standard image in the world coordinate, calculate the second movement vector of the robotic arm in the world coordinate through the conversion of the first movement vector, and solve the position and attitude parameters of the robotic arm for trial acquisition movement, and adjust the position and attitude parameters of the robotic arm to move for trial acquisition to ensure that d ≤ D.
[0120] The specific calculation process is as follows: The coordinate of the feature point F in the standard image is (x1, y1), and the coordinate in the trial acquisition image to be matched is (x2, y2). Project (x1, y1) onto the trial acquisition image to be matched. The calculation formula for the Euclidean distance d between the two points (x1, y1) and (x2, y2) is as follows:
[0121]
[0122] Then subtract the coordinate of the feature point of the trial acquisition image to be matched from the coordinate of the feature point of the standard image to obtain the difference vector:
[0123] (Δx, Δy) = (x1 - x2, y1 - y2).
[0124] Given two points (x1, y1) and (x2, y2), represent them as homogeneous coordinates:
[0125]
[0126] Using the internal parameters A of the camera, convert the image coordinates to camera coordinates:
[0127]
[0128] Use the external parameter matrix R of the camera t , where R is the rotation vector and t is the translation vector, to convert the camera coordinates to world coordinates:
[0129]
[0130] where, is the coordinate of the point in the world coordinate system, is the coordinate of the point in the camera coordinate system. Try to collect the first movement vector of two points (x1, y1) and (x2, y2) in the test image and the standard image in the real physical world as
[0131] Assume that the current position and orientation of the character detection robotic arm are represented by a six-dimensional vector: T = [X Y Z α β γ] T , where X, Y, and Z represent the position of the end of the character detection robotic arm in the base coordinate system, and α, β, and γ represent the orientation of the end of the character detection robotic arm.
[0132] Assume that the rotation matrix from the world coordinate system to the robotic arm coordinate system is R wb and the translation vector t wb :
[0133]
[0134] Calculate the second movement vector Δp in the robotic arm coordinate system b = R wb ·Δp w , since it is only a translation operation, calculate the new position of the character detection robotic arm through the movement vector Δp b , and keep the orientation of the robotic arm unchanged. The position and orientation of the moved robotic arm can be obtained as:
[0135]
[0136] That is, the new position and orientation of the character detection robotic arm are: T new = [X' Y' Z' α' β' γ'] T = [X + Δx b Y + Δy b Z + Δz b α β γ] T , where Δx b , Δy b and Δz b are obtained through Δpb Calculated movement amount.
[0137] After obtaining the new position and posture of the robotic arm, drive the joints of the robotic arm to rotate or move until d is less than the set threshold D.
[0138] Refer to again Figure 7 , for the structural features between the trial acquisition image and the standard image, calculate the average values μ1 and μ2, standard deviations σ1 and σ2, and covariance σ of the standard image and the trial acquisition image to be matched 12 . Measure the brightness, contrast, and structure of the two images to obtain the SSIM index.
[0139] Specifically, the brightness measurement of the trial acquisition image and the standard image is performed by comparing the average brightness of the two images, and the formula is as follows:
[0140]
[0141] In the formula, C1 is a constant used to avoid the denominator being 0, and usually C1 = (K1L) 2 , where L is the dynamic range of pixel values. For 8-bit images, L = 255, and K1 is a very small constant, usually taken as 0.01.
[0142] The contrast measurement of the trial acquisition image and the standard image is performed by comparing the standard deviations of the images, and the formula is as follows:
[0143]
[0144] In the formula, C2 is a constant used to avoid the denominator being 0, and usually C2 = (K2L) 2 , where L is the dynamic range of pixel values. For 8-bit images, K = 255, and K2 is a constant, taken as 0.03.
[0145] The structure measurement of the trial acquisition image and the standard image is performed by comparing the covariance between the images, and the formula is as follows:
[0146]
[0147] In the formula, C3 is a constant, usually taken as C3 = C2 / 2.
[0148] Obtain the SSIM index between the two images through the brightness, contrast, and structure parameters of the trial acquisition image and the standard image by the following formula:
[0149] SSIM(x, y) = l(x, y) α ·c(x, y) β ·s(x, y) γ, l(x, y) represents the luminance parameter, which is obtained by comparing the average luminance of two images. c(x, y) represents the contrast parameter, which is obtained by comparing the standard deviations of two images. s(x, y) represents the structure parameter, which is obtained by comparing the covariance between two images. (x, y) represents the position coordinates for extracting the luminance, contrast, and structure parameters from the test acquisition image and the standard image. α, β, and γ are weight parameters, which are adjusted according to the largest difference between images and are generally taken as 1. The value range of the SSIM index is [-1, 1]. The SSIM index approaching 1 indicates that the two images are exactly the same, and the SSIM index approaching -1 indicates a poor similarity between the two images.
[0150] In this embodiment, in order to make the test acquisition image to be matched more similar to the standard image, a loss function Loss is designed. The optimization objective is to minimize the difference between the current SSIM index and the target SSIM index between the test acquisition image and the standard image. The constraint objective is that after the character targets in the two images roughly overlap, the acquisition module driven by the robotic arm moves within a spherical space with a radius of T. The specific expression of the loss function Loss is as follows:
[0151] Loss = (SSIM current - SSIM target ) 2 + λ·Penalty(x, T),
[0152] where SSIM current is the current SSIM index between the test acquisition image and the standard image, SSIM target is the target SSIM index between the test acquisition image and the standard image, SSIM target takes a value of 0.8 to 0.9, λ is the weight for adjusting the penalty term, Penalty(x, T) is the penalty function that constrains the current position x of the robotic arm within the spherical region T. If the current position x of the robotic arm is within the spherical region with a radius of T P , the penalty function is 0, otherwise the penalty function is the Euclidean distance from the robotic arm position x to the spherical region T P . T P is equal to the Euclidean distance threshold D of the feature points set when the test acquisition image and the standard image are input into the image matching model for matching. The position of the robotic arm x is adjusted through the gradient descent optimization algorithm so that the current SSIM index between the test acquisition image and the standard image reaches the set target SSIM index, or the optimization stops when the loss function Loss is minimized. At this time, the robotic arm stops moving, and the current position x of the robotic arm is the optimal acquisition pose of the irregularly deformed character of the product.
[0153] Thus, the optimal acquisition pose adjustment of the character detection robotic arm in step S1 is completed. The character detection robotic arm acquires images of the irregular characters of the product to be detected on the transportation conveyor belt according to this optimal acquisition pose, and uploads the acquired character images of the product to be detected to the device control host.
[0154] S2. After denoising and preprocessing the character images of the product to be detected collected in step S1, confirm the region of interest where the characters of the product to be detected are located in the image through normalized cross-correlation, and perform affine transformation on the region of interest and the standard image to correct the character distortion and tilt in the character image of the product to be detected, so as to obtain the corrected character image of the region of interest.
[0155] The device control host receives the character images of the product to be detected collected in step S1, and uses the non-local means denoising method to denoise the collected character images of the product to be detected, weakening the influence of exposure or reflection caused by different lighting conditions and different curved surface materials on the character images of the product to be detected.
[0156] Specifically, the basic idea of non-local means denoising is to use the similarity between each pixel point in the image for filtering operations. Non-local means denoising performs weighted averaging by calculating the similarity of pixel points throughout the image, so as to better remove noise and preserve the texture detail information of the image. The specific reasoning process of this non-local means denoising is as follows:
[0157] The size of the character image I of the product to be detected is M×N, and the pixel value of the i-th row and j-th column is defined as I ij . For each pixel I in the character image of the product to be detected ij , define a window of size (2W + 1)×(2W + 1), where W is the window radius. For the current pixel I ij , calculate the similarity of the pixel values within each surrounding window. Let the pixel value of the k-th window in the image I be I k , then the similarity between the current window and the k-th window is:
[0158]
[0159] In the formula, the meanings of p and q are the identifiers for counting in the summation formula. Calculate the weight of each pixel according to the similarity. For the current pixel I ij , its weight W ij,k is:
[0160]
[0161] Among them, h is a parameter that controls the similarity weight. Use the weight to perform weighted averaging on the current pixel to obtain the denoised pixel value:
[0162]
[0163] Repeat the above steps for all pixels in the image, and the denoised image is obtained.
[0164] Then, use the template matching method based on normalized cross - correlation to locate the core region of the characters in the character image of the product to be detected, that is, the region where the irregular characters on the product surface are located. Given the character image I of the product to be detected M×N (i.e., the above - denoised image ), and a template target image T m×n , the template target image T m×n is the character image to be matched, which is obtained by intercepting the character region from the standard image in step S1. The purpose is to find the region in the character image I of the product to be detected M×N that is most similar to the template target image T m×n . Slide and match the template target image T m×n in the character image I of the product to be detected M×N . During the sliding process, use the normalized cross - correlation similarity to measure the similarity between the template target image T m×n and the character image I of the product to be detected M×N . Specifically, it includes:
[0165] First, calculate the average gray - level values of the template image T m×n and the character image I of the product to be detected M×N , which are denoted as and respectively. m×n and I M×N Perform cross - correlation calculation to obtain the average gray - level cross - correlation term corr(x, y). The calculation formula of the cross - correlation term is as follows:
[0166]
[0167] where (u, v) are the coordinates in the template target image T m×n , and (x, y) are the coordinates in the character image I of the product to be detected M×N .
[0168] Then, calculate the standard - deviation term σ m×n of the template target image T T and the standard - deviation term σ M×N of the character image I of the product to be detected I . The calculation formula of the standard - deviation term is as follows:
[0169]
[0170] Divide the average gray - level cross - correlation term corr(x, y) by the product of the standard - deviation term σ m×nThe standard deviation term σ T and the character image I of the product to be detected M×N The standard deviation term σ I Perform cross - correlation term normalization through the following formula to obtain the template target image T m×n and the character image I of the product to be detected M×N The NCC similarity of:
[0171]
[0172] In the character image of the product to be detected, select a sliding window with a step size of 1 and a sliding size of Q, calculate the NCC similarity of the template target image within each sliding window, regard the area where the NCC similarity is higher than the set similarity threshold as the matching area, find the area with the largest NCC similarity, which is the region of interest (ROI) that is matched, and generate a bounding box according to the position and size of the region of interest to identify the region of interest.
[0173] Construct an affine transformation matrix using the standard image used in step S1 and the region of interest (ROI) obtained by matching in this step. Correct the slight distortion and tilt of the character image through the constructed affine transformation matrix to obtain the corrected character image of the region of interest. Considering that there may be transformations such as translation, rotation, and scaling in the local area, resulting in the geometric shape and pose between the matched region of interest and the template image may not be exactly the same. Through affine transformation, further adjust the matched region of interest to a shape more similar to the template image, so that the subsequent curved - surface character recognition task can be accurately performed.
[0174] S3. Send the corrected character image of the region of interest into a deep - learning character recognition model based on image - edge segmentation. Through image segmentation, effectively separate the character text region and the background region in the image, and generate a bounding box of the character text region, that is, a text box.
[0175] Considering that curved - surface characters usually have obvious edge features, the deep - learning character recognition model in this embodiment adopts the Canny edge - segmentation algorithm, which specifically includes the following sub - steps:
[0176] Perform Gaussian blur on the corrected character image of the region of interest to reduce noise. Gaussian blur can be achieved by convolving the corrected character image of the region of interest with a Gaussian kernel G:
[0177] I blur (x, y)=I(x, y)*G(x, y).
[0178] Calculate the horizontal - direction gradient G x and the vertical - direction gradient G y ,
[0179] G x = I blur (x, y) * S x 、G y = I blur (x, y) * S y ,
[0180] where S x and S y are the Sobel operators in the horizontal and vertical directions respectively.
[0181] The gradient magnitude M and direction θ are calculated by the following formula:
[0182]
[0183] Non-maximum suppression is performed on the gradient magnitude M. The gradient magnitude of each pixel in the corrected character image is compared with the gradient magnitudes of two adjacent pixels along the gradient direction, and the local maximum value is retained. The process of suppressing non-maximum values can be expressed as:
[0184]
[0185] where Δx and Δy represent the small increments in the x and y directions respectively. Using the double-threshold edge connection method, the corrected character image after non-maximum suppression is thresholded. Set the high threshold T high and the low threshold T low . Pixel points greater than the high threshold are marked as strong edges, those less than the low threshold are marked as weak edges, and pixels in between are marked as intermediate edges. Select the strong edge pixels greater than the high threshold T high for edge tracking. The direction θ obtained by calculation guides the extension direction of the edge to determine the connected edge lines. Finally, the connected edge lines are thinned. This step is to further refine the edge shape to make it more accurate. The internal area of the edge line is the character text area of the region of interest. By performing the above steps, the Canny edge detection algorithm can detect clear and accurate edges in the image of interest, which helps to separate the characters from the background.
[0186] Cluster the character text areas in the region of interest, identify and merge local areas belonging to the same text, filter the clustered character text areas, remove too small or irregular areas, and retain larger and continuous character text areas. The size of the text character area is judged by the set area size threshold S a . Areas greater than or equal to S a are larger areas, and those less than S aThose with too small areas. Generate a bounding box for the filtered text area as a text box to identify the position and scope of this text. The bounding box is represented by the circumscribed rectangle of the text area.
[0187] S4. Input the text box into the CRNN text recognition model trained offline. The CRNN model in this embodiment is specifically used to recognize variable-length sequence objects in images. The character lines within the text box are recognized as a unit, and after character recognition, the text sequence of the character lines is returned in sequence form. Then, the recognized character information is uploaded to the device control host and saved in the database for subsequent product qualification judgment.
[0188] The above steps effectively separate the character text area in the irregular character image of the product to be detected from the background and generate a text box for the text area, providing accurate input for the CRNN text recognition model in this step. As Figure 8 shown in the figure, the CRNN text recognition model in this embodiment outputs the text sequence of the character lines within the text box through the following sub-steps:
[0189] First step: Preprocess the character image within the text box as the input image, convert it to a format suitable for the CNN model, adjust the character image within the text box to a fixed size and perform normalization.
[0190] Second step: Use the VGG16 model to perform convolution and pooling operations on the input character image within the text box to extract the feature map. Specifically, let the input character image within the text box be X and the VGG16 model parameters be W. Then the forward propagation process of the CNN can be expressed as:
[0191] F = M1(X; W),
[0192] where F is the feature map extracted by the CNN and M1 is the VGG16 model.
[0193] Third step: Serialize the extracted feature map to obtain the serialized feature sequence S:
[0194] S = Flatten(F).
[0195] Fourth step: Process with the recurrent neural network. Input the serialized feature sequence S into the RNN model, combine the feature sequence by recursively integrating the text character context information, and output a vector at each time step, representing the prediction result at that time step. The final output form of the RNN model is the prediction result sequence composed of the output vectors at all time steps. This sequence is composed of the vectors output by the RNN model at each time step, and each vector represents the prediction result at that time step. This sequence will be used for the next decoding operation.
[0196] Taking the BiLSTM model as an example, assume the parameters of the BiLSTM model are θ and the output of the BiLSTM model is Y. Then the forward propagation process of BiLSTM can be expressed as:
[0197] BiLSTMY = M2(S; θ),
[0198] where M2 is the BiLSTM model.
[0199] Step 5: Decode the predicted result sequence output by the RNN model, and use the CTC loss function to convert the output sequence into the final text sequence and output it. The CTC loss function is usually used for sequence labeling tasks. Assume the input sequence x = (x1, x2,..., x w ) of length w is output in the above steps, where each x w represents a feature vector, and a label sequence y = (y1, y2,..., y U ) of length U represents the target sequence, where each y u represents a label. The calculation of the CTC loss is divided into two steps: the forward algorithm and the backward algorithm.
[0200] (1) Forward algorithm. First, define a (w×U) matrix y', where each element y' t,u represents the probability of generating label u at time step t. For each time step t and each label u, define:
[0201] α t (u) = p(y[1:u]|x[1:t]),
[0202] where α t (u) represents the probability of generating the label sequence y[1:u] at time step t.
[0203] Using the method of dynamic programming, the recurrence formula for calculating α t (u) can be calculated:
[0204]
[0205] blank represents a special label for the blank symbol. Finally, the CTC loss is the negative logarithm of the probability of generating the entire label sequence at the end of time step T:
[0206]
[0207] (2) Backward algorithm. To calculate the derivative, the backward algorithm is needed to calculate the gradient. The backward algorithm is the inverse process of the forward algorithm, and it calculates β t (u) through dynamic programming, where β t(u) represents the probability of generating the label sequence y[u+1:U] at the end of time step t. The recurrence formula of the backward algorithm is as follows:
[0208]
[0209] The finally obtained text sequence is the recognition result of the CRNN text recognition model for recognizing the character lines within the text box, representing the character text sequence detected in the image.
[0210] S5. Match the text sequence recognized in step S4 with the product standard character text sequence in the device control host database, determine whether the character of the product to be detected belongs to unqualified (NG) or qualified (OK), and output the product information on whether the product to be detected is qualified according to the matching result. The information includes whether it is a qualified product (marked as a defective product if unqualified) and the product position information.
[0211] This embodiment performs detection on Figure 9 the product irregular deformation character samples shown in, and the detected result character sequence is as Figure 10 shown. The defective curved surface characters cannot be successfully detected, while other normal curved surface characters are successfully detected. The sorting robotic arm downstream of the character detection robotic arm sorts the products that have completed the detection according to the detection product information. The unqualified character products are sorted by the sorting robotic arm using the position signal in the product information to the defective product material tray, and the qualified character products are sorted by the sorting robotic arm to the qualified product material tray; finally, the material tray containing the qualified character products after sorting is transported to the blanking rack.
[0212] In this article, the orientation or positional relationship indicated by the terms "upper", "lower", "front", "rear", "left", "right", "top", "bottom", "inner", "outer", "vertical", "horizontal", etc. is the orientation or positional relationship based on the orientation or positional relationship shown in the drawings, only for the sake of clear expression of the technical solution and convenient description, so it cannot be understood as a limitation to the present invention.
[0213] In this article, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion. In addition to including the listed elements, it may also include other elements not specifically listed.
[0214] The above is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of changes or substitutions, which should all be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.
Claims
1. A method for detecting irregularly deformed characters of products based on deep learning, characterized in that Including the following steps: S1. Use a robotic arm equipped with a camera with hand-eye coordination to perform image trial acquisition of irregularly deformed characters on the product. The camera is fixed at the end effector of the robotic arm and moves with the robotic arm. Perform hand-eye calibration on the camera and the robotic arm through the checkerboard calibration method. Obtain the acquisition pose deviation of the robotic arm by comparing the position features and structural features between the trial acquisition image of the irregularly deformed character and the standard image, optimize and adjust the optimal acquisition pose of the robotic arm, and optimize the SSIM index between the trial acquisition image and the standard image through the following loss function Loss : , in, is the current SSIM index between the trial image and the standard image, is the target SSIM index between the test image and the standard image, Take 0.8~0.9, is the weight used to adjust the penalty term, Is to constrain the current position of the robot arm In the spherical area T The penalty function within, if the current position of the robot In T P If the area is within the spherical region with a radius of , the penalty function is 0, otherwise the penalty function is the position of the robot arm. To the spherical area T P The Euclidean distance is used to adjust the robot arm through the gradient descent optimization algorithm. The position of the test image makes the current SSIM index between the test image and the standard image reach the set target SSIM index, or the loss function Loss Stop optimization when minimized. At this time, the robot stops moving and the current position of the robot The optimal acquisition posture of the irregularly deformed characters of the product, the robot arm performs image acquisition of the characters of the product to be inspected through the optimal acquisition posture; S2. After denoising and preprocessing the character image of the product to be detected collected in step S1, confirm the region of interest where the character of the product to be detected is located in the image through normalized cross-correlation, perform an affine transformation on the region of interest and the standard image to correct the character distortion and tilt in the character image of the product to be detected, and obtain the corrected character image of the region of interest; S3. Send the corrected character image of the region of interest into a deep learning character recognition model based on image edge segmentation to separate the character text region and the background region in the image, and generate a text box for the character text region; S4. Input the text box into the CRNN text recognition model, recognize the character line in the text box as a unit, and return the text sequence of the character line in sequence form after character recognition; S5. Match the text sequence recognized in step S4 with the product standard character text sequence, judge whether the recognized character of the product to be detected is unqualified or qualified, and output the product information indicating whether the product to be detected is qualified according to the matching result.
2. The method for detecting irregularly deformed characters of a product based on deep learning according to claim 1, characterized in that: In step S1, input the trial acquisition image of the irregularly deformed character of the product and the standard image into an image matching model for matching, and extract the position features and structural features of the two images. First, according to the position features, select the feature point F with common features in the trial acquisition image and the standard image, calculate the Euclidean distance d between the feature points F of the two images, compare the Euclidean distance d with the set threshold D. If d > D, convert the feature point coordinates of the trial acquisition image into the world coordinates where the robotic arm is located through the internal and external parameters of the camera, obtain the first movement vector of the feature points F in the trial acquisition image and the standard image in the world coordinates, calculate the second movement vector of the robotic arm in the world coordinates through the conversion of the first movement vector, solve the position and attitude parameters for the robotic arm to perform trial acquisition movement, and adjust the position and attitude parameters of the robotic arm to move for trial acquisition to ensure d ≤ D. Then, for the structural features, obtain the SSIM index between the two images through the following formula for the brightness, contrast, and structural parameters of the trial acquisition image and the standard image: , l(x,y) represents the brightness parameter, which is obtained by comparing the average brightness of two images. c(x,y) is the contrast parameter, which is obtained by comparing the standard deviations of two images. s(x,y) represents the structure parameter, which is obtained by comparing the covariance between two images. ( x,y ) represents the position coordinates for extracting the brightness, contrast, and structure parameters from the test acquisition image and the standard image. , , are the weight parameters. The value range of the SSIM index is [-1, 1]. When the SSIM index approaches 1, it indicates that the two images are exactly the same. When the SSIM index approaches -1, it indicates a poor similarity between the two images.
3. A method for detecting irregularly deformed characters of a product based on deep learning according to claim 1, characterized in that: In step S2, use the non-local means denoising method to denoise the collected character image of the product to be detected.
4. A method for detecting irregularly deformed characters of a product based on deep learning according to claim 1, characterized in that: In the step S2, a template matching method based on normalized cross-correlation is used to locate the core area of the characters in the character image of the product to be detected, and a template target image of the character area is intercepted from the standard image in the step S1. Using the template target image to perform sliding matching in the character image of the product to be detected , and the normalized cross-correlation similarity is used to measure the similarity between the template target image and the character image of the product to be detected during the sliding process. Specifically, it includes: The average gray value of the template target image and the average gray value of the character image of the product to be detected are subjected to cross-correlation calculation to obtain the average gray cross-correlation term The average gray cross-correlation term is normalized with the standard deviation term of the template target image and the standard deviation term of the character image of the product to be detected through the following formula to obtain the NCC similarity between the template target image and the character image of the product to be detected : The standard deviation term The NCC similarity of the template target image and the character image of the product to be detected is as follows: ; Select a sliding window with a step size of 1 and a sliding size of Q in the character image of the product to be detected, calculate the NCC similarity of the template target image in each sliding window, regard the region with an NCC similarity higher than the set similarity threshold as the matching region, find the region with the largest NCC similarity as the matched region of interest, and generate a bounding box according to the position and size of the region of interest to identify the region of interest.
5. A method for detecting irregularly deformed characters of a product based on deep learning according to claim 1, characterized in that: In step S3, the deep learning character recognition model uses the Canny edge segmentation algorithm, which specifically includes the following sub-steps: Perform Gaussian blur on the corrected character image of the region of interest to reduce noise; Calculate the horizontal gradient of the blurred and corrected character image using the Sobel operator and the vertical gradient , to obtain the gradient magnitude . Perform non-maximum suppression on the gradient magnitude . Compare the gradient magnitude of each pixel in the corrected character image with the gradient magnitudes of two adjacent pixels along the gradient direction, and retain the local maximum value; Using the double-threshold edge connection method, threshold the corrected character image after non-maximum suppression, and set a high threshold , select the strong edge pixels greater than the high threshold for edge tracking to determine the connected edge lines; Finally, perform thinning processing on the connected edge lines, and the internal region of the edge lines is the character text region of the region of interest.
6. The method for detecting irregularly deformed characters of a product based on deep learning according to claim 5, wherein: Cluster the character text regions in the region of interest, identify and merge local regions belonging to the same text, filter the character text regions after clustering, and retain the character text regions where the text character regions are continuous and larger than the set region size threshold S a For the character text regions retained after filtering, generate text boxes using bounding rectangles.
7. A method for detecting irregularly deformed characters of products based on deep learning according to claim 4, characterized in that: In the step S4, the CRNN text recognition model outputs the text sequence of the character line in the text box through the following sub-steps: First step, convert the character image in the text box into the CNN model format, adjust the character image in the text box to a fixed size and normalize it; Second step, use the VGG16 model to perform convolution and pooling operations on the input character image in the text box to extract the feature map; Third step, serialize the extracted feature map to obtain the feature sequence; Fourth step, input the serialized feature sequence into the RNN model, combine the feature sequence by recursively combining the text character context information, and output a vector at each time step to represent the prediction result of that time step; Fifth step, perform a decoding operation on the prediction result sequence output by the RNN model, use the CTC loss function to convert the output sequence into the final text sequence and output it. The finally obtained text sequence is the text sequence recognized by the CRNN text recognition model for the character line in the text box.
8. An irregular deformation character detection device for products based on deep learning, characterized in that: It includes a character detection robotic arm. The execution end of the character detection robotic arm is equipped with a camera component for image acquisition of the irregularly deformed characters of the product. The character detection robotic arm performs image acquisition on the irregularly deformed characters of the product to be detected and transmits the acquired image to the device control host. The device control host determines whether the irregularly deformed characters of the product are qualified by the detection method according to any one of claims 1-7.
9. The device for detecting irregularly deformed characters of a product based on deep learning according to claim 8, wherein: It further includes a transport conveyor belt, a product loading module, and a product sorting module; The product loading module is located upstream of the transport conveyor belt and includes a loading robotic arm for grasping the product to be detected onto the transport conveyor belt; The product sorting module is located downstream of the transport conveyor belt and includes a sorting robotic arm for removing the detected product from the transport conveyor belt, and a material tray for classifying and placing qualified products and unqualified products. The sorting robotic arm sorts the detected products into the corresponding material trays according to the detection result of whether the product to be detected is qualified by the device control host; The character detection robotic arm is located between the loading robotic arm and the sorting robotic arm, and performs image acquisition on the deformed characters of the product to be detected grasped by the loading robotic arm onto the transport conveyor belt.
Citation Information
Patent Citations
Character detection method and device, electronic equipment and storage medium
CN116168395A
Character defect detection method and device, electronic equipment and storage medium
CN117727052A
Highly reflective metal surface character recognition method based on deep learning
CN117275010A