Product image detection method and system for bending machine based on machine vision

By using black and white cameras to collect grayscale photos in bending machine product quality inspection, perform binarization processing, and extracting multimodal features in combination with artificial intelligence and machine vision technology, the problem of low detection accuracy in complex environments in the existing technology is solved, and more flexible and efficient product quality inspection is achieved.

CN119205781BActive Publication Date: 2025-06-06WUXI HUADE AUTOMATION CONTROL TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411721837.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-28
Publication Date
2025-06-06
Estimated Expiration
2044-11-28

AI Technical Summary

Technical Problem

The existing bending machine product quality detection methods based on image recognition have low detection accuracy in complex or polluted environments, and have poor flexibility and adaptability when dealing with products of different shapes and hollow structures.

Method used

Black and white cameras are used to collect grayscale photos of bent products, and after binary processing, an image processing algorithm based on artificial intelligence and machine vision is introduced to extract multimodal features of binarized images and grayscale photos, and multimodal joint perception features are obtained through bidirectional fine-grained interaction to evaluate product quality.

Benefits of technology

It realizes more accurate product quality inspection in complex environments, improves the flexibility and adaptability of the inspection methods, and can effectively deal with products of different shapes and hollow structures.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119205781B_ABST
    Figure CN119205781B_ABST
Patent Text Reader

Abstract

The present application discloses a product image detection method and system of a bending machine based on machine vision, which collects grayscale photos of bending products through a black and white camera, and binarizes the grayscale photos of the bending products to obtain binary images of the bending products, and then introduces an image processing and analysis algorithm based on artificial intelligence and machine vision at the back end to analyze the binary images and grayscale photos of the bending products, thereby learning and capturing the binary modal representation information of the bending product state and the grayscale modal representation information of the bending product state, as well as the multimodal joint perception features based on the bending product state between the two information, and uses this to evaluate the fit, thereby judging whether the product quality is qualified. A more intelligent bending machine product quality detection method can be realized, which improves the flexibility and adaptability of detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of intelligent detection, and more specifically, to a product image detection method and system of a bending machine based on machine vision. Background Art

[0002] A press brake is a mechanical device used for bending metal sheets and other materials, and is widely used in the manufacturing industry. Ensuring the quality of bent products is crucial to the performance of the final product, and traditionally relies on the experience of skilled operators to judge the eligibility of products processed by press brakes. With the development of image recognition technology and hardware, automated detection methods have gradually replaced traditional manual inspections, aiming to reduce labor intensity and improve detection efficiency and accuracy. For example, in some automatic detection methods, the initial data of the bent product image is obtained through a grayscale camera, and then the product quality is automatically determined based on the recognition algorithm to see whether it meets the requirements. Despite this, the image recognition-based press brake product quality detection method still faces challenges, especially in complex or polluted environments, where the lens of the grayscale camera may be unstable or dirt may appear in the background, resulting in "tiny stains" in the initial data, which in turn affects the recognition accuracy.

[0003] In response to the above technical problems, Chinese patent CN117197128A proposes a product image detection method, system and bending machine for a bending machine, which obtains a grayscale photo of the product, binarizes it, and obtains the position information of the centralized black spot through centralization. Then, the deviation value of the centralized black spot is calculated, and the centralized black spot is calibrated according to the deviation value. Finally, the mathematical model of the product shape is obtained by the fitting algorithm, and compared with the predetermined standard model to evaluate the product quality. This method helps to reduce detection errors caused by equipment instability or background dirt by setting a centralized deviation judgment and correction mechanism.

[0004] However, in the above-mentioned product image detection method of the bending machine, the centralized black spot is obtained by simple binarization and centralization processing, and the product function model is obtained based on the deviation value and calibration information of the centralized black spot. This not only increases the complexity of the algorithm, resulting in low detection efficiency, but also has a low level of intelligence. When processing products of different shapes, the calculation method of the centralized deviation needs to be adjusted. Especially when facing products with hollow structures, the calculation of the center value requires additional consideration, and parameters or even algorithms may need to be readjusted, which makes the detection flexibility and adaptability of this method poor.

[0005] Therefore, an optimized product image detection solution for a bending machine is desired. Summary of the invention

[0006] In order to solve the above technical problems, the present application is proposed. The embodiment of the present application provides a product image detection method and system of a bending machine based on machine vision, which collects grayscale photos of bending products through a black and white camera, and binarizes the grayscale photos of the bending products to obtain binary images of the bending products, and then introduces an image processing and analysis algorithm based on artificial intelligence and machine vision at the back end to analyze the binary images and grayscale photos of the bending products, so as to learn and capture the binary modal representation information of the bending product state and the grayscale modal representation information of the bending product state, as well as the multimodal joint perception features based on the bending product state between the two information, and use this to evaluate the fit, so as to judge whether the product quality is qualified. In this way, artificial intelligence and machine vision technology can be used to automatically detect the quality of products processed by bending machines, thereby realizing a more intelligent bending machine product quality detection method and improving the flexibility and adaptability of the detection method.

[0007] According to one aspect of the present application, a method for detecting a product image of a bending machine based on machine vision is provided, which comprises:

[0008] Obtaining a grayscale photo of the bent product to be inspected captured by a black and white camera;

[0009] Performing a threshold-based binarization process on the grayscale photo of the bent product to be detected to obtain a binary image of the bent product;

[0010] Performing feature extraction based on the bending state on the binary image of the bent product and the grayscale photo of the bent product to be detected to obtain a binary modal representation of the bending product state and a grayscale modal representation of the bending product state;

[0011] The binary modal representation of the bending product state and the gray modal representation of the bending product state are subjected to bidirectional fine-grained deconstruction interaction to obtain a multi-modal joint perception feature of the bending product state;

[0012] Among them, the binary modal representation of the bending product state and the gray modal representation of the bending product state are subjected to bidirectional fine-grained deconstruction interaction to obtain the multimodal joint perception feature of the bending product state, including: fine-grained deconstruction of the binary modal representation of the bending product state and the gray modal representation of the bending product state to obtain a set of local features of the channel dimension of the binary modal representation of the bending product state and a set of local features of the channel dimension of the gray modal representation of the bending product state; feature attention interaction aggregation analysis is performed on the set of local features of the channel dimension of the binary modal representation of the bending product state and the set of local features of the channel dimension of the gray modal representation of the bending product state to obtain the multimodal joint perception feature of the bending product state;

[0013] Based on the multi-modal joint perception features of the bending product state, determine whether the product quality is qualified.

[0014] According to another aspect of the present application, a product image detection system for a bending machine based on machine vision is provided, which includes:

[0015] A grayscale photo acquisition module is used to acquire a grayscale photo of the bent product to be inspected collected by a black and white camera;

[0016] A binarization module, used for performing a threshold-based binarization process on the grayscale photo of the bent product to be detected to obtain a binarized image of the bent product;

[0017] A bending state feature extraction module is used to extract features based on the bending state from the binary image of the bent product and the grayscale photo of the bent product to be detected to obtain a binary modal representation of the bending product state and a grayscale modal representation of the bending product state;

[0018] A bending state multimodal joint perception module, used for performing bidirectional fine-grained deconstruction interaction on the binary modal representation of the bending product state and the grayscale modal representation of the bending product state to obtain a bending product state multimodal joint perception feature;

[0019] The product quality detection result generation module is used to determine whether the product quality is qualified based on the multi-modal joint perception characteristics of the bending product state.

[0020] Compared with the prior art, the present application provides a product image detection method and system for a bending machine based on machine vision, which uses a black and white camera to collect grayscale photos of the bending product, and binarizes the grayscale photos of the bending product to obtain a binary image of the bending product, and then introduces an image processing and analysis algorithm based on artificial intelligence and machine vision at the back end to analyze the binary image and grayscale photo of the bending product, thereby learning and capturing the binary modal representation information of the bending product state and the grayscale modal representation information of the bending product state, as well as the multimodal joint perception features based on the bending product state between the two information, and uses this to evaluate the fit, thereby judging whether the product quality is qualified. In this way, artificial intelligence and machine vision technology can be used to automatically detect the quality of products processed by the bending machine, thereby realizing a more intelligent bending machine product quality detection method and improving the flexibility and adaptability of the detection method. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] By describing the embodiments of the present application in more detail in conjunction with the accompanying drawings, the above and other purposes, features and advantages of the present application will become more apparent. The accompanying drawings are used to provide a further understanding of the embodiments of the present application and constitute a part of the specification. Together with the embodiments of the present application, they are used to explain the present application and do not constitute a limitation of the present application. In the accompanying drawings, the same reference numerals generally represent the same components or steps.

[0022] Figure 1 It is a flow chart of a method for detecting a product image of a bending machine based on machine vision according to an embodiment of the present application;

[0023] Figure 2 A data flow diagram of a method for detecting product images of a bending machine based on machine vision according to an embodiment of the present application;

[0024] Figure 3 It is a flowchart of sub-step S4 of the product image detection method of a bending machine based on machine vision according to an embodiment of the present application;

[0025] Figure 4 4 is a block diagram of a product image detection system for a bending machine based on machine vision according to an embodiment of the present application. DETAILED DESCRIPTION

[0026] Below, the exemplary embodiments according to the present application will be described in detail with reference to the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments of the present application, and it should be understood that the present application is not limited to the exemplary embodiments described here.

[0027] As shown in this application and claims, unless the context clearly indicates an exception, the words "a", "an", "an" and / or "the" do not refer to the singular and may also include the plural. Generally speaking, the terms "include" and "comprise" only indicate the inclusion of the steps and elements that have been clearly identified, and these steps and elements do not constitute an exclusive list. The method or device may also include other steps or elements.

[0028] Although the present application makes various references to certain modules in the system according to the embodiments of the present application, any number of different modules can be used and run on the user terminal and / or server. The modules are only illustrative, and different aspects of the system and method can use different modules.

[0029] Flowcharts are used in the present application to illustrate the operations performed by the system according to the embodiments of the present application. It should be understood that the preceding or following operations are not necessarily performed accurately in order. On the contrary, various steps may be processed in reverse order or simultaneously as required. Meanwhile, other operations may also be added to these processes, or a certain step or several steps of operations may be removed from these processes.

[0030] Below, the exemplary embodiments according to the present application will be described in detail with reference to the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments of the present application, and it should be understood that the present application is not limited to the exemplary embodiments described here.

[0031] In the technical solution of the present application, a product image detection method for a bending machine based on machine vision is proposed. Figure 1 A flowchart of a method for detecting a product image of a bending machine based on machine vision according to an embodiment of the present application. Figure 2 FIG. 1 is a data flow diagram of a method for detecting product images of a bending machine based on machine vision according to an embodiment of the present application. Figure 1 and Figure 2 As shown, according to the embodiment of the present application, the product image detection method of the bending machine based on machine vision includes the steps of: S1, obtaining a grayscale photo of the bent product to be detected collected by a black and white camera; S2, performing threshold-based binarization processing on the grayscale photo of the bent product to be detected to obtain a binary image of the bent product; S3, performing feature extraction based on the bending state on the binary image of the bent product and the grayscale photo of the bent product to be detected to obtain a binary modal representation of the bending product state and a grayscale modal representation of the bending product state; S4, performing a bidirectional fine-grained deconstruction interaction on the binary modal representation of the bending product state and the grayscale modal representation of the bending product state to obtain a multi-modal joint perception feature of the bending product state; S5, determining whether the product quality is qualified based on the multi-modal joint perception feature of the bending product state.

[0032] In particular, the S1 and S2 obtain a grayscale photo of the bent product to be detected collected by a black and white camera; and perform a threshold-based binarization process on the grayscale photo of the bent product to be detected to obtain a binarized image of the bent product. It should be understood that the binarization process is to divide the pixels in the image into foreground (usually white) and background (usually black) by setting a threshold, so that the edges and shapes of the objects in the photo can be highlighted. However, in real scenes, background noise is inevitable, such as dust, stains, etc. These factors may produce additional noise points in the image, especially when using a grayscale camera. These "tiny stains" may be mistaken for part of the object after binarization, thereby affecting the accuracy of subsequent feature extraction. In addition, if the camera or shooting device is unstable, such as slight vibration or dirty lens, the collected image will have additional noise or blur, which will also affect the binarization result. In this case, even if the optimal threshold is selected, these effects cannot be completely removed, which leads to deviations in the detection results. Therefore, in the technical solution of the present application, the accuracy of product quality detection is further improved by performing feature interactive joint analysis on the grayscale photo of the bent product to be detected and the binary image of the bent product.

[0033] In particular, S3 performs feature extraction based on the bending state on the binary image of the bent product and the grayscale photo of the bent product to be detected to obtain a binary modal representation of the bending product state and a grayscale modal representation of the bending product state. In a specific example of the present application, the binary image of the bent product is first input into a bending state feature extractor based on a convolutional neural network model to obtain a feature map of the binary modal representation of the bending product state as the binary modal representation of the bending product state; the binary image of the bent product is input into a bending state feature extractor based on a convolutional neural network model for feature mining to extract the binary modal implicit representation information about the bending product state in the binary image of the bent product, thereby obtaining a feature map of the binary modal representation of the bending product state. Similarly, the grayscale photo of the bent product to be detected is input into the bending state feature extractor based on a convolutional neural network model to obtain a feature map of the grayscale modal representation of the bending product state as the grayscale modal representation of the bending product state. By inputting the grayscale photo of the bent product to be detected into the bending state feature extractor based on the convolutional neural network model for feature mining, the grayscale modal implicit representation information about the state of the bent product in the grayscale photo of the bent product to be detected is extracted, thereby obtaining a grayscale modal representation feature map of the bent product state.

[0034] It is worth noting that Convolutional Neural Network (CNN) is a deep learning model specifically designed to process data with grid structures, such as images and audio. It has achieved great success in the fields of computer vision and pattern recognition. The core idea of ​​CNN is to use convolutional layers and pooling layers to extract the features of input data, and perform tasks such as classification or regression through fully connected layers. CNN gradually extracts high-level features of input data through the stacking of multiple convolutional layers and pooling layers, thereby achieving effective modeling and analysis of complex data.

[0035] In particular, in S4, the binary modal representation of the bending product state and the gray modal representation of the bending product state are subjected to bidirectional fine-grained deconstruction interaction to obtain the multi-modal joint perception feature of the bending product state. It should be understood that since the binary modal representation feature map of the bending product state and the gray modal representation feature map of the bending product state respectively contain the binary modal representation feature and the gray modal representation feature of the bending product state, in order to integrate the semantic information of the bending product state in these two feature maps and improve the model's understanding and perception ability of the bending product state, in the technical solution of the present application, the binary modal representation of the bending product state and the gray modal representation of the bending product state are further subjected to bidirectional fine-grained deconstruction interaction to obtain the multi-modal joint perception feature of the bending product state. The process based on bidirectional fine-grained deconstruction interaction aims to improve the semantic joint perception ability of the model through feature fine-grained deconstruction, global fine-grained interaction based on attention mechanism, feature coupling and feature fusion, so as to capture more detailed features to improve the accuracy and reliability of the final product quality inspection. In a specific example of this application, Figure 3 As shown, the S4 includes: S41, performing fine-grained deconstruction on the binary modal representation of the bending product state and the grayscale modal representation of the bending product state to obtain a set of local features of the channel dimension of the binary modal representation of the bending product state and a set of local features of the channel dimension of the grayscale modal representation of the bending product state; S42, performing feature attention interaction aggregation analysis on the set of local features of the channel dimension of the binary modal representation of the bending product state and the set of local features of the channel dimension of the grayscale modal representation of the bending product state to obtain the multimodal joint perception features of the bending product state.

[0036] Specifically, the S41 performs fine-grained deconstruction on the binary modal representation of the bending product state and the grayscale modal representation of the bending product state to obtain a set of local features of the channel dimension of the binary modal representation of the bending product state and a set of local features of the channel dimension of the grayscale modal representation of the bending product state. In a specific example of the present application, the feature map of the binary modal representation of the bending product state and the feature map of the grayscale modal representation of the bending product state are subjected to fine-grained feature deconstruction to obtain a set of local feature matrices of the channel dimension of the binary modal representation of the bending product state as a set of local features of the channel dimension of the binary modal representation of the bending product state and a set of local feature matrices of the channel dimension of the grayscale modal representation of the bending product state as a set of local features of the channel dimension of the grayscale modal representation of the bending product state. That is, the process of fine-grained deconstruction along the channel dimension can allow the model to perform more detailed segmentation and semantic understanding of the binary modal representation feature map of the bending product state and the grayscale modal representation feature map of the bending product state. This allows the model to analyze the local information in the feature map more carefully, thereby capturing more detailed features, which helps to better perceive the state characteristics of the bending product for quality inspection.

[0037] Specifically, the S42 performs feature attention interaction aggregation analysis on the set of local features of the channel dimension of the binary modal representation of the bending product state and the set of local features of the channel dimension of the grayscale modal representation of the bending product state to obtain the multi-modal joint perception features of the bending product state. That is, after the fine-grained deconstruction of the feature graph, the one-way global attention interaction module based on the transformer structure can measure the correlation between different feature matrices in the feature graph by utilizing the interaction of the multi-head attention mechanism in the Transformer structure, strengthen the relationship between the binary modal representation information and the grayscale modal representation information of the bending product state, enable the model to more comprehensively capture the semantic association information between different feature graphs, and pay more attention to the feature areas with rich information, ignoring irrelevant background information, so that the model can more accurately identify key information in a complex environment, such as the state and defects of the bending product, so as to provide a basis for subsequent quality inspection.

[0038] In an embodiment of the present application, a feature attention interaction aggregation analysis is performed on a set of local features of the channel dimension of the binary modal representation of the bending product state and a set of local features of the channel dimension of the grayscale modal representation of the bending product state to obtain the multimodal joint perception feature of the bending product state, including: first, each local feature matrix of the channel dimension of the binary modal representation of the bending product state in the set of local feature matrices of the channel dimension of the binary modal representation of the bending product state is used as a query feature matrix, and a set of local feature matrices of the channel dimension of the grayscale modal representation of the bending product state is used as a set of key feature matrices, and the query feature matrix and the set of the key feature matrix are input into a unidirectional global attention interaction module based on the first converter structure to obtain a set of local feature matrices of the channel dimension of the binary modal representation of the bending product state optimized by unidirectional global attention; similarly, Each local feature matrix of the grayscale modal representation channel dimension of the bending product state in the set of local feature matrices of the grayscale modal representation channel dimension of the bending product state is used as a query feature matrix, and the set of local feature matrices of the binary modal representation channel dimension of the bending product state is used as a set of key feature matrices, and the set of the query feature matrix and the set of the key feature matrix are input into a one-way global attention interaction module based on the second converter structure to obtain a set of local feature matrices of the grayscale modal representation channel dimension of the bending product state optimized by one-way global attention; it should be understood that the effect of the one-way global attention interaction is to enhance the fine-grained channel relationship within the feature graph, so that the model can pay more attention to the key information related to the bending product state and quality inspection tasks in the feature graph, improve the quality and discrimination of the feature representation, and thus help improve the performance of the model in subsequent tasks. Next, the set of local feature matrices of the channel dimension of the binary modal representation of the state of the bending product optimized by the one-way global attention and the set of local feature matrices of the channel dimension of the gray modal representation of the state of the bending product optimized by the one-way global attention are coupled along the channel dimension to obtain the feature graph of the binary modal representation of the state of the bending product optimized by the one-way global interaction and the feature graph of the gray modal representation of the state of the bending product optimized by the one-way global interaction; through the coupling processing along the channel dimension, the binary modal representation of the state of the bending product optimized by the attention and the deep local feature matrix can be recombined along the channel dimension to form a complete feature graph, providing a complete feature representation for the final fusion. Finally, the feature aggregation of the binary modal representation feature graph of the state of the bending product optimized by the one-way global interaction and the gray modal representation feature graph of the state of the bending product optimized by the one-way global interaction is performed to obtain the multi-modal joint perception feature of the state of the bending product.

[0039] Among them, each local feature matrix of the binary modal representation channel dimension of the bending product state in the set of the local feature matrix of the binary modal representation channel dimension of the bending product state is used as a query feature matrix, and the set of the local feature matrix of the grayscale modal representation channel dimension of the bending product state is used as a set of key feature matrices, and the query feature matrix and the set of the key feature matrix are input into a one-way global attention interaction module based on the first converter structure to obtain a set of local feature matrices of the binary modal representation channel dimension of the bending product state by one-way global attention optimization. The process includes: selecting a predetermined local feature matrix of the binary modal representation channel dimension of the bending product state from the set of local feature matrices of the binary modal representation channel dimension of the bending product state as the query feature matrix; calculating the product between the predetermined local feature matrix of the binary modal representation channel dimension of the bending product state and the transposed matrix of each local feature matrix of the grayscale modal representation channel dimension of the bending product state in the set of local feature matrices of the grayscale modal representation channel dimension of the bending product state to obtain a bending matrix; A set of local semantic interaction feature matrices of the channel dimension of the binary modal-grayscale modal representation of the state of the bending product; divide the scale square root of the predetermined local feature matrix of the channel dimension of the binary modal representation of the state of the bending product by the position point of each local semantic interaction feature matrix of the channel dimension of the binary modal-grayscale modal representation of the state of the bending product, and then use the softmax function to perform soft maximum normalization processing on each feature matrix in the set of the obtained feature matrices to obtain a set of local weight matrices of the channel dimension of the binary modal representation of the state of the bending product; use the set of local weight matrices of the channel dimension of the binary modal representation of the state of the bending product as the weighted weight, and calculate the weighted sum of the position between the local feature matrices of the channel dimension of the binary modal representation of the state of the bending product in the set of the local feature matrix of the channel dimension of the binary modal representation of the state of the bending product to obtain the local feature matrix of the channel dimension of the binary modal representation of the state of the bending product optimized by the unidirectional global attention. In the technical solution of the present application, the position-weighted sum between the one-way global interactive optimization bending product state binary modal representation feature map and the one-way global interactive optimization bending product state grayscale modal representation feature map is calculated to obtain the bending product state multimodal joint significant perception feature map as the bending product state multimodal joint perception feature. That is, the weighted sum operation integrates the key information in the two feature maps and generates a bending product state multimodal joint perception feature representation containing rich interaction information.In particular, the final generated "multi-modal joint significant perception feature map of the bending product status" is the result of a fusion of the interactive information between the binary modal representation features and the grayscale modal representation features of the bending product status. It contains the complementary information of the two modalities, which helps the model to more accurately identify and evaluate the status of the bending product, thereby improving the accuracy and reliability of the final quality inspection.

[0040] Similarly, each local feature matrix of the grayscale modal representation channel dimension of the bending product state in the set of local feature matrices of the grayscale modal representation channel dimension of the bending product state is used as a query feature matrix, and the set of local feature matrices of the binary modal representation channel dimension of the bending product state is used as a set of key feature matrices, and the query feature matrix and the set of key feature matrices are input into a one-way global attention interaction module based on a second converter structure to obtain a set of local feature matrices of the grayscale modal representation channel dimension of the bending product state optimized by one-way global attention. The process includes: selecting a predetermined local feature matrix of the grayscale modal representation channel dimension of the bending product state from the set of local feature matrices of the grayscale modal representation channel dimension of the bending product state as the query feature matrix; calculating the product between the predetermined local feature matrix of the grayscale modal representation channel dimension of the bending product state and the transposed matrix of each local feature matrix of the binary modal representation channel dimension of the bending product state in the set of local feature matrices of the binary modal representation channel dimension of the bending product state to obtain to a set of local semantic interaction feature matrices of the bending product state grayscale modality-binary modality representation channel dimension; divide each bending product state grayscale modality-binary modality representation channel dimension local semantic interaction feature matrix in the set of the bending product state grayscale modality-binary modality representation channel dimension local semantic interaction feature matrix by the scale square root of the predetermined bending product state grayscale modality representation channel dimension local feature matrix according to the position point, and then use the softmax function to perform soft maximum value normalization processing on each feature matrix in the obtained set of feature matrices to obtain a set of local weight matrices of the bending product state grayscale modality representation channel dimension; use the set of local weight matrices of the bending product state grayscale modality representation channel dimension as the weighted weight, calculate the position-weighted sum between each bending product state grayscale modality representation channel dimension local feature matrix in the set of the bending product state grayscale modality representation channel dimension local feature matrix to obtain the unidirectional global attention optimized bending product state grayscale modality representation channel dimension local feature matrix.

[0041] In summary, in the above embodiment, the binary modal representation of the bending product state and the gray modal representation of the bending product state are subjected to bidirectional fine-grained deconstruction interaction to obtain the multimodal joint perception feature of the bending product state, including: the binary modal representation of the bending product state and the gray modal representation of the bending product state are subjected to bidirectional fine-grained deconstruction interaction to obtain the multimodal joint perception feature of the bending product state based on the following bidirectional fine-grained feature joint formula; wherein the bidirectional fine-grained feature joint formula is:

[0042]

[0043]

[0044]

[0045]

[0046]

[0047]

[0048]

[0049] in, and represent the binary modal characterization feature graph of the bending product state and the grayscale modal characterization feature graph of the bending product state respectively, It is a feature fine-grained deconstruction operation. They are the first, second, and third dimensions along the channel dimension in the binary modal representation feature graph of the bending product state. and The local feature matrix of the channel dimension is represented by the binary modal representation of the bending product state. They are the first, second, and third dimensions along the channel dimension in the grayscale modal representation feature map of the bending product state. and The grayscale modal representation of the bending product state channel dimension local feature matrix, For the The scale of the local feature matrix of the channel dimension of the binary modal representation of the bending product state, is matrix multiplication, for function, is the number of local feature matrices of the channel dimension of the binary modal representation of the bending product state in the binary modal representation feature graph of the bending product state, For the said The one-way global attention optimization of the local feature matrix of the channel dimension corresponding to the binary modal representation of the state of the bending product, For the The one-way global attention optimization of the local feature matrix of the channel dimension corresponding to the grayscale modal representation of the bending product state grayscale modal representation, Indicates that the features are coupled along the channel dimension, and They are the binary modal characterization feature map of the state of the one-way global interactive optimization bending product and the grayscale modal characterization feature map of the state of the one-way global interactive optimization bending product. and are weighted hyperparameters of the binary modal characterization feature map of the one-way global interactive optimization bending product state and the grayscale modal characterization feature map of the one-way global interactive optimization bending product state, respectively. It is a multi-modal joint salient perception feature map of the bending product state.

[0050] It is worth mentioning that in other specific examples of the present application, the binary modal representation of the bending product state and the grayscale modal representation of the bending product state can also be subjected to bidirectional fine-grained deconstruction interaction in other ways to obtain the multimodal joint perception feature of the bending product state, for example: input the binary modal representation of the bending product state and the grayscale modal representation of the bending product state; decompose the binary modal representation into a series of fine-grained binary feature maps; decompose the grayscale modal representation into a series of fine-grained grayscale feature maps; for each binary feature map, interact with all grayscale feature maps, and the interaction operation can be dot product, convolution or other nonlinear transformation; weighted summation is used to fuse the feature maps generated by the bidirectional interaction to form a multimodal joint perception feature; multimodal joint perception features are extracted from the fused features to obtain the multimodal joint perception feature of the bending product state.

[0051] In particular, the S5 determines whether the product quality is qualified based on the multimodal joint perception features of the bending product state. In a specific example of the present application, the multimodal joint significant perception feature map of the bending product state is input into a fit evaluation module based on a decoder to obtain a fit decoding evaluation value; that is, the multimodal joint interactive perception features of the bending product state are used to perform decoding regression, so as to perform fit evaluation, thereby determining whether the product quality is qualified based on the comparison between the fit decoding evaluation value and the preset threshold. In this way, artificial intelligence and machine vision technology can be used to automatically detect the quality of products processed by the bending machine, thereby realizing a more intelligent bending machine product quality detection method and improving the flexibility and adaptability of the detection method. Furthermore, based on the comparison between the fit decoding evaluation value and the preset threshold, it is determined whether the product quality is qualified. Specifically, in response to the fit decoding evaluation value being less than the predetermined threshold, the product quality is determined to be unqualified.

[0052] In a preferred example, inputting the multimodal joint significant perceptual feature map of the bending product state into a decoder-based fit evaluation module to obtain a fit decoding evaluation value includes:

[0053] Determine the feature mean and feature variance of the multi-modal joint significant perception feature map of the bending product state, and divide the feature mean by the feature variance to obtain the multi-modal joint significant perception probability statistic of the bending product state:

[0054]

[0055] in, and They respectively represent the feature mean and feature variance of the multimodal joint significant perceptual feature map of the bending product state, Indicates the probability statistics of multi-modal joint significant perception of the bending product status;

[0056] The multimodal joint significant perception feature map of the bending product state is multiplied by the inverse of the maximum eigenvalue in the multimodal joint significant perception feature map of the bending product state to obtain a multimodal joint significant perception probability constraint map of the bending product state:

[0057]

[0058] in, represents the multi-modal joint salient perceptual feature map of the bending product state, represents the maximum eigenvalue in the multimodal joint salient perceptual feature map of the bending product state, It means point multiplication by position. Represents the multi-modal joint saliency perception probability constraint graph of the bending product status;

[0059] After adding the bending product state multimodal joint significant perception probability constraint graph and the bending product state multimodal joint significant perception probability statistical value, the logarithmic value with base 2 is calculated to obtain the bending product state multimodal joint significant perception information interaction graph:

[0060]

[0061] in, Represents the multi-modal joint saliency perception probability constraint graph of the bending product state, It represents the probability statistics of multi-modal joint significant perception of the bending product status, Indicates adding by position point. Represents the multi-modal joint salient perception information interaction diagram of the bending product status;

[0062] After performing point subtraction on the multimodal joint significant perception probability constraint graph of the bending product state with the multimodal joint significant perception probability statistics of the bending product state, the inverse of each eigenvalue is calculated to obtain the multimodal joint significant perception sequence constraint graph of the bending product state:

[0063]

[0064] in, Represents the multi-modal joint saliency perception probability constraint graph of the bending product state, It represents the probability statistics of multi-modal joint significant perception of the bending product status, It means to subtract by position point. Represents the multi-modal joint saliency perception sequence constraint graph of the bending product status;

[0065] The multimodal joint significant perception information interaction graph of the bending product state is interpolated with the multimodal joint significant perception sequence constraint graph of the bending product state to obtain an optimized multimodal joint significant perception feature graph of the bending product state;

[0066] The optimized multi-modal joint significant perceptual feature map of the bending product state is input into a decoder-based fit evaluation module to obtain a fit decoding evaluation value.

[0067] Here, in the preferred example, since the bending product state binary modal representation feature map and the bending product state grayscale modal representation feature map respectively represent the image semantic features of the bending product binary image and the image semantic features of the grayscale photo of the bent product to be detected, when they are input into the bidirectional global attention joint perception module based on fine-grained deconstruction, the cross-modal image semantic feature differences will have different attention weights based on the fine-grained decoupling of the spatial distribution of image semantic features. Therefore, the bending product state multimodal joint salient perception feature map will also have a diversified set expression distribution of cross-modal global aggregation features. Therefore, when the bending product state multimodal joint salient perception feature map is decoded by the decoder, it will affect the accuracy of the decoding result.

[0068] Therefore, considering that when the weight matrix of the decoder acts on the multimodal jointly significant perceptual feature vector of the bending product state after the multimodal jointly significant perceptual feature map of the bending product state is expanded, the distribution diversity of the multimodal jointly significant perceptual feature map of the bending product state causes the uncertainty of the weight stimulus parameters of the weight matrix of the decoder, resulting in the lack of a posteriori probability density inference of the multimodal jointly significant perceptual feature map of the bending product state through the action of the weight matrix, which will affect the accuracy of the decoding results.

[0069] In this case, the probability statistical characteristics of the multimodal joint significant perception feature map of the bending product state are used to simulate the mesoscale interaction structure under probability constraints between the feature value scale and the feature map scale of the multimodal joint significant perception feature map of the bending product state, so as to construct a bidirectional latent variable motif based on the mesoscale short sequence relative to the probability statistical value to perform mesoscale bidirectional migration, and perform posterior recovery based on the sequence constraints on the interactive information, thereby improving the convergence effect in the probability density domain and improving the accuracy of the fit decoding evaluation value obtained by the decoder-based fit evaluation module of the multimodal joint significant perception feature map of the bending product state. In this way, the fit evaluation can be performed more accurately, thereby judging whether the product quality is qualified, realizing more intelligent bending machine product quality detection, and improving the flexibility and adaptability of the detection process.

[0070] In summary, according to the embodiment of the present application, the product image detection method of the bending machine based on machine vision is explained, which uses a black and white camera to collect grayscale photos of the bending product, and binarizes the grayscale photos of the bending product to obtain a binary image of the bending product, and then introduces an image processing and analysis algorithm based on artificial intelligence and machine vision at the back end to analyze the binary image and grayscale photo of the bending product, so as to learn and capture the binary modal representation information of the bending product state and the grayscale modal representation information of the bending product state, as well as the multimodal joint perception features based on the bending product state between the two information, and use this to evaluate the fit, so as to determine whether the product quality is qualified. In this way, artificial intelligence and machine vision technology can be used to automatically detect the quality of products processed by the bending machine, thereby realizing a more intelligent bending machine product quality detection method and improving the flexibility and adaptability of the detection method.

[0071] Furthermore, a product image detection system for a bending machine based on machine vision is also provided.

[0072] Figure 4 FIG. 1 is a block diagram of a product image detection system for a bending machine based on machine vision according to an embodiment of the present application. Figure 4 As shown, according to the embodiment of the present application, the product image detection system 300 of the bending machine based on machine vision includes: a grayscale photo acquisition module 310, which is used to obtain a grayscale photo of the bent product to be detected collected by a black and white camera; a binarization module 320, which is used to perform threshold-based binarization processing on the grayscale photo of the bent product to be detected to obtain a binary image of the bent product; a bending state feature extraction module 330, which is used to perform bending state-based feature extraction on the binary image of the bent product and the grayscale photo of the bent product to be detected to obtain a binary modal representation of the bending product state and a grayscale modal representation of the bending product state; a bending state multimodal joint perception module 340, which is used to perform bidirectional fine-grained deconstruction interaction on the binary modal representation of the bending product state and the grayscale modal representation of the bending product state to obtain a multimodal joint perception feature of the bending product state; a product quality detection result generation module 350, which is used to determine whether the product quality is qualified based on the multimodal joint perception feature of the bending product state.

[0073] As described above, the product image detection system 300 for a bending machine based on machine vision according to an embodiment of the present application can be implemented in various wireless terminals, such as a server having a product image detection algorithm for a bending machine based on machine vision. In a possible implementation, the product image detection system 300 for a bending machine based on machine vision according to an embodiment of the present application can be integrated into a wireless terminal as a software module and / or a hardware module. For example, the product image detection system 300 for a bending machine based on machine vision can be a software module in the operating system of the wireless terminal, or can be an application developed for the wireless terminal; of course, the product image detection system 300 for a bending machine based on machine vision can also be one of the many hardware modules of the wireless terminal.

[0074] Alternatively, in another example, the product image detection system 300 of the bending machine based on machine vision and the wireless terminal may also be separate devices, and the product image detection system 300 of the bending machine based on machine vision may be connected to the wireless terminal via a wired and / or wireless network and transmit interactive information in accordance with an agreed data format.

[0075] The embodiments of the present disclosure have been described above, and the above description is exemplary, not exhaustive, and is not limited to the disclosed embodiments. Many modifications and changes will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described embodiments. The selection of terms used herein is intended to best explain the principles of the embodiments, practical applications, or improvements to the technology in the market, or to enable other persons of ordinary skill in the art to understand the embodiments disclosed herein.

Claims

1. A product image detection method for a bending machine based on machine vision, characterized in that: include: Obtaining a grayscale photo of the bent product to be inspected captured by a black and white camera; Performing a threshold-based binarization process on the grayscale photo of the bent product to be detected to obtain a binary image of the bent product; Performing feature extraction based on the bending state on the binary image of the bent product and the grayscale photo of the bent product to be detected to obtain a binary modal representation of the bending product state and a grayscale modal representation of the bending product state; The binary modal representation of the bending product state and the gray modal representation of the bending product state are subjected to bidirectional fine-grained deconstruction interaction to obtain a multi-modal joint perception feature of the bending product state; Among them, the binary modal representation of the bending product state and the gray modal representation of the bending product state are subjected to bidirectional fine-grained deconstruction interaction to obtain the multimodal joint perception feature of the bending product state, including: fine-grained deconstruction of the binary modal representation of the bending product state and the gray modal representation of the bending product state to obtain a set of local features of the channel dimension of the binary modal representation of the bending product state and a set of local features of the channel dimension of the gray modal representation of the bending product state; feature attention interaction aggregation analysis is performed on the set of local features of the channel dimension of the binary modal representation of the bending product state and the set of local features of the channel dimension of the gray modal representation of the bending product state to obtain the multimodal joint perception feature of the bending product state; Determine whether the product quality is qualified based on the multi-modal joint perception characteristics of the bending product state; Among them, the feature attention interaction aggregation analysis is performed on the set of local features of the channel dimension of the binary modal representation of the bending product state and the set of local features of the channel dimension of the grayscale modal representation of the bending product state to obtain the multimodal joint perception features of the bending product state, including: Using each local feature matrix of the channel dimension of the binary modal representation of the bending product state in the set of the local feature matrix of the channel dimension of the binary modal representation of the bending product state as a query feature matrix, and using the set of local feature matrices of the channel dimension of the grayscale modal representation of the bending product state as a set of key feature matrices, inputting the query feature matrix and the set of key feature matrices into a unidirectional global attention interaction module based on the first converter structure to obtain a set of local feature matrices of the channel dimension of the binary modal representation of the bending product state optimized by unidirectional global attention; Using each local feature matrix of the bending product state grayscale modal representation channel dimension in the set of the bending product state grayscale modal representation channel dimension local feature matrix as a query feature matrix, using the set of the bending product state binary modal representation channel dimension local feature matrix as a set of key feature matrices, inputting the query feature matrix and the set of the key feature matrix into a unidirectional global attention interaction module based on a second converter structure to obtain a set of unidirectional global attention optimized bending product state grayscale modal representation channel dimension local feature matrices; The set of local feature matrices of the one-way global attention optimization bending product state binary modal representation channel dimension and the set of local feature matrices of the one-way global attention optimization bending product state grayscale modal representation channel dimension are respectively coupled along the channel dimension to obtain a one-way global interactive optimization bending product state binary modal representation feature map and a one-way global interactive optimization bending product state grayscale modal representation feature map; The binary modal characterization feature map of the one-way global interactive optimization bending product state and the grayscale modal characterization feature map of the one-way global interactive optimization bending product state are feature aggregated to obtain the multi-modal joint perception feature of the bending product state.

2. The product image detection method of a bending machine based on machine vision according to claim 1 is characterized in that: The binary image of the bent product and the grayscale photo of the bent product to be detected are subjected to feature extraction based on the bending state to obtain a binary modal representation of the bending product state and a grayscale modal representation of the bending product state, including: Inputting the bent product binary image into a bending state feature extractor based on a convolutional neural network model to obtain a bent product state binary modal representation feature map as the bent product state binary modal representation; The grayscale photo of the bent product to be detected is input into the bending state feature extractor based on the convolutional neural network model to obtain a bending product state grayscale modal representation feature map as the bending product state grayscale modal representation.

3. The product image detection method of a bending machine based on machine vision according to claim 2 is characterized in that: The binary modal representation of the bending product state and the grayscale modal representation of the bending product state are fine-grained deconstructed to obtain a set of local features of the channel dimension of the binary modal representation of the bending product state and a set of local features of the channel dimension of the grayscale modal representation of the bending product state, including: fine-grained feature deconstruction of the feature graph of the binary modal representation of the bending product state and the feature graph of the grayscale modal representation of the bending product state to obtain a set of local feature matrices of the channel dimension of the binary modal representation of the bending product state as a set of local features of the channel dimension of the binary modal representation of the bending product state and a set of local feature matrices of the channel dimension of the grayscale modal representation of the bending product state as a set of local features of the channel dimension of the grayscale modal representation of the bending product state.

4. The product image detection method of a bending machine based on machine vision according to claim 3 is characterized in that: Each local feature matrix of the binary modal representation channel dimension of the bending product state in the set of the local feature matrix of the binary modal representation channel dimension of the bending product state is used as a query feature matrix, and the set of the local feature matrix of the grayscale modal representation channel dimension of the bending product state is used as a set of key feature matrices, and the set of the query feature matrix and the key feature matrix are input into a unidirectional global attention interaction module based on the first converter structure to obtain a set of local feature matrices of the binary modal representation channel dimension of the bending product state optimized by unidirectional global attention, including: Selecting a predetermined local feature matrix of the binary modal representation channel dimension of the bending product state from the set of the local feature matrix of the binary modal representation channel dimension of the bending product state as a query feature matrix; Calculate the product of the predetermined bending product state binary modal representation channel dimension local feature matrix and the transposed matrix of each bending product state grayscale modal representation channel dimension local feature matrix in the set of bending product state grayscale modal representation channel dimension local feature matrix to obtain a set of bending product state binary modal-grayscale modal representation channel dimension local semantic interaction feature matrices; After dividing the scale square root of the predetermined local feature matrix of the binary modal representation channel dimension of the bending product state by each local semantic interaction feature matrix of the binary modal representation channel dimension of the bending product state in the set of the local semantic interaction feature matrix of the binary modal representation channel dimension of the bending product state at the position point, each feature matrix in the obtained set of feature matrices is subjected to soft maximum normalization processing by using a softmax function to obtain a set of local weight matrices of the binary modal representation channel dimension of the bending product state; Taking the set of local weight matrices of the bending product state binary modal representation channel dimension as weighted weights, the position-weighted sum of the local feature matrices of the bending product state binary modal representation channel dimension in the set of local feature matrices of the bending product state binary modal representation channel dimension is calculated to obtain the unidirectional global attention optimized bending product state binary modal representation channel dimension local feature matrix.

5. The product image detection method of a bending machine based on machine vision according to claim 4 is characterized in that: Each local feature matrix of the bending product state grayscale modal representation channel dimension in the set of the bending product state grayscale modal representation channel dimension local feature matrix is ​​used as a query feature matrix, and the set of the bending product state binary modal representation channel dimension local feature matrix is ​​used as a set of key feature matrices, and the query feature matrix and the set of the key feature matrix are input into a one-way global attention interaction module based on a second converter structure to obtain a set of one-way global attention optimized bending product state grayscale modal representation channel dimension local feature matrices, including: Selecting a predetermined local feature matrix of the channel dimension of grayscale modal representation of the state of the bent product from the set of local feature matrices of the channel dimension of grayscale modal representation of the state of the bent product as a query feature matrix; Calculate the product of the predetermined bending product state grayscale modal representation channel dimension local feature matrix and the transposed matrix of each bending product state binary modal representation channel dimension local feature matrix in the set of bending product state binary modal representation channel dimension local feature matrix to obtain a set of bending product state grayscale modal-binary modal representation channel dimension local semantic interaction feature matrices; After dividing the scale square root of the predetermined local feature matrix of the bending product state grayscale modality representation channel dimension by each of the local semantic interaction feature matrices of the bending product state grayscale modality representation channel dimension in the set of the local semantic interaction feature matrices of the bending product state grayscale modality representation channel dimension at the position point, each feature matrix in the obtained set of feature matrices is subjected to soft maximum normalization processing using a softmax function to obtain a set of local weight matrices of the bending product state grayscale modality representation channel dimension; Taking the set of local weight matrices of the bending product state grayscale modal representation channel dimension as weighted weights, the position-weighted sum of the local feature matrices of the bending product state grayscale modal representation channel dimension in the set of local feature matrices of the bending product state grayscale modal representation channel dimension is calculated to obtain the unidirectional global attention optimized bending product state grayscale modal representation channel dimension local feature matrix.

6. The product image detection method of a bending machine based on machine vision according to claim 5 is characterized in that: The one-way global interactive optimization bending product state binary modal characterization feature map and the one-way global interactive optimization bending product state grayscale modal characterization feature map are feature aggregated to obtain the bending product state multimodal joint perception feature, including: calculating the position-weighted sum between the one-way global interactive optimization bending product state binary modal characterization feature map and the one-way global interactive optimization bending product state grayscale modal characterization feature map to obtain the bending product state multimodal joint significant perception feature map as the bending product state multimodal joint perception feature.

7. The product image detection method of a bending machine based on machine vision according to claim 6 is characterized in that: Based on the multi-modal joint perception features of the bending product state, determining whether the product quality is qualified includes: Inputting the multi-modal joint significant perceptual feature map of the bending product state into a decoder-based fit evaluation module to obtain a fit decoding evaluation value; Based on the comparison between the fit decoding evaluation value and a preset threshold, it is determined whether the product quality is qualified.

8. The product image detection method of a bending machine based on machine vision according to claim 7 is characterized in that: In response to the fit decoding evaluation value being less than the predetermined threshold, it is determined that the product quality is unqualified.

9. A product image detection system for a bending machine based on machine vision, characterized in that: include: A grayscale photo acquisition module is used to acquire a grayscale photo of the bent product to be inspected collected by a black and white camera; A binarization module, used for performing a threshold-based binarization process on the grayscale photo of the bent product to be detected to obtain a binarized image of the bent product; A bending state feature extraction module is used to extract features based on the bending state from the binary image of the bent product and the grayscale photo of the bent product to be detected to obtain a binary modal representation of the bending product state and a grayscale modal representation of the bending product state; A bending state multimodal joint perception module, used for performing bidirectional fine-grained deconstruction interaction on the binary modal representation of the bending product state and the grayscale modal representation of the bending product state to obtain a bending product state multimodal joint perception feature; A product quality detection result generation module, used to determine whether the product quality is qualified based on the multi-modal joint perception characteristics of the bending product state; Among them, the feature attention interaction aggregation analysis is performed on the set of local features of the channel dimension of the binary modal representation of the bending product state and the set of local features of the channel dimension of the grayscale modal representation of the bending product state to obtain the multimodal joint perception features of the bending product state, including: Using each local feature matrix of the channel dimension of the binary modal representation of the bending product state in the set of the local feature matrix of the channel dimension of the binary modal representation of the bending product state as a query feature matrix, and using the set of local feature matrices of the channel dimension of the grayscale modal representation of the bending product state as a set of key feature matrices, inputting the query feature matrix and the set of key feature matrices into a unidirectional global attention interaction module based on the first converter structure to obtain a set of local feature matrices of the channel dimension of the binary modal representation of the bending product state optimized by unidirectional global attention; Using each local feature matrix of the bending product state grayscale modal representation channel dimension in the set of the bending product state grayscale modal representation channel dimension local feature matrix as a query feature matrix, using the set of the bending product state binary modal representation channel dimension local feature matrix as a set of key feature matrices, inputting the query feature matrix and the set of the key feature matrix into a unidirectional global attention interaction module based on a second converter structure to obtain a set of unidirectional global attention optimized bending product state grayscale modal representation channel dimension local feature matrices; The set of local feature matrices of the one-way global attention optimization bending product state binary modal representation channel dimension and the set of local feature matrices of the one-way global attention optimization bending product state grayscale modal representation channel dimension are respectively coupled along the channel dimension to obtain a one-way global interactive optimization bending product state binary modal representation feature map and a one-way global interactive optimization bending product state grayscale modal representation feature map; The binary modal characterization feature map of the one-way global interactive optimization bending product state and the grayscale modal characterization feature map of the one-way global interactive optimization bending product state are feature aggregated to obtain the multi-modal joint perception feature of the bending product state.

Citation Information

Patent Citations

  • Multi-modal high-frame-rate frame insertion method based on edge enhancement

    CN117097858A

  • Product image detection method and system of bending machine and bending machine

    CN117197128A