Multi-mode-based steel coil surface defect detection method, device, equipment and medium

Through the multimodal detection method, the surface defect of the steel coil is detected by using the encoder decoder model, which solves the problem of insufficient detection accuracy of a single mode and achieves higher detection accuracy and application scope.

CN120047750APending Publication Date: 2025-05-27CISDI INFORMATION TECH CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510212229.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-25
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

In the prior art, steel coil detection relies on single-modal image data, making it difficult to fully capture the detailed information of product defects, resulting in insufficient defect recognition accuracy, poor adaptability to complex defect types, and poor stability under different environmental conditions.

Method used

Using a multimodal steel coil surface defect detection method, by obtaining the input image containing the steel coil surface and its corresponding task description, encode and converting, determining the discrete sequence, and inputting it into the encoder decoder model (such as the Pix2SeqV2 model), to obtain the task output sequence, and then perform inverse quantization inference to determine the detection box, mask and defect description information.

Benefits of technology

It improves the accuracy and scope of application of steel coil surface defect detection, enhances the ability to identify complex defect types, and maintains stability under different environmental conditions, and improves the performance of industrial image processing systems and the automatic detection of steel coil surface defects.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120047750A_ABST
    Figure CN120047750A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of steel coil detection, and discloses a multi-mode-based steel coil surface defect detection method, device and equipment and a medium, and the method comprises the following steps: obtaining an input image containing a steel coil surface and a task description corresponding to the input image; performing code conversion on the input image and the task description, and determining a discrete sequence; inputting the discrete sequence into an encoder and decoder model to obtain a task output sequence; and performing inverse quantization reasoning according to the task output sequence, and determining a detection frame, a mask and defect description information of an input image so as to complete surface defect detection of the steel coil. According to the method, samples of different defect types can be rapidly obtained through the sample pair formed by the input image and the task description, rapid iteration of model training is facilitated, and meanwhile, through the multi-modal image set, the detection frame, the mask and the defect description information are directly output, the steel coil surface defect detection precision is improved, and the detection application range is widened.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of steel coil detection, and in particular to a multi-modal-based steel coil surface defect detection method, device, equipment and medium. Background Art

[0002] In the field of modern industrial manufacturing, especially in precision manufacturing and automated production lines, the control of product quality has become increasingly stringent. In order to ensure the high quality standards of products, manufacturers need to implement efficient defect detection during the production process.

[0003] However, in the related art, steel coils are often visually inspected based on image data that relies on a single modality, such as visible light images or infrared images. This method is difficult to fully capture detailed information about product defects, and can easily lead to problems such as insufficient defect recognition accuracy, poor adaptability to complex defect types, and weak stability under different environmental conditions. Summary of the invention

[0004] The present invention provides a multi-modal-based steel coil surface defect detection method, device, equipment and medium to solve the technical problems of low steel coil detection accuracy and single detection type.

[0005] In the first aspect, the present application provides a multimodal steel coil surface defect detection method, comprising: obtaining an input image including a steel coil surface and a task description corresponding to the input image; encoding and converting the input image and the task description to determine a discrete sequence, wherein the task descriptions of different task types use different encoding methods; inputting the discrete sequence into an encoder-decoder model to obtain a task output sequence; performing inverse quantization reasoning based on the task output sequence to determine the detection box, mask and defect description information of the input image to complete the steel coil surface defect detection.

[0006] In some embodiments of the first aspect, the task descriptions of different task types adopt different encoding methods, including: if the task description is a target detection task, the bounding box of the target is obtained by quantizing continuous image coordinates in the input image, and the bounding box and the target description are converted into a series of discrete sequences; if the task description is an instance segmentation task, the polygons corresponding to the instance mask in the input image are predicted in the form of an image coordinate sequence, and the coordinates corresponding to the polygons are quantized into a discrete sequence; if the task description is an image description task, the task description is converted into a text sequence.

[0007] In some embodiments of the first aspect, the encoder-decoder model is determined by:

[0008] A data set consisting of multiple groups of sample pairs is obtained, each group of the sample pairs includes the input image and the corresponding task description; the data set is divided into a training set, a test set and a validation set, and the encoder-decoder model is a Pix2SeqV2 model; the Pix2SeqV2 model is trained using the training set, and each initial parameter is adjusted by back-propagation to obtain a first encoder-decoder model; the validation set is used to adjust the hyperparameters of the first encoder-decoder model, and a second encoder-decoder model is obtained by fitting; the test set is used to evaluate the performance of the second encoder-decoder model, and a third encoder-decoder model with satisfactory generalization performance is output, and the third encoder-decoder model is used as the final encoder-decoder model.

[0009] In some embodiments of the first aspect, the encoder-decoder model includes an encoder, which is based on a Transformer architecture and / or a ConvNeXt architecture; a Head is used to receive feature maps passed from a backbone network or other layers and convert them into task outputs, and predict the output category labels, bounding boxes, and segmentation masks; a Neck is an intermediate layer connecting the backbone network and the Head; a decoder, conditioned on a prompt, is used to generate an output sequence for a single target detection task. During training, the prompt and the expected output are connected into a single sequence, and sequence weighting is used to ensure that the decoder is only trained to predict the expected output.

[0010] In some embodiments of the first aspect, the prompt is given and used to guide the model to process and generate input instructions containing different types of data, wherein the different types of data include images, text and voice.

[0011] In some embodiments of the first aspect, performing inverse quantization reasoning according to the task output sequence to determine the detection frame, mask, and defect description information of the input image includes:

[0012] For the target detection task, the predicted output sequence of the task is decomposed into multiple token tuples to obtain coordinate tokens and category tokens, and the coordinate tokens are dequantized to obtain the detection box; for the instance segmentation task, the coordinate token corresponding to each polygon is dequantized and converted into a dense mask, and the mask is mean-calculated to obtain a binary mask; for the image description task, the predicted category token is mapped to text to form defect description information.

[0013] In some embodiments of the first aspect, the method also includes: mapping features of different modalities to a preset feature domain through feature conversion by comparing an image sequence corresponding to the input image with a text sequence corresponding to the task description to form a fused feature; constructing a defect detection network based on the fused feature, and training the defect detection network with a multimodal data set consisting of the fused features represented by the input image and the task description, wherein the defect detection network is used to simultaneously perform target detection, instance segmentation and defect description; and using the trained defect detection network to detect the acquired image to be tested to output a detection frame, mask and defect description information of the image to be tested, wherein the image to be tested is the input image.

[0014] In the second aspect, the present application provides a multi-modal steel coil surface defect detection device, including: an acquisition module, used to acquire an input image containing a steel coil surface and a task description corresponding to the input image; a coding conversion module, used to perform coding conversion on the input image and the task description, and determine a discrete sequence, wherein the task descriptions of different task types use different coding methods; a task output module, used to input the discrete sequence into an encoder-decoder model to obtain a task output sequence; a defect detection module, used to perform inverse quantization reasoning based on the task output sequence, and determine the detection frame, mask and defect description information of the input image to complete the steel coil surface defect detection.

[0015] In a third aspect, the present application provides a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the above-mentioned multimodal-based steel coil surface defect detection method when executing the computer program.

[0016] In a fourth aspect, the present application provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the steps of the above-mentioned multi-modal steel coil surface defect detection method are implemented.

[0017] In the scheme implemented by the above-mentioned multimodal steel coil surface defect detection method, device, equipment and medium, the method obtains a sample pair formed by an input image containing the steel coil surface and a task description corresponding to the input image; encodes and converts the multimodal sample pair to determine a discrete sequence; inputs the discrete sequence into an encoder-decoder model to obtain a task output sequence, and the encoder-decoder model is a Pix2SeqV2 model; performs inverse quantization reasoning based on the task output sequence to determine the detection frame, mask and defect description information of the input image; on the one hand, the present application can not only improve the performance of the industrial image processing system, but also improve the automatic detection capability of steel coil surface defects; on the other hand, through the sample pairs formed by the input image and the task description, samples of different defect types can be quickly obtained, which is conducive to the rapid iteration of model training. At the same time, through a multimodal image set, the detection frame, mask and defect description information are directly output, which also improves the accuracy of steel coil surface defect detection and the scope of detection application. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] Figure 1 It is a schematic diagram of a process of a method for detecting surface defects of a steel coil based on multi-mode in one embodiment of the present invention;

[0019] Figure 2 1 is a schematic structural diagram of a multi-modal steel coil surface defect detection device according to an embodiment of the present invention;

[0020] Figure 3 is a schematic diagram of a processing structure of a multi-modal steel coil surface defect detection device in one embodiment of the present invention;

[0021] Figure 4 is a schematic diagram of a structure of a computer device in one embodiment of the present invention;

[0022] Figure 5 It is another structural schematic diagram of a computer device in one embodiment of the present invention. DETAILED DESCRIPTION

[0023] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0024] The present invention is described in detail below through specific embodiments. Figure 1 As shown, Figure 1 A schematic flow chart of a multi-modal steel coil surface defect detection method provided in an embodiment of the present invention includes the following steps:

[0025] Step S101, obtaining an input image including the surface of a steel coil and a task description corresponding to the input image;

[0026] Specifically, a high-resolution camera or image acquisition device is used to obtain a clear image of the steel coil surface. According to the detection requirements, a corresponding task description is generated for each input image, such as "detecting cracks on the surface of the steel coil" and "identifying rust on the surface of the steel coil".

[0027] The input image and the task description correspond one to one to form a sample training pair. For example, multiple images of the surface of a steel coil are collected at the same time to improve the detection efficiency.

[0028] Step S102, performing encoding conversion on the input image and the task description to determine a discrete sequence, wherein different encoding methods are used for task descriptions of different task types;

[0029] Specifically, the input image is preprocessed, such as normalization and denoising, to improve the image quality. According to the type of task description, the corresponding encoding method is selected to convert the image and task description into a discrete sequence. For example, for the target detection task, a specific encoding rule can be used to encode the target object into a discrete sequence. It should be noted that the tasks include but are not limited to target detection tasks, instance segmentation tasks, and image description tasks. Different tasks use different encoding methods for encoding conversion.

[0030] Step S103, inputting the discrete sequence into the encoder-decoder model to obtain the task output sequence, for example, the encoder-decoder model is a Pix2SeqV2 model;

[0031] Specifically, the Pix2SeqV2 model can convert the input discrete sequence into an output sequence, which contains key information such as detection boxes, masks, and defect description information. The encoder of the Pix2SeqV2 model is responsible for extracting image features, and the decoder is responsible for generating output sequences based on the features. For example, the Pix2SeqV2 model is selected as the encoder-decoder model, the discrete sequence is input into the model, and the task output sequence is obtained through the processing of the encoder and decoder.

[0032] In addition, this application can also adopt other advanced encoder-decoder models, such as Faster R-CNN, YOLO, etc., to adapt to different detection tasks.

[0033] Step S104, performing inverse quantization reasoning according to the task output sequence to determine the detection frame, mask and defect description information of the input image to complete the surface defect detection of the steel coil.

[0034] Specifically, the dequantization reasoning can convert the task output sequence back to the original numerical form, thereby extracting key information; the detection box is used to locate the defect position, the mask is used to identify the defect area, and the defect description information is used to describe the characteristics and type of the defect. For example, the task output sequence is dequantized and converted back to the original numerical form. According to the dequantized result, the detection box, mask and defect description information of the input image are determined.

[0035] Through the above methods, this application reduces manual intervention and time costs by automating the image acquisition and processing process; among them, the Pix2SeqV2 model and inverse quantization reasoning technology are used to accurately identify and describe defects on the surface of steel coils; at the same time, it supports multiple task types and encoding methods, and can adapt to different detection needs and scenarios; in addition, it also provides an intelligent solution for steel coil surface defect detection, which promotes technological progress and industrial upgrading in related industries.

[0036] In some embodiments, task descriptions of different task types use different encoding methods, including:

[0037] If the task description is a target detection task, the bounding box of the target is obtained by quantizing the continuous image coordinates in the input image, and the bounding box and target description are converted into a series of discrete sequences;

[0038] If the task is described as an instance segmentation task, the polygon corresponding to the instance mask in the input image is predicted in the form of an image coordinate sequence, and the coordinates corresponding to the polygon are quantized into a discrete sequence;

[0039] If the task description is an image description task, the task description is converted into a text sequence.

[0040] For example, if the task is described as a target detection task, a predefined quantization step is used to convert continuous coordinates in the image into discrete integer coordinates; the target objects in the image are traversed, and the bounding box coordinates of each object are quantized; based on the quantized coordinates, the bounding box of each target is determined. The bounding box is usually represented as a four-tuple (x_min, y_min, x_max, y_max), which represents the coordinates of the upper left corner and the lower right corner, respectively. The coordinates of the bounding box and the description of the target object (such as the category label) are converted into a series of discrete symbols or digital sequences; this is implemented through an encoding algorithm, such as using one-hot encoding or embedding vectors to represent discrete sequences.

[0041] For another example, if the task is described as an instance segmentation task, a deep learning model (such as yolov11seg) is used to predict the mask of each instance in the input image, and the mask is converted into a polygonal representation, usually through an approximation algorithm (such as the Douglas-Peucker algorithm); the vertex coordinates of the polygon are quantized, similar to the quantization process in the target detection task; and the quantized coordinate sequence is used as the output of the model.

[0042] For example, the image description task requires the generation of text descriptions related to the image content. Natural language processing technology is used to convert the text descriptions in the image description task into a series of discrete symbols or word sequences, including but not limited to preprocessing steps such as word segmentation, stop word removal, and stem extraction.

[0043] Through the above methods, the present application can handle different types of image tasks (object detection, instance segmentation, image description) and has strong adaptability. Through preprocessing steps such as quantization and polygonal approximation, the complexity of data processing is reduced and the processing efficiency is improved.

[0044] In some embodiments, the encoder-decoder model is determined by:

[0045] Obtain a data set consisting of multiple groups of sample pairs, each group of sample pairs includes an input image and a corresponding task description;

[0046] The dataset is divided into training set, test set and validation set. The training set is used to train the Pix2SeqV2 model, and the initial parameters are adjusted through back propagation to fit the first encoder-decoder model. The validation set is used to adjust the hyperparameters of the first encoder-decoder model to fit the second encoder-decoder model. The test set is used to evaluate the performance of the second encoder-decoder model, and the third encoder-decoder model with satisfactory generalization performance is output. The third encoder-decoder model is used as the final encoder-decoder model.

[0047] Specifically, a large amount of image data is collected, and for each image, a corresponding task description is provided, such as the specific requirements of target detection, image segmentation, etc. The image and the corresponding task description are combined into sample pairs to form a data set, which is randomly divided into three parts: training set, test set, and validation set. The usual division ratio is 70% of the data for training, 20% for testing, and 10% for validation.

[0048] Initialize the parameters of the Pix2SeqV2 model; input the images in the training set into the model and output the corresponding task description prediction results; calculate the loss between the prediction results and the actual task description, and adjust the model parameters through the back propagation algorithm to minimize the loss. For example, the Pix2SeqV2 model is based on the encoder-decoder architecture, which converts images into discrete labeled sequences for prediction. The back propagation algorithm updates the model parameters by calculating gradients to optimize the prediction results.

[0049] During the training process, the validation set is used regularly to evaluate the performance of the model. Based on the performance on the validation set, the model's hyperparameters (such as learning rate, batch size, etc.) are adjusted. The training process is repeated until the best hyperparameter combination is found. For example, the validation set is used to evaluate the performance of the model under different hyperparameters to select the best hyperparameter combination.

[0050] Use the trained model to make predictions on images in the test set. For example, the test set is used to evaluate the generalization ability of the model on unknown data. Based on the performance evaluation results on the test set, select the model with the best generalization performance. Save and deploy the model as the final encoder-decoder model. For example, select the best performing model to save and deploy to ensure accuracy in real applications.

[0051] Through the above methods, through reasonable data set division and training process, the training efficiency of the model and the accuracy of performance evaluation are ensured; the encoder-decoder architecture and back propagation algorithm of the Pix2SeqV2 model are used to successfully learn the mapping relationship between images and task descriptions; by adjusting hyperparameters and performance evaluation on the test set, the performance of the model is further optimized and the generalization ability of the model is improved.

[0052] The final output model has excellent generalization ability and can accurately meet the requirements of task description in practical applications, providing strong support for applications in the fields of image processing and computer vision.

[0053] In some embodiments, the encoder-decoder model includes an encoder based on a Transformer architecture and / or a ConvNeXt architecture; a Head is used to receive feature maps passed from a backbone network or other layers and convert them into task outputs, predicting output category labels, bounding boxes, and segmentation masks; a Neck is an intermediate layer connecting the backbone network and the Head; a decoder, conditioned on a prompt, is used to generate an output sequence for a single target detection task. During training, the prompt and the desired output are connected into a single sequence, and sequence weighting is used to ensure that the decoder is only trained to predict the desired output.

[0054] The encoder is based on the Transformer architecture or / and the ConvNeXt architecture. For the Transformer architecture, the encoder converts each word in the input sequence into a feature vector and generates an encoding output, which contains the semantic information of the input sequence. For the ConvNeXt architecture, the encoder captures the features of the spatial hierarchy through convolutional layers.

[0055] For example, the Transformer architecture uses the self-attention mechanism to capture long-distance dependencies in the input sequence and enhances the expressiveness of features through multi-layer stacked encoder blocks. The ConvNeXt architecture adopts the design concept of modern convolutional neural networks and gradually extracts features from images by stacking multiple convolutional blocks. The encoder can efficiently extract feature representations of input data (such as images, text, etc.), providing a solid foundation for subsequent task predictions.

[0056] For example, the Head is used to receive feature maps passed from the backbone network or other layers and convert them into task outputs. This includes predicting the output category labels, bounding boxes, and segmentation masks. The Head usually contains a series of convolutional layers, pooling layers, and fully connected layers, which map feature maps to task-related output spaces. For example, for classification tasks, the Head uses a softmax function to map features to category distributions; for detection tasks, the Head includes a classification head (for predicting target categories) and a regression head (for predicting the location of target boxes). The Head can accurately predict outputs based on feature maps, improving the prediction accuracy and robustness of the task.

[0057] For example, Neck is the middle layer between the backbone network and the Head, which is responsible for further processing or adjusting the features extracted by the backbone network. The structure of the Neck module is diverse, and the common ones are feature pyramid network and path aggregation network. These modules enhance the expressiveness of features by fusing feature maps of different scales, and improve the model's ability to detect objects of different scales. The Neck module can enhance the expressiveness of features and improve the detection accuracy and generalization ability of the model for objects of different scales.

[0058] The decoder is conditioned on a prompt and used to produce an output sequence for a single object detection task. During training, the prompt and the desired output are concatenated into a single sequence, and sequence weighting is used to ensure that the decoder is only trained to predict the desired output.

[0059] For example, the decoder is usually based on the Transformer architecture or similar structures, using self-attention mechanism and position encoding to generate output sequences. In the object detection task, the decoder receives the feature map and prompt from the encoder as input, and gradually generates output information such as the location and category of the object.

[0060] Adjust the structure and parameters of the decoder according to the task requirements and model complexity. For example, in the object detection task, a Transformer-based decoder can be used to generate output information such as bounding boxes and category labels.

[0061] In this embodiment, the decoder can accurately generate the output sequence of the target detection task based on the prompt and the feature map output by the encoder, improving the prediction accuracy and efficiency of the task. At the same time, through methods such as sequence weighting, it can ensure that the decoder only focuses on the desired output, further improving the performance of the model.

[0062] In some embodiments, a prompt is given to guide the model to process and generate input instructions containing different types of data, including images, text, and voice.

[0063] In some embodiments, performing inverse quantization reasoning based on the task output sequence to determine the detection frame, mask, and defect description information of the input image includes:

[0064] For the target detection task, the predicted task output sequence is decomposed into multiple token tuples to obtain coordinate tokens and category tokens, and the coordinate tokens are dequantized to obtain the detection box; for the instance segmentation task, the coordinate token corresponding to each polygon is dequantized and converted into a dense mask, and the mask is mean-calculated to obtain a binary mask; for the image description task, the predicted category token is mapped to text to form defect description information.

[0065] For example, the predicted task output sequence is usually generated by a neural network and contains the location and category information of the target. Through a specific decoding method, the output sequence is decomposed into multiple token tuples, each of which contains a coordinate token and a category token. The coordinate token is usually represented in a quantized form to reduce storage and transmission overhead. Through dequantization processing, the coordinate token is converted into an actual coordinate value to obtain the location of the detection box.

[0066] In instance segmentation tasks, targets are usually represented in the form of polygons. The coordinate tokens of each polygon are dequantized to obtain the actual coordinate values. A dense mask is generated based on the dequantized polygon coordinates. The mask is averaged, and pixels with mask values ​​greater than a threshold are set to 1, and the remaining pixels are set to 0, thereby obtaining a binary mask. A dense mask is generated based on the dequantized polygon coordinates. The mask is averaged, and pixels with mask values ​​greater than a threshold are set to 1, and the remaining pixels are set to 0, thereby obtaining a binary mask. In image description tasks, neural networks usually predict category tokens related to the image content. Defect description information is formed by mapping these category tokens to predefined text descriptions.

[0067] Through the above method, by adopting key technologies such as token tuple decomposition, inverse quantization processing, binary mask generation and text description generation, accurate detection and description of targets in images are achieved; it can be adjusted and optimized according to specific application scenarios and needs. This solution provides new ideas and methods for target detection, instance segmentation and image description tasks in the field of computer vision, which helps to promote the development and application of related technologies.

[0068] In some embodiments, the method also includes: mapping features of different modalities to a preset feature domain through feature conversion by comparing an image sequence corresponding to an input image with a text sequence corresponding to a task description to form a fused feature; constructing a defect detection network based on the fused feature, and training the defect detection network with a multimodal data set composed of the fused features represented by the input image and the task description, wherein the defect detection network is used to simultaneously perform target detection, instance segmentation and defect description; and using the trained defect detection network to detect the acquired image to be tested to output a detection frame, mask and defect description information of the image to be tested, wherein the image to be tested is the input image.

[0069] Exemplarily, the input image is first converted into an image sequence, which involves segmenting the image into small blocks or extracting key features from the image; at the same time, the task description is converted into a text sequence, including converting the natural language description into a format that can be processed by a computer, such as a word vector or an embedded representation. Then, the features of the image sequence and the text sequence are mapped to a preset feature domain using a feature conversion method (such as a convolutional neural network CNN, a recurrent neural network RNN, etc.) to form a fusion feature.

[0070] Exemplarily, a defect detection network is constructed based on fusion features. The network generally includes an input layer, a feature extraction layer, a classification layer, and an output layer. The input layer receives a multimodal data set, i.e., a data set consisting of an input image and fusion features characterized by the task description. Through the training process, the parameters of the defect detection network are optimized so that it can accurately identify defects in the image.

[0071] Exemplarily, the image to be tested is input into a trained defect detection network, and the network processes the input image and outputs a detection box, a mask, and defect description information; the detection box is used to locate the defect position in the image, the mask is used to represent the precise shape and range of the defect, and the defect description information provides a detailed text description of the defect.

[0072] In the above way, by fusing image and text features, full use is made of multimodal information, the accuracy and robustness of defect detection are improved, and the constructed defect detection network can simultaneously perform target detection, instance segmentation and defect description, realizing the integration and unification of multi-task processing. In addition, this solution can be widely used in industrial automation, intelligent manufacturing, quality inspection and other fields, providing strong technical support for defect detection and quality control.

[0073] Next, another specific embodiment is used to exemplify the multi-modal steel coil surface defect detection method provided in the above embodiment. The multi-modal steel coil surface defect detection method provided in this embodiment includes the following main steps:

[0074] 1) Collect the surface image of the steel coil;

[0075] 2) Convert the input image and task description into input tokens;

[0076] 3) Obtain the task output token through the pix2seqv2 encoder-decoder model;

[0077] 4) Infer the required detection box, mask and defect description information.

[0078] In 2), the conversion to input tokens is characterized in that both the image input and the task description are encoded into discrete token sequences, and different encoding methods are used for different visual tasks: for target detection tasks, this patent converts bounding boxes and object descriptions into a series of discrete tokens by quantizing continuous image coordinates; for instance segmentation tasks, this patent predicts the polygon corresponding to the instance mask in the form of an image coordinate sequence (i.e., taking a given polygon vertex as the starting point and determining the sequence of polygon vertices in a clockwise direction), rather than a pixel-by-pixel mask. Similarly, this patent quantizes the coordinates into discrete tokens. For image description tasks, this patent directly converts into text tokens.

[0079] In 3), the pix2seqv2 encoder is characterized in that the image encoder perceives pixels on the image and maps them to hidden representations, which can be specifically instantiated as ConvNeXt, Transformer, or a combination thereof. The Transformer-based sequence decoder is widely used in modern language modeling, which generates a token at a time, conditional on the previous token and the encoded image representation. This eliminates the complexity and customization of modern neural network architectures for these visual tasks (such as a head or neck specific to each task).

[0080] In 4), the pix2seqv2 decoder is characterized in that the decoder directly generates output tokens for a single object detection task conditioned on a prompt so that the model can generate outputs adapted to the task of interest. During training, the model concatenates the prompt and the desired output into a single sequence, using a token weighting scheme to ensure that the decoder is only trained to predict the desired output. During inference, the prompt is given and fixed, so the decoder only needs to generate the rest of the sequence. The training goal is to maximize the likelihood of the image-based token and the previous token, that is:

[0081]

[0082] Where X represents the input image and Y is a sequence of length L related to X. As mentioned earlier, the initial part of the sequence Y is a task description, for which the weight W is set to zero so that it is not included in the loss.

[0083] Furthermore, prompts are used to guide the model to process and generate input instructions containing different types of data (such as text, images, audio, etc.). By designing appropriate prompts, users can specify how the model associates and processes these different modal data to generate multimodal content.

[0084] In the reasoning process described, it is characterized in that the task output token is decoded, and different decoding methods are used for different visual tasks: for the target detection task, the predicted sequence is decomposed into a tuple with 5 tokens to obtain the coordinate token and the category token, and the coordinate token is dequantized to obtain the bounding box; for the instance segmentation task, the coordinate token corresponding to each polygon is dequantized and then converted into a dense mask. Since the model itself is not trained using any geometry-specific regularizer, the output polygon mask may be noisy. In order to reduce the noise, multiple sequences can be sampled, and then the mask is averaged and a binary mask is obtained by a simple threshold; for image description, the predicted discrete token is directly mapped to the text.

[0085] It can be seen that in the above scheme, a sample pair is formed by obtaining an input image including the surface of a steel coil and a task description corresponding to the input image; the multimodal sample pair is encoded and converted to determine a discrete sequence; the discrete sequence is input into the encoder-decoder model to obtain a task output sequence, and the encoder-decoder model is a Pix2SeqV2 model; inverse quantization reasoning is performed according to the task output sequence to determine the detection frame, mask and defect description information of the input image; on the one hand, the present application can not only improve the performance of the industrial image processing system, but also improve the automatic detection capability of steel coil surface defects; on the other hand, through the sample pairs formed by the input image and the task description, samples of different defect types can be quickly obtained, which is conducive to the rapid iteration of model training. At the same time, through a multimodal image set, the detection frame, mask and defect description information are directly output, which also improves the detection accuracy of steel coil surface defects and the scope of detection.

[0086] See Figure 3 , is a schematic diagram of a processing structure of a multi-modal steel coil surface defect detection device in one embodiment of the present invention; comprising:

[0087] The image information and text information are input separately, the image is divided into blocks and combined with the position encoding to be input into the encoder network; at the same time, the text information is encoded and combined with the image information in the encoder. Finally, the information of the encoder is passed through the decoder to obtain the predicted token value.

[0088] It should be understood that the order of execution of the steps in the above embodiment does not necessarily mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiment of the present invention.

[0089] In one embodiment, a multi-modal steel coil surface defect detection device is provided, and the multi-modal steel coil surface defect detection device corresponds to the multi-modal steel coil surface defect detection method in the above embodiment. Figure 2 As shown, the multi-modal steel coil surface defect detection device includes an acquisition module 201, a code conversion module 202, a task output module 203 and a defect detection module 204. The functional modules are described in detail as follows:

[0090] An acquisition module 201 is used to acquire an input image including a surface of a steel coil and a task description corresponding to the input image;

[0091] A coding conversion module 202 is used to perform coding conversion on the input image and the task description to determine a discrete sequence, wherein task descriptions of different task types use different coding methods;

[0092] The task output module 203 is used to input the discrete sequence into the encoder-decoder model to obtain the task output sequence;

[0093] The defect detection module 204 is used to perform inverse quantization reasoning according to the task output sequence, determine the detection frame, mask and defect description information of the input image, so as to complete the surface defect detection of the steel coil.

[0094] The embodiment of the present invention provides a steel coil surface defect detection device based on multimodality, which obtains a sample pair formed by an input image including a steel coil surface and a task description corresponding to the input image; encodes and converts the multimodal sample pair to determine a discrete sequence; inputs the discrete sequence into an encoder-decoder model to obtain a task output sequence, and the encoder-decoder model is a Pix2SeqV2 model; performs inverse quantization reasoning based on the task output sequence to determine a detection frame, a mask and defect description information of the input image; on the one hand, the present application can not only improve the performance of the industrial image processing system, but also improve the automatic detection capability of steel coil surface defects; on the other hand, through the sample pairs formed by the input image and the task description, samples of different defect types can be quickly obtained, which is conducive to the rapid iteration of model training. At the same time, through a multimodal image set, the detection frame, mask and defect description information are directly output, which also improves the detection accuracy of steel coil surface defects and the scope of application of detection.

[0095] For the specific definition of the multi-modal steel coil surface defect detection device, please refer to the definition of the multi-modal steel coil surface defect detection method above, which will not be repeated here. Each module in the above-mentioned multi-modal steel coil surface defect detection device can be implemented in whole or in part by software, hardware and their combination. The above-mentioned modules can be embedded in or independent of the processor in the computer device in the form of hardware, or can be stored in the memory of the computer device in the form of software, so that the processor can call and execute the operations corresponding to the above modules.

[0096] In one embodiment, a computer device is provided. The computer device may be a server, and its internal structure diagram may be as follows: Figure 4 As shown. The computer device includes a processor, a memory, a network interface and a database connected via a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile and / or volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external client via a network connection. When the computer program is executed by the processor, it realizes the functions or steps of a service-side method for detecting surface defects of a steel coil based on multimodality.

[0097] In one embodiment, a computer device is provided. The computer device may be a client, and its internal structure diagram may be as follows: Figure 5 As shown. The computer device includes a processor, a memory, a network interface, a display screen and an input device connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external server through a network connection. When the computer program is executed by the processor, the client-side functions or steps of a multi-modal steel coil surface defect detection method are implemented.

[0098] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the following steps when executing the computer program: obtaining an input image including a surface of a steel coil and a task description corresponding to the input image; encoding and converting the input image and the task description to determine a discrete sequence, wherein different encoding methods are used for task descriptions of different task types; inputting the discrete sequence into an encoder-decoder model to obtain a task output sequence, wherein the encoder-decoder model is a Pix2SeqV2 model; performing inverse quantization inference based on the task output sequence to determine a detection frame, a mask, and defect description information of the input image to complete surface defect detection of the steel coil.

[0099] The computer device provided in the above embodiment obtains a sample pair formed by an input image including the surface of a steel coil and a task description corresponding to the input image; encodes and converts the multimodal sample pairs to determine a discrete sequence; inputs the discrete sequence into an encoder-decoder model to obtain a task output sequence, and the encoder-decoder model is a Pix2SeqV2 model; performs inverse quantization reasoning based on the task output sequence to determine the detection frame, mask and defect description information of the input image; on the one hand, the present application can not only improve the performance of the industrial image processing system, but also improve the automatic detection capability of steel coil surface defects; on the other hand, through the sample pairs formed by the input image and the task description, samples of different defect types can be quickly obtained, which is conducive to the rapid iteration of model training. At the same time, through a multimodal image set, the detection frame, mask and defect description information are directly output, which also improves the detection accuracy of steel coil surface defects and the scope of detection.

[0100] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented: obtaining an input image including the surface of a steel coil and a task description corresponding to the input image; encoding and converting the input image and the task description to determine a discrete sequence, wherein different encoding methods are used for task descriptions of different task types; inputting the discrete sequence into an encoder-decoder model to obtain a task output sequence, wherein the encoder-decoder model is a Pix2SeqV2 model; performing inverse quantization reasoning based on the task output sequence to determine the detection box, mask and defect description information of the input image to complete the surface defect detection of the steel coil.

[0101] It should be noted that the above functions or steps that can be implemented by the computer-readable storage medium or computer device can refer to the relevant descriptions on the server side and the client side in the aforementioned method embodiment. To avoid repetition, they will not be described one by one here.

[0102] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiments can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in the present application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. As an illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).

[0103] Those skilled in the art can clearly understand that for the convenience and simplicity of description, only the division of the above-mentioned functional units and modules is used as an example. In actual applications, the above-mentioned functions can be distributed and completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.

[0104] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that the technical solutions described in the aforementioned embodiments may still be modified, or some of the technical features may be replaced by equivalents. Such modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention, and should be included in the protection scope of the present invention.

Claims

1. A multi-modal steel coil surface defect detection method, characterized in that: include: Obtaining an input image including a surface of a steel coil and a task description corresponding to the input image; Performing encoding conversion on the input image and the task description to determine a discrete sequence, wherein the task descriptions of different task types use different encoding methods; Inputting the discrete sequence into the encoder-decoder model to obtain a task output sequence; Dequantization reasoning is performed according to the task output sequence to determine the detection frame, mask and defect description information of the input image to complete the surface defect detection of the steel coil.

2. The multi-modal steel coil surface defect detection method according to claim 1, characterized in that: The task descriptions of different task types use different encoding methods, including: If the task description is a target detection task, a bounding box of the target is obtained by quantizing continuous image coordinates in the input image, and the bounding box and the target description are converted into a series of discrete sequences; If the task is described as an instance segmentation task, a polygon corresponding to the instance mask in the input image is predicted in the form of an image coordinate sequence, and the coordinates corresponding to the polygon are quantized into a discrete sequence; If the task description is an image description task, the task description is converted into a text sequence.

3. The multi-modal steel coil surface defect detection method according to claim 1, characterized in that: The method for determining the encoder-decoder model includes: Acquire a data set consisting of multiple groups of sample pairs, each group of the sample pairs including the input image and the corresponding task description; The data set is divided into a training set, a test set and a validation set. The encoder-decoder model is a Pix2SeqV2 model. The Pix2SeqV2 model is trained using the training set, and the initial parameters are adjusted by back-propagation to obtain a first encoder-decoder model. The validation set is used to adjust the hyperparameters of the first encoder-decoder model to obtain a second encoder-decoder model. The test set is used to evaluate the performance of the second encoder-decoder model, and a third encoder-decoder model with satisfactory generalization performance is output. The third encoder-decoder model is used as the final encoder-decoder model.

4. The multi-modal steel coil surface defect detection method according to any one of claims 1 to 3, characterized in that: The encoder-decoder model includes an encoder, which is based on the Transformer architecture and / or the ConvNeXt architecture; the Head is used to receive the feature map passed from the backbone network or other layers, and convert it into task output, predict the output category label, bounding box and segmentation mask; the Neck is the intermediate layer connecting the backbone network and the Head; the decoder, with a prompt as a condition, is used to generate an output sequence for a single target detection task. During training, the prompt and the expected output are connected into a single sequence, and sequence weighting is used to ensure that the decoder is only trained to predict the expected output.

5. The multi-modal steel coil surface defect detection method according to claim 4, characterized in that: The prompt is given and used to guide the model to process and generate input instructions containing different types of data, including images, text and voice.

6. The multi-modal steel coil surface defect detection method according to any one of claims 1 to 3, characterized in that: Dequantization reasoning is performed according to the task output sequence to determine the detection frame, mask and defect description information of the input image, including: For the target detection task, the predicted output sequence of the task is decomposed into multiple token tuples to obtain coordinate tokens and category tokens, and the coordinate tokens are dequantized to obtain the detection box; for the instance segmentation task, the coordinate token corresponding to each polygon is dequantized and converted into a dense mask, and the mask is mean-calculated to obtain a binary mask; for the image description task, the predicted category token is mapped to text to form defect description information.

7. The multi-modal steel coil surface defect detection method according to any one of claims 1 to 3, characterized in that: Also includes: The image sequence corresponding to the input image and the text sequence corresponding to the task description are mapped to a preset feature domain through feature conversion to form a fusion feature; Based on the fusion features, a defect detection network is constructed, and the defect detection network is trained using a multimodal data set consisting of the input image and the fusion features represented by the task description, wherein the defect detection network is used to simultaneously perform target detection, instance segmentation and defect description; the trained defect detection network is used to detect the acquired image to be tested, so as to output a detection frame, a mask and defect description information of the image to be tested, wherein the image to be tested is the input image.

8. A multi-modal steel coil surface defect detection device, characterized in that: include: An acquisition module, used to acquire an input image including a surface of a steel coil and a task description corresponding to the input image; A coding conversion module, used for performing coding conversion on the input image and the task description to determine a discrete sequence, wherein the task descriptions of different task types use different coding methods; A task output module, used for inputting the discrete sequence into an encoder-decoder model to obtain a task output sequence; The defect detection module is used to perform inverse quantization reasoning according to the task output sequence to determine the detection frame, mask and defect description information of the input image to complete the surface defect detection of the steel coil.

9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.

Citation Information

Cited By

  • Training-free multi-light-field substrate glass defect detection method and system based on visual language model

    CN121120632A

  • Multimodal-based steel coil surface defect detection method and apparatus, device, and medium

    WO2026179159A1