A cross-attention-based image-text modal scrap metal element detection system

By integrating image and text features, the cross-attention-based image-text modal scrap metal element detection system solves the problems of difficulty in distinguishing similar metals and insufficient element detection in existing technologies. It achieves efficient and accurate scrap steel classification and element composition detection, thereby improving smelting efficiency and steel quality.

CN119380873BActive Publication Date: 2025-12-30ANHUI UNIVERSITY OF TECHNOLOGY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411518015.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-10-29
Publication Date
2025-12-30
Estimated Expiration
2044-10-29

AI Technical Summary

Technical Problem

Existing machine vision-based scrap steel detection methods have difficulty distinguishing between metals with similar appearances and cannot detect the specific metal element composition, resulting in insufficient detection accuracy and reliability.

Method used

A cross-attention-based image-text modal scrap metal element detection system is adopted. It acquires images through high-definition camera equipment and writes detailed description text. The cross-attention mechanism is used to fuse image and text features to generate cross-modal high-dimensional feature vectors, which are then combined with pre-trained convolutional neural networks and language models for classification.

Benefits of technology

It significantly improves the accuracy of similar metal identification, can detect specific metal element composition, enhances the robustness and generalization ability of the model, improves smelting efficiency and steel quality, and reduces reliance on manual identification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119380873B_ABST
    Figure CN119380873B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of metal element detection, and discloses a picture-text modal scrap steel metal element detection system based on cross attention, which comprises an image acquisition module, a text acquisition module, a data preprocessing module, a data set division module, an image feature extraction module, a text feature extraction module, a cross attention mechanism module, a feature fusion module, a classifier module, a model evaluation and optimization module, a metal element database, a result output module and a system deployment and application module. By introducing the cross attention mechanism, visual and text features can be effectively fused, so that when the model processes metals with similar appearances, the model can rely on image features and utilize text features for differentiation at the same time, the recognition accuracy is significantly improved, and the system can fully utilize the information of the image modal and the text modal, so that the model can more comprehensively and deeply analyze the scrap steel characteristics, and the classification effect is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of metal element detection technology, specifically to a graphic modal metal element detection system for scrap steel based on cross-attention. Background Technology

[0002] The steel smelting industry faces challenges, including pressures in technology, cost, and management. Scrap steel, as a green and renewable resource, can replace iron ore in steelmaking, helping to reduce energy consumption, environmental pollution, and achieve green and low-carbon emissions.

[0003] Due to the large volume and variety of scrap steel used, improving steel production and smelting efficiency requires the classification and testing of scrap steel. An important part of scrap steel testing is the analysis of metal elements in scrap steel. The presence of impurities in different scrap steels can affect the performance of steel. Traditional scrap steel metal element testing relies on manual identification, which has the disadvantages of being highly subjective and inefficient.

[0004] The rapid development of artificial intelligence technology has led to the rise of machine vision-based scrap steel inspection methods, demonstrating significant advantages. By installing high-definition imaging equipment at scrap metal unloading points to acquire high-resolution images and using deep learning models for analysis, inspection efficiency and accuracy can be effectively improved. However, existing machine vision-based scrap steel inspection methods still have limitations, as follows:

[0005] Existing machine vision-based scrap steel detection methods mainly rely on image features, which have limitations when dealing with metals with similar appearances. For example, titanium alloys and aluminum alloys have similar colors and are difficult to distinguish through image features. The insufficiency of single-modal information reduces the detection accuracy and reliability of the model. Moreover, most existing scrap steel detection methods can identify the type of metal, but cannot detect the specific metal element composition.

[0006] Therefore, those skilled in the art provide a cross-attention-based graphic modal scrap metal element detection system to solve the problems mentioned in the background art. Summary of the Invention

[0007] To address the shortcomings of existing technologies, this invention provides a cross-attention-based graphic modal scrap metal element detection system to solve the problems in the background technology.

[0008] To achieve the above objectives, the present invention is implemented through the following technical solution: a cross-attention-based image-text modal scrap metal element detection system, comprising an image acquisition module, a text acquisition module, a data preprocessing module, a dataset partitioning module, an image feature extraction module, a text feature extraction module, a cross-attention mechanism module, a feature fusion module, a classifier module, a model evaluation and optimization module, a metal element database, a result output module, and a system deployment and application module.

[0009] Preferably, the image acquisition module is equipped with a high-definition camera for capturing images of the scrap steel unloading site; the text acquisition module is responsible for writing detailed descriptive text for the images; the data preprocessing module includes image preprocessing and text preprocessing; the dataset partitioning module divides the dataset into training, validation, and test sets, and performs K-fold cross-validation; the image feature extraction module uses a pre-trained convolutional neural network to extract visual features from the images; the text feature extraction module uses a pre-trained language model to encode the text and generate text feature representations; the cross-attention mechanism module is responsible for calculating attention weights between image and text features; and the feature fusion module applies the weights calculated by the cross-attention mechanism to the image... The system uses image and text features to generate a unified cross-modal high-dimensional feature vector. The classifier module processes the cross-modal high-dimensional feature vector to generate a probability distribution of scrap steel metal types. The model evaluation and optimization module uses a validation set to evaluate and fine-tune the model performance, and also uses an independent test set to evaluate the model performance. The metal element database contains information on different metal types and their elemental compositions, and supports querying elemental composition based on metal type. The result output module determines the scrap steel type and queries the corresponding metal elemental composition based on the model output, and outputs the metal type and elemental composition information. The system deployment and application module transmits image and text descriptions to the trained scrap steel detection model and outputs the scrap steel type and metal element determination results.

[0010] Preferably, the data preprocessing module includes an image preprocessing unit and a text preprocessing unit; the image preprocessing unit is used to scale and normalize the image; the text preprocessing unit is used for word segmentation, removal of useless information, and standardization of text format;

[0011] The dataset partitioning module includes a training set partitioning unit, a test set partitioning unit, and a K-fold cross-validation unit; the K-fold cross-validation unit is used to perform K-fold cross-validation on the training set.

[0012] Preferably, the model evaluation and optimization module includes a performance evaluation unit, a model tuning unit, and a final testing unit; the performance evaluation unit uses a validation set to evaluate model performance; the model tuning unit performs model tuning based on the evaluation results; and the final testing unit uses an independent test set to evaluate the final model performance.

[0013] The metal element database includes a data collection unit, a data entry unit, and a query unit. The data collection unit collects information on different metals and their elemental compositions from sources such as scientific literature, metal materials handbooks, and chemical databases. The data entry unit enters the collected metal types and their elemental compositions into the database. The query unit searches the database based on the predicted metal types to obtain the corresponding elemental composition information.

[0014] Preferably, the result output module includes a result generation unit, an element composition query unit, and an output unit; the result generation unit determines the type of scrap steel based on the model output results; the element composition query unit obtains the corresponding metal element composition information from the metal element database; and the output unit outputs the metal type and element composition information.

[0015] The system deployment and application module includes an image and text input unit and a result return unit. The image and text input unit inputs image and text descriptions into the trained scrap steel detection model. The result return unit returns the determination results of scrap steel type and metal element based on the model output results.

[0016] Preferably, the image feature extraction module uses a pre-trained convolutional neural network to extract image features. Specifically, the formulas for convolution operations, pooling operations, and fully connected layers are as follows:

[0017] ,in, It is the value at position (i,j) of the k-th feature map. It is the input image. It is a convolution kernel. It's a bias. It is an activation function;

[0018] Pooling operations: ,in, The input image is a pooling window of size m*n.

[0019] Preferably, the cross-attention mechanism module is used to calculate attention weights between image and text features to achieve feature alignment and fusion between modalities. The attention mechanism formula is as follows:

[0020] ,in, It is a query matrix. It is a key matrix. It is a value matrix. It is the dimension of the key vector.

[0021] Preferably, the feature fusion module applies the weights calculated by the cross-attention mechanism to image and text features to generate a unified cross-modal high-dimensional feature vector, with the following weighted fusion formula:

[0022] ,in, and It is attention weight. and These are image and text features.

[0023] Preferably, the classifier module processes the cross-modal high-dimensional feature vectors to generate probability distributions for different scrap metal types, and the classifier formula is as follows:

[0024] ,in, and It is a weight matrix. and It is a bias vector. It is the fused feature vector.

[0025] This invention provides a graph-and-text modal metal element detection system for scrap steel based on cross-attention. It offers the following advantages:

[0026] 1. By introducing a cross-attention mechanism, this invention can effectively integrate visual and textual features, enabling the model to distinguish between metals with similar appearances by relying on image features while also utilizing textual features, thus significantly improving recognition accuracy. Furthermore, the system can fully utilize information from both image and textual modalities, allowing the model to analyze the characteristics of scrap steel more comprehensively and deeply, thereby improving classification performance.

[0027] 2. This invention overcomes the problem of insufficient information in a single modality by introducing text features and cross-attention mechanism, thereby enhancing the robustness and generalization ability of the model. While identifying the type of scrap steel, it can also detect the specific metal element composition. By establishing a metal element database, the system can query the corresponding element composition information based on the predicted metal type, providing an important reference for effectively removing harmful elements and improving steel production and quality during the smelting process.

[0028] 3. By accurately detecting the type and specific elemental composition of scrap steel, the system can help the smelting process better control the raw material composition, effectively remove elements harmful to steel, improve smelting efficiency and steel quality, and utilize artificial intelligence technology to automate the detection and classification of scrap steel, reducing reliance on manual identification, improving detection efficiency and consistency, and reducing the risk of misjudgment caused by subjective factors.

[0029] 4. Through the consistency and complementarity of multimodal data, the system can provide more accurate classification and element detection results when facing scrap steel of different types, states and uses. It is highly adaptable and can operate stably in complex and ever-changing real-world application scenarios. Furthermore, by continuously collecting and accumulating scrap steel image and text data, the system can continuously optimize model performance and gradually improve the accuracy and efficiency of scrap steel detection by utilizing advanced deep learning and cross-attention mechanisms. Attached Figure Description

[0030] Figure 1 This is a flowchart of the present invention;

[0031] Figure 2 This is a structural diagram of the present invention;

[0032] Figure 3 This is a schematic diagram of the data preprocessing module of the present invention;

[0033] Figure 4 This is a schematic diagram of the model evaluation and optimization module of the present invention;

[0034] Figure 5 This is a schematic diagram of the metal element database of the present invention;

[0035] Figure 6 This is a schematic diagram of the result output module of the present invention;

[0036] Figure 7 This is a schematic diagram illustrating the system deployment and application of the present invention;

[0037] Figure 8 This is a schematic diagram of the model structure of the present invention;

[0038] Figure 9 This is a schematic diagram of the model structure unit of the present invention;

[0039] Figure 10 This is a schematic diagram of the connection of the model structure units of the present invention. Detailed Implementation

[0040] To enable those skilled in the art to understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. Other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort should fall within the scope of protection of the present invention.

[0041] The present invention will now be described in detail with reference to the accompanying drawings:

[0042] Example:

[0043] Please see the appendix Figure 1- Appendix Figure 6 This invention provides a cross-attention-based image-text modal scrap metal element detection system, comprising an image acquisition module, a text acquisition module, a data preprocessing module, a dataset partitioning module, an image feature extraction module, a text feature extraction module, a cross-attention mechanism module, a feature fusion module, a classifier module, a model evaluation and optimization module, a metal element database, a result output module, and a system deployment and application module. The image acquisition module is equipped with a high-definition camera for acquiring images of the scrap unloading site. The text acquisition module is responsible for writing detailed descriptive text for the images. The data preprocessing module includes image preprocessing and text preprocessing. The dataset partitioning module divides the dataset into training, validation, and test sets and performs K-fold cross-validation. The image feature extraction module uses a pre-trained convolutional neural network to extract visual features from the images. The text feature extraction module uses a pre-trained language model to analyze the text. The system encodes and generates text feature representations; the cross-attention mechanism module is responsible for calculating attention weights between image and text features; the feature fusion module applies the weights calculated by the cross-attention mechanism to image and text features to generate a unified cross-modal high-dimensional feature vector; the classifier module processes the cross-modal high-dimensional feature vector to generate a probability distribution of scrap metal types; the model evaluation and optimization module uses a validation set to evaluate model performance and perform tuning, and uses an independent test set to evaluate model performance; the metal element database contains information on different metal types and their elemental composition, supporting querying elemental composition based on metal type; the result output module determines the scrap steel type and queries the corresponding metal elemental composition based on the model output, and outputs the metal type and elemental composition information; the system deployment and application module transmits image and text descriptions to the trained scrap steel detection model and outputs the scrap steel type and metal element determination results.

[0044] The image acquisition module benefits from high-definition camera equipment ensuring image quality and providing a reliable data foundation for subsequent feature extraction and recognition. The text acquisition module supplements insufficient image information, providing rich text features for multimodal fusion and helping to improve classification accuracy. The data preprocessing module standardizes data formats, improving the efficiency and effectiveness of model training and feature extraction. The dataset partitioning module rationally partitions data, ensuring the generalization ability and reliability of model training, avoiding overfitting, and improving model performance. The image feature extraction module effectively extracts local and global features of images, improving the understanding and recognition of scrap steel images. The text feature extraction module captures the semantic information of text, generating high-quality text feature vectors to support multimodal fusion. The cross-attention mechanism module bridges the gap between visual and textual information. The module bridges the semantic gap between state and language modalities, improving the fusion effect of multimodal data and enhancing the model's recognition accuracy. The feature fusion module integrates image and text information to generate more comprehensive feature representations, improving the accuracy of scrap steel classification. The classifier module achieves accurate classification of scrap steel types through a multilayer perceptron. The model evaluation and optimization module ensures the model's performance in practical applications, continuously improving its accuracy and robustness through tuning. The metal element database provides a comprehensive metal element information base, providing a basis for the detection and query of specific elemental compositions. The result output module provides an intuitive display of scrap steel detection results, facilitating user understanding and application. The system deployment and application module ensures the practical deployment and application of the system, enabling the detection model to operate effectively in real-world environments.

[0045] The data preprocessing module includes an image preprocessing unit and a text preprocessing unit. The image preprocessing unit is used for image scaling and normalization. The text preprocessing unit is used for word segmentation, removal of useless information, and standardization of text format. The dataset partitioning module includes a training set partitioning unit, a test set partitioning unit, and a K-fold cross-validation unit. The training set partitioning unit divides 80% of the dataset into the training set for model training. The test set partitioning unit divides 20% of the dataset into the test set for model evaluation. The K-fold cross-validation unit performs K-fold cross-validation on the training set. The model evaluation and optimization module includes a performance evaluation unit, a model tuning unit, and a final testing unit. The performance evaluation unit uses the validation set to evaluate model performance. The model tuning unit tunes the model based on the evaluation results. The final testing unit uses an independent test set to evaluate the final model performance. The metal element database includes a data collection unit, a data entry unit, and a query unit. The system comprises the following modules: a data collection unit, a metal materials handbook, and a chemical database, which collect information on different metals and their elemental composition; a data entry unit, which enters the collected metal types and their elemental composition information into the database; a query unit, which queries the database based on the predicted metal type to obtain the corresponding elemental composition information; a result output module, which includes a result generation unit, an elemental composition query unit, and an output unit; a result generation unit, which determines the scrap steel type based on the model output; an elemental composition query unit, which obtains the corresponding metal elemental composition information from the metal element database; and an output unit, which outputs the metal type and elemental composition information; and a system deployment and application module, which includes an image and text input unit and a result return unit; an image and text input unit, which inputs image and text descriptions into the trained scrap steel detection model; and a result return unit, which returns the determination results of the scrap steel type and metal elements based on the model output.

[0046] The image preprocessing unit benefits from standardizing image size and pixel values, improving consistency and efficiency when inputting images into the model, and enhancing the model's learning effect. The text preprocessing unit improves the quality and consistency of text data, ensuring the model can correctly understand and process text features. The training set partitioning unit provides sufficient training data, improving the model's learning ability and accuracy. The test set partitioning unit provides an independent dataset for evaluating the model's performance in real-world applications, ensuring the model's generalization ability. K-fold cross-validation reduces overfitting through multiple training and validation iterations, improving the model's robustness and generalization ability. The performance evaluation unit monitors the model's performance during training, promptly identifying and resolving problems to ensure model stability and performance. The model tuning unit optimizes model parameters, improving model accuracy and performance. The final testing unit comprehensively evaluates the model's final performance. The system ensures the reliability and effectiveness of the model in a real-world environment. The data collection unit provides comprehensive and authoritative data sources, ensuring the accuracy and completeness of the database information. The data entry unit establishes and maintains a metal element database, providing data support for subsequent queries and testing. The query unit quickly and accurately provides the required elemental composition information, supporting scrap steel detection and classification. The result generation unit provides intuitive classification results, facilitating user understanding and application. The elemental composition query unit supplements scrap steel type information, providing detailed elemental composition information to support subsequent smelting and processing. The output unit provides a complete test report for user reference and decision-making. The image and text input units ensure smooth data input into the model, initiating the scrap steel detection process. The result return unit provides real-time test results, supporting rapid decision-making and processing, improving the efficiency of scrap steel classification and processing.

[0047] The image feature extraction module uses a pre-trained convolutional neural network to extract image features. Specifically, the formulas for convolution operations, pooling operations, and fully connected layers are as follows:

[0048] ,in, It is the value at position (i,j) of the k-th feature map. It is the input image. It is a convolution kernel. It's a bias. It is an activation function;

[0049] Pooling operations: ,in, The input image is a pooling window of size m*n.

[0050] The cross-attention mechanism module is used to calculate attention weights between image and text features, achieving feature alignment and fusion between modalities. The attention mechanism formula is as follows:

[0051] ,in, It is a query matrix. It is a key matrix. It is a value matrix. It is the dimension of the key vector.

[0052] The feature fusion module applies the weights calculated by the cross-attention mechanism to image and text features, generating a unified cross-modal high-dimensional feature vector. Its weighted fusion formula is as follows:

[0053] ,in, and It is attention weight. and These are image and text features.

[0054] The classifier module processes cross-modal high-dimensional feature vectors to generate probability distributions for different scrap metal types, and the classifier formula is as follows:

[0055] ,in, and It is a weight matrix. and It is a bias vector. It is the fused feature vector.

[0056] Please see the appendix Figure 7 and attached Figure 8 The specific model architecture of the visual language task model based on cross-attention is as follows: The model includes an image input embedding layer, a text input embedding layer, a visual encoder, a language encoder, and a cross-modal encoder;

[0057] The model has images and text: pictures and words. The input embedding layer converts images and text into word-level sentence embeddings and target-level graph embeddings.

[0058] The text input embedding layer consists of a tokenizer, a position encoder, a word encoder, and a normalization layer;

[0059] After the statement is input into the text input embedding layer, it is first divided into n words by the tokenizer: The word encoder converts the segmented words into fixed-length vectors, using the following formula:

[0060] The position encoder obtains the absolute position of each word in the sentence and converts it into a fixed-length vector. Then, the vectors obtained from position encoding and word encoding are added together and normalized. ;

[0061] The image input embedding layer consists of an object detector and a feature extractor;

[0062] Object detectors are typically pre-trained object detection models that can identify objects of most categories. They detect targets in an image. Each target is represented as a location feature. and ROL features The target represents the following output after the fully connected layer and the normalized layer:

[0063] , ,

[0064] ,

[0065] The model's visual encoder and language encoder are both made by It consists of stacked single-mode encoders, each of which, as shown in the figure, includes the following: self-attention layer, residual connection, normalization layer, and feedforward layer.

[0066] For a visual encoder, the input is For a language encoder, the input is In the self-attention layer, attention weights are calculated between features to capture the relationships between them. In the feedforward layer, the output of the attention layer undergoes a non-linear transformation to extract high-level features. In the normalization layer, a residual connection is established between the output of the sub-layer and the input that has not passed through the sub-layer.

[0067] ;

[0068] Cross-modal encoders are made by It is composed of stacked bit structures, that is, using the first bit structure. The output of the layer is used as the first The input to the layer, a separate cross-modal encoder, contains a bidirectional cross-attention layer, a self-attention layer, and a feedforward layer, as shown in the appendix. Figure 9 As shown;

[0069] The core of a cross-modal encoder is the cross-attention layer, which transitions between language and vision, and vice versa. The attention mechanism searches for information related to the query vector within the context vector. The (k-1)th layer outputs both the query vector and the context vector (e.g., the context vector). Query vector ), enter into the first The cross-attention layer of the layer is obtained as follows: ,

[0070] ,

[0071] Cross-attention layers are used to exchange information and align modalities to learn cross-modal representations. To further establish the internal relationships between modalities, the output of the cross-attention layer is fed into the self-attention layer, which then obtains: ,

[0072] ,

[0073] Finally, the The output of the multi-modal encoder is generated by the feedforward layer, and residual connections and normalization layers are added after each sub-layer, similar to a single-modal encoder, as shown in the attached figure. Figure 10 As shown.

[0074] The cross-modal encoder has outputs: language, visual, and multimodal outputs. The language and visual outputs are generated by the cross-modal encoder. A special tag is added before the word segmentation operation. The corresponding feature vector of this special tag in the language feature sequence is used as the cross-modal output. Finally, the text output of the cross-modal encoder will obtain the metal type.

[0075] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.

Claims

1. A cross-attention-based image-text modal scrap metal element detection system, characterized in that, The system comprises an image acquisition module, a text acquisition module, a data preprocessing module, a data set division module, an image feature extraction module, a text feature extraction module, a cross-attention mechanism module, a feature fusion module, a classifier module, a model evaluation and optimization module, a metal element database, a result output module, and a system deployment and application module. The image acquisition module is equipped with a high-definition camera device for collecting on-site images of unloaded scrap steel materials; the text acquisition module is responsible for writing description texts for the images; the data preprocessing module includes image preprocessing and text preprocessing; the data set division module divides the data set into a training set, a validation set, and a test set, and performs K-fold cross-validation on the training set; the image feature extraction module uses a pre-trained convolutional neural network to extract visual features of the images; the text feature extraction module uses a pre-trained language model to encode the texts to generate text feature representations; the cross-attention mechanism module is responsible for calculating attention weights between image features and text features; the feature fusion module applies the weights calculated by the cross-attention mechanism to the image features and the text features to generate a unified cross-modal high-dimensional feature vector; the classifier module processes the cross-modal high-dimensional feature vector to generate a probability distribution of the scrap steel metal type; the model evaluation and optimization module evaluates the model performance using the validation set and performs tuning, and evaluates the model performance using an independent test set; the metal element database contains information of different metal types and their element compositions, and supports querying element compositions according to metal types; the result output module determines the scrap steel type according to the model output result, queries the corresponding element composition, and outputs the information of the metal type and the element composition; The system deployment and application module transmits the image and text description to the trained scrap steel detection model to output the scrap steel type and metal element determination result; The image feature extraction module uses a pre-trained convolutional neural network to extract image features, and the formula is as follows: wherein, is a value of the (i, j) position of the kth feature map in the feature map, is an input image, is a convolution kernel, is a bias, is an activation function; wherein, is the input image, and the size of the pooling window is m*n; The feature fusion module applies the weights calculated by the cross-attention mechanism to the image features and the text features, and generates a unified cross-modal high-dimensional feature vector through a weighted fusion formula, wherein the weighted fusion formula is as follows: wherein, and are attention weights, and are image features and text features; The classifier module processes the cross-modal high-dimensional feature vector to generate a probability distribution of the scrap steel metal type, wherein the classifier formula is as follows: wherein, and are a first weight matrix and a second weight matrix, respectively, and are a first bias vector and a second bias vector, respectively, is the fused feature vector.

2. The cross-attention-based image-text modal scrap metal element detection system according to claim 1, wherein The data preprocessing module includes an image preprocessing unit and a text preprocessing unit; the image preprocessing unit is used to perform scaling and normalization processing on the image; The text preprocessing unit is used to perform word segmentation, remove useless information, and standardize the text format; The data set division module includes a training set division unit, a test set division unit, and a K-fold cross-validation unit; the K-fold cross-validation unit is used to perform K-fold cross-validation on the training set.

3. The cross-attention-based image-text modal scrap metal element detection system according to claim 1, wherein, The model evaluation and optimization module comprises a performance evaluation unit, a model tuning unit and a final test unit; the performance evaluation unit evaluates the model performance using a validation set; the model tuning unit tunes the model according to the evaluation result; and the final test unit evaluates the final model performance using an independent test set; The metal element database comprises a data collection unit, a data entry unit and a query unit; the data collection unit collects information of different metals and element compositions thereof through scientific literature, metal material manuals and chemical databases; the data entry unit enters the collected information of metal types and element compositions into the database; and the query unit queries the database according to the predicted metal type to obtain the corresponding element composition information.

4. The cross-attention-based image-text modal scrap metal element detection system according to claim 1, wherein, The result output module comprises a result generation unit, an element composition query unit and an output unit; the result generation unit determines the scrap steel type according to the model output result; The element composition query unit obtains the corresponding metal element composition information from the metal element database; and the output unit outputs the information of metal types and element compositions; The system deployment and application module comprises an image and text input unit and a result return unit; the image and text input unit inputs images and text descriptions into the trained scrap steel detection model; and the result return unit returns the determination results of scrap steel types and metal elements according to the model output result.

5. The cross-attention based image-text modal scrap metal element detection system according to claim 1, wherein, The cross-attention mechanism module is used for calculating attention weights between image and text features to realize feature alignment and fusion between modalities, wherein the attention mechanism formula is: wherein, is a query matrix, is a key matrix, is a value matrix, is the dimension of the key vector.

Citation Information

Patent Citations

  • Feature level fusion multi-mode-based rolling bearing fault detection method

    CN116821840A

  • Intelligent nursing decision support system based on deep learning

    CN118629620A