A method and system for accurate recognition of gear images using a multimodal model

By fusing image and text information through multimodal model, the problem of insufficient accuracy of gear recognition in complex environments of single-modal recognition method is solved, and high-precision gear image recognition and type recommendation are achieved.

CN118537705BActive Publication Date: 2025-07-11BEIHANG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410762047.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-06-13
Publication Date
2025-07-11
Estimated Expiration
2044-06-13

AI Technical Summary

Technical Problem

The existing single-modal image recognition methods are susceptible to factors such as light, pollution and wear in gear recognition, and it is difficult to achieve high-precision recognition in complex environments, especially for the classification and identification of high-complex gears.

Method used

Using a multimodal model, combined with the multimodal Transformer model of ResNet and Llama2, the visual and semantic features of the gear are extracted by fusion image processing and text information processing, the positive and negative sample pairs are generated using knowledge graphs, and the model is trained by comparative learning and mixed loss functions.

Benefits of technology

It significantly improves the accuracy of gear image recognition and the stability of the recognition system, and can accurately identify gear types and attributes in complex environments, improving the recognition efficiency and accuracy of industrial automation systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118537705B_ABST
    Figure CN118537705B_ABST
Patent Text Reader

Abstract

The present invention discloses a method and system for accurately identifying gear images using a multimodal model. This method realizes gear image recognition through a multimodal Transformer model that fuses ResNet and Llama2; the method includes the following steps: S100: Data collection and preprocessing; collecting image data and text data from various open-source part libraries and / or part standard files, S200: Dynamic sample pair selection strategy; S300: Establishing a multimodal model; the multimodal model includes an image processing branch and a text processing branch; the image processing branch uses a ResNet model for image embedding and model training; the text processing branch uses a Llama2 model to obtain embeddings of text descriptions and deep text learning; S400: Advanced fusion strategy; S500: Model training and evaluation; the present invention not only optimizes the gear image recognition process, but also improves the overall engineering efficiency and data security by integrating it into an industrial automation system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of computer vision and multimodal learning, and in particular to a method and system for accurately identifying gear images using a multimodal model. Background Art

[0002] In the current fields of industrial automation and computer vision, accurately identifying images of complex mechanical components, especially in the classification and recognition of high-complexity components such as gears, poses significant technical challenges. Traditional image recognition methods are single-modal methods that make judgments only through one type of recognition method, i.e., images. Because they are easily affected by factors such as lighting, contamination, and wear, they usually cannot cope with diverse industrial environments. In addition, when dealing with mechanical components with complex geometric shapes or high similarity, the recognition accuracy of single-modal recognition methods is often not ideal. Especially in the field of gear recognition, since it is difficult to accurately distinguish various types of gears only from images, there is an urgent need to develop a gear image recognition method with high-precision recognition capabilities. Summary of the Invention

[0003] Aiming at the problems existing in the prior art, the purpose of the present invention is to provide a method and system for accurately identifying gear images using a multimodal model, which is realized by a multimodal Transformer model that fuses ResNet and Llama2. This technology combines image processing and text information processing, and through a deep learning model, it realizes the fusion and modeling of multimodal features, and can effectively improve the recognition accuracy of gear images.

[0004] To achieve the above purpose, the present invention provides a method for accurately identifying gear images using a multimodal model. This method realizes gear image recognition through a multimodal Transformer model that fuses ResNet and Llama2; the method includes the following steps:

[0005] Step S100: Data collection and preprocessing; collect image data and text data from various open-source part libraries and part standard files, and process the image data and text data through normalization embedding and data augmentation methods;

[0006] Step S200: Dynamic sample pair selection strategy; utilize the knowledge and data stored in the knowledge graph to integrate positive and negative sample pairs, and apply the hard negative mining method to generate negative sample pairs;

[0007] Step S300: Establish a multimodal model; the multimodal model includes an image processing branch and a text processing branch. The image processing branch uses a ResNet model for image embedding and model training; the text processing branch uses a Llama2 model to obtain text embedding and perform deep text learning;

[0008] Step S400: Advanced fusion strategy; fuse the feature vectors of the image and the embedding vectors of the text by vector concatenation to form vector pairs, and then input the vector pairs into a multi-modal model for training;

[0009] Step S500: Model training and evaluation; adopt the method of contrastive learning, use contrastive loss and triplet loss to train the model to achieve gear image recognition; at the same time, evaluate the model performance on an independent test set.

[0010] Furthermore, step S100 includes:

[0011] Step S101: Collect gear image data; collect various types of gear images from various open-source part libraries and part standard documents, and store them in the knowledge graph;

[0012] Step S102: Use GPT-4V to recognize gear images; preliminarily recognize the collected gear images through GPT-4V and generate corresponding descriptive texts;

[0013] Step S103: Text quality inspection; check the descriptive texts generated by GPT-4V to determine whether they meet the standard requirements; if the texts do not meet the requirements, return and modify the prompt in the API to ensure that the generated descriptive texts are accurate and meet the requirements.

[0014] Furthermore, step S200 includes:

[0015] Step S201: Build a knowledge graph; establish a knowledge graph to store detailed attributes of gears and knowledge in the engineering field by integrating data from various part libraries and standard documents;

[0016] Step S202A: Generate positive sample pairs; generate positive sample pairs of correctly paired images and text descriptions according to the data in the knowledge graph for training the model;

[0017] Step S202B: Generate negative sample pairs; based on the hard negative mining method, generate negative sample pairs that meet the similarity requirements but have different labels from the positive sample pairs to enhance the challenge of model training, where step S202B is executed in parallel and synchronously with step S202A.

[0018] Furthermore, step S300 includes an image processing branch and a text processing branch; among them, the image processing branch includes:

[0019] Step S301A: Vectorize gear images using the ResNet model; vectorize the gear image data through a pre-trained ResNet model to generate image feature embeddings;

[0020] Step S302A: Extract image features; continue to use the pre-trained ResNet model to extract the gear image features after vectorization. The image features refer to the visual information crucial for the gear recognition task, specifically including: the edge and shape information of the gear (such as contour, tooth shape, size, and ratio), the texture and pattern on the gear surface (such as material texture and surface wear condition), the color and reflection features under light, and local features such as tooth gaps, keyway holes, and threaded holes;

[0021] The text processing branch includes:

[0022] Step S301B: Use the Llama2 model for text vectorization; input the description text into the Llama2 model for vectorization processing to generate text embeddings;

[0023] Step S302B: Extract text features; use the pre-trained Llama2 model to process the description text and extract the key features from the text embeddings. The key features include semantic information (such as the type, use of the gear, and the cooperation method with other mechanical components), syntactic structure (such as the construction of sentences and the use of syntactic elements), technical keywords and phrases (covering the proprietary terms and key descriptions of the gear), and dimensional data and process information (describing the specific dimensions and manufacturing processes of the gear).

[0024] Furthermore, step S400 includes:

[0025] Step S401: Fusion of vector features; fuse the image feature vector and the text embedding vector through vector concatenation to form a multimodal feature representation;

[0026] Step S402: Dimensional transformation of the fused vector; to ensure consistency, perform dimensional transformation on the fused vector to make it compatible with the subsequent model structure;

[0027] Step S403: Design a multimodal Transformer model; design a multimodal Transformer model and input the fused feature vector into it to combine the gear image and the gear description text for gear image recognition.

[0028] Furthermore, step S500 includes:

[0029] Step S501: Deliver positive and negative sample pairs to the model in a controlled variable manner; in a controlled variable manner, deliver positive and negative sample pairs to the model progressively in order and proportion;

[0030] Step S502: Design the learning ratio of positive and negative sample pairs; through experiments and model training experience, design an appropriate learning ratio of positive and negative sample pairs to ensure that the model learns between positive and negative sample pairs;

[0031] Step S503: Design a loss function to distinguish positive and negative sample pairs; adopt a contrastive loss function to distinguish positive and negative sample pairs.

[0032] Step S504: Recognition of gear images; use the trained model to achieve the accurate recognition task of gear images.

[0033] Furthermore, in step S202A, the specific operation steps are as follows: Extract text descriptions and corresponding images from the knowledge graph, and based on the attribute data and pairing relationships, determine the correct positive sample pairs; then, perform normalized embedding processing on these positive sample pairs. The text descriptions are vectorized through the Llama2 model, and the images are feature-extracted using the ResNet model to ensure the consistency of the two embedding dimensions; on this basis, use the vector concatenation method to fuse the image feature vectors and text embedding vectors to form the final representation of the positive sample pairs for use as a reference standard in subsequent model training.

[0034] Furthermore, in step S202B, the specific operation steps are as follows: When constructing negative sample pairs, first obtain the attribute data about gears from the knowledge graph and vectorize these data; by calculating the Euclidean distance of the feature vectors, determine sample pairs that meet the similarity requirements but have different labels; for two feature vectors χ1 and χ2, the calculation formula for the Euclidean distance is: where x1i and x2i are the values of the two feature vectors in the i-th dimension respectively, and n is the dimension of the feature vectors.

[0035] Furthermore, in step S403, the construction of the multimodal Transformer model includes:

[0036] First, define the self-attention module. Self-attention allows the model to allocate different attention weights among features at different positions in the sequence.

[0037] Subsequently, adopt the standard multi-head self-attention mechanism and focus on different subsets of representations through scaled dot-product attention calculation.

[0038] Then, construct a Transformer module with self-attention as the core; it contains 12 Transformer modules, and each Transformer module contains a multi-head self-attention mechanism and a feed-forward network.

[0039] In the model integration stage, form a deep network structure by stacking the said Transformer blocks.

[0040] Finally, construct the output layer; the output layer for the classification task is customized to identify the gear type, while for the tasks of gear attribute and parameter prediction, it is achieved by constructing an additional output layer.

[0041] On the other hand, the present invention provides a system for accurately identifying gear images using a multimodal model, and the system includes:

[0042] Data integration and optimization module: Collect gear image data and text data from multiple sources, and process the image data and text data through normalized embedding and data augmentation.

[0043] Dynamic sample scheduling module: Integrate positive and negative sample pairs through a knowledge graph, and apply the hard negative mining strategy to generate highly challenging negative sample pairs.

[0044] Multimodal architecture design module: Establish a framework with two branches: one branch is the image processing branch, which uses the ResNet model to extract gear image features; the other branch is the text processing branch, which uses the Llama2 model to vectorize text descriptions and further perform deep training.

[0045] Advanced feature fusion module: Fuse the image features and text embedding vectors through vector concatenation, and then input them into the multimodal Transformer model to achieve gear recognition.

[0046] Deep training and performance evaluation module: Train the model through a contrastive learning method, combining triplet loss and contrastive loss.

[0047] Beneficial effects:

[0048] The present invention proposes a new multimodal learning model for accurately identifying gear images by combining the vision-language model GPT-4V, the large language model Llama2, and the pre-trained convolutional neural network model ResNet. This model effectively fuses image data and text data, and through advanced feature extraction and multimodal fusion strategies, effectively improves the recognition accuracy for complex gear images. Description of the drawings

[0049] Figure 1 Shows a flowchart of the method for accurately identifying gear images according to the present invention;

[0050] Figure 2 Shows a schematic diagram of dynamically generating positive and negative sample pairs using the knowledge and data stored in the KG according to the present invention;

[0051] Figure 3 Shows a schematic diagram of advanced fusion of image and text feature vectors by the multimodal Transformer model according to the present invention. Detailed implementation manners

[0052] The technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0053] In the description of the present invention, it should be noted that the orientation or positional relationship indicated by the terms "center", "upper", "lower", "left", "right", "vertical", "horizontal", "inner", "outer", etc. is based on the orientation or positional relationship shown in the accompanying drawings. It is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be construed as a limitation to the present invention. In addition, the terms "first", "second", and "third" are only used for descriptive purposes and should not be construed as indicating or implying relative importance.

[0054] In the description of the present invention, it should be noted that unless otherwise clearly specified and limited, the terms "mounted", "connected", and "connected" should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be directly connected, or indirectly connected through an intermediate medium, and it can be the communication inside two elements. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.

[0055] The following is combined with Figures 1 - 3 The specific implementation manners of the present invention will be described in detail. It should be understood that the specific implementation manners described herein are only used to illustrate and explain the present invention, and are not used to limit the present invention.

[0056] The present invention provides a method and system for accurately identifying gear images using a multimodal Transformer model that integrates ResNet and Llama2, which is used for accurate identification and classification of gear types in industrial automation systems, and gear selection for specific engineering tasks.

[0057] The present invention demonstrates a multimodal Transformer model that combines ResNet and Llama2 technologies. This model is particularly suitable for industrial automation systems, and the specific implementation is as follows: First, use the ResNet model to extract key visual features from gear images, including but not limited to information such as the size, shape, and surface texture of the gears. The deep network structure of the ResNet model allows for effective analysis of these complex image data, thereby identifying the unique features of various gears.

[0058] Subsequently, the Llama2 model processes the gear-related descriptive text, which contains information about gear materials, load-bearing capacity, and applicable ranges. Through its advanced text embedding technology, Llama2 converts semantic information into vector forms that can be learned by the model, thus supplementing the data provided by image analysis.

[0059] The data from these two technologies are input into a multimodal Transformer model. The model efficiently integrates the feature vectors of images and text through an innovative fusion strategy. Through the self-attention mechanism, the model can assign different weights to different features of the gear, optimizing the information processing during recognition. Ultimately, the system can accurately identify the gear type, greatly improving the selection efficiency and accuracy.

[0060] A system for accurately identifying gear images using a multimodal model according to the present invention includes the following five functional modules:

[0061] Data Integration and Optimization Module: Responsible for collecting gear image and related text data from multiple sources, and processing these data through normalization embedding and data augmentation (such as rotation, scaling, cropping, text synonym replacement, etc.) to provide standardized input for subsequent steps.

[0062] Dynamic Sample Scheduling Module: Integrates positive and negative sample pairs through a knowledge graph, and applies the hard negative mining strategy to generate highly challenging negative sample pairs, thereby increasing the difficulty and effect of model training.

[0063] Multimodal Architecture Design Module: Establishes a framework with two branches: the image processing branch uses the ResNet model to extract gear image features, while the text processing branch uses the Llama2 model to vectorize text descriptions and further conducts in-depth training.

[0064] Advanced Feature Fusion Module: Fuses image features and text embedding vectors through vector concatenation, and then inputs them into a multimodal Transformer model to improve the accuracy and consistency of recognition.

[0065] Deep Training and Performance Evaluation Module: Trains the model through a contrastive learning method, combining triplet loss and contrastive loss. During the training process, the model alternately processes positive and negative sample pairs, and ensures the generalization ability and stability of the model through a strict evaluation process. Ultimately, the design of this multi-module system enables the multimodal learning model to have high accuracy in identifying gear images and can be used in practical application scenarios such as gear type recommendation and classification.

[0066] Such as Figure 1As shown, the method for accurately identifying gear images using a multimodal model according to the present invention is implemented through a multimodal Transformer model that fuses ResNet and Llama2. The method includes the following steps:

[0067] Step S100: Data collection and preprocessing; collect gear images and related text data from various sources, including various part libraries and / or part standard documents, and apply a series of standardized embedding and data augmentation techniques for processing. These techniques include image rotation, scaling, cropping, and synonym replacement for text, aiming to enhance the diversity and quality of the data, ensuring that in the subsequent multimodal learning process, the model can adapt to various different inputs and maintain high-precision recognition and classification performance. Specifically, it includes the following steps S101 - S103:

[0068] Step S101: Collect gear image data; collect various types of gear images from various open-source part libraries and part standard documents and store them in the knowledge graph.

[0069] Step S102: Use GPT-4V to identify gear images; preliminarily identify the collected gear images through GPT-4V and generate corresponding description texts.

[0070] Specifically, in step S102, use GPT-4V to identify gear images, and enhance the accuracy and comprehensiveness of gear text descriptions through a prompt strategy specially customized for industrial applications. Specifically, it includes:

[0071] 1. Refine the description requirements: Clearly specify the gear characteristics to be described in the prompt, such as material, size, type, usage conditions (new or old, wear condition), etc.;

[0072] 2. Add relevant context information: Add more context about the gear usage environment in the prompt, such as common mechanical applications or specific industry backgrounds, to guide the generation of more professional descriptions;

[0073] 3. Use structured prompts: When constructing the prompt, the present invention adopts a phased strategy to guide GPT-4V to deeply describe the physical characteristics of the gear, and then expand to its functions and uses in specific industrial applications, ensuring that the generated description is accurate and meets the actual industrial needs;

[0074] 4. Introduce a comparison and selection mechanism: Set the model to generate three different description texts at one time, and then select the most accurate and comprehensive answer through a scoring mechanism and manual screening.

[0075] Step S103: Text quality inspection; Check the descriptive text generated by GPT-4V to determine whether it meets the standard requirements. If the text does not meet the standard, return and modify the prompt in the API to ensure that the generated descriptive text is accurate and meets the requirements.

[0076] Step S103: Text quality inspection; Check the descriptive text generated by GPT-4V to determine whether it meets the standard requirements. The judgment criteria are divided into two stages: First, the system automatically compares the generated text with the gear types, attribute parameters, and dimensional data stored in the knowledge graph. This step ensures that the key numbers contained in the text match the verified and accurate information sources. Second, on the basis of ensuring the accuracy of the key information, a scoring mechanism and manual review are combined to examine the expression method and logical relationship of the text, and the descriptive text that can most accurately express the gear characteristics is selected. If the generated text does not meet the standard or is of too poor quality, modify the prompt in the API to ensure that the generated descriptive text is accurate and meets the requirements.

[0077] Step S200: The dynamic sample pair selection strategy constructs positive and negative sample pairs based on the knowledge and data stored in the knowledge graph. This strategy introduces the hard negative mining method into the sample generation process. By identifying and selecting negative samples with high similarity but different labels from the positive sample pairs, a set of difficult and challenging samples is formed. This method aims to increase the difficulty and diversity of model training, ensure that the training process has sufficient challenges, thereby improving the model's ability to distinguish between different samples, and ultimately enhancing the model's generalization performance and accuracy. It specifically includes the following steps S201, S202A, and S202B. Among them:

[0078] Step S201: Construct a knowledge graph; By integrating data from various part libraries and standard documents, the present invention establishes a knowledge graph to store detailed attributes of gears and knowledge in the engineering field. The knowledge graph contains image data of gears and also integrates text data corresponding to the images, including descriptions such as technical specifications, material types, and dimensional information.

[0079] Step S202A: Generate positive sample pairs; Extract text descriptions and corresponding images from the knowledge graph (KG). Based on the attribute data and pairing relationships, determine the correct positive sample pairs. Then, perform normalization embedding processing on these positive sample pairs. The text descriptions are vectorized through the LlaMa2 model, and the images are feature-extracted using the ResNet model to ensure the consistency of the two embedding dimensions. On this basis, use the vector splicing method to fuse the image feature vectors and text embedding vectors to form the final representation of the positive sample pairs for use as a reference standard in subsequent model training.

[0080] Step S202B: Generate negative sample pairs; as Figure 2 shown, based on the hard negative mining method, when constructing negative sample pairs, first obtain the attribute data of gears from the knowledge graph and generate the feature vectors of the attributes. The feature vector refers to converting each attribute parameter of the gear into a numerical representation, and these numerical values can be processed by computer programs. For example, the number of teeth can be directly represented as an integer, the material attribute can be represented by the index number of the material type, the module can be represented as a real number, and the accuracy grade can be encoded as an ordinal grade. These numericalized features constitute a point in a multi-dimensional space, that is, the feature vector. By calculating the Euclidean distance of the feature vectors, sample pairs with high similarity but different labels are determined. For two feature vectors χ1 and χ2, the calculation formula of the Euclidean distance is: where x1i and x2i are the values of the two feature vectors in the i-th dimension respectively, and n is the dimension of the feature vector. The smaller the Euclidean distance, the higher the similarity between the samples. For example, two gears may be very similar in appearance features such as module, number of teeth, gear shape, and tooth surface profile, but one has a special hole position processed while the other does not, and this belongs to a negative sample pair with a very high similarity. By comparing the feature vectors of these attributes, if their similarity reaches a certain threshold, negative sample pairs can be generated.

[0081] In this process, the key lies in identifying sample combinations with high similarity to the positive samples, but there are subtle differences in the key features between these samples. This way of generating negative sample pairs provides more challenging data for model training, ensuring that the model can correctly distinguish samples with different labels during the learning process and improving the generalization ability and accuracy of the model.

[0082] In this process, the key lies in identifying sample combinations with high similarity to the positive samples but with subtle differences in the key features. The key features are determined by referring to expert knowledge and industry practice experience. Among them, the key features include tooth shape, number of teeth, size, tooth surface processing type, profile accuracy, surface treatment, and other special processing details. The tooth shape includes straight teeth, helical teeth, and bevel teeth, etc.; the size includes specific measurements such as addendum circle diameter, dedendum circle diameter, pitch circle diameter, and tooth width; the tooth surface processing type includes no processing, polishing, and hardening treatment; the profile accuracy includes accuracy grade or tolerance range; the surface treatment includes plating and coating; the processing details include hub design, size and type of threaded holes and keyways. This way of generating negative sample pairs provides more challenging data for model training, ensuring that the model can correctly distinguish samples with different labels during the learning process and improving the generalization ability and accuracy of the model.

[0083] Step S300: Construct a multimodal model: The multimodal model framework includes an image processing branch and a text processing branch. Among them, the image processing branch uses the ResNet model to extract embedded features from gear images and trains a deep learning model on this basis; the text processing branch uses the Llama2 model to embed the text description to obtain a vectorized representation of the text features, and then conducts in-depth learning and training on the text data. Specifically, the image processing branch includes the following steps:

[0084] Step S301A: Vectorize the gear image using the ResNet model; Vectorize the gear image data through the pre-trained ResNet model to generate image feature embeddings.

[0085] Step S302A: Extract image features; Continue to use the pre-trained ResNet model to extract the gear image features after vectorization processing; The image features refer to the visual information for gear recognition tasks, including: the edge and shape information of the gear, the texture and pattern on the gear surface, the reflection characteristics under color and light, and the local features of tooth gaps, keyway holes, and threaded holes. Among them, the edge and shape information of the gear includes contour, tooth shape, size, and proportion; the texture and pattern on the gear surface include material texture and surface wear conditions.

[0086] The text processing branch includes the following steps:

[0087] Step S301B: Vectorize the text using the Llama2 model; Input the descriptive text into the Llama2 model for vectorization processing to generate text embeddings.

[0088] Step S302B: Extract text features; Use the pre-trained Llama2 model to process the descriptive text and extract the key features in the text embeddings. The "key features" here include: semantic information (such as the type, use, and cooperation method with other mechanical components of the gear), grammatical structure (such as the construction of sentences and the use of grammatical elements), technical keywords and phrases (covering the proprietary terms and key descriptions of gears), and dimensional data and process information (describing the specific dimensions and manufacturing processes of the gear).

[0089] Step S400: Advanced fusion strategy; As Figure 3 shown, through vector concatenation, integrate the image feature vector and the text embedding vector to form a multimodal feature representation. Then, input these fused vectors into a carefully designed multimodal Transformer model for in-depth training. This process aims to improve the recognition accuracy of the model for gear images and ensure the consistency of the recognition results, specifically including the following steps S401, S402, S403.

[0090] Step S401: Fusion of vector features; The image feature vector and the text embedding vector are fused through vector concatenation to form a multi-modal feature representation.

[0091] Step S402: Dimension transformation of the fused vector; To ensure consistency, the dimension of the fused vector is transformed to be compatible with the subsequent model structure.

[0092] Step S403: Design a multi-modal Transformer model; First, define the self-attention module, which is the basis of the Transformer architecture. Self-attention allows the model to allocate different attention weights among features at different positions in the sequence. Among them, features at different positions include: (1) Tooth profile area: including the number, shape, size, and clearance of teeth; (2) Gear edge treatment: including the edge machining of gears, such as chamfering or sharpening; (3) Surface texture: Different machining techniques (such as grinding, rolling) will leave different textures on the gear surface, and these textures may appear as subtle gray-scale or color changes in the image. In this process, the concepts of query (Q), key (K), and value (V) are introduced, and they are automatically adjusted through the learning process to assist the model in capturing complex relationships in the data.

[0093] Subsequently, adopt the standard multi-head self-attention mechanism, which enhances the model's attention ability by processing different subspace representations in parallel. Through scaled dot-product attention calculation, it is possible to focus on different subsets of representations, thereby promoting the integration and abstraction of information.

[0094] Then, construct a Transformer block with self-attention as the core. This design contains 12 such modules, each of which not only includes a multi-head self-attention mechanism but also a feed-forward network (FFN). To stabilize the training process and improve learning efficiency, layer normalization and residual connections are also embedded within each block. In addition, normalization layers are integrated when necessary to prevent overfitting.

[0095] In the model integration stage, a deep network structure is formed by stacking the above-defined Transformer blocks. Initially, it is recommended to start experiments with a model of 6 to 12 layers, and this scale can usually handle tasks of medium complexity. The preliminary evaluation of the model performance will determine whether it is necessary to increase the number of layers to capture more complex data structures.

[0096] Finally, the design of the output layer focuses on the specific requirements of the task. The output layer for the classification task is customized to identify the gear type, while for tasks related to gear attribute and parameter prediction, additional output layers are designed. This design ensures the adaptability and flexibility of the model in different tasks.

[0097] The design of the entire multi-modal Transformer model takes into account both efficiency and accuracy, thus providing powerful learning capabilities for the accurate recognition and type recommendation of gear images.

[0098] Step S500: Model training and evaluation; In the model training and evaluation stage, a contrastive learning strategy is adopted, and contrastive loss and triplet loss are used as the core training methods to strengthen the model's ability to distinguish between positive and negative sample pairs. During the training process, by introducing diverse positive and negative sample pairs, the learning effect of the model in a multi-modal environment is enhanced. After the model training is completed, an independent test set is used to comprehensively evaluate its performance, measuring the generalization ability and stability of the model in practical applications. Such evaluation ensures that the model can accurately identify gear images, including the following steps S501 - S504.

[0099] Step S501: Deliver positive and negative sample pairs to the model in a controlled variable manner; By means of controlling variables, the core is to establish a training sequence from simple to complex to ensure the model's progressive learning and generalization ability when identifying gear images. For positive samples, although the text description matches the image, they are not randomly delivered to the model, but are trained and learned in a progressive manner according to complexity. For example, among a series of gears, there may be three shapes: Shape A without a boss, Shape B with a boss, and Shape BK with a boss and threads. It is necessary to ensure that the selected gear sets are consistent in other attributes. During training, the simplest Shape A samples will be delivered first, then the medium-complexity Shape B samples, and finally the most complex Shape BK samples. The purpose is to let the model master the basics first and then gradually understand the differences in complex structures.

[0100] When setting the complexity sequence of gear images, it is mainly distinguished according to the design characteristics, machining accuracy, and functional complexity of the gears.

[0101] Simple gear images usually show basic gear shapes and rarely have additional processing, being suitable for general transmission applications without excessive performance customization requirements; Medium-complexity images contain one detailed feature among bosses, keyways, and threaded holes on the basic gear structure, being used in mechanical load and working condition occasions with higher performance requirements; While more complex gear images contain two or more detailed features among bosses, keyways, and threaded holes, and also contain multiple engineering characteristics and fine manufacturing processes, being suitable for special application scenarios with high performance and high load. Through the above sequence of progressive training and learning according to complexity, the progressive learning and generalization ability of the model when identifying gear images are ensured.

[0102] Negative sample pairs are samples where the text description and the image do not match. For negative samples, the model's delivery strategy also follows the idea of controlling variables. The difference is that for negative samples, the similarity between sample pairs is controlled. By delivering in descending order of similarity, the model can gradually understand the subtle differences between different features and learn to distinguish similar but mismatched samples.

[0103] Through such a method of controlling variables, the model can gradually accumulate knowledge during the training process, avoiding abrupt changes and overly complex inputs. This method helps the model learn at a stable pace, ensuring its sufficient adaptability when facing complex data, while maintaining the coherence and regularity of training.

[0104] Step S502: Design the learning ratio of positive and negative sample pairs; Through experiments and model training experience, design an appropriate learning ratio of positive and negative sample pairs to ensure that the model learns sufficient information between positive and negative sample pairs.

[0105] The present invention preferably maintains the ratio of positive and negative samples between 1:1 and 1:3. Setting the ratio of positive and negative samples within this range can help the model effectively identify and exclude incorrect gear types while maintaining sufficient recognition of positive samples, ensuring the overall performance of the model and the wide range of applications. This balance strategy helps avoid overfitting or learning bias, improving the stability and accuracy of the model in the actual industrial environment.

[0106] Step S503: Design a loss function to distinguish positive and negative sample pairs; The goal of designing the loss function is to precisely distinguish positive and negative sample pairs, thereby guiding the model to correctly separate them in the feature space. To achieve this goal, an effective strategy is to adopt a composite loss function that combines triplet loss and contrastive loss.

[0107] The triplet loss function focuses on the relative distances between an anchor sample, a positive sample, and a negative sample. Its purpose is to make the distance between samples of the same class closer than the distance between samples of different classes. In this context, the contrastive loss aims to directly push the sample towards its corresponding positive sample and away from the negative sample. Combining these two losses can provide stronger discrimination ability. Therefore, the designed loss function L can be expressed as:

[0108] L = α·L triplet + β·L contrastive ;

[0109] where L triplet is the triplet loss, L contrastive is the contrastive loss, and α and β are coefficients that adjust the weights of the two parts of the loss.

[0110] The formula for the triplet loss is:

[0111]

[0112] The formula for the contrastive loss is as follows:

[0113]

[0114] Here, f(x) represents the feature vector, a represents the anchor sample, p represents the positive sample, n represents the negative sample, and m is a preset boundary value used to define the minimum distance between positive and negative samples. y i is the label of the sample pair, and N is the number of samples.

[0115] This composite loss function allows the model to simultaneously learn the absolute and relative distances between samples in the multi-modal space, enhancing its ability to distinguish between positive and negative sample pairs. By adjusting the values of α and β in the experiment, the balance between the two losses can be achieved, thereby finding the optimal combination suitable for a specific dataset and model architecture.

[0116] Step S504: Gear evaluation and accurate identification of images; During the training of the multi-modal learning model, the adjustment of key parameters includes the batch size, learning rate, and the ratio of positive to negative samples. During training, the model adopts a multi-modal fusion strategy to extract features from image and text data respectively. After each training, the performance of the model is evaluated through an independent test set. The model evaluation metrics include precision, recall, F1-score, Normalized Discounted Cumulative Gain (NDCG), contrastive loss, and triplet loss to ensure the generalization ability and stability of the model. In addition, the method of cross-validation is adopted to measure the performance of the model with multiple evaluation criteria, and finally the optimal model configuration is determined. The calculation methods of each index are shown in Table 1:

[0117]

[0118] According to Table 1, the formula for precision is:

[0119]

[0120] where TP is the number of samples correctly classified in the test results; FP is the number of samples misclassified as positive. The formula for recall is

[0121]

[0122] Among them, FN is the number of actual positive samples omitted by the model.

[0123] The calculation formula of F1 score is:

[0124]

[0125] The calculation formula of normalized discounted cumulative gain (NDCG) is:

[0126]

[0127] Where k is the length of the recommendation list output by the model.

[0128] The formula for calculating contrastive loss is:

[0129]

[0130] Among them, m is the boundary value or threshold of the contrast loss. The goal of the loss function is to minimize the distance between positive sample pairs and to make the distance between negative sample pairs at least greater than the boundary value m.

[0131] The calculation formula for triplet loss is:

[0132]

[0133] Among them, m is used to control the threshold of the distance between positive and negative samples. The model needs to be adjusted to shorten the distance between the anchor point and the positive sample.

[0134] The best model determined after evaluation is used to accurately identify the gear images in the test set.

[0135] The following is an embodiment of the present invention.

[0136] Step S100: Data collection and preprocessing: Gear images and text data are collected from industrial libraries, engineering design documents, and standard documents. The text data sources are divided into two parts: one part comes directly from documents such as technical manuals, which record the physical properties of gears in detail; the other part is the description generated by GPT-4V after identifying the gear image, which not only covers the basic properties of gears, but also contains more extensive information such as application scenarios and functions.

[0137] The image is resized to a uniform size of 224x224 pixels during normalization, and then the image data is processed using the formula:

[0138] Among them, μ is the mean of the image data, and σ is the standard deviation. The diversity of image samples is increased by methods such as rotation, scaling, and cropping. The text part is subjected to synonym replacement and denoising processing to ensure data quality.

[0139] Step S200: Dynamic sample pair selection strategy. Utilize the attribute data in the knowledge graph to generate positive and negative sample pairs.

[0140] Use GPT-4V to perform preliminary recognition on the gear image and generate corresponding text descriptions.

[0141] Two types of text data are integrated and stored in the knowledge graph. The text descriptions generated by GPT-4V will be further subjected to text embedding through the Llama model for in-depth model training and analysis.

[0142] To ensure the accuracy of the text description, the following process is adopted:

[0143] First, use GPT-4V to generate preliminary text descriptions and conduct text quality checks. Judge whether the content and length of the text description meet the requirements, which are divided into two stages: First, the system automatically compares the description text generated by CPT-4V with the gear types, attribute parameters, and dimension data stored in the knowledge graph. This step ensures that the key numbers included in the text match the verified and accurate information sources. Second, on the basis of ensuring the accuracy of the key information, a combination of a scoring mechanism and manual review is used to examine the expression method and logical relationship of the text, and the description text that can most accurately express the gear characteristics is screened out. If the generated text does not meet the standards or is of too poor quality, modify the prompt in the API to ensure that the generated description text is accurate and meets the requirements. If not satisfied, return and modify the prompt of GPT-4V for regeneration. Ensure that the text description has no redundant noise information and retains the core content.

[0144] To ensure the accuracy of the text description, design a prompt to generate high-quality text descriptions.

[0145] To ensure the accuracy of the text description, a prompt strategy specifically customized for the gear field enhances the accuracy and comprehensiveness of the gear text description. Specific techniques include:

[0146] 1. Refine the description requirements: Clearly specify the gear characteristics to be described in the prompt, such as material, size, type, usage status (new or old, wear condition), etc.;

[0147] 2. Add relevant context information: Add more context about the gear usage environment in the prompt, such as common mechanical applications or specific industry backgrounds, to guide the generation of more professional descriptions;

[0148] 3. Use structured prompts: When constructing the prompt, we adopted a phased strategy to guide GPT-4V to deeply describe the physical characteristics of the gear, and then extended to its functions and uses in specific industrial applications, ensuring that the generated description is accurate and meets the actual industrial needs;

[0149] 4. Introduce a comparison and selection mechanism: Set the model to generate three different descriptive texts at once, and then select the most accurate and comprehensive answer through a scoring mechanism and manual screening. The following Table 2 shows the process of adjusting the prompt to generate text descriptions and the corresponding quality inspection results:

[0150]

[0151]

[0152] Generate positive and negative sample pairs using the attribute data in the knowledge graph.

[0153] Generation of positive sample pairs: First, obtain the attribute data of the gear from the knowledge graph, and then pair the image and text description. The positive sample pair consists of the correctly paired image and text description.

[0154] Generation of negative sample pairs: First, use the hard negative mining method to select samples with high similarity but different labels from the positive sample pairs. Then calculate the similarity using the Euclidean distance formula: Calculate the difference between negative sample pairs, and then sort the sample pairs according to the calculated distance to ensure that the selection of negative sample pairs can enhance the challenge of model training.

[0155] Step S300: Multimodal model design; Design a multimodal learning model, including two branches: image processing and text processing. The image processing branch uses the ResNet model to extract image features and normalize them. The text processing branch uses Llama2 to embed the text description and then perform deep training.

[0156] Step S400: Advanced fusion strategy; Concatenate the image features and text features into vectors, and then input them into the multimodal Transformer model for fusion. This step shows the fusion process through a formula:

[0157] V = [V image , V text ;

[0158] Among them, V is the fused multimodal feature vector, V image represents the feature vector extracted from the gear image, and V text represents the feature vector extracted from the relevant text description.

[0159] Step S500: Model training and evaluation. The model is trained using contrastive loss and triplet loss. The calculation formula of the loss function is:

[0160] L = α·L triplet + β·L contrastive ;

[0161] where L triplet is the triplet loss, L contrastive is the contrastive loss, and α and β are coefficients for adjusting the weights of the two parts of the loss.

[0162] The formula for the triplet loss is:

[0163]

[0164] The formula for the contrastive loss is:

[0165]

[0166] Here, f(x) represents the feature vector, a represents the anchor sample (selected from the dataset as the comparison reference point, usually highly representative and able to represent the typical features of a certain category or group. In the scenario of gear recognition, it means that the selected anchor sample should typically reflect the general attributes and features of this type of gear), p represents the positive sample pair, n represents the negative sample pair, m is a preset boundary value used to define the minimum distance between positive and negative samples. y i is the label of the sample pair, and N is the number of samples.

[0167] During the training process, an independent test set is used to evaluate the model performance, and accuracy, recall, F1-score, and Normalized Discounted Cumulative Gain (NDCG) are measured to determine the optimal model configuration.

[0168] After the model training is completed, it is applied to the industrial production line for gear image recognition. The key points of this step are:

[0169] 1. Use the multi-modal learning model for gear image recognition to ensure high accuracy.

[0170] 2. Use performance metrics to evaluate the model to ensure the generalization ability of the model in the actual environment.

[0171] The key points of the present invention are:

[0172] Multimodal Fusion Architecture Integrating Vision and Language Models: The present invention realizes the deep fusion of visual information and text information by integrating cutting-edge vision and language models, such as GPT-4V and Llama2, to create a multimodal Transformer model. This fusion strategy enriches the context dimension of the model's information processing and significantly enhances the model's interpretation and reasoning capabilities. The key innovation lies in that through this multimodal architecture, the model can integrate and utilize complex knowledge that cannot be provided independently by each modality, thereby greatly improving the recognition accuracy and processing efficiency of gear images. This architecture ensures that information from different data sources can be effectively fused and interacted within a unified framework, bringing unprecedented accuracy and depth to gear image recognition.

[0173] Refined Text Generation Enhances the Precision of Gear Recognition System: The present invention adopts a set of prompt strategies specifically designed for the gear field. This strategy is achieved through the following four core steps: First, clearly specify the gear characteristics to be described in the prompt, such as material, size, and type, etc., to ensure that the text information exactly matches the actual attributes of the gear. Second, add detailed context information about the gear's usage environment to guide the generation of more professional and specific descriptions. In addition, we adopt a phased strategy to construct the prompt, starting from the description of the gear's physical characteristics and gradually expanding to its functions and values in specific industrial applications. Finally, by generating multiple text options and using a scoring mechanism and manual review to select the optimal text description, this not only improves the quality of the text but also ensures the high accuracy and reliability of the model when processing professional and technical content. This innovative strategy significantly improves the performance and practicality of the entire system.

[0174] High-Quality Sample Generation and Progressive Training Strategy: The present invention utilizes a knowledge graph and a hard negative mining strategy to generate high-quality positive and negative sample pairs. This method ensures the diversity and sufficient challenge of the training data, providing rich learning materials for model training.

[0175] During the model training process, a progressive training strategy is adopted. By controlling the ratio of positive and negative sample pairs, more difficult negative sample pairs are gradually introduced. This strategy allows the model to contact positive sample pairs in the early stage to help the model establish basic feature recognition capabilities; in the later stage, the ratio of negative sample pairs is increased to strengthen the model's learning ability for complex samples. Through the progressive training method, it is ensured that the model can effectively distinguish positive and negative samples and improve the generalization ability.

[0176] Fine-grained Contrastive Learning and Hybrid Loss Function Design: The present invention optimizes the model training process through a fine-grained contrastive learning method and an innovative combination of contrastive loss and triplet loss. This hybrid loss function design is tailored to the characteristics of the multi-modal feature space and can accurately distinguish positive and negative sample pairs. The key innovation lies in that this loss function not only considers the direct differences between samples (contrastive loss), but also considers the relative positions of samples in the multi-modal space (triplet loss). Through this composite strategy, the boundary values and weight parameters of the model can be adjusted more effectively. This method enhances the model's sensitivity to subtle feature differences, improves the stability of training and the accuracy of the model, thus significantly enhancing the performance in practical applications.

[0177] According to the multi-modal gear recognition method of the present invention, compared with traditional single-modal recognition, the multi-modal model has stronger interpretation and processing capabilities. Under this technical framework, the system can simultaneously analyze information from images and text descriptions, thus better understanding the characteristics of complex components. The combination of deep learning and large language models provides higher accuracy and flexibility for the recognition of complex components.

[0178] The present invention provides effective fusion of data from different modalities, design of an adaptive feature extraction algorithm, and development of an efficient data preprocessing strategy, which has the technical advantage of high recognition accuracy in industrial environments, especially in the field of gear recognition. The development of this technology will promote the further improvement of industrial intelligence and bring higher efficiency and accuracy to the manufacturing industry.

[0179] Any process or method description in the flowchart of the present invention or described in other ways herein can be understood as representing a module, segment, or part of code including one or more executable instructions for implementing a specific logical function or process, which can be implemented on any computer scale medium for an instruction execution system, apparatus, or device. The computer-readable medium can be any medium including storage, communication, propagation, or transmission of a program for use by an instruction execution system, apparatus, or device. Including read-only memory, magnetic disks, or optical discs, etc.

[0180] In the description of this specification, the description referring to terms such as "embodiment", "example", etc. means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. In addition, those skilled in the art can combine or combine different embodiments or examples described in this specification and the features therein without contradiction.

[0181] Although the above content has shown and described embodiments of the present invention, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those of ordinary skill in the art can perform update operations such as changes, modifications, substitutions, and variations on the above embodiments within the scope of the present invention.

Claims

1. A method for precise recognition of gear images using a multimodal model, characterized in that, This method realizes gear image recognition through a multimodal Transformer model that fuses ResNet and Llama2; the method includes the following steps: Step S100: Data collection and preprocessing; collect image data and text data from various open-source part libraries and part standard files, and process the image data and text data through normalization embedding and data augmentation methods; Step S200: Dynamic sample pair selection strategy; Integrate positive and negative sample pairs using the knowledge and data stored in the knowledge graph, and apply the hard negative mining method to generate negative sample pairs; Step S300: Establish a multimodal model; the multimodal model includes an image processing branch and a text processing branch. The image processing branch uses the ResNet model for image embedding and model training; the text processing branch uses the Llama2 model to obtain text embeddings and perform in-depth text learning; Step S400: Advanced fusion strategy; fuse the feature vectors of the image and the embedding vectors of the text by vector concatenation to form a vector pair, and then input the vector pair into the multimodal model for training; Step S500: Model training and evaluation; adopt the method of contrastive learning, use contrastive loss and triplet loss to train the model to achieve gear image recognition; at the same time, evaluate the model performance on an independent test set; Step S200 includes: Step S201: Construct a knowledge graph; establish a knowledge graph to store detailed attributes of gears and knowledge in the engineering field by integrating data from various part libraries and standard files; Step S202A: Generate positive sample pairs; generate positive sample pairs of correctly paired images and text descriptions according to the data in the knowledge graph for training the model; Step S202B: Generate negative sample pairs; based on the hard negative mining method, generate negative sample pairs with a similarity to the positive sample pairs that meets the requirements but different labels to enhance the challenge of model training. Among them, Step S202B is executed in parallel and synchronously with Step S202A; In step S202B, the specific operation steps are as follows: when constructing negative sample pairs, first obtain the attribute data about gears from the knowledge graph and vectorize this data; determine sample pairs with different labels but meeting the similarity requirements by calculating the Euclidean distance of the feature vectors; for two feature vectors , the calculation formula for the Euclidean distance is: , where are the values of the two feature vectors in the th dimension respectively, is the dimension of the feature vector.

2. The method for accurate recognition of gear images according to claim 1, characterized in that Step S100 includes: Step S101: Collect gear image data; collect various types of gear images from various open-source part libraries and part standard files and store them in the knowledge graph; Step S102: Use GPT-4V to recognize gear images; preliminarily recognize the collected gear images through GPT-4V and generate corresponding description texts; Step S103: Text quality inspection; check the description texts generated by GPT-4V to determine whether they meet the standard requirements; if the texts do not meet the requirements, return and modify the prompt in the API to ensure that the generated description texts are accurate and meet the requirements.

3. The method for accurately identifying a gear image according to claim 1, characterized in that, Step S300 includes an image processing branch and a text processing branch; Among them, the image processing branch includes: Step S301A: Vectorize gear images using the ResNet model; vectorize the gear image data through a pre-trained ResNet model to generate image feature embeddings; Step S302A: Extract image features; continue to use the pre-trained ResNet model to extract the gear image features after vectorization; the image features refer to the visual information for gear recognition tasks, including: the edge and shape information of the gear, the texture and pattern on the gear surface, the reflection features under color and illumination, and the local features of tooth gaps, keyway holes, and threaded holes; The text processing branch includes: Step S301B: Perform text vectorization using the Llama2 model; input the description text into the Llama2 model for vectorization processing to generate text embeddings; Step S302B: Extract text features; use the pre-trained Llama2 model to process the description text and extract the key features in the text embeddings; the key features include semantic information, syntactic structure, technical keywords and phrases, as well as dimensional data and process information.

4. The method for accurately identifying a gear image according to claim 1, wherein Step S400 includes: Step S401: Fusion of vector features; fuse the image feature vector and the text embedding vector through vector concatenation to form a multi-modal feature representation; Step S402: Dimension transformation of the fused vector; to ensure consistency, perform dimension transformation on the fused vector to make it compatible with the subsequent model structure; Step S403: Design a multi-modal Transformer model; design a multi-modal Transformer model and input the fused feature vector into it to combine the gear image and the gear description text to achieve gear image recognition.

5. The method for accurately identifying a gear image according to claim 1, wherein Step S500 includes: Step S501: Deliver positive and negative sample pairs to the model in a controlled variable manner; in a controlled variable manner, deliver the positive and negative sample pairs to the model progressively in order and proportion; Step S502: Design the learning ratio of positive and negative sample pairs; through experiments and model training experience, design an appropriate learning ratio of positive and negative sample pairs to ensure that the model learns between positive and negative sample pairs; Step S503: Design a loss function to distinguish positive and negative sample pairs; adopt a contrastive loss function to distinguish positive and negative sample pairs; Step S504: Recognition of gear images; use the trained model to achieve the accurate recognition task of gear images.

6. The method for accurately identifying a gear image according to claim 1, wherein, In Step S202A, the specific operation steps are as follows: extract the text description and the corresponding image from the knowledge graph, and determine the correct positive sample pairs based on the attribute data and pairing relationships; then, perform normalized embedding processing on these positive sample pairs. The text description is vectorized through the Llama2 model, and the image is feature-extracted using the ResNet model to ensure the consistency of the two embedding dimensions; on this basis, use the vector concatenation method to fuse the image feature vector and the text embedding vector to form the final representation of the positive sample pair for use as a reference standard in subsequent model training.

7. The method for accurately identifying a gear image according to claim 4, wherein, In Step S403, the construction of the multi-modal Transformer model includes: First, define the self-attention module. Self-attention allows the model to allocate different attention weights among features at different positions in the sequence; Subsequently, a standard multi-head self-attention mechanism is adopted to focus on different subsets of representations through scaled dot-product attention calculation; Subsequently, a Transformer module with self-attention as the core is constructed; it contains 12 Transformer modules, and each Transformer module contains a multi-head self-attention mechanism and a feed-forward network In the model integration stage, a deep network structure is formed by stacking the Transformer blocks; Finally, an output layer is constructed; the output layer for the classification task is customized to identify the gear type, while for the tasks of gear attribute and parameter prediction, it is achieved by constructing an additional output layer.

8. A system for accurately identifying gear images using a multimodal model, characterized in that, The system is used to implement the method according to any one of claims 1-7, and the system includes: Data integration and optimization module: Collect gear image data and text data from multiple sources, and process the image data and text data through normalization embedding and data augmentation; Dynamic sample scheduling module: Integrate positive and negative sample pairs through a knowledge graph, and apply the hard negative mining strategy to generate highly challenging negative sample pairs; Multi-modal architecture design module: A framework with two branches is established: one branch is the image processing branch, which uses a ResNet model to extract gear image features; the other branch is the text processing branch, which uses the Llama2 model to vectorize text descriptions and further perform deep training; Advanced feature fusion module: Fuse the image features and text embedding vectors through vector concatenation, and then input them into a multi-modal Transformer model to achieve gear recognition; Deep training and performance evaluation module: Train the model through a contrastive learning method, combining triplet loss and contrastive loss.

Citation Information

Patent Citations

  • Open domain relation extraction method and system based on adaptive clustering

    CN116662457A

  • Multi-modal model training method, target object detection method and system

    CN117935019A