Text guidance-based mechanical tool image target detection method
By constructing a text object detection model and introducing the SPD module, EMA attention mechanism and SA spatial attention module, the problem of traditional object detection methods detecting mechanical tools in complex environments in the field of mechanical manufacturing is solved, and high-precision mechanical tool detection and lightweight object detection solutions are realized.
Patent Information
- Application Number
- CN202510193626.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-21
- Publication Date
- 2025-06-10
AI Technical Summary
Traditional object detection methods based on closed category sets are difficult to effectively detect mechanical tools in complex environments in the field of mechanical manufacturing, and existing text-guided object detection methods have shortcomings in the alignment of image areas and text fine-grained sizes.
A text-guided mechanical tool image object detection method is adopted. By constructing a text object detection model, the visual detection module, text guidance module, text visual fusion module and prediction module are used, combined with the SPD module, EMA attention mechanism and SA spatial attention module, to achieve high-precision detection of mechanical tools.
The model's processing and detection accuracy of low-resolution images and small and medium-sized target objects is improved, and the perception of specific target areas of text is enhanced, providing a lightweight object detection solution suitable for robot control platforms with limited computing resources.
Smart Images

Figure CN120125804A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of object detection, and particularly relates to a method for detecting mechanical tool images guided by text. Background Art
[0002] In the field of mechanical manufacturing, using vision to quickly and accurately detect tools in mechanical manufacturing is a key link in realizing mechanical manufacturing. Traditional object detection methods based on intelligent algorithms such as deep learning convolutional neural networks rely on predefined categories and labeled data, and can achieve high detection accuracy and speed in specific object detections with clear task requirements, which has been widely applied in application scenarios such as industrial component detection and service industry face recognition.
[0003] However, in the scenarios of the mechanical manufacturing field, the object detection environment is complex and changeable, and the types of detection objects are numerous, which makes it difficult for traditional object detection methods based on closed category sets to effectively meet the requirements for reliable detection of tools in the mechanical manufacturing field. In recent years, text-guided object detection methods have emerged, which describe the object through natural language and dynamically adjust the detection target, and have the ability to detect objects in complex and changeable environments, providing a new technical solution for reference in object detection in the mechanical manufacturing field. At present, a large number of scholars have conducted research.
[0004] Zareian et al. proposed the OVR-CNN model in the paper “Open-vocabulary object detection using captions, in Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition, pages: 14393-14402, 2021.” The model uses the pre-trained visual encoder of the image-text pair to achieve global alignment between the image and the text description to complete the object detection task. However, the OVR-CNN model mainly focuses on the alignment between the image as a whole and the text description, ignoring the fine-grained alignment between the image region and the text, which limits its performance in practical application scenarios. Carion et al. proposed the DETR model method in the paper “End-to-end object detection with transformers, in European conference on computer vision, pages: 213-229, 2020.” to achieve end-to-end object detection. This method uses the excellent sequence modeling ability of the attention mechanism to detect the target object, and directly generates the detection result through the encoder-decoder structure of the Transformer to improve the target positioning ability, but the algorithm complexity of this method is high and the real-time performance is poor. Summary of the invention
[0005] In order to solve the problems existing in the background technology, the purpose of the present invention is to provide a text-guided machine tool image target detection method.
[0006] The technical solution adopted by the present invention includes:
[0007] S1. Use a camera to capture pictures of mechanical tools, and create text information for each mechanical tool picture. Each mechanical tool picture and the corresponding text information form an image-text pair. The image-text pairs obtained from all mechanical tool pictures form a mechanical tool dataset.
[0008] S2. Build a text object detection model in a computer, input the machine tool data set into the text object detection model for training, and obtain a trained text object detection model.
[0009] S3. Obtain the image of the mechanical tool to be tested and the text information corresponding to the image of the mechanical tool to be tested respectively, input the image of the mechanical tool to be tested and the corresponding text information into the trained text target detection model for detection, and obtain the mechanical tool detection result.
[0010] The text target detection model includes a visual detection module, a text guidance module, a text visual fusion module and a prediction module; the mechanical tool pictures in the mechanical tool data set are input into the visual detection module for processing, the text information corresponding to each mechanical tool picture in the mechanical tool data set is input into the text guidance module for processing to obtain a first text embedding feature, the result of the processing by the visual detection module and the first text embedding feature are input into the text visual fusion module for processing together, and the result of the processing by the text visual fusion module is input into the prediction module for processing to obtain a mechanical tool detection result.
[0011] The visual detection module includes a convolutional layer one, a convolutional layer two, an SPD module one, a C2f module one, a convolutional layer three, an SPD module two, a C2f module two, a convolutional layer four, an SPD module three, a C2f module three, a convolutional layer five, a C2f module four, an SPD module four and an SPPF pooling module which are connected in series in sequence; the input end of the convolutional layer one serves as the input end of the visual detection module, and the output end of the C2f module two, the output end of the C2f module three and the output end of the SPPF pooling module serve as the output end of the visual detection module at the same time.
[0012] The convolution layer 1, the convolution layer 2, the convolution layer 3, the convolution layer 4 and the convolution layer 5 all adopt 3×3 convolution layers.
[0013] The text-guided module uses a contrastive language-image pre-training model.
[0014] The text visual fusion module includes a first spatial feature extraction module, a second spatial feature extraction module, a third spatial feature extraction module, a fourth spatial feature extraction module, a first image pooling attention mechanism module, a second image pooling attention mechanism module, a first visual text feature extraction attention module, a second visual text feature extraction attention module and a third visual text feature extraction attention module; the first spatial feature extraction module, the second spatial feature extraction module, the third spatial feature extraction module and the fourth spatial feature extraction module all have an image feature input terminal and a text feature input terminal.
[0015] The output result of the C2f module 2 is transmitted to the image feature input end of the first spatial feature extraction module for processing, the output result of the C2f module 3 is transmitted to the image feature input end of the second spatial feature extraction module for processing, the output result of the SPPF pooling module is transmitted to the input end of the first image pooling attention mechanism module for processing, the first text embedding feature is respectively input to the text feature input end of the second spatial feature extraction module and the text feature input end of the first spatial feature extraction module for processing, the output result of the first spatial feature extraction module and the output result of the second spatial feature extraction module are both transmitted to the input end of the first image pooling attention mechanism module for processing, the output result of the first image pooling attention mechanism module is subjected to the first residual connection with the first text embedding feature to obtain the second text embedding feature, the second text embedding feature is respectively input to the text feature input end of the third spatial feature extraction module and the text feature input end of the fourth spatial feature extraction module for processing, the first image pooling attention mechanism module The output results of the attention mechanism module are also transmitted to the image feature input end of the third spatial feature extraction module, the image feature input end of the fourth spatial feature extraction module and the input end of the first visual text feature extraction attention module for processing, the output result of the third spatial feature extraction module is input into the input end of the second visual text feature extraction attention module for processing, the output result of the fourth spatial feature extraction module is input into the input end of the third visual text feature extraction attention module for processing, the output results of the first visual text feature extraction attention module, the second visual text feature extraction attention module and the third visual text feature extraction attention module are all input into the input end of the second image pooling attention mechanism module for processing, the output result of the second image pooling attention mechanism module is residually connected with the second text embedding feature for the second time to obtain the image perception embedding feature, and the image perception embedding feature and the input result of the second image pooling attention mechanism module are used together as the output result of the text visual fusion module.
[0016] The first spatial feature extraction module, the second spatial feature extraction module, the third spatial feature extraction module and the fourth spatial feature extraction module all adopt the same structure, which specifically includes a bottleneck module, an SA spatial attention module and an activation module; the result of the previous layer input is first feature segmented to obtain a feature Figure 1 and Features Figure 2 , the characteristics Figure 1 The input is processed in the bottleneck module. The result of the bottleneck module is input into the SA spatial attention module for processing. The result of the SA spatial attention module and the result of the bottleneck module are processed for feature merging. The result of the feature merging is input into the activation module for processing. The result of the activation module and the feature Figure 2 Perform residual connection, and the result of residual connection is used as the output of the structure.
[0017] The bottleneck module is the Dark Bottleneck module and the activation module is Max-Sigmoid; the first image pooling attention mechanism module and the second image pooling attention mechanism module both adopt the I-PoolingAttention module in YOLO-World; the first visual text feature extraction attention module, the second visual text feature extraction attention module and the third visual text feature extraction attention module are all EMA attention mechanisms.
[0018] The prediction module includes a text comparison head and a box head; the result output by the second image pooling attention mechanism module is input into the box head to obtain a prediction box in the image; the result output by the second image pooling attention mechanism module is also input into the text comparison head together with the obtained prediction box to obtain a target embedding feature of each prediction box, the target embedding feature is matched with the image perception embedding feature to obtain a regional text matching result, and the regional text matching result is combined with the prediction box to obtain a final mechanical tool detection result.
[0019] The prediction box only frames the position of the target in the image; the regional text matching result is only the category information of the target in the prediction box. The regional text matching result and the prediction box are combined to obtain a detection result that contains both position information and category information.
[0020] The text comparison header adopts the text comparison header in YOLO-World; the frame header adopts the frame header in YOLO-World.
[0021] Compared with the prior art, the beneficial effects of the present invention are:
[0022] 1. The present invention introduces the SPD module into the visual detection module to improve the low-resolution image processing capability and the feature representation capability of small and medium-sized target objects, thereby improving the detection accuracy of the model for actual target objects.
[0023] 2. The present invention introduces the EMA attention mechanism in the text vision fusion module to improve the text-guided visual feature representation capability, and introduces the SA spatial attention module to enable the model to pay more attention to the text-related areas in the image, enhance the model's perception of the specific target areas described in the text, and improve the target detection accuracy.
[0024] 3. The method provided by the present invention adopts a lightweight text-guided target detection model, which can be deployed on a robot control platform with limited computing resources. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] Figure 1 It is the overall framework diagram of the method of the present invention.
[0026] Figure 2 It is a framework diagram of the visual detection module in the method of the present invention.
[0027] Figure 3 It is a framework diagram of the spatial feature extraction module in the method of the present invention. DETAILED DESCRIPTION
[0028] like Figure 1 As shown, this embodiment is implemented by using the following steps:
[0029] S1. Use a camera to capture pictures of mechanical tools, and create text information for each mechanical tool picture. Each mechanical tool picture and the corresponding text information form an image-text pair. The image-text pairs obtained from all mechanical tool pictures form a mechanical tool dataset.
[0030] The text information is information containing the target to be detected in the machine tool image, which may be a sentence or several words.
[0031] In a specific implementation, the established text information may be "the picture contains scissors, a cross screwdriver and pliers", and the picture and the text information are combined to form an image-text pair.
[0032] S2. constructing a text object detection model in a computer, inputting the machine tool data set into the text object detection model for training, and obtaining a trained text object detection model;
[0033] The text information is text information of an object to be detected in the image of the machine tool to be detected.
[0034] The text target detection model includes a visual detection module, a text guidance module, a text visual fusion module and a prediction module; the mechanical tool pictures in the mechanical tool data set are input into the visual detection module for processing, the text information corresponding to the mechanical tool pictures in the mechanical tool data set is input into the text guidance module for processing to obtain the first text embedding feature, the result of the processing by the visual detection module and the first text embedding feature are input into the text visual fusion module for processing together, and the result of the processing by the text visual fusion module is input into the prediction module for processing to obtain the mechanical tool detection result.
[0035] In the specific implementation, before the input machine tool image enters the model, data preprocessing operations are required to adapt to the model input requirements. The preprocessing process includes image scaling and normalization, proportion preservation and padding, and image conversion.
[0036] The main function of the visual detection module is to extract multi-scale features from the input image, provide necessary semantic and spatial information for subsequent target detection tasks, and generate multi-scale feature information.
[0037] likeFigure 2 As shown, the visual detection module includes a convolutional layer one, a convolutional layer two, an SPD module one, a C2f module one, a convolutional layer three, an SPD module two, a C2f module two, a convolutional layer four, an SPD module three, a C2f module three, a convolutional layer five, a C2f module four, an SPD module four and an SPPF pooling module which are connected in series in sequence; the input end of the convolutional layer one serves as the input end of the visual detection module, and the output end of the C2f module two, the output end of the C2f module three and the output end of the SPPF pooling module serve as the output end of the visual detection module at the same time.
[0038] Convolutional layer 1, convolutional layer 2, convolutional layer 3, convolutional layer 4 and convolutional layer 5 all use 3×3 convolutional layers. More specifically, the scale of the feature map output by the output end of C2f module 2 is 80×80, the scale of the feature map output by the output end of C2f module 3 is 40×40, and the scale of the feature map output by the output end of SPPF pooling module is 20×20.
[0039] The processing process of the visual inspection module is as follows: the mechanical tool image is sequentially input into the convolution layer 1, convolution layer 2, SPD module 1, C2f module 1, convolution layer 3, SPD module 2, C2f module 2, convolution layer 4, SPD module 3, C2f module 3, convolution layer 5, C2f module 4, SPD module 4 and SPPF pooling module for processing, and the output result of C2f module 2, the output result of C2f module 3 and the output result of SPPF pooling module are simultaneously used as the output end of the visual inspection module. In the specific implementation, the output result of C2f module 2, the output result of C2f module 3 and the output result of SPPF pooling module are feature maps of three scales.
[0040] The processing process of the SPD module is as follows: First, the input feature map X(S,S,C 1 ) is transformed through the SPD layer, sliced and reorganized into multiple sub-graphs, each sub-graph has a size of (S / scale, S / scale, C 1 ). These subgraphs are connected in the channel dimension to form an intermediate feature representation X'(S / scale,S / scale,scale2C 1 ), keeping the spatial dimension unchanged and increasing the number of channels. Next, the feature map is processed with a non-strided convolutional layer with a stride of 1. This convolutional layer does not change the spatial size of the feature map, but extracts important feature information in the increased channel and converts the intermediate representation X' into the final output X" (S / scale,S / scale,C 2 ), where C 2 Represents the number of filters in the convolutional layer (C 2 <scale2C 1 ).
[0041] The text guidance module adopts a contrastive language-image pre-training (CLIP) model.
[0042] The text-vision fusion module includes a first spatial feature extraction module, a second spatial feature extraction module, a third spatial feature extraction module, a fourth spatial feature extraction module, a first image pooling attention mechanism module, a second image pooling attention mechanism module, a first visual-text feature extraction attention module, a second visual-text feature extraction attention module, and a third visual-text feature extraction attention module; the first spatial feature extraction module, the second spatial feature extraction module, the third spatial feature extraction module, and the fourth spatial feature extraction module all have an image feature input end and a text feature input end;
[0043] The output result of the C2f module two is transmitted to the image feature input end of the first spatial feature extraction module for processing, the output result of the C2f module three is transmitted to the image feature input end of the second spatial feature extraction module for processing, the output result of the SPPF pooling module is transmitted to the input end of the first image pooling attention mechanism module for processing, the first text embedding feature is respectively input to the text feature input ends of the second spatial feature extraction module and the first spatial feature extraction module for processing, the output results of the first spatial feature extraction module and the second spatial feature extraction module are both transmitted to the input end of the first image pooling attention mechanism module for processing, the output result of the first image pooling attention mechanism module and the first text embedding feature are subjected to a first residual connection to obtain a second text embedding feature, the second text embedding feature is respectively input to the text feature input ends of the third spatial feature extraction module and the fourth spatial feature extraction module for processing, the output result of the first image pooling attention mechanism module is also respectively transmitted to the image feature input ends of the third spatial feature extraction module, the fourth spatial feature extraction module, and the input end of the first visual-text feature extraction attention module for processing, the output result of the third spatial feature extraction module is input to the input end of the second visual-text feature extraction attention module for processing, the output result of the fourth spatial feature extraction module is input to the input end of the third visual-text feature extraction attention module for processing, the output results of the first visual-text feature extraction attention module, the second visual-text feature extraction attention module, and the third visual-text feature extraction attention module are all input to the input end of the second image pooling attention mechanism module for processing, the output result of the second image pooling attention mechanism module and the second text embedding feature are subjected to a second residual connection to obtain an image perception embedding feature, and the image perception embedding feature and the input result of the second image pooling attention mechanism module are jointly used as the output result of the text-vision fusion module.
[0044] The text vision fusion module fuses the feature maps of three scales output by the visual detection module and the first text embedding features output by the text guidance module, and outputs feature maps containing semantic information (the input result of the second image pooling attention mechanism module) and image-perceived embedding features.
[0045] like Figure 3 As shown in the figure, the first spatial feature extraction module, the second spatial feature extraction module, the third spatial feature extraction module and the fourth spatial feature extraction module all adopt the same structure (SA-T-CSPLayer), which specifically includes a bottleneck module, an SA spatial attention module and an activation module; the result of the previous layer input is first feature segmented to obtain features of the same size Figure 1 and Features Figure 2 ,feature Figure 1 The input is processed in the bottleneck module. The result of the bottleneck module is input into the SA spatial attention module for processing. The result of the SA spatial attention module and the result of the bottleneck module are processed for feature merging. The result of the feature merging is input into the activation module for processing. The result of the activation module and the feature Figure 2 Perform residual connection, and the result of residual connection is used as the output of the structure.
[0046] In the specific implementation, the feature map of the previous layer input is 40×40×512. After feature segmentation, the obtained feature map Figure 1 and Features Figure 2 Both are 40×40×256.
[0047] In a specific implementation, feature merging is performed on the channel dimension.
[0048] The bottleneck module is the Dark Bottleneck module and the activation module is Max-Sigmoid; the first image pooling attention mechanism module and the second image pooling attention mechanism module both adopt the I-Pooling Attention module in YOLO-World; the first visual text feature extraction attention module, the second visual text feature extraction attention module and the third visual text feature extraction attention module are all EMA attention mechanisms.
[0049] The prediction module includes a text comparison head and a box head; the result output by the second image pooling attention mechanism module is input into the box head to obtain all prediction boxes in the image; the result output by the second image pooling attention mechanism module is also input into the text comparison head together with all the obtained prediction boxes to obtain the target embedding features of each prediction box, the target embedding features are matched with the image perception embedding features to obtain the regional text matching results, and the regional text matching results are combined with the prediction boxes to obtain the final mechanical tool detection results.
[0050] The prediction box only frames the position of the target in the image; the region text matching result is only the category information of the target in the prediction box. Combining the region text matching result and the prediction box gives the detection result that contains both position information and category information. The target embedding feature is a point in a high-dimensional space that encodes the feature information of the target (each object contains semantic information such as "scissors", "cross screwdriver", "pliers"), enabling comparison and differentiation between different objects. The box head predicts the position of each target in the image, which is represented in the form of bounding boxes, and each bounding box contains the position and size information of the target. When the target embedding feature is matched with the image perception embedding feature, the one with the highest matching degree is selected as the category of each prediction box.
[0051] The text contrast head adopts the Text Constrastive Head in YOLO-World;
[0052] The box head adopts the Box Head in YOLO-World.
[0053] The text contrast head adopts the cosine similarity calculation method.
[0054] The training configuration of this embodiment is as follows: the operating system is Linux (Ubuntu 20.04), the graphics processing unit GPU is RTX3080Ti, the Cud version is 11.3, and pytorch is 1.10.1.
[0055] The AdamW optimizer is used to iterate the model for 100 epochs, and the initial learning rate is set to 0.0002.
[0056] S3. Respectively obtain the image of the mechanical tool to be measured and the text information corresponding to the image of the mechanical tool to be measured, and input the image of the mechanical tool to be measured and the corresponding text information into the trained text target detection model for detection to obtain the mechanical tool detection result.
[0057] In this embodiment, the detection accuracy of the present invention is evaluated using three average precision metrics, namely AP, AP0.5, and AP0.75. The calculation formula is as follows:
[0058]
[0059] In the formula, P represents precision, which is the ratio of the number of positive samples predicted by the model to the number of all detected samples, and R represents recall, which is the ratio of the number of positive samples correctly predicted by the model to the number of positive samples actually appearing.
[0060] In this embodiment, a dataset of seven types of mechanical tools is constructed (G01: Phillips screwdriver, G02: slotted screwdriver, G03: knife, G04: tape measure, G05: scissors, G06: pliers, G07: wire strippers). Experiments are conducted on this dataset in this embodiment and compared with advanced algorithms. The comparisons are shown in Tables 1 and 2. Table 1 Comparison of the precision AP0.5 and AP0.75 between the method proposed in this invention and advanced algorithms
[0061]
[0062] Table 2 Comparison of the precision AP between the method proposed in this invention and advanced algorithms
[0063]
[0064] Four advanced methods are selected in Tables 1 and 2: RegionCLIP, Detic, OVR-CNN, YOLO-World-S. Among them, the category represents the category of evaluation indicators, and the evaluation indicators adopted are AP0.5 and AP0.75. It can be seen from the information in the table that compared with the comparison methods, the AP0.5 and AP0.75 of this invention are the highest among several comparison methods. Among the seven types of tools, for G03 (knife), the highest detection precision AP0.75 in the current advanced methods is 89%, and the AP0.75 of the proposed method reaches 93.3%, which is significantly increased by 4.4%.
[0065] In this invention, by introducing the SPD module into the visual detection module, the ability to process low-resolution images and represent the features of medium and small target objects is improved, and the detection precision of the model for actual target objects is enhanced. Specifically, the detection precision AP on the dataset of this embodiment increases from 77% to 78.9% compared with the model precision without introducing the SPD module.
[0066] In this invention, the EMA attention mechanism is introduced into the text-visual fusion module to improve the ability to represent text-guided visual features, and the SA spatial attention module is introduced to make the model pay more attention to the regions related to the text in the image, enhance the model's perception ability of text-specific target regions, and improve the target detection precision. Specifically, the detection precision AP on the dataset of this embodiment increases from 78.9% to 79.5% compared with the model detection precision without introducing the EMA attention mechanism.
[0067] Finally, it should be emphasized that the above embodiments are only a preferred specific implementation manner of the present invention. In order to help understand the method and core idea of the present invention, detailed descriptions are provided through embodiments in this article. However, the protection scope of the present invention is not limited to these specific embodiments. Any modification to the technical solutions in the foregoing embodiments, or a variant obtained by equivalently replacing some technical indicators, should be regarded as being included within the protection scope of the present invention.
Claims
1. A text-guided machine tool image target detection method, characterized in that: The following steps are involved: S1. Use a camera to capture pictures of mechanical tools, and create text information for each mechanical tool picture. Each mechanical tool picture and the corresponding text information form an image-text pair. The image-text pairs obtained from all mechanical tool pictures form a mechanical tool dataset. S2. Build a text object detection model, input the machine tool data set into the text object detection model for training, and obtain a trained text object detection model; S3. Obtain images of the mechanical tool to be tested and corresponding text information respectively, input the images of the mechanical tool to be tested and the corresponding text information into a trained text object detection model for detection, and obtain a mechanical tool detection result.
2. The text-guided machine tool image target detection method according to claim 1, characterized in that: The text target detection model includes a visual detection module, a text guidance module, a text visual fusion module and a prediction module; the mechanical tool pictures in the mechanical tool data set are input into the visual detection module for processing, the text information in the mechanical tool data set is input into the text guidance module for processing to obtain a first text embedding feature, the result of the processing by the visual detection module and the first text embedding feature are input into the text visual fusion module for processing together, and the result of the processing by the text visual fusion module is input into the prediction module for processing to obtain a mechanical tool detection result.
3. The text-guided machine tool image target detection method according to claim 2, characterized in that: The visual detection module includes a convolutional layer one, a convolutional layer two, an SPD module one, a C2f module one, a convolutional layer three, an SPD module two, a C2f module two, a convolutional layer four, an SPD module three, a C2f module three, a convolutional layer five, a C2f module four, an SPD module four and an SPPF pooling module which are connected in series in sequence; the input end of the convolutional layer one serves as the input end of the visual detection module, and the output end of the C2f module two, the output end of the C2f module three and the output end of the SPPF pooling module serve as the output end of the visual detection module at the same time.
4. The text-guided machine tool image target detection method according to claim 3, characterized in that: The text-guided module uses a contrastive language-image pre-training model.
5. The text-guided machine tool image target detection method according to claim 2, characterized in that: The text vision fusion module includes a first spatial feature extraction module, a second spatial feature extraction module, a third spatial feature extraction module, a fourth spatial feature extraction module, a first image pooling attention mechanism module, a second image pooling attention mechanism module, a first visual text feature extraction attention module, a second visual text feature extraction attention module and a third visual text feature extraction attention module; the first spatial feature extraction module, the second spatial feature extraction module, the third spatial feature extraction module and the fourth spatial feature extraction module all have an image feature input terminal and a text feature input terminal; The output result of the C2f module 2 is transmitted to the image feature input end of the first spatial feature extraction module for processing, the output result of the C2f module 3 is transmitted to the image feature input end of the second spatial feature extraction module for processing, the output result of the SPPF pooling module is transmitted to the input end of the first image pooling attention mechanism module for processing, the first text embedding feature is respectively input to the text feature input end of the second spatial feature extraction module and the text feature input end of the first spatial feature extraction module for processing, the output result of the first spatial feature extraction module and the output result of the second spatial feature extraction module are both transmitted to the input end of the first image pooling attention mechanism module for processing, the output result of the first image pooling attention mechanism module is subjected to the first residual connection with the first text embedding feature to obtain the second text embedding feature, the second text embedding feature is respectively input to the text feature input end of the third spatial feature extraction module and the text feature input end of the fourth spatial feature extraction module for processing, the first image pooling attention mechanism module The output results of the attention mechanism module are also transmitted to the image feature input end of the third spatial feature extraction module, the image feature input end of the fourth spatial feature extraction module and the input end of the first visual text feature extraction attention module for processing, the output result of the third spatial feature extraction module is input into the input end of the second visual text feature extraction attention module for processing, the output result of the fourth spatial feature extraction module is input into the input end of the third visual text feature extraction attention module for processing, the output results of the first visual text feature extraction attention module, the second visual text feature extraction attention module and the third visual text feature extraction attention module are all input into the input end of the second image pooling attention mechanism module for processing, the output result of the second image pooling attention mechanism module is residually connected with the second text embedding feature for the second time to obtain the image perception embedding feature, and the image perception embedding feature and the input result of the second image pooling attention mechanism module are used together as the output result of the text visual fusion module.
6. The method for machine tool image target detection based on text guidance according to claim 5, characterized in that: The first spatial feature extraction module, the second spatial feature extraction module, the third spatial feature extraction module and the fourth spatial feature extraction module all adopt the same structure, which specifically includes a bottleneck module, an SA spatial attention module and an activation module; the result of the previous layer input is first feature segmented to obtain feature map one and feature map two, the feature map one is input into the bottleneck module for processing, the result of the bottleneck module processing is input into the SA spatial attention module for processing, the result of the SA spatial attention module processing and the result of the bottleneck module processing are feature merged, the result of the feature merging is input into the activation module for processing, the result of the activation module processing and the feature map two are residually connected, and the result of the residual connection is used as the output of the structure.
7. The method for machine tool image target detection based on text guidance according to claim 6, characterized in that: The bottleneck module is the Dark Bottleneck module and the activation module is Max-Sigmoid; the first image pooling attention mechanism module and the second image pooling attention mechanism module both adopt the I-PoolingAttention module in YOLO-World; the first visual text feature extraction attention module, the second visual text feature extraction attention module and the third visual text feature extraction attention module are all EMA attention mechanisms.
8. The text-guided machine tool image target detection method according to claim 2, characterized in that: The prediction module includes a text comparison head and a box head; the result output by the second image pooling attention mechanism module is input into the box head to obtain a prediction box in the image; The output result of the second image pooling attention mechanism module is also input into the text comparison head together with the obtained prediction box to obtain the target embedding feature of each prediction box. The target embedding feature is matched with the image-perceived embedding feature to obtain the regional text matching result. The regional text matching result is combined with the prediction box to obtain the final mechanical tool detection result.
9. The text-guided machine tool image target detection method according to claim 8, characterized in that: The text comparison header adopts the text comparison header in YOLO-World; the frame header adopts the frame header in YOLO-World.