Bar detection method and device and storage medium

Through the method of channel space dual attention and deformable convolution combined with anchor frame library, the problem of unstable rod detection is solved, efficient and accurate detection and tracking of rods is achieved, and the degree of automation of production and the accuracy of material tracking is improved.

CN120510089APending Publication Date: 2025-08-19CERI DIGITAL TECHNOLOGY (BEIJING) CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510476594.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-16
Publication Date
2025-08-19

AI Technical Summary

Technical Problem

When detecting rods, especially in steel production, existing visual inspection technology has unstable detection and is prone to "rear-finishing" or "leftover" phenomena, resulting in reduced production efficiency and quality problems, making it difficult to meet the high-precision material tracking needs of rods.

Method used

The method of channel space dual attention and deformable convolution combined with anchor frame library is adopted to obtain bar image, scale perception and feature fusion, anchor frame prior information is constructed, and the category, position and motion state of bar are decoded.

Benefits of technology

It realizes efficient and accurate detection and tracking of rods in complex production environments, improves the accuracy and reliability of material tracking, and reduces errors and accidents in production.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120510089A_ABST
    Figure CN120510089A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a bar detection method and device and a storage medium, and relates to the technical field of steel and detection, and the method comprises the steps: obtaining a bar image of a target bar; performing scale perception on the bar image to obtain image features; carrying out weight distribution on the image features through channel space double-path attention to obtain a feature map containing texture information and semantic information, and carrying out feature fusion on the feature map to obtain multi-level features; an anchor frame library of the target bar is constructed through deformable convolution, and anchor frame prior information is determined according to the anchor frame library; and decoding the multi-level features according to the anchor frame prior information to obtain a detection result, wherein the detection result comprises the category, the position and the motion state of the bar. According to the method, the bar can be efficiently and accurately detected.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of steel and detection technology, and in particular to a method, device and storage medium for detecting bar materials. Background Art

[0002] Bars (especially steel bars) are widely used in key industries such as metallurgy, construction, and manufacturing. Due to the enormous demand, the scale and efficiency of steel production directly impact the industry's supply chain. The production of steel bars typically involves multiple steps, including heating, rolling, shearing, sizing, and bundling. From heating the steel billet to packaging the finished product, each step is crucial. To ensure efficient production and reduce waste and errors, an accurate material tracking system is crucial.

[0003] With the advancement of industrial automation, vision-based bar inspection technology is playing an increasingly important role in steel production. Traditional material tracking methods rely on PLC signals, using logic and data calculations to track the position and status of bars. However, these methods are susceptible to hardware limitations, signal instability, and human intervention. In contrast, visual inspection systems use cameras to capture real-time images of bars during production. Combined with advanced object detection algorithms, they enable precise material tracking and improve production automation.

[0004] However, existing visual inspection technology still faces some technical bottlenecks when dealing with rod-shaped objects. First, rods have an extreme aspect ratio, a relatively simple shape, and a simple surface texture, which makes general target detection algorithms ineffective when detecting rods. In addition, the surface of rods is easily reflective, surface oxidation produces a large amount of iron sheets, they are densely arranged, have variable sizes, and the production environment is complex and may be blocked, all of which place higher demands on the robustness of the algorithm. Current detection technology is unstable when dealing with these problems, especially in material tracking in the sawing area, where "tailgating" or "falling behind" often occurs, leading to information errors, rhythm confusion, and even production accidents in downstream processes.

[0005] Therefore, in response to the practical problems in steel bar production, there is an urgent need for a specially optimized bar detection technology to improve the accuracy and reliability of material tracking throughout the entire production process and avoid reduced production efficiency and quality problems caused by inaccurate detection. Summary of the Invention

[0006] The purpose of the embodiments of the present invention is to provide a method, device and storage medium for detecting bar materials, which can realize efficient and accurate detection of bar materials.

[0007] To achieve the above objectives, an embodiment of the present invention provides a method for detecting a bar, the method comprising:

[0008] Acquire a bar image of the target bar;

[0009] Performing scale perception on the bar image to obtain image features;

[0010] The image features are weighted by using dual-path attention in channel space to obtain a feature map containing texture information and semantic information, and the feature map is fused to obtain multi-level features;

[0011] Constructing an anchor frame library of the target bar through deformable convolution, and determining anchor frame prior information according to the anchor frame library;

[0012] The multi-level features are decoded according to the anchor frame prior information to obtain a detection result, where the detection result includes the category, position, and motion state of the bar.

[0013] Optionally, performing scale perception on the bar image to obtain image features includes:

[0014] Aligning rod images of different scales using deformable convolution, and performing upsampling and downsampling operations on the aligned rod images to obtain transformed images;

[0015] The converted image is subjected to feature extraction through multiple convolutional layers to obtain image features.

[0016] Optionally, extracting features from the converted image using multiple convolutional layers to obtain image features includes:

[0017] Extracting resolution features in the converted image through different layer networks;

[0018] Image features of bars with different diameters are extracted according to the resolution features.

[0019] Optionally, the weight distribution of the image features by channel-space dual-path attention to obtain a feature map containing texture information and semantic information includes:

[0020] The channel-space dual-path attention includes channel attention and spatial attention;

[0021] Using the channel attention to pool the image to generate a channel weight vector, and performing full connection layer learning on the image to determine the dependency relationship between channels;

[0022] Extracting contextual information of the image using the spatial attention pooling and depthwise separable deformable convolution to generate a spatial weight matrix;

[0023] The channel weight vectors, inter-channel dependencies and spatial weight matrices are element-wise added and channel-wise concatenated to obtain a feature map containing texture information and semantic information.

[0024] Optionally, constructing an anchor frame library of the target bar through deformable convolution, and determining anchor frame prior information according to the anchor frame library includes:

[0025] Determine an anchor frame library of the target bar based on a deformable convolution and offset prediction network, wherein the anchor frame library includes a rotation angle and an aspect ratio, wherein the rotation angle is 0°-180° and the aspect ratio is 1:5-1:50;

[0026] The center coordinates, major axis length, minor axis length and rotation angle of the anchor frame library are encoded as anchor frame prior information through geometric constraints.

[0027] Optionally, encoding the center coordinates, major axis length, minor axis length, and rotation angle of the anchor frame library as anchor frame prior information through geometric constraints includes:

[0028] Q i =MLP(x c ,y c ,w,h,θ)

[0029] Among them, Q i is the anchor box prior information,

[0030] (x c ,y c ) is the center coordinate of the anchor box,

[0031] w,h are the length of the major axis and the length of the minor axis respectively,

[0032] θ is the rotation angle,

[0033] MLP stands for Multi-layer Perceptron.

[0034] Optionally, decoding the multi-level features according to the anchor box prior information to obtain a detection result includes:

[0035] Matching the multi-level features with the anchor box prior information through a decoder to obtain matching information;

[0036] The feature differences of adjacent frames in the matching information are predicted by temporal expansion and optical flow residual as the detection result.

[0037] Optionally, the method further includes:

[0038] In the self-attention layer of the decoder, an attention radius along the long axis of the target bar is set to be four times that of the short axis.

[0039] On the other hand, the present invention also provides a device for detecting a bar, the device comprising:

[0040] An acquisition module, used for acquiring a bar image of a target bar;

[0041] A first processing module is used to perform scale perception on the bar image to obtain image features;

[0042] The second processing module is used to weight the image features through channel space dual-path attention to obtain a feature map containing texture information and semantic information, and perform feature fusion on the feature map to obtain multi-level features;

[0043] A third processing module is configured to construct an anchor frame library of the target bar through deformable convolution, and determine anchor frame prior information according to the anchor frame library;

[0044] The fourth processing module is used to decode the multi-level features according to the anchor frame prior information to obtain a detection result, where the detection result includes the category, position and motion state of the bar.

[0045] On the other hand, the present invention further provides a processor configured to execute the above-mentioned method for detecting bar materials.

[0046] On the other hand, the present invention further provides a machine-readable storage medium having instructions stored thereon. When the instructions are executed by a processor, the processor is configured to execute the above-mentioned method for detecting bar materials.

[0047] On the other hand, the present invention further provides a computer program product, comprising a computer program, which implements the above-mentioned method for detecting bar materials when executed by a processor.

[0048] A method for detecting rods of the present invention includes: obtaining a rod image of a target rod; performing scale perception on the rod image to obtain image features; weighting the image features using channel-space dual-path attention to obtain a feature map containing texture information and semantic information, and performing feature fusion on the feature map to obtain multi-level features; constructing an anchor frame library of the target rod using deformable convolution, and determining anchor frame prior information based on the anchor frame library; decoding the multi-level features based on the anchor frame prior information to obtain a detection result, wherein the detection result includes the category, position, and motion state of the rod. This method performs channel-space dual-path attention on the image, combines the constructed anchor frame library with its multi-level features to decode the category, position, and motion state of the rod, and can efficiently and accurately complete the detection and tracking of rods in complex production environments.

[0049] Other features and advantages of the embodiments of the present invention will be described in detail in the subsequent detailed description. BRIEF DESCRIPTION OF THE DRAWINGS

[0050] The accompanying drawings are used to provide a further understanding of the embodiments of the present invention and constitute a part of the specification. Together with the following detailed description, they are used to explain the embodiments of the present invention, but do not constitute a limitation of the embodiments of the present invention. In the accompanying drawings:

[0051] Figure 1 This is a schematic flow chart of a method for detecting bar materials according to the present invention;

[0052] Figure 2 is a schematic diagram of an embodiment of the present invention;

[0053] Figure 3 is a schematic diagram of a feature adaptation module of the present invention;

[0054] Figure 4 is a schematic diagram of a second embodiment of the present invention;

[0055] Figure 5 It is a schematic diagram of a device for detecting rods of the present invention;

[0056] Figure 6 It is a diagram of a computer device.

[0057] Description of Reference Numerals

[0058] 100- Device for detecting bars;

[0059] 200-Get module;

[0060] 300-first processing module;

[0061] 400-second processing module;

[0062] 500-third processing module;

[0063] 600-Fourth processing module. DETAILED DESCRIPTION

[0064] The following describes the specific implementation of the embodiment of the present invention in detail with reference to the accompanying drawings. It should be understood that the specific implementation described herein is only used to illustrate and explain the embodiment of the present invention and is not used to limit the embodiment of the present invention.

[0065] It should be noted that the acquisition, transmission, storage, use, and processing of data in the technical solution of this application are in compliance with the relevant provisions of laws and regulations. In the embodiments of this application, certain software, components, models, and other existing solutions in the industry may be mentioned. These should be considered as exemplary. Their purpose is only to illustrate the feasibility of implementing the technical solution of this application, but it does not mean that the applicant has or will necessarily use such solutions.

[0066] Example 1

[0067] Figure 1 This is a flow chart of a method for detecting rods of the present invention. Figure 1 As shown, a method for detecting a bar material of the present invention includes: step S101 is to obtain a bar material image of a target bar material.

[0068] According to one specific embodiment, a camera is positioned within the sawing area to capture real-time images of the target bar. Specifically, the camera utilizes a high-resolution camera to capture high-definition images of the bar during production. The camera is positioned appropriately within the sawing area to ensure full visibility of the target bar. The target bar can be a long object, such as a steel bar.

[0069] Step S102 is to perform scale perception on the bar image to obtain image features.

[0070] According to a specific embodiment, performing scale perception on the rod image to obtain image features includes: using deformable convolution to align rod images of different scales, performing upsampling and downsampling operations on the aligned rod image to obtain a converted image; and performing feature extraction on the converted image through multiple convolution layers to obtain image features.

[0071] The method of extracting features from the converted image through multiple convolutional layers to obtain image features includes: extracting resolution features from the converted image through different layers of networks; and extracting image features of rods of different diameters based on the resolution features.

[0072] Specifically, multiple convolutional layers are introduced to extract features at different scales from the bar image, resulting in feature maps (i.e., bar images) of different scales. High-resolution features (stride = 8) are extracted in the shallow network (e.g., conv3_x of ResNet-50) to detect thin bars with diameters less than 50 mm, resulting in a feature map for thin bars. Mid-scale features (stride = 16) are extracted in the middle network (conv4_x) to adapt to standard bars with diameters of 50-150 mm, resulting in a feature map for standard bars. Low-resolution features (stride = 32) are extracted in the deep network (conv5_x) to detect thick bars with diameters greater than 150 mm, resulting in a feature map for thick bars. Deformable convolutions (Deformable ConvNets) are used to align feature maps of different scales to address cross-layer misalignment caused by bar tilt.

[0073] The transformed image is obtained by performing scale transformation (such as upsampling and downsampling operations) on the feature maps of different scales. In order to adapt to slender objects of different sizes, dynamic weights are assigned to the output of each scale branch, and the calculation formula is: where s iis the confidence score of each scale branch, calculated by lightweight MLP.

[0074] By performing scale perception on the bar image, the scale of the feature map can be automatically adjusted to better handle the diversity of slender objects and improve the accuracy and efficiency of detection.

[0075] Step S103 is to perform weight distribution on the image features through channel space dual-path attention to obtain a feature map containing texture information and semantic information, and perform feature fusion on the feature map to obtain multi-level features.

[0076] According to a specific implementation method, the image features are weighted by channel-space dual-path attention to obtain a feature map containing texture information and semantic information, including: the channel-space dual-path attention includes channel attention and spatial attention; the channel attention is used to pool the image to generate a channel weight vector, and the image is subjected to full-connected layer learning to determine the channel dependency; the spatial attention pooling and depth-wise separable deformable convolution are used to extract the contextual information of the image for generating a spatial weight matrix; the channel weight vector, the channel dependency and the spatial weight matrix are element-wise added and channel-wise spliced to obtain a feature map containing texture information and semantic information.

[0077] Specifically, such as Figure 2 As shown, this application introduces an attention mechanism to assign different weights to different areas of the feature map to highlight the key feature areas of slender objects. Features from different scales and levels are fused to capture the multi-level features of slender objects. Local enhancement is to focus on the edge area of the bar through spatial attention (CBAM module); global association is to use channel attention (SE module) to strengthen cross-channel semantic association and automatically balance the two-way output through the gating mechanism; the parameters of the adaptive feature module are optimized through deep learning technology to enable it to operate stably in various complex environments.

[0078] like Figure 3 As shown in the figure, scale refers to the size of the feature map. Scale 1 is 4 times the input image downsampled, scale 2 is 8 times the input image downsampled, and scale 3 is 16 times the input image downsampled. The acquisition method is to downsample the previous scale through a convolution operation.

[0079] This application uses a channel-space dual-path attention module to dynamically assign weights to the image features, wherein: the channel attention branch generates a channel weight vector through pooling and uses a fully connected layer to learn the dependencies between channels; the spatial attention branch uses pooling and depth-separable deformable convolution to extract global context information and generate a spatial weight matrix; a combination of element-by-element addition, channel splicing, and branch gating mechanisms is used to highlight the key feature areas of slender objects and construct a feature map containing fine-grained texture information and semantic information. This method can dynamically adjust internal parameters according to different input features to better adapt to the feature representation of slender objects.

[0080] Step S104 is to construct an anchor frame library of the target bar through deformable convolution, and determine anchor frame prior information according to the anchor frame library.

[0081] According to a specific embodiment, the anchor frame library of the target bar is constructed by deformable convolution, and the anchor frame prior information is determined according to the anchor frame library, including: determining the anchor frame library of the target bar based on deformable convolution and offset prediction network, the anchor frame library including rotation angle and aspect ratio, the rotation angle is 0°-180°, and the aspect ratio is 1:5-1:50; encoding the center coordinates, major axis length, minor axis length and rotation angle of the anchor frame library as anchor frame prior information through geometric constraints.

[0082] The method of encoding the center coordinates, major axis length, minor axis length and rotation angle of the anchor frame library into anchor frame prior information through geometric constraints includes:

[0083] Q i =MLP(x c ,y c ,w,h,θ)

[0084] Among them, Q i is the anchor box prior information, (x c ,y c ) is the center coordinate of the anchor frame, w, h are the major axis length and minor axis length respectively, θ is the rotation angle, and MLP is a multi-layer perceptron.

[0085] Specifically, based on the statistical analysis of historical production line data, a library of anchor frames with aspect ratios ranging from 1:5 to 1:20 (such as 50×1000px, 60×1200px) is constructed, and the shape of the anchor frames is dynamically adjusted through deformable convolution; prior geometric information is injected into the decoder (Object Queries). The specific formula is: Q i =MLP(x c ,y c ,w,h,θ). Where (x c ,y c) are the center coordinates of the anchor frame, w, h are the width and height, and θ is the inclination angle of the rod. The multi-layer perceptron MLP maps the five-dimensional vector to a high-dimensional feature space. In the decoder self-attention layer, an attention radius four times that of the short axis (Y axis) is set along the long axis (X axis) of the rod to reduce the computational complexity of the invalid area.

[0086] Specifically, based on the offset prediction network of deformable convolution, an anchor frame parameter set that adapts to the aspect ratio characteristics of the rod is dynamically generated. The anchor frame includes a rectangular frame with a rotation angle of 0°-180° and an aspect ratio of 1:5 to 1:50; the center coordinates, major axis / minor axis length and rotation angle of the anchor frame are encoded into a position prior vector of the decoder through a geometric constraint module (multi-layer perceptron).

[0087] This application specifically integrates anchors with extreme aspect ratios to achieve a balance between algorithm speed and accuracy. To address the extreme aspect ratios of bars, this application reintroduces anchors into the model to improve the performance of the decoder module. This improved object query construction significantly improves the decoder's efficiency when handling slender objects.

[0088] Step S105 is to decode the multi-level features according to the anchor frame prior information to obtain a detection result, and the detection result includes the category, position and motion state of the bar.

[0089] According to a specific embodiment, decoding the multi-level features based on the anchor frame prior information to obtain a detection result includes: matching the multi-level features with the anchor frame prior information via a decoder to obtain matching information; and predicting the feature differences of adjacent frames in the matching information using temporal expansion and optical flow residuals to obtain the detection result. The method also includes setting an attention radius along the long axis of the target bar four times that of the short axis in the self-attention layer of the decoder.

[0090] Specifically, a cascaded Transformer decoder architecture is adopted to match multi-level features with anchor box prior information through a multi-head cross-attention mechanism; an optical flow residual prediction branch is introduced in the temporal expansion module, and the motion vector of the rod is calculated by differentially analyzing the features of adjacent frames; the output layer generates three sets of prediction parameters in parallel: the classification branch uses the Sigmoid function to output the probability of the rod material category; the regression branch outputs the five-dimensional parameters of the rotated bounding box (center coordinates x, y, major axis l, minor axis s, angle θ); and the motion state branch outputs the velocity vector and motion direction angle.

[0091] This application builds a dynamic aspect ratio anchor frame library (1:5 to 1:20) based on historical production line data, and realizes adaptive adjustment of anchor frame shape through deformable convolution; injects geometric prior information into Object Queries, including center coordinates (x c ,yc ), width, height w, h and tilt angle θ are mapped to a high-dimensional feature space through MLP; an asymmetric attention mechanism is adopted to set an attention radius four times that of the short axis (Y axis) along the long axis (X axis) of the bar to reduce the computational complexity.

[0092] In addition, the present application also uses multimodal dynamic feature fusion technology and progressive multi-scale perception technology. The multimodal dynamic feature fusion technology specifically includes a local enhancement path and a global association path. The local enhancement path focuses on the edge area of the rod based on spatial attention (CBAM module); the global association path strengthens cross-channel semantic association through channel attention (SE module); the present application also balances the weights of the two outputs through a dynamic gating mechanism. The progressive multi-scale perception technology includes, according to the rod diameter classification detection: shallow high-resolution features (stride = 8) detect thin rods with a diameter of <50mm; middle-level features (stride = 16) adapt to 50-150mm standard rods; deep low-resolution features (stride = 32) detect thick rods >150mm; cross-scale feature alignment is achieved through deformable convolution to solve the inter-layer misalignment problem of tilted rods; dynamic weighted multi-scale output based on scale confidence score, calculation formula:

[0093]

[0094] By performing channel-space dual-path attention processing on the image and decoding its multi-level features in combination with the constructed anchor frame library, the category, position and motion state of the rods can be obtained, which can efficiently and accurately complete the detection and tracking of rods in complex production environments.

[0095] Example 2

[0096] The present invention proposes a target detection method optimized for bar materials, such as Figure 4 As shown, the first step is image acquisition. After the system starts, the image acquisition module collects images of the bar to be inspected in real time and transmits them to the data processing module. Next, data processing is performed. After preprocessing, the image is sent to the feature extraction submodule to extract multi-level features. The feature data is then passed to the target detection submodule, which uses the trained model to identify the target and obtain the test results. Next, the results are displayed. The bar inspection results, including category, location, and status information, are fed back to the operator in real time through the display module for immediate tracking and processing. Finally, subsequent processing is carried out. Based on the inspection results, the system can perform corresponding operations according to production needs, such as alarming, recording, or automatically adjusting production parameters.

[0097] The image processing module uses a high-resolution camera to capture real-time images of the bars during production. The camera is positioned appropriately within the sawing area to ensure a full view of the object being inspected. The data processing module includes three submodules: the image preprocessing submodule removes noise, enhances, and normalizes the acquired images to improve the accuracy of subsequent inspections. The feature extraction submodule utilizes an adaptive feature module to dynamically adjust internal parameters to better extract the features of slender objects. The object detection submodule, based on an improved decoder module, uses a trained object detection model to identify objects and obtain detection results, including the category, position, and motion status of each bar. The display module displays the inspection results in real time, providing a visual interface for relevant information, allowing operators to easily monitor the production process and material status.

[0098] Specifically, this method involves dataset preparation: collecting images of bar materials with various aspect ratios to construct diverse training and validation sets. The dataset ensures coverage of a variety of complex backgrounds and lighting conditions to enhance the model's generalization capabilities. Feature extraction then proceeds. Using an adaptive feature module, a convolutional neural network (CNN) is employed to extract high-level image features. Feature maps at different scales are then fused to capture multi-level information about slender objects. Model optimization then follows: an improved decoder module is employed to optimize features for slender objects, making object queries more adaptable to bar-shaped objects with extreme aspect ratios. An attention mechanism is introduced to achieve dynamic weight adjustment, improving the accuracy of feature representation. Finally, training and evaluation are performed: a loss function is used to train the model, and detection accuracy is improved by continuously optimizing model parameters. Model performance is evaluated using a validation set, and hyperparameter adjustments are performed to ensure stable operation under various environments. Table 1 shows a comparison of the performance of this method with existing state-of-the-art methods on the bar material dataset.

[0099] Table 1:

[0100] algorithm MS-COCO (objects with aspect ratio greater than 10) BarSence Datasets DINO 38.1 92.6 Ours 54.8 97.8

[0101] Using the DINO algorithm as a baseline, the detection accuracy for objects with an aspect ratio greater than 10 is 38.1% AP50 on the open-source detection dataset MS-COCO, and 92.6% on the self-built BarSence dataset. This achievement, building on the DINO algorithm, improves the detection accuracy to 54.8% AP50 on MS-COCO and 97.8% AP50 on the self-built BarSence dataset.

[0102] This method, combining an adaptive feature module with a scale perception module, has achieved remarkable results in detecting slender objects. By dynamically adjusting feature extraction and model parameters, the system can efficiently and accurately detect and track bar materials in complex production environments. Material tracking accuracy has increased from 98.7% to 99.0%, and system response time has been reduced from 130ms to 90ms, meeting the real-time requirements of high-speed production lines.

[0103] Example 3

[0104] On the other hand, the present invention also provides a device for detecting rods, such as Figure 5 As shown, the device 100 for detecting rods includes: an acquisition module 200, used to acquire a rod image of a target rod; a first processing module 300, used to perform scale perception on the rod image to obtain image features; a second processing module 400, used to weight the image features through channel space dual-way attention to obtain a feature map containing texture information and semantic information, and perform feature fusion on the feature map to obtain multi-level features; a third processing module 500, used to construct an anchor frame library of the target rod through deformable convolution, and determine anchor frame prior information based on the anchor frame library; a fourth processing module 600, used to decode the multi-level features based on the anchor frame prior information to obtain a detection result, wherein the detection result includes the category, position and motion state of the rod.

[0105] A method for detecting rods of the present invention includes: obtaining a rod image of a target rod; performing scale perception on the rod image to obtain image features; weighting the image features using channel-space dual-path attention to obtain a feature map containing texture information and semantic information, and performing feature fusion on the feature map to obtain multi-level features; constructing an anchor frame library of the target rod using deformable convolution, and determining anchor frame prior information based on the anchor frame library; decoding the multi-level features based on the anchor frame prior information to obtain a detection result, wherein the detection result includes the category, position, and motion state of the rod. This method performs channel-space dual-path attention on the image, combines the constructed anchor frame library with its multi-level features to decode the category, position, and motion state of the rod, and can efficiently and accurately complete the detection and tracking of rods in complex production environments.

[0106] An embodiment of the present application provides a storage medium having a program stored thereon, which implements the above-mentioned method for detecting rods when the program is executed by a processor.

[0107] An embodiment of the present application provides a processor, which is used to run a program, wherein the program executes the above-mentioned method for detecting rods when running.

[0108] In one embodiment, a computer device is provided. The computer device may be a terminal, and its internal structure diagram may be as follows: Figure 6 As shown. The computer device includes a processor A01, a network interface A02, a display screen A04, an input device A05 and a memory (not shown in the figure) connected via a system bus. Among them, the processor A01 of the computer device is used to provide computing and control capabilities. The memory of the computer device includes an internal memory A03 and a non-volatile storage medium A06. The non-volatile storage medium A06 stores an operating system B01 and a computer program B02. The internal memory A03 provides an environment for the operation of the operating system B01 and the computer program B02 in the non-volatile storage medium A06. The network interface A02 of the computer device is used to communicate with an external terminal via a network connection. When the computer program is executed by the processor A01, a method for evaluating and detecting rods is implemented. The display screen A04 of the computer device can be a liquid crystal display or an electronic ink display, and the input device A05 of the computer device can be a touch layer covering the display screen, or a button, trackball or touchpad provided on the computer device housing, or an external keyboard, touchpad or mouse, etc.

[0109] An embodiment of the present application provides a device, which includes a processor, a memory, and a program stored in the memory and executable on the processor. When the processor executes the program, the method for detecting a bar material according to any embodiment of the present invention is implemented.

[0110] The present application also provides a computer program product, which, when executed on a data processing device, is suitable for executing a program that initiates the method for detecting a bar material according to any embodiment of the present invention.

[0111] Those skilled in the art will appreciate that the embodiments of the present application can be provided as methods, systems, or computer program products. Therefore, the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.

[0112] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the steps in the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0113] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.

[0114] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.

[0115] In a typical configuration, a computing device includes one or more processors (CPUs), input / output interfaces, network interfaces, and memory.

[0116] The memory may include non-permanent memory in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. The memory is an example of a computer-readable medium.

[0117] Computer-readable media includes permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology for information storage. Information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic disk storage or other magnetic storage devices, or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory computer-readable media such as modulated data signals and carrier waves.

[0118] It should be noted that in the embodiments of the present application, certain software, components, models and other existing solutions in the industry may be mentioned. They should be regarded as exemplary. Their purpose is only to illustrate the feasibility of implementing the technical solution of the present application, but it does not mean that the applicant has or will necessarily use the solution.

[0119] It should also be noted that the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, commodity, or apparatus that includes a series of elements includes not only those elements but also other elements not explicitly listed, or includes elements inherent to such process, method, commodity, or apparatus. In the absence of further limitations, an element defined by the phrase "comprises a ..." does not exclude the presence of other identical elements in the process, method, commodity, or apparatus that includes the element.

[0120] The above are merely embodiments of the present application and are not intended to limit the present application. For those skilled in the art, the present application may have various changes and variations. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application should all be included within the scope of the claims of the present application.

Claims

1. A method for detecting a bar, characterized in that: The method includes: Acquire a bar image of the target bar; Performing scale perception on the bar image to obtain image features; The image features are weighted by using dual-path attention in channel space to obtain a feature map containing texture information and semantic information, and the feature map is subjected to feature fusion to obtain multi-level features; Constructing an anchor frame library of the target bar through deformable convolution, and determining anchor frame prior information according to the anchor frame library; The multi-level features are decoded according to the anchor frame prior information to obtain a detection result, where the detection result includes the category, position, and motion state of the bar.

2. The method according to claim 1, characterized in that The performing scale perception on the bar image to obtain image features includes: Aligning rod images of different scales using deformable convolution, and performing upsampling and downsampling operations on the aligned rod images to obtain transformed images; The converted image is subjected to feature extraction through multiple convolutional layers to obtain image features.

3. The method according to claim 2, characterized in that The step of extracting features from the converted image through multiple convolutional layers to obtain image features includes: Extracting resolution features from the converted image through different layer networks; Image features of bars with different diameters are extracted according to the resolution features.

4. The method according to claim 1, wherein The weight distribution of the image features by the channel space dual-path attention to obtain a feature map containing texture information and semantic information includes: The channel-space dual-path attention includes channel attention and spatial attention; Using the channel attention to pool the image to generate a channel weight vector, and performing fully connected layer learning on the image to determine the dependency between channels; Extracting contextual information of the image using the spatial attention pooling and depthwise separable deformable convolution to generate a spatial weight matrix; The channel weight vectors, inter-channel dependencies and spatial weight matrices are element-wise added and channel-wise concatenated to obtain a feature map containing texture information and semantic information.

5. The method according to claim 1, wherein The constructing of the anchor frame library of the target bar by deformable convolution and determining anchor frame prior information according to the anchor frame library include: Determine an anchor frame library of the target bar based on a deformable convolution and offset prediction network, wherein the anchor frame library includes a rotation angle and an aspect ratio, wherein the rotation angle is 0°-180° and the aspect ratio is 1:5-1:50; The center coordinates, major axis length, minor axis length and rotation angle of the anchor frame library are encoded as anchor frame prior information through geometric constraints.

6. The method according to claim 5, characterized in that The method of encoding the center coordinates, major axis length, minor axis length and rotation angle of the anchor frame library into anchor frame prior information through geometric constraints includes: Q i =MLP(x c ,y c ,w,h,θ) Among them, Q i is the anchor box prior information, (x c ,y c ) is the center coordinate of the anchor box, w,h are the length of the major axis and the length of the minor axis respectively, θ is the rotation angle, MLP stands for Multi-layer Perceptron.

7. The method according to claim 1, characterized in that The decoding of the multi-level features according to the anchor frame prior information to obtain a detection result includes: Matching the multi-level features with the anchor box prior information through a decoder to obtain matching information; The feature differences of adjacent frames in the matching information are predicted by temporal expansion and optical flow residual as the detection result.

8. The method according to claim 7, characterized in that The method further includes: In the self-attention layer of the decoder, an attention radius along the long axis of the target bar is set to be four times that of the short axis.

9. A device for detecting bar material, characterized in that: The device includes: An acquisition module, used for acquiring a bar image of a target bar; A first processing module is used to perform scale perception on the bar image to obtain image features; The second processing module is used to weight the image features through channel space dual-path attention to obtain a feature map containing texture information and semantic information, and perform feature fusion on the feature map to obtain multi-level features; A third processing module is configured to construct an anchor frame library of the target bar through deformable convolution, and determine anchor frame prior information according to the anchor frame library; The fourth processing module is used to decode the multi-level features according to the anchor frame prior information to obtain a detection result, where the detection result includes the category, position and motion state of the bar.

10. A processor, characterized in that: The apparatus is configured to perform the method for inspecting a bar material according to any one of claims 1 to 8.

11. A machine-readable storage medium having instructions stored thereon, characterized in that: When the instructions are executed by a processor, the processor is configured to perform the method for detecting a bar material according to any one of claims 1 to 8.

12. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the computer program implements the method for detecting a bar material according to any one of claims 1 to 8.