An adaptive multi-task visual measurement method

CN122265272BActive Publication Date: 2026-09-29CHINA TOWER CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202610702732.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-05-21
Publication Date
2026-09-29
Estimated Expiration
2046-05-21

AI Technical Summary

Technical Problem

但仍存在一些针对性不足的问题:一是特征共享与任务专属特征平衡难,易出现“任务干扰”;二是特征融合为静态固定权重,无法适配图像内容动态调整;三是缺乏物理一致性约束,测量结果可靠性不足

Benefits of technology

(1)提升复杂场景适配能力:本发明基于CSPDarknet网络优化,保留其高效特征提取优势,新增可变形卷积模块组与动态特征校准模块,可变形卷积通过自适应感受野适配变形目标,动态特征校准强化弱纹理目标特征;

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122265272B_ABST
    Figure CN122265272B_ABST
Patent Text Reader

Abstract

The present application relates to a kind of self-adapting multitask vision measurement method, belong to vision measurement technical field.The method uses self-adapting multitask vision measurement network model;The self-adapting multitask vision measurement network model includes input layer, feature extraction layer, task perception separation layer, content dynamic fusion layer and collaborative inference output layer;The present application is based on CSPDarknet network optimization, retains its efficient feature extraction advantage, adds deformable convolution module group and dynamic feature calibration module, and deformable convolution is adapted to deformation target by adaptive receptive field, and dynamic feature calibration strengthens weak texture target feature;Innovative design dynamic fusion and collaborative inference mechanism, compared with single-task neural network model series connection scheme, improve detection efficiency, while avoiding double task result conflict.The scene adaptation rate of the present application reaches 96.2%, realizes the precise adaptation of model in specific scene.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of visual measurement technology, specifically relating to an adaptive multi-task visual measurement method. Background Technology

[0002] With the intelligent upgrading of the manufacturing industry, AI quality inspection, with its advantages of high efficiency, accuracy, and scalability, is gradually replacing traditional manual quality inspection and becoming a core technical means to ensure product quality. Visual measurement, as a key component of AI quality inspection, mainly uses image acquisition and processing technology to automatically detect key indicators such as product geometry, conformity markings, and printed markings (such as digital scales). It is widely used in quality inspection scenarios in fields such as tower construction, machining, and building components. In infrastructure projects, three-tube towers, as core supporting structures in the communications and power sectors, directly determine the overall stability of their foundation steel reinforcement engineering. It is crucial to measure indicators such as the length, spacing, and quantity of the steel reinforcement, as accurate measurement of these parameters is essential for project safety.

[0003] Currently, visual measurement technology has gradually evolved from the traditional machine vision stage to the deep learning-driven intelligent visual measurement stage. Traditional machine vision methods rely on manually designed features (such as SIFT and HOG features) and fixed rules, which are applicable to structured scenarios, but have poor robustness and adaptability when facing complex industrial environments (such as changes in lighting, target deformation, occlusion, etc.). The introduction of deep learning technology has significantly improved the performance of visual measurement, and methods based on models such as convolutional neural networks have made breakthrough progress in tasks such as target detection and image recognition. However, general-purpose algorithm models are difficult to adapt well to specific visual measurement tasks, especially in the specific scenario of rebar dimension measurement, where complex conditions such as lighting, occlusion, and corrosion amplify the limitations of insufficient model adaptability.

[0004] In recent years, multi-task learning has become a research hotspot in the field of visual measurement. Its core idea is to achieve information complementarity among multiple tasks by sharing feature extraction layers, thereby improving the overall performance and efficiency of the model. However, some problems still exist: first, it is difficult to balance feature sharing with task-specific features, which can easily lead to "task interference"; second, feature fusion uses static fixed weights, which cannot be dynamically adjusted to adapt to image content; and third, there is a lack of physical consistency constraints, resulting in insufficient reliability of measurement results. Therefore, developing multi-task deep collaborative visual measurement technology for specific scenarios and achieving scenario-based adaptation and optimization of general models can not only solve the practical needs of rebar measurement but also provide technical support for the measurement of similar infrastructure components, which is of great significance for improving the efficiency of industrial and infrastructure quality inspection and reducing production costs.

[0005] Traditional machine vision measurement technologies, such as CN111189387A "A Machine Vision-Based Method for Dimensional Inspection of Industrial Parts," focus on the precise geometrical dimension inspection of industrial parts. The core steps include: acquiring workpiece images using a high-resolution camera; eliminating image noise interference through median filtering; and achieving illumination equalization through grayscale correction to improve image quality. A standard calibration board is used to establish the mapping relationship between the pixel coordinate system and the physical coordinate system, determining the pixel equivalent and providing a basis for dimension quantization. The Canny edge detection algorithm is used to extract the workpiece contour edges, and the RANSAC algorithm is used to fit the geometric shapes such as lines and arcs in the contour, eliminating edge noise points and accurately locating the feature contour. Based on the fitted geometric features, the length, width, aperture, angle, and other dimensional parameters of the part are calculated, and the results are compared with preset standard thresholds to output the pass / fail inspection result. Its core is a "manually designed feature + step-by-step processing" technical approach, relying on fixed algorithm parameters and preset feature templates to extract target information. It is only suitable for structured industrial scenarios with stable lighting, regular workpiece shapes, and no deformation.

[0006] Deep learning-based visual measurement techniques, such as CN121121542A "A Method for Target Detection in UAV Images Based on Improved YOLO11," employ the following process: adapting the improved YOLO11 model to geometric target detection in industrial scenarios. Core optimizations include: redefining anchor box sizes using K-means clustering for small-scale geometric targets on industrial parts; introducing a lightweight attention mechanism to the native C2f module of YOLO11 to enhance target feature response; and improving target discrimination capabilities in complex backgrounds through multi-scale feature fusion. During detection, the industrial image undergoes anti-interference preprocessing (denoising and contrast enhancement), then is input into the improved YOLO11 model to obtain geometric target detection boxes. The number of detection boxes meeting the confidence threshold is counted, and the geometric measurement results are output based on a pre-defined size conversion relationship. While its core principle is "single-task independent modeling," and it has undergone basic optimization for general industrial scenarios, it struggles to achieve good adaptation in specific domains involving multi-index measurements and lacks targeted design for multi-task collaboration.

[0007] Simple fusion-based multi-task deep learning techniques, such as CN121304671A "A Method and System for Visual Defect Detection of Industrial Products Based on Multi-Task Reverse Knowledge Distillation," achieve collaborative defect detection and feature localization of industrial products through a multi-task learning framework. The core steps include: first, using a lightweight backbone network as a shared feature extractor to extract multi-scale features from the input industrial product image, outputting a general feature map; then, designing two parallel task branches: a defect detection branch and a defect localization branch, with both branches sharing the general features extracted by the backbone network; finally, constructing a multi-task total loss function, using a weighted sum with fixed weights, to jointly train the shared backbone network and the dual-task branches, with the weight coefficients preset before training and remaining unchanged during training; and outputting the defect detection and localization results in parallel, only displaying the results by overlay, without setting up consistency verification logic between tasks. Its core is "shared backbone + fixed weight fusion," but this simple fusion model is difficult to adapt well to specific domains and lacks a dynamic adjustment mechanism for specific tasks.

[0008] Therefore, overcoming the shortcomings of existing technologies is an urgent problem to be solved in the field of visual measurement technology. Summary of the Invention

[0009] The purpose of this invention is to address the shortcomings of existing technologies and provide an adaptive multi-task visual measurement method.

[0010] To achieve the above objectives, the technical solution adopted by the present invention is as follows: An adaptive multi-task visual measurement method is proposed, which employs an adaptive multi-task visual measurement network model; the adaptive multi-task visual measurement network model includes an input layer, a feature extraction layer, a task perception separation layer, a content dynamic fusion layer, and a collaborative reasoning output layer. The adaptive multi-task vision measurement method includes the following steps: S1. Image Input and Standardization Preprocessing: The input layer receives the original RGB image I0, performs preprocessing on it, and obtains the preprocessed image data I. S2. Feature Extraction: Input image data I into the feature extraction layer for feature extraction, and output feature maps at three different scales. , and This constitutes a multi-scale feature pyramid. Among them, C3 is a feature map downsampled by 8 times, C4 is a feature map downsampled by 16 times, and C5 is a feature map downsampled by 32 times; S3. Task-aware separation: Separating multi-scale feature pyramids The input task-aware separation layer extracts and enhances task-specific features through parallel digit-aware and geometry-aware branches, respectively, and outputs a digit-specific feature pyramid. and geometric features of the pyramid ; S4. Dynamic Content Integration: Integrating Digital-Specific Feature Pyramids and geometric features of the pyramid The input content is dynamically fused using a gated fusion layer, and the output is a collaborative feature pyramid. ; S5. Collaborative Reasoning Output: Output the collaborative feature pyramid. The input collaborative reasoning output layer performs digital and geometric calculations through parallel digital and geometric measurement heads, respectively, and outputs the measurement results.

[0011] Furthermore, in S1, preprocessing includes denoising, illumination normalization, and size normalization, which uniformly scales the image to a resolution of 640×640.

[0012] Furthermore, in S2, the input to the feature encoder in the feature extraction layer comprises five sequentially connected stages. ,in: : Including those connected in series Convolutional layer, first CSP module; wherein, the first CSP1 module consists of two sequentially connected layers. Composed of convolutional layers and residuals; : Including those connected in series Convolutional layer, second CSP module; wherein, the structure of the second CSP module is similar to... The first CSP module has the same structure but different parameters, and the number of output feature map channels is 128. : Including those connected in series The convolutional layer, specifically the first deformable convolutional module group DCMG, contains eight cascaded standard deformable convolutional blocks, with an output feature map of 256 channels. This output serves as the low-scale feature of the feature pyramid. That is, a feature map downsampled by 8 times; It includes a 3×3 convolutional layer connected in series, a second deformable convolutional module group DCMG, and a first dynamic feature calibration module DFCM; the second deformable convolutional module group DCMG has the same structure as the first deformable convolutional module group DCMG. The first dynamic feature calibration module (DFCM) employs a dual-branch attention collaboration mechanism, combining local structural attention and global channel attention, resulting in an output feature map with 512 channels. This output serves as the mesoscale feature of the feature pyramid. That is, a feature map downsampled by 16 times.

[0013] : Including those connected in series The network structure consists of convolutional layers, a third deformable convolutional module group (DCMG), and a second dynamic feature calibration module (DFCM). The third deformable convolutional module group (DCMG) contains four cascaded deformable convolutional blocks. The network structure of the second dynamic feature calibration module (DFCM) is similar to... The first dynamic feature calibration module (DFCM) has the same structure but different parameters, and its output feature map has 1024 channels. This output serves as the high-scale feature of the feature pyramid. That is, a feature map downsampled by 32 times.

[0014] Furthermore, in S3, the digital perception branch and the geometric perception branch process the multi-scale feature pyramid output from the feature extraction layer in parallel. ; In the digital sensing branch, channel standardization is performed first. Execute separately Convolution is used to uniformly adjust the number of channels in the feature maps at all scales to 256, resulting in... Then, multi-scale feature fusion is performed to obtain the reconstructed multi-scale features. For the three enhanced features in the multi-scale feature pyramid {M3, M4, M5}, a digital attention module is applied to strengthen the digit-specific features, and the final output is a digit-specific feature pyramid. ; In the geometric perception branch, channel standardization is performed first, and then... Execute separately Convolution is used to uniformly adjust the number of channels in the feature maps at all scales to 256, resulting in... Then, multi-scale feature fusion is performed to obtain the reconstructed multi-scale features. For the three enhanced features in the multi-scale feature pyramid {M3, M4, M5}, geometric attention modularization is applied to strengthen the geometry-specific features, and finally, a geometry-specific feature pyramid is output. .

[0015] Furthermore, in S4, the content dynamic fusion layer needs to process the digital-specific feature pyramid output by the task-aware separation layer. and geometric features of the pyramid The fusion process is as follows: In the pyramid of digital-specific features Perform global max pooling to obtain a digital globally salient vector. ; ; In the geometric-specific feature pyramid Perform global max pooling to obtain the geometrically global saliency vector. ; Then, cross-attention weight prediction is performed, and a cross-attention heatmap is calculated; the cross-attention heatmap includes the following four graphs: Digital self-assessment chart :calculate With its own vector The result is obtained by multiplying the product along each channel and summing the products along the channels. The numbers are evaluated by geometric graphs :calculate and the other vector The result is obtained by multiplying each channel dot product and summing them. Geometric self-evaluation diagram :calculate and The result is obtained by multiplying the product along each channel and summing the products along the channels. Geometric evaluation graph :calculate and The result is obtained by multiplying each channel dot product and summing them. Based on the output cross-attention heatmap, a dynamic gating weight map is generated, which includes the following two maps: Numerical Feature Weighting Map: ; Geometric feature weight map: ;in, and These are learnable positive scalar parameters at this scale, controlling the strength of self-reinforcement and cross-inhibition, respectively. Finally, gated feature fusion is performed, specifically by using the generated weight map to modulate and fuse the features using a weighted summation method. The modulation and fusion formula is as follows: Output the final collaborative feature pyramid .

[0016] Furthermore, in S5, the specific process of the digital measuring head performing digital calculations is as follows: Obtaining the Collaborative Feature Pyramid Then, input it into the digital measuring head; The detection network of the digital measurement head deepens and refines the features of the input collaborative feature pyramid, and then outputs detection boxes adapted to this slender shape. At the same time, it outputs the confidence of each detection box, and uses a non-maximum suppression algorithm to filter the output detection boxes and remove overlapping detection boxes. Next, OCR recognition is performed to obtain a string of numbers.

[0017] Finally, the numeric string is subjected to physical rule validation and length validation. The result with the highest confidence level that passes the validation is selected, its value is parsed and output.

[0018] Furthermore, in S5, the specific process of the geometric measurement head performing geometric calculations is as follows: The geometric measurement head uses the input co-feature pyramid The algorithm finds all intersection regions in the image and outputs the detection boxes of each intersection region and the confidence score of each detection box. Then, it uses a non-maximum suppression algorithm to filter the output detection boxes and finally outputs the rectangles of the detected intersection regions and their confidence scores. The number of detection boxes with confidence scores higher than the threshold is counted and output.

[0019] Furthermore, in S1, the original RGB image is the image of the three-tube tower reinforcement; the digital measuring head outputs the reinforcement size measurement result; and the geometric measuring head outputs the spacing measurement result between the reinforcement bars.

[0020] This invention also provides an adaptive multi-task vision measurement system, employing the aforementioned adaptive multi-task vision measurement method, comprising: The input module is used to receive the original RGB image I0, preprocess it, and obtain the preprocessed image data I; The feature extraction module, connected to the input module, is used to input image data I into the feature extraction layer for feature extraction and output feature maps at three different scales. , and This constitutes a multi-scale feature pyramid. Among them, C3 is a feature map downsampled by 8 times, C4 is a feature map downsampled by 16 times, and C5 is a feature map downsampled by 32 times; The task-aware separation module, connected to the feature extraction module, is used to process multi-scale feature pyramids. The input task-aware separation layer extracts and enhances task-specific features through parallel digit-aware and geometry-aware branches, respectively, and outputs a digit-specific feature pyramid. and geometric features of the pyramid ; The content dynamic fusion module, connected to the task awareness separation module, is used to integrate the digital-specific feature pyramid. and geometric features of the pyramid The input content is dynamically fused using a gated fusion layer, and the output is a collaborative feature pyramid. ; The collaborative inference output layer module, connected to the content dynamic fusion module, is used to integrate the collaborative feature pyramid. The input collaborative reasoning output layer performs digital and geometric calculations through parallel digital and geometric measurement heads, respectively, and outputs the measurement results.

[0021] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the adaptive multi-task vision measurement method as described herein.

[0022] Traditional machine vision measurement technology suffers from limitations in scene adaptability and accuracy (CN111189387A): This technology relies on manually designed fixed feature extraction algorithms and parameters, which have significant limitations in the specific field of rebar measurement: First, it has poor robustness, and the accuracy of edge extraction and contour fitting drops significantly when faced with sudden changes in lighting or slight deformation of the target; second, it has weak generalization ability, and for different types and specifications of industrial parts, it is necessary to readjust the edge detection threshold, feature fitting parameters, etc., resulting in long adaptation cycles and high costs. The feature encoder structure designed in this invention can improve the adaptability of the vision measurement system to complex industrial scenes, reduce adaptation costs, realize the collaborative processing of digital recognition and geometric measurement, and improve detection performance.

[0023] The shortcomings of basic deep learning technology (CN121121542A) are as follows: This technology revolves around a single task of "geometric object detection," and the model architecture lacks adaptability to different types of measurement tasks. Reinforcing bar measurement tasks are diverse (e.g., geometric measurements such as intersection counting, and numerical measurements such as digital scale reading), and existing technologies cannot flexibly adapt to multiple measurement needs. Secondly, although a lightweight attention mechanism is introduced to enhance feature response, it does not perform differentiated extraction of features specific to each measurement task, making it difficult to accurately capture the core features of different measurement tasks. This leads to missed detections and false detections in complex scenarios, thus affecting the accuracy of the measurement results. This invention improves the model's adaptability to multiple types of measurements in reinforcing bar tasks through a multi-task integrated architecture, enhances generalization ability, improves measurement accuracy in complex scenarios, and ensures the reliability of the results.

[0024] The fusion strategy and feature extraction deficiencies of the simple multi-task fusion technology (CN121304671A): The "fixed weight fusion" design of this technology has key shortcomings. First, the fusion strategy is rigid and cannot dynamically adjust feature weights according to the differences in the content of the input images, which easily leads to "task interference" and causes a decline in the performance of a certain task. Second, it lacks physical consistency constraints and has not established association rules for the output results of the two tasks, making it impossible to avoid contradictory results. Third, the feature extraction adaptability is insufficient, and the shared lightweight backbone network has limited ability to capture features of targets with low contrast and weak texture defects. This invention designs an adaptive dynamic fusion strategy to avoid task interference; introduces physical consistency constraints to improve the reliability of the results; and optimizes the feature extraction capability of the backbone network to achieve accurate adaptation of the model in specific scenarios.

[0025] This invention, through the design of "input layer → feature extraction layer → task perception separation layer → content dynamic fusion layer → collaborative reasoning output layer", is adapted to the three-tube tower rebar measurement scenario, realizes deep collaboration between heterogeneous tasks of digital recognition and geometric measurement, and effectively solves the problem of parameter redundancy in single-task deep models and result conflict in dual-task models. This invention is based on a feature encoder improved from CSPDarknet, which integrates deformable convolutional modules and dynamic feature calibration modules. It improves the feature extraction capability of complex scenes of steel reinforcement in three-tube towers by adapting the receptive field to deformable targets and strengthening the features of weakly textured targets through bi-branch attention. This invention provides a content dynamic fusion module based on a cross-attention gating mechanism. It extracts the global salient vector of the task through global pooling, adaptively predicts the fusion weights of the prediction space, and then combines gated weighted summation to achieve dynamic collaborative fusion of dual task features, thus solving the "task interference" defect of existing static fusion strategies.

[0026] The main applications of this invention are concentrated on the construction site of three-tube towers, for the identification of rebar parameters. One of the measurement tasks that can be performed is the length of the rebar. For example, for a rebar cage, it is necessary to know the length or width of the rebar cage. By inputting a well-taken on-site photo (the photo needs to be taken with a measuring tape as a reference object, and the length measurement must be taken from the photo), the length or width of the rebar cage can be output through the digital measuring head of this invention.

[0027] Measurement Task 2: Measuring the spacing between reinforcing bars. For example, if you need to know the spacing between the reinforcing bars in a reinforcing cage, input a well-taken on-site photo (the photo needs to be taken with a measuring tape as a reference object, and a reference object is necessary for length measurement). The digital measuring head and geometric measuring head of this invention can be used to measure the total length of the spacing between several reinforcing bars in the photo, as well as the spacing between the reinforcing bars.

[0028] Compared with the prior art, the beneficial effects of this invention are as follows: (1) Improve the adaptability to complex scenes: This invention is based on the optimization of the CSPDarknet network, retains its advantages of efficient feature extraction, and adds a deformable convolution module group and a dynamic feature calibration module. The deformable convolution adapts to deformable targets through adaptive receptive field, and the dynamic feature calibration enhances the features of weak texture targets. (2) Achieve deep collaboration and efficiency improvement between dual tasks: Construct a multi-task integrated architecture on the basis of YOLO network, reduce parameter redundancy by sharing backbone network; Innovatively design dynamic fusion and collaborative reasoning mechanism, which improves detection efficiency and avoids conflict between dual task results compared with the "single task neural network model chaining" scheme; (3) Ensure the physical reliability of measurement results: By using an innovative physical consistency loss function and digital verification mechanism, the dual-task output is constrained to conform to physical laws. Through the multi-task scheme of consistency constraint, the accuracy of measurement results is improved, and the requirements of high-precision quality inspection are met.

[0029] (4) A three-tube tower construction site was selected for rebar measurement. The percentage of samples that were qualified in the overall sample was used as the scene accuracy index. The scene accuracy of traditional machine learning algorithms was only 52.6%, which is difficult to cope with interference such as lighting and simple occlusion. The scene accuracy of the single-task deep learning general model (YOLO model) was 92.8%, and the recognition effect under different scenes was also good. The scene adaptation rate of the present invention reached 96.2%, which is 3.4 percentage points higher than the highest level of the prior art. This highlights the adaptation design of the present invention for complex scenes in the rebar recognition task of three-tube towers. It can effectively cope with various interference factors under different scenes, thereby significantly improving the measurement accuracy of the model in a specific field. Attached Figure Description

[0030] Figure 1 This is a flowchart of the adaptive multi-task vision measurement method of the present invention; Figure 2 This is a schematic diagram of the feature extraction layer structure; Figure 3 This is a schematic diagram of the Dynamic Feature Calibration Module (DFCM). Figure 4 A schematic diagram of the task-aware separation layer; Figure 5 This is a structural diagram of the digital sensing branch; Figure 6 This is a schematic diagram of the structure of the geometric perception branch; Figure 7 This is a structural diagram of the content dynamic fusion layer and the collaborative reasoning output layer. Detailed Implementation

[0031] The present invention will now be described in further detail with reference to the embodiments.

[0032] Those skilled in the art will understand that the following embodiments are for illustrative purposes only and should not be construed as limiting the scope of the invention. Where specific techniques or conditions are not specified in the embodiments, they are performed in accordance with the techniques or conditions described in the literature in the field or according to the product instructions. Materials or equipment whose manufacturers are not specified are all conventional products that can be obtained by purchase.

[0033] The core technical solution of this invention lies in constructing an adaptive multi-task visual measurement method, which employs an Adaptive Multi-task Visual Measurement Network (AMVM-Net). This network adopts an architecture of "input layer → feature extraction layer → task-aware separation layer → content dynamic fusion layer → collaborative inference output layer." Through end-to-end training, it achieves deep collaboration between two heterogeneous tasks—digit recognition and geometric measurement—and joint output under physical consistency constraints, thereby achieving integrated and automated visual measurement. The flowchart is shown below. Figure 1 As shown. The following section elaborates on the overall architecture, core module design, and training strategy: 1. Overall network architecture components: AMVM-Net is composed of four interconnected layers: Input Layer → Feature Extraction Layer → Task-Aware Separation Layer → Content Dynamic Fusion Layer → Collaborative Inference Output Layer. The hierarchical relationship is as follows: (1) Input layer: Receives the image to be detected. Before input, the image can be preprocessed (denoising → light correction → size standardization) and uniformly scaled to 640×640 resolution. The scaled image is adapted to the subsequent network input requirements.

[0034] (2) Feature extraction layer: The feature encoder is used as the backbone network to receive the image after preprocessing in the input layer and output feature maps at three different scales. , , This constitutes a multi-scale feature pyramid. ,in, This is a feature map downsampled by 8 times. For feature maps downsampled by 16 times, The feature map is downsampled by 32 times and used as input for subsequent task-oriented separation.

[0035] (3) Task-aware separation layer: Set up a task-aware pyramid structure to receive the output of the feature extraction layer. General features are extracted and enhanced through two parallel and functionally independent branches (digit perception branch and geometry perception branch), respectively, to extract and enhance the specific features required for digit recognition and geometry detection, ultimately outputting two task-specific feature pyramids: the output of the digit perception branch. Geometric perception branch output .

[0036] (4) Content dynamic fusion layer: Deploy an adaptive routing module, based on the cross-attention gating mechanism, receive two task-specific feature pyramids, predict spatial adaptive fusion weights, perform gated fusion, and output a collaborative feature pyramid. .

[0037] (5) Collaborative reasoning output layer: Configure dual-branch measurement heads (digital measurement head and geometric measurement head), receive collaborative feature pyramids in parallel, perform digital measurement and geometric measurement respectively, and output measurement results.

[0038] 2. Detailed Process Steps S1. Image Input and Standardization Preprocessing: Receive input RGB raw image For images Preprocessing steps include denoising, illumination normalization, and finally size normalization of the image. The images are uniformly scaled to a resolution of 640×640, and after preprocessing, the image data is obtained. .

[0039] S2, Feature Extraction: Preprocessed image data The input is fed into the feature encoder, which is an improvement on the CSPDarknet architecture. A multi-scale adaptive feature extractor is designed, which improves the extraction capability of deformable targets and weakly textured targets through deformable receptive fields and adaptive feature calibration. It includes 5 serially connected stages. ), each The network outputs feature maps at different scales at each stage. To adapt to the multi-scale information requirements of object detection tasks, the network selects... The final output serves as the effective feature layer, ultimately producing three feature maps of different scales, from shallow to deep, forming a multi-scale feature pyramid. The network structure diagram is as follows: Figure 2 As shown. The specific structure of each stage is as follows: : Connected in series Convolutional layer, first CSP module (Cross Stage Partial); the first CSP1 module consists of 2 It consists of two convolutional layers and residual connections. The convolutional layers and residuals are sequentially connected; : Connected in series Convolutional layer, second CSP module; wherein, the structure of the second CSP module is similar to... The first CSP module has the same structure but different parameters, and the number of output feature map channels is 128. : Connected in series The convolutional layer, specifically the first Deformable Convolution Module Group (DCMG), comprises eight cascaded standard Deformable Convolution Blocks (DCBs), with an output feature map of 256 channels. This output serves as the low-scale feature of the feature pyramid. That is, a feature map downsampled by 8 times; The structure consists of a 3×3 convolutional layer, a second deformable convolutional module group (DCMG), and a first dynamic feature calibration module (DFCM) connected in series. The second deformable convolutional module group (DCMG) has the same structure as the first deformable convolutional module group (DCMG), containing eight cascaded standard deformable convolutional blocks (DCBs). The first dynamic feature calibration module (DFCM) employs a dual-branch attention collaboration mechanism, combining local structural attention and global channel attention. The structure diagram is shown below. Figure 3 As shown. The local attention branch affects the input feature map. implement Depthwise separable convolution extracts local spatial features, and then... Convolution generates a single-channel spatial attention map Then, the Sigmoid function is applied for normalization. Global channel attention branch: This involves processing the input feature map... First, global average pooling is performed, then two fully connected layers (a first fully connected layer and a second fully connected layer) are used. The first fully connected layer reduces the dimensionality, and the second fully connected layer increases the dimensionality to generate the channel attention vector. Spatial attention map and channel attention vector To perform a collaborative fusion operation, the elements are multiplied one by one via a broadcast mechanism: ,in, Indicates will After being copied multiple times in the spatial dimension, and Element-by-element multiplication It involves fusing a joint attention map. Then, it's scaled using a learnable parameter. (Initial value 1.0) Weighted: , This is the weighted attention map. (The original input feature map is then used.) With weighted attention map Element-by-element multiplication: , This is the attention-enhanced output feature map, where ⊙ represents element-wise multiplication. The final output feature map has 512 channels and serves as the mesoscale feature of the feature pyramid. That is, a feature map downsampled by 16 times.

[0040] : Connected in series The network structure consists of convolutional layers, a third deformable convolutional module group (DCMG), and a second dynamic feature calibration module (DFCM). The third deformable convolutional module group (DCMG) contains four cascaded deformable convolutional blocks (DCBs). The network structure of the second dynamic feature calibration module (DFCM) is similar to... The first dynamic feature calibration module (DFCM) has the same structure but different parameters, and its output feature map has 1024 channels. This output serves as the high-scale feature of the feature pyramid. That is, a feature map downsampled by 32 times; S3, Task Awareness Separation: The digital perception branch and the geometric perception branch process the multi-scale feature pyramid output from the feature extraction layer in parallel. In the digital sensing branch, channel standardization is performed first. Execute separately Convolution is used to uniformly adjust the number of channels in the feature maps at all scales to 256, resulting in... .

[0041] Then, multi-scale feature fusion is performed, such as... Figure 4 As shown, the first step is top-down path aggregation (enhancing semantic information), which... By upsampling, and adjusting with 1×1 convolution... Feature concatenation is performed to obtain fused features. Subsequently, Upsampling was performed again, compared with the adjusted Feature concatenation is performed to obtain fused features. This step transfers strong semantic information from deep features to shallower layers. Then, bottom-up path aggregation (enriching detailed information) is performed: pass Convolution performs downsampling, and Feature concatenation is performed to obtain enhanced features. Then, Similarly, downsampling is performed, and... Feature concatenation is performed to obtain enhanced features. This step transmits the semantically enhanced shallow detail information back to the deeper layers, achieving bidirectional complementarity between details and semantics. Ultimately, we obtain the reconstructed multi-scale features. ,in, That is They correspond to the original The scale is 256, but each feature map incorporates information from other scales. The geometry-aware branch performs the exact same operation but uses independent convolution parameters.

[0042] Next, multi-scale task-specific attention enhancement is performed. A task-specific attention module is independently applied to each scale feature after reconstruction, and parameters are not shared across scales. In the digital perception branch, this will be done... The Digital Attention Module (DAM) is applied independently at each scale. The processing flow of the DAM is as follows: Figure 5 As shown, the input features are reduced in dimensionality by a 1×1 convolution to decrease the number of channels and reduce computational complexity. Then, a 3×3 convolution is used to extract local features related to digital textures. Next, a 1×1 convolution is used to compress the number of feature channels to 1, and a single-channel spatial attention mask is generated by the Sigmoid activation function. Finally, the single-channel attention mask is expanded to the same number of channels as the original input features through a broadcast mechanism. The expanded attention mask is multiplied element-wise with the original input features, and the result of the multiplication is added to the original input features through a residual connection to obtain the weighted features as the module output.

[0043] The digital attention module is applied independently to each of the three features in the multi-scale feature pyramid {M3, M4, M5}, ultimately outputting a digital-specific feature pyramid. .

[0044] In the geometric perception branch, it will be used for... The Geometric Attention Module (GAM) is applied independently at each scale. The GAM processing flow is as follows: Figure 6 As shown, first through Convolution reduces the number of channels in the input feature map to 1, decreasing the computational complexity of subsequent edge detection. The horizontal direction is then applied to the reduced feature map. Operator ( ) and vertical direction Operator ( This process yields horizontal and vertical edge response maps. The absolute values ​​of both edge response maps are taken, and then element-wise summed to obtain a composite edge response map, enhancing the saliency of edge features. This is achieved through 3×3 convolution. Convolution and The activation function converts the edge response map into a geometric attention mask. The geometric attention mask is then expanded to the same dimension as the original input features through a broadcast mechanism. The expanded attention mask is then multiplied element-wise with the original input features. The result of the multiplication is then added to the original input through a residual connection, and finally the weighted features are obtained as the module output.

[0045] For each of the three features in the multi-scale feature pyramid {M3, M4, M5}, a geometric attention module is applied independently, ultimately outputting a geometry-specific feature pyramid. The size is the same as that of the digital perception branch.

[0046] S4. Dynamic Content Integration: The content dynamic fusion layer needs to obtain the digital-specific feature pyramid output by the task-aware separation layer. and geometric features of the pyramid The fusion process is as follows: First, global pooling is performed on the features at each scale in the input digit-specific feature pyramid and geometry-specific feature pyramid to obtain context vectors representing the global saliency information of their respective tasks. Specifically, for the feature pyramid at each scale... scale : The pyramid of digital-specific features Perform global max pooling (GMP) to obtain a digital globally salient vector. ; for geometrically specific features in the pyramid Perform global max pooling (GMP) to obtain the geometrically global saliency vector. .

[0047] Then, cross-attention weight prediction is performed for the first feature in the feature pyramid output by the task perception layer. scale Calculate the cross-attention heatmap; the cross-attention heatmap includes the following 4 graphs: Digital self-assessment chart :calculate With its own vector The dot product of each channel is calculated and summed along the channels.

[0048] The numbers are evaluated by geometric graphs :calculate and the other vector Multiply each channel by a dot product and sum them.

[0049] Geometric self-evaluation diagram :calculate and The dot product of each channel is calculated and summed along the channels.

[0050] Geometric evaluation graph :calculate and Multiply each channel by a dot product and sum them.

[0051] Based on the output cross-attention heatmap, a dynamic gating weight map is generated, which includes the following two maps: Numerical Feature Weighting Map: ; Geometric feature weight map: .in, and The learnable positive scalar parameters at this scale ( and These are end-to-end learnable parameters that do not require manual setting of their final values. The initial values ​​of both parameters are uniformly set to 1.0. During the training process, the network model will adaptively learn based on the data distribution of the training set and the importance of image features at different scales, and then automatically obtain the optimal values ​​at each scale, thereby controlling the strength of self-reinforcement and cross-inhibition respectively.

[0052] Gated feature fusion: Using the generated weight map, features are modulated and fused using a weighted summation method. Output the final collaborative feature pyramid It will serve as the common input for subsequent dual measuring heads.

[0053] S5, Collaborative Reasoning Output: S5.1, Digital Measurement Head Inference: Obtaining the Collaborative Feature Pyramid Then, it is input into the digital measuring head. The workflow of the digital measuring head is "detection → recognition → verification"; the detection network consists of three 3×3 convolutional layers, such as... Figure 7 This system is used to deepen and refine the input collaborative feature pyramid, enhancing the discriminative power of digit region features. Based on the elongated shape of digit regions, and combining the refined collaborative features' ability to accurately represent these regions, the detection network outputs bounding boxes adapted to this elongated shape. It also outputs the confidence score for each bounding box and uses a widely used non-maximum suppression algorithm (an existing method) to filter the output bounding boxes, removing overlapping ones. The confidence score is automatically calculated and output by the digit measurement head's detection network, reflecting the probability that a candidate region is a true digit region, providing a basis for subsequent verification mechanisms.

[0054] OCR (Optical Character Recognition) recognition: The digit region is cropped according to the digit region bounding box parameters and corrected to horizontal by affine transformation; then adaptive image enhancement (such as deblurring and binarization) is performed; finally, the enhanced image patch is input into a lightweight OCR network for sequence recognition to obtain the digit string.

[0055] Physical rule verification and length verification: The identified numeric string is sent to the verification unit and verified according to the pre-set verification rules. The identified numeric string must conform to the numerical rules of the measuring tape scale, and the value must be a non-negative integer and within the measuring tape's range. The recognition results of adjacent numeric regions must show a monotonically increasing trend, conforming to the distribution logic of the measuring tape scale. The length of the numeric string is fixed. Outliers must be excluded from the recognition results. For results that fail verification, a "confidence decay" or "parameter rollback reprocessing" mechanism is triggered based on the confidence level output by the detection network. Confidence decay refers to reducing the weight of the result when the recognition confidence level is at a moderate level (e.g., 0.5 ≤ confidence level < 0.8). The parameter rollback reprocessing mechanism refers to rolling back to the OCR recognition stage when the recognition confidence level is below a threshold (e.g., confidence level < 0.5), readjusting the image enhancement parameters (e.g., increasing the deblurring intensity) and OCR network recognition parameters, and performing a second recognition of the numeric region. Finally, the result with the highest confidence level that passes the verification is selected, its value is parsed, and output.

[0056] S5.2, Geometric Measurement Head Reasoning: A simplified "detection-counting" paradigm is adopted. The geometric measurement head network consists of three 3×3 convolutional layers. These layers are used to optimize and refine the input collaborative feature pyramid, enhancing the feature discrimination of the intersection region of the measuring tape and rebar. Input collaborative feature pyramid Based on the characteristics of intersection regions, the algorithm accurately finds all intersection regions in the image and outputs the detection boxes of each intersection region and the confidence score of each detection box. Then, the widely used non-maximum suppression algorithm is used to filter the output detection boxes, and finally outputs the rectangular boxes and confidence scores of the detected intersection regions.

[0057] Count output: Directly count the number of detection boxes with a confidence level higher than the threshold (e.g., 0.5) and output an integer N.

[0058] Relevant personnel can directly use the results output by this network: depending on actual measurement needs, they can directly use the rebar length results from the digital measuring head, the rebar spacing results from the geometric measuring head, or both results simultaneously for verification. For example, to measure the rebar length of a three-tube tower, input an image of the on-site rebar length using a measuring tape as a reference; the digital measuring head can be used directly to view only the length measurement result. To identify the rebar spacing of independent columns in a three-tube tower, input an image of the on-site measured rebar spacing using a measuring tape as a reference; the digital measuring head automatically recognizes the measuring tape value and outputs the start and end lengths of the rebar arrangement area (e.g., measuring tape reading "500mm"), and the spacing between adjacent rebars using the geometric measuring head (e.g., 100mm). Relevant personnel can use the measurement results from the measuring head according to task requirements. Three-tube tower rebar quality inspection tasks can be performed based on the output rebar length and spacing.

[0059] This invention improves and extends the CSPDarknet backbone of the YOLO series of networks, constructing an integrated architecture for multi-task collaboration. Therefore, this invention utilizes the basic architecture of the YOLO network and achieves more powerful multi-task collaborative measurement capabilities than the original YOLO by adding deformable convolutions, dynamic feature calibration, task-aware separation, and dynamic content fusion.

[0060] 3. Training Strategies A training loop of "forward propagation - multi-task loss calculation - backpropagation update" is adopted, combined with a three-stage progressive training strategy to ensure stable network convergence, as detailed below: Forward propagation: For a batch of training samples (the training samples must be real-life images containing measuring tapes and steel bars of a three-tube tower, and simultaneously labeled with two types of tags: the position of the numerical scale and the coordinates of the intersection of the scales; the samples must cover actual usage scenarios and meet the training requirements of both numerical recognition and geometric counting tasks), execute the complete inference process described in Part 1, and record all intermediate results (such as prediction boxes, recognition results, quantity N, etc.).

[0061] Based on the forward feed and labeled data, calculate the multi-task loss function. Total loss. It is the weighted sum of the following items: in, It is the coefficient of digital detection loss. It is the coefficient of digit recognition loss. It is the coefficient of geometric detection loss. It is the coefficient of the fusion supervision loss. This is the coefficient for physical consistency loss; the coefficient value can be set as follows: The sum of all coefficients must be 1, which reflects the importance of different loss terms; It is digital detection loss. It is digital recognition loss, It is geometric detection loss. It employs a fusion-supervised loss mechanism, utilizing region masks generated by digital and geometric detectors as weak supervisory signals to construct a contrastive loss, thereby encouraging the fusion of weights. and It exhibits a high response in the corresponding task region. The detector head gradient needs to be truncated when calculating this loss to avoid interfering with its training. It is a loss of physical consistency. , It is the output value of the digital measuring head. This is the output value of the geometric measuring head. The training image label values ​​(in deep learning, network models need to be trained using labeled data; these labels are manually added only during training, and no further labeling is needed after training is complete). The physical consistency loss is used to constrain the outputs of the two tasks to conform to physical laws. Finally, the total loss is obtained. The weights of all network components are updated using the backpropagation algorithm.

[0062] Three-stage training strategy: To ensure stable convergence of the network, progressive training is employed: Phase 1: Feature Encoder Initialization. The feature encoder is initialized using weights pre-trained on a large image dataset (such as ImageNet).

[0063] Phase Two: Task-Aware Training. Real-world images of the steel reinforcement bars of the three-tube tower are used as training data. The feature encoder parameters are frozen, and only the task separation layer, dynamic fusion layer, and collaborative inference output layer are trained.

[0064] Phase 3: End-to-end fine-tuning. Unfreeze all parameters, perform joint fine-tuning using the full loss, and gradually increase sample difficulty using a course learning strategy to improve overall robustness.

[0065] Joint fine-tuning is a fixed term in deep learning. It refers to further training an existing pre-trained model using data from a specific task or domain to optimize the model's performance on that task. In this context, joint fine-tuning means that all models in the network participate in the training of the network.

[0066] The learning strategy involves first teaching the model to simple measurement samples, and then gradually introducing more complex and challenging samples. In the steel rebar tape measure measurement task: standard images with clear scales, sufficient lighting, proper placement, and simple backgrounds are used for initial training, allowing the model to learn basic length recognition and spacing counting. Then, challenging samples from real construction sites, such as blurred scales, dim lighting, overexposed lighting, tilted or bent tape measures, and densely packed and messy rebar cages, are gradually introduced to progressively increase the model's difficulty, making the model more stable and robust.

[0067] Using real-life images of the rebar cage from the three-tube tower project as training data, and placing a measuring tape on the rebar as a reference, the length of a certain rebar segment can be identified from the image, or the spacing between multiple rebars, or the number of rebars.

[0068] 4. Experimental verification To verify the adaptability and technical superiority of this invention for the rebar measurement scenario of three-tube towers, experiments were conducted in various scenarios. Rebar length and spacing images taken at the three-tube tower site in multiple different scenarios were used as training samples. The coefficients of the loss function during network model training were set as follows: .

[0069] This invention's performance was compared with existing technologies such as traditional machine vision and single-task deep learning. In terms of measurement accuracy, Mean Absolute Error (MAE) was used for comparison. Traditional machine vision had an average error of ±10cm for spacing measurement and ±7cm for length measurement, both exceeding the allowable range of engineering specifications. The single-task deep learning general model had an average error of ±2cm for spacing measurement and ±2cm for length measurement, meeting basic specifications, but still with significant room for improvement in accuracy. In contrast, this invention, through domain-specific adaptation and optimization, achieved an average error of only ±1cm for spacing measurement and ±2cm for length measurement, demonstrating improved accuracy compared to existing technologies.

[0070] A three-tube tower construction site was selected for rebar measurement. The percentage of samples that passed overall verification in the identified scenario was used as the scene accuracy indicator. Traditional machine learning algorithms achieved a scene accuracy of only 52.6%, struggling to handle interference from lighting conditions and simple occlusions. The single-task deep learning general model (YOLO model) achieved a scene accuracy of 92.8%, demonstrating good recognition performance across various scenarios. This invention achieves a scene adaptability rate of 96.2%, a 3.4 percentage point improvement over the highest existing technology. This highlights the invention's adaptability design for complex scenarios in three-tube tower rebar identification tasks, effectively addressing various interference factors in different scenarios and significantly improving the model's measurement accuracy in specific fields.

[0071] Simultaneously, experiments were designed to verify the synergistic effect of each module in the network model. Specifically, the deformable convolution module, dynamic fusion module, and physical consistency loss were removed from the present invention, and the changes in the overall pass rate were tested. Experimental results show that the overall pass rate of the complete scheme is 96.2%. After removing the deformable convolution module, the overall pass rate drops to 94.5%, with a performance decrease of 1.7%, indicating that the deformable convolution plays a crucial role through its adaptive receptive field and is the foundation for accurate feature extraction. After removing the dynamic fusion module and replacing it with fixed-weight fusion, the overall pass rate drops to 93.2%, with a performance decrease of 3.0%, proving that the dynamic fusion strategy can effectively avoid task interference and achieve adaptive feature optimization. After removing the physical consistency loss, the overall pass rate drops to 94.1%, with a performance decrease of 2.1%, indicating that the physical consistency constraint can ensure the logical rationality of the dual-task output and avoid contradictory results. The above results fully demonstrate that the high performance of this invention cannot be achieved by optimizing a single module, but is the result of the synergistic effect of all modules in the entire process of "input layer → feature extraction layer → task perception separation layer → content dynamic fusion layer → collaborative reasoning output layer". Each module is indispensable, which confirms the rationality and excellence of the technical solution of this invention.

[0072] The foregoing has shown and described the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The embodiments and descriptions in the specification are merely illustrative of the principles of the invention. Various changes and modifications can be made to the invention without departing from its spirit and scope, and all such changes and modifications fall within the scope of the present invention as claimed. The scope of protection of this invention is defined by the appended claims and their equivalents.

Claims

1. An adaptive multi-task visual measurement method, characterized in that, An adaptive multi-task visual measurement network model is adopted; the adaptive multi-task visual measurement network model includes an input layer, a feature extraction layer, a task perception separation layer, a content dynamic fusion layer, and a collaborative reasoning output layer; The adaptive multi-task vision measurement method includes the following steps: S1. Image Input and Standardization Preprocessing: The input layer receives the original RGB image I0, performs preprocessing on it, and obtains the preprocessed image data I. S2. Feature Extraction: Input image data I into the feature extraction layer for feature extraction, and output feature maps at three different scales. , and This constitutes a multi-scale feature pyramid. Among them, C3 is a feature map downsampled by 8 times, C4 is a feature map downsampled by 16 times, and C5 is a feature map downsampled by 32 times; S3. Task-aware separation: Separating multi-scale feature pyramids The input task-aware separation layer extracts and enhances task-specific features through parallel digit-aware and geometry-aware branches, respectively, and outputs a digit-specific feature pyramid. and geometric features of the pyramid ; S4. Dynamic Content Integration: Integrating Digital-Specific Feature Pyramids and geometric features of the pyramid The input content is dynamically fused using a gated fusion layer, and the output is a collaborative feature pyramid. ; S5. Collaborative Reasoning Output: Output the collaborative feature pyramid. The input collaborative reasoning output layer performs digital and geometric calculations through parallel digital and geometric measurement heads, respectively, and outputs the measurement results. The content dynamic fusion layer needs to obtain the digital-specific feature pyramid output by the task-aware separation layer. and geometric features of the pyramid The fusion process is as follows: In the pyramid of digital-specific features Perform global max pooling to obtain a digital globally salient vector. ; ; In the geometric-specific feature pyramid Perform global max pooling to obtain the geometrically global saliency vector. ; Then, cross-attention weight prediction is performed, and a cross-attention heatmap is calculated; the cross-attention heatmap includes the following four graphs: Digital self-assessment chart :calculate with its own vector The result is obtained by multiplying the product along each channel and summing the products along the channels. The numbers are evaluated by geometric graphs :calculate and the other vector The result is obtained by multiplying each channel dot product and summing them. Geometric self-evaluation diagram :calculate and The result is obtained by multiplying the product along each channel and summing the products along the channels. Geometric evaluation graph :calculate and The result is obtained by multiplying each channel dot product and summing them. Based on the output cross-attention heatmap, a dynamic gating weight map is generated, which includes the following two maps: Numerical Feature Weighting Map: ; Geometric feature weight map: ;in, and These are learnable positive scalar parameters at this scale, controlling the strength of self-reinforcement and cross-inhibition, respectively. Finally, gated feature fusion is performed, specifically by using the generated weight map to modulate and fuse the features using a weighted summation method. The modulation and fusion formula is as follows: Output the final collaborative feature pyramid .

2. The adaptive multi-task vision measurement method according to claim 1, characterized in that, In S1, preprocessing includes noise reduction, illumination normalization, and size normalization, which uniformly scales the image to a resolution of 640×640.

3. The adaptive multi-task vision measurement method according to claim 1, characterized in that, In S2, the input-to-feature encoder process in the feature extraction layer comprises five sequentially connected stages. ,in: : Including those connected in series Convolutional layer, first CSP module; wherein, the first CSP1 module consists of two sequentially connected layers. Composed of convolutional layers and residuals; : Including those connected in series Convolutional layer, second CSP module; wherein, the structure of the second CSP module is similar to... The first CSP module has the same structure but different parameters, and the number of output feature map channels is 128. : Including those connected in series The convolutional layer, specifically the first deformable convolutional module group DCMG, contains eight cascaded standard deformable convolutional blocks, with an output feature map of 256 channels. This output serves as the low-scale feature of the feature pyramid. That is, a feature map downsampled by 8 times; It includes a 3×3 convolutional layer connected in series, a second deformable convolutional module group DCMG, and a first dynamic feature calibration module DFCM; the second deformable convolutional module group DCMG has the same structure as the first deformable convolutional module group DCMG. The first dynamic feature calibration module (DFCM) employs a dual-branch attention collaboration mechanism, combining local structural attention and global channel attention, resulting in an output feature map with 512 channels. This output serves as the mesoscale feature of the feature pyramid. That is, a feature map downsampled by 16 times; : Including those connected in series The network structure consists of convolutional layers, a third deformable convolutional module group (DCMG), and a second dynamic feature calibration module (DFCM). The third deformable convolutional module group (DCMG) contains four cascaded deformable convolutional blocks. The network structure of the second dynamic feature calibration module (DFCM) is similar to... The first dynamic feature calibration module (DFCM) has the same structure but different parameters, and its output feature map has 1024 channels. This output serves as the high-scale feature of the feature pyramid. That is, a feature map downsampled by 32 times.

4. The adaptive multi-task vision measurement method according to claim 1, characterized in that, In S3, the digital perception branch and the geometric perception branch process the multi-scale feature pyramid output from the feature extraction layer in parallel. ; In the digital sensing branch, channel standardization is performed first. Execute separately Convolution is used to uniformly adjust the number of channels in the feature maps at all scales to 256, resulting in... Then, multi-scale feature fusion is performed to obtain the reconstructed multi-scale features. For the three enhanced features in the multi-scale feature pyramid {M3, M4, M5}, a digital attention module is applied to strengthen the digit-specific features, and the final output is a digit-specific feature pyramid. ; In the geometric perception branch, channel standardization is performed first, and then... Execute separately Convolution is used to uniformly adjust the number of channels in the feature maps at all scales to 256, resulting in... Then, multi-scale feature fusion is performed to obtain the reconstructed multi-scale features. For the three enhanced features in the multi-scale feature pyramid {M3, M4, M5}, geometric attention modularization is applied to strengthen the geometry-specific features, and finally, a geometry-specific feature pyramid is output. .

5. The adaptive multi-task vision measurement method according to claim 1, characterized in that, In S5, the specific process of digital measurement by the digital measuring head is as follows: Obtaining the Collaborative Feature Pyramid Then, input it into the digital measuring head; The detection network of the digital measurement head deepens and refines the features of the input collaborative feature pyramid, and then outputs detection boxes adapted to this slender shape. At the same time, it outputs the confidence of each detection box, and uses a non-maximum suppression algorithm to filter the output detection boxes and remove overlapping detection boxes. Next, OCR recognition is performed to obtain the numeric string; Finally, the numeric string is subjected to physical rule validation and length validation. The result with the highest confidence level that passes the validation is selected, its value is parsed and output.

6. The adaptive multi-task vision measurement method according to claim 1, characterized in that, In S5, the specific process of geometric measurement by the geometric measuring head is as follows: The geometric measurement head uses the input co-feature pyramid The algorithm finds all intersection regions in the image and outputs the detection boxes of each intersection region and the confidence score of each detection box. Then, it uses a non-maximum suppression algorithm to filter the output detection boxes and finally outputs the rectangles of the detected intersection regions and their confidence scores. The number of detection boxes with confidence scores higher than the threshold is counted and output.

7. The adaptive multi-task vision measurement method according to claim 1, characterized in that, In S1, the original RGB image is the image of the three-tube tower reinforcement; the digital measuring head outputs the reinforcement size measurement results; the geometric measuring head outputs the spacing measurement results between the reinforcement bars.

8. An adaptive multi-task vision measurement system, employing the adaptive multi-task vision measurement method according to any one of claims 1 to 7, characterized in that, include: The input module is used to receive the original RGB image I0, preprocess it, and obtain the preprocessed image data I; The feature extraction module, connected to the input module, is used to input image data I into the feature extraction layer for feature extraction and output feature maps at three different scales. , and This constitutes a multi-scale feature pyramid. Among them, C3 is a feature map downsampled by 8 times, C4 is a feature map downsampled by 16 times, and C5 is a feature map downsampled by 32 times; The task-aware separation module, connected to the feature extraction module, is used to process multi-scale feature pyramids. The input task-aware separation layer extracts and enhances task-specific features through parallel digit-aware and geometry-aware branches, respectively, and outputs a digit-specific feature pyramid. and geometric features of the pyramid ; The content dynamic fusion module, connected to the task awareness separation module, is used to integrate the digital-specific feature pyramid. and geometric features of the pyramid The input content is dynamically fused using a gated fusion layer, and the output is a collaborative feature pyramid. ; The collaborative inference output layer module, connected to the content dynamic fusion module, is used to integrate the collaborative feature pyramid. The input collaborative reasoning output layer performs digital and geometric calculations through parallel digital and geometric measurement heads, respectively, and outputs the measurement results.

9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the adaptive multi-task vision measurement method as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Industrial part size detection method based on machine vision

    CN111189387A

  • Unmanned aerial vehicle image target detection method based on improved YOLO11

    CN121121542A

  • Industrial product visual defect detection method and system based on multitask reverse knowledge distillation

    CN121304671A

  • SAR directed target detection method based on multi-scale context sensing

    CN121582554A

  • Acoustic array adaptive calibration and correction system applied to underwater moving target

    CN121763268A