Metering asset management system and method based on multi-modal image instance segmentation
The metrological asset management system, which utilizes multimodal image instance segmentation and dynamic weight fusion, overcomes the limitations of single feature detection and enables efficient and accurate detection and management of metrological assets.
Patent Information
- Application Number
- CN202511129426.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-13
- Publication Date
- 2025-11-18
AI Technical Summary
Existing technologies rely on single-feature detection in measurement asset management, which cannot comprehensively evaluate multimodal features and lacks scenario adaptability, resulting in high false negative rates and low verification efficiency.
A metrology asset management system and method based on multimodal image instance segmentation is adopted. Through a multimodal feature extraction module (appearance, quality, and on-site feature branches) and a dynamic weight fusion module, inspection reports are generated, realizing fully automatic and intelligent verification of metrology assets.
It improves the accuracy, automation level and efficiency of measurement asset inspection, reduces the missed detection rate, and adapts to the needs of different business scenarios.
Smart Images

Figure CN120976546A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of image recognition, in particular to a metering asset management system and method based on multi-modal image instance segmentation. BACKGROUND
[0002] With the rapid development of the power, oil, and natural gas industries and the increasing number of metering assets, asset operation and maintenance management has become a complex and highly efficient work. In these industries, the normal operation of metering devices (such as electricity meters, gas meters, etc.) is crucial to the stability of the entire system. Therefore, how to efficiently and accurately manage and maintain these metering assets has become an important issue in the industry.
[0003] Currently, many enterprises rely on traditional manual inspection or single image processing technology to monitor the status of metering devices. For example, some existing systems rely on simple image classification technology to inspect the appearance of metering devices, but these methods usually have limitations and cannot comprehensively identify various state characteristics of the devices, such as appearance problems, quality problems, and problems in the field environment. In addition, most of these technologies rely on static thresholds and cannot adapt to different actual scenarios, and are easily affected by factors such as light and occlusion, leading to false positives or false negatives. SUMMARY
[0004] The present application provides a metering asset management system and method based on multi-modal image instance segmentation, aiming to solve the technical problems of traditional metering asset image verification methods relying only on single feature detection, being unable to comprehensively evaluate multi-modal features, and lacking scene adaptability, resulting in high false negative rate and low verification efficiency. The technical effect of improving the accuracy, automation level, and efficiency of metering asset detection is achieved through multi-modal image instance segmentation and dynamic weight optimization.
[0005] In view of the above problems, the present application provides a metering asset management system and method based on multi-modal image instance segmentation.
[0006] In a first aspect, a metrology asset management system based on multi-modal image instance segmentation is provided, which comprises: a system setting unit configured to set a system overall architecture with an input layer, a multi-modal feature extraction module, a dynamic weight fusion module, and an output layer based on multi-modal image instances of metrology assets; a feature extraction unit configured to use an appearance feature branch in the multi-modal feature extraction module to obtain a material result, use a quality feature branch in the multi-modal feature extraction module to obtain a moire result, and use a scene feature branch in the multi-modal feature extraction module to obtain a bounding box result; a dynamic fusion unit configured to embed a weight configuration table in the dynamic weight fusion module, dynamically adjust weight values according to metrology asset business scenarios, and then perform dynamic weight fusion in combination with the material result, the moire result, and the bounding box result to generate a detection report; and a metrology asset management unit configured to perform metrology asset management based on the detection report.
[0007] In another aspect, a metrology asset management method based on multi-modal image instance segmentation is provided, which comprises: setting a system overall architecture with an input layer, a multi-modal feature extraction module, a dynamic weight fusion module, and an output layer based on multi-modal image instances of metrology assets; using an appearance feature branch in the multi-modal feature extraction module to obtain a material result, using a quality feature branch in the multi-modal feature extraction module to obtain a moire result, and using a scene feature branch in the multi-modal feature extraction module to obtain a bounding box result; embedding a weight configuration table in the dynamic weight fusion module, dynamically adjusting weight values according to metrology asset business scenarios, and then performing dynamic weight fusion in combination with the material result, the moire result, and the bounding box result to generate a detection report; and performing metrology asset management based on the detection report.
[0008] The one or more technical solutions provided in the present application have at least the following technical effects or advantages: The system overall architecture is provided with an input layer, a multi-modal feature extraction module, a dynamic weight fusion module, and an output layer due to the adoption of the multi-modal image instance based on the measured assets; the multi-modal feature extraction module includes an appearance feature branch, a quality feature branch, and a scene feature branch, the material result is obtained by using the appearance feature branch in the multi-modal feature extraction module, the moire result is obtained by using the quality feature branch in the multi-modal feature extraction module, and the frame result is obtained by using the scene feature branch in the multi-modal feature extraction module; the dynamic weight fusion module is internally provided with a weight configuration table, the weight value is dynamically adjusted according to the measured asset business scene, then, the dynamic weight fusion is performed in combination with the material result, the moire result, and the frame result to generate a detection report; the measured asset management is performed by using the detection report. The technical problem of high omission rate and low verification efficiency caused by the fact that the traditional measured asset image verification method only relies on single feature detection, cannot comprehensively evaluate multi-modal features, and lacks scene adaptability is solved, and the technical effect of improving the accuracy, automation level, and efficiency of measured asset detection through multi-modal image instance segmentation and dynamic weight optimization is achieved.
[0009] The above description is only a summary of the technical solutions of the present application, in order to more clearly understand the technical means of the present application, the specific embodiments of the present application can be implemented according to the content of the description, and in order to make the above and other purposes, features and advantages of the present application more obvious and easy to understand, the following specific embodiments of the present application are described. BRIEF DESCRIPTION OF DRAWINGS
[0010] Figure 1 A structural schematic diagram of a measured asset management system based on multi-modal image instance segmentation is provided for the embodiments of the present application.
[0011] Figure 2 A flowchart of a measured asset management method based on multi-modal image instance segmentation is provided for the embodiments of the present application.
[0012] Explanation of reference numerals: system setting unit 11, feature extraction unit 12, dynamic fusion unit 13, measured asset management unit 14. DETAILED DESCRIPTION
[0013] The present application provides a measured asset management system and method based on multi-modal image instance segmentation, which solves the technical problem of high omission rate and low verification efficiency caused by the fact that the traditional measured asset image verification method only relies on single feature detection, cannot comprehensively evaluate multi-modal features, and lacks scene adaptability, and achieves the technical effect of improving the accuracy, automation level, and efficiency of measured asset detection through multi-modal image instance segmentation and dynamic weight optimization.
[0014] The technical solutions of this application will now be clearly and completely described with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. It should be understood that this application is not limited to the exemplary embodiments described herein. All other embodiments obtained by those skilled in the art based on the embodiments of this application without creative effort are within the scope of protection of this application. It should also be noted that, for ease of description, only the parts related to this application are shown in the accompanying drawings, not all of them.
[0015] Example 1, as Figure 1 As shown in the embodiment of this application, a metering asset management system based on multimodal image instance segmentation is provided. The system includes: System setup unit 11 is used to set up the overall system architecture based on multimodal image instances of measurement assets, including an input layer, a multimodal feature extraction module, a dynamic weight fusion module, and an output layer.
[0016] Specifically, in system setup unit 11, based on multimodal image instances, efficient verification and management of metrological assets are achieved through different module combinations. The overall system architecture includes an input layer, a multimodal feature extraction module, a dynamic weight fusion module, and an output layer. These modules work together to ensure overall efficiency and accuracy. The input layer receives externally input multimodal image instances from metrological assets in actual applications, which may include images of devices such as meter boxes and smart meters. Input image formats can include common formats such as JPEG and PNG, and support different resolutions, such as 800×600 to 1920×1080. Through the input layer, image data can be received and preprocessed to ensure that subsequent processing steps are not affected by the quality of the input data. The multimodal feature extraction module is one of the core modules of the system, consisting of three main branches: appearance feature branch, quality feature branch, and on-site feature branch. Each branch processes different types of features to perform comprehensive analysis of the metrological assets in the images. The dynamic weight fusion module has a built-in weight configuration table that can dynamically adjust the weight values of each feature according to different metrology asset business scenarios (such as equipment filing, on-site inspection, equipment rotation, etc.). For example, when filing equipment, appearance features may have a higher weight, while on-site features may be more important during on-site inspection. By dynamically adjusting the weights, it can better adapt to different application scenarios, thereby achieving accurate defect detection and assessment. The output layer is used to output the generated inspection report. Through the design of the above overall architecture, fully automated and intelligent verification of metrology assets can be achieved, reducing manual intervention, improving detection accuracy and efficiency, and dynamically adjusting feature weights according to the needs of different operation and maintenance scenarios, thereby achieving higher adaptability and accuracy.
[0017] The feature extraction unit 12 is configured to use the multi-modal feature extraction module to extract appearance feature branch, quality feature branch, and scene feature branch, use the appearance feature branch in the multi-modal feature extraction module to obtain material result, use the quality feature branch in the multi-modal feature extraction module to obtain moire result, and use the scene feature branch in the multi-modal feature extraction module to obtain frame result.
[0018] Specifically, in the feature extraction unit 12, the stored multi-modal feature extraction module is used to extract various types of key feature information from the input metrology asset multi-modal image. The multi-modal feature extraction module includes an appearance feature branch, a quality feature branch, and a scene feature branch. The appearance feature branch is responsible for analyzing the appearance information of the metrology asset in the image, mainly including the integrity of the label and the classification of the material. By segmenting the label, it can be determined whether the label is complete, whether there is a missing or damaged part, and thus the appearance state of the metrology asset can be evaluated. The quality feature branch focuses on analyzing the quality problems of the image, especially the moire phenomenon and the sharpness of the image. This branch uses frequency domain analysis techniques (such as FFT) to detect possible moire in the image, and at the same time, it scores the sharpness of the image to obtain a sharpness value between 0 and 100, which helps the system to determine whether the image quality meets the analysis requirements. The scene feature branch is used to analyze the shooting angle and frame state of the image. Since the shooting of the metrology asset may be affected by the on-site environment (such as improper shooting angle or existence of obstructions around the equipment), this branch corrects the tilt angle of the image to ensure the standardization of the image data. Through the cooperative work of the three branches, the multi-modal feature extraction module can extract key information such as material result, moire result, and frame result from the input metrology asset image, providing accurate basic data for subsequent image analysis, and helping to realize intelligent management and efficient maintenance of metrology assets.
[0019] Further, the feature extraction unit 12 includes: The appearance feature branch is based on an improved model of Mask R-CNN, and outputs material result including label segmentation mask and material classification information.
[0020] In a preferred implementation, the appearance feature branch is based on an improved model built on Mask R-CNN, aiming to perform accurate appearance analysis on metrology asset images, especially label segmentation and material classification. Mask R-CNN is a classic image segmentation algorithm that combines convolutional neural networks (CNN) and region convolutional neural networks (RoIAlign) to achieve pixel-level accurate segmentation of targets in images. The appearance feature branch uses an improved version of Mask R-CNN, with a ResNet-101 backbone and a feature pyramid network (FPN) to enhance the ability to extract multi-scale features. First, the model extracts key information from the input image through multi-level convolution processing, and then performs label segmentation based on this information. The label segmentation mask generation process can accurately identify the label region on the metrology asset in the image and assign a precise mask to each label region, indicating the specific location and shape of the label. In this way, it can be determined whether the label is intact, whether there is a missing or damaged situation. In addition, the improved Mask R-CNN model also has the ability of material classification, that is, the identification of the surface material of the metrology asset. The model analyzes the image region around the label and can classify different materials of the metrology asset (such as metal, plastic, etc.), thereby generating a material result that includes label segmentation masks and material classification information, providing a comprehensive analysis of the appearance integrity and material state of the metrology asset for the entire system, ensuring the intelligent detection and management capabilities of the entire metrology asset management system.
[0021] Further, the feature extraction unit 12 includes: The quality feature branch detects moire patterns using frequency domain analysis and combines a Laplacian operator to determine a moire result including image sharpness and moire information.
[0022] In one possible implementation, the quality feature branch detects moire phenomenon in the image through frequency domain analysis techniques and calculates the sharpness of the image in combination with the Laplacian operator. Specifically, frequency domain analysis is a method of detecting moire by converting the image from the spatial domain to the frequency domain. Moire is an artifact caused by periodic interference during scanning or shooting, which is usually manifested as fine and repetitive patterns in the image. These patterns have a negative impact on image quality. In order to detect this phenomenon, the quality feature branch performs a fast Fourier transform (FFT) on the image, converting it to the frequency domain. In the frequency domain, moire phenomenon is usually manifested as frequency peaks in a specific frequency range. By identifying these peaks, moire in the image is detected. On the basis of moire detection, the Laplacian operator is applied to the spatial domain of the image to calculate the sharpness of the image. The Laplacian operator is a second-order derivative operation that can sensitively capture edge information in the image, reflecting the sharpness of the image. By applying the Laplacian operator to the image, the sharpness score of the image can be calculated. The numerical range of the sharpness score is from 0 to 100. The higher the score, the sharper the image. The lower the score, the blurrier the image. Finally, the quality feature branch will combine moire detection and sharpness score to output a moire result, which includes moire information and sharpness score of the image, providing an important basis for subsequent image processing and metrology asset analysis to ensure accurate metrology asset management.
[0023] Further, the feature extraction unit 12 includes: The on-site feature branch corrects the image tilt angle through the spatial transformation network, and outputs a frame result including angle difference value and frame anomaly detection information.
[0024] In an implementable embodiment, the on-site feature branch is constructed based on a spatial transformation network (STN) to correct the tilt angle of the image. Specifically, the STN is a deep learning model that can learn how to perform spatial transformation on the input image. The STN first analyzes the structural features of the image and automatically calculates the tilt angle of the current image according to the content of the image. Then, the STN generates a transformation matrix according to the tilt angle to perform geometric transformation on the image, eliminating or reducing the impact of the tilt. Finally, the corrected image is output. After correction, it is detected whether there is a frame abnormality, such as frame damage, obstruction or incompleteness, etc. The frame abnormality detection is realized by an image edge detection algorithm, which can accurately identify the edges in the image and judge whether there are defects. If the frame of the image has irregular morphology (for example, some parts are blocked or damaged), these abnormal areas can be accurately identified, and the frame result is generated, including the angle difference and the frame abnormality detection information. The angle difference represents the difference in tilt angle between the original image and the corrected image, which provides specific information about the geometric correction of the image, helping to understand the degree of image deviation. The frame abnormality detection information indicates whether the image has frame damage or other types of abnormalities. Through this on-site feature branch, the tilt angle of the image can be automatically corrected to ensure the correct viewing angle of the image and effectively detect frame abnormalities, thereby improving the image quality and detection accuracy in the process of measuring asset management.
[0025] Further, the feature extraction unit 12 includes: The appearance feature branch, the quality feature branch, and the on-site feature branch share the first three convolutional layers of ResNet-101; at the same time, the work order number is identified by OCR, and the measurement asset business scene is automatically matched.
[0026] In a feasible implementation, the appearance feature branch, the quality feature branch and the scene feature branch are all based on ResNet-101 as the backbone network. In order to reduce the computational overhead, the appearance feature branch, the quality feature branch and the scene feature branch share the first three convolutional layers of ResNet-101, so as to realize more efficient calculation while ensuring that each branch can extract rich underlying features such as edges, textures, colors and the like, thereby providing important support for the subsequent extraction of appearance, quality and scene features. Since the three feature branches share the same convolutional layers, they can be trained on the same network basis, avoiding repeated calculation and resource waste. In addition, through optical character recognition (OCR) technology, the work order number in the image is recognized and automatically matched with the measurement asset business scene. The OCR technology can extract text information from the image of the measurement asset and identify the work order number and other key information. The work order number is usually associated with a specific business scene (such as equipment filing, on-site inspection, equipment rotation, etc.). Through the identification of the work order number, the image can be automatically matched with the corresponding business scene, thereby providing a basis for subsequent feature weight adjustment and defect detection. For example, in the equipment filing scene, the weight of the appearance feature can be set higher, while in the on-site inspection scene, more attention can be paid to the weight of the scene feature, so as to ensure that the system can dynamically adjust the detection parameters according to different business scenes, thereby realizing more accurate measurement asset management.
[0027] Further, the feature extraction unit 12 further comprises: The inference calculation graph of the appearance feature branch, the quality feature branch and the scene feature branch in the multi-modal feature extraction module is fused by using a TensorRT optimization model. The dynamic shape inference function of the TensorRT optimization model is used to adapt to the input of multi-resolution images, and the weight is hard-coded by combining a rule engine. The initial weight reference value is allocated to the appearance feature branch, the quality feature branch and the scene feature branch according to the priority of the measurement asset business scene.
[0028] In a feasible implementation, in order to improve the inference efficiency and adaptability, the TensorRT optimization model is used to optimize each branch in the multi-modal feature extraction module. Specifically, first, the inference calculation graph of the appearance feature branch, the quality feature branch and the scene feature branch in the multi-modal feature extraction module is analyzed using TensorRT, and adjacent operations that can be combined are identified, such as convolution layers and activation layers, or combinations of multiple convolution operations. By combining these calculation layers, unnecessary calculation overhead in the model is reduced, thereby improving the inference speed. Subsequently, in order to adapt to different resolution image inputs, the dynamic shape inference function of the TensorRT model is enabled, that is, through the dynamic shape inference function, TensorRT can dynamically adjust the calculation graph of the model according to the different sizes of the input image, so that multi-resolution images can be input and inferred at the same time, ensuring that different resolution images can be effectively processed. Then, the weights in the model are hard-coded in combination with the rule engine, so as to divide the appropriate initial weight reference value for each feature branch. In this process, the rule engine sets the initial weight value for different feature branches according to the priority of the metering asset business scenario. These weight values are obtained through business rules and demand analysis, to ensure that the model can prioritize processing the most important feature branch, thereby improving decision-making efficiency in actual application.
[0029] Further, the feature extraction unit 12 further comprises: The multi-modal feature extraction module uses a cross-branch feature interaction mechanism for joint training, including shared backbone network setting, weighted loss function and parallel branch architecture, and realizes the interaction and enhancement of different branch features in multi-scale levels through a feature pyramid network; ViT is used to replace ResNet-101 as the backbone network, the multi-modal image of the metering asset is divided into blocks and converted into sequence features, and the global feature correlation is captured through a multi-head self-attention mechanism.
[0030] In an implementable embodiment, the multi-modal feature extraction module adopts a joint training strategy to achieve the effect of cross-branch feature interaction and enhanced feature learning. Specifically, in order to fully exploit the relevance between different branches, a cross-branch feature interaction mechanism is adopted for joint training, which includes shared backbone network setting, weighted loss function and parallel branch architecture, so that the features of each branch can not only learn their own tasks, but also effectively interact and fuse with the features of other branches. This interaction helps to improve the learning ability of each branch, so that different features can complement each other, thereby improving the recognition and judgment ability of complex scenes. The shared backbone network setting is that all feature branches share the first few layers of the backbone network (such as the first three convolutional layers of ResNet-101), so that feature information can be shared during model training, reducing redundant calculations and improving training speed and efficiency. The weighted loss function is to calculate the loss of the three branches by weighting, and the weight of the appearance feature branch can be 0.5, the weight of the quality feature branch can be 0.3, and the weight of the scene feature branch can be 0.2. The parallel branch architecture is to process each branch in parallel to speed up the feature extraction process, so that each branch can focus on its specific type of feature, and ensure compatibility and complementarity between different features when fused. In order to better capture multi-scale features, a feature pyramid network (FPN) is introduced in the feature extraction module. FPN can aggregate features of different scales layer by layer, thereby enhancing the perception ability of the model under multiple scales. In the feature extraction process of different branches, FPN fuses low-level and high-level features at different levels, so that each branch can obtain rich context information when processing images of different resolutions and scales, further improving the adaptability of the model to complex scenes. In addition, a visual Transformer (ViT) is introduced in the backbone network to replace the traditional ResNet-101. ViT divides the image into multiple small blocks and converts these small blocks into sequence features. Through the multi-head self-attention mechanism, ViT can effectively capture the global correlation between different regions in the image and enhance the efficiency of information transmission, further improving the feature extraction capability.
[0031] The dynamic fusion unit 13 is used for the dynamic weight fusion module to build a weight configuration table, dynamically adjust the weight value according to the metering asset business scene, and then dynamically fuse the material result, moire result and frame result to generate a detection report.
[0032] Specifically, in the dynamic fusion unit 13, the dynamic weight fusion module is built-in with a weight configuration table containing weight distribution rules for various metering asset business scenarios. Each business scenario (such as metering asset filing, equipment rotation, on-site inspection, etc.) has different feature requirements. The system will automatically call the corresponding weight configuration table according to the different business scenarios. After the input image passes through the multi-modal feature extraction module and outputs different feature results (such as material results, moire results, and frame results), the dynamic weight fusion module automatically adjusts the weight of each feature according to the current business scenario, thereby automatically assigning different weight values to the material, quality, and on-site features. Once the feature weight is dynamically adjusted, the material results, moire results, and frame results are weighted and averaged according to their corresponding weights, and the final detection report is generated based on the fused results. This detection report not only includes the detection results of each feature, but also presents the weighted scores and comprehensive detection scores of each feature, helping operation and maintenance personnel quickly identify the health status of metering assets and facilitating timely adoption of necessary maintenance measures.
[0033] The metering asset management unit 14 is used to manage metering assets based on the detection report.
[0034] Specifically, in the metering asset management unit 14, after obtaining the detection report, the detection report is used to manage metering assets, i.e., based on the comprehensive score in the detection report, the health status of assets can be accurately evaluated, and each asset is labeled as "normal", "abnormal", or "maintenance required". Operation and maintenance personnel can take appropriate management measures based on the health status evaluation results in the detection report, such as increasing the monitoring frequency of key assets, performing regular maintenance based on the detection report, or appropriately extending the detection and maintenance cycle of assets to reduce management costs, thereby achieving more refined and efficient asset management.
[0035] In summary, the metering asset management system based on multi-modal image instance segmentation provided by the embodiments of the present application has the following technical effects: The system setting unit 11 is used for setting the overall architecture of the system with an input layer, a multi-modal feature extraction module, a dynamic weight fusion module and an output layer based on the multi-modal image instance of the measurement asset; the feature extraction unit 12 is used for the multi-modal feature extraction module including an appearance feature branch, a quality feature branch and a scene feature branch, obtaining the material result by using the appearance feature branch in the multi-modal feature extraction module, obtaining the moire result by using the quality feature branch in the multi-modal feature extraction module, and obtaining the bounding box result by using the scene feature branch in the multi-modal feature extraction module; the dynamic fusion unit 13 is used for the dynamic weight fusion module with a built-in weight configuration table, dynamically adjusting the weight value according to the measurement asset business scenario, and then dynamically fusing the material result, the moire result and the bounding box result to generate a detection report; and the measurement asset management unit 14 is used for managing the measurement asset by using the detection report. Through the above steps, the technical problem that the traditional measurement asset image checking method only relies on single feature detection, cannot comprehensively evaluate multi-modal features and lacks scene adaptability, resulting in high missed detection rate and low checking efficiency is solved, and the technical effect of improving the accuracy, automation level and efficiency of measurement asset detection through multi-modal image instance segmentation and dynamic weight optimization is achieved.
[0036] In the embodiment two, based on the same inventive concept as the measurement asset management system based on multi-modal image instance segmentation in the foregoing embodiments, as shown in the embodiment two, the embodiment of the present application provides a measurement asset management method based on multi-modal image instance segmentation, which comprises the following steps: Figure 2 The embodiment two, based on the same inventive concept as the measurement asset management system based on multi-modal image instance segmentation in the foregoing embodiments, as shown in the embodiment two, the embodiment of the present application provides a measurement asset management method based on multi-modal image instance segmentation, which comprises the following steps: The overall architecture of the system with an input layer, a multi-modal feature extraction module, a dynamic weight fusion module and an output layer is set based on the multi-modal image instance of the measurement asset; the multi-modal feature extraction module includes an appearance feature branch, a quality feature branch and a scene feature branch, the material result is obtained by using the appearance feature branch in the multi-modal feature extraction module, the moire result is obtained by using the quality feature branch in the multi-modal feature extraction module, and the bounding box result is obtained by using the scene feature branch in the multi-modal feature extraction module; the dynamic weight fusion module has a built-in weight configuration table, the weight value is dynamically adjusted according to the measurement asset business scenario, then the dynamic weight fusion is performed in combination with the material result, the moire result and the bounding box result to generate a detection report; and the measurement asset is managed by using the detection report.
[0037] Further, the method comprises: The appearance feature branch is based on an improved model of Mask R-CNN, and outputs the material result including a label segmentation mask and material classification information.
[0038] Further, the method comprises: The quality feature branch adopts frequency domain analysis to detect moire, and combines a Laplacian operator to determine a moire result including image sharpness and moire information.
[0039] Further, the method comprises: The on-site feature branch corrects an image tilt angle through a spatial transformation network, and outputs a frame result including an angle difference value and frame anomaly detection information.
[0040] Further, the method comprises: The appearance feature branch, the quality feature branch, and the on-site feature branch share the first three convolutional layers of a ResNet-101; meanwhile, an order number is recognized through OCR, and a metrology asset business scenario is automatically matched.
[0041] Further, the method comprises: A TensorRT optimization model is used to perform layer fusion on inference calculation graphs of the appearance feature branch, the quality feature branch, and the on-site feature branch in the multi-modal feature extraction module; a dynamic shape inference function of the TensorRT optimization model is used to adapt to multi-resolution image input, and a rule engine is used to hard-code weights, and initial weight reference values are allocated to the appearance feature branch, the quality feature branch, and the on-site feature branch according to priorities of metrology asset business scenarios.
[0042] Further, the method comprises: The multi-modal feature extraction module joint training adopts a cross-branch feature interaction mechanism, including a shared backbone network setting, a weighted loss function, and a parallel branch architecture, and realizes interactive enhancement of different branch features at multi-scale levels through a feature pyramid network; a ViT is used to replace a ResNet-101 as a backbone network, multi-modal images of metrology assets are divided and converted into sequence features, and global feature correlations are captured through a multi-head self-attention mechanism.
[0043] Any step of the above-described method can be stored as computer instructions or programs in an unrestricted computer memory, and can be called and recognized by an unrestricted computer processor to implement any method in the embodiments of the present application, and no redundant limitations are made herein.
[0044] Further, the above-described first or second possibility does not only represent a sequential relationship, but also can represent a certain specific concept, and / or refers to a single or all selection between multiple elements. Obviously, those skilled in the art can make various modifications and variations to the present application without departing from the scope of the present application. Thus, if these modifications and variations of the present application belong to the scope of the present application and its equivalent technologies, the present application intends to include these modifications and variations.
Claims
1. A metering asset management system based on multimodal image instance segmentation, characterized in that, include: The system setup unit is used to set up the overall system architecture based on multimodal image instances of measurement assets, including an input layer, a multimodal feature extraction module, a dynamic weight fusion module, and an output layer. The feature extraction unit is used in the multimodal feature extraction module, which includes an appearance feature branch, a quality feature branch, and a field feature branch. The appearance feature branch in the multimodal feature extraction module is used to obtain the material result, the quality feature branch in the multimodal feature extraction module is used to obtain the moiré pattern result, and the field feature branch in the multimodal feature extraction module is used to obtain the border result. The dynamic fusion unit is used to dynamically adjust the weight values according to the measurement asset business scenario by using the built-in weight configuration table of the dynamic weight fusion module. Then, it combines the material results, moiré pattern results, and border results to perform dynamic weight fusion and generate a test report. The measurement asset management unit is used to manage measurement assets based on the aforementioned test report.
2. The metering asset management system based on multimodal image instance segmentation as described in claim 1, characterized in that, The feature extraction unit includes: The appearance feature branch is based on an improved model of Mask R-CNN, and outputs material results including label segmentation masks and material classification information.
3. The metering asset management system based on multimodal image instance segmentation as described in claim 2, characterized in that, The feature extraction unit includes: The quality feature branch uses frequency domain analysis to detect moiré patterns, and combines the Laplacian operator to determine the moiré pattern result, which includes image sharpness and moiré pattern information.
4. The metering asset management system based on multimodal image instance segmentation as described in claim 2, characterized in that, The feature extraction unit includes: The on-site feature branch corrects the image tilt angle through a spatial transformation network and outputs a border result that includes angle difference and border anomaly detection information.
5. The metering asset management system based on multimodal image instance segmentation as described in claim 4, characterized in that, The feature extraction unit includes: The appearance feature branch, quality feature branch, and field feature branch share the first three convolutional layers of ResNet-101. At the same time, the work order number is identified by OCR, and the measurement asset business scenario is automatically matched.
6. The metering asset management system based on multimodal image instance segmentation as described in claim 5, characterized in that, The feature extraction unit also includes: The TensorRT optimization model is used to perform layer fusion on the inference computation graphs of the appearance feature branch, quality feature branch, and field feature branch in the multimodal feature extraction module; The TensorRT-optimized model's dynamic shape inference function adapts to multi-resolution image input. At the same time, the weights are hard-coded using a rule engine, and initial weight benchmark values are assigned to the appearance feature branch, quality feature branch, and on-site feature branch according to the priority of the asset measurement business scenario.
7. The metering asset management system based on multimodal image instance segmentation as described in claim 6, characterized in that, The feature extraction unit also includes: The multimodal feature extraction module adopts a cross-branch feature interaction mechanism for joint training, including a shared backbone network setting, a weighted loss function, and a parallel branch architecture. The feature pyramid network is used to enhance the interaction of features from different branches at multiple scale levels. Using ViT instead of ResNet-101 as the backbone network, the multimodal images of the quantitative assets are divided into blocks and converted into sequential features, and the global feature associations are captured through a multi-head self-attention mechanism.
8. A method for managing measurement assets based on multimodal image instance segmentation, characterized in that, The method is executed by the metering asset management system based on multimodal image instance segmentation as described in any one of claims 1 to 7, comprising: Based on multimodal image instances of measurable assets, a system architecture is set up with an input layer, a multimodal feature extraction module, a dynamic weight fusion module, and an output layer. The multimodal feature extraction module includes an appearance feature branch, a quality feature branch, and a field feature branch. The appearance feature branch in the multimodal feature extraction module is used to obtain the material result, the quality feature branch in the multimodal feature extraction module is used to obtain the moiré pattern result, and the field feature branch in the multimodal feature extraction module is used to obtain the border result. The dynamic weight fusion module has a built-in weight configuration table, which dynamically adjusts the weight values according to the asset measurement business scenario. Then, it combines the material results, moiré pattern results, and border results to perform dynamic weight fusion and generate a test report. The aforementioned test report is used for metrological asset management.