Inspection agent collaborative awareness system based on semantic driving
By using a semantically driven collaborative perception system for inspection agents, the problems of insufficient semantic alignment accuracy and inadequate dynamic framing guidance in multi-agent collaborative inspections are solved. This enables timely acquisition of high-value information and accurate diagnosis of equipment status, thereby improving the collaborative efficiency and accuracy of the inspection system.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- 西安圣瞳科技有限公司
- Filing Date
- 2026-01-26
- Publication Date
- 2026-05-05
AI Technical Summary
Existing inspection technologies suffer from insufficient semantic alignment accuracy and inadequate dynamic framing guidance in multi-agent collaborative inspection and cross-source information fusion, resulting in untimely and inaccurate acquisition of high-value information.
A semantically driven collaborative perception system for inspection agents is adopted. The prototype mapping module generates a semantic prototype set and mapping table for inspection tasks. Combined with the semantic graph module, pixel-level visual feature matching is performed. Sparse selection and directional mutual transmission of semantic supply and demand coupling are performed. Position-level semantic attention fusion and measurable sensitivity recalibration are carried out to generate a fusion probability map and instance table, thereby realizing equipment health status diagnosis and fault detection.
It enables the identification and directional transmission of high-value candidate regions among multiple inspection agents, ensuring the accuracy and reliability of semantic feature alignment and fusion results across agents, and improving the real-time performance and precision of equipment status diagnosis.
Smart Images

Figure CN121982609A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of intelligent inspection technology, and in particular to a semantically driven collaborative perception system for intelligent inspection agents. Background Technology
[0002] In the field of industrial equipment operation and maintenance and inspection, intelligent inspection agents are widely used in tasks such as equipment component condition detection, defect identification, and surface quality assessment. Conventional methods typically rely on pre-set detection models and task lists. The intelligent inspection agent collects on-site image and video data, performs feature extraction and target recognition after unified preprocessing, and combines the category information in the task order with the detection standards to complete the status determination. In the information collection and processing process, these methods are mostly based on static matching of equipment category, defect type, and related monitoring indicators. They use visual feature comparison, template detection, and rule-based judgment to identify and record the equipment operating status, generating inspection reports to support maintenance decisions.
[0003] However, conventional methods have limitations in multi-agent collaborative inspection and cross-source information fusion. On the one hand, their data matching is mostly based on the local features and static rules of a single agent, lacking unified encoding and high-precision alignment of multi-source semantic information, which limits real-time semantic communication and supplementation between multiple agents. On the other hand, in terms of task-driven and framing guidance, conventional methods are unable to couple the strength of semantic demand with the strength of evidence information for dynamic perspective planning, resulting in the acquisition of some high-value information not being timely or accurate enough. Summary of the Invention
[0004] In view of the aforementioned existing problems, the present invention is proposed.
[0005] Therefore, this invention provides a semantically driven collaborative perception system for inspection agents, which solves the problems of insufficient semantic alignment accuracy and inadequate dynamic framing guidance among multiple inspection agents.
[0006] To solve the above-mentioned technical problems, the present invention provides the following technical solution: This invention provides a semantically driven collaborative perception system for inspection intelligent agents, comprising, The prototype mapping module collects information on the categories of inspection targets and the monitoring indicators of the inspection targets, and generates a semantic prototype set and a semantic mapping table for inspection tasks through semantic transformation. The semantic graph module collects inspection image data and inspection video data through the inspection agent, and performs pixel-level visual feature matching based on the semantic prototype set of the inspection task to generate semantic confidence graph and semantic request graph. The coupled mutual transmission module, under the joint constraints of the semantic confidence graph and the semantic request graph, performs sparse selection and directional mutual transmission of semantic supply and demand coupling, and generates a multi-source sparse semantic feature queue. The fusion relabeling module performs position-level semantic attention fusion and measurable sensitivity relabeling on a multi-source sparse semantic feature queue, generating a fusion probability map and a fusion instance table. The diagnostic framing module performs equipment health status diagnosis and fault detection based on the fusion instance table, and generates a diagnostic list and inspection framing instructions.
[0007] As a preferred embodiment of the semantic-driven collaborative perception system for inspection intelligent agents described in this invention, the inspection target category information includes equipment component categories, defect types, surface quality problems, and structural problems. Each inspection target category information contains a unique category identifier and a complete text description. After collection, the information is standardized and organized according to the unique category identifier to generate an inspection target category information set. Based on the inspection target category information set, the monitoring indicators of the inspection targets are collected one by one. Establish a one-to-one correspondence between the inspection target category information and the monitoring indicators of the inspection target, and generate a set of corresponding inspection target categories and monitoring indicators.
[0008] As a preferred embodiment of the semantic-driven collaborative perception system for inspection intelligent agents described in this invention, the semantic conversion includes encoding the textual description of each inspection target category information using a text encoder to generate a text semantic embedding vector. Samples corresponding to the inspection target category information are extracted from historical inspection image data and historical inspection video data. Visual features are extracted using a visual encoder to generate visual embedding vectors. By fusing textual semantic embedding vectors and visual embedding vectors, a semantic prototype vector of the inspection target category is generated; The monitoring indicators of the inspection targets are encoded using the same text and visual encoding as the category information of the inspection targets, and then fused to generate semantic prototype vectors of the monitoring indicators. The semantic prototype vectors of inspection target categories and monitoring indicator semantic prototype vectors are aggregated to generate a semantic prototype set of inspection tasks. Based on the one-to-one correspondence between the inspection target categories and the monitoring indicator sets of the inspection targets, a semantic mapping table of inspection tasks is generated.
[0009] As a preferred embodiment of the semantic-driven collaborative perception system for inspection agents described in this invention, the steps for generating a semantic confidence map and a semantic request map by performing pixel-level visual feature matching based on the semantic prototype set of the inspection task are as follows: Under a unified preprocessing standard, the inspection image data and inspection video data are sequentially subjected to denoising, resolution standardization, and contrast enhancement. Visual embedding vectors are calculated for pixels in the preprocessed inspection image data and inspection video data. The semantic similarity is obtained by calculating the cosine similarity between the visual embedding vector of the pixel and the semantic prototype vector in the semantic prototype set of the inspection task, and taking the maximum value. The text readability score is obtained by performing optical character recognition confidence normalization on the pixel neighborhood, and the image blur score is obtained by monotonic mapping using spatial gradient class sharpness metric. By positively accumulating and suppressing ambiguity in the semantic similarity and text readability scores through imaging ambiguity scoring, a comprehensive semantic confidence score is generated. The semantic confidence map is generated by filling the two-dimensional coordinate grid with the comprehensive semantic confidence value according to the pixel coordinate position, and the semantic request map is generated by filling the same pixel coordinate position with the pixel value of one minus the comprehensive semantic confidence value.
[0010] As a preferred embodiment of the semantic-driven collaborative perception system for inspection agents described in this invention, the sparse selection and directional mutual transmission of semantic supply and demand coupling comprises the following steps: Regions with strong supply and strong demand are identified based on semantic confidence maps and semantic request maps. Candidate regions are divided according to fixed grid and connectivity rules, and the domain confidence integral, region request mean, and transmission cost of the candidate regions are calculated. Based on the regional confidence score, the regional request mean, and the transmission cost, a single fractional scoring method is used to evaluate the value of candidate regions and calculate their benefit scores. The benefit scores are sorted from high to low, and a second screening is performed to obtain the candidate regions to be adopted. The candidate regions to be adopted are cropped into the smallest recognizable image blocks and bound to compact embeddings and initial semantic category judgments, along with timestamps, pixel coordinate ranges, and source inspection agent identifiers; For each adopted candidate region, a sub-region with the same pixel coordinate range is extracted from the semantic request graph of each inspection agent other than the source inspection agent, and the average region request value is calculated. The inspection agents with a regional request average greater than zero are grouped into a target inspection agent set, sorted from high to low according to the regional request average, and then included in the sending list in order according to the time window bandwidth budget and the maximum sending volume limit. Targeted mutual transmission is performed according to the sending list, and a sending record is generated at the same time.
[0011] As a preferred embodiment of the semantic-driven collaborative perception system for inspection agents described in this invention, the multi-source sparse semantic feature queue includes a minimum identifiable image block, compact embedding, initial semantic category judgment, and transmission record.
[0012] As a preferred embodiment of the semantically driven collaborative perception system for inspection agents described in this invention, the location-level semantic attention fusion and measurable sensitivity recalibration steps are as follows: Spatially align the multi-source sparse semantic feature queue on a unified reference pixel grid. After spatial alignment, a probability graph of the source inspection agent and descriptive information of the source inspection agent are established. Based on the smallest recognizable image patch in the multi-source sparse semantic feature queue, a candidate instance mask is generated for the smallest recognizable image patch; Calculate the average semantic confidence of the source inspection agent within the candidate instance mask, and generate prior attention weights according to the normalized relation; The source inspection agent probability map is weighted and synthesized according to the prior attention weight to generate an initial fusion probability expression, and the measurable sensitivity is measured on the candidate instance mask pixel set by the leave-one source inspection agent sensitivity verification method. The prior attention weights are shrunk and normalized based on the measurable sensitivity to obtain the recalibrated attention weights, and the fusion probability value is calculated according to the recalibrated attention weights.
[0013] As a preferred embodiment of the semantic-driven collaborative perception system for inspection agents described in this invention, the fusion probability map and fusion instance table include: filling the fusion probability map with fusion probability values according to the reference pixel grid based on pixel coordinates; and performing binarization, connectivity analysis, and geometric refinement on the generated fusion probability map to form the fusion instance table.
[0014] As a preferred embodiment of the semantic-driven collaborative perception system for inspection intelligent agents described in this invention, the health status diagnosis and fault detection of the execution equipment includes establishing an instance time series based on a fusion instance table within an adjacent time index, calculating normalized diagnostic indicators and uncertainties, and generating grade labels and diagnostic lists based on the comparison results of diagnostic indicators.
[0015] As a preferred embodiment of the semantic-driven collaborative perception system for inspection agents described in this invention, the inspection framing instruction includes generating corresponding inspection framing instructions for the inspection agent on the candidate viewpoint set of the inspection agent based on the semantic request graph and the diagnostic list.
[0016] The beneficial effects of this invention are as follows: by performing sparse selection and directional mutual transmission of semantic supply and demand coupling under the dual constraints of semantic confidence graph and semantic request graph, the identification of high-value candidate regions and directional transmission across inspection agents are realized; by performing position-level semantic attention fusion and measurable sensitivity recalibration on a unified reference pixel grid, the spatial alignment, confidence optimization and uncertainty quantification of multi-source sparse semantic features are realized, ensuring the accuracy and reliability of the fusion results. Attached Figure Description
[0017] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the following description of the embodiments will be briefly introduced. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0018] Figure 1 This is a schematic diagram of a semantically driven collaborative perception system for inspection agents.
[0019] Figure 2 The flowchart for generating semantic confidence graphs and semantic request graphs.
[0020] Figure 3 This is a flowchart of sparse selection and directed exchange for semantic supply and demand coupling.
[0021] Figure 4 A flowchart for position-level semantic attention fusion and measurable sensitivity recalibration. Detailed Implementation
[0022] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings.
[0023] Many specific details are set forth in the following description in order to provide a full understanding of the invention. However, the invention may also be practiced in other ways different from those described herein, and those skilled in the art can make similar extensions without departing from the spirit of the invention. Therefore, the invention is not limited to the specific embodiments disclosed below.
[0024] Secondly, the term "one embodiment" or "embodiment" as used herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The phrase "in one embodiment" appearing in different places in this specification does not necessarily refer to the same embodiment, nor is it a single or selective embodiment that is mutually exclusive with other embodiments.
[0025] Reference Figures 1-4 This is one embodiment of the present invention, which provides a semantically driven collaborative perception system for inspection agents, comprising the following steps: The prototype mapping module collects information on the categories of inspection targets and the monitoring indicators of the inspection targets, and generates a semantic prototype set and a semantic mapping table for inspection tasks through semantic transformation.
[0026] Furthermore, at the beginning of the inspection task, inspection target category information is collected from the inspection task work order, including equipment component category, defect type, surface quality problem and structural problem, that is, all object categories that need to be inspected or monitored. Equipment component category includes but is not limited to pump body, pipe and valve, defect type includes but is not limited to crack, corrosion and wear, surface quality problem includes but is not limited to oil stain, scratch and weld defect, and structural problem includes but is not limited to misalignment, detachment and crack, etc., and ensure that each inspection target category information includes a unique category identifier (e.g. equipment number or UUID) and a complete text description in the inspection task. After collection, all inspection target category information is standardized and organized according to the unique category identifier to generate an inspection target category information set. Among them, the inspection task work order refers to a document or electronic file used in the process of industrial or equipment maintenance to indicate the content, objectives, task standards and related information of the inspection task.
[0027] After standardizing and organizing the information set of inspection target categories, the monitoring indicators for each inspection target are determined and collected one by one based on the information set of inspection target categories. The names, quantification methods, units of measurement and value ranges of the monitoring indicators are clarified, and a unique mapping relationship is established. Specifically, monitoring indicators for equipment component categories include, but are not limited to, dimensional measurements (e.g., length, diameter, and thickness), surface quality inspections (e.g., surface roughness, wear depth), and performance evaluations (e.g., flow rate, pressure). Monitoring indicators for defect types include, but are not limited to, crack length, corrosion area ratio, and wear depth. Monitoring indicators for surface quality issues include, but are not limited to, weld joint quality, oil stain coverage area ratio, and scratch length. Monitoring indicators for structural issues include, but are not limited to, misalignment deviations, detachment area, and crack length and depth. Each monitoring indicator must have a unique correspondence with the corresponding equipment component category, defect type, surface quality issue, or structural issue, and detailed measurement standards and data formats must be provided according to the inspection task requirements. Each inspection target category information is mapped to its corresponding inspection target monitoring indicator to form a "target-monitoring indicator" relationship, and a set of corresponding inspection target categories and inspection target monitoring indicators is generated.
[0028] Furthermore, based on the set of corresponding inspection target categories and their monitoring indicators, semantic conversion technology is used to transform each inspection target category and its corresponding monitoring indicators into a computable and matchable semantic representation. Specifically, a text encoder (such as BERT, GPT, etc.) is used to encode the description of each inspection target category, transforming the text description into a text semantic embedding vector. For example, the text description of a crack is "a crack appears on the surface of the equipment, which may lead to structural failure". The text encoder generates a text semantic embedding vector by analyzing the words and context relationships in the text description. The text semantic embedding vector can fully express the semantic features of the inspection target category and can capture the key meaning in the text description. Visual encoders (such as ResNet, Vision Transformer, etc.) are used to extract visual features of the inspection target category from historical inspection images or video data. For example, for crack targets, the visual encoder analyzes the crack features in historical inspection images, including crack length features, crack morphology features, and crack distribution features. It then transforms the crack length features, morphology features, and distribution features into visual embedding vectors. The visual embedding vectors contain all the visual features of the crack in the historical inspection images and can accurately describe the appearance features of the crack, including the shape, size, and location of the crack. Next, the text semantic embedding vector and the visual embedding vector are fused to generate the semantic prototype vector of the inspection target category. The fusion method can be weighted average, concatenation or multilayer perceptual fusion network, etc., with the aim of ensuring that the text description and visual features of the inspection target category can be effectively combined. Similarly, text encoders and visual encoders are used to semantically transform the monitoring indicators of each inspection target to obtain the semantic prototype vector of the monitoring indicators. The semantic prototype vectors of each inspection target category and the semantic prototype vectors of the monitoring indicators are aggregated to generate a semantic prototype set for inspection tasks. Based on the mapping relationship between each inspection target category and its monitoring indicator in the corresponding set of inspection target categories and monitoring indicators, the semantic prototype vectors of each inspection target category and the semantic prototype vectors of the monitoring indicators are matched to generate a semantic mapping table for inspection tasks.
[0029] The semantic graph module collects inspection image data and inspection video data through the inspection agent, and performs pixel-level visual feature matching based on the semantic prototype set of inspection tasks to generate semantic confidence graphs and semantic request graphs.
[0030] Furthermore, during the inspection task, the inspection agent collects inspection image data and inspection video data through its own configured video equipment, including but not limited to scene information such as the appearance of equipment components, surface quality and defect types. The inspection image data is a single-frame pixel matrix with timestamp, resolution, encoding format and camera identifier, and the inspection video data is a sequence of pixel frames arranged in chronological order with frame rate, duration, encoding format and camera identifier. Denoising, resolution standardization, and contrast enhancement are performed on inspection image data and inspection video data under a unified preprocessing standard. Pixel-level visual feature matching is performed on the preprocessed inspection image data and inspection video data based on the semantic prototype set of the inspection task. Pixel coordinates Defined as the pixel coordinates in the preprocessed inspection image data, or the time index of the preprocessed inspection video data. Pixel coordinates in an image frame (the set of pixel coordinates is denoted as...) For pixel coordinates Calculate the local visual embedding vector of a pixel. (Generated by the visual encoding process), and using the semantic prototype vectors in the inspection task semantic prototype set as semantic references, the visual embedding vectors of pixel coordinates are compared with each semantic prototype vector in the inspection task semantic prototype set. The semantic similarity is obtained by taking the maximum value after calculating the cosine similarity. To ensure readability for text-related monitoring indicators, a text readability scoring channel is enabled for pixel neighborhoods belonging to text reading monitoring indicators in the inspection task semantic mapping table. Specifically, firstly, a publicly available text detection method, such as a deep convolutional network-based method, is used within the pixel neighborhood to locate rectangular text regions containing readings. Secondly, a text recognition method, such as a convolutional recurrent neural network or an open-source text recognition tool, is used within the detected text regions to obtain the reading content and simultaneously generate recognition confidence scores. Next, the recognition confidence scores of all text regions are weighted and averaged according to the pixel area of the text regions to obtain candidate reading signals. Finally, a unified dimensional mapping rule is applied to the candidate reading signals and the range is restricted to 0 to 1 to obtain the text readability score. The text readability scoring channel is disabled for pixel neighborhoods not belonging to text reading monitoring indicators in the inspection task semantic mapping table. Among them, the unified dimension mapping rule is as follows: when the confidence dimension of optical character recognition is 0 to 1, the candidate reading signal is directly used; when the confidence dimension of optical character recognition is 0 to 100, the candidate reading signal is scaled proportionally; when the confidence dimension of optical character recognition is inconsistent with the distribution of the historical annotation set and there is cross-scale, the 5th percentile and 95th percentile of the historical annotation set are used as the endpoints of the linear mapping for linear mapping, and the values below the 5th percentile are set to 0 and the values above the 95th percentile are set to 1. The image blur score is obtained by monotonic mapping using a spatial gradient-based sharpness metric. Specifically, the spatial gradient magnitudes in the horizontal and vertical directions are calculated within the pixel neighborhood, and edge-preserving median filtering is performed to suppress noise. The mean square of the gradient magnitudes is then used as the sharpness metric. Within the same frame, based on the distribution of the sharpness metric, it is linearly mapped from the 5th to the 95th percentile to the interval between 0 and 1 and truncated to obtain the sharpness score. The sharpness scores are then complemented to obtain the image blur score, with a higher image blur score indicating a higher degree of blur. For each pixel in the pixel set By positively accumulating and suppressing the semantic similarity and text readability scores through imaging blurriness scores, the comprehensive semantic confidence score is calculated and expressed as: ; in, Indicates the time index is The pixel coordinate index in the image frame is The comprehensive semantic confidence score, with example values ranging from 0 to 1, is used to express the degree of confidence that a pixel belongs to the target semantics of the inspection task semantic prototype set in the semantic confidence map. Indicates the time index is The pixel coordinate index in the image frame is The semantic similarity comes from and The normalized result of the maximum cosine similarity, with example values ranging from 0 to 1, is used to express the degree of fit between the pixel and the semantic prototype set of the inspection task. Indicates the time index is The pixel coordinate index in the image frame is The text readability score is derived from the monotonic normalization result of the optical character recognition confidence score. The example value ranges from 0 to 1 and is used to measure the degree of recognizability related to text-related monitoring indicators. Indicates the time index is The pixel coordinate index in the image frame is The imaging blur score is derived from the monotonic mapping result of the sharpness measure based on the spatial gradient. The example value ranges from 0 to 1. An increase in the value indicates an increase in the degree of blur. It is used as a blur suppression factor in the comprehensive semantic confidence. Indicates the time index is The pixel coordinate index in the image frame is Visual embedding vectors, The first in the semantic prototype set of inspection tasks A semantic prototype vector.
[0031] Furthermore, in the time index The set of pixels in an image frame Above, the coordinates of each pixel The comprehensive semantic confidence score is filled into a two-dimensional coordinate grid according to the pixel coordinate position to generate a semantic confidence map, and each pixel coordinate is... The pixel values of the semantic confidence score minus one are filled into a two-dimensional coordinate grid according to the same pixel coordinate position to generate a semantic request map. The semantic confidence map and the semantic request map, which are arranged continuously by time index, are time aligned according to the timestamp order.
[0032] It should be noted that after preprocessing, the target set for pixel-level visual feature matching, the measurement caliber, unit of measurement, and value range of the monitoring indicators are limited according to the semantic mapping table of the inspection task. The limited target set only includes the equipment component categories, defect types, surface quality problems, and structural problems registered in the semantic mapping table of the inspection task. A text readability scoring channel is enabled for pixel neighborhoods belonging to text reading monitoring indicators in the semantic mapping table of the inspection task, and disabled for pixel neighborhoods not belonging to text reading monitoring indicators, ensuring consistency between the semantic confidence map calculation and the monitoring indicator type. A pixel-level constraint mask is constructed based on the field of view constraints, shape constraints, and minimum area constraints of each monitoring indicator in the semantic mapping table of the inspection task. Specifically, this is done based on the camera corresponding to the monitoring indicator. Based on the installation angle range, camera imaging geometry parameters, and effective imaging distance interval, observable areas are marked on a reference pixel grid to obtain an initial pixel set. This initial pixel set is then compared with geometric templates (e.g., rectangles, circles, and cracks) of equipment components or defects, retaining only pixel regions similar to the templates. Connectivity analysis is performed on the compared pixel regions to remove small, discrete regions, forming continuous and stable effective regions. Mask markers are generated on these effective regions, and the pixel set formed by these mask markers serves as a pixel-level constraint mask. Semantic similarity and comprehensive semantic confidence are calculated within the pixel-level constraint mask, and calculations are skipped outside the mask, ensuring that the semantic confidence map only covers effective regions consistent with the inspection task's semantic mapping table.
[0033] The coupled mutual transmission module, under the joint constraints of the semantic confidence graph and the semantic request graph, performs sparse selection and directional mutual transmission of semantic supply and demand coupling, and generates a multi-source sparse semantic feature queue.
[0034] Furthermore, the semantic confidence graph is used to quantify the semantic matching reliability of pixel regions with the inspection target category and monitoring indicators. High-value regions in the semantic confidence graph represent reliable semantic evidence, belonging to "supply". The semantic request graph is used to quantify the degree of observation demand for pixel regions from the current perspective. High-value regions in the semantic request graph represent the demand for semantic evidence, belonging to "demand". Within the same time index, the sparse selection and targeted mutual transmission of semantic supply and demand coupling, by aligning "supply" and "demand", prioritizes regions with both strong supply and strong demand. Through bandwidth constraints and value ranking, the region-level transmission order is arranged so that the inspection agent only sends the most valuable evidence regions and points them to the most needed target inspection agents, thereby avoiding bandwidth waste and latency accumulation caused by indiscriminate transmission.
[0035] Sparse selection and targeted inter-transmission in semantic supply and demand coupling includes candidate region generation and metric aggregation, value assessment and sparse selection, and targeted inter-transmission and reception determination.
[0036] Candidate region generation and metric aggregation include, within each time index, dividing the pixel set into candidate region sets according to a fixed grid step size and connectivity rules (e.g., a grid step size of eight pixels, connectivity rules of eight-connectivity, and a minimum connected area of sixty-four pixels). For each candidate region, three types of quantities are calculated: first, the region confidence integral, defined as the sum of the comprehensive semantic confidence scores pixel by pixel within the candidate region pixel set, with the number of pixels as the upper bound; second, the region request mean, defined as the average of the semantic request scores pixel by pixel within the candidate region pixel set, maintaining a normalized range of 0 to 1; and third, the transmission cost, defined as the transmission cost under the combined effects of the candidate region clipping box area, encoding format, and metadata field length. The value of the transmission cost varies with data volume and network conditions. When using bytes as the unit, the value ranges to a positive integer; when using milliseconds as the unit, the value ranges to a positive real number. The region confidence integral is used to quantify the strength of available evidence, the region request mean is used to quantify the strength of the demand for external evidence, and the transmission cost is used to quantify the transmission cost.
[0037] Value assessment and sparse selection involve evaluating the value of each candidate region using a single fractional scoring method within the same time index, based on the region confidence score, the region mean request value, and the transmission cost. The benefit score is calculated for each candidate region and expressed as follows: ; in, The inspection intelligent agent is identified as The candidate region index is In time index The efficiency score indicates that the higher the efficiency score, the higher the value for the same sending cost. The inspection intelligent agent is identified as The candidate region index is In time index The region confidence integral, with values ranging from 0 to the number of pixels in the candidate region. The inspection intelligent agent is identified as The candidate region index is In time index The regional average request value. The inspection intelligent agent is identified as The candidate region index is In time index The transmission cost.
[0038] The regions are sorted from highest to lowest based on their benefit scores, and then a second screening is performed using a sparse selection benefit threshold and time window bandwidth budget to obtain a set of candidate regions that are adopted.
[0039] The secondary screening process includes: after generating a candidate region set, calculating the benefit score for each candidate region and sorting them from highest to lowest benefit score to generate a ranking sequence; traversing the candidate regions one by one in the ranking sequence and performing sparse selection benefit screening; marking a candidate region as passed when its benefit score is not less than the sparse selection benefit threshold, and directly eliminating a candidate region when its benefit score is less than the sparse selection benefit threshold; forming a screening subsequence from the candidate regions that have passed the sparse selection benefit screening in the ranking sequence, and accumulating the transmission cost in the screening subsequence according to the ranking order; adding candidate regions whose accumulated value does not exceed the time window bandwidth budget to the set of adopted candidate regions; and adding a new item will cause the accumulated transmission cost to increase. If the cost exceeds the time window bandwidth budget, the addition process stops and the selection process for the current time window ends. When different candidate regions overlap in the pixel coordinate range, a parallel conflict resolution rule is executed, prioritizing the retention of candidate regions with higher average region request values. If the average region request values are the same, the candidate region with lower transmission cost is retained. If the transmission costs are also the same, the candidate region with an earlier timestamp is retained. When the screening result based on the sparse selection benefit threshold is empty, an empty set guarantee rule is enabled. Starting from the first item of the original sorting sequence, the data is accumulated according to the time window bandwidth budget until the accumulated transmission cost reaches a small portion (e.g., 10%) of the time window bandwidth budget or the original sorting sequence is exhausted, in order to avoid zero transmission in the current time window. The sparse selection benefit threshold, with an example value range of 0.3 to 1.5, is determined based on robust statistics of the unit cost benefit distribution. The sample distribution of benefit scores is collected over multiple consecutive time windows, and the median and upper quartile are obtained after removing the extreme values at both ends of the distribution. The sparse selection benefit threshold is set as a compromise between the median and the upper quartile, thereby maintaining a balance between pass rate and benefit in different scenarios. The example value range and basis ensure that the sparse selection benefit threshold is neither too low, causing a large number of low-value areas to pass, nor too high, causing high-value areas to be falsely rejected. The time window bandwidth budget, with an example value range of 64,000 bytes to 520,000 bytes for a 1-second time window, is based on the product of link rate, protocol overhead ratio, control margin ratio, and time window length. This quantifies the field communication capabilities and real-time targets into the time window bandwidth budget, thereby providing a clear upper limit for bandwidth selection in the secondary screening.
[0040] The targeted transmission and reception determination process includes: cropping the adopted candidate region into the smallest identifiable image block and binding it with compact embedding and initial semantic category judgment, along with a timestamp, pixel coordinate range, and source inspection agent identifier; for each adopted candidate region, extracting a sub-region with the same pixel coordinate range from the semantic request map of each inspection agent other than the source inspection agent and calculating the region request average; forming a target inspection agent set by inspection agents with a region request average greater than zero, sorting them from high to low region request average, and adding them to the transmission list in order according to the time window bandwidth budget and the total transmission limit; target inspection agents not included in the transmission list do not transmit; when multiple target inspection agents have the same region request average, the transmission order is determined according to the following order: more recent timestamp, larger pixel coordinate range, and higher initial semantic category risk level. Among them, the initial semantic category judgment means that within the candidate region pixel set, the category with the highest semantic similarity is selected pixel by pixel according to the semantic prototype set of the inspection task, the frequency of occurrence is counted and the category with the most occurrences is taken as the initial semantic category judgment of the candidate region. The initial risk level of the semantic category is determined by combining the defect severity of the candidate region's semantic category with the safety level of the equipment component. For example, when the semantic category of the candidate region is crack (high defect severity) and the component it belongs to is a pressure vessel (high equipment component safety level), the initial risk level of the semantic category is determined to be the highest level; when the semantic category of the candidate region is coating peeling (medium defect severity) and the component it belongs to is a load-bearing beam (high equipment component safety level), the initial risk level of the semantic category is determined to be medium-high level; when the semantic category of the candidate region is surface stain (low defect severity) and the component it belongs to is an outer shell (low equipment component safety level), the initial risk level of the semantic category is determined to be the lowest level. Compact embedding refers to the dimension vector obtained by weighting the pixel-level visual embedding vectors according to the comprehensive semantic confidence within the set of candidate region pixels.
[0041] Directed mutual transmission is performed according to the sending list. At the same time, a sending record is generated and written with a timestamp, source inspection agent identifier, target inspection agent identifier, pixel coordinate range and region request mean. The smallest recognizable image block, compact embedding, and initial semantic category judgment together with the sending record constitute a multi-source sparse semantic feature queue.
[0042] Furthermore, the sparse selection and targeted inter-transmission of semantic supply and demand are simultaneously applied to limit the total number of transmissions within each time index using the time window bandwidth budget. In the set of adopted candidate regions, a merging strategy is performed on the smallest identifiable image blocks with overlapping pixels in the same time index to reduce redundancy. The merging rule is that when the semantic categories are initially judged to be consistent, the union of the pixel coordinate range is calculated and the transmission cost is updated synchronously. When the semantic categories are initially judged to be inconsistent, the smallest identifiable image block with the higher benefit score is retained and the smallest identifiable image block with the lower benefit score is removed. When the transmission list generated by sorting and filtering reaches the upper limit of the time window bandwidth budget, the addition of new regions stops. Candidate regions not included in the current transmission list are retained in the local candidate region list and participate in the calculation and sorting again in the candidate region generation and measurement aggregation stage of the next time index.
[0043] The fusion relabeling module performs position-level semantic attention fusion and measurable sensitivity relabeling on a multi-source sparse semantic feature queue, generating a fusion probability map and a fusion instance table.
[0044] Furthermore, the position-level semantic attention fusion includes spatial alignment, source inspection agent credibility quantification, and position-level weighted fusion. The measurable sensitivity recalibration includes leave-one source inspection agent sensitivity verification and weight recalibration. Specifically, within the same time index, spatial alignment, source inspection agent credibility quantification, position-level weighted fusion, leave-one source inspection agent sensitivity verification, and weight recalibration are sequentially performed on the multi-source sparse semantic feature queue to generate a fusion probability map and a fusion instance table.
[0045] The multi-source sparse semantic feature queue is spatially aligned on a unified reference pixel grid. The spatial alignment process is to select a reference pixel grid and crop all the smallest recognizable image blocks to the corresponding pixel coordinate range, then use bilinear interpolation to scale to the reference pixel grid resolution, and finally use boundary padding to process the boundary uncovered area generated after scaling. After spatial alignment, a source inspection agent probability map and source inspection agent description information are established for each source inspection agent. The source inspection agent probability map is a pixel-level probability matrix consistent with the reference pixel grid. The source inspection agent description information includes the source inspection agent identifier, timestamp, pixel coordinate range, and preliminary semantic category judgment. The spatial alignment process only performs three types of operations: cropping, scaling, and boundary padding to avoid semantic distortion caused by unnecessary deformation and ensure that the subsequent position-level semantic attention fusion is completed within a strict co-domain.
[0046] Based on the smallest recognizable image patch in the multi-source sparse semantic feature queue, a candidate instance mask is generated for each smallest recognizable image patch. The candidate instance mask is a binary matrix consistent with the reference pixel grid. The value of the cell whose pixel coordinates belong to the smallest recognizable image patch is 1, and the value of the cell whose pixel coordinates do not belong to the smallest recognizable image patch is 0. The candidate instance mask is used to limit the set of pixels for position-level semantic attention fusion and measurable sensitivity recalibration, and is used to calculate the average semantic confidence and sensitivity of the source inspection agent. At the same time, the candidate instance mask serves as the starting region constraint for fusion instance generation and geometric refinement. Within each candidate instance mask, the average semantic confidence of the source inspection agent is calculated, and the prior attention weight is obtained according to the normalization relation. The higher the average semantic confidence of the source inspection agent, the larger the prior attention weight, as expressed as: ; in, Indicates the prior attention weights. Indicates the number of source inspection intelligent agents. Indicates the source inspection agent index. This represents a placeholder index in the summation, used to traverse all source inspection agents within the normalized denominator. The source inspection agent index is indicated as The average semantic confidence of the source inspection agent; Subsequently, position-level weighting is performed on the probabilistic map of the source inspection agent using prior attention weights to obtain the initial fusion probability expression. Simultaneously, for the pixel mask corresponding to each candidate region, the measurable sensitivity of the source inspection agent on the candidate region is measured by successively eliminating source inspection agents and recalculating the fusion probability expression, as expressed as: ; in, The source inspection agent index is indicated as The measurable sensitivity, ranging from 0 to 1, is derived from the initial fusion probability expression and the removal of the source index. The pixel-level differences between the recalculated probability representations. Represents the set of candidate instance mask pixels. Indicates the number of pixels in the candidate instance mask. This represents the initial fusion probability expression obtained by weighting using prior attention weights. The index of the inspection agent excluding the source is... The fusion probability expression obtained by reweighting; The prior attention weights are shrunk and normalized based on measurable sensitivity to obtain recalibrated attention weights. The fusion probability map is then calculated at pixel coordinates based on these recalibrated attention weights. The fusion probability value is expressed as: ; ; in, Represented as a fusion probability map at pixel coordinates The fusion probability value, ranging from 0 to 1, represents the probability that a pixel belongs to the target semantics of the inspection task semantic prototype set. This indicates that attention weights are being recalibrated. The source inspection agent index is indicated as The probability map of the source inspection agent is located at pixel coordinates. The probability value at that location; Furthermore, the fusion probability map is filled with fusion probability values according to the reference pixel grid based on the pixel coordinates. When multiple candidate instance masks overlap at the same time index, the fusion probability map selects the value with the larger fusion probability value at the overlapping pixel coordinates to ensure that only the strongest semantic evidence is retained at the same pixel coordinate. When there are uncovered areas at the pixel coordinates of adjacent candidate instance masks, the fusion probability map keeps the value of the pixel coordinates corresponding to the uncovered area as zero, and the pixel coordinates corresponding to the uncovered area do not participate in the subsequent connectivity analysis. The fusion probability map is arranged in ascending order of timestamps in the time index dimension and is consistent with the timestamp of the source inspection agent probability map. After the fusion probability graph is generated, binarization, connectivity analysis, and geometric refinement are performed. Binarization employs a foreground determination criterion, calculating the average neighborhood fusion probability centered on each pixel. The pixel fusion probability value is compared with the average neighborhood fusion probability value. If the pixel fusion probability value is greater than the average neighborhood fusion probability value, it is marked as a foreground pixel; otherwise, it is marked as a background pixel. After traversing all pixel coordinates, a preliminary foreground mask is obtained. The connectivity analysis uses the eight-connectivity rule to extract the set of connected components from the initial foreground mask, and generates the minimum bounding rectangle for each connected component by calculating the minimum and maximum values of the horizontal and vertical coordinates. When the minimum bounding rectangles of two connected components overlap, they are identified as related regions and merged. After merging, the pixel coordinate range and timestamp of the merged region are updated. Geometric refinement generates a fusion instance mask for each connected component and calculates the fusion instance bounding box (minimum bounding rectangle). Skeleton refinement is performed inside the fusion instance mask, and the skeleton length is calculated. The length is measured by combining the pixel-to-physical scale calibration factor. The area is measured by multiplying the number of pixels in the fusion instance mask by the square of the pixel-to-physical scale calibration factor. In the case of character reading scenarios, character recognition is performed inside the fusion instance mask, and the character reading confidence is recorded. Simultaneously, for each fusion instance mask, the initial semantic category judgment and corresponding recalibration attention weights from the source inspection agent's description information are collected to construct a candidate category set as the deduplicated set of the initial semantic category judgment. The total weight votes are calculated for each candidate category. Specifically, among all source inspection agents, those whose initial semantic category judgment is equal to the current candidate category are selected, and the recalibration attention weights corresponding to the selected source inspection agents are added one by one. The accumulated result is the total weight votes for the current candidate category. The candidate category with the largest total weight votes is selected as the semantic category. When there are ties in the total weight votes, the initial semantic category judgment corresponding to the source inspection agent with the largest recalibration attention weight is selected as the semantic category. When the total weight votes are still tied, the candidate category with the smaller encoding order is selected as the semantic category according to the category identifier encoding order comparison rule. Finally, the length measurement, area measurement, character reading, timestamp, source inspection agent weight record (the recalibrated attention weight saved in each fusion instance), and semantic category are bound to the corresponding fusion instance to form a fusion instance table.
[0047] The diagnostic framing module performs equipment health status diagnosis and fault detection based on the fusion instance table, and generates a diagnostic list and inspection framing instructions.
[0048] Furthermore, within adjacent time indices, a one-to-one association is performed on fused instances according to the fused instance table to form an instance time series. Specifically, within the same semantic category, the intersection-union ratio (IUR) of the pixel coordinate ranges of two fused instance records is calculated and sorted from high to low. When the IURs are tied, a pairing object is selected based on the Euclidean distance between the center points from small to large. After a pairing is completed, the two records no longer participate in the current pairing. Within the current time index, if the IUR of the pixel coordinate ranges of all fused instance records of the same semantic category in the previous time index is equal to zero with that of the current fused instance, a new instance time series is created. When the instance time series in the previous time index does not have a non-zero IUR among any fused instance records in the current time index, the instance time series is determined to be in a terminated state and will not continue in subsequent time indices.
[0049] For each instance time series, at each time index, length measurement, area measurement, and character reading deviation are calculated using the fused instance mask and fused instance bounding box. Length measurement is obtained by performing skeleton refinement within the fused instance mask to obtain the skeleton length pixel value, and then performing scale conversion based on the pixel-to-physical scale calibration factor to form the length measurement. Area measurement is obtained by counting the number of pixels in the fused instance mask and performing area scale conversion based on the pixel-to-physical scale calibration factor to form the area measurement. Character reading deviation is obtained by calculating the absolute difference between the character reading and the nearest boundary of the allowable reading range given by the inspection task semantic mapping table, and then performing bandwidth normalization mapping using the allowable deviation bandwidth to form a dimensionless deviation. Then, normalization is performed. Length normalization maps the length measurement to the length safety upper limit in the inspection task semantic mapping table proportionally and truncates it in the range of 0 to 1. Area normalization maps the area measurement to the area safety upper limit in the inspection task semantic mapping table proportionally and truncates it in the range of 0 to 1. Reading normalization maps the character reading deviation to the allowable deviation bandwidth proportionally and truncates it in the range of 0 to 1. Uncertainty is obtained by applying a complement mapping after weighted averaging of the pixel-level semantic confidence corresponding to the weight records of the source inspection agent on the fusion instance mask pixel set. The value range is 0 to 1.
[0050] The diagnostic benchmark is obtained by calculating the arithmetic mean of length normalization, area normalization, and reading normalization within the same time index. The number of indicators among the three normalized quantities that exceed the diagnostic benchmark is counted and marked. When the number mark is three, the judgment level label is red, indicating that the safe operating boundary has been exceeded and immediate action is required. When the number mark is two, the judgment level label is yellow, indicating that the safe operating boundary is approaching and close monitoring is required. When the number mark is less than or equal to one, the judgment level label is green, indicating that the safe operating boundary is approaching and close monitoring is required. Subsequently, an uncertainty check is performed, comparing the uncertainty with the diagnostic benchmark. If the uncertainty is greater than the diagnostic benchmark, the level is downgraded by one grade (red level downgraded to yellow level, yellow level downgraded to green level, green level remains unchanged). If the uncertainty is less than or equal to the diagnostic benchmark, the original level is maintained.
[0051] For each time index, a diagnostic record is generated for each instance time series. The diagnostic record includes instance identifier, semantic category, fused instance bounding box, fused instance mask, length measurement, area measurement, character reading deviation, length normalization, area normalization, reading normalization, uncertainty, grade label, source inspection agent weight record, and timestamp; where the instance identifier is a unique identifier for the instance time series. All diagnostic records within the same time index are aggregated to form a diagnostic list. The diagnostic list is grouped by instance identifier and sorted in ascending order by time index within each group. At the same time, it is sorted in ascending order by timestamp at the overall level to facilitate comparison and traceability, and to maintain consistency with the field naming of the fused instance table.
[0052] Furthermore, within the same time index, based on the semantic request graph and diagnostic list, corresponding inspection framing instructions are generated for each inspection agent on the candidate viewpoint set of the inspection agent. The inspection framing instructions are used to guide the shooting orientation, pitch, focal length, and position adjustment of the inspection agent in the next time index. The specific steps are as follows: A candidate view set is constructed, which consists of discrete combinations of four parameters: azimuth increment, pitch increment, focal length ratio, and camera position offset. The candidate view coverage area is established with the fusion instance center of the diagnostic checklist marked as red and yellow as the target center. Based on the optical parameters of the video equipment configured by the inspection agent, the azimuth increment and pitch increment are converted into pixel horizontal displacement and pixel vertical displacement, respectively. The coverage width and coverage height are calculated according to the focal length ratio. After translation on the reference pixel grid according to the camera position offset, the candidate view coverage area is obtained. Subsequently, three metrics are calculated: the coverage average is equal to the arithmetic mean of the semantic request graph over the pixel set of the candidate view coverage area, used to quantify the intensity of information demand; the expected coverage area ratio is equal to the ratio of the intersection area of the candidate view coverage area and the high-priority target area to the area of the candidate view coverage area, where the high-priority target area is obtained by the union of the bounding boxes of the fused instances labeled red and yellow in the diagnostic checklist; and the view rotation and translation cost refers to the cost generated by the combined angular change and position offset of the candidate view relative to the reference view. This is obtained by calculating the absolute value of the azimuth change, the absolute value of the pitch change, and the camera position offset, and then combining the angular change and position offset. The candidate viewpoints are sorted and selected in a multi-level sorting manner based on the average coverage, the cost of viewpoint rotation and translation, and the expected coverage area ratio. Specifically, they are sorted from high to low according to the average coverage. When the average coverage is the same, they are sorted from small to large according to the cost of viewpoint rotation and translation. If the sorting is still the same, they are sorted from large to small according to the expected coverage area ratio. A preset number of candidate viewpoints are selected based on the sorting results, and inspection framing instructions are generated. The inspection and framing instruction records the inspection agent's identifier, time index, target location, azimuth increment, pitch increment, focal length ratio, camera position offset, average coverage, cost of viewpoint rotation and translation, and expected coverage area ratio.
[0053] In summary, this invention achieves the identification of high-value candidate regions and directional transmission across inspection agents by performing sparse selection and directional transmission of semantic supply and demand coupling under the dual constraints of semantic confidence graph and semantic request graph; and achieves spatial alignment, confidence optimization and uncertainty quantification of multi-source sparse semantic features by performing position-level semantic attention fusion and measurable sensitivity recalibration on a unified reference pixel grid, ensuring the accuracy and reliability of the fusion results.
[0054] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.
Claims
1. A semantically driven collaborative perception system for inspection intelligent agents, characterized in that: include, The prototype mapping module collects information on the categories of inspection targets and the monitoring indicators of the inspection targets, and generates a semantic prototype set and a semantic mapping table for inspection tasks through semantic transformation. The semantic graph module collects inspection image data and inspection video data through the inspection agent, and performs pixel-level visual feature matching based on the semantic prototype set of the inspection task to generate semantic confidence graph and semantic request graph. The coupled mutual transmission module, under the joint constraints of the semantic confidence graph and the semantic request graph, performs sparse selection and directional mutual transmission of semantic supply and demand coupling, and generates a multi-source sparse semantic feature queue. The fusion relabeling module performs position-level semantic attention fusion and measurable sensitivity relabeling on a multi-source sparse semantic feature queue, generating a fusion probability map and a fusion instance table. The diagnostic framing module performs equipment health status diagnosis and fault detection based on the fusion instance table, and generates a diagnostic list and inspection framing instructions.
2. The semantic-driven collaborative perception system for inspection agents as described in claim 1, characterized in that: The inspection target category information includes equipment component category, defect type, surface quality problem and structural problem, and each inspection target category information contains a unique category identifier and a complete text description. After collection, the information is standardized and organized according to the unique category identifier to generate an inspection target category information set, and the monitoring indicators of the inspection targets are collected one by one based on the inspection target category information set. Establish a one-to-one correspondence between the inspection target category information and the monitoring indicators of the inspection target, and generate a set of corresponding inspection target categories and monitoring indicators.
3. The semantic-driven collaborative perception system for inspection agents as described in claim 2, characterized in that: The semantic transformation includes encoding the textual description of each inspection target category information using a text encoder to generate a text semantic embedding vector. Samples corresponding to the inspection target category information are extracted from historical inspection image data and historical inspection video data. Visual features are extracted using a visual encoder to generate visual embedding vectors. By fusing textual semantic embedding vectors and visual embedding vectors, a semantic prototype vector of the inspection target category is generated; The monitoring indicators of the inspection targets are encoded using the same text and visual encoding as the category information of the inspection targets, and then fused to generate semantic prototype vectors of the monitoring indicators. The semantic prototype vectors of inspection target categories and monitoring indicator semantic prototype vectors are aggregated to generate a semantic prototype set of inspection tasks. Based on the one-to-one correspondence between the inspection target categories and the monitoring indicator sets of the inspection targets, a semantic mapping table of inspection tasks is generated.
4. The semantic-driven collaborative perception system for inspection agents as described in claim 3, characterized in that: The steps for generating semantic confidence maps and semantic request maps by performing pixel-level visual feature matching based on the semantic prototype set of the inspection task are as follows. Under a unified preprocessing standard, the inspection image data and inspection video data are sequentially subjected to denoising, resolution standardization, and contrast enhancement. Visual embedding vectors are calculated for pixels in the preprocessed inspection image data and inspection video data. The semantic similarity is obtained by calculating the cosine similarity between the visual embedding vector of the pixel and the semantic prototype vector in the semantic prototype set of the inspection task, and taking the maximum value. The text readability score is obtained by performing optical character recognition confidence normalization on the pixel neighborhood, and the image blur score is obtained by monotonic mapping using spatial gradient class sharpness metric. By positively accumulating and suppressing ambiguity in the semantic similarity and text readability scores through imaging ambiguity scoring, a comprehensive semantic confidence score is generated. The semantic confidence map is generated by filling the two-dimensional coordinate grid with the comprehensive semantic confidence value according to the pixel coordinate position, and the semantic request map is generated by filling the same pixel coordinate position with the pixel value of one minus the comprehensive semantic confidence value.
5. The semantic-driven collaborative perception system for inspection agents as described in claim 4, characterized in that: The sparse selection and directed mutual transmission of semantic supply and demand coupling are described in the following steps. Regions with strong supply and strong demand are identified based on semantic confidence maps and semantic request maps. Candidate regions are divided according to fixed grid and connectivity rules, and the domain confidence integral, region request mean, and transmission cost of the candidate regions are calculated. Based on the regional confidence score, the regional request mean, and the transmission cost, a single fractional scoring method is used to evaluate the value of candidate regions and calculate their benefit scores. The benefit scores are sorted from high to low, and a second screening is performed to obtain the candidate regions to be adopted. The candidate regions to be adopted are cropped into the smallest recognizable image blocks and bound to compact embeddings and initial semantic category judgments, along with timestamps, pixel coordinate ranges, and source inspection agent identifiers; For each adopted candidate region, a sub-region with the same pixel coordinate range is extracted from the semantic request graph of each inspection agent other than the source inspection agent, and the average region request value is calculated. The inspection agents with a regional request average greater than zero are grouped into a target inspection agent set, sorted from high to low according to the regional request average, and then included in the sending list in order according to the time window bandwidth budget and the maximum sending volume limit. Targeted mutual transmission is performed according to the sending list, and a sending record is generated at the same time.
6. The semantic-driven collaborative perception system for inspection agents as described in claim 5, characterized in that: The multi-source sparse semantic feature queue includes the smallest recognizable image block, compact embedding, initial semantic category judgment, and transmission record.
7. The semantic-driven collaborative perception system for inspection agents as described in claim 6, characterized in that: The location-level semantic attention fusion and measurable sensitivity recalibration steps are as follows: Spatially align the multi-source sparse semantic feature queue on a unified reference pixel grid. After spatial alignment, a probability graph of the source inspection agent and descriptive information of the source inspection agent are established. Based on the smallest recognizable image patch in the multi-source sparse semantic feature queue, a candidate instance mask is generated for the smallest recognizable image patch; Calculate the average semantic confidence of the source inspection agent within the candidate instance mask, and generate prior attention weights according to the normalized relation; The source inspection agent probability map is weighted and synthesized according to the prior attention weight to generate an initial fusion probability expression, and the measurable sensitivity is measured on the candidate instance mask pixel set by the leave-one source inspection agent sensitivity verification method. The prior attention weights are shrunk and normalized based on the measurable sensitivity to obtain the recalibrated attention weights, and the fusion probability value is calculated according to the recalibrated attention weights.
8. The semantic-driven collaborative perception system for inspection agents as described in claim 7, characterized in that: The fusion probability map and fusion instance table include filling the fusion probability map with the fusion probability values according to the reference pixel grid based on the pixel coordinates. After the fusion probability map is generated, binarization, connectivity analysis and geometric refinement are performed to form the fusion instance table.
9. The semantic-driven collaborative perception system for inspection agents as described in claim 8, characterized in that: The health status diagnosis and fault detection of the execution equipment includes establishing an instance time series based on the fusion instance table within the adjacent time index, calculating normalized diagnostic indicators and uncertainties, and generating level labels and diagnostic lists based on the comparison results of diagnostic indicators.
10. The semantic-driven collaborative perception system for inspection agents as described in claim 9, characterized in that: The inspection framing instruction includes generating corresponding inspection framing instructions for the inspection agent on the candidate view set of the inspection agent based on the semantic request graph and the diagnostic list.
Citation Information
Cited By
Method and system for mitigating relationship hallucination of multimodal large model based on relationship-aware visual augmentation
CN122176807A