Sky-ground comprehensive operation inspection method and system based on multi-source cooperation
By calculating the modal redundancy consistency and semantic transition distortion index to evaluate the quality of multi-source data fusion, the problem of insufficient fusion quality evaluation in existing systems is solved, and the effective judgment of the fusion effect of multi-source data is realized, ensuring the reliability and stability of operation and maintenance tasks.
Patent Information
- Application Number
- CN202511046152.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-29
- Publication Date
- 2025-11-18
AI Technical Summary
Existing multi-source collaborative operation and maintenance systems lack effective evaluation and feedback mechanisms for fusion quality, making it difficult to identify when modal data fusion fails, thus affecting the reliability and stability of the overall operation and maintenance tasks.
The fusion quality of multi-source data is evaluated by calculating the modal redundancy consistency index and the semantic transition distortion index. The fusion quality index is compared with a preset threshold to determine the fusion effect of the modal data, and the data is re-fused when it is unqualified.
To ensure the effectiveness of multi-source data fusion, reduce misjudgments and omissions, and guarantee the reliability and stability of operation and maintenance tasks.
Smart Images

Figure CN120974402A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of collaborative inspection technology, in particular to a multi-source collaborative sky-ground comprehensive operation inspection method and system. BACKGROUND
[0002] With the continuous development of multi-source perception technology, through multi-source collaborative comprehensive operation inspection mechanism is becoming an important technical path for complex space target monitoring, analysis and decision. For example, by collaboratively integrating data resources from the "sky-ground" three types of observation platforms, a multi-scale, multi-modal, cross-platform integrated perception system is constructed. Among them, "sky" mainly relies on satellite remote sensing means, with macro observation ability of large range and high timeliness; "air" mainly refers to low-altitude platforms such as unmanned aerial vehicles and helicopters, which have high spatial resolution and flexible deployment characteristics; "ground" includes fixed sensors, mobile devices, manual patrol and other ways, which can provide fine-grained, close-range state perception data. By fusing multi-modal data (such as images, videos, infrared, radar, structural signals, etc.) collected by the three types of platforms, global coverage, stereoscopic perception and dynamic analysis of the target area or object are realized, which provides more comprehensive and intelligent support for state monitoring, risk assessment, task scheduling and other operation inspection tasks.
[0003] Although this mechanism shows good perception ability and task collaboration in practical scenarios, the current multi-source collaborative operation inspection system generally lacks effective evaluation and feedback mechanism for fusion quality. Specifically, the existing method often defaults that multi-source data can be naturally complementary in the fusion processing stage, but lacks quantitative judgment of the fusion effect of the three types of modal data of sky-ground, and does not establish an automatic detection and self-checking mechanism for fusion anomalies. Once the modal data fusion fails, the system is difficult to identify the occurrence of fusion failure, thereby causing misjudgment, missed judgment and even systematic bias in the downstream analysis, identification and decision-making links, which seriously affects the reliability and stability of the overall operation inspection task. SUMMARY
[0004] The purpose of the present application is to solve the above-mentioned problems, and to provide a multi-source collaborative sky-ground comprehensive operation inspection method and system.
[0005] In the first aspect of the present application, a multi-source collaborative sky-ground comprehensive operation inspection method is first proposed, which comprises:
[0006] Collecting data of sky modal, air modal and ground modal, and extracting corresponding feature information, and calculating modal redundancy consistency index based on feature information between different modalities;
[0007] Analyzing the feature transition path between the feature information of each modal, and calculating the semantic transition distortion degree index;
[0008] According to the modal redundancy consistency index and the semantic transition distortion degree index, a fusion quality index is calculated;
[0009] The fusion quality index is compared with a preset threshold, and whether the fusion quality of the three modal data of the sky, the air and the ground is qualified is determined according to a comparison result, if qualified, a running inspection result is output according to the fused data.
[0010] Optionally, the step of calculating the modal redundancy consistency index based on the feature information between different modalities is:
[0011] Normalized feature vectors representing the same target object in the sky modality, the air modality and the ground modality are collected, and are denoted as a sky modality feature vector F T , an air modality feature vector F A and a ground modality feature vector F G , respectively.
[0012] A modal difference matrix R is constructed, and an element R ij in the modal difference matrix R represents a normalized difference degree between modal i and modal j, and is defined as: In the formula, i and j are both in {T, A, G}, and ||·||2 represents a two-norm of a vector.
[0013] The difference between the difference degrees between two modes in the modal difference matrix is squared and summed to obtain an information structure asymmetry error of each modality.
[0014] The information structure asymmetry errors of all modalities are squared and accumulated to obtain a structure conflict degree index S of the modalities as a whole.
[0015] The cosine of the angle between the three groups of feature vectors is calculated, and an average value thereof is taken to obtain a redundancy cooperative direction consistency index C between the modalities.
[0016] The modal redundancy consistency index is calculated according to the overall structure conflict degree index S and the redundancy cooperative direction consistency index C, and the formula for calculation is: In the formula, MRI is the modal redundancy consistency index.
[0017] Optionally, the step of calculating the semantic transition distortion degree index is:
[0018] Normalized feature vectors representing the same target object in the sky modality, the air modality and the ground modality are collected, and are denoted as a sky modality feature vector F T , an air modality feature vector F A and a ground modality feature vector F G .
[0019] A semantic transition difference vector between modalities is calculated, ΔTA=F A -F T , and ΔAG=FG -F A , wherein, ΔTA is a semantic transition difference vector from the sky mode to the air mode, and ΔAG is a semantic transition difference vector from the air mode to the ground mode;
[0020] calculating a semantic transition included angle tension coefficient J, , wherein, <·,·> is a vector dot product, and ||·||2 represents a two-norm of a vector;
[0021] calculating a semantic path curvature coefficient K,
[0022] calculating a semantic drift coefficient D, and the calculation step is ΔTG=F G -F T , wherein, ΔTG represents a semantic transition difference total vector from the sky mode to the ground mode directly,
[0023] multiplying the semantic transition included angle tension coefficient J, the semantic path curvature coefficient K, and the semantic drift coefficient D to obtain a semantic transition distortion degree index.
[0024] Optionally, according to the modal redundancy consistency index and the semantic transition distortion degree index, the step of calculating the fusion quality index is:
[0025] normalizing the modal redundancy consistency index and the semantic transition distortion degree index, and performing weighted summation on the normalized modal redundancy consistency index and the semantic transition distortion degree index to obtain the fusion quality index.
[0026] Optionally, the step of comparing the fusion quality index with a preset threshold and determining whether the fusion quality of the sky-ground-three-mode data is qualified according to a comparison result includes:
[0027] If the fusion quality index is less than the preset threshold, it indicates that the fusion quality of the sky-ground-three-mode data is qualified, and a result of operation and inspection is directly output according to data fused by the sky-ground-three-mode data;
[0028] If the fusion quality index is not less than the preset threshold, it indicates that the fusion quality of the sky-ground-three-mode data is unqualified, and the sky-ground-three-mode data needs to be fused again until the fusion quality of the sky-ground-three-mode data is qualified.
[0029] In the second aspect of the embodiment of the present application, a multi-source collaborative sky-ground integrated operation and inspection system is provided, and the system includes:
[0030] a modal redundancy consistency module: collecting data of a sky mode, an air mode and a ground mode, extracting corresponding feature information, and calculating a modal redundancy consistency index based on feature information between different modes;
[0031] semantic transition distortion module: analyzing the feature transition path between the feature information of each modality, and calculating the semantic transition distortion index;
[0032] fusion quality module: calculating the fusion quality index according to the modality redundancy consistency index and the semantic transition distortion index;
[0033] operation detection module: comparing the fusion quality index with a preset threshold, and determining whether the fusion quality of the three modalities of sky, air and ground is qualified according to the comparison result, and if qualified, outputting the operation detection result according to the fused data.
[0034] Optionally, the modality redundancy consistency module comprises:
[0035] feature vector module: collecting normalized feature vectors representing the same target object in the sky modality, the air modality and the ground modality, respectively denoted as sky modality feature vector F T , air modality feature vector F A and ground modality feature vector F G ;
[0036] construction module: constructing a modality difference matrix R, whose element R ij represents the normalized difference degree between modality i and modality j, defined as: wherein, i and j are both ∈{T, A, G}, and ||·||2 represents the two-norm of a vector;
[0037] information structure asymmetry error module: squaring and summing the difference degree difference between the two modalities in the modality difference matrix to obtain the information structure asymmetry error of each modality;
[0038] structure conflict degree module: squaring and accumulating the information structure asymmetry errors of all modalities to obtain the structure conflict degree index S of the whole modality;
[0039] direction consistency module: calculating the cosine of the angle between each two of the three groups of feature vectors, and taking the average value to obtain the redundancy collaborative direction consistency index C between the modalities;
[0040] modality redundancy consistency index module: calculating the modality redundancy consistency index according to the whole structure conflict degree index S and the redundancy collaborative direction consistency index C, and the formula is: wherein, MRI is the modality redundancy consistency index.
[0041] Optionally, the semantic transition distortion module comprises:
[0042] feature vector module: collecting normalized feature vectors representing the same target object in the sky modality, the air modality and the ground modality, respectively denoted as sky modality feature vector F T, an air mode feature vector F A and a ground mode feature vector F G ;
[0043] a semantic transition difference vector module: calculating an inter-modal semantic transition difference vector, ΔTA=F A -F T , ΔAG=F G -F A , wherein ΔTA is a semantic transition difference vector from the sky mode to the air mode, and ΔAG is a semantic transition difference vector from the air mode to the ground mode;
[0044] a transition angle tension module: calculating a semantic transition angle tension coefficient J, , wherein <·,·> is a vector dot product, and ||·||2 represents a two-norm of a vector;
[0045] a semantic path curvature module: calculating a semantic path curvature coefficient K,
[0046] a semantic drift module: calculating a semantic drift coefficient D, the calculation being performed according to ΔTG=F G -F T , wherein ΔTG represents a total semantic transition difference vector from the sky mode to the ground mode,
[0047] a semantic transition distortion index module: multiplying the semantic transition angle tension coefficient J, the semantic path curvature coefficient K, and the semantic drift coefficient D to obtain a semantic transition distortion index.
[0048] Optionally, the fusion quality module specifically includes:
[0049] normalizing the modal redundancy consistency index and the semantic transition distortion index, and performing weighted summation on the normalized modal redundancy consistency index and the semantic transition distortion index to obtain a fusion quality index.
[0050] Optionally, the operation inspection module includes:
[0051] a first judgment module: if the fusion quality index is less than a preset threshold, it indicates that the fusion quality of the sky-ground-three-mode data is qualified, and an operation inspection result is directly output according to the data fused by the sky-ground-three-mode data;
[0052] a second judgment module: if the fusion quality index is not less than the preset threshold, it indicates that the fusion quality of the sky-ground-three-mode data is unqualified, and the sky-ground-three-mode data needs to be re-fused until the fusion quality of the sky-ground-three-mode data is qualified.
[0053] Advantages of the present application:
[0054] The present application provides a multi-source collaborative sky-ground comprehensive operation and inspection method and system, by collecting data of sky modal, air modal and ground modal, and respectively extracting feature information representing the same target object; and based on the feature distribution relationship between different modalities, calculating a modal redundancy consistency index to measure the consistency degree of each modal in semantic collaboration and information structure; by analyzing the transition path between modal features, further calculating a semantic transition distortion degree index to evaluate the transmission continuity and stability of modal data at the semantic level; fusing the modal redundancy consistency index and the semantic transition distortion degree index to obtain a fusion quality index; comparing the fusion quality index with the quality threshold preset by the system, and judging whether the fusion quality of the sky-ground three modalities meets the standard according to the comparison result, if the fusion quality index is less than or equal to the threshold, it is considered that the fusion result is qualified, and the final operation and maintenance detection result can be directly output; in this way, the fusion effect of the sky-ground three modal data can be judged, whether the fusion effect of the sky-ground three modal data is available, the effectiveness of data fusion is ensured, the misjudgment, missed judgment and even systematic deviation caused in the downstream analysis, recognition and decision-making links are reduced, and the reliability and stability of the overall operation and inspection task are ensured. BRIEF DESCRIPTION OF DRAWINGS
[0055] The present application will be further described below with reference to the accompanying drawings.
[0056] Figure 1 A flowchart of a multi-source collaborative sky-ground comprehensive operation and inspection method;
[0057] Figure 2 A framework diagram of a multi-source collaborative sky-ground comprehensive operation and inspection system. DETAILED DESCRIPTION
[0058] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor fall within the scope of protection of the present application.
[0059] Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor fall within the scope of protection of the present application.
[0060] The embodiments of the present application provide a multi-source collaborative sky-ground comprehensive operation and inspection method. Referring to Figure 1 , Figure 1A flowchart of a multi-source collaborative sky-ground comprehensive operation inspection method is provided for an embodiment of the present application. The method comprises the following steps:
[0061] S1: Collect data of sky modal, air modal and ground modal, and extract corresponding feature information, and calculate a modal redundancy consistency index based on feature information between different modes, for measuring the redundancy coordination degree of multi-modal data in the representation space;
[0062] S2: Analyze the feature transition path between the modal feature information, and calculate a semantic transition distortion degree index, for measuring the transition continuity and stability of modal data at the semantic level;
[0063] S3: Calculate a fusion quality index according to the modal redundancy consistency index and the semantic transition distortion degree index;
[0064] S4: Compare the fusion quality index with a preset threshold, and determine whether the fusion quality of the sky-ground three modal data is qualified according to the comparison result, and if qualified, output an operation inspection result according to the fused data.
[0065] Based on the multi-source collaborative comprehensive operation inspection method provided by the embodiment of the present application, through the above-mentioned manner, the fusion effect of the sky-ground three modal data can be judged, whether the fusion effect of the sky-ground three modal data is available can be judged, the effectiveness of data fusion is ensured, the misjudgment, omission and even systematic deviation caused in the downstream analysis, recognition and decision-making links are reduced, and the reliability and stability of the overall operation inspection task are ensured.
[0066] In one embodiment, data of sky modal, air modal and ground modal are collected, and corresponding feature information is extracted, and a modal redundancy consistency index is calculated based on feature information between different modes, for measuring the redundancy coordination degree of multi-modal data in the representation space;
[0067] In one implementation, the step of calculating the modal redundancy consistency index based on feature information between different modes is:
[0068] The normalized feature vectors representing the same target object in the sky modal, the air modal and the ground modal are collected, and are respectively denoted as a sky modal feature vector F T , an air modal feature vector F A and a ground modal feature vector F G .
[0069] A 3x3 asymmetric modal difference matrix R is constructed, and an element R ij of the matrix represents the normalized difference degree between modal i and modal j, and is defined as: In the formula, i and j are both in {T, A, G}, and ||·||2 represents the two norm of the vector; R ijThe smaller, the more similar modal i and modal j are; the matrix is an asymmetric matrix, which preserves the directional difference information between modes;
[0070] Square the difference value of the difference degree between the two modes in the modal difference matrix and sum it up to get the information structure asymmetry error Q of each mode i i , Q i =∑ i≠j (R ij -R ji ) 2 , Q i represents the structural asymmetry error of modal i; for measuring the symmetry error of information exchange between modal i and other modes, Q i The larger, the more unbalanced the information transmission between modal i and other modes is;
[0071] Square and accumulate the information structure asymmetry error of all modes to get the overall structural conflict index of the mode, which reflects the asymmetry conflict degree of the information redundancy relationship between the three modes; calculate the overall structural conflict index S, S = Q T 2 +Q A 2 +Q G 2 ;
[0072] Calculate the cosine of the angle between the three groups of eigenvectors, and take the average value to get the redundancy collaborative direction consistency index between the modes; the calculation steps are: represents the feature difference vector between the sky mode and the air mode, whether it is close to the difference direction between the air mode and the ground mode in direction, reflecting whether there is "continuous consistency"; represents the difference direction between the air mode and the sky mode, whether it is aligned with the direction from the sky mode to the ground mode, to evaluate whether there is "semantic coordination of intermediate modal"; represents the direction from the ground mode to the sky mode, whether it is aligned with the transition direction from the sky to the air, for checking whether the ground is consistent with the air-sky semantics in direction; the formula for calculating the redundancy collaborative direction consistency index is: In the formula, C is the redundancy collaborative direction consistency index;
[0073] According to the overall structural conflict index S and the redundancy collaborative direction consistency index, the modal redundancy consistency index is calculated, and the formula is: In the formula, MRI is the modal redundancy consistency index.
[0074] It should be noted that the data acquisition of the sky mode, the air mode and the ground mode is usually based on an existing multi-source heterogeneous perception system, and the same target object is observed by deploying perception devices at different spatial levels to obtain multi-modal raw data. The data sources and their acquisition methods are as follows: first, the sky mode data usually comes from remote sensing payloads (such as optical imaging sensors, multispectral imagers, SAR synthetic aperture radars, etc.) carried by high-orbit or medium-orbit satellites, which are used to periodically image regional targets at a larger spatial scale. For example, the roof structure, thermal features, and building boundaries of a transformer facility in a certain area can be obtained by high-resolution optical satellites (such as Gaofen-1, WorldView). The satellite data acquisition method is usually plan-based periodic imaging, which can also be triggered by ground control systems according to preset tasks; second, the air mode data is mainly obtained by visual, infrared or laser radar sensors carried by low-altitude unmanned aerial vehicles or vertical take-off and landing aircraft, which have the characteristics of high resolution, close range and flexible scheduling, and can capture the details of the target texture, structure outline and even local thermal imaging information. The unmanned aerial vehicle is controlled by the ground station in real time, and flies according to the set route, and automatically collects images or point clouds when flying over the target. For example, when detecting a power transmission line, the unmanned aerial vehicle can observe the insulator and hardware corrosion of the conductor from multiple angles in the air; finally, the ground mode data is usually obtained by deploying fixed cameras, mobile inspection robots or portable detection devices (such as infrared thermal imagers, vibration sensors) on the scene. These devices can obtain real-time surface defects, running state data, sound or thermal response characteristics of the target object. For example, at a transformer device, a ground thermal imager obtains its temperature distribution map, a ground robot scans the device outline with a laser radar, or a handheld instrument captures a close-up image.
[0075] After data acquisition, the three modal data need to be registered and paired, that is, the target objects described in the sky, air and ground data are spatially aligned and semantically matched to ensure that they are indeed observations of the same physical target object. Common methods include spatial registration based on geographic coordinates, image matching based on feature points, or time synchronization based on task ID, etc. For example, if the same device data is to be fused, the geographic location information of the device is used as a reference to align the images or sensor data from the three modes and extract their feature vectors to ensure that subsequent calculation steps are based on the same target. After registration, normalized feature vectors of different modal data can be extracted based on deep convolutional network (CNN), Transformer or multi-modal contrast learning network algorithms, which are used to represent the semantics, structure, texture or state information of the target. Normalization processing can use L2 norm standardization method to ensure that the feature vectors of each mode can be compared and calculated at the same scale.
[0076] It should be noted that the modal redundancy consistency index is a comprehensive index for measuring the relationship between semantic coordination and structural redundancy of multi-modal data in the representation space, and the core purpose is to evaluate whether the data from the three perception levels of sky modal, air modal and ground modal can achieve efficient fusion, collaborative expression and semantic complementation at the feature level. In a multi-modal perception system, different modalities often present different information representation methods for the same target object due to differences in perception perspective, resolution, sensing mechanism, etc. The calculation of the modal redundancy consistency index comprehensively considers two key factors: one is the asymmetric conflict degree of information structure between modalities, that is, whether there is directional bias and structural mismatch in the semantic expression between different modalities; the other is the collaborative consistency between modalities in the feature direction, that is, whether they describe target differences in a close way, whether there is a continuous transmission and alignment ability in semantics. When the index value is larger, it means that the structural conflict between different modalities is stronger (i.e., information exchange is biased or unbalanced), and there is a lack of continuity and semantic alignment ability in directional expression. In this case, multi-modal data is prone to semantic fragmentation, information redundancy overlap or representation distortion in the fusion process, thereby leading to a decline in overall fusion performance. For example, when dealing with a ground vehicle target, if the satellite image (sky modality) can only identify the outline of the vehicle, the aerial drone image (air modality) identifies the color and shadow, but the texture information obtained by the ground video image (ground modality) is significantly inconsistent with the former two in the feature direction, or the structural expression is unbalanced (such as the absence of a key attribute in a modality), then redundant or misleading features are easily introduced in the fusion representation, ultimately affecting the recognition accuracy or the accuracy of behavior understanding. Therefore, as a negative correlation index for measuring the "modality collaboration degree", the smaller the value of the modal redundancy consistency index, the higher the directional consistency and structural symmetry between the three types of modalities, and the more semantic integrity and expression efficiency of the joint features constructed after fusion; on the contrary, the larger the index, the more semantic incoordination and misalignment between modalities, which is not conducive to obtaining stable and reliable fusion expression results. Through the monitoring and regulation of the index, the fusion quality evaluation basis can be provided for the multi-modal fusion system, and the dynamic adjustment of the perception strategy or feature optimization strategy can also be guided.
[0077] It should be noted that the advantages of calculating the modal redundancy consistency index in the above manner are: compared with directly using traditional methods such as Euclidean distance, simple cosine similarity average or information entropy, the above manner can simultaneously consider the two dimensions of “structural symmetry” and “semantic direction consistency” between modes, thereby realizing more comprehensive and in-depth characterization of the quality of multi-modal data fusion. Traditional methods often only focus on a single similarity measure between modes (such as vector distance, cosine angle, etc.), ignoring the imbalance between modes in structural expression, information transmission path and semantic transition process. The present method can capture the directional difference in information exchange between modes by introducing an asymmetric modal difference matrix, and further quantify whether the information share assumed by the modes in joint representation is balanced.
[0078] In one embodiment, the feature transition paths between the feature information of each mode are analyzed, and a semantic transition distortion degree index is calculated to measure the transition continuity and stability of the modal data at the semantic level.
[0079] In one implementation, the step of calculating the semantic transition distortion degree index is:
[0080] Normalized feature vectors representing the same target object in the sky mode, the air mode and the ground mode are collected, and are denoted as the sky mode feature vector F T , the air mode feature vector F A and the ground mode feature vector F G , respectively.
[0081] The inter-modal semantic transition difference vectors are calculated as ΔTA=F A -F T , and ΔAG=F G -F A , where ΔTA is the semantic transition difference vector from the sky mode to the air mode, and ΔAG is the semantic transition difference vector from the air mode to the ground mode. This is the basis for constructing the semantic change path, and the two vectors are connected to form a “T→A→G” semantic conduction chain.
[0082] The semantic transition angle tension coefficient J is calculated as where <·,·> is the vector dot product, and ||·||2 represents the two-norm of the vector. J is the semantic transition angle tension coefficient, which is used to measure whether the two semantic vectors are directionally consistent.
[0083] The semantic path curvature coefficient K is calculated as In the formula, K is a semantic path curvature coefficient, the numerator ||ΔTA-ΔAG||2 represents the "offset" or "curvature" between the two difference vectors; the denominator ||ΔTA||2·||ΔAG||2 means the total length of the semantic path; K is used to measure whether the whole path is smooth. If the semantic expression from the sky to the earth shows a smooth and continuous trend, the two directions should be close, then the difference is small, and the curvature is small; on the contrary, if there is a bend or mutation in the middle, the curvature is large;
[0084] A semantic drift coefficient D is calculated, and the calculation steps are: ΔTG=F G -F T , and ΔTG represents the total difference vector of the semantic transition from the sky mode to the ground mode, The semantic drift coefficient D is used to measure whether the real semantic path is consistent with the step-by-step path; if the semantic conduction is naturally continuous, the synthesized path and the actual path should be equal, that is, the drift is 0; if the drift is large, it means that the semantic expression "seems to be transmitted in stages", but in fact it jumps on the whole space;
[0085] The semantic transition angle tension coefficient J, the semantic path curvature coefficient K, and the semantic drift coefficient D are multiplied to obtain the semantic transition distortion degree index STD, STD=J·K·D.
[0086] It should be noted that the semantic transition distortion index is a quantitative index for measuring the transition continuity and stability of multi-source modal data at the semantic level. The semantic transition distortion index reflects whether the modal feature expression remains a natural, smooth and consistent transition process in the semantic conduction chain from the sky modality to the air modality and then to the ground modality. The lower the index, the more coherent and coordinated the semantic conversion process between the three modal data, which reflects a good fusion effect; on the contrary, a higher semantic transition distortion index value means that there is a significant break, deviation or mutation in the semantic expression between the modes, which shows the semantic distortion problem in the fusion process, thereby leading to an unsatisfactory fusion effect. The index describes the quality of semantic transition through three core measures: the semantic transition angle tension coefficient measures the consistency of the semantic change direction between the modes. If the angle between the two semantic change vectors is large, it indicates that the information transmission deviates or mutates in direction. Secondly, the semantic path curvature coefficient reflects the smoothness of the semantic conduction path. If the path has obvious bending or turning back, it means that the semantic conversion is unnatural. Finally, the semantic drift coefficient measures the matching degree of the overall semantic change and the local step change path. If the drift is large, it means that the overall semantic expression has a jump or fault and lacks continuous transmission. The product of the three forms the semantic transition distortion index. Only when the direction, path and smoothness are all abnormal, the index will significantly increase, accurately capturing the semantic distortion in modal fusion. For example, if the features extracted from the sky modality through remote sensing images and the image features captured by the air unmanned aerial vehicle have a large difference in semantic expression, and the features collected by the unmanned aerial vehicle and the ground sensor show a jump or incoherent change, the semantic transition distortion index value will be significantly high, indicating that the fusion model has defects in maintaining semantic consistency. On the contrary, if the features of the three modalities form a smooth and directionally consistent change trajectory along the semantic space, the semantic transition distortion index value tends to zero, indicating that the fusion process successfully realizes the organic coordination of multi-modal data. Therefore, the semantic transition distortion index is not only an important index for evaluating the fusion quality, but also provides a clear direction and reference for multi-modal fusion optimization.
[0087] It should be noted that the combination of "semantic transition included angle tension coefficient + semantic path curvature coefficient + semantic drift coefficient" is used to calculate the semantic transition distortion index, in order to comprehensively and carefully reveal the multi-dimensional distortion characteristics of modal data in the semantic space, and not just limited to local differences or overall deviations. The advantage of this method is that it combines direction consistency (included angle), path smoothness (curvature) and total consistency (drift) of the three key aspects, which can effectively depict the complex situation of "structural disorder" in semantic transition, and has stronger explanatory power and applicability. By using a single index such as an included angle or an Euclidean distance to measure the difference between modes, the semantic conduction relationship between modes is easily ignored. For example, although the difference between each pair of modes may not be large, the overall path may have significant bending or semantic mutation between the sky mode (such as remote sensing image), the air mode (such as unmanned aerial vehicle shooting) and the ground mode (such as vehicle-mounted camera). However, these can only be identified when the overall relationship of the three semantic transition paths is considered. Therefore, the included angle tension coefficient captures the consistency between the semantic change directions, revealing whether there is a "transition jump"; the curvature coefficient detects whether the modal path is "bent" or nonlinearly folded in the semantic space, capturing the smoothness of the conduction; and the drift coefficient further measures whether the actual semantic path is really equivalent to directly transitioning from the sky mode to the ground mode, that is, the "equivalence" of the entire path. The advantage of the above calculation is that it is more sensitive to complex multi-modal semantic fusion. For example, in an intelligent traffic application, the sky mode captures the city macro structure, the air mode focuses on traffic density, and the ground mode reflects the license plate and driving behavior characteristics - the three have different semantic levels, different scales, and different focuses. Traditional measurement methods may not accurately express the complex semantic migration from "structural layer" to "behavior layer"; and STD can make a detailed judgment on the quality of semantic transition from the direction (whether it is concentrated towards a semantic target), the path (whether it is smooth), and the result (whether it is a logical closed loop).
[0088] In one embodiment, according to the modal redundancy consistency index and the semantic transition distortion index, the step of calculating the fusion quality index is:
[0089] The modal redundancy consistency index and the semantic transition distortion index are normalized, and the normalized modal redundancy consistency index and the normalized semantic transition distortion index are weighted and summed to obtain the fusion quality index, and the formula for calculation is: GHY = a1 x zx + a2 x zc, wherein GHY is the fusion quality index, zx and zc are the normalized modal redundancy consistency index and the normalized semantic transition distortion index respectively, a1 and a2 are respectively the preset weight coefficients of the normalized modal redundancy consistency index and the normalized semantic transition distortion index, and a1 and a2 are both greater than 0;
[0090] It should be noted that the above normalization processing dimension removal method includes Min-Max normalization, Z-Score standardization, etc., which will not be repeated here; a1 and a2 are set according to actual conditions, generally a1 and a2 are equal and the sum is 1, for example, a1 and a2 can be 0.5, 0.5.
[0091] In one embodiment, the fusion quality index is compared with a preset threshold, and it is determined whether the fusion quality of the three modal data of sky and ground is qualified according to the comparison result, if qualified, the operation and inspection result is output according to the fused data.
[0092] In one embodiment, the step of comparing the fusion quality index with a preset threshold and determining whether the fusion quality of the three modal data of sky and ground is qualified according to the comparison result includes:
[0093] If the fusion quality index is less than the preset threshold, it indicates that the fusion quality of the three modal data of sky and ground is qualified, and the operation and inspection result is directly output according to the data fused by the three modal data of sky and ground;
[0094] If the fusion quality index is not less than the preset threshold, it indicates that the fusion quality of the three modal data of sky and ground is not qualified, and the three modal data of sky and ground needs to be fused again until the fusion quality of the three modal data of sky and ground is qualified.
[0095] It should be noted that if the fusion quality index is less than or equal to the threshold, it indicates that the redundancy consistency between the modes is good, the semantic transition is smooth, and the distortion degree is within a controllable range, and the fusion result has high reliability and stability, so the final operation and inspection result can be generated based on the fusion data without further correction or intervention; on the contrary, if the fusion quality index is greater than the preset threshold, it indicates that there is significant semantic break, structural incoordination or redundancy conflict between the three modes, the current fusion result cannot reliably reflect the real characteristics of the target object, and there is a risk of misjudgment, so the re-execution mechanism of the data fusion process should be triggered. The mechanism needs to adjust or optimize the current modal feature fusion strategy, such as re-extracting features, selecting different fusion algorithms, optimizing weight distribution or adjusting the order of modal alignment, etc., until the calculated fusion quality index falls within the preset threshold range again, ensuring that the fused data has sufficient accuracy, integrity and stability, thereby supporting the subsequent operation and inspection tasks to realize intelligent identification and processing with high reliability.
[0096] Based on the same inventive concept, the embodiments of the present application also provide a multi-source collaborative sky-ground comprehensive operation and inspection system. Referring to Figure 2 , Figure 2 A framework diagram of a multi-source collaborative sky-ground comprehensive operation and inspection system is provided for the embodiments of the present application, and the system includes:
[0097] A modal redundancy consistency module: data of the sky modal, the air modal and the ground modal are collected, and corresponding feature information is extracted, and a modal redundancy consistency index is calculated based on the feature information between different modes;
[0098] A semantic transition distortion module: a feature transition path between modal feature information is analyzed, and a semantic transition distortion degree index is calculated;
[0099] A fusion quality module: a fusion quality index is calculated according to the modal redundancy consistency index and the semantic transition distortion degree index;
[0100] An operation detection module: the fusion quality index is compared with a preset threshold, and whether the fusion quality of the sky-ground three modal data is qualified is determined according to a comparison result, if qualified, an operation detection result is output according to the fused data.
[0101] Based on the multi-source collaborative comprehensive operation detection system provided in the embodiment of the application, the fusion effect of the sky-ground three modal data can be judged, whether the fusion effect of the sky-ground three modal data is available can be judged, the effectiveness of data fusion is ensured, the misjudgment, omission and even systematic deviation caused in the downstream analysis, identification and decision-making links are reduced, and the reliability and stability of the overall operation detection task are ensured.
[0102] In one embodiment, the modal redundancy consistency module comprises:
[0103] A feature vector module: normalized feature vectors representing the same target object in the sky modal, the air modal and the ground modal are collected, and are respectively denoted as a sky modal feature vector F T , an air modal feature vector F A and a ground modal feature vector F G .
[0104] A construction module: a modal difference matrix R is constructed, and an element R ij of the modal difference matrix R represents a normalized difference degree between modal i and modal j, and is defined as: In the formula, i and j are both included in {T, A, G}, and ||·||2 represents a two-norm of a vector;
[0105] An information structure asymmetry error module: a difference between the difference degrees of two modes in the modal difference matrix is squared and summed to obtain an information structure asymmetry error of each mode;
[0106] A structure conflict degree module: the information structure asymmetry errors of all modes are squared and accumulated to obtain a structure conflict degree index S of the modes as a whole;
[0107] A direction consistency module: the included angle cosine between each two of the three groups of feature vectors is calculated, and an average value thereof is taken to obtain a redundancy collaborative direction consistency index C between the modes.
[0108] Modality Redundancy Consistency Index Module: Calculate the modality redundancy consistency index according to the overall structure conflict index S and the redundancy collaborative direction consistency index C, and the formula is: In the formula, MRI is the modality redundancy consistency index.
[0109] In an embodiment, the semantic transition distortion module includes:
[0110] Feature Vector Module: Collect normalized feature vectors representing the same target object in the sky modality, the aerial modality, and the ground modality, denoted as the sky modality feature vector F T , the aerial modality feature vector F A , and the ground modality feature vector F G , respectively.
[0111] Semantic Transition Difference Vector Module: Calculate the inter-modality semantic transition difference vector, ΔTA=F A -F T , ΔAG=F G -F A , where ΔTA is the semantic transition difference vector from the sky modality to the aerial modality, and ΔAG is the semantic transition difference vector from the aerial modality to the ground modality.
[0112] Transition Angle Tension Module: Calculate the semantic transition angle tension coefficient J, In the formula, <·,·> is the vector dot product, and ||·||2 represents the two-norm of the vector.
[0113] Semantic Path Curvature Module: Calculate the semantic path curvature coefficient K,
[0114] Semantic Drift Module: Calculate the semantic drift coefficient D, and the steps are: ΔTG=F G -F T , where ΔTG represents the total semantic transition difference vector from the sky modality directly to the ground modality.
[0115] Semantic Transition Distortion Index Module: Multiply the semantic transition angle tension coefficient J, the semantic path curvature coefficient K, and the semantic drift coefficient D to obtain the semantic transition distortion index.
[0116] In an embodiment, the fusion quality module specifically includes:
[0117] Normalize the modality redundancy consistency index and the semantic transition distortion index, and perform weighted summation on the normalized modality redundancy consistency index and the semantic transition distortion index to obtain the fusion quality index.
[0118] In one embodiment, the operation inspection module comprises:
[0119] The first judging module: if the fusion quality index is less than the preset threshold, it indicates that the fusion quality of the three modal data of sky and ground is qualified, and the operation inspection result is directly output according to the data fused by the three modal data of sky and ground;
[0120] The second judging module: if the fusion quality index is not less than the preset threshold, it indicates that the fusion quality of the three modal data of sky and ground is unqualified, and the three modal data of sky and ground need to be fused again until the fusion quality of the three modal data of sky and ground is qualified
[0121] The above has carried out the detailed description to one embodiment of the application, but the content is only the preferred embodiment of the application, cannot be considered for limiting the implementation range of the application. All equivalent changes and improvements made according to the application scope should still belong to the patent coverage range of the application.
Claims
1. A method based on multi-source collaborative sky-ground comprehensive operation and inspection, characterized in that, The method comprises the following steps: Collecting data of the sky mode, the air mode and the ground mode, extracting corresponding feature information, and calculating a modal redundancy consistency index based on the feature information between different modes; Analyzing feature transition paths between feature information of each mode, and calculating a semantic transition distortion degree index; Calculating a fusion quality index according to the modal redundancy consistency index and the semantic transition distortion degree index; Comparing the fusion quality index with a preset threshold, and determining whether the fusion quality of the three modes of the sky, the air and the ground is qualified according to a comparison result, and if qualified, outputting an operation inspection result according to fused data.
2. The multi-source collaborative sky-ground integrated operation and inspection method according to claim 1, characterized in that, The step of calculating the modal redundancy consistency index based on the feature information between different modes comprises: Collect normalized feature vectors representing the same target object in the sky modality, the aerial modality, and the ground modality, denoted as sky modality feature vector F T , aerial modality feature vector F A , and ground modality feature vector F G , respectively; A modal difference matrix R is constructed, whose elements R ij denotes the normalized difference between modal i and modal j, defined as: where i and j ∈ {T, A, G}, and ||·||2 denotes the 2-norm of a vector. Squaring and summing up difference values of difference degrees between corresponding two modes in a modal difference matrix to obtain information structure asymmetry errors of each mode; Squaring and accumulating information structure asymmetry errors of all modes to obtain a structure conflict degree index S of the whole mode; Calculating a cosine of an included angle between two of the three groups of feature vectors, and taking an average value to obtain a redundancy cooperative direction consistency index C between the modes; The modal redundancy consistency index is calculated according to the overall structure conflict index S and the redundancy cooperative direction consistency index C, and the formula is: In the formula, MRI is the modal redundancy consistency index.
3. The multi-source collaborative sky-ground integrated operation and inspection method according to claim 1, characterized in that, The step of calculating the semantic transition distortion degree index comprises: Collect normalized feature vectors representing the same target object in the sky modality, the aerial modality, and the ground modality, denoted as sky modality feature vector F T , aerial modality feature vector F A , and ground modality feature vector F G , respectively. computing the inter-modal semantic transition difference vector, ΔTA = F A -F T , ΔAG = F G -F A , where ΔTA is the semantic transition difference vector from the sky modality to the air modality, and ΔAG is the semantic transition difference vector from the air modality to the ground modality. computing the semantic transition included angle tension coefficient J, where <·, ·> is the vector dot product and ||·||2 denotes the two-norm of a vector. calculating a semantic path curvature coefficient K, A semantic drift coefficient D is calculated, the steps of which are: ΔTG = F G -F T ΔTG represents a semantic transition difference total vector from the sky modality directly to the ground modality, Multiplying a semantic transition included angle tension coefficient J, a semantic path curvature coefficient K and a semantic drift coefficient D to obtain the semantic transition distortion degree index.
4. The multi-source collaborative sky-ground integrated operation and inspection method according to claim 1, characterized in that, The step of calculating the fusion quality index according to the modal redundancy consistency index and the semantic transition distortion degree index comprises: Normalizing the modal redundancy consistency index and the semantic transition distortion degree index, and performing weighted summation on the normalized modal redundancy consistency index and the semantic transition distortion degree index to obtain the fusion quality index.
5. The multi-source collaborative sky-ground integrated operation and inspection method according to claim 1, characterized in that, The step of comparing the fusion quality index with the preset threshold and determining whether the fusion quality of the three modes of the sky, the air and the ground is qualified according to a comparison result comprises: If the fusion quality index is less than the preset threshold, it indicates that the fusion quality of the three modes of the sky, the air and the ground is qualified, and an operation inspection result is directly output according to fused data of the three modes of the sky, the air and the ground; If the fusion quality index is not less than the preset threshold, it indicates that the fusion quality of the three modes of the sky, the air and the ground is not qualified, and the three modes of the sky, the air and the ground need to be fused again until the fusion quality of the three modes of the sky, the air and the ground is qualified.
6. A multi-source collaborative sky-ground integrated operation and inspection system based on, characterized in that, The system comprises: A modal redundancy consistency module: collecting data of the sky mode, the air mode and the ground mode, extracting corresponding feature information, and calculating a modal redundancy consistency index based on the feature information between different modes; A semantic transition distortion module: analyzing feature transition paths between feature information of each mode, and calculating a semantic transition distortion degree index; A fusion quality module: calculating a fusion quality index according to the modal redundancy consistency index and the semantic transition distortion degree index; An operation inspection module: comparing the fusion quality index with a preset threshold, and determining whether the fusion quality of the three modes of the sky, the air and the ground is qualified according to a comparison result, and if qualified, outputting an operation inspection result according to fused data.
7. The multi-source collaborative sky-ground integrated operation and inspection system according to claim 6, characterized in that, The modal redundancy consistency module comprises: Feature vector module: collect normalized feature vectors representing the same target object in the sky modality, the aerial modality and the ground modality, denoted as sky modality feature vector F T , aerial modality feature vector F A , and ground modality feature vector F G , respectively; Constructing the module: construct the modal difference matrix R, whose elements R ij denote the normalized difference degree between modal i and modal j, defined as: where i and j ∈ {T, A, G}, ||·||2 denotes the two-norm of the vector. An information structure asymmetry error module: squaring and summing up difference values of difference degrees between corresponding two modes in a modal difference matrix to obtain information structure asymmetry errors of each mode; The structure conflict degree module squares and accumulates the information structure asymmetry errors of all modes to obtain a structure conflict degree index S of the modes as a whole; The direction consistency module calculates the cosine of the angle between each two of the three sets of eigenvectors and obtains a redundant cooperative direction consistency index C between the modes by averaging the cosine of the angle; The modal redundancy consistency index module: according to the overall structure conflict index S and the redundancy cooperative direction consistency index C, the modal redundancy consistency index is calculated, and the formula is: In the formula, MRI is the modal redundancy consistency index.
8. The multi-source collaborative sky-ground integrated operation and inspection system according to claim 6, characterized in that, The semantic transition distortion module includes: Feature vector module: collect normalized feature vectors representing the same target object in the sky modality, the aerial modality and the ground modality, denoted as sky modality feature vector F T , aerial modality feature vector F A , and ground modality feature vector F G , respectively; Semantic transition difference vector module: calculate inter-modality semantic transition difference vector, ΔTA = F A -F T , ΔAG = F G -F A , in the formula, ΔTA is the semantic transition difference vector from the sky modality to the air modality, and ΔAG is the semantic transition difference vector from the air modality to the ground modality. Transition angle tension module: calculate the semantic transition angle tension coefficient J, where <·, ·> is the vector dot product and ||·||2 denotes the two-norm of a vector. semantic path curvature module: calculating the semantic path curvature coefficient K, semantic drift module: calculating a semantic drift coefficient D, the steps of the calculation being: ΔTG = F G -F T ΔTG represents the semantic transition differential total vector from the sky modality directly to the ground modality, The semantic transition distortion degree index module multiplies the semantic transition angle tension coefficient J, the semantic path curvature coefficient K and the semantic drift coefficient D to obtain a semantic transition distortion degree index.
9. The multi-source collaborative sky-ground integrated operation and inspection system according to claim 6, characterized in that, The fusion quality module specifically includes: The modal redundant consistency index and the semantic transition distortion degree index are normalized, and the normalized modal redundant consistency index and the semantic transition distortion degree index are weighted and summed to obtain a fusion quality index.
10. The multi-source collaborative sky-ground integrated operation and inspection system according to claim 6, characterized in that, The operation and inspection module includes: The first judgment module: if the fusion quality index is less than a preset threshold, it indicates that the fusion quality of the three modes of sky, ground and sky-ground is qualified, and the operation and inspection result is directly output according to the data fused by the three modes of sky, ground and sky-ground; The second judgment module: if the fusion quality index is not less than the preset threshold, it indicates that the fusion quality of the three modes of sky, ground and sky-ground is not qualified, and the three modes of sky, ground and sky-ground need to be fused again until the fusion quality of the three modes of sky, ground and sky-ground is qualified.
Citation Information
Patent Citations
Ultra-short-term photovoltaic power prediction method and system based on sky-ground multi-modal data fusion
CN119312995A
Improved multi-modal information fusion extraction method and system
CN119885061A
Multi-source data fusion method and system based on cloud computing
CN119989267A
Big data intelligent processing method and system based on space-air-ground integration
CN120067645A
Data processing method and system for multi-source complex biological information data
CN120148619A