Inspection method, device and equipment based on digital airspace system and medium
By using pre-trained neural network models, MoE architecture, variational autoencoder and multimodal transformer decoder in digital airspace systems, multi-source perceptual data is processed and integrated, and data errors and inefficiency in drone inspections are solved, and more efficient and intelligent inspection decisions are achieved.
Patent Information
- Application Number
- CN202510451229.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-11
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-04-11
AI Technical Summary
Digital airspace systems face problems such as multi-source perceived data errors, conflicts and sensor failures during UAV inspections, resulting in inefficient inspections.
The pre-trained neural network model is used to process multi-source perception monitoring data, generate structured semantic descriptions and timing measurements, and perform multimodal fusion through a digital airspace expert model based on MoE architecture. Then, the characterization is enhanced using a variational autoencoder, and a digital airspace knowledge graph is constructed through a multimodal transformer decoder, and finally the conditions are used to generate patrol decision-making suggestions.
It improves the utilization rate of multi-source perceived data, enhances the richness and accuracy of data representation, and improves the efficiency of drone patrol and the intelligence level of decision-making.
Smart Images

Figure CN119989283A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of information technology, and in particular to an inspection method, device, equipment and medium based on a digital airspace system. Background Art
[0002] The digital airspace system is an indispensable component of drone power inspections. It is a digital platform that updates air information in real time through sensors on the ground and on drones. It is the basis for realizing automated and intelligent drone inspections.
[0003] At present, the digital airspace system provides drones with all-round environmental perception capabilities by integrating and processing multi-source perception data from heterogeneous sensors and external sources. However, these information sources are different, and when used directly, there will be problems such as perception data errors, conflicts, and even sensor failures, which can easily mislead the inspection tasks and reduce the inspection efficiency of drones.
[0004] It can be seen that how to improve the utilization rate of multi-source perception data used by digital airspace systems for drone inspections in order to improve the inspection efficiency of drones has become a technical problem that technical personnel in this field need to solve urgently. Summary of the invention
[0005] The present invention provides an inspection method, device, equipment and medium based on a digital airspace system, which solves the problem of how to improve the utilization rate of multi-source perception data adopted by the digital airspace system when inspecting unmanned aerial vehicles, so as to improve the inspection efficiency of unmanned aerial vehicles.
[0006] In order to solve the above technical problems, the first aspect of the present invention provides an inspection method based on a digital airspace system, comprising: Obtain multi-source sensing monitoring data and external meteorological space data of the area to be inspected; According to the type of the multi-source sensing monitoring data, a plurality of pre-trained neural network models are used to process the multi-source sensing monitoring data to obtain a structured semantic description and a time series measurement value; Inputting the external meteorological spatial data, the structured semantic description and the time series measurement value into a digital airspace expert model based on the MoE architecture for fusion processing to obtain a multimodal fusion representation; The multimodal fusion representation is enhanced by a variational autoencoder, and the obtained multimodal fusion enhanced representation is input into a multimodal transformer decoder for decoding to obtain a set of entity triples to construct a digital spatial domain knowledge graph; The digital airspace knowledge graph is processed using an attention enhancement mechanism based on a graph neural network to obtain attention-enhanced node representations that are input into a conditional generative adversarial network model for task prediction, generating inspection decision recommendations for the area to be inspected so that the drone can execute them.
[0007] As one of the preferred solutions, the multi-source sensing monitoring data includes unstructured multi-source sensing monitoring data, and the unstructured multi-source sensing monitoring data includes unstructured point cloud data and unstructured image data; wherein, According to the type of the multi-source sensing monitoring data, a plurality of pre-trained neural network models are used to process the multi-source sensing monitoring data to obtain a structured semantic description and a time series measurement value, including: Extracting features from the unstructured point cloud according to a pre-trained 3D convolutional neural network model and a spatial attention mechanism to obtain geometric features; Extracting features from the unstructured image data based on a pre-trained residual neural network model to obtain visual features; Using a cross-modal encoder to align and fuse the geometric features and the visual features to generate a fused feature; The fused features are converted into the structured semantic description through a pre-trained multimodal transformer decoder.
[0008] As one of the preferred solutions, the multi-source sensing monitoring data also includes structured multi-source sensing time series monitoring data; wherein, The method further comprises: processing the multi-source sensing monitoring data using a plurality of pre-trained neural network models according to the type of the multi-source sensing monitoring data to obtain a structured semantic description and a time series measurement value; Using a sliding time window to convert the structured multi-source sensing time series monitoring data into a time series structure sequence; Inputting the temporal structure sequence into a pre-trained deep learning network model to calculate the similarity between the temporal structure sequence and the historical temporal structure sequence through an attention mechanism to obtain a similarity distribution; Generate a data deviation correction matrix based on the error characteristics of the sensor corresponding to the structured multi-source sensing time series monitoring data and the similarity distribution; The data deviation correction matrix is applied to the structured multi-source sensing time series monitoring data to obtain corrected time series measurement values.
[0009] As one of the preferred solutions, the digital airspace expert model includes several expert network models and a gating network; wherein, The training process of the digital airspace expert model includes: The historical external meteorological spatial data, the historical structured semantic description and the historical time series measurement values are used as training set data, and the corresponding multi-task labels are annotated for the training set data; the multi-task labels include airspace status, environmental risk level, regional weather classification and inspection decision suggestions; Align and enhance the training set data with multiple task labels to form an input data set to be passed to the gating network to obtain the output probability; Selecting a corresponding expert network model according to the output probability to process the input data set, obtaining a corresponding output result, and performing weighted summation based on the output probability assigned by the gating network to form a multimodal fusion representation prediction result; A loss function is used to evaluate the loss between the multimodal fusion representation prediction result and the multi-task label, and the gradient of the loss relative to the model parameters is calculated through a back-propagation algorithm to update the parameters of each of the expert network models and the gating network.
[0010] As one of the preferred solutions, the multimodal transformer decoder includes an enhanced self-attention layer, a hierarchical entity recognition layer and a relationship extraction network layer; wherein, The obtained multimodal fusion enhanced representation is input into the multimodal transformer decoder for decoding to obtain a set of entity triples to construct a digital spatial domain knowledge graph, including: Inputting the multimodal fusion enhanced representation into the enhanced self-attention layer for processing to capture the long-range dependencies in the multimodal fusion enhanced representation; The long-range dependency relationship is subjected to entity recognition and extraction processing by the hierarchical entity recognition layer to obtain basic entities; the basic entities include meteorological elements, geographical features and equipment status; Extracting the relationship between the basic entities based on the relationship extraction network layer, and generating a standardized entity triple set and its corresponding confidence score; the relationship includes spatial position relationship, causal relationship and temporal relationship; The entity triple set is converted into a natural language description through a preset prompt template, and a query prompt is constructed in combination with domain knowledge; Inputting the preset prompt template and the query prompt into a large language model for knowledge reasoning, and performing quality control on each reasoning result using each confidence score as a weight; The reasoning results output by the large language model are integrated to form a unified knowledge system to construct a digital spatial knowledge graph.
[0011] As one of the preferred solutions, the graph neural network includes a gated attention layer, a temporal attention layer and a residual connection layer; wherein, The attention enhancement mechanism based on the graph neural network is used to process the digital spatial domain knowledge graph to obtain the attention-enhanced node representation, including: Standardizing the edges and nodes in the digital spatial knowledge graph, and constructing a multi-scale neighborhood feature aggregator for each node to aggregate the structural feature representations of each node from local to global at multiple scales; Calculating the task attention score of each of the structural feature representations through the gated attention layer, and performing weighted summation on each of the structural feature representations based on each of the task attention scores to obtain a fused scale feature representation; Calculating the temporal attention score of each of the structural feature representations based on the temporal attention layer, and performing weighted summation on each of the structural feature representations according to each of the temporal attention scores to obtain a fused temporal feature representation; The fused scale feature representation and the fused time feature representation are processed by the residual connection layer to realize online updating of the node representation, and the processing steps of the gated attention layer, the temporal attention layer and the residual connection layer are iteratively executed to obtain the node representation after attention enhancement.
[0012] As one of the preferred solutions, the conditional generative adversarial network model includes a generator having an encoding layer, a decoding layer and a multi-task branch layer, and a Transformer-based multi-task discriminator; wherein, The node representation with enhanced attention is input into the conditional generative adversarial network model for task prediction, and an inspection decision suggestion for the area to be inspected is generated, including: The encoding layer is used to encode the node representation to obtain an encoding result, and the encoding result and random noise are mapped to a target output space through the decoding layer to obtain a decoding output result; The decoding output result is processed by the multi-task branch layer to obtain the airspace state prediction result, regional weather prediction result, environmental risk prediction result and inspection decision suggestion prediction result of the area to be inspected; The airspace status prediction results, the regional weather prediction results, the environmental risk prediction results and the inspection decision recommendation prediction results are evaluated according to the multi-task discriminator, and the inspection decision recommendation prediction results are adjusted based on the evaluation results to obtain inspection decision recommendations for the area to be inspected.
[0013] A second aspect of the present invention provides an inspection device based on a digital airspace system, comprising: A data acquisition module is used to acquire multi-source sensing monitoring data and external meteorological space data of the area to be inspected; A data processing module, used to process the multi-source sensing monitoring data using a plurality of pre-trained neural network models according to the type of the multi-source sensing monitoring data to obtain a structured semantic description and a time series measurement value; A feature fusion module, used for inputting the external meteorological spatial data, the structured semantic description and the time series measurement value into the digital airspace expert model based on the MoE architecture for fusion processing to obtain a multimodal fusion representation; A graph construction module is used to enhance the multimodal fusion representation through a variational autoencoder, and input the obtained multimodal fusion enhanced representation into a multimodal transformer decoder for decoding to obtain a set of entity triples to construct a digital spatial domain knowledge graph; The decision generation module is used to process the digital airspace knowledge graph using an attention enhancement mechanism based on a graph neural network, obtain attention-enhanced node representations to input into a conditional generative adversarial network model for task prediction, and generate inspection decision recommendations for the area to be inspected so that the drone can execute them.
[0014] A third aspect of the present invention provides an electronic device, comprising a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor, wherein the processor implements the inspection method based on the digital airspace system as described above when executing the computer program.
[0015] A fourth aspect of the present invention provides a computer-readable storage medium, which includes a stored computer program, wherein when the device where the computer-readable storage medium is located executes the computer program, the inspection method based on the digital airspace system as described above is implemented.
[0016] Compared with the prior art, the embodiments of the present invention have the following advantages: (1) Pre-trained neural network models are used to process different types of data, which can efficiently and accurately extract structured semantic descriptions and time series measurements, avoiding the time-consuming and complex process of training models from scratch. At the same time, meteorological data is taken into consideration, making the decision-making process closer to the actual environment and improving the accuracy and practicality of decision-making. (2) Through the MoE architecture, the most appropriate expert model can be selected for processing according to the characteristics of different data types, achieving effective fusion of multimodal data and improving the richness and accuracy of data representation. The use of variational autoencoders for enhanced processing can further improve the robustness and generalization ability of data representation, providing a stable and reliable foundation for the subsequent decoding process. (3) The enhanced multimodal fusion representation is decoded by the multimodal transformer decoder, which can accurately extract entity triplets, providing strong technical support for the construction of a digital spatial knowledge graph, and also helping to explore the intrinsic connections between data and provide a basis for subsequent decision-making; the digital spatial knowledge graph is processed through the attention mechanism, which can highlight key nodes and relationships and improve the accuracy and pertinence of node representation. Combined with the attention-enhanced node representation, the conditional generative adversarial network model can generate inspection decision recommendations that are closer to reality, avoiding the limitations of human experience and improving the level of intelligent decision-making. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] In order to more clearly illustrate the technical solution of the present invention, the drawings required for use in the implementation mode will be briefly introduced below. Obviously, the drawings described below are only some implementation modes of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0018] Figure 1 It is a flow chart of an inspection method based on a digital airspace system provided by a certain embodiment of the present invention; Figure 2 It is a structural diagram of a patrol device based on a digital airspace system provided by a certain embodiment of the present invention; Figure 3 It is a structural diagram of an electronic device provided by a certain embodiment of the present invention. DETAILED DESCRIPTION
[0019] The following is a clear and complete description of the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings and embodiments. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all of the embodiments. The purpose of providing these embodiments is to make the disclosure of the present invention more thorough and comprehensive. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0020] In the description of this application, the terms "first", "second", "third", etc. are used for descriptive purposes only and are not to be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Thus, a feature defined as "first", "second", "third", etc. may explicitly or implicitly include one or more of the feature. In the description of this application, unless otherwise specified, "plurality" means two or more.
[0021] In the description of the present application, it should be noted that, unless otherwise clearly specified and limited, the terms "installed", "connected", and "connected" should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection, or it can be indirectly connected through an intermediate medium, or it can be a connection between the two elements. The terms "vertical", "horizontal", "left", "right", "upper", "lower" and similar expressions used herein are only for illustrative purposes, and do not indicate or imply that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as a limitation on the present invention. The term "and / or" used herein includes any and all combinations of one or more related listed items. For those of ordinary skill in the art, the specific meanings of the above terms in this application can be understood according to specific circumstances.
[0022] In the description of this application, it should be noted that, unless otherwise defined, all technical and scientific terms used in the present invention have the same meaning as those commonly understood by those skilled in the art. The terms used in the specification of the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention. For those of ordinary skill in the art, the specific meanings of the above terms in this application can be understood by specific circumstances.
[0023] In one embodiment, if Figure 1 As shown, the first aspect of the present invention provides an inspection method based on a digital airspace system, comprising: S1. Obtain multi-source sensing monitoring data and external meteorological space data of the area to be inspected; S2. According to the type of the multi-source sensing monitoring data, use a plurality of pre-trained neural network models to process the multi-source sensing monitoring data to obtain a structured semantic description and a time series measurement value; S3, inputting the external meteorological spatial data, the structured semantic description and the time series measurement value into the digital airspace expert model based on the MoE architecture for fusion processing to obtain a multimodal fusion representation; S4, enhancing the multimodal fusion representation through a variational autoencoder, and inputting the obtained multimodal fusion enhanced representation into a multimodal transformer decoder for decoding to obtain a set of entity triples to construct a digital spatial domain knowledge graph; S5. Use the attention enhancement mechanism based on graph neural network to process the digital airspace knowledge graph, obtain the attention-enhanced node representation to input into the conditional generative adversarial network model for task prediction, and generate inspection decision suggestions for the area to be inspected so that the drone can execute.
[0024] Specifically, the present invention updates the multi-source heterogeneous perception monitoring data of the area to be inspected in real time through various sensors carried by the ground and unmanned aerial vehicles (such as ground equipment including electromagnetic radiation meters, wireless network cards, climate observation instruments, anemometers, etc., and airborne equipment including wireless network cards, ADS-B (Automatic Dependent Surveillance - Broadcast) equipment, lidars, and cameras); wherein the multi-source perception monitoring data includes unstructured multi-source perception monitoring data (and unstructured multi-source perception monitoring data includes unstructured point cloud data and unstructured image data, i.e., point cloud data collected by lidars and high-definition images / video streams acquired by cameras) and structured multi-source perception time series monitoring data (i.e., time series data collected by electromagnetic radiation meters, wireless network cards, climate observation instruments, anemometers, lidars, etc.). The present invention also introduces external meteorological spatial data accessed through the Geographic Information System (GIS), meteorological platform, and ground station: meteorological knowledge data (climate knowledge base built into the digital airspace system and real-time weather data provided by the meteorological service API, etc.), electronic map data, remote sensing data, airspace restriction information (including management information such as no-fly zones, restricted airspaces, and temporary flight airspaces), terrain and geomorphic data (including information such as height, slope, and water distribution), etc. The present invention also takes multi-source heterogeneous perception data and meteorological data into consideration, making the decision-making process closer to the actual environment to improve the accuracy and practicality of decision-making.
[0025] According to the type of multi-source perception monitoring data (structured and unstructured), the corresponding pre-trained neural network model is selected for processing. For example, for image data, a convolutional neural network (CNN) can be used for feature extraction and classification; for video data, a three-dimensional convolutional neural network (3D CNN) or a recurrent neural network (RNN) can be used for processing; and the processed data is subjected to structured semantic description and time series measurement value extraction to provide a basis for subsequent data fusion. The pre-trained neural network model has high recognition accuracy and generalization ability, and can accurately process multi-source perception monitoring data.
[0026] The external meteorological spatial data, structured semantic descriptions and time series measurements are preprocessed to ensure the consistency of data format and dimension, and these preprocessed data are input into the digital airspace expert model based on the MoE architecture (Mixture of Experts). The model consists of multiple expert network models and a gating network. Each expert network model is good at processing a specific type of data. The appropriate expert network model is selected through the gating network to process the data to obtain a multimodal fusion representation. Among them, the MoE architecture improves the flexibility and scalability of the model and can process various types of data, while the multimodal fusion representation integrates information from multiple data sources, improves the utilization rate of multi-source perception data, and also improves the accuracy and reliability of subsequent decision-making.
[0027] The multimodal fusion representation is enhanced by using a variational autoencoder (VAE): that is, the encoder network of the VAE is used to map it to the probability distribution of the latent space, and the latent variables are sampled from the distribution. The decoder network of the VAE takes the latent variables as conditions and simultaneously reconstructs the original features of multiple modalities, including sensor numerical sequences, image features, point cloud features, etc.; then, by minimizing the joint loss function of the reconstruction error and the KL divergence, the fusion representation is forced to learn the essential features and intrinsic associations of different modal data, and the multimodal fusion enhanced representation can be obtained. The present invention uses this self-supervised learning method to significantly enhance the integrity and robustness of the representation, provide a more reliable feature basis for subsequent semantic parsing and knowledge extraction, so as to extract the potential features of the data and improve the robustness and generalization ability of the multimodal fusion representation; wherein, in the autoencoder selection part, β-VAE can be used instead of the standard VAE, and VQ-VAE can be used for discrete representation learning. Preferably, conditional VAE is used to realize representation learning of specific scenarios.
[0028] The enhanced multimodal fusion representation is input into the multimodal transformer decoder, and the enhanced self-attention layer, hierarchical entity recognition layer and relationship extraction network layer are used to gradually decode and refine the input data to obtain a set of entity triples for constructing a digital spatial knowledge graph; and after the digital spatial knowledge graph is represented as graph structure data, the graph structure data is processed using an attention enhancement mechanism based on a graph neural network to extract node representations, and then the node representations are weighted by the attention mechanism to obtain attention-enhanced node representations to be input into the conditional generative adversarial network model to generate inspection decision suggestions for the inspection area, and finally the inspection decision suggestions are input into the drone control system to guide the drone to perform inspection tasks. The present invention improves the accuracy and reliability of node representations through the attention enhancement mechanism, provides a basis for subsequent task predictions, and uses the conditional generative adversarial network model to generate task prediction results related to the input data, thereby improving the accuracy and diversity of predictions.
[0029] The present invention integrates multi-source perception monitoring data and external meteorological space data, and adopts advanced neural network models, MoE architecture, variational autoencoder, multimodal transformer decoder, graph neural network attention enhancement mechanism, and conditional generative adversarial network model and other technical means to achieve comprehensive monitoring, accurate decision-making and autonomous inspection of the inspection area, which has significant beneficial effects in improving inspection efficiency, reducing labor costs, and ensuring inspection safety.
[0030] In one embodiment, step S2 includes: Extracting features from the unstructured point cloud according to a pre-trained 3D convolutional neural network model and a spatial attention mechanism to obtain geometric features; Extracting features from the unstructured image data based on a pre-trained residual neural network model to obtain visual features; Using a cross-modal encoder to align and fuse the geometric features and the visual features to generate a fused feature; The fused features are converted into the structured semantic description through a pre-trained multimodal transformer decoder.
[0031] Specifically, for unstructured multi-source perception monitoring data, the present invention uses a pre-trained 3D convolutional neural network model combined with a spatial attention mechanism to perform feature extraction on unstructured point cloud data. The spatial attention mechanism is used to guide the model to focus on important geometric structure information in the point cloud, thereby obtaining geometric features; and a pre-trained residual neural network model (ResNet, Residual Neural Network) is used to perform feature extraction on unstructured image data, so as to solve the gradient vanishing and gradient explosion problems of deep neural networks during the training process by introducing a residual structure, thereby extracting deep visual features in the image; and then a cross-modal encoder is used to align and fuse the two generated features to learn the intrinsic connection between features of different modalities and fuse them into a unified representation, namely, fused features; finally, a pre-trained multimodal transformer decoder is used to convert the fused features into structured semantic descriptions, namely, by pre-training the multimodal transformer decoder on large-scale text data, the deep bidirectional representation of the language is captured to generate accurate and coherent semantic descriptions. Among them, ViT can be used instead of ResNet in the multimodal feature extraction part, or SwinTransformer can be used for hierarchical feature extraction, such as ConvNeXt and other new convolutional network architectures; for the processing of point cloud data, PointNet++ can also be used instead of 3D convolutional network. Optionally, DGCNN can be used for dynamic graph convolution processing. Furthermore, Point Transformer can be used to process point cloud data.
[0032] The present invention uses pre-trained 3D convolutional neural networks and residual neural networks to efficiently extract geometric features and visual features from unstructured point clouds and image data respectively, and uses the generalization ability of pre-trained models to improve the accuracy and efficiency of feature extraction; the fused features generated by processing geometric features and visual features through a cross-modal encoder provide a rich information basis for subsequent semantic description; the fused features are converted into structured semantic descriptions through a pre-trained multimodal transformer decoder, realizing the conversion of unstructured data to structured data, which is convenient for subsequent analysis and application.
[0033] In one embodiment, step S2 further includes: Using a sliding time window to convert the structured multi-source sensing time series monitoring data into a time series structure sequence; Inputting the temporal structure sequence into a pre-trained deep learning network model to calculate the similarity between the temporal structure sequence and the historical temporal structure sequence through an attention mechanism to obtain a similarity distribution; Generate a data deviation correction matrix based on the error characteristics of the sensor corresponding to the structured multi-source sensing time series monitoring data and the similarity distribution; The data deviation correction matrix is applied to the structured multi-source sensing time series monitoring data to obtain corrected time series measurement values.
[0034] Specifically, for structured multi-source perception time series monitoring data, the present invention constructs a deep learning network model, LSTM (Long Short Term Memory) model, and establishes a normal distribution pattern of sensor data by learning the time series characteristics and distribution laws of historical data. First, the structured multi-source perception time series monitoring data is converted into a time series sequence using a sliding time window and input into the LSTM model, so that the model calculates the similarity between the current input data and the historical data distribution through the attention mechanism, and analyzes the error range and law of each sensor in combination with the error characteristics of the sensor, and then evaluates the degree to which each time series structure sequence is affected by the error, and generates a data deviation correction matrix to correct the error in the time series measurement value; finally, the correction matrix is applied to the structured multi-source perception time series monitoring data, and the time series measurement values are corrected one by one according to the correction coefficients in the correction matrix, and the corrected time series measurement values are output for subsequent analysis; wherein, Temporal Convolutional Network can be used instead of LSTM, and optionally, Informer is used to process long sequence time series data, and preferably, Transformer-XL is used to achieve longer-term dependency modeling.
[0035] Through the sliding time window and deep learning network model, the present invention can more effectively capture the dynamic changes of time series data, improve the accuracy of data similarity calculation, thereby generating a more accurate data deviation correction matrix and improving the accuracy of time series measurement values; using the pre-trained deep learning network model, it can make full use of the learning results of existing data, quickly adapt to the processing requirements of new data, and enhance the generalization ability of the model; the entire processing flow is highly automated, which reduces manual intervention and improves processing efficiency and accuracy.
[0036] In one embodiment, the digital airspace expert model includes a plurality of expert network models and a gating network; wherein the training process of the digital airspace expert model includes: The historical external meteorological spatial data, the historical structured semantic description and the historical time series measurement values are used as training set data, and the corresponding multi-task labels are annotated for the training set data; the multi-task labels include airspace status, environmental risk level, regional weather classification and inspection decision suggestions; Align and enhance the training set data with multiple task labels to form an input data set to be passed to the gating network to obtain the output probability; Selecting a corresponding expert network model according to the output probability to process the input data set, obtaining a corresponding output result, and performing weighted summation based on the output probability assigned by the gating network to form a multimodal fusion representation prediction result; A loss function is used to evaluate the loss between the multimodal fusion representation prediction result and the multi-task label, and the gradient of the loss relative to the model parameters is calculated through a back-propagation algorithm to update the parameters of each of the expert network models and the gating network.
[0037] Specifically, for the training of the digital airspace expert model, it needs to be initialized first. The expert network model is the core component of the MoE model. Each expert network model can be an independent small neural network or a more complex structure. In the initialization stage, it is necessary to select a suitable network structure and parameters for each expert network model; the gating network is responsible for deciding which expert or experts should process the input data. In the initialization stage, it is necessary to set appropriate structures and parameters for the gating network to ensure that it can accurately select suitable expert network models for different inputs.
[0038] Next, the training data is constructed, which includes the following: historical structured semantic description (such as wind speed, air pressure, network information, electromagnetic field information, etc.), such as: { "sensor_data": { "temperature": 25.6, "humidity": 65.3, "wind_speed": 5.2, "wind_direction": 180, "pressure": 1013.2, "electromagnetic_field": 0.5 }, "timestamp":"2024-01-20 14:30:00" }; Scene semantic description, such as: { "scene_description": { "obstacles":"There is a 30-meter-high transmission tower 200 meters ahead, and a forest 100 meters to the left, with the highest tree about 15 meters high", "environment":"Current visibility is good, with no obvious meteorological obstacles", "risk_factors":"There is strong electromagnetic interference around the transmission lines" } }; Historical external meteorological spatial data, such as: { "weather_forecast": { "short_term":"There will be thunderstorms in the next 2 hours", "warning":"Severe convective weather warning" }, "airspace_restrictions": { "no_fly_zones": ["Zone A","Zone B"], "height_limits":"Maximum flight altitude 300 meters", "time_restrictions":"14:00-16:00 flight restrictions" } }; There are also historical time series measurement data, which are not listed here one by one. Multi-task annotation data needs to cover the spatial state, such as: { "flight_recommendations": { "route_adjustment":"It is recommended to shift 200 meters east to avoid strong electromagnetic interference areas", "height_adjustment":"It is recommended to lower the flight altitude to 150 meters", "timing_suggestion":"It is recommended to complete the task in advance and avoid thunderstorms", "risk_level":"Medium risk", "priority_actions": [ "Adjust flight path to avoid transmission towers", "Strengthen real-time meteorological monitoring" ] } }; Environmental risk warning and Regional weather forecasts, such as: { "risk_assessment": { "electromagnetic_forecast": { "intensity_map":"Electromagnetic field intensity prediction within 50 meters around the transmission line", "interference_zones": ["A zone","B zone"], "safe_corridors":"Recommended flight corridor coordinates" }, "obstacle_analysis": { "dynamic_obstacles":"Mobile crane operation warning", "temporary_restrictions":"Temporary construction areas", "bird_activity":"Bird migration activity forecast" }, "signal_coverage": { "weak_signal_areas":"Communication signal coverage blind area prediction", "backup_channels":"Backup communication scheme" } } }; Flight decision advice, such as: { "flight_recommendations": { "route_adjustment":"It is recommended to shift 200 meters east to avoid strong electromagnetic interference areas", "height_adjustment":"It is recommended to lower the flight altitude to 150 meters", "timing_suggestion":"It is recommended to complete the task in advance and avoid thunderstorms", "risk_level":"Medium risk", "priority_actions": [ "Adjust flight path to avoid transmission towers", "Strengthen real-time meteorological monitoring" ] } }.
[0039] The collected data is labeled with the above labels to provide the supervision information required for model training, and the data labeled with many task labels is preprocessed by cleaning, normalization, alignment and enhancement (such as grayscale transformation, spatial filtering, etc.) to improve data quality and model training efficiency, forming an input data set to pass to the gating network, so that the gating network outputs a probability distribution according to the characteristics of the input data, indicating the importance of each expert network model in processing the current input.
[0040] According to the output of the gating network, the corresponding expert network model is selected to process the input data. Each expert processes the input data in his or her field of expertise and generates corresponding output results. Finally, the output results of each expert are weighted and summed according to the probability assigned by the gating network to form the final output, that is, the multimodal fusion representation prediction result, which integrates the knowledge and information of multiple experts.
[0041] Use standard loss functions (such as cross entropy loss, mean square error, etc.) to evaluate the gap loss between the model prediction and the true label. The gradient of the loss relative to the model parameters is calculated through the back propagation algorithm, and the parameters of the expert network and the gating network are updated. The training process is iterated until the loss output by the digital spatial domain expert model is less than the preset loss threshold or the preset number of iterations is reached. In addition, during the training process, in order to reduce the computational cost, usually only a part of the experts will be activated. For example, in the top-k routing strategy, only the k experts with the highest scores will be activated. The use of this sparse activation mechanism helps to improve the training efficiency and generalization ability of the model. In order to avoid the problem of some experts being overloaded while other experts are idle, it is necessary to consider how to achieve load balancing during training. This means that the gating network should not only focus on accuracy, but also ensure that the workload of each expert is relatively balanced. Load balancing can be achieved by introducing additional load balancing loss (LBL). Since MoE contains multiple expert network models, it may face the risk of overfitting. Therefore, appropriate regularization techniques (such as Dropout, L2 regularization, etc.) need to be used during training to prevent overfitting.
[0042] The digital spatial expert model adopts an expert selection strategy based on attention routing, so that each expert can selectively process knowledge in different fields during the training stage. During reasoning, the most relevant expert output is selected according to the attention score and weighted fusion is performed to obtain a multimodal fusion representation. Among them, the digital spatial expert model can use Switch Transformer to replace the MoE architecture, optionally, Dynamic Routing Networks, and preferably, Hierarchical MoE is used to achieve multi-level expert division of labor. In the routing strategy part, Top-K routing can also be used instead of attention routing, and reinforcement learning can be used to dynamically select experts.
[0043] By fusing historical external meteorological spatial data, historical structured semantic descriptions and historical time series measurements, as well as corresponding multi-task labels, the model can more comprehensively understand the airspace status, thereby improving the accuracy of the prediction; by using multimodal fusion to characterize the prediction results, the complementarity between different modal data can be fully utilized, and the generalization ability of the model can be enhanced, so that it can better adapt to different scenarios and conditions; by annotating the training set data with corresponding multi-task labels, the model can learn multiple related tasks at the same time, thereby improving the overall performance and efficiency.
[0044] In one embodiment, the multimodal transformer decoder includes an enhanced self-attention layer, a hierarchical entity recognition layer, and a relationship extraction network layer; wherein, The obtained multimodal fusion enhanced representation is input into the multimodal transformer decoder for decoding to obtain a set of entity triples to construct a digital spatial domain knowledge graph, including: Inputting the multimodal fusion enhanced representation into the enhanced self-attention layer for processing to capture the long-range dependencies in the multimodal fusion enhanced representation; The long-range dependency relationship is subjected to entity recognition and extraction processing by the hierarchical entity recognition layer to obtain basic entities; the basic entities include meteorological elements, geographical features and equipment status; Extracting the relationship between the basic entities based on the relationship extraction network layer, and generating a standardized entity triple set and its corresponding confidence score; the relationship includes spatial position relationship, causal relationship and temporal relationship; The entity triple set is converted into a natural language description through a preset prompt template, and a query prompt is constructed in combination with domain knowledge; Inputting the preset prompt template and the query prompt into a large language model for knowledge reasoning, and performing quality control on each reasoning result using each confidence score as a weight; The reasoning results output by the large language model are integrated to form a unified knowledge system to construct a digital spatial knowledge graph.
[0045] Specifically, the multimodal transformer decoder in the present invention can also be considered as a hierarchical structure integrating multiple BERT improved networks; wherein the enhanced self-attention layer is a RoBERTa network layer, the hierarchical entity recognition layer includes a SpanBERT network layer and a BERT-CRF network layer, and the relationship extraction network layer includes an R-BERT network layer and a TDEER network layer. Then, the decoding process of the multimodal transformer decoder for the multimodal fusion enhanced representation is essentially the gradual decoding and refinement of the multimodal fusion enhanced representation using multiple BERT improved networks, including: using a multi-head self-attention mechanism through the RoBERTa network layer to calculate different The attention scores between positions are used to capture the long-range dependencies in the multimodal fusion enhanced representation, and then the abstract features are gradually transformed into specific semantic concepts through the multi-layer Transformer structure. That is, the SpanBERT network layer is used to identify the entity fragments in the long-range dependencies and perform entity boundary detection, and the BERT-CRF network layer is used to further refine, classify and extract the entities to obtain the basic entities; then the relationship between the basic entities is extracted based on the R-BERT network layer and the TDEER network layer, and the extracted relationship is normalized to generate a normalized set of entity triples and their corresponding confidence scores, which are used for quality control in the subsequent knowledge graph construction.
[0046] Design specific prompt templates to convert entity triples into natural language descriptions. These templates should be concise and clear, and be able to accurately convey the information in the entity triples; combine the natural language description with the pre-domain knowledge to construct query prompts, so as to guide the large language model to perform knowledge reasoning in the subsequent steps; input these prompt templates and the constructed query prompts into the large language model that has been fine-tuned for domain adaptability. Based on its powerful knowledge reasoning ability, the model supplements missing relations, eliminates semantic conflicts, and unifies conceptual expressions, that is, synonymous entities are unified through entity linking technology, and implicit relation edges are supplemented through relational reasoning. The contradictory relationships are eliminated through consistency checks. At the same time, the confidence scores of each entity triple are used as weights to control the quality of each reasoning result during the reasoning process. Then, the reasoning results of the large language model are integrated to form a unified knowledge system. During the integration process, it is necessary to deal with information conflicts and duplications to ensure the accuracy and consistency of knowledge. Finally, the integrated knowledge is stored in the knowledge base in an appropriate format. Common storage methods include graph databases (such as Neo4j) and RDF storage systems (such as Jena). At the same time, the stored knowledge is optimized to improve query efficiency and accuracy, thereby forming a digital spatial domain knowledge graph. In addition, in the knowledge graph relationship reasoning part, rule reasoning can be used instead of neural network reasoning, and knowledge distillation can be used to simplify the reasoning process. Further, probabilistic graph model reasoning can be realized.
[0047] The present invention uses a multimodal transformer decoder to more effectively perform entity recognition and relationship extraction, thereby improving the accuracy and completeness of the digital spatial domain knowledge graph; wherein, the enhanced self-attention layer can capture the long-range dependencies in the representation, which helps to more accurately understand the relationship between entities when constructing the knowledge graph; and the introduction of multimodal fusion enhanced representation enables the solution to make full use of various types of data and improve the richness and practicality of the knowledge graph; through the preset prompt template, the entity triple set is converted into a natural language description, and the query prompt is constructed in combination with the domain knowledge, and then input into the large language model for knowledge reasoning, which can improve the efficiency and accuracy of reasoning, and at the same time, the confidence score is used as the weight to control the quality of the reasoning result, further ensuring the reliability of the knowledge graph.
[0048] In one embodiment, the graph neural network includes a gated attention layer, a temporal attention layer, and a residual connection layer; wherein the attention enhancement mechanism based on the graph neural network is used to process the digital spatial domain knowledge graph to obtain an attention-enhanced node representation, including: Standardizing the edges and nodes in the digital spatial knowledge graph, and constructing a multi-scale neighborhood feature aggregator for each node to aggregate the structural feature representations of each node from local to global at multiple scales; Calculating the task attention score of each of the structural feature representations through the gated attention layer, and performing weighted summation on each of the structural feature representations based on each of the task attention scores to obtain a fused scale feature representation; Calculating the temporal attention score of each of the structural feature representations based on the temporal attention layer, and performing weighted summation on each of the structural feature representations according to each of the temporal attention scores to obtain a fused temporal feature representation; The fused scale feature representation and the fused time feature representation are processed by the residual connection layer to realize online updating of the node representation, and the processing steps of the gated attention layer, the temporal attention layer and the residual connection layer are iteratively executed to obtain the node representation after attention enhancement.
[0049] Specifically, based on the knowledge graph, the present invention uses a dynamic graph attention network enhancement mechanism to achieve adaptive representation optimization based on drone scenarios: The edges and nodes in the digital spatial knowledge graph are standardized to ensure the consistency and comparability of the data, and a multi-scale neighborhood feature aggregator is constructed for each node, namely 1. Neighbor sampling: for each node, a certain number of neighbor nodes are sampled according to its position in the graph; 2. Feature aggregation: the features of the sampled neighbor nodes are aggregated to form a local feature representation of the node. This operation can be achieved through summation, averaging, maximum pooling and other operations; 3. Multi-scale processing: the above-mentioned neighbor sampling and feature aggregation process is repeated at different scales to capture structural information of different ranges, and then the structural feature representation of each node from local to global at multiple scales is aggregated, and the multi-level information of the node is captured to improve the richness of the node representation.
[0050] Through the gated attention mechanism, the importance weights of features of different scales are dynamically adjusted according to the current task. That is, the task attention score of each structural feature representation is calculated through the gated attention layer, that is, for the feature representation of each scale, its attention score related to the current task is calculated. This operation can be achieved by performing a dot product or a bilinear transformation on the feature representation and the query vector related to the task, and then applying the softmax function to reflect the importance of different structural features for the current task. Based on the attention score of each task, the weighted sum of the corresponding structural feature representations is performed to obtain a fused scale feature representation, which helps to highlight important features and suppress minor features.
[0051] First, the structural feature representation of each node is converted into a time series, and then based on the temporal attention layer, for each time point in the time series, its attention score related to the current time point is calculated. This can be achieved by performing a dot product or bilinear transformation on the state representation of the current time point and the state representation of the past time point, and then applying the softmax function to reflect the influence of the structural features of different time points on the current task. The weighted summation of the corresponding structural feature representations is performed according to the attention scores of each time to obtain a fused temporal feature representation, which helps to capture the temporal dependency of graph data, model the temporal evolution of node states, and improve the adaptability of the model to dynamic changes.
[0052] The fused scale feature representation and the fused time feature representation are added together through the residual connection layer to retain the input information and prevent the gradient from disappearing or exploding. The processing result is then normalized to have a stable mean and variance to ensure the stability of feature transfer, so as to achieve online update of node representation and alleviate the gradient disappearance and over-smoothing problems in deep graph neural networks.
[0053] Finally, the processing of the above-mentioned gated attention layer, temporal attention layer and residual connection layer is iterated multiple times to gradually optimize the node representation until the predetermined number of iterations or convergence conditions are reached, and the node representation after attention enhancement is obtained, which not only contains the local and global structural information of the node, but also incorporates the time evolution law and task-related dynamic adjustment information, which can be used for subsequent drone scene analysis and decision-making tasks. In addition, in each iteration, the importance weights of features of different scales and the time evolution feature representation can be dynamically adjusted according to the requirements of the current task and the structural information of the graph, so as to gradually optimize the node representation and improve the accuracy and generalization ability of the model. Furthermore, for graph neural networks, the present invention can use GraphSAGE instead of GAT, optionally, use GIN for graph isomorphism representation learning, preferably, use spatiotemporal graph neural networks; the attention mechanism part can also use multi-head attention instead of gated attention, and sparse attention can be used to improve efficiency.
[0054] The present invention uses the attention mechanism to enable the solution to focus on important nodes and edges in the graph, thereby generating more representative node representations; the multi-scale neighborhood feature aggregator is used to capture the structural characteristics of nodes at different scales, which helps to more comprehensively understand the role and position of nodes in the graph; the temporal attention layer is used to capture the time dependency of the structural feature representation, which helps to process dynamically changing graph data; the residual connection layer is used to alleviate the gradient vanishing and over-smoothing problems in deep graph neural networks.
[0055] In one embodiment, the conditional generative adversarial network model includes a generator having an encoding layer, a decoding layer, and a multi-task branch layer, and a Transformer-based multi-task discriminator; wherein the node representation with enhanced attention is input into the conditional generative adversarial network model for task prediction, and an inspection decision suggestion for the area to be inspected is generated, including: The encoding layer is used to encode the node representation to obtain an encoding result, and the encoding result and random noise are mapped to a target output space through the decoding layer to obtain a decoding output result; The decoding output result is processed by the multi-task branch layer to obtain the airspace state prediction result, regional weather prediction result, environmental risk prediction result and inspection decision suggestion prediction result of the area to be inspected; The airspace status prediction results, the regional weather prediction results, the environmental risk prediction results and the inspection decision recommendation prediction results are evaluated according to the multi-task discriminator, and the inspection decision recommendation prediction results are adjusted based on the evaluation results to obtain inspection decision recommendations for the area to be inspected.
[0056] Specifically, the present invention uses a multi-task prediction generation framework based on a conditional generative adversarial network to generate a high-quality solution; wherein the generator adopts a Transformer-based encoder-decoder structure, including an encoding layer, a decoding layer, and a multi-task branch layer, which encodes the attention-enhanced node representation through the encoding layer to extract the global dependency, and uses the decoding layer to map the encoding result and random noise to the target output space to obtain the decoded output result, and then uses the multi-task branch layer to process the decoded output result to obtain the airspace state prediction result of the inspection area (that is, the multidimensional state sequence in the future time window, including key indicators such as meteorological conditions, electromagnetic environment, and communication quality), the regional weather forecast result (combining multi-source data to generate refined local The system generates the following prediction results: partial weather forecast, including wind field distribution, precipitation probability, visibility change, etc.), environmental risk prediction results (environmental risk level of the inspection area) and inspection decision suggestion prediction results (that is, generating multi-level decision suggestions based on the current state and prediction results, including route planning, altitude adjustment, speed control and other specific instructions). Each task has an independent fully connected header to generate the prediction results of a specific task. Finally, the Transformer-based multi-task discriminator is used to evaluate the authenticity and task relevance of these prediction results, and output the authenticity score and auxiliary classification results. If the authenticity score is lower than the evaluation score threshold, the decision suggestion is optimized according to the auxiliary classification result. Otherwise, the inspection decision suggestion prediction result is output as the inspection decision suggestion for the inspection area. For example, the input includes: node representation: fusion features of region A (meteorological conditions, electromagnetic environment, communication quality), current conditions: the current period is the noon peak, and region B is temporarily banned from flying; output: airspace status: meteorological conditions of region A in the next hour (wind speed↑, precipitation probability↑), electromagnetic environment (interference↑), communication quality (↓); regional weather: wind field distribution of region A in the next hour (northwest wind 15m / s), precipitation probability (80%), visibility (<1km); environmental risk: the risk level of drone conflict in region A is high; inspection decision: it is recommended to adjust the route to avoid region A, reduce the flight altitude to 100m, and control the speed at 10m / s. In addition, the generation architecture can use Diffusion Models instead of GAN, optionally, use Flow-based Models, preferably, use autoregressive generation model; in the multi-task learning part, soft parameter sharing can be used instead of hard parameter sharing, and progressive learning strategy can be used, and further, dynamic task weight adjustment can be achieved. It should be noted that all models directly used in the present invention can be directly applied only after being trained with historical data, such as the conditional generative adversarial network model, the multimodal transformer decoder and its various layers.
[0057] The present invention adopts multi-task collaborative prediction. Through the multi-task branching layer, the model can simultaneously generate airspace status, regional weather, environmental risks and inspection decision suggestions, providing comprehensive airspace management support; adopting the Transformer-based encoder-decoder structure and multi-task discriminator, it can capture complex airspace dynamic characteristics and generate highly realistic and task-related prediction results; through the evaluation and feedback of the discriminator, the output of the generator is dynamically adjusted to improve the rationality and executability of the decision suggestions; random noise and adversarial training mechanisms are introduced to enhance the robustness of the model to noise and abnormal data, and generate diversified decision-making solutions.
[0058] In the embodiment of the present application, based on the problem of how to improve the utilization rate of multi-source perception data used by the digital airspace system for drone inspections to improve the inspection efficiency of drones, a patrol method based on the digital airspace system is designed, which combines multimodal learning with data distribution monitoring technology to build an end-to-end data processing framework, and realizes the full-process intelligent processing from the collection, enhancement, knowledge extraction, and decision generation of multi-source heterogeneous perception data: multimodal feature extraction and time series data calibration are performed through pre-trained neural network models to improve data quality; an expert model based on the MoE architecture is developed, and an attention routing strategy is adopted to realize the dynamic fusion of knowledge in different fields; Variational autoencoders are introduced for representation enhancement, and combined with a multimodal transformer decoder to achieve accurate semantic parsing; large language models are innovatively applied to knowledge graph construction, and the integrity of knowledge representation is improved through entity linking and relational reasoning; a dynamic graph attention network is designed to achieve scene-adaptive representation optimization; finally, high-quality multi-task prediction and decision generation are achieved based on conditional generative adversarial networks; data quality and decision accuracy are significantly improved, and end-to-end processing from raw perception data to intelligent decision-making is achieved, providing reliable digital airspace support for drone power inspections, significantly improving the intelligence level of drone power inspections, and has important engineering application value.
[0059] It should be noted that although the steps in the above flowchart are shown in sequence according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise specified in this document, there is no strict order restriction for the execution of these steps, and these steps can be executed in other orders.
[0060] In another embodiment, if Figure 2 As shown, the second aspect of the present invention provides an inspection device based on a digital airspace system, comprising: The data acquisition module 10 is used to acquire multi-source sensing monitoring data and external meteorological space data of the area to be inspected; A data processing module 20 is used to process the multi-source sensing monitoring data using a plurality of pre-trained neural network models according to the type of the multi-source sensing monitoring data to obtain a structured semantic description and a time series measurement value; A feature fusion module 30 is used to input the external meteorological spatial data, the structured semantic description and the time series measurement value into a digital airspace expert model based on the MoE architecture for fusion processing to obtain a multimodal fusion representation; A graph construction module 40 is used to enhance the multimodal fusion representation through a variational autoencoder, and input the obtained multimodal fusion enhanced representation into a multimodal transformer decoder for decoding to obtain a set of entity triples to construct a digital spatial domain knowledge graph; The decision generation module 50 is used to process the digital airspace knowledge graph using an attention enhancement mechanism based on a graph neural network, obtain attention-enhanced node representations to input into a conditional generative adversarial network model for task prediction, and generate inspection decision recommendations for the area to be inspected so that the drone can execute them.
[0061] It should be noted that each module in the above-mentioned inspection device based on a digital airspace system can be fully or partially implemented by software, hardware and a combination thereof. The above-mentioned modules can be embedded in or independent of the processor in the computer device in the form of hardware, or can be stored in the memory of the computer device in the form of software, so that the processor can call and execute the operations corresponding to the above modules. For the specific definition of an inspection device based on a digital airspace system, please refer to the definition of an inspection method based on a digital airspace system above. The two have the same functions and effects and will not be repeated here.
[0062] A third aspect of the present invention provides an electronic device, the electronic device comprising: processor, memory, and bus; The bus is used to connect the processor and the memory; The memory is used to store operation instructions; The processor is used to call the operation instruction, and the executable instruction enables the processor to perform operations corresponding to the inspection method based on the digital airspace system as shown in the first aspect of the present application.
[0063] In an alternative embodiment, an electronic device is provided, such as Figure 3 As shown, Figure 3The electronic device 5000 shown includes: a processor 5001 and a memory 5003. The processor 5001 and the memory 5003 are connected, such as through a bus 5002. Optionally, the electronic device 5000 may also include a transceiver 5004. It should be noted that in actual applications, the transceiver 5004 is not limited to one, and the structure of the electronic device 5000 does not constitute a limitation on the embodiments of the present application.
[0064] Processor 5001 may be a CPU, a general-purpose processor, a DSP, an ASIC, an FPGA or other programmable logic device, a transistor logic device, a hardware component or any combination thereof. It may implement or execute various exemplary logic blocks, modules and circuits described in conjunction with the disclosure of this application. Processor 5001 may also be a combination that implements computing functions, such as a combination of one or more microprocessors, a combination of a DSP and a microprocessor, etc.
[0065] The bus 5002 may include a path to transmit information between the above components. The bus 5002 may be a PCI bus or an EISA bus, etc. The bus 5002 may be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 3 Only one thick line is used in the diagram, but this does not mean that there is only one bus or only one type of bus.
[0066] The memory 5003 may be a ROM or other type of static storage device that can store static information and instructions, a RAM or other type of dynamic storage device that can store information and instructions, or an EEPROM, a CD-ROM or other optical disk storage, an optical disk storage (including a compressed optical disk, a laser disk, an optical disk, a digital versatile disk, a Blu-ray disk, etc.), a magnetic disk storage medium or other magnetic storage device, or any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited to these.
[0067] The memory 5003 is used to store application code for executing the solution of the present application, and the execution is controlled by the processor 5001. The processor 5001 is used to execute the application code stored in the memory 5003 to implement the content shown in any of the above method embodiments.
[0068] Among them, electronic devices include but are not limited to: mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), vehicle-mounted terminals (such as vehicle-mounted navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc.
[0069] The fourth aspect of the present invention provides a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, an inspection method based on a digital airspace system shown in the first aspect of the present application is implemented.
[0070] Another embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored. When the computer-readable storage medium is run on a computer, the computer can execute the corresponding content in the aforementioned method embodiment.
[0071] In addition, an embodiment of the present invention further provides a computer-readable storage medium on which a computer program is stored, and when the program is executed by a processor, the steps of the above method are implemented.
[0072] In summary, the present invention relates to the field of information technology, and discloses an inspection method, device, equipment and medium based on a digital airspace system, which uses a pre-trained neural network model to extract and calibrate features of multi-source perception monitoring data of the inspection area, obtains structured semantic descriptions and time series measurement values, and combines them with external meteorological space data to input into a digital airspace expert model based on the MoE architecture for multimodal fusion; enhances the generated multimodal fusion representation through a variational autoencoder, inputs it into a multimodal transformer decoder for decoding, and obtains a set of entity triples to construct a digital airspace knowledge graph; uses a graph neural network-based attention enhancement mechanism to enhance the node representation in the knowledge graph, and inputs it into a conditional generative adversarial network model for multi-task prediction, generates inspection decision suggestions for drone execution, and effectively improves the utilization rate of multi-source perception data, thereby improving the inspection efficiency of drones.
[0073] Each embodiment in this specification is described in a progressive manner, and the same or similar parts of each embodiment can be directly referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the device embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment. It should be noted that the technical features of the above-mentioned embodiments can be combined arbitrarily. In order to make the description concise, all possible combinations of the technical features in the above-mentioned embodiments are not described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0074] The above-mentioned embodiments only express several preferred implementation modes of the present application, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of the invention patent. It should be pointed out that, for ordinary technicians in the technical field, several improvements and substitutions can be made without departing from the technical principles of the present invention, and these improvements and substitutions should also be regarded as the protection scope of the present application. Therefore, the protection scope of the patent of the present application shall be based on the protection scope of the claims.
Claims
1. A patrol inspection method based on a digital airspace system, characterized in that: include: Obtain multi-source sensing monitoring data and external meteorological space data of the area to be inspected; According to the type of the multi-source sensing monitoring data, a plurality of pre-trained neural network models are used to process the multi-source sensing monitoring data to obtain a structured semantic description and a time series measurement value; Inputting the external meteorological spatial data, the structured semantic description and the time series measurement value into a digital airspace expert model based on the MoE architecture for fusion processing to obtain a multimodal fusion representation; The multimodal fusion representation is enhanced by a variational autoencoder, and the obtained multimodal fusion enhanced representation is input into a multimodal transformer decoder for decoding to obtain a set of entity triples to construct a digital spatial domain knowledge graph; The digital airspace knowledge graph is processed using an attention enhancement mechanism based on a graph neural network to obtain attention-enhanced node representations that are input into a conditional generative adversarial network model for task prediction, generating inspection decision recommendations for the area to be inspected so that the drone can execute them.
2. The inspection method based on a digital airspace system according to claim 1, characterized in that: The multi-source sensing monitoring data includes unstructured multi-source sensing monitoring data, and the unstructured multi-source sensing monitoring data includes unstructured point cloud data and unstructured image data; wherein, According to the type of the multi-source sensing monitoring data, a plurality of pre-trained neural network models are used to process the multi-source sensing monitoring data to obtain a structured semantic description and a time series measurement value, including: Extracting features from the unstructured point cloud according to a pre-trained 3D convolutional neural network model and a spatial attention mechanism to obtain geometric features; Extracting features from the unstructured image data based on a pre-trained residual neural network model to obtain visual features; Using a cross-modal encoder to align and fuse the geometric features and the visual features to generate a fused feature; The fused features are converted into the structured semantic description through a pre-trained multimodal transformer decoder.
3. The inspection method based on a digital airspace system according to claim 1, characterized in that: The multi-source sensing monitoring data also includes structured multi-source sensing time series monitoring data; wherein, The method further comprises: processing the multi-source sensing monitoring data using a plurality of pre-trained neural network models according to the type of the multi-source sensing monitoring data to obtain a structured semantic description and a time series measurement value; Using a sliding time window to convert the structured multi-source sensing time series monitoring data into a time series structure sequence; Inputting the temporal structure sequence into a pre-trained deep learning network model to calculate the similarity between the temporal structure sequence and the historical temporal structure sequence through an attention mechanism to obtain a similarity distribution; Generate a data deviation correction matrix based on the error characteristics of the sensor corresponding to the structured multi-source sensing time series monitoring data and the similarity distribution; The data deviation correction matrix is applied to the structured multi-source sensing time series monitoring data to obtain corrected time series measurement values.
4. The inspection method based on a digital airspace system according to claim 1, characterized in that: The digital airspace expert model includes several expert network models and a gated network; wherein, The training process of the digital airspace expert model includes: The historical external meteorological spatial data, the historical structured semantic description and the historical time series measurement values are used as training set data, and the corresponding multi-task labels are annotated for the training set data; the multi-task labels include airspace status, environmental risk level, regional weather classification and inspection decision suggestions; Align and enhance the training set data with multiple task labels to form an input data set to be passed to the gating network to obtain the output probability; Selecting a corresponding expert network model according to the output probability to process the input data set, obtaining a corresponding output result, performing weighted summation based on the output probability assigned by the gating network, and forming a multimodal fusion representation prediction result; A loss function is used to evaluate the loss between the multimodal fusion representation prediction result and the multi-task label, and the gradient of the loss relative to the model parameters is calculated through a back-propagation algorithm to update the parameters of each of the expert network models and the gating network.
5. The inspection method based on a digital airspace system according to claim 1, characterized in that: The multimodal transformer decoder includes an enhanced self-attention layer, a hierarchical entity recognition layer, and a relationship extraction network layer; wherein, The obtained multimodal fusion enhanced representation is input into the multimodal transformer decoder for decoding to obtain a set of entity triples to construct a digital spatial domain knowledge graph, including: Inputting the multimodal fusion enhanced representation into the enhanced self-attention layer for processing to capture the long-range dependencies in the multimodal fusion enhanced representation; The long-range dependency relationship is subjected to entity recognition and extraction processing by the hierarchical entity recognition layer to obtain basic entities; the basic entities include meteorological elements, geographical features and equipment status; Extracting the relationship between the basic entities based on the relationship extraction network layer, and generating a standardized entity triple set and its corresponding confidence score; the relationship includes spatial position relationship, causal relationship and temporal relationship; The entity triple set is converted into a natural language description through a preset prompt template, and a query prompt is constructed in combination with domain knowledge; Inputting the preset prompt template and the query prompt into a large language model for knowledge reasoning, and performing quality control on each reasoning result using each confidence score as a weight; The reasoning results output by the large language model are integrated to form a unified knowledge system to construct a digital spatial knowledge graph.
6. The inspection method based on a digital airspace system according to claim 1, characterized in that: The graph neural network includes a gated attention layer, a temporal attention layer and a residual connection layer; wherein, The attention enhancement mechanism based on the graph neural network is used to process the digital spatial domain knowledge graph to obtain the attention-enhanced node representation, including: Standardizing the edges and nodes in the digital spatial knowledge graph, and constructing a multi-scale neighborhood feature aggregator for each node to aggregate the structural feature representations of each node from local to global at multiple scales; Calculating the task attention score of each of the structural feature representations through the gated attention layer, and performing weighted summation on each of the structural feature representations based on each of the task attention scores to obtain a fused scale feature representation; Calculating the temporal attention score of each of the structural feature representations based on the temporal attention layer, and performing weighted summation on each of the structural feature representations according to each of the temporal attention scores to obtain a fused temporal feature representation; The fused scale feature representation and the fused time feature representation are processed by the residual connection layer to realize online updating of the node representation, and the processing steps of the gated attention layer, the temporal attention layer and the residual connection layer are iteratively executed to obtain the node representation after attention enhancement.
7. The inspection method based on a digital airspace system according to claim 1, characterized in that: The conditional generative adversarial network model includes a generator having an encoding layer, a decoding layer and a multi-task branch layer, and a Transformer-based multi-task discriminator; wherein, The node representation with enhanced attention is input into the conditional generative adversarial network model for task prediction, and an inspection decision suggestion for the area to be inspected is generated, including: The encoding layer is used to encode the node representation to obtain an encoding result, and the encoding result and random noise are mapped to a target output space through the decoding layer to obtain a decoding output result; The decoding output result is processed by the multi-task branch layer to obtain the airspace state prediction result, regional weather prediction result, environmental risk prediction result and inspection decision suggestion prediction result of the area to be inspected; The airspace status prediction results, the regional weather prediction results, the environmental risk prediction results and the inspection decision recommendation prediction results are evaluated according to the multi-task discriminator, and the inspection decision recommendation prediction results are adjusted based on the evaluation results to obtain inspection decision recommendations for the area to be inspected.
8. A patrol device based on a digital airspace system, characterized in that: include: A data acquisition module is used to acquire multi-source sensing monitoring data and external meteorological space data of the area to be inspected; A data processing module, used to process the multi-source sensing monitoring data using a plurality of pre-trained neural network models according to the type of the multi-source sensing monitoring data to obtain a structured semantic description and a time series measurement value; A feature fusion module, used for inputting the external meteorological spatial data, the structured semantic description and the time series measurement value into the digital airspace expert model based on the MoE architecture for fusion processing to obtain a multimodal fusion representation; A graph construction module is used to enhance the multimodal fusion representation through a variational autoencoder, and input the obtained multimodal fusion enhanced representation into a multimodal transformer decoder for decoding to obtain a set of entity triples to construct a digital spatial domain knowledge graph; The decision generation module is used to process the digital airspace knowledge graph using an attention enhancement mechanism based on a graph neural network, obtain attention-enhanced node representations to input into a conditional generative adversarial network model for task prediction, and generate inspection decision recommendations for the area to be inspected so that the drone can execute them.
9. An electronic device, characterized in that: It includes a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor, and when the processor executes the computer program, it implements the inspection method based on the digital airspace system as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium includes a stored computer program, wherein when the device where the computer-readable storage medium is located executes the computer program, the inspection method based on the digital airspace system as described in any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Rainfall prediction method and system based on artificial intelligence algorithm and knowledge graph
CN114385611A
Knowledge graph driven semantic governance scheme dynamic generation method and system
CN118114758A
Fishery monitoring method and device, terminal equipment and storage medium
CN118607719A
Hybrid expert target detection system and method
CN118675030A
Weather prediction system based on time sequence knowledge graph reasoning and static multivariate space-time attention fusion
CN119471859A
Cited By
Intelligent data analysis method and system based on industry large model
CN120086266A
Intelligent Data Analysis Method and System Based on Industry Large Model
CN120086266B
Surgical robot quality control and fault prediction method based on artificial intelligence
CN120183644A
Road area environment safety evaluation method and system in highway engineering construction period
CN120355239A
Logistics trajectory optimization method and device for industrial chain and electronic equipment
CN120447590A