Digital airspace multi-source sensing data processing method, system, equipment and medium

The multi-modal representation learning method is used to process multi-source perceived data in drone power inspection and build a knowledge graph, which solves the problem of low utilization rate of multi-source perceived data and achieves high efficiency and accuracy of data processing.

CN120219986AActive Publication Date: 2025-06-27STATE GRID ECONOMIC TECH RES INST CO LTD +1
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510208922.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-25
Publication Date
2025-06-27
Estimated Expiration
2045-02-25

AI Technical Summary

Technical Problem

The utilization rate of multi-source perceived data used by digital airspace systems in drone power inspection is low, resulting in limited data efficiency and accuracy.

Method used

By obtaining multi-source perceived real-time data and historical data during the power inspection of drone, the multi-modal representation learning method is used to process the real-time data, generate the first knowledge graph, and filter it based on the differences between historical data and real-time data, build the second knowledge graph, and finally input the two into the machine learning model for processing.

Benefits of technology

It significantly improves the utilization rate of multi-source perceived data in drone power inspection, enhances the accuracy and efficiency of data processing, and ensures the comprehensiveness and accuracy of data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120219986A_ABST
    Figure CN120219986A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of information processing, and discloses a digital airspace multi-source sensing data processing method, system and device and a medium, and the method comprises the steps: obtaining multi-source sensing real-time data and multi-source sensing historical data collected by a digital airspace system in an unmanned aerial vehicle power inspection process; processing the multi-source sensing real-time data through a multi-modal representation learning method to obtain label data, and combining the label data with the multi-source sensing historical data to generate a first knowledge graph; performing difference calculation on the multi-source sensing real-time data based on the multi-source sensing historical data, and screening the multi-source sensing real-time data according to a calculation result, so as to construct a second knowledge graph based on the screened multi-source sensing real-time data; inputting the first knowledge graph and the second knowledge graph into a machine learning model for processing to obtain multi-source sensing processing data; the comprehensive, accurate and efficient utilization of the multi-source sensing data is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of information processing, and particularly to a method, system, device and medium for processing multi-source perception data in a digital airspace. Background Art

[0002] At present, the digital airspace system can provide necessary flight environment information for unmanned aerial vehicles (UAVs), and has gradually become an indispensable part of UAV power line inspection. The data updated in real time by the digital airspace system to ensure the safe operation of automated UAV inspection comes from numerous heterogeneous sensors. However, these perception data often have errors and conflicts when used, resulting in very low utilization rates of these data.

[0003] It can be seen that how to solve the problem of low utilization rate of multi-source perception data used in UAV power line inspection by the digital airspace system has become a technical problem that needs to be urgently solved by those skilled in the art. Summary of the Invention

[0004] The present invention provides a method, system, device and medium for processing multi-source perception data in a digital airspace, so as to solve the problem of how to improve the utilization rate of multi-source perception data collected during UAV power line inspection.

[0005] To solve the above technical problem, a first aspect of the present invention provides a method for processing multi-source perception data in a digital airspace, including:

[0006] Obtaining multi-source perception real-time data and multi-source perception historical data during UAV power line inspection collected through the digital airspace system;

[0007] Processing the multi-source perception real-time data by using a multi-modal representation learning method to obtain label data, and combining it with the multi-source perception historical data to generate a first knowledge graph;

[0008] Based on the multi-source perception historical data, calculating the difference of the multi-source perception real-time data, and screening the multi-source perception real-time data according to the calculation result, so as to construct a second knowledge graph based on the screened multi-source perception real-time data;

[0009] Inputting the first knowledge graph and the second knowledge graph into a machine learning model for processing to obtain multi-source perception processed data.

[0010] As a preferred solution, the processing the multi-source perception real-time data by using a multi-modal representation learning method to obtain label data, and combining it with the multi-source perception historical data to generate a first knowledge graph includes:

[0011] Extract features from the multi-source perception real-time data to obtain multi-modal feature representations, and input them into a deep learning model for processing to obtain labeled data; among them, the deep learning model is trained using the multi-source perception historical data with labeled data.

[0012] Fuse and align the labeled data and the multi-source perception historical data to obtain labeled fusion data, and build an ontology library based on the labeled fusion data.

[0013] Train an entity recognition model to identify the entities corresponding to the ontology library from the labeled fusion data, and train a relation extraction model to extract the relationships between the entities from the labeled fusion data, and build and optimize a knowledge graph based on the entities and the relationships to obtain the first knowledge graph.

[0014] As one of the preferred solutions, the first knowledge graph includes a climate knowledge sub-graph and a network signal strength knowledge sub-graph; among them,

[0015] Building and optimizing the knowledge graph based on the entities and the relationships to obtain the first knowledge graph includes:

[0016] Extract climate entities from the entities, use the weather type in the climate entities as the climate central entity, and connect relevant climate attributes through the climate central entity to obtain the climate knowledge sub-graph.

[0017] Extract network signal entities from the entities, use the network signal strength in the network signal entities as the network signal central entity, and connect relevant network signal attributes through the network signal central entity to obtain the network signal strength knowledge sub-graph.

[0018] Optimize the climate knowledge sub-graph and the network signal strength knowledge sub-graph to obtain the first knowledge graph.

[0019] As one of the preferred solutions, based on the multi-source perception historical data, calculate the difference of the multi-source perception real-time data, and screen the multi-source perception real-time data according to the calculation result, including:

[0020] Convert the multi-source perception historical data and the multi-source perception real-time data into historical data blocks and real-time data blocks respectively through a sliding window.

[0021] Calculate the data distributions between the historical data blocks and the real-time data blocks respectively according to the KL divergence and make a comparison, so as to obtain the difference value between the multi-source perception historical data and the multi-source perception real-time data according to the comparison result.

[0022] Compare the difference value with a preset difference threshold, and filter the multi-source perception real-time data according to the comparison result.

[0023] As one of the preferred solutions, constructing the second knowledge graph based on the filtered multi-source perception real-time data includes:

[0024] Extract multi-modal data features from the filtered multi-source perception real-time data and perform feature fusion to obtain cross-modal feature data as sample data to train a cross-modal association model;

[0025] Process the cross-modal feature data through the trained cross-modal association model to construct and optimize a knowledge graph to obtain the second knowledge graph.

[0026] As one of the preferred solutions, inputting the first knowledge graph and the second knowledge graph into a machine learning model for processing to obtain multi-source perception processed data includes:

[0027] Perform entity matching on the first knowledge graph and the second knowledge graph through a neural network model to obtain an entity matching result, and perform conflict detection according to the entity matching result;

[0028] If there is no conflict in the entity matching result, output the successfully matched entity pairs as the multi-source perception processed data;

[0029] If there is a conflict in the entity matching result, calculate the similarity between the multi-source perception historical data and the first knowledge graph and the second knowledge graph respectively for comparison, and output the data in the knowledge graph with higher similarity as the multi-source perception processed data.

[0030] As one of the preferred solutions, performing entity matching on the first knowledge graph and the second knowledge graph through a neural network model to obtain an entity matching result includes:

[0031] Convert the entities and relationships in the first knowledge graph and the entities and relationships in the second knowledge graph into a first natural language description and a second natural language description respectively, and perform embedding transformation on the first natural language description and the second natural language description respectively to obtain a first high-dimensional semantic representation and a second high-dimensional semantic representation;

[0032] Input the entities of the first knowledge graph and the entities of the second knowledge graph into a graph neural network model for feature extraction to obtain a first entity structure feature and a second entity structure feature;

[0033] Perform fusion processing on the first natural language description and the first entity structure feature, and the second natural language description and the second entity structure feature respectively, to obtain a first fusion feature and a second fusion feature;

[0034] Based on the first fusion feature and the second fusion feature, extract a number of candidate entity pairs from the first knowledge graph and the second knowledge graph, quantify the similarity between the candidate entity pairs, and output the candidate entity pairs with a similarity exceeding a preset similarity threshold as successfully matched entity pairs.

[0035] The second aspect of the present invention provides a processing system for multi-source perception data in a digital airspace, including:

[0036] A multi-source data acquisition module for acquiring multi-source perception real-time data and multi-source perception historical data during the drone power inspection collected by the digital airspace system;

[0037] A first graph generation module for processing the multi-source perception real-time data by a multi-modal representation learning method to obtain label data, and combining it with the multi-source perception historical data to generate a first knowledge graph;

[0038] A second graph construction module for calculating the difference of the multi-source perception real-time data based on the multi-source perception historical data, screening the multi-source perception real-time data according to the calculation result, and constructing a second knowledge graph based on the screened multi-source perception real-time data;

[0039] A multi-source data processing module for inputting the first knowledge graph and the second knowledge graph into a machine learning model for processing to obtain multi-source perception processed data.

[0040] The third aspect of the present invention provides an electronic device, including a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor. When the processor executes the computer program, the processing method of the multi-source perception data in the digital airspace as described above is implemented.

[0041] The fourth aspect of the present invention provides a computer-readable storage medium, which includes a stored computer program. When the device where the computer-readable storage medium is located executes the computer program, the processing method of the multi-source perception data in the digital airspace as described above is implemented.

[0042] Compared with the prior art, the beneficial effects of the embodiments of the present invention are at least one of the following:

[0043] (1) By acquiring multi-dimensional data during the power inspection of drones, it provides a comprehensive and accurate information basis for subsequent data processing and knowledge graph construction; and uses multi-modal representation learning method to process these data to extract the correlation information and common features between different modal data, so as to effectively utilize the data information of multiple modalities and improve the accuracy and efficiency of data processing;

[0044] (2) Combining label data with multi-source perception historical data can generate a first knowledge graph covering multi-source perception data, providing strong support for subsequent data analysis; and screening real-time data based on the differences between historical data and real-time data to remove redundant and abnormal data, and retain real-time data with large differences from historical data, so as to construct a more accurate and targeted second knowledge graph;

[0045] (3) Through the data fusion processing of the knowledge graph by a machine learning model, it not only considers the relevance and complementarity between different modal data, but also utilizes the powerful computing power of the machine learning model to screen out data that does not conform to the current situation to reduce error data, thereby significantly improving the efficiency of drone power inspection. BRIEF DESCRIPTION OF THE DRAWINGS

[0046] In order to more clearly illustrate the technical solutions of the present invention, the drawings required for implementation will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0047] Figure 1 is a flowchart of a method for processing multi-source perception data of a digital airspace provided by an embodiment of the present invention;

[0048] Figure 2 is a schematic structural diagram of a climate knowledge sub-graph provided by an embodiment of the present invention;

[0049] Figure 3 is a structural diagram of a system for processing multi-source perception data of a digital airspace provided by an embodiment of the present invention;

[0050] Figure 4 is a structural diagram of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0051] The following describes the technical solutions in the embodiments of the present invention clearly and completely with reference to the accompanying drawings and embodiments. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. The purpose of providing these embodiments is to make the disclosure of the present invention more thorough and comprehensive. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the scope of protection of the present invention.

[0052] In the description of the present application, the terms "first", "second", "third", etc. are only used for descriptive purposes, and cannot be understood as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, features defined with "first", "second", "third", etc. may explicitly or implicitly include one or more of such features. In the description of the present application, unless otherwise stated, the meaning of "a plurality" is two or more.

[0053] In the description of the present application, it should be noted that unless otherwise clearly defined and limited, the terms "installed", "connected", "connected" should be understood in a broad sense. For example, it may be a fixed connection, a detachable connection, or an integral connection; it may be a mechanical connection or an electrical connection; it may be directly connected, or indirectly connected through an intermediate medium, and it may be the communication inside two components. The terms "vertical", "horizontal", "left", "right", "up", "down" and similar expressions used herein are only for the purpose of illustration, rather than indicating or implying that the system or component referred to must have a specific orientation, be constructed and operated in a specific orientation, and thus cannot be understood as a limitation to the present invention. The term "and / or" used herein includes any and all combinations of one or more of the related listed items. For those of ordinary skill in the art, the specific meanings of the above terms in the present application can be understood according to specific circumstances.

[0054] In the description of the present application, it should be noted that unless otherwise defined, all technical and scientific terms used in the present invention have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present invention belongs. The terms used in the specification of the present invention are only for the purpose of describing specific embodiments, and are not intended to limit the present invention. For those of ordinary skill in the art, the specific meanings of the above terms in the present application can be understood according to specific circumstances.

[0055] In one embodiment, as Figure 1 shown, the first aspect of the present invention provides a method for processing multi-source perception data in a digital airspace, including:

[0056] S1. Obtain the multi-source perception real-time data and multi-source perception historical data during the UAV power inspection collected by the digital airspace system;

[0057] Specifically, the present invention collects and transmits various data during the power inspection of drones through various sensors, communication devices, input devices, etc. contained in the digital airspace system, such as electromagnetic radiation meters, wireless network cards, climate observatories, anemometers set near the ground, and climate data (such as temperature, humidity, wind speed, etc.) and network signal strength data (such as interference sources, interference intensity, signal strength, etc.) collected by the wireless network card, lidar, and camera carried by the drone. At the same time, historical climate data and historical network signal strength data, etc. are obtained as references and comparisons. Among them, the data collected by the electromagnetic radiation meter, wireless network card, climate observatory, anemometer, and lidar are structured data, while the data collected by the camera are picture or video data (i.e., unstructured data), and the wireless network card is used to transmit data to achieve information transmission between the air and the ground. And the multi-source perception real-time data and the multi-source perception historical data are essentially of the same data type, both being the above-mentioned data, only the time points to which they belong are different. The time point and specific duration of the multi-source perception historical data are not limited and can be determined according to requirements. In addition, after the data is obtained, it is usually cleaned, and preprocessing processes such as removing noise, outliers, and duplicate data, normalization, and denoising are performed to improve the data quality. Since the multi-source perception data described in the present invention has data with different types of structures, it is also necessary to unify and convert its format for subsequent processing.

[0058] S2. Process the multi-source perception real-time data through the multi-modal representation learning method to obtain label data, and combine it with the multi-source perception historical data to generate a first knowledge graph;

[0059] In one embodiment, the process of processing the multi-source perception real-time data through the multi-modal representation learning method to obtain label data and combining it with the multi-source perception historical data to generate a first knowledge graph includes:

[0060] Extract features from the multi-source perception real-time data to obtain a multi-modal feature representation for input into a deep learning model for processing to obtain label data; among them, the deep learning model is trained with the multi-source perception historical data with labeled label data.

[0061] Fuse and align the label data and the multi-source perception historical data to obtain label fusion data, and build an ontology library based on the label fusion data;

[0062] Train an entity recognition model to identify the entities corresponding to the ontology library from the label fusion data, and train a relationship extraction model to extract the relationships between the entities from the label fusion data, and construct and optimize a knowledge graph according to the entities and the relationships to obtain the first knowledge graph.

[0063] Specifically, the present invention utilizes multimodal representation learning technology to convert data of different modalities such as text, images, and audio into vector forms that can be understood and processed by computers, captures the potential relationships between data from multiple modalities, and establishes a connection between image content and labels, namely, multimodal feature representation. As for the feature extraction process in multimodal representation learning, the corresponding feature extraction method is used for each sensor data. For example, for image data, image feature extraction methods such as HOG and SIFT can be used; for text data (such as data from climate observation instruments), text feature extraction methods such as TF-IDF and Word2Vec can be used; for audio data, audio feature extraction methods such as MFCC can be used, that is, appropriate extraction algorithms are adopted according to the corresponding data types to extract single-modal features, thereby reflecting the essential attributes and mutual relationships of the data; and for the fusion of single-modal features, feature fusion techniques such as weighted average and product kernel can be used to fuse the features of different modalities into a whole to form a multimodal feature representation. For example, for climate data, the joint representation of image modality (such as the full map of the real scene) and text modality (such as weather report) can be used; for network signal strength data, the joint representation of audio modality (such as audio features of signal interference) and text modality (such as text description of signal strength) can be used.

[0064] Then, a suitable multimodal representation learning model is selected, such as a deep learning model (such as a convolutional neural network CNN, a recurrent neural network RNN, a Transformer, etc.), and the multimodal feature representation corresponding to the multi-source perception historical data with labeled data is used as input for training, so that the model can learn the association and representation between multimodal data, and the multimodal feature representation of the multi-source perception real-time data obtained after processing is input into the trained deep learning model for processing, and the multi-source perception real-time data carrying the label data is output; wherein the label data can be represented as the category, attribute or other information of the multi-source perception real-time data, that is, the label data can be a classification label (such as climate type, signal strength level) or a regression label (such as a specific temperature value, signal strength value).

[0065] Then, the multi-source perception real-time data carrying labeled data and the multi-source perception historical data are fused through simple splicing, weighted averaging, feature selection, etc., and then aligned using rule-based matching, similarity-based matching, and machine learning-based matching methods to obtain labeled fused data to ensure that data from different sources remain consistent in format, content, etc.; then, an ontology library is constructed based on the fused labeled fusion data. The ontology library is the core component of the knowledge graph, which defines the entities, attributes, and relationships between entities in the knowledge graph.

[0066] By training an entity recognition model and a relation extraction model to identify the entities defined in the ontology library from the label fusion data and extract the relations between these entities, the training of the models is all carried out using the labeled training set, which will not be elaborated here. The entity recognition model can adopt rule-based methods, machine learning-based methods (such as Hidden Markov Model HMM, Conditional Random Field CRF, etc.) and deep learning-based methods (such as Recurrent Neural Network RNN, Long Short-Term Memory Network LSTM, Transformer, etc.) to learn the features and rules of entities in the label fusion data; the relation extraction model can adopt machine learning-based methods: using statistical models (such as Support Vector Machine SVM, Conditional Random Field CRF, etc.) or deep learning models (such as Recurrent Neural Network RNN, Convolutional Neural Network CNN, Transformer, etc.) to learn the features of entities and relations from the labeled data, or using distant supervision to automatically generate labeled data using the structured information in the knowledge base to train the relation extraction model, and then learn the features and rules between entities and relations in the label fusion data, such as the relations between climate entities (the correlation between temperature and humidity), the relations between network signal entities (the relation between interference sources and signal strength), etc.

[0067] Finally, based on the entities identified from the label fusion data and the relations extracted therefrom, a knowledge graph can be constructed and optimized, including removing redundant information, correcting incorrect relations, etc., to obtain the first knowledge graph. The present invention realizes efficient, accurate and real-time data processing and label generation by making full use of multi-source perception data, accurately extracting the single-modal features of the data and fusing them to form a multi-modal feature representation, and using a deep learning model to process and classify the multi-modal feature representation, which helps to improve the accuracy of subsequent processing and analysis; at the same time, by fusing real-time label data with historical data, the complementarity of data from different sources can be fully utilized, improving the comprehensiveness and accuracy of the data. By training an entity recognition model and a relation extraction model to automatically identify and extract the entities corresponding to the ontology library and their relations from the fusion data, the efficiency and automation degree of knowledge graph construction are greatly improved.

[0068] In one embodiment, the first knowledge graph includes a climate knowledge sub-graph and a network signal strength knowledge sub-graph; wherein,

[0069] Constructing and optimizing the knowledge graph according to the entities and the relations to obtain the first knowledge graph includes:

[0070] Extracting climate entities from the entities, taking the weather type in the climate entities as the climate central entity, and connecting relevant climate attributes through the climate central entity to obtain the climate knowledge sub-graph;

[0071] Extract network signal entities from the entities, use the network signal strength in the network signal entities as the network signal center entity, and connect relevant network signal attributes through the network signal center entity to obtain the network signal strength knowledge sub-graph;

[0072] Optimize the climate knowledge sub-graph and the network signal strength knowledge sub-graph to obtain the first knowledge graph.

[0073] Specifically, the present invention uses natural language processing (NLP) technology or rule-based methods to extract climate-related entities from entities, such as climate types, weather phenomena, geographical locations, etc., and selects the weather type that can summarize and represent the main climate characteristics of a region as the climate center entity among the climate entities, and connects climate attributes related to the climate center entity, such as the ranges of temperature, humidity, wind speed, and air pressure, to form a climate knowledge sub-graph. Taking light rain as an example, the structure of this climate knowledge sub-graph is as Figure 2 shown, which associates the ranges of four attributes: temperature, humidity, wind speed, and air pressure.

[0074] Similarly, the construction of the network signal strength knowledge sub-graph is achieved by using natural language processing (NLP) technology or rule-based methods to extract network signal-related entities from entities, such as network signal strength, base station location, network type, etc., and selecting the network signal strength that can directly reflect the network signal quality as the network signal center entity among the network signal entities, and connecting network signal attributes related to the network signal center entity, such as interference sources, interference intensities, network signal strength ranges, etc., thereby forming a network signal strength knowledge sub-graph.

[0075] Finally, optimize the climate knowledge sub-graph and the network signal strength knowledge sub-graph through steps including entity recognition, relationship extraction, redundant information removal, etc., to obtain the first knowledge graph. By constructing a knowledge graph, the present invention represents climate and network signal information in a structured manner, facilitating machine understanding and processing; integrating climate knowledge and network signal knowledge in one knowledge graph realizes cross-domain information sharing and correlation analysis; the knowledge graph supports efficient querying and reasoning, and can quickly obtain information and relationships related to climate and network signals.

[0076] S3. Based on the multi-source perception historical data, perform a difference calculation on the multi-source perception real-time data, and screen the multi-source perception real-time data according to the calculation result, so as to construct a second knowledge graph based on the screened multi-source perception real-time data;

[0077] In one embodiment, the performing a difference calculation on the multi-source perception real-time data based on the multi-source perception historical data and screening the multi-source perception real-time data according to the calculation result includes:

[0078] Convert the multi-source perception historical data and the multi-source perception real-time data into historical data blocks and real-time data blocks respectively through a sliding window;

[0079] Calculate the data distributions between the historical data blocks and the real-time data blocks respectively according to the KL divergence and make a comparison, so as to obtain the difference value between the multi-source perception historical data and the multi-source perception real-time data according to the comparison result;

[0080] Compare the difference value with a preset difference threshold, and screen the multi-source perception real-time data according to the comparison result.

[0081] Specifically, the present invention uses a sliding window with the same time length to convert the multi-source perception historical data and the multi-source perception real-time data to obtain historical data blocks and real-time data blocks. Each data block contains all data points within the window; for each historical data block and the corresponding real-time data block, calculate the KL divergence (Kullback-Leibler Divergence) between them, and then compare the data distribution differences between the historical data blocks and the real-time data blocks. The larger the KL divergence value, the greater the distribution difference between the two data blocks; set a reasonable difference threshold according to actual requirements and data characteristics to judge whether the difference between the real-time data block and the historical data block is significant, and compare the difference value of each real-time data block with the preset difference threshold. If the difference value is greater than the threshold, it is considered that there is a significant difference between the real-time data block corresponding to the difference value and the historical data block, which may be caused by anomalies, noise or data quality problems. Therefore, it is excluded, otherwise the real-time data is retained, and the screened real-time data is output for subsequent analysis, prediction or decision-making.

[0082] The present invention processes and analyzes data streams in real time through the sliding window technology to adapt to a dynamically changing environment; uses the KL divergence to measure the data distribution difference, can effectively identify and exclude abnormal or low-quality real-time data, and improve the reliability and accuracy of the data; the entire solution realizes an automated processing process, reduces manual intervention, and improves the processing efficiency and accuracy; can process perception data from multiple sources, and improves the comprehensiveness and accuracy of the data.

[0083] In one embodiment, constructing a second knowledge graph based on the screened multi-source perception real-time data includes:

[0084] Extract multi-modal data features from the screened multi-source perception real-time data and perform feature fusion to obtain cross-modal feature data as sample data to train a cross-modal association model;

[0085] Process the cross-modal feature data through the trained cross-modal association model to construct and optimize a knowledge graph to obtain the second knowledge graph.

[0086] Specifically, according to corresponding feature extraction algorithms, such as natural language processing technology (NLP), computer vision technology, image and audio processing technology, etc., the present invention extracts multi-modal data features from the filtered multi-source perception real-time data, and fuses the features of different modalities to form a unified feature vector, that is, cross-modal feature data, which is used as sample data to train a convolutional neural network to obtain a cross-modal association model; then, using the trained cross-modal association model, the cross-modal feature data is fused to construct a knowledge graph, and the corresponding entities, attributes and relationships are determined during the construction process and represented in a structured manner, and the constructed second knowledge graph is optimized through steps including entity recognition, relationship extraction, redundant information removal, etc. to improve the accuracy and integrity of the graph.

[0087] The present invention can process data of multiple modalities, including text, images, audio and video, etc., so as to make full use of the rich information in the multi-source perception real-time data; through feature fusion technology, the features of different modalities are fused to form a more comprehensive and accurate feature representation. At the same time, the cross-modal association model can associate data of different modalities and extract useful information from them; using the fused cross-modal feature data to construct a knowledge graph makes the information represented in a structured manner, which is convenient for subsequent query, analysis and reasoning.

[0088] S4. Input the first knowledge graph and the second knowledge graph into a machine learning model for processing to obtain multi-source perception processing data;

[0089] In one embodiment, step S4 includes:

[0090] Performing entity matching on the first knowledge graph and the second knowledge graph through a neural network model to obtain an entity matching result, and performing conflict detection according to the entity matching result;

[0091] In one embodiment, the performing entity matching on the first knowledge graph and the second knowledge graph through a neural network model to obtain an entity matching result includes:

[0092] Converting the entities and relationships in the first knowledge graph and the entities and relationships in the second knowledge graph into a first natural language description and a second natural language description respectively, and performing embedding transformation on the first natural language description and the second natural language description respectively to obtain a first high-dimensional semantic representation and a second high-dimensional semantic representation;

[0093] Inputting the entities of the first knowledge graph and the entities of the second knowledge graph into a graph neural network model for feature extraction to obtain a first entity structure feature and a second entity structure feature;

[0094] Perform fusion processing on the first natural language description and the first entity structure feature, and the second natural language description and the second entity structure feature respectively to obtain a first fusion feature and a second fusion feature;

[0095] Based on the first fusion feature and the second fusion feature, extract a number of candidate entity pairs from the first knowledge graph and the second knowledge graph to quantify the similarity between the candidate entity pairs, and output the candidate entity pairs with similarity exceeding a preset similarity threshold as the successfully matched entity pairs.

[0096] Specifically, in the present invention, the entities and relationships in the first knowledge graph and the second knowledge graph are respectively converted into natural language descriptions through natural language generation technology and input into a pre-trained embedding model (such as BERT, GPT, etc.) to convert the natural language text into a high-dimensional semantic representation; then, a graph convolutional network (GCN), a graph attention network (GAT), etc. are used to extract structural features according to the positions of the entities in each knowledge graph and their relationships with other entities; and methods such as splicing, weighted summation, and attention mechanism are used to fuse the high-dimensional semantic representation of the natural language description and the structural features extracted by the graph neural network to obtain the corresponding fusion features, which contain both the semantic information of the entities and their structural information in the graph; finally, based on the obtained fusion features, a number of candidate entity pairs are extracted from the first knowledge graph and the second knowledge graph by calculating metrics such as cosine similarity and Euclidean distance between the fusion features, the entity pairs with higher similarity are selected as candidates, and a similarity measurement method such as a similarity function based on learning is used to further quantify the similarity of the candidate entity pairs, so as to output the candidate entity pairs with similarity exceeding the preset similarity threshold as the successfully matched entity pairs. The present invention captures rich semantic information of entities and relationships through natural language descriptions and pre-trained embedding models, improving the accuracy of matching; uses graph neural network models to extract the structural features of entities in the graph to distinguish similar but different entities; and fuses semantic representations and structural features, which can make full use of the advantages of both and improve the robustness and accuracy of matching.

[0097] If there is no conflict in the entity matching result, output the successfully matched entity pairs as the multi-source perception processing data;

[0098] If there is a conflict in the entity matching result, calculate the similarity between the multi-source perception historical data and the first knowledge graph and the second knowledge graph respectively for comparison, and output the data in the knowledge graph with higher similarity as the multi-source perception processing data.

[0099] Specifically, the entity matching process of the present invention for the knowledge graph is essentially a screening process for multi-source perception data. By using the relationships between attributes, data that does not conform to the current situation is screened out to reduce incorrect data. Specifically: check whether there are conflicts in attributes or relationships between the successfully matched entity pairs. When there are no conflicts in the entity matching results, that is, there are no inconsistencies in attributes or relationships between all the successfully matched entity pairs, it indicates that all these data are valuable and can be retained for use. That is, these successfully matched entity pairs can be directly used as the output of multi-source perception processed data. When there are conflicts in the entity matching results, calculate the similarity between the multi-source perception historical data and the first knowledge graph and the second knowledge graph respectively based on multiple dimensions such as the overlap degree of entity attributes, the similarity of relationships, and the proximity of timestamps. And according to the similarity calculation results, select the data in the knowledge graph with high similarity and retain it as the output of multi-source perception processed data for subsequent use. The present invention automatically processes the conflict problems in entity matching through conflict detection and similarity calculation, ensures the accuracy of the selected data, can make full use of the complementarity of multi-source perception data, cleans and calibrates the multi-source heterogeneous perception information for the digital airspace system for UAV inspection, solves problems such as abnormal multi-source heterogeneous perception data, conflicts with each other, and data errors caused by sensor failures, and thus improves the utilization rate of multi-source perception data.

[0100] In the embodiments of the present application, based on the problem of how to improve the utilization rate of multi-source perception data collected during UAV power inspection, a processing method for multi-source perception data in a digital airspace is designed. It obtains multi-dimensional data during UAV power inspection and uses the multi-modal representation learning method to process these data to extract the association information and common features between different modal data, so as to effectively utilize the data information of multiple modalities and improve the accuracy and efficiency of data processing. Then, the obtained label data is combined with the multi-source perception historical data to generate a first knowledge graph covering multi-source perception data, and the real-time data is screened based on the differences between the historical data and the real-time data to remove redundant and abnormal data and retain the real-time data with a large difference from the historical data, thereby constructing a more accurate and targeted second knowledge graph. A machine learning model is used to process the first and second knowledge graphs to achieve real-time processing of data, and thus significantly improve the efficiency of UAV power inspection.

[0101] It should be noted that although the steps in the above flowcharts are shown in sequence according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear description in this article, the execution of these steps has no strict order limit, and these steps can be executed in other orders.

[0102] In another embodiment, as Figure 3As shown in the figure, the second aspect of the present invention provides a processing system for multi-source perception data in a digital airspace, including:

[0103] A multi-source data acquisition module 10, configured to acquire real-time multi-source perception data and historical multi-source perception data during the UAV power inspection collected by the digital airspace system;

[0104] A first knowledge graph generation module 20, configured to process the real-time multi-source perception data through a multi-modal representation learning method to obtain labeled data, and combine it with the historical multi-source perception data to generate a first knowledge graph;

[0105] A second knowledge graph construction module 30, configured to perform a difference calculation on the real-time multi-source perception data based on the historical multi-source perception data, and screen the real-time multi-source perception data according to the calculation result, so as to construct a second knowledge graph based on the screened real-time multi-source perception data;

[0106] A multi-source data processing module 40, configured to input the first knowledge graph and the second knowledge graph into a machine learning model for processing to obtain multi-source perception processed data.

[0107] It should be noted that each module in the above-mentioned processing system for multi-source perception data in a digital airspace can be implemented in whole or in part by software, hardware, and their combination. The above-mentioned modules can be embedded in the processor of the computer device in hardware form or independent of it, or stored in the memory of the computer device in software form, so that the processor can call and execute the operations corresponding to the above-mentioned modules. For the specific limitations of the processing system for multi-source perception data in a digital airspace, refer to the limitations of the processing method for multi-source perception data in a digital airspace in the above text. The two have the same functions and effects, and will not be elaborated here.

[0108] The third aspect of the present invention provides an electronic device, which includes:

[0109] A processor, a memory, and a bus;

[0110] The bus is used to connect the processor and the memory;

[0111] The memory is used to store operation instructions;

[0112] The processor is configured to execute the corresponding operations of the processing method for multi-source perception data in a digital airspace as shown in the first aspect of the present application by calling the operation instructions.

[0113] In an optional embodiment, an electronic device is provided, as Figure 4 shown Figure 4The electronic device 5000 shown includes: a processor 5001 and a memory 5003. Among them, the processor 5001 and the memory 5003 are connected, such as connected through a bus 5002. Optionally, the electronic device 5000 may further include a transceiver 5004. It should be noted that in practical applications, the transceiver 5004 is not limited to one, and the structure of the electronic device 5000 does not constitute a limitation to the embodiments of the present application.

[0114] The processor 5001 may be a CPU, a general-purpose processor, a DSP, an ASIC, an FPGA or other programmable logic devices, transistor logic devices, hardware components or any combination thereof. It can implement or execute various exemplary logic blocks, modules and circuits described in connection with the disclosure of the present application. The processor 5001 may also be a combination that implements computing functions, such as a combination including one or more microprocessors, a combination of a DSP and a microprocessor, etc.

[0115] The bus 5002 may include a path for transmitting information between the above components. The bus 5002 may be a PCI bus or an EISA bus, etc. The bus 5002 may be divided into an address bus, a data bus, a control bus, etc. For the sake of representation, Figure 4 only a thick line is used to represent it in the figure, but it does not mean that there is only one bus or one type of bus.

[0116] The memory 5003 may be a ROM or other types of static storage devices that can store static information and instructions, a RAM or other types of dynamic storage devices that can store information and instructions, or an EEPROM, a CD-ROM or other optical disc storage, optical disc storage (including compact discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), magnetic disk storage media or other magnetic storage devices, or any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto.

[0117] The memory 5003 is used to store the application program code for executing the solution of the present application, and is controlled by the processor 5001 to execute. The processor 5001 is used to execute the application program code stored in the memory 5003 to implement the content shown in any of the foregoing method embodiments.

[0118] Among them, the electronic device includes but is not limited to: mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Tablet Computers), PMPs (Portable Multimedia Players), vehicle terminals (such as vehicle navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc.

[0119] In a fourth aspect of the present invention, there is provided a computer-readable storage medium having a computer program stored thereon, and when the program is executed by a processor, it implements a method for processing multi-source perception data in a digital airspace as shown in the first aspect of the present application.

[0120] Another embodiment of the present application provides a computer-readable storage medium having a computer program stored thereon, and when it runs on a computer, it enables the computer to execute the corresponding content in the foregoing method embodiments.

[0121] In addition, an embodiment of the present invention further provides a computer-readable storage medium having a computer program stored thereon, and when the program is executed by a processor, it implements the steps of the above method.

[0122] In summary, the present invention relates to the technical field of information processing, and discloses a method, a system, a device and a medium for processing multi-source perception data in a digital airspace. The method includes obtaining multi-source perception real-time data and multi-source perception historical data during the UAV power inspection through a digital airspace system; processing the multi-source perception real-time data by a multi-modal representation learning method to obtain label data, and combining it with the multi-source perception historical data to generate a first knowledge graph; performing a difference calculation on the multi-source perception real-time data based on the multi-source perception historical data, and screening the multi-source perception real-time data according to the calculation result to construct a second knowledge graph based on the screened multi-source perception real-time data; inputting the first knowledge graph and the second knowledge graph into a machine learning model for processing to obtain multi-source perception processed data; realizing the comprehensive, accurate and efficient utilization of multi-source perception data during the UAV power inspection.

[0123] Each embodiment in this specification is described in a progressive manner. For parts that are the same or similar in each embodiment, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and reference can be made to the partial description of the method embodiment for related parts. It should be noted that the technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.

[0124] The above-described embodiments merely represent several preferred embodiments of the present application. The description thereof is relatively specific and detailed, but it should not be construed as a limitation on the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the technical principle of the present invention, several improvements and substitutions can be made, and these improvements and substitutions should also be regarded as the protection scope of the present application. Therefore, the protection scope of the patent of the present application shall be subject to the protection scope of the claims described above.

Claims

1. A method for processing digital spatial multi-source perception data, characterized in that: include: Acquire multi-source sensing real-time data and multi-source sensing historical data collected by the digital airspace system during the UAV power inspection process; Processing the multi-source perception real-time data through a multimodal representation learning method to obtain label data, and combining it with the multi-source perception historical data to generate a first knowledge graph; Based on the multi-source perception historical data, performing difference calculation on the multi-source perception real-time data, and filtering the multi-source perception real-time data according to the calculation result, so as to construct a second knowledge graph based on the filtered multi-source perception real-time data; The first knowledge graph and the second knowledge graph are input into a machine learning model for processing to obtain multi-source perception processing data.

2. The method for processing digital spatial multi-source perception data according to claim 1, characterized in that: The multi-source perception real-time data is processed by a multimodal representation learning method to obtain label data, and the label data is combined with the multi-source perception historical data to generate a first knowledge graph, including: Extracting features from the multi-source perception real-time data to obtain a multi-modal feature representation for input into a deep learning model for processing to obtain label data; wherein the deep learning model is trained using multi-source perception historical data with labeled data; The label data and the multi-source perception historical data are fused and aligned to obtain label fusion data, so as to construct an ontology library based on the label fusion data; The first knowledge graph is obtained by training an entity recognition model to identify entities corresponding to the ontology library from the label fusion data, and by training a relationship extraction model to extract the relationships between the entities from the label fusion data, so as to construct and optimize a knowledge graph based on the entities and the relationships.

3. The method for processing digital spatial multi-source perception data according to claim 2, characterized in that: The first knowledge graph includes a climate knowledge sub-graph and a network signal strength knowledge sub-graph; wherein, The step of constructing and optimizing a knowledge graph according to the entity and the relationship to obtain the first knowledge graph includes: Extracting a climate entity from the entity, taking the weather type in the climate entity as a climate center entity, and connecting related climate attributes through the climate center entity to obtain the climate knowledge subgraph; Extracting a network signal entity from the entity, taking the network signal strength in the network signal entity as the network signal center entity, and connecting related network signal attributes through the network signal center entity to obtain the network signal strength knowledge subgraph; The climate knowledge sub-graph and the network signal strength knowledge sub-graph are optimized to obtain the first knowledge graph.

4. The method for processing digital spatial multi-source perception data according to claim 1, characterized in that: The performing difference calculation on the multi-source perception real-time data based on the multi-source perception historical data, and screening the multi-source perception real-time data according to the calculation result, includes: Converting the multi-source perception historical data and the multi-source perception real-time data into historical data blocks and real-time data blocks respectively through a sliding window; Calculating data distribution between the historical data block and the real-time data block according to KL divergence and comparing them, so as to obtain a difference value between the multi-source perception historical data and the multi-source perception real-time data according to the comparison result; The difference value is compared with a preset difference threshold, and the multi-source perception real-time data is screened according to the comparison result.

5. The method for processing digital spatial multi-source perception data according to claim 1, characterized in that: The step of constructing a second knowledge graph based on the filtered multi-source perception real-time data includes: Extract multimodal data features from the filtered multi-source perception real-time data and perform feature fusion to obtain cross-modal feature data as sample data for training a cross-modal association model; The cross-modal feature data is processed by the trained cross-modal association model to construct and optimize the knowledge graph to obtain the second knowledge graph.

6. The method for processing digital spatial multi-source perception data according to claim 1, characterized in that: The step of inputting the first knowledge graph and the second knowledge graph into a machine learning model for processing to obtain multi-source perception processing data includes: Performing entity matching on the first knowledge graph and the second knowledge graph through a neural network model to obtain entity matching results, and performing conflict detection based on the entity matching results; If there is no conflict in the entity matching results, the successfully matched entity pairs are output as the multi-source perception processing data; If there is a conflict in the entity matching results, the similarity between the multi-source perception historical data and the first knowledge graph and the second knowledge graph is calculated for comparison, and the data in the knowledge graph with high similarity is output as the multi-source perception processed data.

7. The method for processing digital spatial multi-source perception data according to claim 6, characterized in that: The performing entity matching on the first knowledge graph and the second knowledge graph by using a neural network model to obtain an entity matching result includes: Converting entities and relationships in the first knowledge graph and entities and relationships in the second knowledge graph into a first natural language description and a second natural language description, respectively, and performing embedding transformation on the first natural language description and the second natural language description, respectively, to obtain a first high-dimensional semantic representation and a second high-dimensional semantic representation; Input the entities of the first knowledge graph and the entities of the second knowledge graph into the graph neural network model for feature extraction, respectively, to obtain the first entity structure feature and the second entity structure feature; fusing the first natural language description and the first entity structure feature, and the second natural language description and the second entity structure feature, respectively, to obtain a first fused feature and a second fused feature; Based on the first fusion feature and the second fusion feature, several candidate entity pairs are extracted from the first knowledge graph and the second knowledge graph to quantify the similarity between the candidate entity pairs, and the candidate entity pairs whose similarity exceeds a preset similarity threshold are output as successfully matched entity pairs.

8. A digital spatial multi-source sensing data processing system, characterized in that: include: A multi-source data acquisition module is used to acquire multi-source sensing real-time data and multi-source sensing historical data collected by the digital airspace system during the UAV power inspection process; A first graph generation module is used to process the multi-source perception real-time data through a multimodal representation learning method to obtain label data, and combine it with the multi-source perception historical data to generate a first knowledge graph; A second graph construction module is used to perform difference calculation on the multi-source perception real-time data based on the multi-source perception historical data, and filter the multi-source perception real-time data according to the calculation result, so as to construct a second knowledge graph based on the filtered multi-source perception real-time data; The multi-source data processing module is used to input the first knowledge graph and the second knowledge graph into the machine learning model for processing to obtain multi-source perception processing data.

9. An electronic device, characterized in that: The method comprises a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor, wherein when the processor executes the computer program, the method for processing digitized spatial multi-source perception data according to any one of claims 1 to 7 is implemented.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium includes a stored computer program, wherein when the device where the computer-readable storage medium is located executes the computer program, the method for processing digitized spatial multi-source perception data as described in any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Unsupervised knowledge graph fusion method and device based on multi-order neighborhood attention network

    CN112784065A

  • Knowledge graph fusion method and system

    CN117251578A

  • Power grid dispatching multi-mode knowledge graph construction method and system

    CN118035463A

  • Media data recommendation

    US20250371089A1

  • Fault knowledge graph construction method and apparatus

    WO2023045417A1