An unmanned aerial vehicle visual reasoning and question-answering method in a near-earth security rescue scene

CN118298335BActive Publication Date: 2026-09-25NORTHWESTERN POLYTECHNICAL UNIV
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202410432750.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-04-11
Publication Date
2026-09-25
Estimated Expiration
2044-04-11

AI Technical Summary

Technical Problem

[0004]上述算法存在两个主要缺陷:第一,特征学习策略过于简单,提取的特征往往粒度粗糙且语义层次单一,这将导致问答系统缺乏对不同尺度、不同层次语义信息的表征能力不足;第二,缺少对显式和隐式关系的推理建模,这将导致问答系统无法感知不同目标间的潜在关系,进而无法回答需要推理的复杂问题

Benefits of technology

[0038]本发明能够学习不同粒度的多模态特征信息,分析不同尺度、不同层次的目标信息,推理不同目标间的隐含关系,有效提高实际救援场景下回答的准确性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118298335B_ABST
    Figure CN118298335B_ABST
Patent Text Reader

Abstract

The application discloses a kind of unmanned aerial vehicle vision reasoning and question and answer method under the scene of security rescue, integrates unmanned aerial vehicle camera system and terminal workstation, wherein unmanned aerial vehicle camera system is responsible for visual information collection and transmission, terminal workstation is deployed vision question and answer system to carry out real-time man-machine interaction.The method of the application can learn multi-modal feature information of different granularity, analyze target information of different scales and different levels, infer the implicit relationship between different targets, and effectively improve the accuracy of the answer in actual rescue scene.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of artificial intelligence technology, specifically relating to a method for visual reasoning and question answering of unmanned aerial vehicles (UAVs) in on-site security and rescue scenarios. Background Technology

[0002] Temporary site security refers to a diversified, cross-domain, three-dimensional, collaborative, and intelligent system addressing the needs for safety, rescue, protection, and production within a temporary site. Within this system, drones play a crucial role in various scenarios such as disaster assessment, disaster relief, and geological exploration. However, current drone interaction systems heavily rely on human intervention, lacking intelligent reasoning systems for extracting and analyzing information from captured video images, and lacking real-time question-and-answer systems for rapid responses to extracted visual information. This significantly reduces information processing efficiency and emergency response speed.

[0003] Visual reasoning and question answering systems are a typical interdisciplinary artificial intelligence task, involving computer vision, natural language processing, knowledge graphs, and other fields. Its basic form is as follows: input an image and a series of questions about the image's content into the question answering system, and the system will provide the correct answers to the corresponding questions. In this process, the visual reasoning and question answering system first extracts feature information from both the image and text modalities, then performs semantic alignment and feature fusion between the two modalities, and finally inputs the information into a classifier to predict the correct answer. Traditional visual question answering systems typically use convolutional neural networks and recurrent neural networks to learn visual and text representations respectively, then use attention mechanisms to calculate attention scores between the two modalities to achieve semantic alignment, then use feature fusion methods such as concatenation and Hadmar product to obtain fused features, and finally use a multilayer perceptron for answer prediction.

[0004] The aforementioned algorithms suffer from two main drawbacks: First, the feature learning strategy is overly simplistic, resulting in coarse-grained features with limited semantic depth. This leads to a lack of representational ability for question-answering systems across different scales and levels of semantic information. Second, the lack of reasoning modeling for explicit and implicit relationships prevents the question-answering system from perceiving potential relationships between different targets, thus hindering its ability to answer complex questions requiring reasoning. Therefore, from an algorithm performance perspective, a technical solution is urgently needed to address these two shortcomings and achieve a breakthrough in algorithmic performance.

[0005] Furthermore, from an application perspective, current visual reasoning and question-answering methods lack integration with UAV camera systems. Current research primarily focuses on the performance of visual reasoning and question-answering systems in general natural scenes, with a few studies investigating their effectiveness in remote sensing data scenarios. Research on close integration with UAV systems is scarce. Therefore, from an application perspective, there is an urgent need to design a technical method to address the issue of integration with UAV platforms. Summary of the Invention

[0006] To overcome the shortcomings of existing technologies, this invention provides a UAV visual reasoning and question-answering method for on-site security and rescue scenarios. It integrates a UAV camera system and a terminal workstation, where the UAV camera system is responsible for visual information acquisition and transmission, and the terminal workstation deploys a visual question-answering system for real-time human-computer interaction. This invention's method can learn multimodal feature information at different granularities, analyze target information at different scales and levels, and infer the implicit relationships between different targets, effectively improving the accuracy of responses in actual rescue scenarios.

[0007] The technical solution adopted by this invention to solve its technical problem is as follows:

[0008] Step 1: Remotely control the drone through the remote terminal workstation. The drone acquires aerial optical images over the disaster relief area. The aerial optical images are used as a real-time video stream to be processed and transmitted by the drone to the remote terminal workstation through the communication data link.

[0009] Step 2: The remote terminal workstation receives the real-time video stream transmitted by the UAV from the communication data link; it uses a fixed-frequency frame extraction method to extract aerial images from the real-time video stream to form a sparse image data stream, and inputs the sparse image data stream into the visual reasoning and question answering system deployed on the remote terminal workstation.

[0010] Step 3: The visual reasoning and question answering system comprises four parts: a multi-scale visual encoder, a multi-level language encoder, an intra-modal-inter-modal graph reasoning unit, and a parallel cross-modal graph reasoning network.

[0011] Step 4: The multi-scale visual encoder uses a ResNet convolutional network as its backbone and employs the feature pyramid method of the Faster-RCNN object detection network to represent image information at three different scales: "target level," "region level," and "image level." The "target level" objects are used to represent vehicles and houses; the "region level" objects are used to represent streets and roads; and the "image level" objects are used to represent farmland, ridges, and lakes.

[0012] Step 5: The multi-level language encoder uses BERT as its basic framework and learns language features at different levels: word level, phrase level, and sentence level. The word level represents each word individually, the phrase level represents phrases by performing convolutions with different phase lengths, and the sentence level represents sentences by clauses and complete sentences.

[0013] Step 6: The intra-modal-inter-modal graph reasoning unit is based on a graph convolutional network. It first calculates the intra-modal attention score, then aggregates all intra-modal information by point aggregation of the graph, and then uses cross-modal graph attention reasoning to perform semantic alignment between different modalities.

[0014] Step 7: The cross-modal graph reasoning network is based on intra-modal and inter-modal graph reasoning units. It aligns the "target-level", "region-level", and "image-level" features of the visual modality with the "word-level", "region-level", and "image-level" features of the language modality, and performs cross-modal graph reasoning in parallel.

[0015] Furthermore, the input to the visual reasoning and question-answering system is an aerial optical image V captured by the UAV and a related query question Q, with the goal of selecting candidate answers. The correct answer is predicted as a. * Its mathematical expression is as follows:

[0016]

[0017] Among them, F θ It is a visual question answering model with learnable parameters, where 'a' is one of the candidate answers. One answer.

[0018] Furthermore, the multi-scale visual encoder first uses a pre-trained ResNet convolutional neural network, and then learns layer by layer according to the feature pyramid method;

[0019] The image is fed into the encoder, and object-level and image-level feature maps {V} are obtained from conv5 and avgpool respectively. o V i}; Then, for V o Applying MaxPool to obtain regional visual information V r Finally, three projection matrices {W} are used. o W r W i Maintaining the same dimensions for features; the above process can be represented as:

[0020] {V o V i} = ResNet(V)

[0021] V r =MaxPool(V o )

[0022]

[0023] Furthermore, the multi-level language encoder, given a sequence of question words T = {t1, t2, ..., t...} m First, the BERT tokenizer is applied to generate word-level language features. Then, for Q wPerform Conv1D operations; calculate the inner product of each word vector using two filters with different window sizes, and apply MaxPool to the corresponding features to obtain phrase-level language features. Then, the AvgPool operation is used to collect information from the entire sentence and encode sentence-level features. Three projection matrices {W} are introduced. w W p W s To maintain the dimensionality of the multi-level representation; the above process is represented as:

[0024] Q w =Tokenizer(T)

[0025] Q p =MaxPool(bi(Q) w ),tri(Q w ))

[0026] Q s =AvgPool(Q w )

[0027]

[0028] Furthermore, the intra-modal-inter-modal graph reasoning unit first constructs both a visual graph and a language graph simultaneously. Where mo∈{v,p} represents a modality, v represents a visual modality, and p represents a linguistic modality; se∈{o,r,i,w,p,s} represents semantics, o represents target-level semantics, r represents region-level semantics, i represents scene-level semantics, w represents word-level semantics, and s represents sentence-level semantics; then, within the same modality, a node aggregation strategy is used on the graph to perform intra-modal interactions and calculate the attention distribution; this process is described as:

[0029]

[0030]

[0031]

[0032] in, I is a learnable parameter matrix, and I is a diagonal matrix of jump connections;

[0033] After obtaining enhanced node features and Then, a cross-modal graph attention mechanism is applied to enable multi-granular semantic interactions between different modalities.

[0034] Furthermore, in the cross-modal graph reasoning network, the graph relationship reasoning process between different semantic levels is constructed in parallel, i.e., from the target level to the word level, from the region level to the phrase level, and from the image level to the sentence level; then, various features are fused using Hadamard product; finally, the final representation is input into a two-layer MLP classifier to predict the corresponding answer; and the parameters are optimized using the cross-entropy loss function. The calculation formula is as follows:

[0035]

[0036] Among them, y i The answer is a. i The tag, f fus It is the final feature, MLP(f) fus ) represents the corresponding probability.

[0037] The beneficial effects of this invention are as follows:

[0038] This invention can learn multimodal feature information at different granularities, analyze target information at different scales and levels, and infer the implicit relationships between different targets, effectively improving the accuracy of responses in actual rescue scenarios. Attached Figure Description

[0039] Figure 1 This is a complete flowchart of the present invention;

[0040] Figure 2 This is an example diagram of the multi-level representation learning of the present invention;

[0041] Figure 3 This is a schematic diagram of the multi-scale visual encoder of the present invention;

[0042] Figure 4 This is a schematic diagram of the multi-level language encoder of the present invention;

[0043] Figure 5 This is a schematic diagram of the intra-modal-inter-modal inference unit of the present invention;

[0044] Figure 6 This is a schematic diagram of the parallel cross-modal graph inference network of the present invention. Detailed Implementation

[0045] The present invention will be further described below with reference to the accompanying drawings and embodiments.

[0046] To overcome the shortcomings of existing technologies, this invention provides a UAV visual reasoning and question-answering method for on-site security and rescue scenarios. It integrates a UAV camera system and a terminal workstation, where the UAV camera system is responsible for visual information acquisition and transmission, and the terminal workstation deploys a visual question-answering system for real-time human-computer interaction. This reasoning and question-answering method can learn multimodal feature information at different granularities, analyze target information at different scales and levels, and infer the implicit relationships between different targets, effectively improving the accuracy of responses in actual rescue scenarios.

[0047] The main steps include:

[0048] S1: The drone is remotely controlled via a remote terminal workstation, acquiring aerial optical images over the disaster relief area. This data, as a real-time video stream to be processed, is transmitted from the drone to the remote terminal workstation via a communication data link.

[0049] S2: The remote terminal workstation receives the real-time video stream transmitted by the UAV from the communication data link. Aerial images are extracted from the video stream using a fixed-frequency frame extraction method to form a sparse image data stream, which is then input into the visual reasoning and question-answering system deployed on the terminal workstation.

[0050] S3: The visual reasoning and question answering system mainly includes four core algorithm steps: multi-scale visual encoder, multi-level language encoder, intra-modal-inter-modal graph reasoning unit, and parallel cross-modal graph reasoning network.

[0051] S4: For the multi-scale visual encoder, a ResNet convolutional network is used as the backbone network, and the feature pyramid method of the Faster-RCNN object detection network is adopted to represent image information according to three different scales: "object level", "region level" and "image level". Specifically, the "object level" object scale is the smallest and is used to represent micro-objects such as vehicles and houses; the "region level" object scale is medium and is used to represent medium-sized objects such as blocks and roads; the "image level" object scale is the largest and is used to represent macro-objects such as farmland, ridges and lakes.

[0052] S5: For multi-level language encoders, using BERT as the basic framework, different levels of language features are learned at the "word level," "phrase level," and "sentence level." Specifically, "word level" represents each word individually, "phrase level" represents phrases through convolution with different phase lengths, and "sentence level" represents sentences by clauses and complete sentences. Combined with a multi-scale visual encoder, "word level" language features are often semantically similar to "target level" visual features, both pointing to micro-objects; "phrase level" language features are often semantically similar to "region level" visual features, both representing medium-sized objects and their relationships; and "sentence level" language features are often semantically similar to "image level" visual features, both representing macro-objects and their relationships.

[0053] S6: For intra-modal and inter-modal graph reasoning units, based on graph convolutional networks, intra-modal attention scores are first calculated to highlight key information within the modality. Then, graphs are aggregated to gather all information within the modality. Finally, cross-modal graph attention reasoning is used to perform semantic alignment between different modalities, so as to achieve the reasoning effect of strengthening the reasoning of relevant information and weakening irrelevant information.

[0054] S7: For cross-modal graph reasoning networks, the feature is that, based on the above graph reasoning units, the "target-level", "region-level", and "image-level" features of the visual modality are semantically aligned with the "word-level", "region-level", and "image-level" features of the language modality, and cross-modal graph reasoning is performed in parallel.

[0055] Example:

[0056] The input to the visual reasoning and question answering system is an aerial optical image V taken by a drone and a related query question Q. The goal is to select candidate answers. The correct answer is predicted as a. * Its mathematical expression is as follows:

[0057]

[0058] Where, F θ It is a visual question answering model with learnable parameters, where 'a' is the candidate answer x. One answer.

[0059] For multi-scale visual encoders, a pre-trained ResNet convolutional neural network is first used, followed by layer-by-layer learning using a feature pyramid approach. Specifically, the image is fed into the encoder, and object-level and image-level feature maps {V} are obtained from conv5 and avgpool, respectively. o V i Then, in order to obtain regional visual information V r We are concerned about V o MaxPool is applied. Finally, to maintain the same dimensions for these features, three projection matrices {W} are used. o W r W i The above process can be represented as:

[0060] {V o V i} = ResNet(V),

[0061] V r =MaxPool(V o ),

[0062]

[0063] For a multi-level language encoder, given a sequence of question words T = {t1, t2, ..., t...} m First, BERT tokenizer is applied to generate word-level language features. Then, in order to extract phrase-level semantics, Q is... w Conv1D operations are performed. Specifically, two filters (with different window sizes: bigram and trigram) are used to calculate the inner product of each word vector, and MaxPool is applied to the corresponding features to obtain phrase-level language features. Then, the AvgPool operation is used to collect information from the entire sentence and encode sentence-level features. In addition, three projection matrices {W} are introduced. w W p W s This is done to maintain the dimensionality of the multi-level representation. The above process can be represented as:

[0064] Q w =Tokenizer(T),

[0065] Q p =MaxPool(bi(Q) w ),tri(Q w )),

[0066] Q s =AvgPool(Q w ),

[0067]

[0068] For the intra-modal-inter-modal graph reasoning unit, visual graphs and language graphs are constructed simultaneously. Where mo∈{v,p} represents a modality, and se∈{o,r,i,w,p,s} represents a semantic. Then, within the same modality, a node aggregation strategy is applied to the graph to perform intra-modal interactions and compute the attention distribution. This process can be described as follows:

[0069]

[0070]

[0071]

[0072] in, is a learnable parameter matrix, and I is the diagonal matrix of jump connections. This is used to obtain the enhanced node features. and Subsequently, a cross-modal graph attention mechanism is applied to enable multi-granular semantic interactions between different modalities. Taking the visual target-language word level as an example, the visual target-level representation process can be described as follows:

[0073]

[0074]

[0075]

[0076] in, It is an aggregation feature in the word graph. It is the attention score of each node in the object graph. It is a trainable parameter matrix. This results in the final object-level visual features. As for word-level graph reasoning, it is a conjugate of the above process, yielding the corresponding features. Furthermore, the operations are similar for the region-phrase level semantic layer and the image-sentence level semantic layer.

[0077] The encoder module obtains multi-granularity features, a necessary foundation for solving complex problems with multiple semantic layers. The inference unit module allows the model to calculate not only the correlation between nodes within the same modality and their neighbors, but also the similarity between nodes in different modalities and those in another modality. Based on these two steps, the model can strengthen semantic features relevant to the current question and weaken irrelevant information, thereby analyzing which node is more important for answering the question.

[0078] To improve the efficiency of the above reasoning process, a parallel graph relation reasoning network was designed. Specifically, the graph relation reasoning process between different semantic levels is constructed in parallel (target level to word level, region level to phrase level, image level to sentence level). Then, a simple Hadamard product is used to fuse various features. Finally, the final representation is input into a two-layer MLP classifier to predict the corresponding answer. The cross-entropy loss function is used to optimize the parameters, and its calculation formula is as follows:

[0079]

[0080] Among them, y i The answer is a. i The tag, f fus It is the final feature, MLP(f) fus ) represents the corresponding probability.

[0081] Table 1. Performance of the present invention on the LR dataset.

[0082]

[0083] Table 2. Performance of the present invention on the HR dataset.

[0084]

[0085] To verify the effectiveness of the designed visual reasoning and question answering algorithm, tests were conducted on the LR and HR datasets. The results in Tables 1 and 2 show that the algorithm significantly improves performance compared to existing methods.

Claims

1. A method for visual reasoning and question answering using unmanned aerial vehicles (UAVs) in a local security and rescue scenario, characterized in that, Includes the following steps: Step 1: Remotely control the drone through the remote terminal workstation. The drone acquires aerial optical images over the disaster relief area. The aerial optical images are used as a real-time video stream to be processed and transmitted by the drone to the remote terminal workstation through the communication data link. Step 2: The remote terminal workstation receives the real-time video stream transmitted by the UAV from the communication data link; it uses a fixed-frequency frame extraction method to extract aerial images from the real-time video stream to form a sparse image data stream, and inputs the sparse image data stream into the visual reasoning and question answering system deployed on the remote terminal workstation. Step 3: The visual reasoning and question answering system comprises four parts: a multi-scale visual encoder, a multi-level language encoder, an intra-modal-inter-modal graph reasoning unit, and a parallel cross-modal graph reasoning network. Step 4: The multi-scale visual encoder uses a ResNet convolutional network as its backbone and employs the feature pyramid method of the Faster-RCNN object detection network to represent image information at three different scales: "target level," "region level," and "image level." The "target level" objects are used to represent vehicles and houses; the "region level" objects are used to represent streets and roads; and the "image level" objects are used to represent farmland, ridges, and lakes. Step 5: The multi-level language encoder uses BERT as its basic framework and learns language features at different levels: word level, phrase level, and sentence level. The word level represents each word individually, the phrase level represents phrases by performing convolutions with different phase lengths, and the sentence level represents sentences by segmentation and whole sentences. Step 6: The intra-modal-inter-modal graph reasoning unit is based on a graph convolutional network. It first calculates the intra-modal attention score, then aggregates all intra-modal information by point aggregation of the graph, and then uses cross-modal graph attention reasoning to perform semantic alignment between different modalities. Step 7: The cross-modal graph reasoning network is based on intra-modal and inter-modal graph reasoning units. It aligns the "target-level", "region-level", and "image-level" features of the visual modality with the "word-level", "region-level", and "image-level" features of the language modality, and performs cross-modal graph reasoning in parallel.

2. The UAV visual reasoning and question-answering method for on-site security and rescue scenarios according to claim 1, characterized in that, The visual reasoning and question-answering system takes as input an aerial optical image V captured by a drone and a related query question Q, and aims to select candidate answers... The correct answer is predicted as a. * Its mathematical expression is as follows: Among them, F θ It is a visual question answering model with learnable parameters, where 'a' is one of the candidate answers. One answer.

3. The UAV visual reasoning and question-answering method for on-site security and rescue scenarios according to claim 1, characterized in that, The multi-scale visual encoder first uses a pre-trained ResNet convolutional neural network, and then learns layer by layer according to the feature pyramid method. The image is fed into the encoder, and object-level and image-level feature maps {V} are obtained from conv5 and avgpool respectively. o V i }; Then, for V o Applying MaxPool to obtain regional visual information V r Finally, three projection matrices {W} are used. o W r W i Maintaining the same dimensions for features; the above process can be represented as: {V o ,V i }=ResNet(V) V r =MaxPool(V o ) 4. The UAV visual reasoning and question-answering method for on-site security and rescue scenarios according to claim 1, characterized in that, The multi-level language encoder, given a sequence of question words T = {t1, t2, ..., t...} m First, BERT tokenizer is applied to generate word-level language features. Then, for Q w Perform Conv1D operations; calculate the inner product of each word vector using two filters with different window sizes, and apply MaxPool to the corresponding features to obtain phrase-level language features. Then, the AvgPool operation is used to collect information from the entire sentence and encode sentence-level features. Three projection matrices {W} are introduced. w W p W s To maintain the dimensionality of the multi-level representation; the above process is represented as: Q w =Tokenizer(T) Q p =MaxPool(bi(Q w ),tri(Q w )) Q s =AvgPool(Q w ) 5. The UAV visual reasoning and question-answering method for on-site security and rescue scenarios according to claim 1, characterized in that, The intra-modal-inter-modal graph reasoning unit first constructs both a visual graph and a language graph simultaneously. Where mo∈{v,p} represents a modality, v represents a visual modality, and p represents a linguistic modality; se∈{o,r,i,w,p,s} represents semantics, o represents target-level semantics, r represents region-level semantics, i represents scene-level semantics, w represents word-level semantics, and s represents sentence-level semantics; then, within the same modality, a node aggregation strategy is used on the graph to perform intra-modal interactions and calculate the attention distribution; this process is described as follows: in, I is a learnable parameter matrix, and I is a diagonal matrix of jump connections; After obtaining enhanced node features and Then, a cross-modal graph attention mechanism is applied to enable multi-granular semantic interactions between different modalities.

6. The UAV visual reasoning and question-answering method for on-site security and rescue scenarios according to claim 1, characterized in that, The cross-modal graph reasoning network constructs graph relationship reasoning processes in parallel across different semantic levels, i.e., from the target level to the word level, from the region level to the phrase level, and from the image level to the sentence level. Then, various features are fused using the Hadamard product. The final representation is then input into a two-layer MLP classifier to predict the corresponding answer. The parameters are optimized using the cross-entropy loss function, the calculation formula of which is: Among them, y i The answer is a. i The tag, f fus It is the final feature, MLP(f) fus ) represents the corresponding probability.

Citation Information

Patent Citations

  • Adaptive, individualized, and contextualized text-to-speech systems and methods

    US20240194178A1