Target damage assessment method based on large model quantification lateral fine tuning

By using large-model quantization lateral fine-tuning technology in target damage assessment, a DAMAGE-VQA fine-grained target damage assessment network was constructed, which solved the problems of training difficulties and evaluation result errors in the existing technology, and achieved high-precision fine-grained damage assessment.

CN120164085APending Publication Date: 2025-06-17NORTHWESTERN POLYTECHNICAL UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510205477.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-24
Publication Date
2025-06-17

AI Technical Summary

Technical Problem

The existing target damage assessment method is based on complex nested deep learning models, resulting in training difficulties and errors in evaluation results.

Method used

Using a target damage assessment method based on large-scale quantization lateral fine-tuning, we can achieve lateral fine-tuning by constructing a DAMAGE-VQA fine-grained target damage assessment network, quantifying the graphic and text fusion encoder and reply decoder, adding a lateral branch network, and performing gradient-independent downsampling and weighted information fusion to achieve lateral fine-tuning.

Benefits of technology

The accuracy of target damage assessment is improved, and the generated fine-grained damage assessment results can provide scientific basis for decision-making and evaluation, and optimize resource allocation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120164085A_ABST
    Figure CN120164085A_ABST
Patent Text Reader

Abstract

The invention particularly relates to a target damage assessment method based on large model quantification lateral fine tuning, which comprises the following steps: acquiring a damage target image, and performing question and answer text annotation on the damage target image to construct a training sample; constructing a DAMAGE-VQA fine-grained target damage assessment network, quantifying each layer of an image-text fusion encoder and a part of a reply decoder in the DAMAGE-VQA fine-grained target damage assessment network, and adding lateral branch networks to the quantized image-text fusion encoder and reply decoder layer by layer; and inputting the training sample into a DAMAGE-VQA fine-grained target damage assessment network for training, and outputting the category of the damage assessment sample, the overall damage degree and the fine-grained damage degree of each part. According to the method, accurate fine tuning is carried out on the image-text fusion encoder and the reply decoder, so that the image text pre-training large model is more suitable for a fine-grained damage assessment task of a target, the assessment accuracy is improved, and a detailed damage assessment result of a part is generated.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of target assessment, and relates to a method for target damage assessment in the form of fine-grained visual question answering, and particularly to a method for target damage assessment based on large model quantization lateral fine-tuning. Background Art

[0002] Traditional target damage assessment methods rely on mathematical statistical models such as fuzzy inference and Markov chains. Although these methods can provide preliminary estimates, they highly depend on prior knowledge and assumptions and are difficult to cope with complex environments, resulting in limited accuracy of assessment results. At the same time, although traditional deep learning models can handle tasks with a certain degree of complexity, most of them require complex coupling and nesting of detection and segmentation models, with low training efficiency and difficulty in rapid deployment. Large-scale pre-trained visual question answering (VQA) models have powerful image understanding and semantic reasoning capabilities, can automatically extract target features from image data, and perform high-precision damage assessment. Compared with traditional models, VQA large models have stronger adaptability. Although the fine-tuning process requires a large amount of computing resources, various efficient fine-tuning methods can be used to reduce resource consumption and quickly adapt to task requirements. How to design an efficient fine-tuning method to introduce image-text large models into target damage research and generate accurate and fine-grained damage assessment results will become the top priority.

[0003] It should be noted that the information disclosed in the above background art section is only used to enhance the understanding of the background of the present invention, and thus may include information that does not constitute the prior art known to those of ordinary skill in the art. Summary of the Invention

[0004] The present invention provides a method for target damage assessment based on large model quantization lateral fine-tuning, which is used to overcome the problems that the existing target damage assessment methods are difficult to train due to multiple complex nested deep learning models and there are errors in the assessment results.

[0005] Other features and advantages of the present invention will become apparent through the following detailed description, or be learned in part through the practice of the present invention.

[0006] According to the first aspect of the present invention, there is provided a method for target damage assessment based on large model quantization lateral fine-tuning, the method comprising:

[0007] Obtain a damaged target image, and perform question-and-answer text annotation on the damaged target image to construct a training sample;

[0008] Construct a DAMAGE-VQA fine-grained target damage assessment network. Among them, the DAMAGE-VQA fine-grained target damage assessment network includes a question text encoder, an image encoder, a text-image fusion encoder, and an answer decoder. Quantize each layer of the text-image fusion encoder and part of the answer decoder; add lateral branch networks layer by layer to the quantized text-image fusion encoder and answer decoder, and dequantize the corresponding backbone network layers. Reduce the number of parameters and align the branch networks through gradient-free downsampling. The backbone and branch networks perform weighted information fusion to form a lateral fine-tuning network. The final output of the lateral network is upsampled and then weighted and fused with the backbone network, and the backbone network layer maintains the quantized state unchanged;

[0009] Input the training samples into the DAMAGE-VQA fine-grained target damage assessment network for training. Input the training samples and the corresponding question-and-answer text annotations into the image encoder and the question text encoder of the Transformer architecture respectively to extract features; the features of the image encoder and the features of the question text encoder interact fully in the text-image fusion encoder, gradually realizing the alignment of the image features and the question text features, and outputting a high-level unified representation of the image features and the question text features under the guidance and optimization of the lateral fine-tuning network; input the output result of the text-image fusion encoder into the answer decoder, and based on the text-image alignment unified representation and the initialized answer representation, fully fuse through self-attention and cross-attention, and output the category of the damage assessment sample, the overall damage degree, and the fine-grained damage degree of each part under the guidance of the lateral fine-tuning network.

[0010] In some exemplary embodiments, the method further includes data augmentation for the training samples, where the data augmentation includes cropping, flipping, and shifting.

[0011] In some exemplary embodiments, quantizing each layer of the text-image fusion encoder and part of the answer decoder specifically means: quantizing all the parameters of the selected network layer to 4-bit.

[0012] In some exemplary embodiments, adding the lateral branch network includes:

[0013] Copy the structure of the backbone network layer as the main body of the lightweight lateral branch network;

[0014] The output of the original quantized backbone network layer Dequantize back to And after downsampling, it is weighted and fused with the output of the previous lateral branch network As the output of the current lateral branch network

[0015] Stack lateral branch networks in sequence among multiple network layers designated for lateral fine-tuning. In the last layer, after upsampling the output result of the lateral branch network, perform weighted fusion with the output of the backbone network to output the final result. After upsampling, it is weighted and fused with the output of the backbone network to output the final result.

[0016] In some exemplary embodiments, the method further includes:

[0017] Fine-tune the DAMAGE-VQA fine-grained target damage assessment network. Select AdamW as the optimizer and use the image-text contrast loss L QIC and the reply result cross-entropy loss L RCE ; Promote the full alignment of image-text feature information during the fine-tuning process and constrain the reply decoder to generate the desired fine-grained target damage result.

[0018] In some exemplary embodiments, the image-text contrast loss L QIC Through the contrast learning method, ensure the maximization of the similarity between the question text features and the image features, and achieve the high-matching alignment of image-text features. It is implemented based on the principle of maximizing information quantity using the contrast loss:

[0019]

[0020] Among them, sim(I,Q) calculates the similarity between the output features I of the image encoder and the output features Q of the question text encoder. I ′ is the negative sample, and τ is the scale adjustment coefficient of the loss function.

[0021] In some exemplary embodiments, the reply result cross-entropy loss L RCE is used to compare the gap between the final result output by the reply encoder and the label reply:

[0022]

[0023] Among them, y i and are the reply result output of the model and the label reply result respectively.

[0024] According to the second aspect of the present invention, there is provided a storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the target damage assessment method based on large model quantization lateral fine-tuning described in the first aspect above.

[0025] According to the third aspect of the present invention, there is provided a computer program product, on which a computer program is stored. When the computer program is executed by a processor, it implements the target damage assessment method based on large model quantization lateral fine-tuning described in the first aspect above.

[0026] According to a fourth aspect of the present invention, there is provided an electronic device, comprising:

[0027] a processor; and

[0028] a memory for storing executable instructions of the processor;

[0029] wherein the processor is configured to implement the target damage assessment method based on large model quantization lateral fine-tuning described in the first aspect above when executing the executable instructions.

[0030] The target damage assessment method based on large model quantization lateral fine-tuning provided by the embodiments of the present invention is designed based on an image-text large model, which can make full use of the feature extraction and text understanding advantages of the large model. There is no need to combine complex detection and recognition semantic segmentation networks for retraining, nor to introduce prior guidance, and it can directly output fine-grained damage level assessment results for the whole target and each part. By combining the provided quantization lateral fine-tuning technology, precise fine-tuning can be performed on the key parts of model fine-tuning, namely the image-text fusion encoder and the response decoder, so that the pre-trained image-text large model is more applicable to the fine-grained damage assessment task of the target, improving the accuracy of the assessment. The generated damage assessment results detailed to the parts can provide a more scientific basis for decision-making, assessment and other work, and optimize the reasonable allocation of resources.

[0031] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0032] The drawings herein are incorporated into the specification and constitute a part of the specification, showing embodiments consistent with the present invention, and are used together with the specification to explain the principles of the present invention. Obviously, the drawings in the following description are only some embodiments of the present invention, and those of ordinary skill in the art can obtain other drawings based on these drawings without creative efforts.

[0033] Figure 1 It is a flowchart of a target damage assessment method based on large model quantization lateral fine-tuning for an exemplary embodiment of the present invention;

[0034] Figure 2 It is a DAMAGE-VQA fine-grained target damage assessment network model diagram designed for an exemplary embodiment of the present invention;

[0035] Figure 3 It is a quantization fine-tuning lateral network model diagram designed for an exemplary embodiment of the present invention;

[0036] Figure 4 Model structure schematic diagram and information gradient propagation schematic diagram;

[0037] Figure 5 The vehicle damage assessment result generated by the method of the exemplary embodiment of the present invention. Detailed implementation manners

[0038] Example embodiments will now be described more fully with reference to the accompanying drawings. However, the example embodiments can be implemented in various forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided so that this invention will be more complete and comprehensive, and will fully convey the concept of the example embodiments to those skilled in the art. The features, structures, or characteristics described may be combined in any suitable manner in one or more embodiments.

[0039] In addition, the accompanying drawings are only schematic illustrations of the present invention and are not necessarily drawn to scale. The same reference numerals in the drawings denote the same or similar parts, and thus repeated descriptions thereof will be omitted. Some of the block diagrams shown in the drawings are functional entities and do not necessarily correspond to physically or logically independent entities. These functional entities may be implemented in the form of software, or implemented in one or more hardware modules or integrated circuits, or implemented in different networks and / or processor devices and / or microcontroller devices.

[0040] In view of the disadvantages and deficiencies of the prior art, in this example embodiment, a target damage assessment method based on large model quantization lateral fine-tuning is provided. Target samples are acquired, and visual question-answering annotation of the data is performed using Labelimg software and self-built annotation codes. Data augmentation is performed on the data to generate a training set and a test set; a DAMAGE-VQA fine-grained target damage assessment network is used, and a quantization fine-tuning lateral network is established based on it; the above overall network is fine-tuned using the training set; the test data is input into the overall network for testing to obtain the fine-grained damage assessment result of the target.

[0041] Refer to Figure 1 As shown, the method may specifically include the following steps:

[0042] Step 1: Deploy drones and ground cameras to collect damage target samples, use Labelimg to label the categories of the targets in the image samples, divide and label the part names according to the damage tree, use the self-built code to label the question content and the damage levels of each part for the image samples, fuse all the data annotations, and divide the above samples into a target group damage assessment training set and a test set at a ratio of 8:2;

[0043] Step 2: Perform data augmentation on the target group damage assessment training set, such as cropping, flipping, and shifting, to generate training samples with a dimension of H×W×3 and associate corresponding text labels, and perform semantic replacement of the same meaning on the question content in the labels;

[0044] Step 3: Construct the DAMAGE-VQA fine-grained target damage assessment network. As shown in Figure 2 , the DAMAGE-VQA fine-grained target damage assessment network includes a question text encoder, an image encoder, a text-image fusion encoder, and an answer decoder. Among them, both the question text encoder and the image encoder adopt the Transformer architecture, including a feed-forward layer and a self-attention layer; the text-image fusion encoder and the answer decoder include a feed-forward layer, a cross-attention layer, and a self-attention layer; the feed-forward layer of the question text encoder is connected to the self-attention layer of the text-image fusion encoder, and the feed-forward layer of the image encoder is connected to the cross-attention layer of the text-image fusion encoder; the feed-forward layer of the text-image fusion encoder is connected to the cross-attention layer of the answer decoder. Selective quantization is performed on the network layers, and only the layers of the text-image fusion encoder and part of the answer decoder are quantized. All the parameters of the selected network layers are quantized to 4-bit.

[0045] Step 4: Add lateral branch networks layer by layer to the text-image fusion encoder and the answer decoder in the quantized text-image fusion encoder, and de-quantize the corresponding main network layers. Reduce the number of parameters and align the branch networks through gradient-free downsampling. The main network and the branch networks are combined through weighted information fusion to form a lateral fine-tuning network. The final output of the lateral network is upsampled and then weighted and fused with the main network, and the main network layer maintains the quantized state unchanged.

[0046] Reference Figure 3 , step 4 of adding lateral branch networks specifically includes the following steps:

[0047] Step 4-1: Copy the structure of the main network layer, reduce the hidden layer dimension by 4-8 times, and use it as the main body of the lightweight lateral branch network;

[0048] Step 4-2: De-quantize the output of the original quantized main network layer back to and after downsampling, perform weighted fusion with the output of the previous lateral branch network as the output of the current lateral branch network

[0049]

[0050] where r is the downsampling multiple, which is implemented through the LoRA matrix, and the specific multiple is consistent with the reduced hidden layer dimension of the lateral branch network, and α is the weighting coefficient;

[0051] Step 4-3: Stack the lateral branch networks in sequence in multiple network layers specified for lateral fine-tuning according to step 4-2. In the last layer, the output result of the lateral branch network After r-fold upsampling, it is weighted and fused with the output of the backbone network to output the final result F o :

[0052]

[0053] Step 5: Input the damage assessment training samples and the corresponding Q&A text annotations into the image encoder and the text encoder of the Transformer architecture respectively. In the image-text fusion encoder, the feature output of the image encoder will interact fully with the features of the text encoder to gradually align the image features with the features of the question text, and under the guidance and optimization of the lateral fine-tuning network, output a high-level unified representation of the image features and the question text features;

[0054] Step 6: Input the output result of the image-text fusion encoder in Step 5 into the reply decoder of the Transformer architecture. Based on the image-text alignment unified representation and the initialized reply representation, through full fusion of self-attention and cross-attention, under the guidance of the lateral fine-tuning network, output the category of the damage assessment sample, the overall damage degree, and the fine-grained damage degree of each part;

[0055] Step 7: According to the prior knowledge of experts, establish a damage metric for the target group. Taking armored vehicles in a vehicle as an example, the damage rate α ∈ (0, 20%) means no damage, the damage rate α ∈ [20%, 50%) means mild damage, the damage rate α ∈ [50%, 80%] means moderate damage, and the damage rate α ∈ [80%, 100%] means severe damage;

[0056] Step 9: Fine-tune the DAMAGE-VQA fine-grained target damage assessment network described in Steps 3 - 6. The optimizer is selected as AdamW, and the loss function uses the image-text contrast loss L QIC and the cross-entropy loss L RCE of the reply result. Promote the full alignment of image-text feature information during the fine-tuning process and constrain the reply decoder to generate the desired fine-grained target damage result;

[0057] Reference Figure 4 In, the gradient propagation and loss function setting in the fine-tuning process of Step 9 specifically include the following steps:

[0058] Step 9-1: During the forward propagation of fine-tuning, the quantized backbone network and the lateral branch network jointly participate in generating the Q&A reply result.

[0059] Step 9-2: During the process of optimizing by gradient backpropagation, the quantized backbone network only propagates the gradient but does not update, and the lateral branch network not only propagates the gradient but also is guided by the backpropagation to optimize and update the parameters.

[0060] Step 9-3: Image-Text Contrastive Loss L LIC Through contrastive learning, the similarity between the question text features and the image features is maximized to achieve a high degree of alignment of the image-text features. The contrastive loss is implemented based on the principle of maximizing information entropy as follows:

[0061]

[0062] Among them, sim(I,Q) calculates the similarity between the output features I of the image encoder and the output features Q of the question text encoder, where I ′ is the negative sample, and τ is the scale adjustment coefficient of the loss function

[0063] Step 9-4: Reply Result Cross-Entropy Loss L RCE It is used to compare the gap between the final result output by the reply encoder and the labeled reply:

[0064]

[0065] Among them, y i and are the output of the model's reply result and the labeled reply result, respectively.

[0066] Reference Figure 5 , the method of the present invention can make full use of the performance and technical advantages of the pre-trained image-text large model, accurately detect the target category, give the overall damage assessment of the target, and make accurate and fine-grained damage level assessments for each part of the target, which is consistent with the expert conclusion. The effect of the method of the present invention is further illustrated by the following simulation experiments.

[0067] 1. Simulation conditions.

[0068] The method of the present invention is simulated based on the Pytorch framework and Anaconda software on a central processing unit of Intel Core i7-9750H CPU, 32G of memory, 4 Nvidia RTX4090 graphics cards, and Ubuntu 18.04 operating system.

[0069] 2. Simulation content.

[0070] The data used in the simulation are 3 randomly collected target images of different categories, denoted as vehicle, tank, and aircraft, respectively.

[0071] To prove the effectiveness of the inventive method, the above 3 target images are processed by the method of the present invention and compared with the expert judgment conclusion. The comparison results are shown in Table 1.

[0072] As can be seen from Table 1, the damage assessment results of the present invention are consistent with those of experts. Applying the present invention to the field will enhance the perception and understanding ability of commanders and soldiers regarding damage scenarios.

[0073] Table 1

[0074]

[0075] It should be noted that, on the other hand, the present application also provides a storage medium. This storage medium can be included in an electronic device; or it can exist independently without being assembled into the electronic device. The above storage medium carries one or more programs. When the above one or more programs are executed by an electronic device, the electronic device realizes the methods described in the following embodiments. For example, the electronic device can realize each step of the method as Figure 1 shown.

[0076] In one embodiment, the present application provides a computer program product, including a computer program that realizes the steps in the above method embodiments when executed by a processor.

[0077] In addition, the above drawings are only schematic illustrations of the processes included in the method according to the exemplary embodiments of the present invention, rather than for limiting purposes. It is easy to understand that the processes shown in the above drawings do not indicate or limit the time sequence of these processes. Additionally, it is also easy to understand that these processes can be executed synchronously or asynchronously, for example, in multiple modules.

[0078] Those skilled in the art will easily think of other embodiments of the present invention after considering the specification and practicing the invention here. The present application aims to cover any variations, uses, or adaptations of the present invention that follow the general principles of the present invention and include common general knowledge or conventional technical means in the technical field not disclosed in the present invention. The specification and embodiments are only regarded as exemplary, and the true scope and spirit of the present invention are pointed out by the claims.

[0079] It should be understood that the present invention is not limited to the exact structure already described and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present invention is only defined by the appended claims.

Claims

1. A target damage assessment method based on large model quantitative lateral fine-tuning, characterized in that: The method comprises: Obtain the damaged target image, perform question-answering text annotation on the damaged target image to construct training samples; Construct a DAMAGE-VQA fine-grained target damage assessment network, which includes a question text encoder, an image encoder, a text-image fusion encoder, and a reply decoder. The layers of the text-image fusion encoder and the reply decoder are quantized. Lateral branch networks are added layer by layer to the quantized text-image fusion encoder and reply decoder, and the corresponding backbone network layers are dequantized. The parameters are reduced by gradient-independent downsampling and the branch networks are aligned. The backbone and branch networks are weightedly fused to form a lateral fine-tuning network. The final output of the lateral network is weightedly fused with the backbone network after upsampling, and the backbone network layer maintains the quantization state unchanged. The training samples are input into the DAMAGE-VQA fine-grained target damage assessment network for training. The training samples and the corresponding question-answer text annotations are respectively input into the image encoder and question text encoder of the Transformer architecture to extract features. The features of the image encoder and the features of the question text encoder are fully interacted in the image-text fusion encoder to gradually realize the alignment of image features with question text features. Under the guidance and optimization of the lateral fine-tuning network, a high-level unified representation of image features and question text features is output. The output result of the image-text fusion encoder is input into the reply decoder. The unified representation based on the image-text alignment and the initialized reply representation are fully fused through self-attention and mutual attention. Under the guidance of the lateral fine-tuning network, the final decoded output damage assessment sample category, overall damage degree and fine-grained damage degree of each part are output.

2. The target damage assessment method based on large model quantitative lateral fine-tuning according to claim 1 is characterized in that: The method further comprises performing data augmentation on the training samples, wherein the data augmentation comprises cropping, flipping and shifting.

3. The target damage assessment method based on large model quantitative lateral fine-tuning according to claim 1 or 2 is characterized in that: The quantization of each layer of the image-text fusion encoder and the part of the reply decoder specifically includes: quantizing all parameters of the selected network layer to 4-bit.

4. The target damage assessment method based on large model quantitative lateral fine-tuning according to claim 1 or 2 is characterized in that: The adding of the lateral branch network comprises: Copy the structure of the backbone network layer as the lightweight lateral branch network body; The output of the original quantized backbone network layer Dequantize back to 16-bit After downsampling, it is compared with the output of the previous lateral branch network Perform weighted fusion as the output of the current lateral branch network The lateral branch network is sequentially stacked in multiple network layers designated for lateral fine-tuning. In the last layer, the output of the lateral branch network is After upsampling, the final result is weightedly fused with the output of the backbone network.

5. The target damage assessment method based on large model quantitative lateral fine-tuning according to claim 1 is characterized in that: The method further comprises: The DAMAGE-VQA fine-grained target damage assessment network is fine-tuned, the optimizer uses AdamW, and the loss function uses the image-text contrast loss L QIC And the response result cross entropy loss L RCE ; Promote the full alignment of image and text feature information during the fine-tuning process and constrain the reply decoder to generate the desired fine-grained target damage results.

6. The target damage assessment method based on large model quantitative lateral fine-tuning according to claim 4 is characterized in that: The image-text contrast loss L QIC Through contrast learning, we ensure that the similarity between the question text features and the image features is maximized, and achieve high matching alignment of image and text features. We use contrast loss to achieve this based on the principle of maximizing information: Among them, sim(I,Q) calculates the similarity between the output feature I of the image encoder and the output feature Q of the question text encoder, I ′ is a negative sample, and τ is the scale adjustment coefficient of the loss function.

7. The target damage assessment method based on large model quantitative lateral fine-tuning according to claim 4 is characterized in that: The response result cross entropy loss L RCE Used to compare the gap between the final result output by the reply encoder and the label reply: Among them, y i and They are the model’s response output and label response results respectively.

8. A storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the target damage assessment method based on large model quantitative lateral fine-tuning is implemented as described in any one of claims 1 to 7.

9. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the target damage assessment method based on large model quantitative lateral fine-tuning according to any one of claims 1 to 7 is implemented.

10. An electronic device, characterized in that: include: processor; as well as A memory, configured to store executable instructions of the processor; The processor is configured to execute the target damage assessment method based on large model quantitative lateral fine-tuning according to any one of claims 1 to 7 by executing the executable instructions.