Trusted multi-source fusion target identity recognition method based on fine-grained semantic guidance

By using a fine-grained semantic guidance method, the degree of opinion conflict and confidence in multi-source fusion target identity recognition are quantified, which solves the problem of performance limitation of multi-source fusion target identity recognition methods in fine-grained recognition and achieves more reliable and interpretable target recognition results.

CN121479713BActive Publication Date: 2026-03-20TSINGHUA UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202610025703.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-01-09
Publication Date
2026-03-20
Estimated Expiration
2046-01-09

AI Technical Summary

Technical Problem

Existing multi-source fusion target identification methods have limited performance in fine-grained identification, cannot provide reliable target identification, and are difficult to quantify the uncertainty of decision-making ability.

Method used

A credible multi-source fusion target identity recognition method based on fine-grained semantic guidance is adopted. By using a target feature extraction network, a target evidence neural network, and a target fine-grained feature aggregation network, the degree of opinion conflict in modal data is quantified, and confidence and unconfidence are calculated based on Dirichlet distribution parameters to optimize the conflict decision and fusion process of multi-source criteria.

Benefits of technology

It improves the reliability of multi-source fusion target identification and the interpretability of identification results. By quantifying the degree of conflict and confidence of viewpoints, it optimizes the conflict decision of multi-source criteria and improves the reliability of identification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121479713B_ABST
    Figure CN121479713B_ABST
Patent Text Reader

Abstract

The application provides a kind of based on fine-grained semantic guidance trusted multi-source fusion target identity recognition method, it is related to target identification technical field, the method includes: the multiple first modal data corresponding to the target to be identified is input into each modal corresponding target feature extraction network and obtains each first modal feature;Each first modal feature is input into each modal corresponding target evidence neural network and target fine-grained feature aggregation network respectively, to obtain first evidence vector and first fine-grained feature vector;Based on each first evidence vector, generate the view of each first modal data, determine the first fusion view based on the view of each first modal data, based on each first fine-grained feature vector, the view conflict degree between the view of each first modal data in first fusion view is quantified;Based on each first evidence vector, determine first fusion evidence, determine the target identity corresponding to the target to be identified and confidence degree.The application improves the reliability of multi-source fusion target identification.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of target recognition, and in particular to a trusted multi-source fusion target identity recognition method based on fine-grained semantic guidance. BACKGROUND

[0002] Due to differences in imaging principles, imaging conditions, and sensor performance parameters, target data obtained by multiple detection platforms differ in modal characteristics and modal quality. How to utilize the consistency and complementary information of multi-source detection data to achieve target identity recognition is a problem that needs to be solved urgently.

[0003] Common multi-source fusion target identity recognition methods generally utilize deep learning architectures such as convolution modules, Transformer modules, attention mechanisms, and knowledge distillation to perform feature-level or decision-level fusion on different modal data, thereby determining target identity based on fused features. However, multi-class target sample features often exhibit large intra-class differences and small inter-class differences, which limits the performance of ordinary target recognition methods in fine-grained identity recognition and makes it difficult to provide reliable target recognition. At the same time, most target recognition solutions can only provide probability distribution predictions for target identity, making it difficult to quantify the uncertainty of different decision-making abilities and providing reliable target recognition. SUMMARY

[0004] To address the deficiencies in the prior art, the present application provides a trusted multi-source fusion target identity recognition method based on fine-grained semantic guidance, which improves the reliability of multi-source fusion target identity recognition.

[0005] In a first aspect, the present application provides a trusted multi-source fusion target identity recognition method based on fine-grained semantic guidance, which comprises the following steps:

[0006] Inputting multiple first modal data corresponding to a target to be recognized into a target feature extraction network corresponding to each modality to obtain first modal features of each of the first modal data; the multiple first modal data are multi-modal data detected by different platforms;

[0007] Inputting each of the first modal features into a target evidence neural network corresponding to each modality to obtain a first evidence vector in which each of the first modal features is assigned to at least two preset class identities, and inputting each of the first modal features into a target fine-grained feature aggregation network corresponding to each modality to obtain a first fine-grained feature vector in which each of the first modal features is assigned to each of the preset class identities;

[0008] generate a view of each of the first modal data based on each of the first evidence vectors, determine a first fusion view based on the view of each of the first modal data, and quantify a view conflict degree between the view of each of the first modal data in the first fusion view based on each of the first fine-grained feature vectors; the quantification of the view conflict degree is used to constrain the consistency of evidence and view of different modalities in the model training process of the target evidence neural network, the target feature extraction network, and the target fine-grained feature aggregation network;

[0009] determine a first fusion evidence based on each of the first evidence vectors, determine the target identity corresponding to the to-be-identified target based on the first fusion evidence, and determine a confidence degree corresponding to the target identity based on the first fusion view.

[0010] According to the fine-grained semantic guidance-based trusted multi-source fusion target identity recognition method provided in the application, the view of each of the first modal data includes a prior probability, a confidence degree, and an untrustworthiness degree of each of the first modal data being assigned to each of the preset category identities;

[0011] The generating of the view of each of the first modal data based on each of the first evidence vectors and the determining of the first fusion view based on the view of each of the first modal data include:

[0012] Based on each of the first evidence vectors, the Dirichlet distribution parameters corresponding to each of the modalities are calculated, and based on the Dirichlet distribution parameters corresponding to each of the modalities, the Dirichlet strengths corresponding to each of the modalities are calculated.

[0013] Based on the number of the preset category identities, the prior probability of each of the first modal data being assigned to each of the preset category identities is determined.

[0014] Based on each of the first evidence vectors and the Dirichlet strengths corresponding to each of the modalities, the confidence degree of each of the first modal data being assigned to each of the preset category identities is determined.

[0015] Based on the Dirichlet strengths corresponding to each of the modalities, the untrustworthiness degree of each of the first modal data being assigned to each of the preset category identities is determined.

[0016] The prior probability of each of the first modal data being assigned to each of the preset category identities, the confidence degree of each of the first modal data being assigned to each of the preset category identities, and the untrustworthiness degree of each of the first modal data being assigned to each of the preset category identities are fused to obtain the first fusion view.

[0017] According to the application, a target identity recognition method based on fine-grained semantic guidance is provided, which quantizes the view conflict degree between views of each first modal data in the first fusion view based on each first fine-grained feature vector, and includes the following steps:

[0018] For any one of the view pairs of each view of the first modal data, the corresponding fine-grained semantic weighted projection distance between the first view and the second view is determined based on the correlation between the modal fine-grained distribution weight vector of the first view and the modal fine-grained distribution weight vector of the second view in the view pair, the Dirichlet distribution parameter corresponding to each modal, and the Dirichlet intensity corresponding to each modal; the modal fine-grained distribution weight vector of the first view and the modal fine-grained distribution weight vector of the second view in the view pair are determined based on each first fine-grained feature vector;

[0019] Based on the untrustworthiness of the first modal data in the first view being assigned to each of the preset category identities and the untrustworthiness of the first modal data in the second view being assigned to each of the preset category identities, the corresponding fusion certainty between the first view and the second view is determined.

[0020] The product of the projection distance and the fusion certainty is determined as the conflict degree of the view pair.

[0021] Based on the conflict degree of each view pair, the conflict degree between views of each first modal data in the first fusion view is determined.

[0022] According to the application, a target identity recognition method based on fine-grained semantic guidance is provided, which determines a first fusion evidence based on each first evidence vector, and determines the target identity corresponding to the target to be recognized based on the first fusion evidence, and includes the following steps:

[0023] The mean of each first evidence vector is calculated, and the first fusion evidence is determined based on the mean;

[0024] The Dirichlet distribution parameter corresponding to the first fusion modal and the Dirichlet intensity corresponding to the first fusion modal are calculated according to the first fusion evidence;

[0025] Based on the Dirichlet distribution parameter corresponding to the first fusion modal and the Dirichlet intensity corresponding to the first fusion modal, the probability distribution of the target to be recognized being each preset category identity is calculated.

[0026] Based on the maximum probability value in the probability distribution, the target identity is determined.

[0027] According to the application, a target identity recognition method based on fine-grained semantic guidance is provided.

[0028] A plurality of second modal data corresponding to a plurality of sample targets is obtained, and each second modal data is input into an initial feature extraction network corresponding to each modal to obtain a second modal feature of each second modal data.

[0029] Each second modal feature is input into an initial evidence neural network corresponding to each modal to obtain a second evidence vector in which each second modal feature is assigned to each preset class identity, and each second modal feature is input into an initial fine-grained feature aggregation network corresponding to each modal to obtain a second fine-grained feature vector in which each second modal feature is assigned to each preset class identity.

[0030] Based on each second evidence vector, a view of each second modal data is generated, a second fusion view is determined based on the views of each second modal data, and a view conflict degree and a decision confidence between the views of each second modal data in the second fusion view are quantified based on the second fine-grained feature vector.

[0031] Based on each second evidence vector, a second fusion evidence is determined, a predicted identity class corresponding to each sample target is determined based on the second fusion evidence, and a confidence corresponding to the predicted identity class corresponding to each sample target is determined based on the second fusion view.

[0032] Based on the evidence vector and the fusion evidence vector of each second modal data corresponding to each sample target, a Dirichlet distribution parameter predicted by each second modal data and a fusion predicted Dirichlet distribution parameter are determined, and a loss value of a first loss function is determined in combination with a real identity class distribution corresponding to each sample target.

[0033] Based on the view conflict degree between the views of each second modal data, a loss value of a second loss function is determined.

[0034] Based on the loss value of the first loss function and the loss value of the second loss function, model parameters corresponding to each of the initial feature extraction network corresponding to each modal, the initial evidence neural network corresponding to each modal, and the initial fine-grained feature aggregation network corresponding to each modal are adjusted to obtain an adjusted model.

[0035] The training of the adjusted model is continued until a training stop condition is reached, and the target feature extraction network, the target evidence neural network, and the target fine-grained feature aggregation network are obtained.

[0036] According to the application, a target identity recognition method based on fine-grained semantic guidance and trusted multi-source fusion is provided. The loss value of the second loss function is determined according to the viewpoint conflict degree between the viewpoints of the second modal data, and the method comprises the following steps:

[0037] The loss value of the second loss function is calculated by using the following formula (1):

[0038] (1)

[0039] Wherein, represents the loss value of the second loss function, represents the number of preset category identities, represents the first modal, represents the second modal, represents the third modal, represents the fourth modal, represents the fifth modal, represents the sixth modal, and represents the weighted projection distance between the fine-grained semantics of and represents the first target sample feature of the first modal, represents the second target sample feature of the first modal, represents the third target sample feature of the first modal, represents the fourth target sample feature of the first modal, represents the fifth target sample feature of the first modal, represents the sixth target sample feature of the first modal, represents the viewpoint modal fine-grained distribution weight vector of the first target sample feature of the first modal, represents the viewpoint modal fine-grained distribution weight vector of the second target sample feature of the first modal, represents the viewpoint modal fine-grained distribution weight vector of the third target sample feature of the first modal, represents the viewpoint modal fine-grained distribution weight vector of the fourth target sample feature of the first modal, represents the viewpoint modal fine-grained distribution weight vector of the fifth target sample feature of the first modal, and represents the viewpoint modal fine-grained distribution weight vector of the sixth target sample feature of the first modal.

[0040] According to the application, a target identity recognition method based on fine-grained semantic guidance and trusted multi-source fusion is provided. The plurality of first modal data comprises at least one of the following: visible light image, synthetic aperture radar image, near-infrared image, panchromatic image, laser radar point cloud data, voice description and text description.

[0041] According to the application, a target identity recognition method based on fine-grained semantic guidance and trusted multi-source fusion is provided. The type of the target feature extraction network is a deep learning model or an encoder, and the encoder is obtained based on multi-modal large model training. The type of the target evidence neural network is the deep learning model. The type of the target fine-grained feature aggregation network is the deep learning model.

[0042] In a second aspect, the application further provides a device for target identity recognition based on fine-grained semantic guidance and trusted multi-source fusion. The device comprises the following modules:

[0043] a feature extraction module, configured to input a plurality of first modality data corresponding to a target to be identified into a target feature extraction network corresponding to each modality to obtain first modality features of each of the first modality data, wherein the plurality of first modality data are multi-modality data detected by different platforms;

[0044] an identity recognition module, configured to input each of the first modality features into a target evidence neural network corresponding to each modality to obtain a first evidence vector in which each of the first modality features is assigned to at least two preset class identities, and input each of the first modality features into a target fine-grained feature aggregation network corresponding to each modality to obtain a first fine-grained feature vector in which each of the first modality features is assigned to each of the preset class identities;

[0045] generate a viewpoint of each of the first modality data based on each of the first evidence vectors, determine a first fusion viewpoint based on the viewpoints of the first modality data, and quantify a viewpoint conflict degree between the viewpoints of the first modality data in the first fusion viewpoint based on each of the first fine-grained feature vectors, wherein the quantification of the viewpoint conflict degree is used to constrain the consistency of evidences and viewpoints of different modalities in a model training process of the target evidence neural network, the target feature extraction network, and the fine-grained feature aggregation network;

[0046] determine a first fusion evidence based on each of the first evidence vectors, determine a target identity corresponding to the target to be identified based on the first fusion evidence, and determine a confidence degree corresponding to the target identity based on the first fusion viewpoint.

[0047] In a third aspect, the present application further provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the target identity recognition method based on fine-grained semantic guidance and trusted multi-source fusion according to any one of the above aspects when executing the computer program.

[0048] In a fourth aspect, the present application further provides a non-transitory computer readable storage medium having a computer program stored thereon, wherein the computer program is executable by a processor to implement the target identity recognition method based on fine-grained semantic guidance and trusted multi-source fusion according to any one of the above aspects.

[0049] In a fifth aspect, the present application further provides a computer program product comprising a computer program, wherein the computer program is executable by a processor to implement the target identity recognition method based on fine-grained semantic guidance and trusted multi-source fusion according to any one of the above aspects.

[0050] The application provides a trusted multi-source fusion target identity recognition method based on fine-grained semantic guidance.

[0051] The first modal feature of each modal data is input into the target fine-grained feature aggregation network and the target evidence neural network, respectively, to obtain the first evidence vector in which each first modal feature is assigned to at least two preset category identities and the first fine-grained feature vector in which each first modal feature is assigned to each preset category identity, and then the view of each first modal data is generated based on the first evidence vector, the first fusion view is determined based on the view of each first modal data, and the view conflict degree between the views in the first fusion view is quantified based on the first fine-grained feature vector, and the quantification of the view conflict degree is used to constrain the consistency of the evidence and the view of different modalities in the model training process of the target evidence neural network, the target feature extraction network and the target fine-grained feature aggregation network, and finally the target identity corresponding to the target to be identified is determined based on the first fusion evidence determined based on the first evidence vector, and the corresponding confidence. BRIEF DESCRIPTION OF DRAWINGS

[0052] In order to more clearly illustrate the technical solutions in the application or the prior art, the following will briefly introduce the drawings needed to be used in the embodiments or the prior art description. Obviously, the drawings in the following description are some embodiments of the application, and other drawings can also be obtained by those skilled in the art without any creative effort on the basis of these drawings.

[0053] Figure 1 is a flowchart of the method for target identity recognition based on fine-grained semantic guidance and trusted multi-source fusion provided by the present application.

[0054] Figure 2 is a flowchart of the model training method provided by the present application.

[0055] Figure 3 is a principle diagram of the method for target identity recognition based on fine-grained semantic guidance and trusted multi-source fusion provided by the present application.

[0056] Figure 4 is a structural diagram of the device for target identity recognition based on fine-grained semantic guidance and trusted multi-source fusion provided by the present application.

[0057] Figure 5 is a structural diagram of the electronic device provided by the present application. DETAILED DESCRIPTION

[0058] In order to make the objectives, technical solutions and advantages of the present application clearer, the technical solutions in the present application will be described clearly and completely below in conjunction with the drawings in the present application. Obviously, the described embodiments are some of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative labor fall within the scope of protection of the present application.

[0059] The method for target identity recognition based on fine-grained semantic guidance and trusted multi-source fusion provided by the present application will be described below in conjunction with Figures 1-5

[0060] Figure 1 is a flowchart of the method for target identity recognition based on fine-grained semantic guidance and trusted multi-source fusion provided by the present application, as shown in Figure 1 , the method comprises the following steps:

[0061] Step 101, inputting the plurality of first modal data corresponding to the to-be-identified target into the target feature extraction network corresponding to each modal to obtain the first modal feature of each first modal data; the plurality of first modal data are multi-modal data detected by different platforms.

[0062] The execution subject of the present embodiment is an electronic device, and a method for target identity recognition based on fine-grained semantic guidance and trusted multi-source fusion is provided, which is used to realize reliable multi-source target identity recognition.

[0063] ​First, acquire multiple first-modal data corresponding to the target to be identified. For example, acquire multimodal data corresponding to the target to be identified from various detection platforms. Multimodal data includes visible light images, synthetic aperture radar (SAR) images, near-infrared images, panchromatic images, lidar point cloud data, voice, text descriptions, and other multimodal data.

[0064] Next, the multiple first modal data corresponding to the target to be identified are input into the target feature extraction network corresponding to the modality to obtain the first modal features of each first modal data. Among them, the target feature extraction network is a trained deep learning model, such as a convolutional neural network.

[0065] Step 102: Input each first modality feature into the target evidence neural network corresponding to each modality to obtain the first evidence vector in which each first modality feature is assigned to at least two preset category identities, and input each first modality feature into the target fine-grained feature aggregation network corresponding to each modality to obtain the first fine-grained feature vector in which each first modality feature is assigned to each preset category identity.

[0066] Specifically, the target evidence neural network and the target fine-grained feature aggregation network are ordinary deep learning models trained from these networks.

[0067] To reduce training costs, the first modality features of the multimodal data are input into the corresponding evidence neural network to obtain a first evidence vector in which each first modality feature is assigned to at least two preset class identities. For example, it can be represented as follows: That is, to obtain the first Modal The modal feature is assigned to the first Evidence of pre-defined identity .

[0068] The first modality features of the multimodal data are input into the corresponding target fine-grained feature aggregation network to obtain the features of each modality. Fine-grained feature vectors The feature dimension of each type of feature vector is d, which means that the feature dimension of the first type is d. The first mode The modal feature corresponding to the first Fine-grained feature vectors .

[0069] Step 103, generating the views of each first modal data based on each first evidence vector, determining a first fusion view based on the views of each first modal data, and quantifying the view conflict degree between the views of each first modal data in the first fusion view based on each first fine-grained feature vector; the quantification of the view conflict degree is used to constrain the consistency of the evidence and the view of different modalities in the model training process of the target evidence neural network, the target feature extraction network and the target fine-grained feature aggregation network.

[0070] Specifically, the views of each first modal data are generated based on the generated first modal features being assigned to the first evidence vector in at least two preset class identities.

[0071] Each view corresponding to each modal feature includes: the prior probability, the confidence and the untrustworthiness of each modal feature being assigned to The views of each first modal data are obtained, and then the first fusion view is generated based on the views of each first modal data.

[0072] Further, the view conflict degree between the views of each first modal data in the first fusion view is quantified based on each first fine-grained feature vector. For example, the view conflict degree between the views of each first modal data in the first fusion view is quantified based on each first fine-grained feature vector.

[0073] The quantification of the view conflict degree is used to constrain the consistency of the evidence and the view of different modalities in the model training process of the target evidence neural network, the target feature extraction network and the fine-grained feature aggregation network, improve the model training effect, and thus improve the reliability of the target identity recognition.

[0074] For example, the correlation of the inter-modal fine-grained salient feature vector (each first fine-grained feature vector) is used as a guide to dynamically adjust the multi-modal conflict distance weight between different categories, and the view conflict degree between each view is calculated.

[0075] Step 104, determining a first fusion evidence based on each first evidence vector, determining a target identity corresponding to the to-be-identified target based on the first fusion evidence, and determining a confidence degree corresponding to the target identity based on the first fusion view.

[0076] Specifically, the first fusion evidence is obtained by fusing each first evidence vector, and the probability distribution of the target identity corresponding to the to-be-identified target is calculated according to the obtained first fusion evidence, that is, the probability of the to-be-identified target being assigned to each preset class identity; then, the target identity corresponding to the to-be-identified target is determined based on the probability of the to-be-identified target being assigned to each preset class identity.

[0077] Correspondingly, the confidence corresponding to the target identity determined based on the first fusion viewpoint in the application, that is, the decision credibility, is used to provide good result explainability.

[0078] The method provided by the embodiment is first to input the plurality of first modal data corresponding to the to-be-identified target into the target feature extraction network corresponding to each modality to obtain the first modal features of each first modal data, and the plurality of first modal data are multi-modal data detected by different platforms. Then, each first modal feature is input into the target evidence neural network corresponding to each modality to obtain a first evidence vector in which each first modal feature is assigned to at least two preset class identities, and the first evidence vector is input into the target fine-grained feature aggregation network corresponding to each modality to obtain a first fine-grained feature vector in which each first modal feature is assigned to each preset class identity. Further, a viewpoint of each first modal data is generated based on each first evidence vector, a first fusion viewpoint is determined based on the viewpoints of each first modal data, and a viewpoint conflict degree between the viewpoints of each first modal data in the first fusion viewpoint is quantified based on each first fine-grained feature vector. The quantification of the viewpoint conflict degree is used to constrain the consistency of the evidence and the viewpoint of different modalities in the model training process of the target evidence neural network, the target feature extraction network, and the target fine-grained feature aggregation network. Furthermore, a first fusion evidence is determined based on each first evidence vector, a target identity corresponding to the to-be-identified target is determined based on the first fusion evidence, and a confidence corresponding to the target identity is determined based on the first fusion viewpoint.

[0079] The first modal features of the to-be-identified target are input into the target fine-grained feature aggregation network and the target evidence neural network respectively, the first evidence vector in which each first modal feature is assigned to at least two preset class identities and the first fine-grained feature vector in which each first modal feature is assigned to each preset class identity are obtained, further, the viewpoint of each first modal data is generated based on each first evidence vector and the first fusion viewpoint is determined based on the viewpoints of each first modal data, and the viewpoint conflict degree between the viewpoints in the first fusion viewpoint is quantified based on the first fine-grained feature vector. The quantification of the viewpoint conflict degree is used to constrain the consistency of the evidence and the viewpoint of different modalities in the model training process of the target evidence neural network, the target feature extraction network, and the fine-grained feature aggregation network. Finally, the target identity corresponding to the to-be-identified target is determined based on the first fusion evidence determined based on each first evidence vector, and the corresponding confidence is determined. The application fully utilizes the inter-class fine-grained feature correlation, optimizes the conflict decision and fusion process of the multi-source criterion, and improves the reliability of multi-source fusion target identification.

[0080] According to the method for identifying target identity based on fine-grained semantic guidance provided by the application, the viewpoint of each first modal data includes the prior probability, the confidence, and the untrustworthiness in which each first modal data is assigned to each preset class identity.

[0081] The viewpoints for generating each first modality of data based on each first evidence vector, and the first fused viewpoints determined based on the viewpoints for each first modality of data, include:

[0082] Based on each first evidence vector, calculate the Dirichlet distribution parameters corresponding to each mode, and based on the Dirichlet distribution parameters corresponding to each mode, calculate the Dirichlet intensity corresponding to each mode.

[0083] Based on the number of preset category identities, determine the prior probability that each first modality data is assigned to each preset category identity;

[0084] Based on each first evidence vector and the Dirichlet intensity corresponding to each modality, the confidence level of each first modality data being assigned to each preset category identity is determined;

[0085] Based on the Dirichlet intensity corresponding to each modality, the unreliability of each first modality data being assigned to each preset category identity is determined;

[0086] The prior probability of each first modality data being assigned to each preset category identity, the confidence level of each first modality data being assigned to each preset category identity, and the unconfidence level of each first modality data being assigned to each preset category identity are fused to obtain the first fused viewpoint.

[0087] Specifically, in some embodiments, the perspective of each first modality data includes the prior probability, confidence level, and unconfidence level of each first modality data being assigned to each preset category identity.

[0088] Step 103 includes the following steps:

[0089] Step 1: Calculate the viewpoints of each first mode data.

[0090] First, based on the first evidence vector (e.g., the first evidence vector that is assigned to at least two predefined class identities for each first modality feature) The first mode Evidence vector corresponding to each modality feature ), calculate the Dirichlet distribution parameters corresponding to each mode (e.g., the 1st mode), The first mode Dirichlet distribution parameters corresponding to each modal feature Based on the Dirichlet distribution parameters corresponding to each mode, the Dirichlet intensity corresponding to each mode is calculated (e.g., the first mode). The first mode (Dirichlet intensity corresponding to each modal feature), k represents the total number of preset category identities.

[0091] For example, the Dirichlet distribution parameters corresponding to each mode can be calculated using the following formula:

[0092]

[0093]

[0094] wherein, represents the first target sample feature of the first modality being assigned to the Dirichlet distribution parameter corresponding to the first preset class identity, represents the first target sample feature of the first modality being assigned to the Dirichlet distribution parameter corresponding to the first preset class identity, represents transposition, represents the first target sample feature of the first modality being assigned to the first evidence vector of the first preset class identity, represents the Dirichlet distribution parameter vector corresponding to the first target sample feature of the first modality, represents the Dirichlet distribution parameter corresponding to the first target sample feature of the first modality being assigned to the first preset class identity, represents the Dirichlet distribution parameter corresponding to the first target sample feature of the first modality being assigned to the first preset class identity.

[0095] For example, the Dirichlet intensity corresponding to each modality is calculated by the following formula:

[0096]

[0097] wherein, represents the Dirichlet intensity corresponding to the first target sample feature of the first modality, is the number of preset class identities, represents the Dirichlet distribution parameter corresponding to the first target sample feature of the first modality being assigned to the first preset class identity.

[0098] ​​​​​​​​​​​​​​​​​​​​​Then, based on the calculated Dirichlet distribution parameters corresponding to each modality and the Dirichlet intensity corresponding to each modality, the prior probability, the confidence and the unconfidence are calculated:

[0099] (1) Prior probability

[0100] Based on the number of preset category identities, the prior probability of each first modality data being assigned to each preset category identity is determined.

[0101] For example, the prior probability of each first modality data being assigned to each preset category identity is calculated by the following formula :

[0102]

[0103] wherein, represents the prior probability of the i-th target sample feature of the k-th modality being assigned to the j-th preset category identity, and k represents the total number of preset category identities.

[0104] (2) Confidence

[0105] Based on each first evidence vector and the Dirichlet intensity corresponding to each modality, the confidence of each first modality data being assigned to each preset category identity is determined.

[0106] For example, the confidence of each first modality data being assigned to each preset category identity is calculated by the following formula :

[0107]

[0108] wherein, represents the confidence of the i-th target sample feature of the k-th modality being assigned to the j-th preset category identity, represents the confidence of the i-th target sample feature of the k-th modality being assigned to the j-th preset category identity, represents the first evidence vector of the i-th target sample feature of the k-th modality being assigned to the j-th preset category identity, represents the Dirichlet intensity corresponding to the i-th target sample feature of the k-th modality.

[0109] (3) Unconfidence

[0110] Based on the Dirichlet intensity corresponding to each modality, the unconfidence of each first modality data being assigned to each preset category identity is determined. ​​​​​​​​​​

[0111] For example, the untrustworthiness of each first modal data being assigned to each preset category identity is calculated by the following formula :

[0112]

[0113]

[0114] wherein, represents the untrustworthiness of the i-th target sample feature of the j-th modal being assigned to the k-th preset category identity, represents the untrustworthiness of the i-th target sample feature of the j-th modal being assigned to the k-th preset category identity, represents the Dirichlet intensity corresponding to the i-th target sample feature of the j-th modal, represents the untrustworthiness of the i-th target sample identity being predicted by the j-th modal data, represents the untrustworthiness of the i-th target sample identity being predicted by the j-th modal data, represents the Dirichlet intensity corresponding to the i-th target sample feature of the j-th modal, represents the untrustworthiness of the i-th target sample identity being predicted by the j-th modal data, represents the untrustworthiness of the i-th target sample identity being predicted by the j-th modal data, represents the number of preset category identities. Thus, the viewpoint of the sample (i-th target sample feature) is represented as:

[0115] ;

[0116] ;

[0117] ;

[0118] .

[0119] wherein, represents the viewpoint corresponding to the i-th modal feature of the j-th modal, represents the prior probability corresponding to the i-th modal feature of the j-th modal, represents the confidence corresponding to the i-th modal feature of the j-th modal, represents the untrustworthiness corresponding to the i-th modal feature of the j-th modal, represents the prior probability of the i-th target sample feature of the j-th modal being assigned to the k-th preset category identity, represents the prior probability of the i-th target sample feature of the j-th modal being assigned to the k-th preset category identity, represents the confidence of the i-th target sample feature of the j-th modal being assigned to the k-th preset category identity, represents the untrustworthiness of the i-th target sample feature of the j-th modal being assigned to the k-th preset category identity, represents the prior probability of the i-th target sample feature of the j-th modal being assigned to the k-th preset category identity, represents the confidence of the i-th target sample feature of the j-th modal being assigned to the k-th preset category identity, represents the untrustworthiness of the i-th target sample feature of the j-th modal being assigned to the k-th preset category identity, represents the prior probability of the i-th target sample feature of the j-th modal being assigned to the k-th preset category identity, represents the confidence of the i-th target sample feature of the j-th modal being assigned to the k-th preset category identity, represents the untrustworthiness of the i-th target sample feature of the j-th modal being assigned to the k-th preset category identity, represents the prior probability of the i-th target sample feature of the j-th modal being assigned to the k-th preset category identity, represents the confidence of the i-th target sample feature of the j-th modal being assigned to the k-th preset category identity, represents the untrustworthiness of the i-th target sample feature of the j-th modal being assigned to the k-th preset category identity, represents the prior probability of the i-th target sample feature of the j-th modal being assigned to the k-th preset category identity, ​​A target sample feature is assigned to a prior probability in the first pre-set class identity, a target sample feature of the first modal is assigned to a prior probability in the first pre-set class identity, a target sample feature of the first modal is assigned to a confidence in the first pre-set class identity, a target sample feature of the first modal is assigned to a confidence in the first pre-set class identity, a target sample feature of the first modal is assigned to a confidence in the first pre-set class identity, a target sample feature of the first modal is assigned to a confidence in the first pre-set class identity, a target sample feature of the first

[0120] Step 2, view fusion

[0121] The prior probability of each first modal data being assigned to each pre-set class identity, the confidence of each first modal data being assigned to each pre-set class identity, and the untrustworthiness of each first modal data being assigned to each pre-set class identity are fused to obtain a first fusion view.

[0122] For example, the views of a sample (a target sample feature) in different modes are fused, including:

[0123]

[0124] , , respectively represent the views of the first modal, the second modal, and the mth modal for the first target sample feature,

[0125]

[0126] In the formula, is the fusion view of the first , m modes, represents the first modal, ​​​​​The perspective of the characteristics of a target sample Indicates the first The th mode The perspective of the characteristics of a target sample For the first The m-th mode The fusion result of the prior probability of each target sample feature for the i-th type of target identity For the first The m-th mode The features of the target sample in the th... The confidence fusion result of the target identity class. For the first The m-th mode The result of fusion of the unreliability of the features of the target sample in the i-th target identity.

[0127] The method provided in this embodiment generates viewpoints for each modality of data based on evidence, fuses them to obtain fused viewpoints of different modalities, which facilitates the subsequent generation of confidence scores for the target identity corresponding to the target to be identified based on the fused viewpoints of different modalities, thereby improving the interpretability of the identification results.

[0128] According to the present invention, a reliable multi-source fusion target identity recognition method based on fine-grained semantic guidance quantifies the degree of viewpoint conflict between viewpoints of each first modality data in a first fused viewpoint based on each first fine-grained feature vector, including:

[0129] For any pair of viewpoints in each first modality of data, the fine-grained semantic weighted projection distance between the first and second viewpoints is determined based on the correlation between the modal fine-grained distribution weight vectors of the first and second viewpoints, the Dirichlet distribution parameters corresponding to each modality, and the Dirichlet intensity corresponding to each modality. The modal fine-grained distribution weight vectors of the first and second viewpoints in the viewpoint pair are determined based on each first fine-grained feature vector.

[0130] Based on the unreliability of the first modal data in the first viewpoint being assigned to each preset category identity, and the unreliability of the first modal data in the second viewpoint being assigned to each preset category identity, the corresponding fusion certainty between the first viewpoint and the second viewpoint is determined;

[0131] The product of the projection distance and the fusion determinism is used to determine the degree of conflict between viewpoint pairs.

[0132] Based on the degree of conflict between each pair of viewpoints, the degree of conflict between the viewpoints of each first modality data in the first fused viewpoint is determined.

[0133] Specifically, in some embodiments, the quantifying the degree of conflict between the viewpoints in step 103 is achieved by the following steps:

[0134] When quantifying the degree of conflict between the viewpoints in the multi-modal data, the inter-modal fine-grained salient feature vector correlation is used as a guide to dynamically adjust the multi-modal conflict distance weight between different categories.

[0135] For any pair of viewpoints in the viewpoints of the first modal data, the modal fine-grained distribution weight vector of the first viewpoint and the modal fine-grained distribution weight vector of the second viewpoint in the viewpoint pair are introduced when calculating the conflict degree of the viewpoint pair, for example, represented as: the modal fine-grained distribution weight vector of the first viewpoint , the modal fine-grained distribution weight vector of the second viewpoint . Wherein, the modal fine-grained distribution weight vector of the first viewpoint and the modal fine-grained distribution weight vector of the second viewpoint are determined based on each of the first fine-grained feature vectors.

[0136] The specific calculation steps include:

[0137] (1) Based on the correlation between the modal fine-grained distribution weight vector of the first viewpoint and the modal fine-grained distribution weight vector of the second viewpoint in the viewpoint pair, the Dirichlet distribution parameter corresponding to each modal, and the Dirichlet intensity corresponding to each modal, the corresponding fine-grained semantic weighted projection distance between the first viewpoint and the second viewpoint is determined.

[0138] For example, the corresponding fine-grained semantic weighted projection distance between the first viewpoint and the second viewpoint is calculated by the following formula:

[0139]

[0140]

[0141]

[0142]

[0143] Wherein, represents the corresponding fine-grained semantic weighted projection distance between the first viewpoint and the second viewpoint, represents the number of preset category identities, represents the Dirichlet distribution parameter corresponding to the first target sample feature of the first modal is assigned to the Dirichlet distribution parameter corresponding to the first preset category identity, represents the Dirichlet intensity corresponding to the first target sample feature of the first modal, Indicates the first The first mode The features of the target sample are assigned to the first... Dirichlet distribution parameters corresponding to each preset category identity. Indicates the first The first mode Dirichlet intensity corresponding to each target sample feature This represents the correlation between the modal fine-grained distribution weight vector of the first viewpoint and the modal fine-grained distribution weight vector of the second viewpoint. This represents the modal fine-grained distribution weight vector representing the first viewpoint. This represents the modal fine-grained distribution weight vector representing the second viewpoint. This indicates the number of feature dimensions in a fine-grained feature vector.

[0144] (2) Based on the unreliability of the first modal data in the first viewpoint being assigned to each preset category identity, and the unreliability of the first modal data in the second viewpoint being assigned to each preset category identity, determine the corresponding fusion certainty between the first viewpoint and the second viewpoint.

[0145] For example, the fusion determinism between the first and second viewpoints can be calculated using the following formula:

[0146]

[0147] in, Expressing the first opinion Second viewpoint The corresponding fusion certainty between them Indicates the first The first mode The unreliability of a target sample feature. Indicates the first The first mode The unreliability of the features of a target sample.

[0148] (3) The product of projection distance and fusion determinism is determined as the degree of conflict between viewpoint pairs.

[0149]

[0150] in, Indicates the degree of conflict between opposing viewpoints. This represents the fine-grained semantically weighted projected distance between the first and second viewpoints. This indicates the certainty of the integration between the first and second viewpoints.

[0151] Further, based on the conflict degree of each view pair, the conflict degree between views of each first modal data in the first fused view is determined.

[0152] In the method provided by the embodiment, when the evidences provided by different modal characteristics conflict, inter-class modal semantic feature correlation is introduced, the projection distance of different evidence views is weighted by the inter-class semantic feature correlation degree, the evidence decision conflict between multi-modal data is refined, the reliability of multi-platform fusion decision is improved, and the reliability of target identity recognition is improved.

[0153] According to the method for target identity recognition provided by the application, the first fused evidence is determined based on each first evidence vector, and the target identity corresponding to the target to be identified is determined based on the first fused evidence.

[0154] The mean of each first evidence vector is calculated, and the first fused evidence is determined based on the mean;

[0155] The Dirichlet distribution parameter corresponding to the first fused modal and the Dirichlet intensity corresponding to the first fused modal are calculated according to the first fused evidence;

[0156] The probability distribution of the target to be identified being each preset class identity is calculated based on the Dirichlet distribution parameter corresponding to the first fused modal and the Dirichlet intensity corresponding to the first fused modal;

[0157] The target identity is determined based on the maximum probability value in the probability distribution.

[0158] Specifically, in some embodiments, step 104 is implemented by the following method:

[0159] First, the mean of each first evidence vector is calculated, and the mean is determined as the first fused evidence . For example, the evidence of each modal sample feature being assigned to the i-th preset class identity is fused , and the first fused evidence is obtained based on the evidence of each modal sample feature being assigned to the i-th preset class identity .

[0160] Then, the Dirichlet distribution parameter corresponding to the first fused modal and the Dirichlet intensity corresponding to the first fused modal are calculated according to the first fused evidence .

[0161] For example, the Dirichlet distribution parameter corresponding to the first fused modal is calculated by the following formula :

[0162]

[0163] wherein, denotes the Dirichlet distribution parameter corresponding to the first fusion modality of the i-th preset category identity, denotes the first fusion evidence of the i-th preset category identity.

[0164] The Dirichlet distribution parameter corresponding to the first fusion modality is calculated by the following formula :

[0165]

[0166] wherein, denotes the Dirichlet distribution parameter corresponding to the first fusion modality, denotes the number of preset category identities, denotes the Dirichlet distribution parameter corresponding to the first fusion modality of the i-th preset category identity.

[0167] Further, based on the Dirichlet distribution parameter corresponding to the first fusion modality and the Dirichlet intensity corresponding to the first fusion modality, the probability distribution of the to-be-identified target being each preset category identity is calculated.

[0168] For example, the probability distribution of the to-be-identified target being each preset category identity is calculated by the following formula:

[0169]

[0170] wherein, is the probability of the to-be-identified target being the i-th preset category identity, denotes the Dirichlet distribution parameter corresponding to the first fusion modality of the i-th preset category identity, is the Dirichlet intensity corresponding to the first fusion modality.

[0171] Further, based on the maximum probability value in the probability distribution, the target identity is determined. For example, the preset category identity corresponding to the maximum probability value in the probability distribution is determined as the target identity of the to-be-identified target.

[0172] The method provided by the embodiment first calculates the mean of each first evidence vector, determines the first fusion evidence based on the mean, calculates the Dirichlet distribution parameter corresponding to the first fusion modality and the Dirichlet intensity corresponding to the first fusion modality according to the first fusion evidence, then calculates the probability distribution of the to-be-identified target being each preset category identity based on the Dirichlet distribution parameter corresponding to the first fusion modality and the Dirichlet intensity corresponding to the first fusion modality, and further determines the target identity based on the maximum probability value in the probability distribution. The embodiment strengthens the multi-modal evidence fusion through mean fusion and probability distribution modeling, and improves the credibility and applicability of the identity recognition method.

[0173] The application provides a fine-grained semantic guidance-based trusted multi-source fusion target identity recognition method, and training steps of a target feature extraction network, a target evidence neural network and a target fine-grained feature aggregation network include:

[0174] Obtaining a plurality of second modal data corresponding to a plurality of sample targets, and inputting each second modal data into an initial feature extraction network corresponding to each modal to obtain second modal features of each second modal data;

[0175] Inputting each second modal feature into an initial evidence neural network corresponding to each modal to obtain a second evidence vector in which each second modal feature is assigned to each preset class identity, and inputting each second modal feature into an initial fine-grained feature aggregation network corresponding to each modal to obtain a second fine-grained feature vector in which each second modal feature is assigned to each preset class identity;

[0176] Generating a view of each second modal data based on each second evidence vector, determining a second fusion view based on the views of each second modal data, and quantifying a view conflict degree and a decision confidence between the views of each second modal data in the second fusion view based on the second fine-grained feature vector;

[0177] Determining a second fusion evidence based on each second evidence vector, determining a predicted identity class corresponding to each sample target based on the second fusion evidence, and determining a confidence degree corresponding to the predicted identity class of each sample target based on the second fusion view;

[0178] Determining a Dirichlet distribution parameter predicted by each second modal data and a fusion Dirichlet distribution parameter predicted based on the evidence vector and the fusion evidence vector of each second modal data corresponding to each sample target, and determining a loss value of a first loss function in combination with a real identity class distribution corresponding to each sample target;

[0179] Determining a loss value of a second loss function based on the view conflict degree between the views of each second modal data;

[0180] Adjusting model parameters corresponding to each of an initial feature extraction network corresponding to each modal, an initial evidence neural network corresponding to each modal and an initial fine-grained feature aggregation network corresponding to each modal based on the loss value of the first loss function and the loss value of the second loss function to obtain an adjusted model;

[0181] Continuing to train the adjusted model until a training stop condition is reached to obtain the target feature extraction network, the target evidence neural network and the target fine-grained feature aggregation network.

[0182] Specifically, in some embodiments, a training step of a cascaded neural network is also included, which includes an initial feature extraction network, an initial evidence neural network, and an initial fine-grained feature aggregation network.

[0183] The training step of the cascaded neural network includes:

[0184] First, sample data is obtained: a plurality of second modal data corresponding to a plurality of sample targets is obtained;

[0185] Then, the initial cascaded neural network is used to determine the output, including:

[0186] Each second modal data is input into the initial feature extraction network corresponding to each modality to obtain the second modal feature of each second modal data; each second modal feature is input into the initial evidence neural network corresponding to each modality to obtain the second evidence vector of each second modal feature being assigned to each preset class identity, and each second modal feature is input into the initial fine-grained feature aggregation network corresponding to each modality to obtain the second fine-grained feature vector of each second modal feature being assigned to each preset class identity.

[0187] Further, the viewpoints of each second modal data are generated based on each second evidence vector, the second fusion viewpoints are determined based on the viewpoints of each second modal data, and the viewpoint conflict degree and the decision confidence between the viewpoints of each second modal data in the second fusion viewpoints are quantified based on the second fine-grained feature vector; the quantification of the viewpoint conflict degree is used to constrain the consistency of the evidence and the viewpoints of different modalities in the model training process of the fine-grained feature aggregation network; the second fusion evidence is determined based on each second evidence vector, the predicted identity class corresponding to each sample target is determined based on the second fusion evidence, and the confidence of the predicted identity class corresponding to each sample target is determined based on the second fusion viewpoints. This process is similar to the steps in the actual recognition process, and will not be described here.

[0188] Further, the Dirichlet distribution parameters of each second modal data and the fusion predicted Dirichlet distribution parameters are determined based on the evidence vectors of each second modal data and the fusion evidence vectors of each sample target, and the loss value of the first loss function is determined in combination with the real identity class distribution corresponding to each sample target.

[0189] For example, the first loss function is the evidence-cross entropy loss and the KL divergence loss. The evidence-cross entropy loss and the KL divergence loss are calculated by the Dirichlet distribution hyperparameters corresponding to the fusion viewpoints of all modalities or single modal viewpoints and the target true value identity distribution gap, and the gap between the Dirichlet distribution parameters of the target sample after removing the correct prediction evidence and the uniform Dirichlet distribution, so as to ensure the uniformity of incorrect class assignment.

[0190] Here, the evidence-cross-entropy loss function utilizes the first... True identity distribution of a sample target Dirichlet distribution corresponding to viewpoints ~Dirichlet( Perform the calculation:

[0191]

[0192] in, The evidence-cross-entropy loss function is represented as follows: Indicates the first The ground truth identity distribution of the i-th identity category of a sample target. Indicates the first The probability corresponding to the viewpoint of the i-th identity category of each sample target, a normalized constant. This represents a multinomial beta function (ensuring the probability density integral is 1). express The probability density function, Indicates the first The probability distribution of predicted identity for each sample target. It is a double gamma function. Indicates the first The Dirichlet distribution hyperparameters corresponding to the target viewpoints of each sample are: , Indicates the first The Dirichlet distribution hyperparameters corresponding to the viewpoints of the i-th identity category of a sample target.

[0193] In addition, the relative entropy (Kullback–Leibler divergence, KL) divergence loss is calculated using the Dirichlet distribution of the target sample after removing evidence of correct prediction and the uniform Dirichlet distribution.

[0194] Furthermore, based on the degree of viewpoint conflict between the viewpoints of each second modality, the loss value of the second loss function is determined, that is, the loss value of the viewpoint conflict degree loss function is determined.

[0195] Then, based on the loss values, the model parameters of the initial feature extraction network, the initial evidence neural network, and the initial fine-grained feature aggregation network for each modality are adjusted to obtain the adjusted model. This completes one model iteration.

[0196] Afterwards, the training of the adjusted model is continued until a training stop condition is reached, obtaining the target feature extraction network, the target evidence neural network and the target fine-grained feature aggregation network. For example, the model parameters of the initial cascaded neural network are fixed when the loss function is minimized, obtaining the target feature extraction network, the target evidence neural network and the target fine-grained feature aggregation network.

[0197] The method provided by the embodiment first acquires the support evidence and the significant feature vector of different modal sample features to a target identity by using the multi-modal evidence neural network and the fine-grained feature aggregation network, thereby acquiring the confidence of the target identity category and constructing the corresponding viewpoint, and finally determines the identity category to which the target belongs by fusing the viewpoints provided by the multi-modal data. Compared with the existing multi-modal fusion scheme, the application introduces the inter-class fine-grained semantic feature correlation when quantifying the evidence conflicts provided by different modal features, weights the projection distance of different evidence viewpoints by the inter-class semantic feature correlation degree, refines the evidence decision conflict between multi-modal data, and improves the reliability of multi-platform fusion decision.

[0198] According to the method for target identity recognition provided by the application, the loss value of the second loss function is determined based on the viewpoint conflict degree between the viewpoints of the second modal data, and the method comprises the following steps:

[0199] The loss value of the second loss function is calculated by using the following formula (1):

[0200] (1)

[0201] Among them, the loss value of the second loss function, the number of preset category identities, the first modal, the first modal, the first modal, the first modal, the first modal, and the fine-grained semantic weighted projection distance between and, the viewpoint of the first target sample feature of the first modal, the viewpoint of the first target sample feature of the first modal, the viewpoint of the first target sample feature of the first modal, the viewpoint modal fine-grained distribution weight vector of the first target sample feature of the first modal, the viewpoint modal fine-grained distribution weight vector of the first target sample feature of the first modal, the viewpoint modal fine-grained distribution weight vector of the first target sample feature of the first modal, the viewpoint modal fine-grained distribution weight vector of the first target sample feature of the first modal, the viewpoint modal fine-grained distribution weight vector of the first target sample feature of the first modal, the viewpoint modal fine-grained distribution weight vector of the first target sample feature of the first modal, the viewpoint modal fine-grained distribution weight vector of the first target sample feature of the first modal, the viewpoint modal fine-grained distribution weight vector of the first target sample feature of the first modal, a view modal fine-grained distribution weight vector of the i-th target sample feature of the j-th modality.

[0202] Specifically, in some embodiments, a second loss function is determined by a view conflict degree between views of each of the first modal data, for constraining consistency of different modal evidence views. An example is as follows:

[0203] (1)

[0204] wherein, represents a loss value of the second loss function, represents a number of preset category identities, represents the i-th target sample feature of the j-th modality, represents the i-th target sample feature of the j-th modality, represents the i-th target sample feature of the j-th modality, represents the i-th target sample feature of the j-th modality, represents and a fine-grained semantic weighted projection distance between and represents the i-th target sample feature of the j-th modality, represents a view of the i-th target sample feature of the j-th modality, represents a view of the i-th target sample feature of the j-th modality, represents a view modal fine-grained distribution weight vector of the i-th target sample feature of the j-th modality, represents a view modal fine-grained distribution weight vector of the i-th target sample feature of the j-th modality. Further, the loss value of the first loss function and the loss value of the second loss function are used to iteratively train the cascaded neural network model (including the initial feature extraction network, the initial evidence neural network and the initial fine-grained feature aggregation network), the model parameters are fixed when the loss function is minimized or a training stop condition is reached, the training of the cascaded neural network model is completed, and then target identity recognition is performed based on the trained cascaded neural network model, thereby improving the reliability of the recognition.

[0205] According to the fine-grained semantic guided trusted multi-source fusion target identity recognition method provided by the application, the plurality of first modal data includes at least one of the following: visible light image, synthetic aperture radar image, near-infrared image, panchromatic image, laser radar point cloud data, voice description and text description.

[0206] According to the fine-grained semantic guided trusted multi-source fusion target identity recognition method provided by the application, the plurality of first modal data includes at least one of the following: visible light image, synthetic aperture radar image, near-infrared image, panchromatic image, laser radar point cloud data, voice description and text description.

[0207] ​​​​​​Specifically, in some embodiments, the plurality of first modality data includes, but is not limited to, visible light images, synthetic aperture radar (SAR) images, near-infrared images, panchromatic images, lidar point cloud data, speech descriptions, and text descriptions.

[0208] The visible light image is the most familiar image form to the human eye, which is used to capture the reflected light information of an object in the visible light band (about 380-750 nanometers) to form a color image with three channels of red, green and blue. The advantage is intuitive, containing rich details and color information; the disadvantage is greatly affected by light and weather. In the night, foggy day, rainy day or smog conditions, the imaging quality will be severely reduced or even unable to image.

[0209] The synthetic aperture radar image is an image data obtained by an active microwave remote sensing technology, which images by emitting microwave signals by itself and receiving signals reflected by ground objects, and does not depend on sunlight. The advantage is that it has all-weather and all-day working ability, and can penetrate clouds, rain, fog and certain vegetation. The disadvantage is that the image is not as intuitive as the optical image, and professional knowledge is needed for interpretation; the image has its own geometric distortion (such as overlap, shadow).

[0210] The near-infrared image is used to capture the reflected light of an object in the near-infrared band (about 750-2500 nanometers). The spectral response of vegetation, water and other objects in this band is very different from that in the visible light band.

[0211] The panchromatic image is usually a single-band, high-resolution black-and-white image obtained by combining the energy of the entire visible light band (which may extend to the near-infrared) of a satellite sensor.

[0212] The speech description is a language description of the target to be identified in the form of audio waveform. The text description is a description of the target in the form of written text.

[0213] The method provided by the embodiment includes at least one of the following: visible light images, synthetic aperture radar images, near-infrared images, panchromatic images, lidar point cloud data, speech descriptions, and text descriptions. The present application can identify the identity of the target to be identified based on the multi-modal data, thereby improving the reliability of the identity recognition.

[0214] According to the method for target identity recognition based on fine-grained semantic guidance provided by the present application, the type of the target feature extraction network is a deep learning model or an encoder, and the encoder is obtained based on multi-modal large model training; the type of the target evidence neural network is a deep learning model; and the type of the target fine-grained feature aggregation network is a deep learning model.

[0215] Specifically, in some embodiments, the type of the target feature extraction network is a deep learning model or an encoder, the encoder is trained based on a multi-modal large model, the type of the target evidence neural network is a deep learning model, and the type of the target fine-grained feature aggregation network is a deep learning model.

[0216] The deep learning model is, for example, a Convolutional Neural Network (CNN), a Transformer network, a Recurrent Neural Network (RNN), or a Long Short-Term Memory (LSTM).

[0217] The encoder trained based on the multi-modal large model is, for example, an Autoencoder (AE) or a Variational Autoencoder (VAE).

[0218] The method provided in the embodiment extracts evidence distribution and a significant feature vector of each modality in different target categories based on different modal evidence networks, introduces a feature distance between the significant feature vectors into difference calculation of multi-platform decision distribution, refines judgment of evidence conflict between platforms by using semantic correlation between categories, and improves reliability of multi-platform fusion target identity recognition based on the deep learning model.

[0219] Figure 2 is a flowchart of a model training method provided by the present application, as shown in Figure 2 The method comprises the following steps.

[0220] In step 201, multi-modal data detected by different platforms on a sample target is obtained, and the multi-modal data is input into a feature extraction network corresponding to each modality to perform feature extraction, so as to obtain features of the multi-modal data.

[0221] In step 202, the features of the multi-modal data are input into a corresponding evidence neural network and a fine-grained feature aggregation network, so as to obtain an evidence vector in which each modality feature is assigned to a multi-class identity and a fine-grained significant feature vector.

[0222] In step 203, a viewpoint after each modality data is fused based on the evidence vector, so as to obtain a fusion viewpoint of different modalities, and the fine-grained significant feature vector is used to quantify a conflict degree between viewpoints and a decision confidence.

[0223] In step 204, the evidence vectors in which each modality feature is assigned to the multi-class identity are fused, so as to obtain a fusion evidence, a probability distribution of the sample target in the multi-class identity is calculated based on the fusion evidence, an identity type of the sample target is determined based on a maximum probability, and a confidence degree corresponding to the identity type of the sample target is determined according to the fusion viewpoint.

[0224] Step 205, iteratively train the parameters of the feature extraction network, the evidence neural network and the fine-grained feature aggregation network to minimize the evidence-cross entropy loss, the divergence and the degree of opinion conflict until the performance converges.

[0225] Figure 3 is the principle schematic diagram of the trusted multi-source fusion target identity recognition method based on fine-grained semantic guidance provided by the application, as shown in the figure, the method comprises the following steps. Figure 3

[0226] First, the multi-modal features of the multi-modal data are input into the corresponding fine-grained feature aggregation network and evidence neural network respectively to obtain the corresponding significant features and evidence.

[0227] The multi-modal features include the output of the visible light image encoder, the output of the SAR image encoder, the output of the near-infrared image encoder and the output of the text encoder; the significant features corresponding to the multi-modal features include: 1 , 2 , 3 , n ; the evidence includes: 1 , 2 , 3 , n ;

[0228] Further, multi-source evidence fusion is performed to infer the target identity; at the same time, multi-source opinion fusion is performed to obtain the degree of opinion conflict and the decision confidence.

[0229] The trusted multi-source fusion target identity recognition device based on fine-grained semantic guidance provided by the application is described below, and the trusted multi-source fusion target identity recognition device based on fine-grained semantic guidance described below can be correspondingly referred to the trusted multi-source fusion target identity recognition method based on fine-grained semantic guidance described above.

[0230] Figure 4 is the structure schematic diagram of the trusted multi-source fusion target identity recognition device based on fine-grained semantic guidance provided by the application, as shown in the figure, the trusted multi-source fusion target identity recognition device 400 based on fine-grained semantic guidance comprises the following modules: Figure 4

[0231] ​​The feature extraction module 410 is configured to input the plurality of first modality data corresponding to the to-be-identified target into a target feature extraction network corresponding to each modality to obtain first modality features of each of the first modality data; the plurality of first modality data are multi-modality data detected by different platforms;

[0232] The identity recognition module 420 is configured to input each of the first modality features into a target evidence neural network corresponding to each modality to obtain a first evidence vector in which each of the first modality features is assigned to at least two preset class identities, and input each of the first modality features into a target fine-grained feature aggregation network corresponding to each modality to obtain a first fine-grained feature vector in which each of the first modality features is assigned to each of the preset class identities;

[0233] Based on each of the first evidence vectors, a viewpoint of each of the first modality data is generated, a first fusion viewpoint is determined based on the viewpoints of each of the first modality data, and a viewpoint conflict degree between the viewpoints of each of the first modality data in the first fusion viewpoint is quantified based on each of the first fine-grained feature vectors; the quantification of the viewpoint conflict degree is used to constrain the consistency of evidences and viewpoints of different modalities in a model training process of the target evidence neural network, the target feature extraction network, and the target fine-grained feature aggregation network;

[0234] Based on each of the first evidence vectors, a first fusion evidence is determined, a target identity corresponding to the to-be-identified target is determined based on the first fusion evidence, and a confidence degree corresponding to the target identity is determined based on the first fusion viewpoint.

[0235] The device provided by the embodiment is characterized by the feature extraction module 410, which is configured to input a plurality of first modal data corresponding to a target to be identified into a target feature extraction network corresponding to each modality to obtain first modal features of each first modal data, and the plurality of first modal data are multi-modal data detected by different platforms.

[0236] The first modal features of the plurality of first modal data are input into the target fine-grained feature aggregation network and the target evidence neural network respectively to obtain the first evidence vector in which each first modal feature is assigned to at least two preset class identities and the first fine-grained feature vector in which each first modal feature is assigned to each preset class identity, and then the viewpoints of each first modal data are generated based on the first evidence vector, the first fusion viewpoint is determined based on the viewpoints of each first modal data, and the viewpoint conflict degree between the viewpoints of each first modal data in the first fusion viewpoint is quantified based on the first fine-grained feature vector, and the quantification of the viewpoint conflict degree is used to constrain the consistency of the evidence and the viewpoint of different modalities in the model training process of the target evidence neural network, the target feature extraction network and the target fine-grained feature aggregation network.

[0237] Figure 5 An example of an entity structure diagram of an electronic device is shown in FIG. 1. Figure 5As shown, the electronic device can include a processor 510, a communications interface 520, a memory 530, and a communications bus 540, wherein the processor 510, the communications interface 520, and the memory 530 complete communications with each other through the communications bus 540. The processor 510 can invoke a logic instruction in the memory 530 to execute the trusted multi-source fusion target identity recognition method based on fine-grained semantic guidance.

[0238] In addition, the logic instruction in the memory 530 described above can be implemented in the form of a software function unit and sold or used as an independent product, and can be stored in a computer-readable storage medium. Based on such understanding, the technical solutions of the present application essentially or the part that contributes to the prior art or part of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium, and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present application. The aforementioned storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk or an optical disk, and various media that can store program codes.

[0239] On the other hand, the present application also provides a computer program product, which includes a computer program, the computer program can be stored on a non-transitory computer readable storage medium, and the computer program is executed by a processor, so that the computer can execute the fine-grained semantic guidance-based trusted multi-source fusion target identity recognition method provided by the above-mentioned methods.

[0240] In yet another aspect, the present application also provides a non-transitory computer readable storage medium, which stores a computer program, and the computer program is executed by a processor to implement the fine-grained semantic guidance-based trusted multi-source fusion target identity recognition method provided by the above-mentioned methods.

[0241] The device embodiments described above are only schematic, wherein the units shown as separate components can or can not be physically separate, and the components shown as units can or can not be physical units, i.e., they can be located in one place, or distributed on multiple network units. Part or all of the modules can be selected to achieve the purpose of the present embodiment scheme according to actual needs. Those skilled in the art can understand and implement it without creative labor.

[0242] Those skilled in the art can clearly understand the technical solutions of the various embodiments from the above description of the embodiments, and the various embodiments can be implemented by means of software with the necessary general hardware platforms, and of course, can also be implemented by hardware. Based on such understanding, the above technical solutions, essentially or in other words, the part of the prior art that makes a contribution, can be embodied in the form of a software product, which can be stored in a computer readable storage medium, such as a ROM / RAM, a magnetic disk, an optical disk, and the like, and includes a number of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.

[0243] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present application, rather than limit them; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that: it can still modify the technical solutions recorded in the foregoing embodiments, or make equivalent replacement for some technical features therein; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. A reliable multi-source fusion target identity recognition method based on fine-grained semantic guidance, characterized in that, include: Multiple first modal data corresponding to the target to be identified are input into the target feature extraction network corresponding to each modality to obtain the first modal features of each first modal data; The various first-mode data are multimodal data detected by different platforms; Each first modality feature is input into the target evidence neural network corresponding to each modality to obtain a first evidence vector in which each first modality feature is assigned to at least two preset category identities. Each first modality feature is then input into the target fine-grained feature aggregation network corresponding to each modality to obtain a first fine-grained feature vector in which each first modality feature is assigned to each preset category identity. The viewpoints of each first modality data are generated based on each first evidence vector, a first fused viewpoint is determined based on each first modality data viewpoint, and the degree of viewpoint conflict between the viewpoints of each first modality data in the first fused viewpoint is quantified based on each first fine-grained feature vector; the quantification of the degree of viewpoint conflict is used to constrain the consistency of evidence and viewpoints of different modalities during the model training process of the target evidence neural network, the target feature extraction network, and the target fine-grained feature aggregation network. First fused evidence is determined based on each of the first evidence vectors, target identity is determined based on the first fused evidence, and confidence level is determined based on the first fused viewpoint.

2. The trusted multi-source fusion target identity recognition method based on fine-grained semantic guidance according to claim 1, characterized in that, The viewpoint of each first modality data includes the prior probability, confidence level, and unconfidence level of each first modality data being assigned to each of the preset category identities; The viewpoints generated based on each of the first evidence vectors to produce each of the first modal data, and the viewpoints determined based on each of the first modal data to determine the first fused viewpoint, include: Based on each of the first evidence vectors, calculate the Dirichlet distribution parameters corresponding to each mode, and based on the Dirichlet distribution parameters corresponding to each mode, calculate the Dirichlet intensity corresponding to each mode. Based on the number of the preset category identities, determine the prior probability that each first modal data is assigned to each preset category identity; Based on each of the first evidence vectors and the Dirichlet strength corresponding to each modality, the confidence level of each of the first modality data being assigned to each of the preset category identities is determined; Based on the Dirichlet intensity corresponding to each modality, the degree of distrust in which each first modality data is assigned to each preset category identity is determined; The prior probability of each first modality data being assigned to each preset category identity, the confidence level of each first modality data being assigned to each preset category identity, and the unconfidence level of each first modality data being assigned to each preset category identity are fused to obtain the first fused viewpoint.

3. The trusted multi-source fusion target identity recognition method based on fine-grained semantic guidance according to claim 1, characterized in that, The step of quantifying the degree of view conflict between viewpoints in the first fused viewpoint based on each of the first fine-grained feature vectors includes: For any pair of viewpoints in each of the first modal data, based on the correlation between the modal fine-grained distribution weight vector of the first viewpoint and the modal fine-grained distribution weight vector of the second viewpoint in the viewpoint pair, the Dirichlet distribution parameters corresponding to each modality, and the Dirichlet intensity corresponding to each modality, the fine-grained semantically weighted projection distance between the first viewpoint and the second viewpoint is determined; the modal fine-grained distribution weight vector of the first viewpoint and the modal fine-grained distribution weight vector of the second viewpoint in the viewpoint pair are determined based on each of the first fine-grained feature vectors; Based on the degree of unreliability of the first modal data being assigned to each of the preset category identities in the first viewpoint, and the degree of unreliability of the first modal data being assigned to each of the preset category identities in the second viewpoint, the corresponding fusion determinism between the first viewpoint and the second viewpoint is determined; The product of the projection distance and the fusion determinism is determined as the degree of conflict between the viewpoint pairs; Based on the degree of conflict between each pair of viewpoints, the degree of conflict between the viewpoints of each of the first modal data in the first fused viewpoint is determined.

4. The trusted multi-source fusion target identity recognition method based on fine-grained semantic guidance according to claim 1, characterized in that, The step of determining the first fused evidence based on each of the first evidence vectors, and determining the target identity corresponding to the target to be identified based on the first fused evidence, includes: Calculate the mean of each of the first evidence vectors, and determine the first fused evidence based on the mean; Calculate the Dirichlet distribution parameters and the Dirichlet intensity corresponding to the first fusion mode based on the first fusion evidence; Based on the Dirichlet distribution parameters corresponding to the first fusion mode and the Dirichlet intensity corresponding to the first fusion mode, the probability distribution of the target to be identified as each of the preset category identities is calculated; The target's identity is determined based on the maximum probability value in the probability distribution.

5. The trusted multi-source fusion target identity recognition method based on fine-grained semantic guidance according to any one of claims 1-4, characterized in that, The training steps for the target feature extraction network, the target evidence neural network, and the target fine-grained feature aggregation network include: Multiple second modal data corresponding to multiple sample targets are acquired, and each second modal data is input into the initial feature extraction network corresponding to each modality to obtain the second modal features of each second modal data; Each second modality feature is input into the initial evidence neural network corresponding to each modality to obtain a second evidence vector in which each second modality feature is assigned to each preset category identity; and each second modality feature is input into the initial fine-grained feature aggregation network corresponding to each modality to obtain a second fine-grained feature vector in which each second modality feature is assigned to each preset category identity. The viewpoints of each second modal data are generated based on each second evidence vector, the second fused viewpoints are determined based on the viewpoints of each second modal data, and the degree of viewpoint conflict and decision credibility among the viewpoints of each second modal data in the second fused viewpoints are quantified based on the second fine-grained feature vectors. Second fusion evidence is determined based on each of the second evidence vectors, the predicted identity category corresponding to each of the sample targets is determined based on the second fusion evidence, and the confidence level corresponding to the predicted identity category corresponding to each of the sample targets is determined based on the second fusion view. Based on the evidence vectors and fused evidence vectors of the second modality data corresponding to each sample target, the Dirichlet distribution parameters predicted by each second modality data and the Dirichlet distribution parameters predicted by the fused data are determined, and the loss value of the first loss function is determined in combination with the distribution of the real identity categories corresponding to each sample target. The loss value of the second loss function is determined based on the degree of viewpoint conflict between the viewpoints of each second modality data. Based on the loss values ​​of the first loss function and the second loss function, adjust the model parameters of the initial feature extraction network, the initial evidence neural network, and the initial fine-grained feature aggregation network corresponding to each modality to obtain the adjusted model; The adjusted model is trained until the training stops, resulting in the target feature extraction network, the target evidence neural network, and the target fine-grained feature aggregation network.

6. The trusted multi-source fusion target identity recognition method based on fine-grained semantic guidance according to claim 5, characterized in that, The determination of the loss value of the second loss function based on the degree of viewpoint conflict between viewpoints in each of the second modal data includes: The loss value of the second loss function is calculated using the following formula (1): (1) in, This represents the loss value of the second loss function. This indicates the number of the preset category identities. Indicates the first One modality, Indicates the first One modality, express and The fine-grained semantically weighted projection distance between them. Indicates the first The first mode The perspective of the characteristics of a target sample Indicates the first The first mode The perspective of the characteristics of a target sample Indicates the first The first mode The viewpoint modality fine-grained distribution weight vector of each target sample feature Indicates the first The first mode A fine-grained distribution weight vector of viewpoint modal features of a target sample.

7. The trusted multi-source fusion target identity recognition method based on fine-grained semantic guidance according to any one of claims 1-4, characterized in that, The multiple first-mode data include at least one of the following: visible light images, synthetic aperture radar images, near-infrared images, panchromatic images, lidar point cloud data, voice descriptions, and text descriptions.

8. The trusted multi-source fusion target identity recognition method based on fine-grained semantic guidance according to any one of claims 1-4, characterized in that, The target feature extraction network is a deep learning model or an encoder, wherein the encoder is trained based on a multimodal large model; the target evidence neural network is a deep learning model; and the target fine-grained feature aggregation network is a deep learning model.

9. A trusted multi-source fusion target identity recognition device based on fine-grained semantic guidance, characterized in that, include: The feature extraction module is used to input multiple first modal data corresponding to the target to be identified into the target feature extraction network corresponding to each modality to obtain the first modal features of each first modal data; The various first-mode data are multimodal data detected by different platforms; The identity recognition module is used to input each first modality feature into the target evidence neural network corresponding to each modality to obtain a first evidence vector in which each first modality feature is assigned to at least two preset category identities, and to input each first modality feature into the target fine-grained feature aggregation network corresponding to each modality to obtain a first fine-grained feature vector in which each first modality feature is assigned to each preset category identity. The viewpoints of each first modality data are generated based on each first evidence vector, a first fused viewpoint is determined based on each first modality data viewpoint, and the degree of viewpoint conflict between the viewpoints of each first modality data in the first fused viewpoint is quantified based on each first fine-grained feature vector; the quantification of the degree of viewpoint conflict is used to constrain the consistency of evidence and viewpoints of different modalities during the model training process of the target evidence neural network, the target feature extraction network, and the target fine-grained feature aggregation network. First fused evidence is determined based on each of the first evidence vectors, target identity is determined based on the first fused evidence, and confidence level is determined based on the first fused viewpoint.

10. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the trusted multi-source fusion target identity recognition method based on fine-grained semantic guidance as described in any one of claims 1 to 8.

Citation Information

Patent Citations

  • Multi-source situation information fusion method

    CN105391694A

  • Multi-modal fine-grained feature hybrid recognition system and method

    CN116258898A