Photoelectric target identification method based on multi-modal knowledge graph reasoning
By constructing a multimodal knowledge graph of optoelectronic targets and using the Transformer model to perform multimodal information interaction, the existing optoelectronic target recognition methods are solved, and the problem that existing optoelectronic target recognition methods are difficult to adapt to multiple data types and complex processing methods is achieved, achieving more accurate and richer optoelectronic target recognition effects.
Patent Information
- Application Number
- CN202411957232.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-29
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2044-12-29
AI Technical Summary
The existing optoelectronic target recognition methods are difficult to adapt to multiple data types, the processing methods are complex, and lack self-iteration and update processes, so they cannot deeply explore and connect fine-grained entity associations of optoelectronic targets.
The multimodal knowledge graph inference method is used to construct a multimodal knowledge graph of photoelectric targets, through image feature extraction, text encoding and numerical vector extraction, and multimodal information interaction and reasoning are carried out in combination with the Transformer model to achieve the identification of photoelectric targets and the acquisition of attribute information.
It improves the accuracy and richness of the photoelectric target recognition task, can more effectively manage and utilize multiple types of photoelectric target data, and supports photoelectric target object recognition and entity query link tasks in different modes.
Smart Images

Figure CN119963882A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of multimodal knowledge graph construction, representation learning and reasoning, and specifically relates to a photoelectric target recognition method based on multimodal knowledge graph reasoning. Background Art
[0002] Optoelectronic target recognition technology plays a vital role in military applications such as combat command and control, battlefield situation awareness, tactical intention recognition, decision support and threat assessment in the field of optoelectronic confrontation. At present, optoelectronic target recognition technology is mainly based on methods such as multi-sensor data fusion or expert systems. For example, Wang Tiehong's "Research on Multi-Sensor Data Fusion Model of Optoelectronic Information System" gives a multi-level sensor information fusion model structure. However, the fusion model based on multi-level processing such as data level, feature level, and decision level makes the algorithm process long and the design complex. Expert system is a branch of artificial intelligence. In the field of military target recognition, recognition knowledge has the characteristics of diversity, timeliness, and regionality. Expert system has problems such as difficulty in acquiring knowledge in the development stage, short validity period after system delivery, and inability to dynamically and autonomously iterate. However, from the perspective of data format and type, commonly used data types for optoelectronic target recognition include structured, semi-structured and unstructured data such as intelligence text, optoelectronic detection images, SQL relational databases, GIS geographic data, etc. These data have the characteristics of many types, large data volume, complex structure, sparse key values, and high data duplication in actual combat or training environments. Conventional methods based on multi-sensor data fusion reasoning are difficult to fully adapt to various data types, the processing methods are complex, and there is a lack of self-iteration and update process.
[0003] With the development of artificial intelligence technology and the Internet, knowledge graphs have powerful semantic expression capabilities, knowledge extraction and fusion capabilities, and reasoning and computing capabilities, providing an effective solution for the knowledge organization and storage of multimodal data and upper-level applications. Multimodal knowledge graphs mainly include {E, R, A, V, T R ,T A}, where E is the entity, R is the relationship, A is the multimodal data attribute, V is the multimodal data attribute value, T R is a multimodal relation triple, which describes the relationship between entities. For example, the relationship between the Paveway series missiles and the laser semi-active guided weapons in the multimodal knowledge graph belongs to; T A It is an attribute triplet that describes the attribute information of the entity itself, such as the country of production, speed, function, etc. of the Paveway series missile entity itself. For example, MMKG uses Wikipedia URI as the query string, crawls images from the search engine, and uses it as an attribute of the entity.
[0004] Existing optoelectronic target recognition methods mainly focus on the detection and recognition of target categories, and do not delve into the fine-grained entity associations of the target, that is, they are unable to conduct in-depth mining and connection of other attributes and additional information of the optoelectronic target.
[0005] Considering the particularity of the optoelectronic confrontation field and the limitation of the pre-trained language model to specific applicable task scenarios, the existing models cannot achieve good results in the field of optoelectronic target recognition. Therefore, it is necessary to propose an optoelectronic target recognition method based on a multimodal pre-trained language model for the optoelectronic confrontation field and the optoelectronic target recognition task, extract the joint entity representation features of optoelectronic targets under different modalities, use efficient representations to describe optoelectronic target objects, fuse the optoelectronic target entity representations of different modalities, and encode the optoelectronic target modal features into the context representation in the model, so as to support the optoelectronic target object recognition tasks and entity query linking tasks of different modalities. Summary of the invention
[0006] 1. Technical issues to be resolved
[0007] The present invention proposes an optoelectronic target recognition method based on multimodal knowledge graph reasoning to solve the technical problems that conventional optoelectronic target recognition methods are difficult to adapt to optoelectronic target data with the characteristics of multiple types, large data volume, complex structure, sparse key values, high data repetition, etc. in actual combat or training environments, and lack effective management and utilization.
[0008] (II) Technical solution
[0009] In order to solve the above technical problems, the present invention proposes a photoelectric target recognition method based on multimodal knowledge graph reasoning, and the photoelectric target recognition method comprises the following steps:
[0010] S1. Constructing a multimodal knowledge graph of optoelectronic targets based on the field of optoelectronic countermeasures
[0011] The sources of optoelectronic target knowledge data are divided into three types according to structure: structured data, semi-structured data and unstructured data. Data preprocessing, information extraction and knowledge fusion are performed on different types of data. The processed entities, relationships, attributes and empirical rules are input into the optoelectronic target multimodal knowledge graph to construct an optoelectronic target multimodal knowledge graph based on the optoelectronic confrontation field.
[0012] S2. Image feature vector extraction
[0013] The photoelectric target image data obtained from the photoelectric target multimodal knowledge graph constructed in step S1 is resized to 224*224 pixels, and then input into the CNN baseline model for image feature extraction;
[0014] S3. Encode text form into a digital vector
[0015] When the photoelectric target knowledge data is in the form of text data, the text information is encoded into a digital code that can be calculated by computer language according to the dictionary;
[0016] S4. Numerical vector extraction
[0017] When the input of optoelectronic target knowledge data is expressed in the form of a {key:value} key-value pair, the value itself is directly embedded as a vector V = {v0,v1,..v q-1}, where V is a numeric vector, v0,…v q-1 is a specific numerical value;
[0018] S5. The image feature vector, the digital vector corresponding to the text form and the numerical vector obtained in the above steps are independently input into the Transformer model to obtain the image mode, text mode and numerical mode embedding representation vector X corresponding to the optoelectronic target multimodal knowledge graph constructed in step S1. M ;
[0019] S6. The different modalities obtained in step S5 are independently embedded into the representation vector X M The decomposition vector K key vector and V value vector of the perform multimodal information interaction;
[0020] S7. Input decoders of different modalities to calculate probability prediction distribution results
[0021] Suppose the feature vectors of the three modal outputs of image modality, text modality, and numerical modality after multimodal interaction are expressed as:
[0022] The entity probability prediction distribution results of different modes and the final probability prediction distribution results are obtained through the following formula:
[0023]
[0024]
[0025] PRO=(1-α-β)PRO T +αPRO P +βPRO V
[0026] ReLu(x)=max(x,0)
[0027] Among them, PRO T 、PRO P 、PRO V, PRO are the text modality, image modality, numerical modality and the final entity probability prediction distribution results; α, β are the hyperparameters for adjusting the proportion of the three modalities, which are obtained during model training; MLP is the multi-layer perceptron calculation
[0028] S8. Photoelectric target recognition based on trained multimodal knowledge graph reasoning model
[0029] The optoelectronic target multimodal knowledge graph dataset constructed in step S1 is input into steps S2-S7 for Transformer model training to obtain the specific values of the model hidden layer parameters, and the values of the hyperparameters α and β in step S7 are set according to the distribution size of the actual dataset; at this point, the inference model based on the multimodal knowledge graph is constructed;
[0030] The optoelectronic target data that needs to be inferred is calculated according to steps S2-S7 to obtain the final probability distribution PRO on the output candidate entities, and the candidate entities are arranged in descending order according to the probability size, and the top K entities with the largest probability are output as the entity results of optoelectronic target recognition; query K entities from the multimodal knowledge graph to obtain attribute information, and the optoelectronic target recognition work is completed.
[0031] Furthermore, in step S1, the structured data includes sql data and xml data, the semi-structured data includes an excel table of a custom data type, and the unstructured data includes photoelectric target image data acquired by the photoelectric sensor and text-based numerical data.
[0032] Furthermore, in step S1, data preprocessing includes word segmentation, named entity recognition, part-of-speech standardization, deletion of duplicate information, correction of invalid values and missing value processing for different types of data.
[0033] Furthermore, in step S1, information extraction includes entity extraction, relationship extraction, and attribute extraction.
[0034] Furthermore, in step S1, knowledge fusion includes eliminating ambiguity and aligning entities of the extracted entities, relationships, attributes and empirical rule contents.
[0035] Furthermore, in step S2, the CNN baseline model includes a VGG, ResNet or DenseNet model.
[0036] Further, in step S4, the optoelectronic target knowledge data includes the azimuth angle, pitch angle, and azimuth speed data of the optoelectronic target.
[0037] Further, in step S5, the embedding representation vectors X of the image modality encoder, the text modality encoder, and the numerical modality encoder are calculated respectively according to the following formulas: M :
[0038] X=X0+X position
[0039] X M =MHA(LN(X))+X
[0040] Among them, X0 is the image feature vector, text encoding vector or numerical vector, X position is the position embedding parameter of the Transformer model, MHA is the multi-head standard calculation formula of the Transformer model, and LN is the normalized calculation of the network layer parameters in the Transformer model using the standard normal distribution function.
[0041] Furthermore, in step S6, when performing multimodal information interaction, the K key vector and V key vector of the modality with the smaller number among the image modality, the text modality and the numerical modality are input into the K key vector and V key vector of the other two modalities.
[0042] Further, in step S6, the embedding representation vector X of the text modality is obtained by the following formula: M K-bond vector in and V value vector The embedded representation vectors X are input to the image mode and the numerical mode respectively. M The K key vector and V value vector are:
[0043]
[0044] Among them, [] represents the concatenation operation, Attention is the calculation formula of the self-attention mechanism of Transformer, and head is the attention calculation after the interaction of modal information; Embedding vector X for text modality M The corresponding decomposition vector; Embedding vector X for image modality M The corresponding decomposition vector; Embedding vector X for the numerical mode M The corresponding decomposition vector; the subscript parameter i is the value of a specific element in the decomposition vector.
[0045] (III) Beneficial effects
[0046] The present invention proposes an optoelectronic target recognition method based on multimodal knowledge graph reasoning. By constructing a multimodal knowledge graph of optoelectronic targets, the optoelectronic targets are connected by relationships, which facilitates the acquisition of relevant information and association relationships, and the input information is dynamically linked and reasoned with the multimodal knowledge graph. The optoelectronic target selects the matching entity with the highest probability from multiple candidate entities of the link and reasoning results, and outputs all attribute information related to the matching entity to provide support for subsequent combat command and control, battlefield situation awareness, auxiliary decision-making and threat assessment and other upper-level military applications. In the present invention, the multimodal knowledge graph integrates different types of data, and performs joint representation learning through multiple modes of image text numerical values, which can provide more comprehensive and multidimensional information, and help improve the accuracy and richness of optoelectronic target recognition tasks; constructing a multimodal knowledge graph based on optoelectronic targets can structure equipment knowledge data, improve knowledge query efficiency, establish information associations, provide intelligence analysis and information services, and support upper-level military applications such as threat assessment and auxiliary decision-making; the reasoning model of the multimodal knowledge graph is based on a neural network, and the intelligentization can continuously expand and update model parameters to continuously adapt to new data and new needs. BRIEF DESCRIPTION OF THE DRAWINGS
[0047] Figure 1 Build a process for multimodal knowledge graph of optoelectronic targets;
[0048] Figure 2 A rendering showing some of the contents of the multimodal knowledge graph of optoelectronic targets;
[0049] Figure 3 Obtain feature vector schematic for feature encoding;
[0050] Figure 4 Implement schematics for multimodal interactions;
[0051] Figure 5 It is a multi-layer perceptron (MLP) neural network graph with 2 layers and one hidden layer. DETAILED DESCRIPTION
[0052] In order to make the purpose, content and advantages of the present invention more clear, the specific implementation methods of the present invention are further described in detail below in conjunction with the drawings and examples.
[0053] This embodiment proposes a photoelectric target recognition method based on multimodal knowledge graph reasoning, and the photoelectric target recognition method specifically includes the following steps:
[0054] S1. Constructing a multimodal knowledge graph of optoelectronic targets based on the field of optoelectronic countermeasures
[0055] The sources of optoelectronic target knowledge data are divided into structured data (such as SQL data, XML data, etc.), semi-structured data (such as custom data type Excel tables, etc.), and unstructured data (such as optoelectronic target image data obtained from various optoelectronic sensors, text-based numerical data, etc.). Figure 1 The construction process of the multimodal knowledge graph of optoelectronic targets shown in the figure is to subject the above different types of data to data preprocessing processes such as word segmentation, named entity recognition, part-of-speech standard, deletion of duplicate information, correction of invalid values and missing values, and then perform information extraction processes such as entity extraction, relationship extraction, attribute extraction, and empirical rules. Then, the extracted entities, relationships, attributes, empirical rules and other contents are subjected to knowledge fusion processes such as disambiguation and entity alignment. Finally, the processed entities, relationships, attributes, empirical rules and other contents can be input into the multimodal knowledge graph of optoelectronic targets to construct a multimodal knowledge graph of optoelectronic targets based on the field of optoelectronic confrontation.
[0056] like Figure 2 As shown in the figure, the multimodal knowledge graph of optoelectronic targets {E, R, A, V, T R ,T A}Detailed content. E entity represents the content, such as: Paveway series, Pyros small tactical ammunition, Hammer modular guided bomb, US AGM-114 Hellfire missile, US AN / AAQ-13Lantirn, US AN / AAQ-14Lantirn, French rubis, US pathfinde, US AN / AAQ-38NITE HAWK, Israeli Litening and other optoelectronic targets. R relationship represents the content, such as: Paveway series, Pyros small tactical ammunition, Hammer modular guided bomb and other optoelectronic targets are laser semi-active guided weapons, and the relationship is belonging to. A is the content represented by multimodal data attributes, such as: the country of production of the Paveway series missile entity itself, hit accuracy, function and other multimodal data attributes. V is the multimodal data attribute value, such as the country of production of the Paveway series missile entity itself is the United States, and the hit accuracy is 1-3 meters and other specific data attribute values. T R is a multimodal relation triple, such as {Paveway series missiles belong to laser semi-active guided weapons}. A It is a triplet of attributes, such as {Paveway series missiles, country of production, United States}.
[0057] S2. Image feature vector extraction
[0058] The photoelectric target image data obtained from the photoelectric target multimodal knowledge graph constructed in step S1 is resized to 224*224 pixels, and then input into a CNN baseline model such as VGG, ResNet or DenseNet for image feature extraction.
[0059] The photoelectric target classification task is evaluated by inputting the image modality data in the photoelectric target multimodal knowledge graph data constructed in step S1 into the VGG, ResNet or DenseNet model. The model with the best comprehensive evaluation indicators of accuracy, precision and recall rate in the photoelectric target classification task is used as the CNN baseline model for image feature vector extraction.
[0060] S3. Encode text form into a digital vector
[0061] When the optoelectronic target knowledge data is in the form of text data, the text information is encoded into a digital code that can be calculated by a computer language according to a dictionary.
[0062] S4. Numerical vector extraction
[0063] When the input of optoelectronic target knowledge data is expressed in the form of {key:value} key-value pairs, because the value itself is an attribute of the entity feature, the value itself is directly embedded as a vector V = {v0,v1,..v q-1}, where V is a numeric vector, v0,…v q-1 It is a series of specific data such as the azimuth angle, pitch angle, azimuth speed, etc. of the optoelectronic target.
[0064] S5. The image feature vector, the digital vector corresponding to the text form and the numerical vector obtained in the above steps are independently input into the Transformer model to obtain the image mode, text mode and numerical mode embedding representation vector X corresponding to the optoelectronic target multimodal knowledge graph constructed in step S1. M .
[0065] According to the following formula, the embedded representation vector X of the image modality encoder (P-EnCoder), text modality encoder (T-EnCoder) and numerical modality encoder (V-EnCoder) are calculated respectively: M :
[0066] X=X0+X position
[0067] X M =MHA(LN(X))+X
[0068] Among them, X0 is the image feature vector, text encoding vector or numerical vector, X positionis the position embedding parameter of the Transformer model, MHA is the multi-head standard calculation formula of the Transformer model, and LN is the normalized calculation of the network layer parameters in the Transformer model using the standard normal distribution function.
[0069] S6. The different modalities obtained in step S5 are independently embedded into the representation vector X M The decomposition vector K key vector and V value vector perform multimodal information interaction.
[0070] Figure 3 For the specific technical principle diagram, the optoelectronic target data type is divided into image mode, text mode and numerical mode according to the modality; the image modality data is extracted through the baseline model of the convolutional neural network CNN, and then the feature is encoded through the encoder of the P-Encoder to obtain the feature vector, the text modality data is encoded into a numerical form through text encoding, and then the feature is encoded through the T-Encoder to obtain the feature vector, and the numerical mode is directly encoded through the V-Encoder to obtain the feature vector; then the feature vectors of the three independent modes are fused and interacted between different modes, so that each mode contains the vector information of other modes; finally, the probability prediction values of the three feature vectors after multi-modal interaction are decoded and output by P-Dncoder, T-Dncoder and V-Dncoder respectively, and the final inference probability value P is obtained.
[0071] The specific technical principles of the multimodal interaction module during the implementation of this method are as follows Figure 4 As shown in the figure, the image modality, text modality and numerical modality are composed of three separate transformer structures. When the number of text modality samples is small, the multimodal information interaction process is to input the K-key vector and V-key vector of the text modality into the K-key vector and V-key vector of the image modality and numerical modality respectively. If the number of samples of other modalities is small, the K-key vector and V-key vector of the modality with the smaller number are input into the K-key vector and V-key vector of the other two modalities in the same way.
[0072] The embedding representation vector X of the text modality is transformed into M The K-bond vector in and V value vector The embedded representation vectors X are input to the image mode and the numerical mode respectively. M The K key vector and V value vector.
[0073]
[0074] Among them, [] represents the concatenation operation, Attention is the calculation formula of the self-attention mechanism of Transformer, and head is the attention calculation after the interaction of modal information; Embedding vector X for text modality M The corresponding decomposition vector; Embedding vector X for image modality M The corresponding decomposition vector; Embedding vector X for the numerical mode M The corresponding decomposition vector; the subscript parameter i is the value of a specific element in the decomposition vector.
[0075] S7. Input decoders of different modalities to calculate probability prediction distribution results
[0076] Suppose the feature vectors of the three modal outputs of image modality, text modality, and numerical modality after multimodal interaction are expressed as:
[0077] The entity probability prediction distribution results of different modes and the final probability prediction distribution results are obtained through the following formula:
[0078]
[0079] PRO=(1-α-β)PRO T +αPRO P +βPRO V
[0080] ReLu(x)=max(x,0)
[0081] Among them, PRO T 、PRO P 、PRO V , PRO are the text modality, image modality, numerical modality and the final entity probability prediction distribution results; α, β are hyperparameters for adjusting the proportion of the three modalities, which are obtained during model training; MLP is a multi-layer perceptron calculation. The specific technical principle of the decoding process DeCoder in the implementation of this method is as follows Figure 5 As shown in the figure, a 2-layer structure diagram is given, x1, x2, .. xn are the input feature vectors that need to be decoded, h1, h2, ... hn are the hidden layer parameters of the network structure obtained by training, and o1, o2, ... on are the final entity probability prediction distribution results of different modes.
[0082] S8. Photoelectric target recognition based on trained multimodal knowledge graph reasoning model
[0083] The photoelectric target multimodal knowledge graph dataset constructed in step S1 is input into steps S2-S7 for Transformer model training to obtain the specific values of the model hidden layer parameters, and the values of the hyperparameters α and β in step S7 are set according to the distribution size of the actual dataset. At this point, the inference model based on the multimodal knowledge graph is completed.
[0084] The optoelectronic target data that needs to be inferred is calculated according to steps S2-S7 to obtain the final probability distribution PRO on the output candidate entities, and the candidate entities are arranged in descending order according to the probability size, and the top K entities with the largest probability are output as the entity results of optoelectronic target recognition. Then, the K entities are queried from the multimodal knowledge graph to obtain all relevant attribute information for the relevant upper-layer application, and the optoelectronic target recognition work is completed.
[0085] The above is only a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the technical principles of the present invention. These improvements and modifications should also be regarded as the scope of protection of the present invention.
Claims
1. A photoelectric target recognition method based on multimodal knowledge graph reasoning, characterized in that: The photoelectric target recognition method comprises the following steps: S1. Constructing a multimodal knowledge graph of optoelectronic targets based on the field of optoelectronic countermeasures The sources of optoelectronic target knowledge data are divided into three types according to structure: structured data, semi-structured data and unstructured data. Data preprocessing, information extraction and knowledge fusion are performed on different types of data. The processed entities, relationships, attributes and empirical rules are input into the optoelectronic target multimodal knowledge graph to construct an optoelectronic target multimodal knowledge graph based on the optoelectronic confrontation field. S2. Image feature vector extraction The photoelectric target image data obtained from the photoelectric target multimodal knowledge graph constructed in step S1 is resized to 224*224 pixels, and then input into the CNN baseline model for image feature extraction; S3. Encode text form into a digital vector When the photoelectric target knowledge data is in the form of text data, the text information is encoded into a digital code that can be calculated by computer language according to the dictionary; S4. Numerical vector extraction When the input of optoelectronic target knowledge data is expressed in the form of a {key:value} key-value pair, the value itself is directly embedded as a vector V = {v0,v1,..v q-1 }, where V is a numeric vector, v0,…v q-1 is a specific numerical value; S5. The image feature vector, the digital vector corresponding to the text form and the numerical vector obtained in the above steps are independently input into the Transformer model to obtain the image mode, text mode and numerical mode embedding representation vector X corresponding to the optoelectronic target multimodal knowledge graph constructed in step S1. M ; S6. The different modalities obtained in step S5 are independently embedded into the representation vector X M The decomposition vector K key vector and V value vector of the perform multimodal information interaction; S7. Input decoders of different modalities to calculate probability prediction distribution results Suppose the feature vectors of the three modal outputs of image modality, text modality, and numerical modality after multimodal interaction are expressed as: The entity probability prediction distribution results of different modes and the final probability prediction distribution results are obtained through the following formula: PRO=(1-α-β)RPO T +aPRO P +βPRO V ReLu(x)=max(x,0) Among them, PRO T 、PRO P 、PRO V , PRO are the text modality, image modality, numerical modality and the final entity probability prediction distribution results; α, β are the hyperparameters for adjusting the proportion of the three modalities, which are obtained during model training; MLP is the multi-layer perceptron calculation S8. Perform photoelectric target recognition based on the trained multimodal knowledge graph inference model. Input the photoelectric target multimodal knowledge graph dataset constructed in step S1 into steps S2-S7 for Transformer model training to obtain the specific values of the model hidden layer parameters, and set the values of the hyperparameters α and β in step S7 according to the distribution size of the actual data set; at this point, the inference model based on the multimodal knowledge graph is constructed; The optoelectronic target data that needs to be inferred is calculated according to steps S2-S7 to obtain the final probability distribution PRO on the output candidate entities, and the candidate entities are arranged in descending order according to the probability size, and the top K entities with the largest probability are output as the entity results of optoelectronic target recognition; query K entities from the multimodal knowledge graph to obtain attribute information, and the optoelectronic target recognition work is completed.
2. The photoelectric target recognition method based on multimodal knowledge graph reasoning according to claim 1, characterized in that: In step S1, the structured data includes sql data and xml data, the semi-structured data includes an excel table of a custom data type, and the unstructured data includes photoelectric target image data acquired by a photoelectric sensor and text-based numerical data.
3. The photoelectric target recognition method based on multimodal knowledge graph reasoning according to claim 1, characterized in that: In step S1, data preprocessing includes word segmentation, named entity recognition, part-of-speech standardization, deletion of duplicate information, correction of invalid values and missing value processing for different types of data.
4. The photoelectric target recognition method based on multimodal knowledge graph reasoning according to claim 1, characterized in that: In step S1, information extraction includes entity extraction, relationship extraction, and attribute extraction.
5. The photoelectric target recognition method based on multimodal knowledge graph reasoning according to claim 1, characterized in that: In step S1, knowledge fusion includes eliminating ambiguity and aligning entities among the extracted entities, relationships, attributes, and empirical rules.
6. The photoelectric target recognition method based on multimodal knowledge graph reasoning according to claim 1, characterized in that: In step B2, the CNN baseline model includes a VGG, ResNet, or DenseNet model.
7. The photoelectric target recognition method based on multimodal knowledge graph reasoning according to claim 1, characterized in that: In step S4, the photoelectric target knowledge data includes the azimuth angle, pitch angle, and azimuth speed data of the photoelectric target.
8. The photoelectric target recognition method based on multimodal knowledge graph reasoning according to claim 1, characterized in that: In step S5, the embedding representation vectors X of the image modality encoder, text modality encoder, and numerical modality encoder are calculated respectively according to the following formulas: M : X=X0+X position X M =MHA(LN(X))+X Among them, X0 is the image feature vector, text encoding vector or numerical vector, X position is the position embedding parameter of the Transformer model, MHA is the multi-head standard calculation formula of the Transformer model, and LN is the normalized calculation of the network layer parameters in the Transformer model using the standard normal distribution function.
9. The photoelectric target recognition method based on multimodal knowledge graph reasoning according to claim 1, characterized in that: In step S6, when performing multimodal information interaction, the K key vector and V key vector of the modality with fewer numbers among the image modality, the text modality and the numerical modality are input into the K key vector and V key vector of the other two modalities.
10. The photoelectric target recognition method based on multimodal knowledge graph reasoning according to claim 9, characterized in that: In step S6, the embedding representation vector X of the text modality is converted into M K-bond vector in and V value vector The embedded representation vectors X are input to the image mode and the numerical mode respectively. M The K key vector and V value vector are: Among them, [] represents the concatenation operation, Attention is the calculation formula of the self-attention mechanism of Transformer, and head is the attention calculation after the interaction of modal information; Embedding vector X for text modality M The corresponding decomposition vector; Embedding vector X for image modality M The corresponding decomposition vector; Embedding vector X for the numerical mode M The corresponding decomposition vector; the subscript parameter i is the value of a specific element in the decomposition vector.
Citation Information
Patent Citations
Large model prompt generation method based on knowledge graph
CN117591663A
Power grid regulation and control knowledge graph construction system, method and program product
CN118211647A
Equipment fault risk monitoring method and system fusing machine vision and knowledge graph
CN118245604A
Method and system for predicting biological entities
GB202402771D0