Knowledge-enhanced aircraft carrier visual target detection and state reasoning method and system
By employing multi-source data processing and knowledge graph-guided methods, the problems of uncontrollable reasoning and difficulty in updating in aircraft carrier target identification have been solved, achieving a leap in capabilities from target discovery to intent assessment, and providing reliable and efficient situational awareness and threat assessment.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-01-30
- Publication Date
- 2026-04-14
AI Technical Summary
Existing technologies lack knowledge-driven deep correlation reasoning and intent understanding capabilities in aircraft carrier target identification and tracking, making it impossible to move from "target discovery" to "intent judgment." Furthermore, they suffer from problems such as uncontrollable reasoning, uninterpretable identification results, and difficulty in updating.
By acquiring multi-source data, standardizing processing, fusion of collaborative detection and localization, knowledge-enhanced identity verification, and context-driven state assessment, combined with knowledge graphs for feature alignment and logical reasoning, a reliable, efficient, and interpretable target detection and state reasoning scheme is constructed.
It enables refined and intelligent situational awareness and threat assessment of large maritime platforms such as aircraft carriers, enhances the depth and foresight of situational understanding, ensures the reliability and interpretability of the reasoning process, meets the requirements of efficient near real-time response in modern battlefields, and reduces the complexity of system iteration and knowledge updates.
Smart Images

Figure CN121859062A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of target recognition and tracking technology, and in particular relates to a knowledge-enhanced method and system for aircraft carrier visual target detection and state reasoning. Background Technology
[0002] In the context of joint all-domain operations, large maritime combat platforms such as aircraft carriers are highly mobile and have a wide deterrent range. Their accurate identification and status assessment are the core of battlefield situational awareness and threat assessment. Current aircraft carrier target identification and tracking technologies lack knowledge-driven deep correlation reasoning and intent understanding capabilities. Existing systems have failed to effectively build and utilize multi-dimensional knowledge associations such as tactics, formations, and functions among aircraft carrier target entities. It is difficult to extract systematic behavioral patterns and combat intentions from discrete spatiotemporal point information. The lack of a domain knowledge-based reasoning framework limits the depth and foresight of situational understanding and prevents the leap from "target discovery" to "intent assessment".
[0003] In the prior art, Chinese invention patent publication number CN106342322B discloses a carrier formation identification method based on confidence rule base reasoning. This method uses the traditional confidence rule base approach for carrier formation reasoning and identification, without considering multi-source heterogeneous intelligence information such as satellites and electronic reconnaissance. Chinese invention patent application publication number CN115393790A discloses a GPU-based all-weather satellite autonomous monitoring system for sea and airspace, and Chinese invention patent application publication number CN119785006A discloses a maritime target detection method based on background information suppression. Both are based on computer vision target detection models for identification, which not only fail to integrate information such as radar, but can only identify the appearance category of the target and cannot obtain the target's identity attributes.
[0004] In the prior art, the target recognition system and method based on large language models and knowledge graphs disclosed in Chinese invention patent application publication number CN120804835A can infer the target's intent. However, this solution still has the following problems: (1) Uncontrollability of the reasoning process and the risk of “black box”: The reasoning process of the large language model is a probabilistic “black box” based on massive corpus training. Although it can make amazing associations and inferences, its internal logic is difficult to trace and guarantee precisely. In fields with high reliability and high security requirements such as military and intelligence, a conclusion that cannot be fully explained and may produce “illusions” (fabricated facts or wrong associations) is unacceptable. (2) Challenges of real-time performance and resource consumption: The reasoning of large language models (especially at the level of hundreds of billions of parameters) requires huge computing resources and time, which is fundamentally contradictory to the requirement of "near real-time" response in modern battlefields. Each reasoning requires inputting a large amount of context (feature package + knowledge graph retrieval fragments) into the model and performing complex attention calculations. (3) Complex knowledge updates and system iterations: The system's capabilities are highly dependent on the large language model. When new tactics, new equipment, or the need to correct reasoning logic emerge, simply updating the knowledge graph may not be enough, because the inherent "worldview" of the large language model may not be compatible with new knowledge. At this time, it may be necessary to readjust or retrain the large language model, which is a very costly, long-term process that may introduce new instabilities.
[0005] Therefore, it is necessary to provide a knowledge-enhanced method and system for aircraft carrier visual target detection and state reasoning to solve the above-mentioned technical problems. Summary of the Invention
[0006] The main objective of this invention is to provide a knowledge-enhanced method for aircraft carrier visual target detection and state reasoning, which solves the core pain points of existing technologies such as uncontrollable reasoning, uninterpretable recognition results, and difficulty in updating. It provides a reliable, efficient, interpretable, and iterative knowledge-enhanced target detection and state reasoning scheme, which is particularly suitable for the urgent need for refined and intelligent situational awareness and threat assessment of key targets of large maritime platforms such as aircraft carriers in the context of joint all-domain operations.
[0007] This invention achieves the above objective through the following technical solution: a knowledge-enhanced method for aircraft carrier visual target detection and state reasoning, comprising the following steps: Step S1: Multi-source data acquisition: Acquire the operating parameters of multiple shooting platforms and acquire multi-source images captured by the multiple shooting platforms; Step S2, Standardization Processing: Perform differential preprocessing, unified spatiotemporal reference registration, and normalization processing on the multi-source images to output standardized multi-source images; Step S3, Collaborative Detection and Localization Fusion: The standardized multi-source images from the same spatiotemporal scene are input into the corresponding target detection sub-models for detection, recognition, and visual representation extraction. The image coordinates are converted into geographic coordinates based on the operating parameters of the shooting platform. The multi-source detection results from spatiotemporally adjacent locations are then fused, and the fused target category and fusion confidence score are output. and geographical location; Step S4, Knowledge-Enhanced Identity Verification: Based on the target category, geographical location, and visual representation output in Step S3, feature alignment guided by knowledge graph query is performed. The visual detection results are verified through feature matching, grouping relationship verification, and temporal logical reasoning. The specific identity information of the target category and the overall confidence level are output. ; Step S5, Context-Driven State Assessment: Based on the specific identity information of the target category, extract context information from the knowledge graph, perform multi-stage state reasoning for the target category, trigger preset state reasoning rules, and output the structured state assessment result of the target category.
[0008] Furthermore, the various shooting platforms include space-based platforms, air-based platforms, and sea-based platforms; The multi-source images include: Satellite optical images, radar images, and multispectral images captured by the space-based platform; Infrared images, photoelectric images, and airborne radar images captured by the airborne platform; The images captured by the sea-based platform include shipborne radar images and shipborne optoelectronic images.
[0009] Furthermore, step S3 includes the following steps: Step S31: Input the multi-source normalized images into the corresponding target detection sub-models. After processing by multiple target detection sub-models, mark the target bounding boxes on the multi-source normalized images, extract the visual representations within the target bounding boxes, and output the image positions of the target bounding boxes on the multi-source normalized images. Determine the suspected category of the target within the target bounding box, and output each suspected category and its corresponding original confidence score. ; Step S32: Based on the operating parameters of the shooting platform, unify the image locations of each suspected category into the same geographic coordinate system and output the geographic location of each suspected category; Step S33: Calculate the initial confidence level for detection results that are in the same geographic coordinate system, within the same time period, and are geographically close. The weighted classification process identifies the suspected category with the highest confidence level after weighting as the target category, and outputs the fused target category and its fusion confidence level. Geographical location and visual representation.
[0010] Furthermore, for N suspected categories that are spatiotemporally adjacent in the same geographic coordinate system, the fusion confidence score for category k is... The calculation formula is: ; in, Let be the original confidence level of the i-th detection result. For reliability weights, This is an indicator function.
[0011] Furthermore, step S4 includes the following steps: Step S41: Using the geographical location of the target category output in step S3 as the center, and combining the typical activity radius and observation time window of the target category, retrieve all entity categories known to be deployed within this spatiotemporal range from the knowledge graph, as a candidate entity set; Step 42: Perform feature matching between the visual representation of the target category output in step S3 and the entity categories in the candidate set, and calculate the similarity between the target category and the entity category. And based on similarity The values are used to initially sort the candidate entities; Step 43: Utilize the grouping relationships between entity categories in the knowledge graph to perform reasoning, verify the rationality of candidate entities, and calculate the association support. ; Step 44: Use temporal logic reasoning to analyze the rationality of the candidate entity's motion timeline and eliminate some candidate entities; Step 45: After excluding some candidate entities, comprehensively evaluate the similarity of the remaining candidate entities. and related support and similarity and related support A weighted fusion process is performed, and the candidate entity with the highest overall confidence score is finally matched with the target category. The specific identity information and overall confidence score of the target category are then provided. .
[0012] Furthermore, in step S43, it is checked whether candidate entity e has related reports in the same spatiotemporal region of the knowledge graph; wherein, there are a total of There are several common candidate entities e, among which If there are reports of this species in the region, then the correlation support is high. The calculation formula is: ; In step S44, the reasoning logic of the temporal logic is as follows: if the minimum speed of the target type is greater than the maximum speed recorded in the knowledge graph, then the candidate entity is determined to be impossible to be the currently observed target and is excluded; wherein, the minimum speed of the target type The calculation formula is: ; in, , For the target type, the most recent report location and time. , The current observation location and time for the target type. Given the maximum speed of candidate entity e, if If so, then candidate entity e is excluded;
[0013] In step 45, the overall confidence level of candidate entity e is determined. The calculation formula is: ; in, The weights are adjustable.
[0014] Furthermore, in step S5, the state reasoning rule is a pre-defined, quantifiable conditional rule; each rule R j Corresponding to a conditional function f j (x), where x is the context feature vector extracted from the knowledge graph; when f j When (x) is true, rule R j It is triggered to deduce or support a specific state.
[0015] Furthermore, the method for constructing the knowledge graph is as follows: Step S10: Domain experts define the knowledge graph architecture, clarifying entity types, relationship types, and attribute structures; Step S20: Based on the large language model, using prompt word templates and sample example sets that incorporate architectural constraints, structured knowledge is efficiently extracted from publicly available information intelligence data to form initial triples; Step S30: Perform semantic vectorization representation and similarity calculation on the extracted entities to achieve entity alignment and fusion across data sources; Step S40: Human experts conduct final verification of the core entities and relationships, and store the confirmed triples into the graph database to form a knowledge graph that can be used for reasoning.
[0016] Furthermore, in step S30, a Transformer-based text encoder is used to map the text descriptions to a unified semantic vector space, calculate the cosine similarity between entity vectors, and when the similarity exceeds a threshold, the system performs entity fusion and integrates their attributes and relationships according to a preset conflict resolution strategy to achieve entity alignment. and their semantic similarity Its vector representation is: ; in, , For the entity vector obtained by the Transformer encoder, when Exceeding the preset threshold At that time, the system performs entity fusion.
[0017] Another objective of this invention is to provide a knowledge-enhanced aircraft carrier visual target detection and state reasoning system, which is used to perform the aforementioned knowledge-enhanced aircraft carrier visual target detection and state reasoning method, comprising: The aircraft carrier knowledge graph construction and management module is used to build, store, and update multi-source intelligence knowledge graphs in a human-machine collaborative manner. The data acquisition and access module is used to collect and classify space-based, air-based, sea-based, and open-source intelligence data; The data processing module is used to process image data; The visual target detection and localization module performs target detection, geographic coordinate transformation, and fusion of multi-source detection results. The knowledge-enhanced verification and identity reasoning module is used to query the knowledge graph, verify the visual results, and infer the specific identity of the target. The aircraft carrier target status assessment and output module is used to assess and output the target's status information based on the aircraft carrier knowledge graph and rule base.
[0018] Compared with existing technologies, the beneficial effects of this invention's knowledge-enhanced aircraft carrier visual target detection and state reasoning method and system are as follows: This solution effectively solves the core pain points of existing technologies, such as uncontrollable reasoning, uninterpretable recognition results, and difficulty in updating. It provides a reliable, efficient, interpretable, and iterative knowledge-enhanced target detection and state reasoning solution, which is particularly suitable for the urgent need for refined and intelligent situational awareness and threat assessment of key targets such as large maritime platforms in the context of joint all-domain operations. Specifically: 1. Enhance the depth and foresight of situational understanding, and achieve a leap in capabilities from "target discovery" to "intent assessment": By constructing and integrating multi-source intelligence knowledge graphs, the discrete aircraft carrier target information detected by vision is associated with deep domain knowledge such as aircraft carrier formation, maritime situation, and historical behavior. This enables reasoning about the systematic behavior patterns and potential combat intentions of targets, overcoming the limitations of existing technologies that only stay at the target identification level and lack in-depth correlation analysis and intent assessment capabilities. 2. Ensure the reliability, interpretability, and controllability of the reasoning process and effectively avoid the risk of "black box": Adopting the framework of "knowledge graph-guided verification + rule-driven reasoning", the core reasoning logic is based on structured knowledge in the fields of aircraft carrier characteristics and aircraft carrier activity range, as well as expert-defined and interpretable reasoning rules. Each step of the reasoning in this method has a clear knowledge basis and logical chain, and the conclusions are reliable, traceable, and auditable, meeting the stringent requirements of high reliability and high security in the military, intelligence and other fields. 3. Achieve efficient near real-time response to meet the needs of dynamic battlefield decision-making: Decompose computationally intensive deep reasoning into efficient knowledge graph retrieval, rule matching, and lightweight logic calculation, avoiding the huge computational overhead and latency caused by large-scale language model reasoning. The system can quickly identify and assess the status of targets, meeting the needs of modern battlefields for "near real-time" situational awareness and response. 4. Reduce the complexity and cost of system iteration and knowledge updates: The system's capabilities mainly rely on a structured knowledge graph and well-defined reasoning rules. When new tactics, new equipment, or corrections to reasoning logic are needed, only additions, deletions, and modifications to the knowledge graph or adjustments to the rule base are required. The update process is intuitive and controllable. This fundamentally solves the problem that updates require retraining, are extremely costly, and are prone to instability when relying on large models, thus giving the system good adaptability and maintainability. 5. Significantly improve the accuracy and robustness of target recognition and identity determination through multi-source heterogeneous data fusion and knowledge-enhanced verification: By fusing multi-view and multi-modal image data from space-based, air-based, and sea-based systems, and through differentiated preprocessing and spatiotemporal registration, complementary information is formed. Prior knowledge in the knowledge graph is used to guide the verification and correlation reasoning of the original visual detection results, which can effectively correct false detections and missed detections of single sensors, and significantly improve the accuracy and confidence of target category determination and specific identity recognition. 6. Construct a closed-loop knowledge system that is collaborative between humans and machines and self-enhancing: The knowledge graph is constructed using a collaborative human-machine model of "expert-defined architecture - large model-assisted extraction - intelligent alignment - manual verification". This model not only leverages the efficiency of machines in processing massive amounts of data, but also incorporates the domain wisdom and quality control of experts. New intelligence generated during system operation can be fed back to the knowledge graph for continuous enrichment and correction, forming a continuously evolving and self-enhancing intelligent intelligence processing closed loop. 7. Output structured and actionable status assessment results to directly support decision-making: The final output is not only the target's identity, but also a structured status assessment that combines the maritime situation, time logic, and rule-based reasoning. This high-level, semantic status information can provide commanders and decision-making systems with direct and clear intelligence support, greatly improving the efficiency of transforming raw data into decision-making basis. Attached Figure Description
[0019] Figure 1 This is a schematic diagram illustrating the steps of a knowledge-enhanced aircraft carrier visual target detection and state reasoning method according to an embodiment of the present invention. Detailed Implementation
[0020] Please refer to Figure 1 This embodiment presents a knowledge-enhanced method for aircraft carrier visual target detection and state reasoning, which includes the following steps: Step S1: Multi-source data acquisition: Acquire the operating parameters of multiple shooting platforms and acquire multi-source images captured by the multiple shooting platforms; Step S2, Standardization Processing: Perform differential preprocessing, unified spatiotemporal reference registration, and normalization processing on the multi-source images to output standardized multi-source images; Step S3, Collaborative Detection and Localization Fusion: The standardized multi-source images from the same spatiotemporal scene are input into the corresponding target detection sub-models for detection, recognition, and visual representation extraction. The image coordinates are converted into geographic coordinates based on the operating parameters of the shooting platform. The multi-source detection results from spatiotemporally adjacent locations are then fused, and the fused target category and fusion confidence score are output. and geographical location; Step S4, Knowledge-Enhanced Identity Verification: Based on the target category, geographical location, and visual representation output in Step S3, feature alignment guided by knowledge graph query is performed. The visual detection results are verified through feature matching, grouping relationship verification, and temporal logical reasoning. The specific identity information of the target category and the overall confidence level are output. ; Step S5, Context-Driven State Assessment: Based on the specific identity information of the target category, extract context information from the knowledge graph, perform multi-stage state reasoning for the target category, trigger preset state reasoning rules, and output the structured state assessment result of the target category.
[0021] Steps S1 to S5 are described in detail below: Regarding step S1, multi-source data acquisition: acquire the operating parameters of multiple shooting platforms and acquire multi-source images captured by multiple shooting platforms.
[0022] Specifically, the multiple imaging platforms include space-based, air-based, and sea-based platforms. Space-based platforms include optical imaging reconnaissance satellites, synthetic aperture radar satellites, and electronic reconnaissance satellites; air-based platforms include reconnaissance aircraft (strategic reconnaissance aircraft, electronic reconnaissance aircraft, etc.), unmanned aerial vehicles (UAVs) (high-altitude long-endurance, tactical, etc.), and early warning aircraft; sea-based platforms include reconnaissance ships, intelligence-gathering ships, surface combat vessels (destroyers, frigates, etc.), and submarines. Specifically, the operating parameters of the imaging platform include spatiotemporal reference parameters, platform attitude parameters, sensor status parameters, and platform motion parameters. Spatiotemporal reference parameters include platform position and timestamps, providing a unified time and space origin for all observation data and forming the basis for data fusion. Platform position refers to precise latitude, longitude, and altitude, while the timestamp refers to the acquired Coordinated Universal Time (UTC). Platform attitude parameters include attitude angles and heading / pointing, used to correct image distortion or signal direction deviation caused by changes in platform attitude. Attitude angles refer to pitch, roll, and yaw angles, while heading / pointing refers to the platform's direction of movement or the sensor's line of sight. Sensor status parameters include sensor operating modes and internal sensor parameters, used for precise geometric correction, radiometric calibration, and feature interpretation to ensure data quality. Sensor operating modes include the imaging mode of synthetic aperture radar (SAR) and the focal length of the optical camera, while internal sensor parameters include camera distortion coefficients and radar wavelength / polarization. Platform motion parameters include velocity vectors, acceleration, and orbit / track data, used for motion compensation, predicting the platform's future position, and calculating Doppler shift.
[0023] Specifically, the multi-source images include: Satellite optical images, radar images, multispectral images, etc., taken from multiple perspectives by the space-based platform; Infrared images, photoelectric images, and airborne radar images taken from multiple perspectives by an airborne platform; Multiple perspectives of shipborne radar images and shipborne optoelectronic images captured by the sea-based platform.
[0024] The shooting platform and its operating parameters may include other platforms or parameters, which will not be listed here. Furthermore, the types of multi-source images correspond to the types of shooting platforms, which will not be listed here either, but can be adjusted according to the actual situation.
[0025] For step S2, standardization processing: perform differential preprocessing, unified spatiotemporal reference registration and normalization processing on the multi-source images to obtain standardized multi-source images.
[0026] Images captured by different imaging platforms at the same time and space are called multi-source images. Because images from different platforms contain different types of interference, differentiated preprocessing is required. For example, optical images may undergo atmospheric correction, geometric fine correction, and color fidelity enhancement to highlight the morphological and textural features of suspected categories; infrared images may undergo non-uniformity correction and temperature calibration to enhance sensitivity to heat sources; and radar images may undergo noise suppression and feature normalization to enhance the ability to express the characteristics of target metal structures and corner reflectors. Other methods can also be used to preprocess different types of images, which will not be listed here. Unified spatiotemporal reference registration refers to unifying the timestamps of all images to the same time system and unifying the geographical locations of all images to the same coordinate system and reference. Normalization processing refers to mapping raw data from different sensors, with different dimensions, and different ranges to a unified, finite numerical range (such as 0~1 or -1~1) through mathematical transformation, or converting it into a standard distribution. All types of images have pixel values with similar numerical ranges and comparable physical meanings. Unifying all information onto a single coordinate plane and a single time axis is the absolute prerequisite for all subsequent intelligent analysis and fusion. It facilitates direct input into subsequent visual object detection models for feature extraction, enabling these models to extract features from images efficiently and stably.
[0027] For step S3, collaborative detection and localization fusion: standardized multi-source images from the same spatiotemporal scene are input into their respective target detection sub-models for detection, recognition, and visual representation extraction. The image coordinates are converted into geographic coordinates using the operating parameters of the imaging platform. The multi-source detection results from spatiotemporally adjacent locations are then fused, and the fused target category and fusion confidence score are output. And geographical location.
[0028] Because images captured by different shooting platforms employ different physical principles, different recognition methods are used during image recognition. Therefore, the visual target detection model includes multiple target detection sub-models, which can correspond to images with different physical principles. Each type of image is configured and trained with a dedicated target detection sub-model. For example, optical images are recognized using YOLOv11, infrared images are recognized using the improved YOLOv8, and radar images are recognized using MDD-YOLOv8. The recognition principles of YOLOv11, the improved YOLOv8, and MDD-YOLOv8 are existing technologies and will not be elaborated here.
[0029] Step S3 includes the following detailed steps: Step S31: Input the multi-source normalized images into the corresponding target detection sub-models. After processing by multiple target detection sub-models, mark the target bounding boxes on the multi-source normalized images, extract the visual representations within the target bounding boxes, and output the image positions of the target bounding boxes on the multi-source normalized images. Determine the suspected category of the target within the target bounding box, and output each suspected category and its corresponding original confidence score. ; Step S32: Based on the operating parameters of the shooting platform, unify the multiple image locations of the target into the same geographic coordinate system and output the geographic location of each suspected category; Step S33: Calculate the initial confidence level for detection results that are in the same geographic coordinate system, within the same time period, and are geographically close. The weighted classification process identifies the suspected category with the highest confidence level after weighting as the target category, and outputs the fused target category and its fusion confidence level. Geographical location and visual representation.
[0030] In step S31, after processing by the object detection sub-model, target bounding boxes are marked on the multi-source normalized images. The principle of marking target bounding boxes is existing technology and will not be described in detail here. Since it is a multi-source normalized image, target bounding boxes are marked on each image. The visual representation within each target bounding box is extracted, the image position of each target bounding box on the image is output, the suspected category of the target within each target bounding box is determined, and each suspected category and its corresponding original confidence score are output. .
[0031] Specifically, visual representations include, but are not limited to, the length of the suspected category. ,contour The dimensions include aspect ratio, deck layout, chimney exhaust status, equipment deployment, and other visual features, which will not be listed here.
[0032] Specifically, suspected categories include, but are not limited to, aircraft carriers, destroyers, frigates, supply ships, merchant ships, and non-ships, and may also include other categories, which are not listed here. The original confidence level of the suspected category will be output. Original confidence level The range is 0 to 1.
[0033] The same object may be captured by different shooting platforms at the same time and from different angles. Therefore, the results of object detection sub-models on images captured by different shooting platforms may differ, such as in category and original confidence. The values may differ. Therefore, the multi-source normalized images are input into the corresponding object detection sub-models. After processing by multiple object detection sub-models, each sub-model will mark the target bounding boxes. The sizes of these bounding boxes may differ. The model will also determine the suspected category and its original confidence level. There may be differences, meaning that multiple suspected categories will be output, and the original confidence score of each suspected category will vary. There may still be differences.
[0034] For example, at the same time and in the same space, for the same object, optical images taken by reconnaissance satellites are standardized to obtain standardized optical images. After being identified by the target detection sub-model (YOLOv11), a target bounding box A is marked on the standardized optical image. The target bounding box A is determined to be of the category of aircraft carrier, and the original confidence level for determining the category as aircraft carrier is given. The confidence level is 0.95. The output image position of the suspected category on the multi-source normalized image is (x=500, y=400, w=150, h=25), where x is the x-axis coordinate (pixel value) of the top-left corner of the target bounding box, representing the number of pixels from the left edge of the target bounding box to the left edge of the image; y is the y-axis coordinate (pixel value) of the top-left corner of the target bounding box, representing the number of pixels from the top edge of the target bounding box to the top edge of the image; w is the width (pixel value) of the target bounding box, representing the span of the target bounding box from left to right; and h is the height (pixel value) of the bounding box. At the same time and space, for the same object, the infrared image captured by the UAV is normalized to obtain a normalized infrared image. After being identified by the target detection sub-model (improved YOLOv8), the target bounding box B is marked on the normalized infrared image. The suspected category within the target bounding box B is determined to be a destroyer, and the original confidence level for determining the category as a destroyer is given. The value is 0.85, and the output image location of the suspected category on the multi-source normalized image is (x=400, y=300, w=100, h=20).
[0035] In step S32, its main function is to convert the image position of the target on the image into a geographical location, so as to find the target's location on Earth.
[0036] In step S33, for N suspected categories that are spatiotemporally adjacent in the same geographic coordinate system, the fusion confidence score of category k is determined. The calculation formula is: ; in, Let be the original confidence level of the i-th detection result. For reliability weights, Let be the indicator function (1 when the detected category is k, 0 otherwise), where the maximum fusion confidence is... The corresponding suspected category is the final target category; that is, when outputting the target category, the corresponding output is the fusion confidence score. .
[0037] For step S4, knowledge-enhanced identity verification: Based on the target category, geographical location, and visual representation output in step S3, feature alignment guided by knowledge graph query is performed. The visual detection results are verified through feature matching, grouping relationship verification, and temporal logical reasoning, outputting specific identity information of the target category and comprehensive confidence level. .
[0038] Specifically, step S4 includes the following steps:
[0039] Step S41: Using the geographical location of the target category output in step S3 as the center, and combining the typical activity radius and observation time window of the target category, retrieve all entity categories known to be deployed within this spatiotemporal range from the knowledge graph as a candidate entity set; for example, if the typical activity radius is 100KM and the observation time window is 72 hours, retrieve all entity categories known to be deployed within this spatiotemporal range from the knowledge graph as a candidate entity set. The value of the typical activity radius and the time of the observation time window are not limited here, and can be adjusted according to the actual situation.
[0040] Step 42: Perform feature matching between the visual representation of the target category output in step S3 and the entity categories in the candidate set, and calculate the similarity between the target category and the entity category. And based on similarity The values are used to initially sort the candidate entities; Step 43: Utilize the grouping relationships between entity categories in the knowledge graph to perform reasoning, verify the rationality of candidate entities, and calculate the association support. ; Step 44: Use temporal logic reasoning to analyze the rationality of the candidate entity's motion timeline and eliminate some candidate entities; Step 45: After excluding some candidate entities, comprehensively evaluate the similarity of the remaining candidate entities. and related support and similarity and related support A weighted fusion process is performed, and the candidate entity with the highest overall confidence score is finally matched with the target category. The specific identity information and overall confidence score of the target category are then provided. .
[0041] In step S42, the length of the target category is... ,contour Key features and specification parameters of candidate entity e in the knowledge graph Compare and calculate similarity. Similarity The similarity range is 0 to 1. The calculation formula is: ; in, , As weight, For scale parameters, This is the contour feature similarity function.
[0042] In step S43, since targets generally appear in the form of aircraft carrier formations, it is checked whether candidate entity e has related reports in the same spatiotemporal region in the knowledge graph. For example, there are a total of There are several common candidate entities e, among which If there are reports of this species in the region, then the correlation support is high. The calculation formula is: ; In step S44, the reasoning logic of the temporal logic is as follows: if the minimum speed of the aircraft carrier target type is greater than the maximum speed recorded in the knowledge graph, then it is determined that the candidate entity cannot be the currently observed target and is excluded; specifically, the minimum speed of the aircraft carrier target type... The calculation formula is: ; like Then, candidate entity e is excluded, where, , For the target type, the most recent report location and time. , The current observation location and time for the aircraft carrier target type. The maximum speed of candidate entity e.
[0043] In step 45, the overall confidence level of candidate entity e is determined. The calculation formula is: ;
[0044] in, Adjustable weights. Select overall confidence level. The highest-ranking candidate entity is selected as the target entity, and its accurate identity and overall confidence level are given. The accurate identity of the target entity, such as the aircraft carrier's serial number, name, and hull number.
[0045] The method for constructing the knowledge graph of multi-source intelligence in "human-machine collaboration" is as follows: Step S10: Domain experts define the knowledge graph architecture, clarifying entity types, relationship types, and attribute structures; Step S20: Based on the large language model, using prompt word templates and sample example sets that incorporate architectural constraints, structured knowledge is efficiently extracted from publicly available information intelligence data to form initial triples; Step S30: Perform semantic vectorization representation and similarity calculation on the extracted entities to achieve entity alignment and fusion across data sources; Step S40: Human experts conduct final verification of the core entities and relationships, and store the confirmed triples into the graph database to form a domain knowledge graph that can be used for reasoning.
[0046] In step S10, the knowledge graph architecture is defined: human experts define the knowledge graph pattern based on the background of multi-source intelligence data. This knowledge graph pattern provides a structured knowledge framework and semantic constraints, and clearly defines the entity types, relation types and attribute structures involved in the knowledge graph pattern. This is equivalent to injecting expert prior knowledge into the domain knowledge to ensure the standardization and professionalism of the knowledge graph pattern construction.
[0047] In step S20, large-scale model information extraction involves using prompts that integrate expert-defined knowledge graph architecture and a small number of examples to guide a large language model to automatically extract structured knowledge fragments from massive amounts of publicly available information intelligence data. This publicly available information intelligence data can be obtained using the imaging platform described in step S1, or from other platforms, such as communication and radar signals intercepted via the electromagnetic spectrum, as well as information from public channels. This information is used for signal analysis, content mining, and intelligence fusion. Public channel information includes radio signals and communication intelligence from electronic reconnaissance satellites, radio signal interception and communication intelligence from reconnaissance aircraft / UAVs, communication and radar signals intercepted by ships / submarines, data from Automatic Identification System (AIS) systems, public news reports, social media information, and so on.
[0048] In step S30, intelligent entity alignment: through semantic vectorization and similarity calculation, descriptions from different data sources that point to the same object in the real world are automatically identified and merged, eliminating information redundancy and conflicts to form a unified knowledge entity. In this embodiment, a Transformer-based text encoder is used to map text descriptions to a unified semantic vector space, calculate the cosine similarity between entity vectors, and when the similarity exceeds a threshold, the system performs entity fusion and integrates their attributes and relationships according to a preset conflict resolution strategy to achieve entity alignment. For example, for two entity descriptions... and their semantic similarity Calculated using its vector representation: ; in, , For the entity vector obtained by the Transformer encoder, when Exceeding the preset threshold At that time, the system performs entity fusion.
[0049] In step S40, the human-machine closed-loop verification is performed: human experts conduct a final quality review of the results automatically processed by the machine, and the correction data generated during the review is fed back to the system to form a dynamic knowledge graph, which is used to continuously optimize the extraction capability of the large model and form a self-reinforcing human-machine collaborative closed loop.
[0050] An example of a knowledge graph construction method: Experts first define the knowledge graph architecture, including entity types such as aircraft carriers, destroyers, ports, and sea areas, as well as relationships such as belonging to, escorting, being located, and heading. Using a large language model, they process AIS historical data, electronic reconnaissance reports, and open-source news in batches using prompt words and a small number of labeled samples, automatically generating triples such as ("CVN-78", "model", "Ford-class") and ("CVN-78", "belonging to", "US Navy"). The system automatically aligns "Ford", "CVN-78", and "USS Gerald R. Ford" mentioned in different reports using semantic vector similarity calculations, merging them into the same entity. Finally, after experts review and confirm key information such as "Ford-class aircraft carrier" and its recent activity reports, the information is imported into the graph database.
[0051] For step S5, context-driven state assessment: Based on the specific identity information of the target category, extract context information from the knowledge graph, perform multi-stage state reasoning for the target category, trigger preset state reasoning rules, and output the structured state assessment result of the target category.
[0052] Specifically, the state reasoning context is extracted from the knowledge graph, including the threat level of the sea area, exercise arrangements, political and diplomatic events, and public reports on the aircraft carrier's status.
[0053] In step S5, the state reasoning rules are pre-defined, quantifiable conditional rules; each rule R j Corresponding to a conditional function f j (x), where x is the context feature vector extracted from the knowledge graph; when f j When (x) is true, rule R j It is triggered to deduce or support a specific state.
[0054] For example, the relevant rule R j Example: (1) If the position change is less than 1 nautical mile for 24 consecutive hours, the "berthing" state is triggered; (2) If the course and speed are stable, the "Sailing" state is triggered; (3) If a candidate entity enters the "exercise area" and carrier-based aircraft take off and land frequently, the "training" state is triggered. (4) If the candidate entity is located in the "hotspot conflict zone" and the electronic display system is activated, the "battle deployment" state is triggered; (5) If a candidate entity is detached from its home port for an extended period of time and patrols in a sensitive area, the “deterrence patrol” status will be triggered.
[0055] This embodiment also provides a knowledge-enhanced aircraft carrier visual target detection and state reasoning system, which is used to implement the aforementioned knowledge-enhanced aircraft carrier visual target detection and state reasoning method, characterized in that it includes: The aircraft carrier knowledge graph construction and management module is used to build, store, and update multi-source intelligence knowledge graphs in a human-machine collaborative manner. The data acquisition and access module is used to collect and classify space-based, air-based, sea-based, and open-source intelligence data; The data processing module is used to process image data; The visual target detection and localization module performs target detection, geographic coordinate transformation, and fusion of multi-source detection results. The knowledge-enhanced verification and identity reasoning module is used to query the knowledge graph, verify the visual results, and infer the specific identity of the target. The aircraft carrier target status assessment and output module is used to assess and output the target's status information based on the aircraft carrier knowledge graph and rule base.
[0056] The above descriptions are merely some embodiments of the present invention. Those skilled in the art can make various modifications and improvements without departing from the inventive concept of the present invention, and these all fall within the scope of protection of the present invention.
Claims
1. A knowledge-enhanced method for aircraft carrier visual target detection and state reasoning, characterized in that: It includes the following steps: Step S1: Multi-source data acquisition: Acquire the operating parameters of multiple shooting platforms and acquire multi-source images captured by the multiple shooting platforms; Step S2, Standardization Processing: Perform differential preprocessing, unified spatiotemporal reference registration, and normalization processing on the multi-source images to output standardized multi-source images; Step S3, Collaborative Detection and Localization Fusion: The standardized multi-source images from the same spatiotemporal scene are input into the corresponding target detection sub-models for detection, recognition, and visual representation extraction. The image coordinates are converted into geographic coordinates based on the operating parameters of the shooting platform. The multi-source detection results from spatiotemporally adjacent locations are then fused, and the fused target category and fusion confidence score are output. and geographical location; Step S4, Knowledge-Enhanced Identity Verification: Based on the target category, geographical location, and visual representation output in Step S3, feature alignment guided by knowledge graph query is performed. The visual detection results are verified through feature matching, grouping relationship verification, and temporal logical reasoning. The specific identity information of the target category and the overall confidence level are output. ; Step S5, Context-Driven State Assessment: Based on the specific identity information of the target category, extract context information from the knowledge graph, perform multi-stage state reasoning for the target category, trigger preset state reasoning rules, and output the structured state assessment result of the target category.
2. The knowledge-enhanced aircraft carrier visual target detection and state reasoning method as described in claim 1, characterized in that: The aforementioned filming platforms include space-based platforms, air-based platforms, and sea-based platforms; The multi-source images include: Satellite optical images, radar images, and multispectral images captured by the space-based platform; Infrared images, photoelectric images, and airborne radar images captured by the airborne platform; The images captured by the sea-based platform include shipborne radar images and shipborne optoelectronic images.
3. The knowledge-enhanced aircraft carrier visual target detection and state reasoning method as described in claim 1, characterized in that: Step S3 includes the following steps: Step S31: Input the multi-source normalized images into the corresponding target detection sub-models. After processing by multiple target detection sub-models, mark the target bounding boxes on the multi-source normalized images, extract the visual representations within the target bounding boxes, and output the image positions of the target bounding boxes on the multi-source normalized images. Determine the suspected category of the target within the target bounding box, and output each suspected category and its corresponding original confidence score. ; Step S32: Based on the operating parameters of the shooting platform, unify the image locations of each suspected category into the same geographic coordinate system and output the geographic location of each suspected category; Step S33: Calculate the initial confidence level for detection results that are in the same geographic coordinate system, within the same time period, and are geographically close. The weighted classification process identifies the suspected category with the highest confidence level after weighting as the target category, and outputs the fused target category and its fusion confidence level. Geographical location and visual representation.
4. The knowledge-enhanced aircraft carrier visual target detection and state reasoning method as described in claim 3, characterized in that: The fusion confidence score for N suspected categories that are spatiotemporally adjacent in the same geographic coordinate system, with category k being the fusion confidence score. The calculation formula is: ; in, Let be the original confidence level of the i-th detection result. For reliability weights, This is an indicator function.
5. The knowledge-enhanced aircraft carrier visual target detection and state reasoning method as described in claim 1, characterized in that: Step S4 includes the following steps: Step S41: Using the geographical location of the target category output in step S3 as the center, and combining the typical activity radius and observation time window of the target category, retrieve all entity categories known to be deployed within this spatiotemporal range from the knowledge graph, as a candidate entity set; Step 42: Perform feature matching between the visual representation of the target category output in step S3 and the entity categories in the candidate set, and calculate the similarity between the target category and the entity category. And based on similarity The values are used to initially sort the candidate entities; Step 43: Utilize the grouping relationships between entity categories in the knowledge graph to perform reasoning, verify the rationality of candidate entities, and calculate the association support. ; Step 44: Use temporal logic reasoning to analyze the rationality of the candidate entity's motion timeline and eliminate some candidate entities; Step 45: After excluding some candidate entities, comprehensively evaluate the similarity of the remaining candidate entities. and related support and similarity and related support A weighted fusion process is performed, and the candidate entity with the highest overall confidence score is finally matched with the target category. The specific identity information and overall confidence score of the target category are then provided. .
6. The knowledge-enhanced aircraft carrier visual target detection and state reasoning method as described in claim 5, characterized in that: In step S43, it is checked whether candidate entity e has related reports in the same spatiotemporal region of the knowledge graph; wherein, there are a total of There are several common candidate entities e, among which If there are reports of this species in the region, then the correlation support is high. The calculation formula is: ; In step S44, the reasoning logic of the temporal logic is as follows: if the minimum speed of the target type is greater than the maximum speed recorded in the knowledge graph, then the candidate entity is determined to be impossible to be the currently observed target and is excluded; wherein, the minimum speed of the target type The calculation formula is: ; in, , For the target type, the most recent report location and time. , The current observation location and time for the target type. Let e be the maximum speed of the candidate entity. If so, then candidate entity e is excluded; In step 45, the overall confidence level of candidate entity e is determined. The calculation formula is: ; in, The weights are adjustable.
7. The knowledge-enhanced aircraft carrier visual target detection and state reasoning method as described in claim 1, characterized in that: In step S5, the state reasoning rule is a pre-defined, quantifiable conditional rule; each rule R j Corresponding to a conditional function f j (x), where x is the context feature vector extracted from the knowledge graph; when f j When (x) is true, rule R j It is triggered to deduce or support a specific state.
8. The knowledge-enhanced aircraft carrier visual target detection and state reasoning method as described in claim 1, characterized in that: The method for constructing the knowledge graph is as follows: Step S10: Domain experts define the knowledge graph architecture, clarifying entity types, relationship types, and attribute structures; Step S20: Based on the large language model, using prompt word templates and sample example sets that incorporate architectural constraints, structured knowledge is efficiently extracted from publicly available intelligence data to form initial triples; Step S30: Perform semantic vectorization representation and similarity calculation on the extracted entities to achieve entity alignment and fusion across data sources; Step S40: Human experts conduct final verification of the core entities and relationships, and store the confirmed triples into the graph database to form a knowledge graph that can be used for reasoning.
9. The knowledge-enhanced aircraft carrier visual target detection and state reasoning method as described in claim 8, characterized in that: In step S30, a Transformer-based text encoder is used to map the text descriptions to a unified semantic vector space. The cosine similarity between entity vectors is calculated. When the similarity exceeds a threshold, the system performs entity fusion and integrates their attributes and relationships according to a preset conflict resolution strategy to achieve entity alignment. For two entity descriptions... and their semantic similarity Its vector representation is: ; in, , For the entity vector obtained by the Transformer encoder, when Exceeding the preset threshold At that time, the system performs entity fusion.
10. A knowledge-enhanced aircraft carrier visual target detection and state reasoning system, characterized in that: It is used to complete the knowledge-enhanced aircraft carrier visual target detection and state reasoning method according to any one of claims 1 to 9, which includes: The aircraft carrier knowledge graph construction and management module is used to build, store, and update multi-source intelligence knowledge graphs in a human-machine collaborative manner. The data acquisition and access module is used to collect and classify space-based, air-based, sea-based, and open-source intelligence data; The data processing module is used to process image data; The visual target detection and localization module performs target detection, geographic coordinate transformation, and fusion of multi-source detection results. The knowledge-enhanced verification and identity reasoning module is used to query the knowledge graph, verify the visual results, and infer the specific identity of the target. The aircraft carrier target status assessment and output module is used to assess and output the target's status information based on the aircraft carrier knowledge graph and rule base.
Citation Information
Patent Citations
A carrier battle group identification method based on confidence rule base reasoning
CN106342322B
Sea-airspace all-weather satellite autonomous monitoring system based on GPU
CN115393790A
A method for detecting marine targets based on background information suppression
CN119785006A
Target recognition system and method based on large language model and knowledge graph
CN120804835A