Radar display control picture analysis method and device
By acquiring and segmenting the displayed images from the radar display and control terminal, using a visual language model for multimodal fusion analysis, and supporting operator interaction, the shortcomings of existing radar target recognition methods in multi-source information integration and interpretability are solved, thereby improving the accuracy and efficiency of radar target recognition.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- BAIYANG TIMES (BEIJING) TECH CO LTD
- Filing Date
- 2026-04-10
- Publication Date
- 2026-05-12
AI Technical Summary
Existing intelligent radar target recognition methods cannot effectively integrate multi-source information, lack natural language understanding and generation capabilities, have poor interpretability in model output, and do not fully utilize radar signal processing principles and target characteristics, resulting in insufficient recognition efficiency and accuracy under complex aerial threat situations.
The system acquires images from the radar display and control terminal, segments them into multiple regions, performs multimodal fusion analysis using a visual language model, and combines task prompts and query text for interactive operation. This supports natural language interaction between the operator and the model, improving information communication and feedback efficiency.
It integrates multi-source information for radar target identification, improves the accuracy and efficiency of identification, enhances the interpretability of model output, helps operators understand and verify judgment logic, and effectively responds to complex aerial threat situations.
Smart Images

Figure CN122018843A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of radar display and control screen analysis technology, and in particular to a radar display and control screen analysis method and apparatus. Background Technology
[0002] In radar display screens, multiple interfaces coexist, including Plan Position Indicator (PPI), Range-Doppler Map (RD), and A-Scope (A-scope). Operators need to observe these interfaces and combine them with track parameters to manually identify radar targets, which demands a high level of professional knowledge and practical experience from the operators. In recent years, deep learning technology has been applied to the field of radar target recognition, with Convolutional Neural Networks (CNNs) being widely used. These CNNs are primarily used for classification and recognition of Synthetic Aperture Radar Images (SAR images), High Resolution Range Profiles (HRRP), and Time-Frequency Maps (TF maps), achieving some success on specific datasets.
[0003] However, its limitations are also significant. First, this technology can only process single-modal data and cannot effectively integrate multi-source information, making it difficult to leverage the complementary advantages between different types of data and proving inadequate in the face of complex and ever-changing real-world situations. Second, the lack of natural language understanding and generation capabilities prevents it from engaging in flexible question-and-answer interactions with operators. In practical scenarios, this significantly reduces the efficiency of information communication and feedback, affecting the rapid and accurate judgment of targets. Third, the model output lacks interpretability; operators struggle to understand its reasoning basis, cannot comprehend or verify the AI's judgment logic, and this reduces the model's credibility and practicality in real-world applications. Finally, general deep learning methods fail to fully utilize expert knowledge in areas such as radar signal processing principles and target characteristics; domain knowledge is difficult to effectively integrate into the model, resulting in insufficient generalization ability and adaptability to different scenarios in practical applications.
[0004] In summary, existing intelligent radar target identification methods face numerous limitations in actual combat environments, with operators still largely relying on traditional manual judgment methods. However, as the aerial threat landscape becomes increasingly complex, this reliance on traditional methods proves insufficient, necessitating a more effective solution. Summary of the Invention
[0005] To address the aforementioned problems, this application provides a radar display and control screen analysis method and apparatus, comprising the following: In a first aspect, this application provides a radar display and control screen analysis method, the method comprising: The radar display and control terminal acquires display images, which include a PPI display area, an RD diagram display area, an A waveform display area, and a track list area. The displayed image is segmented to obtain four independent sub-images corresponding to the PPI display area, RD diagram display area, A waveform display area, and track list area, respectively. The four independent sub-images are processed according to preset rules to obtain the image to be analyzed; The image to be analyzed, the preset task prompts, and the query text are input into a pre-built visual language model for analysis and processing. The image to be analyzed is input into the visual language model in a preset input order. The human-computer interaction interface is used to perform interactive operations on the analysis results output by the visual language model.
[0006] Optionally, the displayed image of the acquisition radar display and control terminal includes: The radar display and control terminal acquires display images according to a preset acquisition sequence and a preset acquisition method. The preset acquisition sequence is real-time acquisition synchronized with the radar scanning cycle or timed acquisition with a set fixed time interval. The preset acquisition method is non-intrusive acquisition or system-integrated acquisition.
[0007] Optionally, the process of processing the four independent sub-images according to preset rules to obtain the image to be analyzed includes: Perform quality inspection on each of the sub-images to determine whether there are any abnormalities such as black screen, frozen screen or missing data. For sub-images with abnormalities, trigger re-acquisition or alarm operation. Normalization processing is performed on qualified sub-images that have completed quality inspection, adjusting the resolution of the qualified sub-images to the input size required by the visual language model, while preserving the original color information of the qualified sub-images.
[0008] Optionally, before inputting the image to be analyzed, the preset task prompts, and the query text into a pre-built visual language model for analysis and processing, the method further includes: The task prompt words are constructed, which include role definition information, task description information, image interpretation information, domain knowledge base information, and output format information. The domain knowledge base information is structured information containing radar target characteristic parameters, radar signal processing principles, and radar tactical interpretation rules.
[0009] Optionally, before inputting the image to be analyzed, the preset task prompts, and the query text into a pre-built visual language model for analysis and processing, the method further includes: Construct a query text, which includes the radar scan cycle number and the text corresponding to the air situation analysis question.
[0010] Optionally, the step of inputting the image to be analyzed, preset task prompts, and query text into a pre-built visual language model for analysis and processing includes: The image to be analyzed is combined in the order of sub-images corresponding to the PPI display area, RD diagram display area, A waveform display area, and track list area, and combined with the task prompt and the query text to form an inference request, which is then input into the visual language model. The visual language model sequentially performs visual encoding, multimodal fusion, and autoregressive generation processes before outputting structured analysis results, and then performs validity verification on the structured analysis results.
[0011] Optionally, the validity verification of the structured analysis results includes: The completeness of fields, the rationality of values, and the logical consistency of the structured analysis results are verified. For structured analysis results that fail verification, re-inference or downgrade processing is triggered. The downgrade processing involves extracting valid fields from the structured analysis results to generate simplified analysis results. The structured analysis results that pass verification are stored in the situation database.
[0012] Optionally, the interactive operation performed on the analysis results output by the visual language model based on the human-computer interaction interface includes: The analysis results are rendered in a visual form on the human-computer interaction interface; the human-computer interaction interface is equipped with a situation analysis panel and a decision suggestion panel. The situation analysis panel is used to display the air situation summary, the identification results of each target and the threat level, and the decision suggestion panel is used to display operation suggestions in the form of a priority list. The human-computer interaction interface also includes an alarm module, which is used to actively pop up reminders for newly emerging threat targets, sudden changes in flight path, and abnormal echoes. Operators can interact with the visual language model using natural language and ask questions or follow-up inquiries about the model's judgment results.
[0013] Optionally, the method further includes: performing track association on the same target using multiple analysis results, performing fusion processing on multiple type judgment results of the same target and updating the confidence of the target type judgment, marking the target's status according to the updated confidence, and triggering anomaly alarms for targets with sudden track changes or contradictory target type judgment results.
[0014] Secondly, this application provides a radar display and control screen analysis device, the device comprising: The acquisition unit is used to acquire the display images of the radar display and control terminal. The display images include a PPI display area, an RD diagram display area, an A waveform display area, and a track list area. The image processing unit is used to perform segmentation processing on the displayed image to obtain four independent sub-images corresponding to the PPI display area, RD diagram display area, A display waveform area and track list area respectively; The image processing unit is also used to process the four independent sub-images according to preset rules to obtain the image to be analyzed; The input unit is used to input the image to be analyzed, the preset task prompt words, and the query text into a pre-built visual language model for analysis and processing. The image to be analyzed is input into the visual language model in a preset input order. The interaction unit is used to perform interactive operations on the analysis results output by the visual language model based on the human-computer interaction interface.
[0015] Optionally, the acquisition unit is specifically used for: The radar display and control terminal acquires display images according to a preset acquisition sequence and a preset acquisition method. The preset acquisition sequence is real-time acquisition synchronized with the radar scanning cycle or timed acquisition with a set fixed time interval. The preset acquisition method is non-intrusive acquisition or system-integrated acquisition.
[0016] Optionally, the image processing unit is specifically used for: Perform quality inspection on each of the sub-images to determine whether there are any abnormalities such as black screen, frozen screen or missing data. For sub-images with abnormalities, trigger re-acquisition or alarm operation. Normalization processing is performed on qualified sub-images that have completed quality inspection, adjusting the resolution of the qualified sub-images to the input size required by the visual language model, while preserving the original color information of the qualified sub-images.
[0017] Optionally, the device further includes a prompt word construction unit for constructing task prompt words. The task prompt words include role definition information, task description information, image interpretation information, domain knowledge base information, and output format information. The domain knowledge base information is structured information containing radar target characteristic parameters, radar signal processing principles, and radar tactical interpretation rules.
[0018] Optionally, the device further includes a query text construction unit for constructing query text, which includes text corresponding to radar scan cycle numbers and air situation analysis questions.
[0019] Optionally, the input unit is specifically used to combine the image to be analyzed in the order of sub-images corresponding to the PPI display area, RD map display area, A waveform display area, and track list area, and combine them with the task prompt words and the query text to form an inference request, and input it into the visual language model; The visual language model sequentially performs visual encoding, multimodal fusion, and autoregressive generation processes before outputting structured analysis results, and then performs validity verification on the structured analysis results.
[0020] Optionally, the validity verification of the structured analysis results includes: The completeness of fields, the rationality of values, and the logical consistency of the structured analysis results are verified. For structured analysis results that fail verification, re-inference or downgrade processing is triggered. The downgrade processing involves extracting valid fields from the structured analysis results to generate simplified analysis results. The structured analysis results that pass verification are stored in the situation database.
[0021] Optionally, the interaction unit is specifically used to render the analysis results in a visual form onto the human-computer interaction interface; The human-computer interaction interface is equipped with a situation analysis panel and a decision suggestion panel. The situation analysis panel is used to display the air situation summary, the identification results of each target and the threat level, and the decision suggestion panel is used to display operation suggestions in the form of a priority list. The human-computer interaction interface also includes an alarm module, which is used to actively pop up reminders for newly emerging threat targets, sudden changes in flight path, and abnormal echoes. Operators can interact with the visual language model using natural language and ask questions or follow-up inquiries about the model's judgment results.
[0022] Optionally, the device further includes a data association unit, which is used to perform track association on the same target using multiple analysis results, perform fusion processing on multiple type judgment results of the same target and update the confidence of the target type judgment, mark the target status according to the updated confidence, and trigger anomaly alarms for targets with sudden track changes or contradictory target type judgment results.
[0023] Thirdly, this application provides an apparatus comprising a memory and a processor, the memory for storing instructions or code, and the processor for executing the instructions or code to cause the apparatus to perform the radar display and control screen analysis method described in any of the implementations of the first aspect.
[0024] Fourthly, this application provides a computer-readable storage medium storing code, wherein when the code is executed, a device running the code implements the radar display and control screen analysis method described in any of the implementations of the first aspect.
[0025] This application provides a method for analyzing radar display and control screens. When executing the method, the display image of the radar display and control terminal is first acquired. The display image includes a PPI display area, an RD diagram display area, an A-waveform display area, and a track list area. Then, the display image is segmented to obtain four independent sub-images corresponding to the PPI display area, RD diagram display area, A-waveform display area, and track list area, respectively. These four independent sub-images are then processed according to preset rules to obtain the image to be analyzed. The image to be analyzed, preset task prompts, and query text are input into a pre-constructed visual language model for analysis. The image to be analyzed is input into the visual language model in a preset input order. Finally, interactive operations are performed on the analysis results output by the visual language model based on a human-computer interaction interface.
[0026] In this way, by acquiring images from radar display and control terminals encompassing multiple regions, segmenting and processing them according to preset rules, and then inputting the processed images, task prompts, and query text into a visual language model, the model can comprehensively analyze information from multiple regions, effectively fusing visual and linguistic information. This achieves a breakthrough in radar target identification, overcoming the limitations of traditional single-modal data processing and enabling the integration of multi-source information. Simultaneously, interactive operations based on a human-computer interface facilitate flexible interaction between the operator and the model, improving information communication and feedback efficiency. This enhances the accuracy and efficiency of radar target identification, strengthens the interpretability of the model output, helps operators better understand and verify judgment logic, effectively responds to complex aerial threat situations, and compensates for the shortcomings of existing intelligent methods. Attached Figure Description
[0027] To more clearly illustrate the technical solutions in this embodiment or the prior art, the drawings used in the description of the embodiment or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0028] Figure 1 A flowchart illustrating a radar display and control screen analysis method provided in this application embodiment; Figure 2 A schematic diagram of a human-computer interaction interface provided in an embodiment of this application; Figure 3This is a schematic diagram of the structure of a radar display and control screen analysis device provided in an embodiment of this application. Detailed Implementation
[0029] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of this application.
[0030] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of related data must comply with the relevant laws, regulations and standards of the relevant countries and regions.
[0031] Figure 1 This is a flowchart illustrating a radar display and control screen analysis method provided in an embodiment of this application. (Combined with...) Figure 1 As shown, the radar display and control screen analysis method provided in this application embodiment may include: S101. Acquire the display image of the radar display and control terminal, wherein the display image includes a PPI display area, an RD diagram display area, an A waveform display area, and a track list area.
[0032] Acquiring the display images of the radar display and control terminal refers to obtaining the real-time display screen on the radar display and control terminal to provide raw visual data for subsequent situation analysis. Among them, PPI display, also known as planar position indicator or B display, is a radar display method that displays the target's azimuth and distance information in polar coordinates. The center of the screen represents the radar position, the radial distance represents the target distance, the angle represents the target azimuth, and the scan lines are refreshed synchronously with the antenna rotation. The RD image is a two-dimensional display image used in radar signal processing. The horizontal axis represents the Doppler frequency channel, corresponding to the radial velocity of the target, and the vertical axis represents the range gate, corresponding to the target range. It is output after moving target detection processing and is used for target detection and velocity measurement. A-type waveform, also known as A-type display, is a radar display method that displays the range-amplitude waveform at a selected azimuth. The horizontal axis represents the range, and the vertical axis represents the echo amplitude. It is mainly used to observe signal-level characteristics such as echo signal-to-noise ratio, clutter environment, and multi-target resolution. The track list is the area in the radar display and control terminal that displays the target track parameters. The track refers to the target motion trajectory formed by associating and filtering the points of multiple scanning cycles in the radar data processing system. It includes the target's position, speed, heading and other status information, and is an important basis for target identification and threat assessment.
[0033] In this step, the displayed images of the radar display and control terminal are acquired according to a preset acquisition sequence and a preset acquisition method. The preset acquisition sequence refers to determining the acquisition time pattern of the radar display and control screen, which can be either real-time acquisition synchronized with the radar scanning cycle or timed acquisition with a fixed time interval. In one implementation of this application, the fixed time interval can be set to acquire images every 1-4 seconds. The preset acquisition method refers to the technical method of acquiring the radar display and control screen, which can be divided into non-intrusive acquisition or system-integrated acquisition. There are two schemes: Scheme A is non-intrusive acquisition, which involves acquiring the video signal output from the display and control terminal through an HDMI / DVI acquisition card without modifying the existing radar hardware and software system. This is suitable for the rapid intelligent transformation of existing radar equipment, with a short implementation cycle, low transformation cost, and no impact on the original system. Scheme B is system-integrated acquisition, which involves adding a screenshot interface to the radar display and control software to directly acquire image data from the system frame buffer. This has lower acquisition latency and higher data transmission efficiency, and is suitable for system-level integration design of new radar equipment.
[0034] By implementing this step, efficient and stable acquisition of radar display and control images can be achieved. The two acquisition methods can be adapted to different radar equipment scenarios, taking into account the needs of upgrading existing equipment and integrating new equipment. At the same time, complete display and control images containing multiple areas are acquired, providing comprehensive raw data for subsequent multimodal information fusion analysis, avoiding analysis errors caused by the lack of image information in a single area, and improving the comprehensiveness of subsequent situation analysis.
[0035] S102. Perform segmentation processing on the displayed image to obtain four independent sub-images corresponding to the PPI display area, RD diagram display area, A waveform display area, and track list area, respectively.
[0036] Image segmentation refers to dividing the complete display image into independent sub-images corresponding to the PPI display area, RD diagram display area, A-waveform area, and track list area according to the preset coordinates of each display area of the radar display and control terminal. This facilitates subsequent targeted processing and analysis of each area. Preset rules refer to the relevant specifications for quality inspection and normalization processing performed on the sub-images to ensure they meet the input requirements of the visual language model. The image to be analyzed refers to standardized image data that, after segmentation, quality inspection, and normalization, can be directly input into the visual language model for inference and analysis.
[0037] S103. The four independent sub-images are processed according to preset rules to obtain the image to be analyzed.
[0038] Processing the four independent sub-images according to preset rules is the core step in achieving standardization of radar display and control screens and adapting to the input requirements of visual language models. The preset rules include two core steps: quality detection and normalization processing, which aim to eliminate invalid image data, unify the image input format, and maximize the retention of feature information in radar images.
[0039] Quality inspection refers to the process of determining whether a sub-image has effective analytical value through image feature analysis. Specifically, the quality of each of the four sub-images obtained from the segmentation—the PPI display area, the RD graph display area, the A-waveform display area, and the track list area—is inspected to determine whether there are any anomalies such as black screens, frozen images, or missing data. For black screen anomalies, the system calculates the average grayscale value of the sub-images. If the average grayscale value is lower than a preset threshold (e.g., 10), it is determined to be a black screen. For image freeze anomalies, the system compares the pixel variance of two consecutive sub-image frames. If the variance is lower than a preset threshold (e.g., 5), it is determined to be an image freeze. For data missing anomalies, the system uses image semantic recognition. If there are no scan lines in the PPI area, no Doppler peaks in the RD map, no waveform curves in the A-display, or no numerical information in the track list, it is determined to be data missing. For sub-images determined to have the above anomalies, the system will automatically trigger a re-acquisition command and re-execute the image acquisition step S101. If three consecutive acquisitions are still unsuccessful, an alarm will be triggered to remind the operator to check for faults in the display and control terminal or acquisition equipment. This processing method can effectively prevent invalid image data from entering the subsequent analysis stage, ensure the validity of the model input data, and improve the reliability of the analysis results.
[0040] After quality inspection, qualified sub-images need to be normalized. Normalization refers to converting sub-images with different resolutions and display formats into standardized images that meet the input size requirements of the Visual Language Model (VLM). In this embodiment, the input size required by the Visual Language Model is 448×448 pixels or 672×672 pixels, which can be selected based on the model's inference efficiency and accuracy requirements: 448×448 pixels is chosen if real-time performance is prioritized, and 672×672 pixels is chosen if analysis accuracy is prioritized. During resolution adjustment, a bilinear interpolation algorithm is used to ensure smooth pixel transitions and avoid edge distortion, while fully preserving the original color information of the sub-images. That is, the pseudo-color encoding of the radar display screen contains key feature information, such as red echoes representing high-threat targets, blue echoes representing clutter, and green echoes representing friendly targets. Preserving the original color information ensures that the Visual Language Model can parse the core features such as echo type and threat level, avoiding the loss of feature information due to color loss. The four sub-images after normalization are the images to be analyzed that can be input into the visual language model. This processing method makes the image data of different types of radar display and control terminals form a unified standard, adapts to the input requirements of the visual language model, and ensures the stability and consistency of model inference.
[0041] S104. Input the image to be analyzed, the preset task prompts and query text into a pre-built visual language model for analysis and processing.
[0042] Before inputting the image to be analyzed into the visual language model, it is necessary to construct the task prompts and query text. This step is the key to injecting radar expertise into the model, clarifying the analysis task, and ensuring that the model outputs structured results that meet the requirements of radar tactical interpretation.
[0043] 1. Task prompt word construction: Task prompts are structured text instructions that guide visual language models to perform radar situational analysis. Their core function is to define the model's role, clarify the analysis task, standardize interpretation rules and output format, and compensate for the knowledge gaps of general visual language models in the radar field. In this embodiment, task prompts contain five types of core information: Role definition information: The model is clearly defined as "a senior radar situation analysis expert, proficient in radar signal processing, target identification and tactical interpretation", for example, "You are a senior analyst with 20 years of radar operation experience, and you need to complete target identification, threat assessment and operation suggestion output based on the radar display screen"; Task Description Information: This section describes the specific tasks that the model needs to complete, namely, analyzing the current air situation based on the input radar display, determining the possible types of each batch of targets, assessing the threat level, and providing operational suggestions. Image interpretation information: guides the model to analyze images of each region according to fixed logic, such as "PPI region focuses on analyzing target azimuth, distance, and number; RD map focuses on analyzing target radial velocity and Doppler characteristics; A display focuses on analyzing echo amplitude and signal-to-noise ratio; track list focuses on extracting parameters such as RCS, speed, and heading"; Domain Knowledge Base Information: The knowledge base is organized in a structured table format, containing parameters such as RCS range, speed range, trajectory characteristics, and threat level for various typical targets. Specifically, it includes: fighter jet targets (A-generation aircraft RCS 3-5m², speed 280-450m / s; B-generation stealth aircraft RCS 0.0001-0.01m², speed 250-400m / s), missile targets (subsonic cruise missile RCS 0.1-0.3m², speed 238-272m / s; supersonic anti-ship missile RCS 0.3-0.5m², speed 680-900m / s), and drone targets (small drones RCS 0.01-0.05m², speed 25-55m / s; individual drone swarms RCS 0.001-0.01m², speed 50-80m / s), etc. The knowledge base also includes explanations of signal processing principles, such as the horizontal axis of the RD graph after MTD processing being the Doppler channel and the vertical axis being the distance gate, and the vertical axis amplitude of the A-display reflecting the echo intensity, etc. Output format information: The model is required to output in JSON format, including four required fields: "Air Situation Summary", "Target Identification Result", "Threat Level", and "Operational Recommendations". For example, "Output format: {'Air Situation Summary':'','Target Identification':[{'Batch Number':'','Type':'','Confidence Level':''}],'Threat Level':[{'Batch Number':'','Level':''}],'Operational Recommendations':[{'Priority':'','Recommendation':''}]}".
[0044] 2. Query text construction: The query text is a specific analysis requirement text for the current scanning cycle, used to convey personalized analysis instructions. In this embodiment, the query text at least includes the radar scanning cycle number and the air situation analysis question, such as "Scanning cycle 15, please analyze the current airspace target distribution, identify each target type and assess the threat level, and provide operational suggestions." At the same time, it can be supplemented with at least one of the following: radar operating mode information (such as "the current radar is in scan-while-follow mode"), friendly forces activity information (such as "the 120°-150° azimuth is friendly forces training airspace"), and intelligence notification information (such as "cruise missile harassment may occur in the current airspace"). The supplementary information can help the model optimize the analysis results in combination with actual combat scenarios. For example, combining friendly forces activity information can avoid misjudging friendly targets as threat targets and improve the identification accuracy.
[0045] 3. Visual language model inference and result verification: The image to be analyzed is combined in a preset order: PPI display area, RD map display area, A waveform display area, and track list area (this order is consistent with the input logic during model training to ensure the stability of feature parsing). This combination, along with the task prompts and query text, forms the inference request, which is then input into a pre-built visual language model. This model is a multimodal large-scale model fine-tuned based on a radar domain dataset, possessing cross-modal understanding capabilities for both images and text. Optimal models include Qwen2-VL-7B and InternVL2-8B, open-source models that perform well in multimodal understanding tasks and support Chinese. Model inference uses domain-fine-tuned weights. The fine-tuning process uses a constructed radar situation analysis dataset, which includes four views of the radar display and corresponding expert annotations (target type, threat level, inference basis, and operational suggestions). Supervised fine-tuning (SFT) is used, keeping the visual encoder parameters frozen or updated with a low learning rate. The main focus is on optimizing the parameters of the language model, enabling the model to learn to understand the semantics of radar professional images and perform inferences according to expert logic. Inference deployment employs quantization acceleration technology, quantizing model parameters from FP16 to INT8 precision. Using vLLM or TensorRT-LLM inference frameworks, it achieves an end-to-end inference latency of 2-3 seconds on a single RTX4090 GPU, meeting real-time requirements.
[0046] The steps for visual language model analysis and processing are as follows: Visual encoding: The four images to be analyzed are converted into high-dimensional visual feature vectors respectively, and the PPI track contour, the Doppler peak position of the RD map, the echo waveform features of the A display, and the numerical features of the track list are extracted. Multimodal fusion: The visual feature vector is aligned and fused with the textual semantic features of task prompts and query text across modalities to form a unified feature vector that includes radar image features and professional knowledge; Autoregressive generation: Based on fusion features, it progressively generates structured analysis results that conform to a preset format, such as {'Air situation summary':'One batch of targets was detected in the current airspace, bearing 135°, distance 80km','Target identification':[{'Batch number':'T-001','Type':'Generation A fighter jet','Confidence level':0.92}],'Threat level':[{'Batch number':'T-001','Level':2}],'Operational recommendations':[{'Priority':1,'Recommendation':'Adjust the radar scanning sector to 130°-140° and continue to track the target'}]}.
[0047] After generating the structured analysis results, their validity needs to be validated. Validation dimensions include: Field completeness: Check whether the required fields such as air situation summary, target identification, threat level, and operation suggestions are included. If any field is missing, the verification will fail. Numerical reasonableness: Check whether the value is within the preset range, such as the confidence level must be between 0 and 1, and the threat level must be an integer between 1 and 4. If it is outside the range, the verification will be deemed unsuccessful. Logical consistency: Check for contradictions between fields. For example, the priority of the operation suggestion corresponding to a target with a threat level of 1 should be 1. If the priority is 3, the verification will fail.
[0048] For results that fail validation, a re-inference process (re-inputting the inference request into the model) or downgrade processing is triggered (extracting valid fields to generate simplified results, such as retaining only basic information like target batch number and location). Validated structured analysis results are stored in a situational awareness database, which stores valid analysis results from each scan cycle, providing data support for subsequent multi-cycle track association. The validity validation process filters out invalid or erroneous data from the model output, ensuring the reliability of the analysis results stored in the database and preventing erroneous data from interfering with subsequent decision-making.
[0049] S105. Perform interactive operations on the analysis results output by the visual language model based on the human-computer interaction interface.
[0050] Presenting the validated structured analysis results to the operator through a human-computer interaction interface, and supporting natural language interaction, is the core link to achieving human-computer collaborative judgment and alleviating the operator's cognitive load.
[0051] Figure 2 This is a schematic diagram of a human-computer interaction interface provided in an embodiment of this application. The interface uses a multi-window split-screen display, including a PPI display area that displays the target position and track in polar coordinates; a range-Doppler graph (RD graph) and an A-waveform display area that display the echo characteristics after signal processing; a track situation table that displays track parameters such as "batch number, bearing, distance, speed, heading, attributes, and threat" in a structured list, intuitively presenting the target distribution and characteristics; and below are a VLM (Visual Language Model) situation analysis output panel and a VLM auxiliary decision suggestion panel. Figure 2 This is merely an illustrative display of the interface layout. In actual applications, different interface states can be presented to correspond to typical air situation scenarios such as "airspace clearance", "single-batch attack", "multiple-batch saturation", and "stealth penetration".
[0052] For example, in the "airspace clearance" scenario, the interface displays the initial situation assessment status of the system, the Plane Position Indicator (PPI) shows no obvious target echoes in the display area, and the visual language model situation analysis panel outputs a basic interpretation of the radar's working status, indicating that the system can understand the various components of the radar display and control interface and perform continuous monitoring. In the "single-batch attack" scenario, the PPI display area shows target echo points marked as batch T-001. The track list on the right displays parameters such as the target's range, speed, and radar cross section (RCS). The VLM situation analysis panel completes a preliminary target type judgment based on RCS features and motion parameters combined with the domain knowledge base. The VLM auxiliary decision suggestion panel provides specific operational suggestions such as conducting secondary interrogation with the target using an IFF (Identification Friend or Foe) device and continuously observing track changes, fully presenting the end-to-end analysis process from image understanding to decision output. In the "multi-batch saturation" scenario, the PPI display area presents multiple batches of targets marked with different colors, the track list displays detailed parameters of multiple targets, the VLM situation analysis panel completes the overall airspace situation assessment and multi-target type identification, and the VLM auxiliary decision suggestion panel outputs multiple operation suggestions in priority order, using colors such as red, orange, and yellow to indicate threat level and urgency, reflecting the system's ability to process multiple targets in parallel and alleviate the cognitive load of operators. In high-threat scenarios such as "stealth penetration", the PPI display area shows the flight paths of multiple batches of targets. The VLM situation analysis panel identifies the T-002 batch as a supersonic missile and raises the warning level. The VLM auxiliary decision suggestion panel highlights the emergency warning information in red, pushes suggestions to take immediate countermeasures, and provides auxiliary prompts to continuously observe other batches, demonstrating the system's real-time alarm and threat assessment capabilities.
[0053] The following is combined Figure 2 This section introduces the human-computer interaction interface and the interaction process based on it. First, the analysis results are rendered in a visual format on the human-computer interaction interface: Situation Analysis Panel: This panel integrates three sub-areas: a PPI planar position display, a range-Doppler chart / A-type display module, and a track situation table. It displays air situation summaries, target identification results, and threat levels in a combination of charts and lists. For example, a polar coordinate graph shows the target bearing and range in the PPI area; a two-dimensional heatmap shows the target velocity and range distribution in the range-Doppler chart; a waveform graph shows the echo amplitude characteristics of the A-type display; and a structured list displays track parameters such as "batch number, bearing, range, velocity, heading, attribute, and threat," intuitively presenting target distribution and characteristics. Under different air situation scenarios, this panel can display different states: in the "airspace clear" scenario, only radar operating parameters and clutter information are displayed; in the "multiple batch saturation" scenario, the distribution of multiple target tracks and threat levels is highlighted.
[0054] Decision Recommendation Panel: Operational recommendations are displayed in a priority list format, with priority 1 being the highest level of urgency. For example, "Priority 1: Switch radar tracking mode to single-target tracking and lock onto batch number T-001; Priority 2: Report the target information to the command center; Priority 3: Activate electronic countermeasures equipment for standby." This helps operators quickly focus on core operations. Operational recommendations can be adjusted in different air situation scenarios: for example, in the "stealth penetration" scenario, it is recommended to prioritize increasing the radar pulse repetition frequency and MTI filter strength, while in the "single-batch attack" scenario, it is recommended to focus on continuous target tracking.
[0055] The human-computer interaction interface also includes an alarm module, which is used to monitor and analyze abnormal situations in real time. For newly emerging threat targets, such as the first detection of a Level 1 threat target, sudden changes in the trajectory, such as the target speed jumping from 200m / s to 500m / s, abnormal echoes, such as the appearance of an unidentified strong echo, an active pop-up reminder is triggered. The pop-up contains key target information and alarm level, such as red for the highest level, yellow for the medium level, and blue for the low level. The human-computer interface also supports natural language interaction between the operator and the visual language model. The operator can ask questions or follow-up questions about the model's judgment results, such as "What is the basis for classifying batch number T-001 as a generation A fighter jet?" and "Does the RCS value of this target match the characteristics of a generation A fighter jet?" The model provides targeted explanations based on the current display screen and the domain knowledge base, such as "Based on: 1. The PPI area shows the target's bearing at 135° and distance at 80km, which matches the typical operating range of a generation A fighter jet; 2. The RD diagram shows a radial velocity of 350m / s, matching the cruise speed of a generation A fighter jet; 3. The RCS value in the track list is 4.2m², which is within the typical RCS range (3-5m²) of a generation A fighter jet." The system retains a complete dialogue history and supports multi-round progressive question and answer, improving the interpretability of the model's analysis results and enabling the operator to understand and verify the model's judgment logic.
[0056] To further improve the accuracy and continuity of target identification, one implementation of this application further includes performing track association and confidence updates using the analysis results of multiple radar scan cycles: First, based on the analysis results of multiple scan cycles stored in the situation database, track association is performed on the same target. By matching the target's batch number, bearing, distance, speed and other motion parameters, the analysis results of the same target in different cycles are associated as continuous tracks. For example, the target data of "bearing 135°-138°, distance 80-78km, speed 350m / s" in the 15th, 16th and 17th scan cycles are associated as continuous tracks of batch number T-001, avoiding the randomness of single-cycle data.
[0057] Secondly, the multiple type judgment results of the same target are fused and the confidence is updated. In this embodiment, the weighted average method is used to calculate the fused confidence. The weight is related to the time sequence of the scanning cycle (the recent cycle has a higher weight). For example, the target type is judged as a generation A fighter jet in the 15th cycle (confidence 0.92), the target type is judged as a generation A fighter jet in the 16th cycle (confidence 0.90), and the target type is judged as a generation A fighter jet in the 17th cycle (confidence 0.95). The weight of the recent cycle is set to 0.4, the weight of the intermediate cycle is 0.3, and the weight of the early cycle is 0.3. Then the fused confidence = 0.92×0.3 + 0.90×0.3 + 0.95×0.4 = 0.926.
[0058] The targets are marked with status based on the updated confidence level: a confidence level ≥ 0.9 is marked as "confirmed target", 0.6-0.9 is marked as "target to be confirmed", and < 0.6 is marked as "suspected target". Different statuses are marked with different colors in the situation analysis panel of the human-computer interaction interface to help operators quickly distinguish the certainty of target type.
[0059] Finally, abnormal alarms are triggered for targets with sudden changes in track (such as a speed change rate exceeding a preset threshold of 100m / s / cycle) or contradictory target type judgment results (such as the type changing from "A-generation fighter jet" to "cruise missile" for two consecutive cycles), reminding operators to conduct key checks. For example, a pop-up window will prompt "Batch number T-001 track changes suddenly, speed increases from 350m / s to 500m / s, please verify the target type", further improving the reliability of radar situation assessment.
[0060] The above are some specific implementations of a radar display and control screen analysis method provided in the embodiments of this application. Based on this, the present application also provides a corresponding device. The device provided in the embodiments of this application will be described below from the perspective of functional modularity.
[0061] Figure 3 This is a schematic diagram of a radar display and control screen analysis device provided in an embodiment of this application. (Combined with...) Figure 3 As shown, the radar display and control screen analysis device 300 provided in this application embodiment includes: The acquisition unit 310 is used to acquire the display images of the radar display and control terminal, the display images including a PPI display area, an RD diagram display area, an A waveform display area and a track list area; Image processing unit 320 is used to perform segmentation processing on the displayed image to obtain four independent sub-images corresponding to the PPI display area, RD diagram display area, A display waveform area and track list area respectively; The image processing unit 320 is also used to process the four independent sub-images according to preset rules to obtain the image to be analyzed; The input unit 330 is used to input the image to be analyzed, the preset task prompt words and the query text into a pre-built visual language model for analysis and processing. The image to be analyzed is input into the visual language model in a preset input order. The interaction unit 340 is used to perform interactive operations on the analysis results output by the visual language model based on the human-computer interaction interface.
[0062] In one implementation of this application embodiment, the acquisition unit is specifically used for: The radar display and control terminal acquires display images according to a preset acquisition sequence and a preset acquisition method. The preset acquisition sequence is real-time acquisition synchronized with the radar scanning cycle or timed acquisition with a set fixed time interval. The preset acquisition method is non-intrusive acquisition or system-integrated acquisition.
[0063] In one implementation of this application embodiment, the image processing unit is specifically used for: Perform quality inspection on each of the sub-images to determine whether there are any abnormalities such as black screen, frozen screen or missing data. For sub-images with abnormalities, trigger re-acquisition or alarm operation. Normalization processing is performed on qualified sub-images that have completed quality inspection, adjusting the resolution of the qualified sub-images to the input size required by the visual language model, while preserving the original color information of the qualified sub-images.
[0064] In one implementation of this application, the device further includes a prompt word construction unit for constructing task prompt words. The task prompt words include role definition information, task description information, image interpretation information, domain knowledge base information, and output format information. The domain knowledge base information is structured information containing radar target characteristic parameters, radar signal processing principles, and radar tactical interpretation rules.
[0065] In one implementation of this application, the apparatus further includes a query text construction unit for constructing query text, wherein the query text includes text corresponding to radar scan cycle number and air situation analysis questions.
[0066] In one implementation of this application, the input unit is specifically used to combine the image to be analyzed in the order of sub-images corresponding to the PPI display area, RD map display area, A waveform display area, and track list area, and combine them with the task prompt word and the query text to form an inference request, and input it into the visual language model; The visual language model sequentially performs visual encoding, multimodal fusion, and autoregressive generation processes before outputting structured analysis results, and then performs validity verification on the structured analysis results.
[0067] In one implementation of this application, the validity verification of the structured analysis result includes: The completeness of fields, the rationality of values, and the logical consistency of the structured analysis results are verified. For structured analysis results that fail verification, re-inference or downgrade processing is triggered. The downgrade processing involves extracting valid fields from the structured analysis results to generate simplified analysis results. The structured analysis results that pass verification are stored in the situation database.
[0068] In one implementation of this application, the interaction unit is specifically used to render the analysis results in a visual form onto the human-computer interaction interface; The human-computer interaction interface is equipped with a situation analysis panel and a decision suggestion panel. The situation analysis panel is used to display the air situation summary, the identification results of each target and the threat level, and the decision suggestion panel is used to display operation suggestions in the form of a priority list. The human-computer interaction interface also includes an alarm module, which is used to actively pop up reminders for newly emerging threat targets, sudden changes in flight path, and abnormal echoes. Operators can interact with the visual language model using natural language and ask questions or follow-up inquiries about the model's judgment results.
[0069] In one implementation of this application, the device further includes a data association unit, which is used to perform track association on the same target using multiple analysis results, perform fusion processing on multiple type judgment results of the same target and update the confidence of the target type judgment, mark the target status according to the updated confidence, and trigger an abnormal alarm for targets with sudden track changes or contradictory target type judgment results.
[0070] This application also provides corresponding devices and computer storage media for implementing the solutions provided in this application.
[0071] The device includes a memory and a processor. The memory stores instructions or code, and the processor executes the instructions or code to cause the device to perform the method described in any embodiment of this application.
[0072] The computer storage medium stores code, and when the code is run, the device running the code implements the method described in any embodiment of this application.
[0073] As can be seen from the above description of the embodiments, those skilled in the art can clearly understand that all or part of the steps in the methods of the above embodiments can be implemented by means of software plus a general-purpose hardware platform. Based on this understanding, the technical solution of this application can be embodied in the form of a software product. This computer software product can be stored in a storage medium, such as a read-only memory (ROM) / RAM, magnetic disk, optical disk, etc., including several instructions to cause a computer device (which may be a personal computer, a server, or a network communication device such as a router) to execute the methods described in various embodiments or some parts of the embodiments of this application.
[0074] It is understood that in the specific embodiments of this application, the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved need to obtain user permission or consent when the above embodiments of this application are applied to specific products or technologies, and the collection, use and processing of related data need to comply with the relevant laws, regulations and standards of relevant countries and regions.
[0075] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0076] It should also be noted that the various embodiments in this specification are described in a progressive manner, and the same or similar parts between the various embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, for the device and apparatus embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can be referred to the description of the method embodiments. The device and apparatus embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components indicated as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of the solution in this embodiment according to actual needs. Those skilled in the art can understand and implement this without creative effort.
[0077] The above description is merely one specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. A method for analyzing radar display and control screens, characterized in that, The method includes: The radar display and control terminal acquires display images, which include a PPI display area, an RD diagram display area, an A waveform display area, and a track list area. The displayed image is segmented to obtain four independent sub-images corresponding to the PPI display area, RD diagram display area, A waveform display area, and track list area, respectively. The four independent sub-images are processed according to preset rules to obtain the image to be analyzed; The image to be analyzed, the preset task prompts, and the query text are input into a pre-built visual language model for analysis and processing. The image to be analyzed is input into the visual language model in a preset input order. The human-computer interaction interface is used to perform interactive operations on the analysis results output by the visual language model.
2. The method according to claim 1, characterized in that, The images displayed on the radar display and control terminal include: The radar display and control terminal acquires display images according to a preset acquisition sequence and a preset acquisition method. The preset acquisition sequence is real-time acquisition synchronized with the radar scanning cycle or timed acquisition with a set fixed time interval. The preset acquisition method is non-intrusive acquisition or system-integrated acquisition.
3. The method according to claim 1, characterized in that, The process of processing the four independent sub-images according to preset rules to obtain the image to be analyzed includes: Perform quality inspection on each of the sub-images to determine whether there are any abnormalities such as black screen, frozen screen or missing data. For sub-images with abnormalities, trigger re-acquisition or alarm operation. Normalization processing is performed on qualified sub-images that have completed quality inspection, adjusting the resolution of the qualified sub-images to the input size required by the visual language model, while preserving the original color information of the qualified sub-images.
4. The method according to claim 1, characterized in that, Before inputting the image to be analyzed, the preset task prompts, and the query text into a pre-constructed visual language model for analysis and processing, the method further includes: The task prompt words are constructed, which include role definition information, task description information, image interpretation information, domain knowledge base information, and output format information. The domain knowledge base information is structured information containing radar target characteristic parameters, radar signal processing principles, and radar tactical interpretation rules.
5. The method according to claim 4, characterized in that, Before inputting the image to be analyzed, the preset task prompts, and the query text into a pre-constructed visual language model for analysis and processing, the method further includes: Construct a query text, which includes the radar scan cycle number and the text corresponding to the air situation analysis question.
6. The method according to claim 1, characterized in that, The step of inputting the image to be analyzed, preset task prompts, and query text into a pre-constructed visual language model for analysis and processing includes: The image to be analyzed is combined in the order of sub-images corresponding to the PPI display area, RD diagram display area, A waveform display area, and track list area, and combined with the task prompt and the query text to form an inference request, which is then input into the visual language model. The visual language model sequentially performs visual encoding, multimodal fusion, and autoregressive generation processes before outputting structured analysis results, and then performs validity verification on the structured analysis results.
7. The method according to claim 6, characterized in that, The validity verification of the structured analysis results includes: The completeness of fields, the rationality of values, and the logical consistency of the structured analysis results are verified. For structured analysis results that fail verification, re-inference or downgrade processing is triggered. The downgrade processing involves extracting valid fields from the structured analysis results to generate simplified analysis results. The structured analysis results that pass verification are stored in the situation database.
8. The method according to claim 1, characterized in that, The interactive operation performed on the analysis results output by the visual language model based on the human-computer interaction interface includes: The analysis results are rendered in a visual form on the human-computer interaction interface; the human-computer interaction interface is equipped with a situation analysis panel and a decision suggestion panel. The situation analysis panel is used to display the air situation summary, the identification results of each target and the threat level, and the decision suggestion panel is used to display operation suggestions in the form of a priority list. The human-computer interaction interface also includes an alarm module, which is used to actively pop up reminders for newly emerging threat targets, sudden changes in flight path, and abnormal echoes. Operators can interact with the visual language model using natural language and ask questions or follow-up inquiries about the model's judgment results.
9. The method according to claim 1, characterized in that, The method further includes: The analysis results are used to perform track association on the same target, and the multiple type judgment results of the same target are fused and the confidence of the target type judgment is updated. The target is marked with a status based on the updated confidence. Anomaly alarm is triggered for targets with sudden track changes or contradictory target type judgment results.
10. A radar display and control screen analysis device, characterized in that, The device includes: The acquisition unit is used to acquire the display images of the radar display and control terminal. The display images include a PPI display area, an RD diagram display area, an A waveform display area, and a track list area. The image processing unit is used to perform segmentation processing on the displayed image to obtain four independent sub-images corresponding to the PPI display area, RD diagram display area, A display waveform area and track list area respectively; The image processing unit is also used to process the four independent sub-images according to preset rules to obtain the image to be analyzed; The input unit is used to input the image to be analyzed, the preset task prompt words, and the query text into a pre-built visual language model for analysis and processing. The image to be analyzed is input into the visual language model in a preset input order. The interaction unit is used to perform interactive operations on the analysis results output by the visual language model based on the human-computer interaction interface.