Automatic sighting telescope based on multi-modal sensing and reinforcement learning

CN122590641APending Publication Date: 2026-08-18潘爱民
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610718357.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-05-23
Publication Date
2026-08-18

AI Technical Summary

Technical Problem

[0005]本发明的目的是提供基于多模态传感与强化学习的自动瞄准镜,以解决现有技术中自动瞄准镜在极端环境下感知能力不足的问题

Benefits of technology

[0024]Compared with the prior art, the automatic aiming scope based on multimodal sensing and reinforcement learning provided by the present invention, by setting up a multimodal sensing module including a target motion detection sensor and a target characteristic detection sensor that resist electromagnetic interference, can continuously and stably acquire the target's motion state information and structural weakness information even in electromagnetic interference or low visibility environments, overcoming the defect of single sensors being prone to failure in extreme environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122590641A_ABST
    Figure CN122590641A_ABST
Patent Text Reader

Abstract

The application discloses an automatic sighting telescope based on multi-modal sensing and reinforcement learning, and relates to the field of weapon aiming, and comprises a multi-modal sensing module, the multi-modal sensing module comprising: an anti-electromagnetic interference target motion detection sensor for continuously acquiring the motion state information of a target in an electromagnetic interference or low-visibility environment; a target characteristic detection sensor; an environment detection sensor; and a dynamic decision module in communication connection with the multi-modal sensing module; the automatic sighting telescope realizes stable target sensing in an electromagnetic interference or low-visibility environment by arranging at least two different types of target detection sensors; and the dynamic decision module based on a reinforcement learning algorithm can simultaneously consider three types of parameters, i.e., target motion, environment trajectory and weapon state, and output a multi-dimensional optimized killing strategy, thereby improving the hit rate and killing success rate in an extreme environment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to weapon aiming technology, specifically to an automatic aiming scope based on multimodal sensing and reinforcement learning. Background Technology

[0002] Automatic sights are an important component of modern individual and vehicle-mounted weapon systems, and their performance directly affects shooting accuracy and combat effectiveness. Existing automatic sights suffer from the following technical problems in practical applications:

[0003] In environments with electromagnetic interference (such as anti-drone devices and communication jamming sources) or low visibility (such as sandstorms, rain, snow, and smoke), the target perception capability of traditional optical sights or single infrared thermal imaging sights is significantly reduced, making it easy for targets to be lost or falsely locked onto. At the same time, some existing automatic sights struggle to simultaneously handle multi-dimensional dynamic variables such as the target's dynamic maneuvering behavior (such as speed changes, direction changes, and serpentine maneuvers), changes in the ballistic environment, and the weapon's own performance degradation (such as barrel overheating). This causes the system's output shooting recommendations to deviate from the actual optimal solution when the target is maneuvering at high speed or the weapon's performance is degraded, resulting in a lower hit rate.

[0004] Therefore, how to achieve stable perception of high-speed moving targets by automatic aiming scopes in environments with electromagnetic interference or low visibility, and how to generate optimized kill strategies based on multi-dimensional dynamic variables, are technical problems that urgently need to be solved in this field. Summary of the Invention

[0005] The purpose of this invention is to provide an auto-aiming scope based on multimodal sensing and reinforcement learning to solve the problem of insufficient perception capability of existing auto-aiming scopes in extreme environments.

[0006] To achieve the above objectives, the present invention provides the following technical solution: an automatic aiming scope based on multimodal sensing and reinforcement learning, comprising:

[0007] A multimodal sensing module, comprising:

[0008] An electromagnetic interference-resistant target motion detection sensor is used to continuously acquire target motion status information in electromagnetic interference or low visibility environments.

[0009] Target characteristic detection sensors are used to acquire structural weaknesses or precise attitude information of targets;

[0010] A dynamic decision-making module, which is communicatively connected to the multimodal sensing module, is configured to:

[0011] Receive and fuse multimodal data collected by the target motion detection sensor and the target characteristic detection sensor;

[0012] Based on reinforcement learning algorithms, the system simultaneously processes target motion state, ballistic environment parameters, environmental influence parameters, and weapon state parameters to generate an optimized kill strategy that includes aiming point and firing timing.

[0013] Furthermore, the electromagnetic interference-resistant target motion detection sensor is a millimeter-wave radar with a detection range of not less than 1500 meters.

[0014] Furthermore, the target characteristic detection sensor includes an infrared polarization imaging sensor for penetrating smoke and identifying weak areas in the target's armor.

[0015] Furthermore, the environmental detection sensor is a wind speed and direction measuring instrument, used to measure the wind direction and wind speed that affect aiming and shooting deviation, and outputs the measured wind speed and wind direction parameters to the dynamic decision module as one of the input parameters of the reinforcement learning algorithm.

[0016] Furthermore, the multimodal sensing module also includes a quantum gyroscope, which has a measurement accuracy on the order of 0.001° and is used to measure the attitude of the sight under vibration-resistant conditions.

[0017] Furthermore, the dynamic decision-making module is trained using a federated learning architecture, which is jointly executed by multiple edge devices and a central server, wherein:

[0018] Each edge device independently updates its local model parameters using local battlefield data and uploads the encrypted parameter gradients to the central server.

[0019] The central server aggregates the parameter gradients from each edge terminal to generate a global model, and then distributes the global model to each edge terminal.

[0020] Furthermore, the optimized kill strategy output by the dynamic decision-making module includes: shooting timing and bullet impact point priority ranking.

[0021] Furthermore, in the priority ranking of impact points, the fuel tank has a higher priority than the tires, and the tires have a higher priority than the cockpit.

[0022] Furthermore, the dynamic decision-making module is also configured to automatically switch to an inertial navigation combined with radar path prediction mode when GPS signal rejection occurs, and output a kill plan for high-speed maneuvering targets.

[0023] Furthermore, it also includes an FPGA signal processing chip, which is connected to the multimodal sensing module and is used to preprocess the sensor data at the microsecond level before sending it to the dynamic decision module.

[0024] Compared with the prior art, the automatic aiming scope based on multimodal sensing and reinforcement learning provided by the present invention, by setting up a multimodal sensing module including a target motion detection sensor and a target characteristic detection sensor that resist electromagnetic interference, can continuously and stably acquire the target's motion state information and structural weakness information even in electromagnetic interference or low visibility environments, overcoming the defect of single sensors being prone to failure in extreme environments.

[0025] By setting up a dynamic decision-making module that communicates with the multimodal perception module and configuring it to simultaneously process target motion state, ballistic environment parameters, and weapon state parameters based on reinforcement learning algorithms, it is possible to generate an optimized kill strategy that includes aiming point and firing timing.

[0026] By simultaneously achieving stable perception in extreme environments and comprehensive decision-making based on multidimensional dynamic variables, this invention can output more timely and targeted aiming points and firing opportunities, resulting in higher hit rates and kill success rates under conditions of high-speed target maneuvering and environmental interference. Attached Figure Description

[0027] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in this invention. For those skilled in the art, other drawings can be obtained based on these drawings.

[0028] Figure 1 This is an overall module block diagram provided for an embodiment of the present invention. Detailed Implementation

[0029] To enable those skilled in the art to better understand the technical solution of the present invention, the present invention will be further described in detail below with reference to the accompanying drawings.

[0030] As attached Figure 1 As shown:

[0031] Example 1:

[0032] This invention provides an automatic aiming scope based on multimodal sensing and reinforcement learning, comprising: a scope body, a multimodal sensing module, a dynamic decision-making module, an FPGA signal processing chip, and a power management module.

[0033] Scope body: Made of carbon fiber, with an overall weight controlled below 800 grams. The outer surface of the scope body is coated with an anti-reflective coating and an electromagnetic pulse shielding layer. The rear of the scope body has an eyepiece interface for connecting to the firearm's rail or interface; the front of the scope body has a mounting base for a multimodal sensing module.

[0034] Multimodal sensing module: Installed on the front or side of the mirror body. This module contains at least two of the following sensors:

[0035] (1) Electromagnetic interference resistant target motion detection sensor. In this embodiment, the sensor uses millimeter-wave radar. The millimeter-wave radar operates at a frequency of 24 GHz or 77 GHz, and its detection range is not less than 1500 meters. The millimeter-wave radar can transmit frequency-modulated continuous waves and receive the echo reflected by the target. By calculating the Doppler frequency shift and flight time, the radial velocity, range, and azimuth information of the target can be obtained. The electromagnetic interference resistance characteristics of the millimeter-wave radar are that its operating frequency band does not overlap with common communication interference frequency bands (such as 2.4 GHz and 5.8 GHz), and its signal processing uses chirped modulation technology, which has a high processing gain and can extract target information under low signal-to-noise ratio conditions. Therefore, this sensor is suitable for continuous target tracking in electromagnetic interference environments.

[0036] (2) Target Characteristic Detection Sensor. In this embodiment, the sensor is an infrared polarization imaging sensor. The infrared polarization imaging sensor includes an infrared focal plane array and a polarization filter assembly, with a working wavelength of 3-5μm or 8-12μm. Unlike ordinary infrared thermal imaging, infrared polarization imaging can detect differences in the polarization state of radiation emitted from the target surface. For example, there are significant differences in the polarization characteristics between the surface of metal armor and the surrounding environment, and the polarization characteristics of structural weaknesses such as weld seams and fuel tank caps of the armor are identifiable. This sensor can penetrate some aerosol obstruction in low-visibility environments such as smoke and dust to identify the armor weakness areas of the target.

[0037] (3) Environmental Detection Sensor. In this embodiment, the sensor is an anemometer, used to measure the wind speed and direction data of the on-site environment. The output of the anemometer is connected to the input of the FPGA signal processing chip, and the measured wind speed and wind direction angle are output in the form of digital signals. After the FPGA chip performs digital filtering on the wind speed and direction data, it is packaged into a structured data frame and sent to the dynamic decision module. The reinforcement learning model of the dynamic decision module uses the wind speed and direction data as part of the ballistic environment parameters, and processes it together with other input parameters (target motion state, weapon state, etc.), and finally outputs aiming point offset suggestions including wind deflection correction.

[0038] Furthermore, the multimodal sensing module in this embodiment also includes a quantum gyroscope. Based on the principle of atomic interference, the quantum gyroscope achieves a measurement accuracy on the order of 0.001°. Installed inside the scope body and fixed to the optical aiming axis, the quantum gyroscope is used to measure the instantaneous attitude angle of the scope under vibration conditions (such as firearm firing or vehicle movement), providing a high-precision attitude reference for ballistic calculation.

[0039] FPGA signal processing chip: Installed inside the mirror housing, it is electrically connected to the output terminals of each sensor in the multimodal sensing module. The FPGA chip has pre-defined parallel data preprocessing logic, including: Fast Fourier Transform of radar echoes, noise filtering and non-uniformity correction of infrared images, and moving average filtering of gyroscope data. The FPGA chip can complete the above preprocessing in microseconds and send the structured data to the dynamic decision module via a high-speed serial interface (such as SPI or LVDS).

[0040] Dynamic Decision Module: In this embodiment, the dynamic decision module employs an embedded AI processing chip (such as an NVIDIA Jetson series chip or a domestically produced equivalent AI chip), which is pre-installed with a reinforcement learning model. The input of the dynamic decision module is connected to the output of the FPGA chip, and its output is connected to the display driver circuit within the mirror or an external tactical terminal interface.

[0041] Power management module: Provides the operating voltage required by the above modules, and can be a rechargeable lithium battery or an external tactical power supply.

[0042] The working process of this automatic sight is as follows:

[0043] Step 1: Target Locking and Multimodal Perception Trigger

[0044] After the operator observes the target through the eyepiece or external display, they half-press the trigger switch (or use a voice command) to activate the millimeter-wave radar in continuous wave detection mode. The radar outputs the target's range, radial velocity, and azimuth data at a frame rate of no less than 50Hz. When the radar detects that the target is within its effective range (e.g., within 1000 meters) and has been stably tracking the target for more than 0.5 seconds, the radar sends a lock confirmation signal to the FPGA chip.

[0045] Upon receiving the lock confirmation signal, the FPGA chip automatically triggers the infrared polarization imaging sensor and quantum gyroscope to start high-precision acquisition mode. The infrared polarization imaging sensor completes multi-frame acquisition and polarization calculation of the target area within 0.05 seconds, outputting a thermal image and polarization distribution map of the target. The FPGA chip performs edge detection and connected component analysis on the polarization distribution map to extract candidate weak point regions (such as fuel tank contours and cockpit gaps). The quantum gyroscope synchronously outputs the current three-axis attitude angles of the aiming scope, with a sampling rate of no less than 200Hz.

[0046] Step 2: The dynamic decision-making module generates ballistic schemes.

[0047] The FPGA chip packages the aforementioned multimodal data (radar motion parameters, infrared weak point markers, and gyroscope attitude) into structured data frames and sends them to the dynamic decision-making module.

[0048] The dynamic decision-making module has a pre-built decision model trained based on a reinforcement learning algorithm. The model takes the current target motion state (velocity, acceleration, rate of change of turning angle), ballistic environmental parameters (the dynamic decision-making module obtains environmental parameters through built-in temperature sensors, barometric pressure sensors, and electronic compasses, or wind speed data through an external anemometer), and weapon state parameters (read from the communication interface, such as the number of rounds fired and barrel temperature) as input state vectors.

[0049] The reinforcement learning model employs a deep Q-network or a proximal policy optimization algorithm architecture. The model outputs an action-value map, including the following decision dimensions:

[0050] Firing timing: Recommended window opening time in milliseconds;

[0051] Aiming point offset: The horizontal and vertical offset suggestion in mils in the scope's field of view;

[0052] Impact point priority: sorting of multiple candidate weak point regions.

[0053] In this embodiment, the dynamic decision-making module outputs three sets of candidate ballistic schemes, each with a confidence score.

[0054] Step 3: Solution Selection and Output

[0055] The dynamic decision-making module selects the highest-scoring option as the recommended option based on the confidence score. Simultaneously, the module reads the weapon's remaining lifespan parameters (e.g., if the barrel temperature exceeds a set threshold, the recommended weight for continuous firing is reduced). Finally, the dynamic decision-making module overlays the aiming point offset and weak point markers onto the eyepiece or external display screen, providing a graphical prompt to the operator.

[0056] If the operator confirms the shot, the dynamic decision module records the actual impact point of the shot (through subsequent sensor data or operator input) for offline or online model updates.

[0057] Working principle: Through the heterogeneous redundancy design of multimodal sensors, effective target information can still be acquired even in environments where a single sensor fails (electromagnetic interference, low visibility). By employing end-to-end decision-making through a reinforcement learning model, multiple sequentially performed tasks (range finding, wind measurement, ballistic calculation, target prediction) are integrated into a parallel decision-making process, thereby shortening the time delay from target lock to firing suggestion output. Simultaneously, by incorporating the weapon's own state (such as barrel temperature) into the decision variables, ineffective firing due to weapon performance degradation can be avoided.

[0058] Example 2:

[0059] This embodiment is basically the same as the previous embodiment, except that in this embodiment, the dynamic decision-making module is trained using a federated learning architecture. Specifically, the federated learning architecture is jointly executed by multiple edge devices and a central server, wherein:

[0060] Each edge device independently updates its local model parameters using local battlefield data and uploads the encrypted parameter gradients to the central server.

[0061] The central server aggregates the parameter gradients from each edge terminal to generate a global model, and then distributes the global model to each edge terminal.

[0062] To implement the federated learning architecture, this embodiment adds the following components to the first embodiment:

[0063] (1) Encrypted communication module: installed inside the mirror body and electrically connected to the dynamic decision module. The encrypted communication module uses military-grade encryption algorithms (such as SM4 or AES-256) to encrypt the gradient of model parameters and communicate with the central server through a tactical data link (such as a software radio module).

[0064] (2) Local model cache: A non-volatile storage area is allocated in the storage space of the dynamic decision module to store the global model parameters received from the central server last time, as well as the gradient accumulation generated by the local cumulative update.

[0065] After each shot, the automatic sight in this embodiment records the following data pair: {state vector before firing, actual firing result (hit / miss, bullet impact point deviation)}. This data pair is stored in the local cache as a training sample.

[0066] When the number of locally cached samples reaches a preset threshold (e.g., 50), the dynamic decision-making module starts the local training process from standby mode. The training process uses the accumulated local samples to perform several gradient descent updates on the current model, calculating the gradient change of the model parameters (i.e., the local update gradient). This gradient is encrypted by the encrypted communication module and then uploaded to the central server via the tactical data link.

[0067] The central server collects encrypted gradients from multiple edge devices (i.e., multiple auto-aiming scopes), performs decryption and aggregation operations (e.g., using the FedAvg algorithm, weighted average based on the number of samples from each edge device). After aggregating to obtain the global model update parameters, the central server distributes the global model back to each edge device. Upon receiving the global model, each edge device replaces its local model with it, completing one federated learning cycle.

[0068] Example 3:

[0069] This embodiment is basically the same as the previous embodiment, except that in this embodiment, the dynamic decision module is also configured to automatically switch to the working mode of inertial navigation combined with radar prediction path under the condition of GPS signal rejection, and output a kill plan for high-speed maneuvering targets.

[0070] Specifically, the dynamic decision-making module integrates a GPS signal monitoring submodule. This submodule monitors the output status of the GPS receiver module in real time. If a valid location solution cannot be resolved for more than 3 consecutive seconds or the confidence level is below a threshold, it is determined to be a GPS signal rejection state. The dynamic decision-making module automatically switches its operating mode.

[0071] In the working mode of inertial navigation combined with radar path prediction:

[0072] (1) Inertial navigation reference maintenance: Dead reckoning is performed using quantum gyroscopes and accelerometers (which can be integrated into a multimodal sensing module). Since the accuracy of the quantum gyroscope in this system reaches the order of 0.001°, the cumulative attitude error in a short period of time (e.g., within 60 seconds) is less than 0.1°, which is sufficient to support GPS-free positioning during a single engagement.

[0073] (2) Radar Prediction Path: The millimeter-wave radar continuously tracks the target's motion and outputs a position sequence. The dynamic decision module incorporates a Kalman filter or Long Short-Term Memory (LSTM) model, which predicts the target's trajectory within the next second based on the target's position sequence of the most recent 2 seconds. The prediction output is a position probability distribution with confidence intervals.

[0074] (3) Ballistic calculation fusion: The dynamic decision module spatially registers the self-aiming line of sight calculated by inertial navigation with the future position of the target predicted by radar, and calculates the lead of the aiming point. Since the absolute coordinates of GPS are lacking, this lead is output to the eyepiece in the form of relative angle (mil).

[0075] Taking a scenario where a target is escaping at 120 km / h using a serpentine maneuver: the dynamic decision module identifies the periodic variation pattern of the target's lateral acceleration (serpentine maneuver characteristic) and predicts the target's most likely position at the end of the bullet's flight time using an LSTM model. Simultaneously, due to GPS signal rejection, the module automatically disables any ballistic correction functions that rely on absolute coordinates, relying entirely on relative measurement data.

[0076] The dynamic decision-making module outputs the following kill plan: It suggests that the lateral offset of the aiming point should no longer be expressed as an absolute distance, but as a micrometer scale on the scope reticle; if the front wheel steering axle of the target vehicle is identified as a weak link in maneuverability (tire lateral force limit), it outputs the suggestion of "prioritizing shooting the front wheel steering axle" with a confidence level (e.g., 92%).

[0077] The foregoing has only described certain exemplary embodiments of the present invention by way of illustration. Undoubtedly, those skilled in the art can modify the described embodiments in various ways without departing from the spirit and scope of the present invention. Therefore, the foregoing drawings and descriptions are illustrative in nature and should not be construed as limiting the scope of protection of the claims of the present invention.

Claims

1. An automatic aiming scope based on multimodal sensing and reinforcement learning, characterized in that, include: A multimodal sensing module, comprising: An electromagnetic interference-resistant target motion detection sensor is used to continuously acquire target motion status information in electromagnetic interference or low visibility environments. Target characteristic detection sensors are used to acquire structural weaknesses or precise attitude information of targets; Environmental detection sensors are used to obtain wind direction and wind speed in the field environment; A dynamic decision-making module, which is communicatively connected to the multimodal sensing module, is configured to: Receive and fuse multimodal data collected by the target motion detection sensor, the target characteristic detection sensor, and the environmental detection sensor; Based on reinforcement learning algorithms, the system simultaneously processes target motion state, ballistic environment parameters, environmental influence parameters, and weapon state parameters to generate an optimized kill strategy that includes aiming point and firing timing.

2. The automatic aiming scope based on multimodal sensing and reinforcement learning according to claim 1, characterized in that, The electromagnetic interference-resistant target motion detection sensor is a millimeter-wave radar with a detection range of not less than 1500 meters.

3. The automatic aiming scope based on multimodal sensing and reinforcement learning according to claim 1, characterized in that, The environmental detection sensor is a wind speed and direction measuring instrument, used to measure wind direction and speed that affect aiming and shooting deviation.

4. The automatic aiming scope based on multimodal sensing and reinforcement learning according to claim 1, characterized in that, The target characteristic detection sensor includes an infrared polarization imaging sensor, used to penetrate smoke and identify weak areas in the target's armor.

5. The automatic aiming scope based on multimodal sensing and reinforcement learning according to claim 1, characterized in that, The multimodal sensing module also includes a quantum gyroscope, which has a measurement accuracy on the order of 0.001° and is used to measure the attitude of a sight under vibration-resistant conditions.

6. The automatic aiming scope based on multimodal sensing and reinforcement learning according to claim 1, characterized in that, The dynamic decision-making module is trained using a federated learning architecture, which is jointly executed by multiple edge devices and a central server, wherein: Each edge device independently updates its local model parameters using local battlefield data and uploads the encrypted parameter gradients to the central server. The central server aggregates the parameter gradients from each edge terminal to generate a global model, and then distributes the global model to each edge terminal.

7. The automatic aiming scope based on multimodal sensing and reinforcement learning according to claim 1, characterized in that, The optimized kill strategy output by the dynamic decision-making module includes: shooting timing and bullet impact point priority ranking.

8. The automatic aiming scope based on multimodal sensing and reinforcement learning according to claim 7, characterized in that, In the priority ranking of impact points, the fuel tank has a higher priority than the tires, and the tires have a higher priority than the cockpit.

9. The automatic aiming scope based on multimodal sensing and reinforcement learning according to claim 1, characterized in that, The dynamic decision-making module is also configured to automatically switch to an inertial navigation combined with radar path prediction mode when GPS signal rejection occurs, and output a kill plan for high-speed maneuvering targets.

10. The automatic aiming scope based on multimodal sensing and reinforcement learning according to claim 1, characterized in that, It also includes an FPGA signal processing chip, which is connected to the multimodal sensing module and is used to preprocess the sensor data at the microsecond level before sending it to the dynamic decision module.