Intelligent mine system and method based on magnetism-sound-light fusion perception and reinforcement learning decision
The intelligent mine system, which integrates magnetic-acoustic-optical perception and reinforcement learning decision-making, solves the problems of inaccurate positioning, imprecise interception, and high false alarm rate in the protection system of cross-sea bridge piers, and achieves efficient, low-cost defense and self-repair against underwater threats around the bridge piers.
Patent Information
- Application Number
- CN202511163535.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-20
- Publication Date
- 2025-12-02
AI Technical Summary
Existing protection systems for the piers of cross-sea bridges cannot effectively identify stealthy blasting targets. Traditional interception methods have large errors and high costs, while mine systems have a high false alarm rate and cannot achieve continuous safety protection.
Employing a fusion of magnetic-acoustic-optical sensing technology and a reinforcement learning decision-making mechanism, the system integrates a fluxgate sensor, a CNN voiceprint classifier, an optical fiber vibration calibration module, and a Q-learning decision-making module to achieve millimeter-level localization and dynamic interception of underwater threats around bridge piers, and repairs damage through a biomimetic self-healing structure.
It achieves full-time, multi-dimensional dynamic protection against underwater threats around bridge piers, reduces false alarm rate, improves interception success rate, and has autonomous repair capability, thus reducing operation and maintenance costs.
Smart Images

Figure CN121048447A_ABST
Abstract
Description
Technical Field
[0001] This patent aims to construct an intelligent mine system specifically designed for defending against underwater attacks on the piers of cross-sea bridges. By integrating multi-source sensing technologies such as fluxgate sensing, voiceprint recognition, and fiber optic positioning with a reinforcement learning dynamic decision-making mechanism, it can capture the metallic properties, acoustic characteristics, and three-dimensional spatial coordinates of underwater targets around the piers in real time, and trigger precise explosive interception and self-repair response. It belongs to the field of underwater bridge defense equipment technology. Background Technology
[0002] This invention relates to the field of safety protection technology for cross-sea bridges, and more particularly to an underwater threat protection system for cross-sea bridge piers. Cross-sea bridge piers face an increasingly severe threat of precision underwater blasting; failure of their protection systems can lead to structural damage and even long-term paralysis of the bridge. Existing protection systems suffer from fundamental technical defects: existing sonar monitoring methods have physical blind spots, making it difficult to identify blasting targets approaching the piers; traditional interception methods have fixed detonation modes and large errors, resulting in a success rate of less than 70%, and the defense process is prone to secondary damage such as falling debris; conventional mine systems have a high false alarm rate of >15%, and their single-use design leads to high maintenance costs, making continuous protection of the piers impossible. Therefore, there is an urgent need to develop an innovative protection system to achieve millimeter-level positioning, dynamic decision-making interception, and self-repair of underwater threats around the piers.
[0003] In existing water security technologies, traditional physical fences or single-sensor monitoring schemes have significant limitations: physical fences are difficult to cover vast water areas, while sonar, infrared, or electromagnetic single-mode detection is easily affected by water attenuation characteristics (such as changes in salinity and turbidity), resulting in a high false alarm rate (see patent CN119920045A). Although multimodal sensing networks based on acoustic-optical-electric fusion (as described in patent CN119920045A) have emerged in recent years, achieving dynamic attenuation correction through the collaborative use of acoustic, optical, and electromagnetic sensors, and introducing wave energy spectral density analysis (S(f)) and intrusion probability prediction models (P(x,t)), their core remains at the level of environmental disturbance identification, lacking an active decision-making mechanism for threats. For example, the anomaly measurement formula in patent CN119920045A... While it can improve anomaly detection sensitivity, it lacks a dynamic response strategy, making real-time interception of targets impossible. Meanwhile, underwater target tracking technologies (such as those described in patent CN118566213B) employ an improved YOLOv5 detection network and adaptive noise Kalman filtering (with the state vector modified to...) It optimizes fish trajectory tracking, but its technical objectives are limited to passively analyzing biological behavior and do not involve the active location and attack decision-making of threatening targets.
[0004] This patent proposes an intelligent mine system based on magnetic-acoustic-optical fusion perception and reinforcement learning decision-making. By integrating high-precision positioning sensing, multi-source data fusion engine, biomimetic self-healing structure and group collaborative decision-making mechanism, it achieves all-time-domain and multi-dimensional dynamic protection against underwater threats around the piers of cross-sea bridges. Summary of the Invention
[0005] An intelligent mine system based on magnetic-acoustic-optical fusion perception and reinforcement learning decision-making consists of a fluxgate sensor 1, a CNN acoustic text classifier 2, an optical fiber vibration calibration module 3, a Q-learning decision-making module 4, a titanium alloy-boron carbide ceramic substrate 5, polyurea microcapsules 6, a DFOS positioning module 7, an LSTM trajectory prediction module 8, a reinforcement learning controller 9, a sensor data fusion module 10, and a cooperative communication interface 11.
[0006] The entire system comprises three core subsystems: a magnetic-acoustic-optical fusion sensing subsystem, a reinforcement learning dynamic decision-making subsystem, and a biomimetic self-healing subsystem. Each subsystem is tightly coupled through a data bus and control signal lines. The three core subsystems work together to achieve full-process autonomy from target detection to damage repair.
[0007] The magnetic-acoustic-optical fusion sensing subsystem includes a fluxgate sensor 1, a CNN acoustic signature classifier 2, an optical fiber vibration calibration module 3, a DFOS positioning module 7, and a sensor data fusion module 10. The DFOS positioning module 7 is directly connected to the sensor data fusion module 10 via an optical fiber data bus, outputting the target's three-dimensional coordinates in real time with a positioning accuracy error ≤0.5m. The fluxgate sensor 1 works in conjunction with the DFOS positioning module 7 to capture characteristic parameters (such as ferromagnetic content) of metal targets with a mass ≥500kg, and is connected to the sensor data fusion module 10 via an independent circuit. The CNN acoustic signature classifier 2 is connected to the sensor data fusion module 10 in parallel to analyze the propeller acoustic signature characteristics in the 20Hz-20kHz frequency band. The optical fiber vibration calibration module 3 is connected in series with the input of the CNN acoustic signature classifier 2 to adaptively filter the acoustic sensing signal, eliminating high-frequency vibration noise (frequency band 20-100Hz) induced by water flow, with a noise reduction efficiency ≥90%. The sensor data fusion module 10 integrates multi-source sensing data, realizes spatial-attribute association, and outputs target state vectors, including target category (enemy / friend identification), movement speed (range 0-30 knots), heading angle (0-360 degrees) and threat level (quantization value 0-1);
[0008] The reinforcement learning dynamic decision-making subsystem includes a Q-learning decision module 4, an LSTM trajectory prediction module 8, a reinforcement learning controller 9, and a cooperative communication interface 11. The LSTM trajectory prediction module 8 is connected to the DFOS positioning module 7 via a high-speed data interface and predicts the target's trajectory in the future time period Δt (Δt∈[2s,15s]) based on real-time three-dimensional coordinates. The Q-learning decision module 4 is connected to the LSTM trajectory prediction module 8 via a high-speed data transmission interface and calculates the optimal detonation point, which must meet three conditions: ① Spatial location: the closest point on the target trajectory to the smart mine at a distance ≤ the effective damage radius of the explosion (R=12m); ② Time window: the time t for the target to arrive at this point ≤ 5s; ③ Value judgment: the action value function Q(s,a)>Qmin=8.5, where the state space s=[target velocity, metal mass, threat level, relative distance], and the action space a includes detonation or standby. The reinforcement learning controller 9 is circuitally connected to the LSTM trajectory prediction module 8. When conditions ①②③ are simultaneously met, a detonation command is triggered. The cooperative communication interface 11 is connected to the reinforcement learning controller 9 via a control signal line, executing cross-system cooperative strategies: interacting with the homogeneous intelligent mine array (the main mine sends detonation coordinates, and the auxiliary mine receives and transmits 20kHz interference waves), interacting with the naval shore-based command and control system (encrypted return of target trajectory data, including DFOS coordinates and acoustic signatures), and interacting with marine environmental monitoring buoys (receiving temperature, salinity, and depth data to correct the acoustic propagation model). All modules (Q-learning decision module 4, LSTM trajectory prediction module 8, and reinforcement learning controller 9) are connected to the sensor data fusion module 10 via a data bus to ensure real-time synchronization of decision data.
[0009] The biomimetic self-healing subsystem comprises a titanium alloy-boron carbide ceramic matrix 5 and polyurea microcapsules 6. The titanium alloy-boron carbide ceramic matrix 5 serves as the core support structure, providing impact resistance (compressive strength ≥ 500 MPa). The polyurea microcapsules 6 (50 μm in diameter) are uniformly distributed within the titanium alloy-boron carbide ceramic matrix 5 through embedded encapsulation. Upon impact (damage depth ≥ 2 mm), they automatically rupture to release the repair agent, filling the cracks and voids, and repairing the outer shell of the titanium alloy-boron carbide ceramic matrix 5 through in-situ polymerization. The repair completion time is ≤ 30 seconds.
[0010] The three subsystems are interconnected:
[0011] The output of the magnetic-acoustic-optical fusion sensing subsystem is directly coupled to the Q-learning decision module 4 of the reinforcement learning dynamic decision-making subsystem via a data bus through the sensor data fusion module 10.
[0012] The reinforcement learning controller 9 of the reinforcement learning dynamic decision-making subsystem is connected to the polyurea microcapsule 6 of the biomimetic self-healing subsystem via a control signal line, triggering the repair mechanism.
[0013] The system operates in a continuous monitoring cycle: when there is no threat, the system enters standby mode: the fluxgate sensor 1, fiber optic vibration calibration module 3, and DFOS positioning module 7 operate at 3W power; the CNN voiceprint classifier 2, reinforcement learning controller 9, LSTM trajectory prediction module 8, and cooperative communication interface 11 are in sleep mode with power consumption ≤0.5W; when a threat is activated (threat level >0.7), the reinforcement learning controller 9 prioritizes resource scheduling (response time ≤0.2ms) and performs cooperative interception; after interception, the sensor data fusion module 10 detects the crack depth δ of the titanium alloy-boron carbide ceramic matrix 5; if δ≥2mm, the reinforcement learning controller 9 triggers the rupture of the polyurea microcapsule 6 to release the repair agent; the repair agent actively fills the crack gaps and completes in-situ polymerization and curing within 30 seconds; after curing, the sensor fusion module 10 re-measures the damage depth δ'; if δ'≥0.5mm, the built-in repair agent reserve unit of the polyurea microcapsule 6 is activated to perform secondary repair. This unit independently encapsulates the repair agent library and can support ≥50 rounds of repair operations to ensure the structural integrity of the titanium alloy-boron carbide ceramic matrix 5. Attached Figure Description
[0014] Appendix Figure 1 This is a schematic diagram of the biomimetic self-healing subsystem structure of this patent.
[0015] Appendix Figure 1 Names of winning bids: 5. Titanium alloy-boron carbide ceramic matrix, 6. Polyurea microcapsules.
[0016] Appendix Figure 2 This is a schematic diagram of the magneto-acoustic-optical fusion sensing subsystem.
[0017] Appendix Figure 2 The winning bids are as follows: 1. Fluxgate sensor, 2. CNN voiceprint classifier, 3. Fiber optic vibration calibration module, 7. DFOS positioning module, and 10. Sensor data fusion module.
[0018] Appendix Figure 3 A schematic diagram of the internal workings of the dynamic decision-making subsystem for enhanced learning.
[0019] Appendix Figure 3 The following are the names of the selected components: 4. Q-learning decision module, 6. Polyurea microcapsule, 8. LSTM trajectory prediction module, 9. Reinforcement learning controller, and 11. Cooperative communication interface. Detailed Implementation
[0020] An intelligent mine system based on magnetic-acoustic-optical fusion perception and reinforcement learning decision-making is implemented as follows: The system consists of a fluxgate sensor 1, a CNN acoustic text classifier 2, an optical fiber vibration calibration module 3, a Q-learning decision-making module 4, a titanium alloy-boron carbide ceramic substrate 5, polyurea microcapsules 6, a DFOS positioning module 7, an LSTM trajectory prediction module 8, a reinforcement learning controller 9, a sensor data fusion module 10, and a cooperative communication interface 11.
[0021] The entire system comprises three subsystems: a magneto-acoustic-optical fusion perception subsystem, a reinforcement learning dynamic decision-making subsystem, and a biomimetic self-repair subsystem.
[0022] During deployment, the mine's outer shell uses a titanium alloy-boron carbide ceramic matrix 5 as a frame, with polyurea microcapsules 6 uniformly embedded inside the titanium alloy-boron carbide ceramic matrix 5 (capsule density ≥1000 capsules / cm³). 3 The core electronic unit integrates sensing and decision-making modules and connects to an external network via a collaborative communication interface 11. The system workflow consists of three stages:
[0023] Target identification stage: DFOS positioning module 7 monitors the underwater environment in real time and outputs the three-dimensional coordinates of the target vessel (error ≤ 0.5m); fluxgate sensor 1 synchronously captures the parameters of the metal target (mass ≥ 500kg); CNN acoustic signature classifier 2 analyzes the propeller acoustic signature (20Hz-20kHz frequency band); fiber optic vibration calibration module 3 eliminates water flow noise (signal-to-noise ratio ≥ 20dB after noise reduction); sensor data fusion module 10 integrates the data, outputs the target status, and calculates the threat level based on speed, heading angle, and metal characteristics.
[0024] Dynamic Decision-Making Phase: When the threat level is ≤0.7, the system enters a low-power standby mode: only the fluxgate sensor 1, fiber optic vibration calibration module 3, and DFOS positioning module 7 operate at 3W power; other modules (CNN voiceprint classifier 2, reinforcement learning controller 9, etc.) are in sleep mode (total power consumption ≤5W). When the threat level is >0.7, the system is activated: the LSTM trajectory prediction module 8 predicts the target's future trajectory (Δt=2s-15s) based on the data from the DFOS positioning module 7; the Q-learning decision module 4 calculates the optimal detonation point, with the following algorithm parameters: state space s=[target velocity, metal mass, threat level, relative distance], reward function R(s,a)=0.7×damage coefficient(s)-0.3×energy cost(a), Q-value update formula Q(s,a)=R(s,a)+γ×max aQ(s,a) (discount factor γ = 0.9). When conditions ① target distance ≤ 12m, ② target arrival time ≤ 5s, and ③ action value function Q(s,a) > 8.5 are met simultaneously, the reinforcement learning controller (9) triggers the detonation command; at the same time, the cooperative communication interface 11 executes the scheduling strategy: the main mine sends the detonation coordinates, and the auxiliary mine emits a 20kHz interference wave. Energy management prioritizes the operation of the reinforcement learning controller 9 (forced resource allocation, response time ≤ 0.2ms).
[0025] Damage Repair Phase: After interception, the sensor data fusion module 10 detects the damage depth of the titanium alloy-boron carbide ceramic substrate 5. If the damage is ≥2mm, the reinforcement learning controller 9 triggers the rupture of the polyurea microcapsules 6, releasing the repair agent. The repair agent fills the crack through a polymerization reaction (repair time ≤30 seconds, response delay ≤5 seconds). The system returns to the monitoring cycle to ensure long-term deployment reliability. Under extreme sea conditions, this self-repair mechanism of the titanium alloy-boron carbide ceramic substrate 5 based on polyurea microcapsules 6 is activated as a backup function with secondary priority (T1 level): the power consumption of this self-repair function is forced to be ≤1.5W (accounting for 15.3% of the system's peak power consumption of 9.8W), and the upper limit of the delay response is ≤5 seconds (measured average 4.2 seconds). When the reinforcement learning controller 9 performs a combat mission (T0 level), the repair process is delayed until the mission ends. When there is no T0 level mission, the repair is activated immediately, and the repair agent of the polyurea microcapsules 6 is polymerized and solidified within 30 seconds. The round counter of polyurea microcapsule 6 records the initial 50 rounds of repair capability. When the retested damage depth δ' ≥ 0.5 mm, the built-in repair agent reserve unit of polyurea microcapsule 6 is activated. After each repair, the round counter updates the remaining available rounds of the built-in repair agent reserve unit. When the total system energy is less than 10%, the self-repair function is suspended, and only the core sensing functions of fluxgate sensor 1, DFOS positioning module 7 and fiber optic vibration calibration module 3 are maintained (power consumption ≤ 3W). The self-repair system automatically restarts the queue of tasks to be repaired after the energy is restored to more than 15%.
[0026] This system achieves a highly efficient and low-carbon underwater defense solution suitable for various marine environments through multi-level collaboration (such as magnetic-acoustic-optical data fusion and reinforcement learning decision-making) and biomimetic design (intelligent repair of polyurea microcapsules 6).
Claims
1. An intelligent mine system based on magnetic-acoustic-optical fusion sensing and reinforcement learning decision-making, characterized in that: It consists of a fluxgate sensor (1), a CNN voiceprint classifier (2), an optical fiber vibration calibration module (3), a Q-learning decision module (4), a titanium alloy-boron carbide ceramic substrate (5), polyurea microcapsules (6), a DFOS positioning module (7), an LSTM trajectory prediction module (8), a reinforcement learning controller (9), a sensor data fusion module (10), and a cooperative communication interface (11); The entire system comprises three subsystems: a magneto-acoustic-optical fusion perception subsystem, a reinforcement learning dynamic decision-making subsystem, and a biomimetic self-repair subsystem. The magnetic-acoustic-optical fusion sensing subsystem includes a fluxgate sensor (1), a CNN acoustic text classifier (2), an optical fiber vibration calibration module (3), a DFOS positioning module (7), and a sensor data fusion module (10); The DFOS positioning module (7) is directly connected to the sensor data fusion module (10) via an optical fiber data bus, and outputs the three-dimensional coordinates of the target in real time; the fluxgate sensor (1) and the DFOS positioning module (7) work together to capture the target's metal characteristics and spatial position, and the fluxgate sensor (1) is independently connected to the sensor data fusion module (10) via a circuit, and outputs the target's metal characteristic parameters; the CNN acoustic text classifier (2) is connected to the sensor data fusion module (10) in parallel and inputs acoustic signals; the optical fiber vibration calibration module (3) is connected in series with the input end of the CNN acoustic text classifier (2) via a circuit to eliminate water flow noise interference; the DFOS positioning module (7) and the fluxgate sensor (1) realize spatial-attribute data association in the sensor data fusion module (10); The reinforcement learning dynamic decision-making subsystem includes a Q-learning decision module (4), an LSTM trajectory prediction module (8), a reinforcement learning controller (9), and a cooperative communication interface (11); The LSTM trajectory prediction module (8) is connected to the DFOS positioning module (7) through a high-speed data interface, and the Q-learning decision module (4) is connected to the LSTM trajectory prediction module (8) through a high-speed data transmission interface; the LSTM trajectory prediction module (8) is connected to the reinforcement learning controller (9) by a circuit connection, and the reinforcement learning controller (9) is connected to the cooperative communication interface (11) through a control signal line; the Q-learning decision module (4), the LSTM trajectory prediction module (8) and the reinforcement learning controller (9) are all connected to the sensor data fusion module (10) through a data bus. The biomimetic self-healing subsystem includes a titanium alloy-boron carbide ceramic matrix (5) and polyurea microcapsules (6); wherein the titanium alloy-boron carbide ceramic matrix (5) and the polyurea microcapsules (6) are connected by an embedded encapsulation to form a composite structure, and the polyurea microcapsules (6) are uniformly distributed inside the titanium alloy-boron carbide ceramic matrix (5).
2. The intelligent mine system based on magneto-acoustic-optical fusion sensing and reinforcement learning decision-making as described in claim 1, characterized in that: The three subsystems are connected by components from their respective systems: The output of the magnetic-acoustic-optical fusion sensing subsystem is closely connected to the Q-learning decision module (4) of the reinforcement learning dynamic decision subsystem via the sensor data fusion module (10); The reinforcement learning controller (9) of the reinforcement learning dynamic decision-making subsystem is connected to the polyurea microcapsule (6) of the biomimetic self-healing subsystem via a control signal line; The linkage strategy of the collaborative communication interface (11) is as follows: Interacting with the isomorphic smart mine array: the main mine sends the detonation coordinates, and the auxiliary mine receives and transmits a 20kHz jamming wave; Interacting with the Navy's shore-based command and control system: Encrypted transmission of target trajectory data (including DFOS positioning coordinates and acoustic signature); Interacting with marine environmental monitoring buoys: Receiving temperature, salinity, and depth data to correct the sound wave propagation attenuation model.
3. The intelligent mine system based on magneto-acoustic-optical fusion sensing and reinforcement learning decision-making as described in claim 1, characterized in that: The DFOS positioning module (7) monitors the underwater target position in real time with a positioning error of ≤0.5m.
4. The intelligent mine system based on magnetic-acoustic-optical fusion sensing and reinforcement learning decision-making as described in claim 1, characterized in that: The fluxgate sensor (1) is used to capture metal targets with a mass ≥ 500 kg; A CNN voiceprint classifier (2) was used to analyze the voiceprints of thrusters in the 20Hz-20kHz frequency band; The fiber optic vibration calibration module (3) is used to eliminate water flow noise.
5. The intelligent mine system based on magneto-acoustic-optical fusion sensing and reinforcement learning decision-making according to claim 1, characterized in that: The titanium alloy-boron carbide ceramic matrix (5) provides structural support, and the polyurea microcapsules (6) with a diameter of 50 μm automatically rupture and release the repair agent after being impacted.
6. The intelligent mine system based on magneto-acoustic-optical fusion sensing and reinforcement learning decision-making according to claim 1, characterized in that: The collaborative communication interface (11) supports cross-system data sharing; the group collaborative strategy includes precise positioning of the main mine and acoustic interference of the auxiliary mine, with an interference frequency of 20kHz.
7. The method for an intelligent mine system based on magneto-acoustic-optical fusion sensing and reinforcement learning decision-making according to claim 1, characterized in that: The method includes a target identification stage, a dynamic decision-making stage, and a damage repair stage; Target recognition stage: DFOS positioning module (7) detects the three-dimensional coordinates of metal-powered ship targets in real time; fluxgate sensor (1) captures the metal characteristic parameters of the target ship hull (mass ≥ 500 kg); CNN voiceprint classifier (2) analyzes voiceprint features; Fiber optic vibration calibration module (3) eliminates environmental noise; sensor data fusion module (10) fuses data and outputs target status: including target category (enemy / friend), speed of movement, heading angle and threat level (what state of the target, in detail); Dynamic decision-making stage: The LSTM trajectory prediction module (8) predicts the motion trajectory within the future time period Δt (Δt∈[2s,15s]) based on the real-time three-dimensional coordinates of the target output by the DFOS positioning module (7); the Q-learning decision-making module (4) calculates the optimal detonation point that satisfies the following conditions: ① Spatial location: The closest point on the target trajectory to the smart mine whose distance is ≤ the effective damage radius of the explosion (R=12m); ②Time window: The time t for the target to arrive at this point is ≤5s; ③ Value determination: Action value function Q(s,a) (s=[target speed, metal mass, threat level, relative distance])>Q min =8,5; The reinforcement learner (9) triggers the detonation command when conditions ①②③ are met simultaneously; Damage repair stage: The polyurea microcapsule (6) automatically ruptures after detecting damage cracks on the surface or inside the titanium alloy-boron carbide ceramic matrix (5), releasing the repair agent to fill the crack gaps and repairing the cracks in the outer shell of the matrix through in-situ polymerization reaction. The repair completion time is ≤30 seconds. After the system starts, it enters a continuous monitoring cycle and maintains standby mode when there is no threat. Low power detection state: the fluxgate sensor (1), fiber optic vibration calibration module (3), and DFOS positioning module (7) operate at 3W power for continuous monitoring. Hibernation components: the CNN voiceprint classifier (2) is turned off. The reinforcement learning controller (9), LSTM trajectory prediction module (8), and cooperative communication interface (11) enter hibernation (power consumption ≤ 0.5W). Threat detection activation: the sensor data fusion module (10) calculates the target threat level and starts interception when the threat level > 0.
7. Cooperative interception execution: the reinforcement learning controller (9) schedules the main mine and auxiliary mine through the cooperative communication interface (11): first, the main mine triggers the Q-learning detonation decision based on the coordinate data of the DFOS positioning module (7); second, the auxiliary mine synchronously emits a 20kHz acoustic interference signal to cover the target acoustic sensor. Damage self-inspection and repair: After the interception is completed, the sensor data fusion module (10) detects the damage depth of the titanium alloy-boron carbide ceramic substrate (5); When the damage depth is ≥2mm, the polyurea microcapsule (6) ruptures and repairs the cracks in the titanium alloy-boron carbide ceramic matrix (5); Energy management strategy: During standby and interception, the system always prioritizes the power supply of the reinforcement learning controller (9); the total power consumption of the system in standby state is ≤5W; when a threat is detected, resources are forcibly allocated to the reinforcement learning controller (9) to ensure its response time is ≤0.2ms; the power supply priority of the self-healing function is secondary, and the response delay is ≤5 seconds.
8. The method for an intelligent mine system based on magneto-acoustic-optical fusion sensing and reinforcement learning decision-making according to claim 7, characterized in that: The state space of the Q-learning decision module (4) includes the target velocity, distance, and angle. The reward function is a combination of damage gain weight 0.7 and energy consumption weight 0.3, and the calculation formula is as follows: R(s,a) = 0.7 × damage coefficient (s) - 0.3 × energy cost (a) Wherein, the state space s contains the DFOS positioning coordinates, target velocity, and the mass of the metal detected by the fluxgate, and a is the action; The update formula for the Q-learning decision module (4) is: Where γ = 0.9 and the threshold Q_min = 8.5.
Citation Information
Patent Citations
Intelligent water area virtual electronic fence based on acousto-optic-electric fusion
CN119920045A