Multimodal perception optimization method and system based on domain knowledge enhancement
By constructing a three-layer military domain knowledge model and knowledge compensation mechanism, the accuracy and robustness issues of modern battlefield multimodal perception technology in complex environments have been solved, achieving high-precision, adaptive battlefield situational awareness and real-time situational assessment.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- XIAMEN YUANTING INFORMATION TECH CO LTD
- Filing Date
- 2026-04-13
- Publication Date
- 2026-05-12
AI Technical Summary
In modern, complex battlefield environments, multimodal perception technology faces problems such as a sharp drop in perception accuracy, system instability, and a lack of integration of military domain knowledge. As a result, the perception results are insufficient in accuracy and robustness in highly confrontational environments, making it difficult to meet the needs of actual combat.
A multimodal perception optimization method based on domain knowledge enhancement is adopted. By constructing a three-layer military domain knowledge model of battlefield structure geometry, combat tactical rules and weapon physical motion consistency, and combining multimodal perception data, spatiotemporal synchronization and military feature extraction are performed to construct a joint optimization objective function. When data is missing, a knowledge compensation mechanism is triggered to achieve high-precision and adaptive perception optimization.
Achieving precise perception and real-time situational assessment across the entire domain in complex battlefield environments improves perception accuracy and system stability, ensuring that perception results conform to military logic and physical laws, and adapts to highly disturbed environments.
Smart Images

Figure CN122020071A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data processing technology in modern complex battlefield combat environments, and in particular to a multimodal perception optimization method and system based on domain knowledge enhancement. Background Technology
[0002] Modern warfare has entered the era of information and intelligence, making battlefield situational awareness a key factor in determining the outcome of operations. With the widespread application of unmanned combat platforms, individual soldier digital equipment, and joint operations command systems, multimodal perception technology has become a core support for battlefield reconnaissance, target identification, and situational assessment. The coordinated use of multiple sensors, such as lidar, visible light cameras, infrared thermal imagers, millimeter-wave radar, and inertial measurement units, provides combat personnel with unprecedented battlefield information acquisition capabilities.
[0003] However, the highly adversarial nature of modern, complex battlefield environments poses severe challenges to multimodal perception technologies. In typical military scenarios such as field maneuver assaults, urban warfare, underground tunnel warfare, border reconnaissance, and joint operations involving multiple services, perception systems face the following prominent problems: I. Strong environmental interference leads to a sharp drop in sensing accuracy Battlefield smoke, gunpowder, and dust severely interfere with the imaging quality of optical sensors, resulting in blurred visible / infrared images and loss of target features. Multipath reflection interference in complex terrain and densely built-up areas causes millimeter-wave radar and lidar to generate numerous false echoes. Severe vibrations from combat platforms (armored vehicles, UAVs) cause IMU attitude drift and point cloud distortion. The combined effect of these interference factors significantly reduces the accuracy of existing sensing systems in real-world combat environments, making it difficult to meet the demands of precision strikes and coordinated operations.
[0004] II. Sudden loss of sensing data leads to system instability In extreme situations such as strong electromagnetic interference suppression, communication link interruption, and partial sensor damage, sensing data may experience sudden and significant loss. Most existing technologies rely on continuous and complete data streams for optimization; once the data loss rate exceeds a threshold, it often leads to algorithm divergence, sensing results failure, and even the complete shutdown of the situational awareness system. This "data-dependent" approach is significantly vulnerable in highly contested battlefield environments.
[0005] Third, existing technologies lack deep integration of military domain knowledge. Current mainstream multimodal perception optimization methods are mainly based on a purely data-driven approach, such as multi-sensor fusion filtering (Kalman filtering, particle filtering), deep learning feature extraction and fusion, and graph optimization. While these methods perform well in certain conventional scenarios, they cannot utilize the inherent geographical structure of the battlefield (standardized features of tunnel cross-sections, symmetry of urban buildings, boundaries of defensive fortifications, etc.) to geometrically correct perception results, leading to severe distortions in structural recognition under obstructed or smoke-filled environments. Perception results may violate basic tactical logic, such as locking without identification, launching attacks without locking on, or state transitions that violate operational procedures. Although the output situational information may be mathematically sound, it is militarily unusable. Furthermore, they cannot utilize the physical limits of weaponry (maximum speed, maximum acceleration, steering overload limits) to verify the perceived motion state, easily resulting in false perception results that violate physical laws, such as ultra-high-speed movement of individual soldiers or sudden ascents and descents of drones without power.
[0006] In summary, existing battlefield multimodal perception technologies have room for improvement in terms of perception accuracy, robustness, and dynamic adaptability under highly contested environments, in order to meet the high-precision perception requirements of modern warfare, which demands all-weather, all-terrain, and strong anti-interference capabilities. Therefore, there is an urgent need for a multimodal perception optimization method that can deeply integrate military domain knowledge, adapt to changes in the battlefield environment, and possess strong anti-interference capabilities, in order to overcome battlefield information fog and achieve accurate all-domain perception and real-time situational assessment in complex combat environments. Summary of the Invention
[0007] In view of this, the purpose of this invention is to propose a multimodal perception optimization method based on domain knowledge enhancement, which can adapt to the modern complex battlefield confrontation environment. It focuses on five core military scenarios: field maneuver, urban warfare, underground tunnel warfare, border reconnaissance, and joint operations of multiple services. It directly addresses the core operational pain points caused by strong confrontation disturbances such as battlefield smoke and dust obscuring, dense obstacle obstruction, severe vibration of combat platforms, suppression of enemy electromagnetic interference, and signal attenuation in complex terrain. These disturbances lead to a sharp drop in multimodal perception accuracy, insufficient system stability, sudden loss of perception data, and failure of optimization algorithm convergence. At the same time, it makes up for the shortcomings of existing battlefield perception technologies, such as the lack of military domain-specific knowledge constraints and the inability to adapt robustness to actual combat needs. It constructs a multimodal perception optimization system with deep military domain knowledge enhancement, providing high-precision perception technology support with all-weather, all-terrain, and strong interference resistance for all-domain situational awareness, autonomous operation of unmanned combat swarms, accurate identification and tracking of battlefield targets, and efficient decision-making in joint operations command. It completely breaks through the fog of battlefield information and realizes accurate all-domain perception and real-time situational assessment in complex combat environments.
[0008] According to one aspect of the present invention, a multimodal perception optimization method based on domain knowledge enhancement is provided, the method comprising: Collect multimodal battlefield perception data, perform spatiotemporal synchronization and military feature extraction, and construct a standardized battlefield state vector; A three-layer military domain knowledge constraint model is constructed, consisting of a battlefield structure geometric knowledge model, a combat tactical rules knowledge model, and a weapon and equipment physical motion consistency knowledge model. The model takes the battlefield state vector as input and outputs battlefield structure geometric constraint terms, combat tactical rules constraint terms, and weapon and equipment physical motion consistency constraint terms. It evaluates the consistency between the geometric structural features in the battlefield state vector and the prior battlefield structure, the consistency between the target state transition in the battlefield state vector and the combat tactical logic, and the consistency between the motion parameters in the battlefield state vector and the physical motion limits of the equipment. The observation error term of the multimodal battlefield perception data is weighted and fused with the three types of constraint terms output by the three-layer military domain knowledge constraint model to construct a joint optimization objective function based on the battlefield state vector; Based on the current battlefield situation, by quantifying the confidence level of each knowledge constraint in the current battlefield environment, and using an exponential function as a regulator, the weight of each constraint in the joint optimization objective function is dynamically allocated, so that the optimization model can adapt to the real-time battlefield environment. The missing proportion of battlefield perception data is calculated. When the missing proportion exceeds a preset threshold, a knowledge compensation mechanism is triggered to reconstruct the battlefield state vector and predict and complete the historical state. The compensated battlefield state vector is then re-input into the joint optimization objective function for optimization. The joint optimization objective function is solved iteratively until the convergence condition is met, and the optimized battlefield state vector is output as the final high-precision battlefield perception result.
[0009] In the above technical solution, the key feature is the "knowledge-driven-data-driven collaborative optimization" closed-loop perception method of this invention. Its core implementation is as follows: First, through multimodal data acquisition and spatiotemporal synchronization, a standardized battlefield state vector is constructed as a unified data carrier for the entire process. Based on this, a three-layer military domain knowledge model is constructed, encompassing battlefield structural geometry, operational tactical rules, and the consistency of weapon and equipment physical motion. This domain knowledge is transformed into computable structural, tactical, and physical constraints. Then, sensor observation errors are weighted and fused with the three types of constraints to form a joint optimization objective function with the state vector as the independent variable. By quantifying the confidence level of each constraint in the current battlefield environment, an exponential function is used to dynamically adjust the weights, enabling the optimization model to adapt to changes in the battlefield situation in real time. When the data loss ratio exceeds a threshold, a knowledge compensation mechanism based on structural template reconstruction and motion prediction is triggered to ensure perception continuity. Finally, a high-precision battlefield state vector conforming to military logic is output through iterative solution.
[0010] Unlike existing data-driven multimodal fusion methods (such as those relying solely on Kalman filtering, deep learning feature stitching, or graph optimization), the key feature of this technology lies in modeling three types of military domain knowledge as quantifiable constraints, which are then integral components of the optimization objective function. The resulting benefits are: the perception results not only numerically fit sensor observations, but also conform to prior battlefield geometry in spatial structure (e.g., building symmetry, tunnel cross-section standardization), behavioral logic in operational tactics (e.g., state transition legitimacy), and motion parameters in accordance with equipment physical limits (e.g., continuity of velocity, acceleration, and angular velocity). This fundamentally eliminates the common "mathematically reasonable but militarily unusable" false perception results found in traditional methods.
[0011] Furthermore, unlike static weights or fixed weighting strategies that rely solely on environment classification, this technique features an exponential weight update mechanism based on dynamic confidence adjustment. Its key characteristic is that it quantifies the reliability of each constraint in the current battlefield environment by calculating the variance of the error in real time, and dynamically allocates weights using an exponential function as a regulator. This ensures that the model is always dominated by the most reliable constraint under different interference conditions (such as increased smoke concentration or enhanced electromagnetic interference). The technical effect of this mechanism is that the system can automatically adjust its optimization strategy when the environment changes abruptly, without manual intervention or scene switching logic, significantly improving its adaptability in highly adversarial environments.
[0012] Unlike existing methods that directly interrupt or simply interpolate when sensor data is missing, this technology introduces a triggered knowledge compensation mechanism. Its key feature is that it uses the proportion of missing data as a trigger condition. When this proportion exceeds a preset threshold, it actively invokes standardized structural templates from the battlefield structural geometric knowledge model for geometric reconstruction. Simultaneously, it performs motion prediction based on historical state vectors and physical motion laws, achieving structured completion of the missing data. The compensated state vectors are then re-inputted into the optimization model to form a closed loop. The beneficial effect of this mechanism is that even in extreme situations such as partial sensor damage, communication interruption, or severe environmental obstruction, the perception system can maintain continuous output, preventing the optimization algorithm from diverging or failing due to data loss, and effectively improving the system's resilience in highly disturbed battlefield environments.
[0013] This technology employs a closed-loop design of "three-layer knowledge model construction—adaptive weight adjustment—trigger-based knowledge compensation," constructing a "knowledge-enhanced" multimodal perception optimization paradigm distinct from traditional data-driven methods. Its core contribution lies in placing military domain knowledge, previously confined or post-processed, into the construction and solution of the optimization objective function, thus deeply integrating knowledge constraints and data observation at the mathematical optimization level. This feature is closely coupled with the overall technical solution's data architecture of "battlefield state vectors throughout the entire process" and "algorithm flow oriented towards real-world scenarios," jointly supporting the invention's technical objective of achieving high-precision, robust, and adaptive perception in complex battlefield environments.
[0014] In some embodiments, the acquisition of multimodal battlefield perception data, spatiotemporal synchronization and military feature extraction, and the construction of a standardized battlefield state vector specifically include: Deploy lidar, visible light cameras, infrared thermal imaging cameras, millimeter-wave radar, and inertial measurement units to collect multimodal battlefield perception data; A composite calibration method combining military-grade timestamp hard matching and anti-interference interpolation soft alignment is used for the multimodal sensing data to uniformly anchor all sensing data to the same operational time node, thereby achieving spatiotemporal synchronization. The terrain slope, accessibility, building and tunnel outlines, and equipment curvature are extracted from the synchronized lidar point cloud as battlefield geographical structure features. Individual soldier combat posture, equipment thermal imaging features, and military target-specific texture features are extracted from visible light or infrared images as combat target features. Target distance, movement speed, and azimuth are extracted from millimeter-wave radar as target motion status features. Attitude change rate, acceleration fluctuation, and maneuver trajectory are extracted from inertial measurement units as maneuver features of combat platforms. All extracted military features are integrated to construct the battlefield state vector.
[0015] In the above technical solution, this technical feature is the preprocessing stage of the perceived data from raw acquisition to structured expression. Its specific implementation is as follows: First, five types of military sensors—LiDAR, visible / infrared cameras, millimeter-wave radar, and inertial measurement units—are deployed to achieve collaborative acquisition of multimodal battlefield data. Then, a composite calibration method combining military-grade timestamp hard matching and anti-interference interpolation soft alignment is adopted to solve the spatiotemporal misalignment problems caused by differences in sampling frequencies of various sensors, platform vibration, and electromagnetic interference in the battlefield environment, uniformly anchoring all data to the same operational time node. Based on this, four types of features with clear military semantics are extracted from each modal data—geographical structure features, operational target features, target motion status features, and operational platform maneuver features. Finally, these features are integrated to construct a standardized battlefield state vector, serving as a unified input for subsequent knowledge modeling and optimization.
[0016] Unlike conventional preprocessing methods that simply use a unified timestamp alignment or rely on a single sensor for feature extraction in existing technologies, the essential feature of this technology lies in the construction of a composite spatiotemporal calibration mechanism and a multi-dimensional military feature extraction framework oriented towards the intense battlefield environment.
[0017] First, at the spatiotemporal synchronization level, unlike conventional linear interpolation or alignment methods that rely solely on hardware triggering, this technology employs a composite calibration strategy of "hard matching + soft alignment." The hard matching layer utilizes a military-grade timing module to achieve nanosecond-level timestamp anchoring, eliminating sampling bias at its source. The soft alignment layer uses high-frequency inertial measurement unit data as a benchmark and employs an anti-interference interpolation algorithm to perform motion compensation on low-frequency sensor data. The technical effect of this composite mechanism is that, even under conditions of severe vibration and electromagnetic interference on the combat platform, it can still ensure accurate alignment of data in the temporal dimension across different modes, suppressing the accumulation of perception errors caused by time misalignment and attitude drift at its source, and providing a spatiotemporally consistent input foundation for subsequent knowledge models.
[0018] Secondly, at the feature extraction level, unlike existing technologies that commonly use image features (such as SIFT and HOG) or point cloud features (such as FPFH), this technology targets specific battlefield needs by transforming raw sensor data into four types of feature modules with clear military semantics: geographical structure features (terrain slope, accessibility, building and tunnel outlines, equipment curvature) directly serve the constraint determination of the battlefield structural geometry model; combat target features (soldier posture, thermal imaging distribution, texture attributes) provide friend-or-foe identification and behavior criteria for the tactical rule model; and target motion status features (distance, speed, azimuth) and platform maneuver features (posture change rate, acceleration, trajectory) provide motion parameter inputs for the physical consistency model. The technical effect of this feature organization method is that it transforms raw sensor data into a structured information carrier directly corresponding to the three-layer knowledge model, avoiding information loss at the military semantic level in general features, and enabling subsequent knowledge constraints to accurately apply to each dimension of the state vector, rather than relying solely on a few key parameters.
[0019] Finally, the four types of features mentioned above are integrated into a standardized battlefield state vector, forming a unified data carrier throughout the entire process. The technical effect of this design is that it eliminates the problem of heterogeneous format of multimodal data in subsequent optimization, compensation, and output stages, and provides a unified mathematical object for the knowledge model input in step two, the construction of the optimization function in step three, and the triggering of the compensation mechanism in step five, so that the entire method forms a complete logical closed loop from data acquisition to result output.
[0020] This technology, through a three-layer preprocessing architecture of "composite spatiotemporal calibration—multi-dimensional military feature extraction—standardized state vector construction," lays a high-quality, structured, and militarily semantically clear data foundation for the subsequent military knowledge modeling and joint optimization of this invention. Its core contribution lies in bringing the practical constraints of battlefield environment adversariality, multi-source sensor heterogeneity, and military information specialization to the preprocessing stage. Through a specifically designed calibration method and feature organization approach, the raw sensor data is transformed into a state vector expression that can directly interface with the three-layer domain knowledge model, thereby ensuring the effective operation of the subsequent "knowledge-driven-data-driven collaborative optimization" mechanism.
[0021] In some embodiments, the battlefield structure geometric knowledge model is as follows:
[0022] in, For battlefield structural geometric constraints, For the real-time attitude angle of the combat platform, The desired attitude angle is based on the battlefield prior structure. This represents the deviation value of the battlefield structure symmetry. , These are dynamic weighting coefficients.
[0023] In the above technical solution, this technical feature is a specific implementation of the battlefield structure geometric knowledge model in this invention, and is the first layer of the three-layer military domain knowledge constraint model. This model uses the battlefield state vector constructed in step one. As input, output battlefield structure geometric constraints. This is used to evaluate the consistency between the geometric structural features in the state vector and the prior battlefield structure. The model is implemented through a combination of two weighted error functions: the first term... Constrain the consistency between the real-time attitude angle of the combat platform and the expected attitude angle calculated based on terrain structure; the second item The symmetry deviation between the constrained perceived structural profile and the prior battlefield geometry template. These two errors are weighted by dynamic weighting coefficients. , Weighted summation is used to form a comprehensive evaluation index for the geometric dimension of the state vector space.
[0024] Unlike existing technologies that rely solely on inertial measurement units for attitude estimation or point cloud data for geometric reconstruction (a single-dimensional processing method), this technology's key feature lies in its construction of a geometric consistency evaluation mechanism based on joint "attitude-structure" constraints. This mechanism integrates platform maneuvering characteristics with geographical structural features through desired attitude angles. Establish coupling relationships and combine various geometric features such as building outlines and equipment shapes through symmetry deviation terms. Incorporate it into a unified constraint framework.
[0025] Specifically, the first item Its essential feature is that it measures the platform's attitude angle. (From the inertial measurement unit) and the desired attitude angle calculated based on geographical structural features such as terrain slope and traffic capacity. A comparison was then performed. The innovation of this design lies in establishing a geometric coupling relationship between the platform's motion state and the battlefield geographical environment: when the platform is traveling on known terrain, if the measured attitude deviates from the expected attitude calculated from the terrain, it indicates that there may be drift in attitude perception or errors in terrain extraction. The resulting technical effect is that this constraint can not only correct the attitude drift caused by vibration of the inertial measurement unit, but also reversely correct the extraction errors of terrain slope and passability in the lidar point cloud, realizing mutual calibration between the two feature modules.
[0026] Second item Its essential characteristic lies in that it measures the battlefield structure symmetry deviation value. As a core evaluation metric, this deviation value comprehensively reflects the degree of difference between geographical structural features such as building / tunnel outlines and equipment curvature and the prior geometric template of the battlefield. Unlike existing technologies that rely solely on point cloud matching or feature point registration, this constraint introduces a structured geometric constraint based on the prior template: when smoke or occlusion causes missing point cloud data or outline distortion, the model forces the perception results to converge towards the prior template by penalizing structural symmetry deviations. The resulting technical effect is that it can effectively correct outline recognition errors caused by environmental interference. For example, when smoke obscures half of a tank body, this constraint can drive the optimization process to use the prior geometric template of military equipment to complete the missing curvature and outline information.
[0027] Dynamic weighting coefficients , The design further enhances the model's scene adaptability. Unlike fixed-weight geometric constraint methods, this technique allows for adaptive adjustment of the weighting of two constraints based on different combat scenarios (such as field warfare, urban warfare, and tunnel warfare). For example, in urban warfare, the weight of the structural symmetry constraint can be increased. Improve attitude-terrain matching weights during field maneuvers. The technical effect of this mechanism is that the model can maintain optimal constraint strength in different battlefield environments, avoiding optimization deviations caused by excessively strong or weak single constraints.
[0028] In summary, this technology, through attitude-structure coupling constraints and a priori template-guided geometric correction mechanism, achieves comprehensive geometric consistency verification of geographical structural features, platform maneuvering features, combat target features, and target motion features in the state vector. Its beneficial effects are manifested in that, even with noise, missing, or distorted sensor data, it can still output perception results that conform to the geometric logic of battlefield space, fundamentally eliminating false perceptions that violate spatial geometry, such as "targets hovering in the air" or "equipment passing through walls."
[0029] The battlefield structural geometric knowledge model constructed by this technical feature is the first pillar of the three-layer military domain knowledge constraint system of this invention. Its core contribution lies in transforming prior knowledge of the battlefield spatial structure into a quantifiable attitude-structure joint constraint function, enabling the optimization process to systematically evaluate and correct the state vector in the geometric dimension. This model complements the subsequent operational tactical rule model and weapon physical motion consistency model: the structural model ensures that the perception results are "spatially reasonable," the tactical model ensures "behavioral compliance," and the physical model ensures "motional accessibility." Together, these three constitute the core of this invention—"knowledge enhancement"—that distinguishes it from traditional data-driven methods, providing geometric constraint guarantees for achieving high-precision, robust multimodal perception in complex battlefield environments.
[0030] In some embodiments, the combat tactical rule knowledge model is verified by establishing a combat state transition logic matrix. When the target state transition in the battlefield state vector conforms to the preset combat state transition logic, the output constraint term is 0; when the state transition does not conform to the preset combat state transition logic, the output constraint term is a preset penalty coefficient.
[0031] In the above technical solution, this technical feature is a specific implementation of the operational tactical rules knowledge model in this invention, and is the second layer of the three-layer military domain knowledge constraint model. This model uses the battlefield state vector constructed in step one. As input, output combat tactical rule constraints. This is used to evaluate the consistency between target state transitions and operational tactical logic in the state vector. The model uses a pre-defined operational state transition logic matrix. Define legal state transition paths (e.g., search → identify → track → lock → attack → withdraw), and base them on the changes in the target state in the current state vector and the matrix. Comparison and verification are performed: when the state transition conforms to the preset tactical logic, the constraint term is set to 0, and no penalty is imposed on the optimization process; when the state transition does not conform to the preset tactical logic, the constraint term is set to the preset penalty coefficient. Force corrections are made to perception results that violate tactical logic.
[0032] Unlike existing technologies that rely solely on target tracking algorithms (such as Kalman filtering and particle filtering) for motion state estimation or simply use simple rules to post-process and verify target behavior, the essential feature of this technology is that it embeds combat tactical rules into the objective function of perception optimization in the form of a state transition logic matrix, thereby achieving hard constraint verification of the target state evolution.
[0033] First, unlike existing technologies where tactical logic is only presented as an "optional" feature in the post-processing stage of perception, this technology directly incorporates constraints into the joint optimization objective function, making tactical logic an active evaluation criterion in the optimization process. Its key feature is that the model uses a pre-defined operational state transition logic matrix. This transforms highly abstract tactical knowledge, which relies heavily on operational doctrine (such as "no locking without identification" and "no attack without locking"), into computable discrete state validity rules. When a target state transition occurs in the battlefield state vector (e.g., jumping directly from "identification" to "attack" and skipping "lock"), the constraint term outputs a penalty coefficient. During the optimization process, a high penalty is imposed on the state vector. The technical effect of this mechanism is that the optimization algorithm will actively avoid state estimation results that may have small errors in sensor observation but violate tactical logic during the iterative solution process, thereby ensuring that the final output perception results conform to the actual combat specifications in terms of behavioral logic, and fundamentally eliminating the false situation judgments that are "mathematically reasonable but tactically unusable" in traditional methods.
[0034] Secondly, this technical feature uses a penalty coefficient. This design achieves a balance between rigid constraints and flexible tolerance for tactical violations. Unlike simple verification methods that rely solely on logical judgments to output a binary "legal / illegal" result, this feature uses a preset penalty coefficient (such as...) This ensures that the violation state is penalized in the optimization objective function by a significantly higher order of magnitude than the sensor observation error, thereby guaranteeing that the optimization process prioritizes the satisfaction of tactical logic constraints; simultaneously, The values can be preset and adjusted according to the combat scenario or mission type (such as urban warfare vs. border patrol), enabling the model to flexibly adapt to different tactical environments. The technical effect of this design is that, while ensuring that the perception results conform to basic tactical specifications, it retains the flexibility to output usable results in extreme cases (such as when sensors fail severely) for the optimization process.
[0035] Furthermore, this technical feature, as the "behavioral logic" dimension in the three-layer constraint system, forms a complementary and synergistic constraint network with the battlefield structural geometry model (spatial dimension) and the physical motion consistency model (motion dimension). For example, when the structural model determines that the target's spatial location conforms to building geometry (not penetrating walls), and the physical model determines that the target's speed conforms to equipment limits (not exceeding speed limits), but the tactical model determines that its state transition is illegal, the optimization process will still reject the state estimate due to the high penalty of tactical constraints. The resulting technical effect is that the three-layer model systematically verifies the compliance of the perception results from different dimensions, ensuring that the final output conforms to the actual battlefield situation at the three levels of spatial geometry, physical motion, and behavioral logic simultaneously, forming a multi-dimensional perception quality assurance mechanism that differs from single-dimensional constraint methods.
[0036] The operational tactical rules knowledge model constructed by this technology is the key behavioral logic dimension of the three-layer military domain knowledge constraint system of this invention. Its core contribution lies in transforming tactical knowledge dependent on operational doctrine into computable state transition legality judgment rules, and embedding them into the objective function of perception optimization in the form of proactive penalties, thus upgrading tactical logic from "post-processing verification" to "proactive constraint during the optimization process." This model, together with the battlefield structure geometry model and the physical motion consistency model, constitutes a complete knowledge constraint system from three dimensions: "spatial rationality—movement accessibility—behavioral compliance," providing tactical compliance assurance for achieving highly reliable, interpretable, and combat-logical battlefield perception in complex battlefield environments.
[0037] In some embodiments, the knowledge model for the consistency of physical motion of the weapon system is as follows:
[0038] in, For the physical motion consistency constraint of weapon equipment, , These represent the current and previous moments' movement speeds of the combat platform, respectively. For the acceleration at the current moment, , These are the angular velocities at the current moment and the previous moment, respectively. Let be the angular acceleration at the current moment. These are the attitude motion constraint weight coefficients.
[0039] In the above technical solution, this technical feature is a specific implementation of the weapon and equipment physical motion consistency knowledge model in this invention, and is the third layer of the three-layer military domain knowledge constraint model. This model uses the battlefield state vector constructed in step one. As input, output the physical motion consistency constraint term of weapon equipment. This is used to evaluate the consistency between the motion parameters in the state vector and the physical motion limits of the equipment. The model achieves systematic constraints on the motion state through a combination of three weighted error functions: the first term... Constraining the temporal continuity of motion velocity, the second term The physical limit of constrained acceleration, the third term The continuity and physical limits of constrained angular velocity and angular acceleration, where This is the attitude motion constraint weight coefficient, used to balance the constraint magnitude of linear and angular motion.
[0040] Unlike existing technologies that merely use Kalman or particle filters to smooth the motion state or simply truncate velocity and acceleration by setting a single threshold, the essential feature of this technology lies in constructing a physical consistency evaluation mechanism with joint constraints of "linear motion and angular motion". It incorporates four types of motion parameters—velocity, acceleration, angular velocity, and angular acceleration—into a unified constraint framework and systematically verifies the compliance of motion features in the state vector through both temporal continuity and physical limits.
[0041] First, in terms of temporal continuity, this model... and Two constraints are implemented on the rate of change of velocity and the rate of change of angular velocity. Unlike existing technologies that rely solely on filtering algorithms for implicit motion smoothing, this technique incorporates temporal continuity as an explicit error term into the optimization objective function, proactively avoiding abrupt changes in motion states during the optimization process. Its key feature is that when there are drastic jumps in velocity or angular velocity in the battlefield state vector (such as a jump from 30 km / h to 80 km / h within 0.1 seconds, or an instantaneous 180-degree change in attitude angle), this constraint term incurs a high penalty, forcing the optimization algorithm to correct such perception results that violate physical continuity. The technical effect of this mechanism is that it fundamentally eliminates "instantaneous" false motion perception caused by electromagnetic interference, data gaps, or algorithmic divergence at the mathematical optimization level, ensuring the physical reliability of target tracking and platform state estimation.
[0042] Secondly, in the physical limit dimension, this model... and Both measures constrain the amplitude of acceleration and angular acceleration. Unlike existing post-processing methods that rely solely on sensor range for hard truncation, this technique limits the physical motion of the equipment (e.g., the maximum acceleration of an armored vehicle does not exceed...). The maximum angular acceleration of the UAV (limited by aerodynamic constraints) is incorporated into the optimization process as an embedded constraint term. Its key feature is that when the acceleration or angular acceleration in the state vector exceeds the physical limit of the corresponding equipment, the constraint term outputs a penalty value proportional to the square of the excess, forcing the optimization results to converge to the physically feasible region. The resulting technical effect is that the perception results will not exhibit phenomena that violate the physical motion laws of the equipment, such as "ultra-high-speed movement of a single soldier," "sudden ascent and descent of an armored vehicle without power," or "abrupt attitude changes of the UAV," significantly improving the authenticity and usability of the perception data.
[0043] Third, this technical feature is implemented through weighting coefficients. This achieves a matching of the magnitudes of linear and angular motion constraints. Unlike existing techniques that handle position / velocity and attitude / angular velocity separately, this feature unifies both types of motion parameters within the same constraint framework, and through… Adjust the weight of angular motion constraints in the total constraints. Its design feature is that, based on the combat scenario (e.g., frequent attitude changes in urban warfare require increased weighting of angular motion constraints) or equipment type (e.g., UAVs require strict constraints on angular velocity, while armored vehicles require focused constraints on acceleration), the weight of angular motion constraints can be adjusted. Pre-set adjustments are made to ensure the model maintains optimal constraint balance in different application scenarios. The technical effect of this mechanism is to avoid the failure of constraints in a certain dimension due to the difference in magnitude between linear and angular motion, and to achieve comprehensive and balanced constraints on the motion state.
[0044] In summary, this technology achieves comprehensive physical consistency verification of platform maneuver characteristics and target motion characteristics in the state vector through four constraints: velocity continuity, acceleration limits, angular velocity continuity, and angular acceleration limits. Furthermore, through extended applications such as the matching relationship between terrain slope and equipment acceleration, and the coupling relationship between individual soldier attitude change rate and human physical limits, it indirectly constrains the rationality of geographical structure characteristics and combat target characteristics. Its beneficial effects are manifested in that the perception results not only conform to continuity and physical limits in the motion trajectory, but also establish an inherent physical logical consistency between motion parameters, equipment type, and battlefield terrain, ensuring the kinematic credibility of the perception data from the source.
[0045] The weapon and equipment physical motion consistency knowledge model constructed by this technology is the key physical dimension of the three-layer military domain knowledge constraint system of this invention. Its core contribution lies in transforming the physical motion limits and motion continuity laws of equipment into quantifiable four-fold constraint functions of velocity, acceleration, angular velocity, and angular acceleration. These are embedded as explicit error terms into the objective function of perception optimization, enabling the optimization process to systematically evaluate and correct the state vector in the kinematic dimension. This model, together with the battlefield structural geometry model (spatial dimension) and the combat tactical rules model (behavioral dimension), constitutes a complete knowledge constraint system from three dimensions: "spatial rationality—behavioral compliance—motion reachability," providing crucial physical motion compliance guarantees for achieving highly reliable and robust multimodal perception in complex battlefield environments.
[0046] In some embodiments, dynamically assigning weights to each constraint in the joint optimization objective function includes: Calculate the first Confidence of class knowledge constraints ,in, For the first The variance of the error of the constraint term; Update the weight coefficients based on confidence level. ,in, As the initial baseline weights, The adjustment factor is constructed using an exponential function, which makes the weights positively correlated with the confidence level.
[0047] In the above technical solution, this technical feature is a specific implementation of the adaptive weight adjustment mechanism in this invention, and is a key link connecting the three-layer military domain knowledge constraint model and the joint optimization objective function. This mechanism is driven by real-time battlefield situation and dynamically adjusts the weight coefficients of each knowledge constraint in the optimization objective function by quantifying the confidence level of each knowledge constraint in the current environment. Specifically, the implementation method is as follows: first, calculate the... Variance of knowledge constraint term error Confidence level is defined accordingly. This allows constraints with smaller and more stable error fluctuations to achieve higher confidence levels; and then, based on the initial baseline weights... Based on this, through the adjustment factor constructed using an exponential function Update real-time weights Furthermore, it explicitly establishes a positive correlation between weights and confidence levels. This mechanism operates continuously during the optimization iteration process, ensuring that the optimization model is always dominated by the most reliable knowledge constraints in the current battlefield environment.
[0048] Unlike existing static weighting methods that use fixed weights or preset weights based solely on scenario type (such as "urban street fighting mode" or "field fighting mode"), the essential feature of this technology lies in constructing a confidence quantification mechanism based on the real-time performance of constraint terms and an exponential dynamic weight update strategy, enabling the optimization model to possess "meta-cognition" capabilities regarding the reliability of its own knowledge constraints.
[0049] First, at the confidence quantification level, this technical feature is achieved through... It transforms the reliability of knowledge constraints into a computable numerical indicator. Its characteristics include: Directly reflects the first The degree of fluctuation in constraint term error: When the battlefield environment is stable, the constraint term error fluctuates little, the variance approaches zero, and the confidence level approaches 1; when strong interference occurs on the battlefield (such as increased smoke concentration causing frequent errors in structural constraints, or electromagnetic interference causing increased errors in physical constraints), the variance of the constraint term error increases accordingly, and the confidence level decreases accordingly. The technical effect of this design is that it enables real-time quantitative assessment of whether knowledge constraints are "currently credible," providing an objective basis for subsequent weight adjustments, rather than relying on prior judgments based on human presets or environmental classifications.
[0050] Secondly, at the weight update level, this technical feature adopts... The dynamic adjustment strategy, in which This paper proposes an exponential function as the adjustment factor, explicitly demonstrating a positive correlation between weights and confidence levels. Unlike existing technologies that rely solely on linear adjustments or threshold switching, the exponential function as an adjustment factor is characterized by its monotonically increasing and smooth nature, ensuring the continuity and stability of weight changes with confidence levels and preventing optimization oscillations caused by abrupt weight changes. Simultaneously, the non-linear nature of the exponential function allows for flexible adjustment of the weights' sensitivity to confidence levels through function parameters, adapting to the dynamic changes in different battlefield environments. The resulting technical effect is that when a certain type of knowledge constraint exhibits high confidence in the current environment (e.g., extremely low variance in structural constraint errors under clear weather conditions), its weight is adaptively increased, allowing this constraint to play a dominant role in optimization. When environmental deterioration leads to a decrease in confidence, the weight automatically decays, preventing unreliable constraints from misleading optimization results. This mechanism achieves a "more reliable, less unreliable" optimization resource allocation, significantly improving the system's adaptability in complex battlefield environments.
[0051] Third, this technical feature serves as a bridge between the three-layer knowledge model and the optimization objective function, and its synergistic effect with the overall technical solution is particularly prominent. Its characteristics lie in: the confidence calculation relies on... Originating from the evaluation process of the state vector by the three-layer model in step two, the weight update result directly affects the proportion of constraint terms in the optimization objective function in step three. The optimized state vector, in turn, influences the confidence calculation at the next moment, forming a closed-loop adaptive architecture of "perception—evaluation—adjustment—optimization." The technical effect of this design is that the optimization model can automatically adjust its optimization strategy when the battlefield environment changes dynamically, without manual intervention or scene switching logic, achieving a leap from "passive adaptation" to "active adaptation."
[0052] The adaptive weight adjustment mechanism constructed by this technical feature is the core enabling element for achieving "knowledge-driven and data-driven collaborative optimization" in this invention. Its core contribution lies in: through the organic combination of confidence quantification and exponential dynamic weight updates, it endows the optimization model with the ability to assess the reliability of knowledge constraints in real time and adjust the optimization strategy accordingly. This upgrades the three-layer military domain knowledge constraint model from a static "scorer" to a dynamically adaptable "intelligent evaluation system" that adapts to the battlefield situation. This mechanism, together with the standardized state vector in step one, the three-layer knowledge model in step two, and the joint optimization objective function in step three, constitutes a complete closed-loop adaptive perception architecture, providing technical support for achieving highly robust and adaptive multimodal perception in complex battlefield environments.
[0053] In some embodiments, triggering a knowledge compensation mechanism includes: Calculate the proportion of missing battlefield data ,in For the amount of missing data, This represents the total data volume for the corresponding modality; when When the threshold is exceeded, the missing geometric perception data is reconstructed by calling a standardized structural template through the battlefield structure geometric knowledge model, and then reconstructed using motion prediction formulas. Predict and complete the battlefield state vector at the current moment, where The time interval between adjacent perceptions.
[0054] In the above technical solution, this technical feature is a specific implementation of the battlefield strong disturbance triggered knowledge compensation mechanism of this invention, and is a key emergency link to ensure the continuous operation of the perception system in extreme environments. This mechanism uses the proportion of missing battlefield data as the trigger condition, and by quantitatively assessing the completeness of perception data for each modality, initiates the compensation process when the data loss exceeds a preset threshold. The specific implementation method is as follows: first, the missing proportion is calculated. ,in For the amount of missing data, This represents the total data volume for the corresponding modality; when When the threshold is exceeded, two types of compensation measures are activated simultaneously: on the one hand, the missing geometric perception data is reconstructed by calling standardized structural templates through the battlefield structure geometric knowledge model; on the other hand, based on historical state vectors and physical motion laws, motion prediction formulas are used to reconstruct the missing geometric perception data. The current battlefield state vector is predicted and completed, and the compensated state vector is re-input into the optimization model to ensure the continuity of the perception and optimization process.
[0055] Unlike conventional methods in existing technologies that simply use zero-value filling, linear interpolation, or direct interruption of the sensing process when sensor data is missing, the essential feature of this technology lies in constructing a four-stage knowledge-driven compensation architecture of "trigger-based judgment - knowledge template reconstruction - motion prediction completion - closed-loop reconnection". It uses the military's prior knowledge and the laws of motion physics as sources of information supplementation under the condition of missing data, realizing a leap from "passive filling" to "active repair".
[0056] First, at the triggering mechanism level, this technical feature is based on the missing ratio. This mechanism achieves quantitative determination and adaptive triggering of compensation timing. Unlike existing technologies that rely solely on a single sensor fault signal or a fixed time interval for triggering, this mechanism is characterized by: and The system can be defined separately for different sensor modalities (such as lidar point clouds, infrared images, and millimeter-wave radar echoes), allowing compensation trigger conditions to precisely match the data characteristics of each modality. Preset thresholds can be flexibly adjusted according to combat scenarios (e.g., 0.2 for urban warfare, 0.3 for open field environments), ensuring the system maintains optimal compensation sensitivity under different conditions. The resulting technical benefits are: avoiding resource waste caused by frequent compensation triggers due to minor data fluctuations (such as 1% random packet loss), and enabling timely intervention when data loss reaches a critical point affecting perception quality, thus achieving optimal control over compensation timing.
[0057] Secondly, at the level of compensation methods, this technology employs a heterogeneous collaborative compensation strategy combining structural template reconstruction and motion prediction completion. Unlike existing technologies that rely solely on time series interpolation, this feature uses two types of compensation methods to address information gaps in different dimensions: The first type of compensation reconstructs the missing geometric perception data by calling standardized structural templates (such as tunnel cross-section templates, urban building outline templates, and typical field terrain templates) from the battlefield structural geometry knowledge model. Its key feature is that the compensation information originates from the battlefield prior knowledge base constructed in step two, rather than simple numerical calculations, ensuring that the compensation results conform to the actual battlefield structure at the spatial geometry level. The second type of compensation utilizes motion prediction formulas... Predictive completion of state vectors is characterized by its prediction being based on the optimized state vectors and physical motion laws of the previous moment, ensuring the compensation results maintain kinematic continuity and physical consistency. The synergistic effect of these two compensation methods is that when lidar point clouds are missing due to smoke, structural template reconstruction can complete the terrain outline and equipment shape; when communication interruptions lead to the loss of target motion data, motion prediction can maintain continuous estimation of the target trajectory. Both methods achieve complementary repair of missing information from the dimensions of "spatial structure" and "motion state," respectively.
[0058] Third, this technical feature, through the design of re-inputting compensated data into the optimization model, forms a closed-loop reconstruction mechanism of "compensation-optimization." Unlike existing technologies where compensated data is merely used as a replacement output and not further processed, this feature integrates the compensated state vector... The joint optimization objective function from step three is re-entered, enabling the compensation data and subsequent observation data to be processed collaboratively within the optimization framework. The resulting technical benefits are: after the compensation result enters the optimization model, it still needs to be verified through three layers of knowledge constraints and balanced with sensor observation error terms, avoiding the cumulative errors that might be introduced by simple compensation; simultaneously, the iterative mechanism of the optimization process can re-evaluate the reliability of the compensation data, gradually correcting compensation biases through subsequent observation data, thus achieving a gradual recovery of sensing accuracy under conditions of missing data.
[0059] The battlefield-triggered knowledge compensation mechanism constructed by this technology is the core guarantee for achieving the "damage resistance" and "resilience" of the perception system in this invention. Its core contributions lie in: upgrading the compensation trigger condition from "sensor fault signal" to "quantitative discrimination of data missing ratio"; upgrading the source of compensation information from "numerical interpolation" to "heterogeneous collaboration of knowledge template reconstruction and motion prediction"; and upgrading compensation data processing from "alternative output" to "closed-loop reconstruction and re-optimization". This mechanism, closely coupled with the standardized state vector in step one, the three-layer knowledge model in step two, and the joint optimization objective function in step three, together constitute a highly robust perception system that can continue to operate under extreme battlefield conditions such as smoke cover, electromagnetic suppression, and partial sensor damage. This provides crucial technical support for achieving the practical capability of "uninterrupted perception and recoverable optimization failures".
[0060] According to another aspect of the present invention, a multimodal perception optimization system based on domain knowledge enhancement is provided, wherein the system comprises, based on the above-described method: The acquisition module is used to collect multimodal battlefield perception data, perform spatiotemporal synchronization and military feature extraction, and construct standardized battlefield state vectors. The constraint module is used to construct a three-layer military domain knowledge constraint model, which includes a battlefield structure geometric knowledge model, a combat tactical rules knowledge model, and a weapon and equipment physical motion consistency knowledge model. The model takes the battlefield state vector as input and outputs battlefield structure geometric constraint terms, combat tactical rules constraint terms, and weapon and equipment physical motion consistency constraint terms. It evaluates the consistency between the geometric structural features in the battlefield state vector and the prior battlefield structure, the consistency between the target state transition in the battlefield state vector and the combat tactical logic, and the consistency between the motion parameters in the battlefield state vector and the physical motion limits of the equipment. The module is used to weight and fuse the observation error terms of multimodal battlefield perception data with the three types of constraint terms output by the three-layer military domain knowledge constraint model to construct a joint optimization objective function based on the battlefield state vector. The weighting module is used to dynamically allocate the weight of each constraint in the joint optimization objective function by quantifying the confidence level of each knowledge constraint in the current battlefield environment based on the current battlefield situation and using an exponential function as a regulator, so that the optimization model can adapt to the real-time battlefield environment. The compensation module is used to calculate the missing proportion of battlefield perception data. When the missing proportion exceeds the preset threshold, the knowledge compensation mechanism is triggered to reconstruct the battlefield state vector and predict and complete the historical state. The compensated battlefield state vector is then re-input into the joint optimization objective function for optimization. The iteration module is used to iteratively solve the joint optimization objective function until the convergence condition is met, and outputs the optimized battlefield state vector as the final high-precision battlefield perception result.
[0061] In the above technical solution, in order to better utilize the above method, this application proposes a multimodal perception optimization system based on domain knowledge enhancement. Each module corresponds to each step of the above method, and its specific principle has been described above and will not be repeated here.
[0062] According to another aspect of the present invention, a multimodal perception optimization device based on domain knowledge enhancement is provided, comprising: At least one processor and a memory communicatively connected to said at least one processor; The memory stores instructions that can be executed by the at least one processor, which, when executed by the at least one processor, enables the at least one processor to perform the method described above.
[0063] In the above technical solution, to better operate and process the method, the method is stored in memory, and the processor executes the stored method. It should be noted that the principle and effect of each step have been described above and will not be elaborated upon here.
[0064] According to another aspect of the present invention, a computer-readable storage medium is provided storing a computer program that, when executed by a processor, implements the above-described method.
[0065] In the above technical solution, to better operate and use the method, the method is stored in a computer-readable storage medium and implemented using a processor. It should be noted that the principle and effect of each step have been described above and will not be elaborated upon here.
[0066] This invention, based on core technological innovation and tailored to the needs of complex combat scenarios, offers the following significant advantages over existing multimodal perception optimization methods, balancing practicality, robustness, and adaptability: 1. Robust and adaptable to highly disturbed environments: It specifically addresses pain points such as smoke, gunpowder, and equipment vibration. Through a three-layer domain knowledge constraint and knowledge-triggered compensation mechanism, it effectively suppresses perception errors, avoids optimization failures caused by data loss, and significantly improves perception stability and reliability. 2. Knowledge-driven optimization to improve perception accuracy: Unlike existing methods that simply fuse data layers or learn neural networks, this approach transforms battlefield structure, tactical rules, and physical motion laws into computable constraints, enabling collaborative optimization through data-driven and knowledge-driven approaches, and significantly improving perception accuracy in complex environments. 3. Strong adaptability and adaptability to multiple scenarios: Through a confidence-based adaptive weight adjustment mechanism, the constraint weights can be dynamically adjusted according to environmental changes in different scenarios without manual intervention, making it more adaptable. 4. Good continuity and outstanding resistance to missing data: The knowledge-triggered compensation mechanism can quickly supplement missing data through structural reconstruction and historical state prediction when the proportion of missing data exceeds the threshold, ensuring the continuity of the perception and optimization process and solving the problem of perception interruption when data is missing in existing methods; 5. Highly practical and easy to promote: The algorithm has a complete process and is easy to operate. It can be directly deployed in multimodal perception systems and can support modern warfare perception technology. 6. Scientific modeling and comprehensive constraints: The constructed three-layer knowledge model of structure, tactics, and physics fully covers structural features, tactical rules, and equipment motion laws, with a complete constraint system that effectively solves the shortcoming of existing technologies lacking systematic knowledge constraints. Attached Figure Description
[0067] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0068] Figure 1 This is a flowchart illustrating an embodiment of a multimodal perception optimization method based on domain knowledge enhancement according to the present invention. Figure 2 This is a schematic diagram of an embodiment of a multimodal perception optimization system based on domain knowledge enhancement according to the present invention. Detailed Implementation
[0069] The present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be particularly noted that the following embodiments are for illustrative purposes only and do not limit the scope of the invention. Similarly, the following embodiments are only some, not all, embodiments of the present invention, and all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0070] Example 1 Please see Figure 1 This embodiment provides a multimodal perception optimization method based on domain knowledge enhancement, applied to a field mobile assault scenario. In this scenario, armored combined arms units and unmanned combat platform clusters maneuver at high speed in the field, facing real-world interference such as battlefield smoke cover, complex terrain obstruction, platform vibration, and mild enemy electromagnetic interference. The method includes: S1. Collect multimodal battlefield perception data, perform spatiotemporal synchronization and military feature extraction, and construct a standardized battlefield state vector; In this embodiment, the process of collecting multimodal battlefield perception data, performing spatiotemporal synchronization and military feature extraction, and constructing a standardized battlefield state vector specifically includes: Deploy lidar, visible light cameras, infrared thermal imaging cameras, millimeter-wave radar, and inertial measurement units to collect multimodal battlefield perception data; A composite calibration method combining military-grade timestamp hard matching and anti-interference interpolation soft alignment is used for the multimodal sensing data to uniformly anchor all sensing data to the same operational time node, thereby achieving spatiotemporal synchronization. The terrain slope, accessibility, building and tunnel outlines, and equipment curvature are extracted from the synchronized lidar point cloud as battlefield geographical structure features. Individual soldier combat posture, equipment thermal imaging features, and military target-specific texture features are extracted from visible light or infrared images as combat target features. Target distance, movement speed, and azimuth are extracted from millimeter-wave radar as target motion status features. Attitude change rate, acceleration fluctuation, and maneuver trajectory are extracted from inertial measurement units as maneuver features of combat platforms. All extracted military features are integrated to construct the battlefield state vector.
[0071] For example, the core principle of this step is multi-domain sensor synergy and complementarity, and battlefield data anti-interference calibration. Drawing on the joint operations logic of multiple branches of service, various reconnaissance sensors are transformed into all-domain perception combat units. By unifying spatiotemporal references and refining military characteristics, the limitations of single-sensor perception are overcome, battlefield interference noise is filtered out, and high-quality battlefield perception data is output, laying a solid foundation for subsequent optimization. Assume the following sensors are deployed on a certain type of armored reconnaissance vehicle: lidar, visible light reconnaissance camera, infrared thermal imaging camera, inertial measurement unit, and millimeter-wave radar. Each sensor synchronously collects battlefield environmental data, specifically including: LiDAR acquires 3D point cloud data To obtain the geometric contours of battlefield terrain, fortifications, and enemy armored targets; Visible light cameras collect battlefield environmental image data To obtain the overall appearance and environment of the target; Infrared thermal imaging cameras acquire infrared image data Identify heat source targets obscured by smoke; Millimeter-wave radar acquires radar echo data To obtain the position and velocity information of distant targets; IMU acquires attitude / accelerometer / angular velocity data. It records the real-time movement of the reconnaissance vehicle while it is traveling on rugged terrain.
[0072] Because the sampling frequencies of the various sensors are different (10Hz for lidar, 30Hz for visible light camera, 20Hz for millimeter-wave radar, and 100Hz for IMU), and vehicle vibration causes sampling deviations, this embodiment uses a composite calibration method combining military-grade timestamp hard matching and anti-interference interpolation soft alignment for spatiotemporal synchronization. Hard matching: Each sensor is connected to a military-grade GPS / BeiDou timing module, with timestamp accuracy reaching the nanosecond level, ensuring that data triggered at the same time are aligned at the hardware level; Soft alignment: For sensor data with mismatched frequencies, an anti-interference interpolation algorithm based on inertial data is used. Using the high-frequency data of the IMU as a benchmark, motion compensation interpolation is performed on other low-frequency data to uniformly anchor all data to the same operational time point. .
[0073] After spatiotemporal synchronization, military-specific features are extracted: Extracting battlefield geographic structure features from point cloud data: terrain slope Traffic Capacity Index Building / tunnel outline Armored equipment shape curvature ; Extract combat target features from visible light / infrared images: individual soldier combat posture (prone / standing), equipment thermal imaging features (such as temperature distribution in tank engine compartment), and military target-specific texture features (such as camouflage pattern matching degree). Extracting target motion characteristics from millimeter-wave radar: target range Movement speed Azimuth ; Extracting combat platform maneuver characteristics from IMU: attitude angles Rate of change of posture Acceleration fluctuation Movement trajectory .
[0074] Integrate all features to construct a standardized battlefield state vector:
[0075] To avoid repetitive comparisons, the constraint attribution of the above state vectors is explained in advance here:
[0076] illustrate: 1. Direct constraints: These refer to parameters explicitly used in the core formulas or evaluation metrics of the model (e.g., ...). Appears in structural model formulas, , (Appearing in physical model formulas), or the model evaluates this parameter through a specific design (e.g.) (Use contour features directly).
[0077] 2. Indirect constraints: These refer to constraints imposed on the model by coupling relationships between parameters (such as matching pose with terrain, or associating target spatial location with contour) or by auxiliary inputs (such as...). The auxiliary inputs affect this parameter, but this parameter is not an explicit independent variable in the model formula.
[0078] 3. All parameters are directly constrained by at least one model (except...) In the structural model, the constraints are indirect, but in the tactical model, they are direct. There are no "unconstrained" parameters.
[0079] 4. The "direct constraints" of the tactical model are manifested in its reliance on the state transition logic matrix. The legality of the target status is verified, and the target status is determined by a combination of factors such as the characteristics of the combat target and the target movement characteristics. Therefore, the tactical model has a direct evaluation function for these characteristics.
[0080] The advantage of this step lies in deploying five major categories of military reconnaissance sensors: lidar, visible light reconnaissance cameras, infrared thermal imaging cameras, inertial measurement units (IMUs), and millimeter-wave radar. This constructs an integrated air-land, active-passive combined military multimodal perception system, fully adaptable to complex combat zones such as field warfare, urban warfare, tunnel warfare, and border reconnaissance. It simultaneously collects battlefield perception data across the entire domain, specifically including: lidar 3D point cloud data. (Battlefield terrain, fortifications, geometric outlines of combat equipment), visible light image data (Overall view of the battlefield environment, target appearance characteristics), infrared image data (Penetrating and identifying personnel and equipment as heat sources, and concealed and camouflaged targets), millimeter-wave radar data. (Smoke and fog penetration, long-range target detection and positioning), IMU attitude / acceleration / angular velocity data (Real-time movement attitude and maneuver status of the combat platform); Subscript To ensure consistent data collection times, rigorously guarantee data spatiotemporal correlation, and prevent situational misalignment, a composite calibration method combining military-grade timestamp hard matching and anti-interference interpolation soft alignment is employed. This method addresses issues such as inconsistent sampling frequencies and triggering mechanisms among reconnaissance sensors in battlefield environments, as well as sampling deviations and data jumps caused by severe vibrations of combat platforms, enemy electromagnetic interference, and natural environmental disturbances. This ensures all multimodal sensing data is uniformly anchored to the same operational time point. To ensure all data corresponds to the battlefield situation at the same time and in the same area, eliminating perception errors caused by time misalignment and attitude drift at the source, and achieving spatiotemporal consistency of perception data across the entire domain. Military-specific features are extracted and normalized from the synchronized calibrated data, discarding redundant general features and accurately extracting core battlefield combat features: battlefield geographical structure features (terrain slope, accessibility, building and tunnel outlines, equipment curvature) are extracted from point cloud data; combat target features (individual soldier combat posture, equipment thermal imaging features, military target-specific texture features) are extracted from visible light / infrared data; target motion status features (target distance, movement speed, azimuth angle) are extracted from millimeter-wave radar; and combat platform maneuver features (attitude change rate, acceleration fluctuation, maneuver trajectory) are extracted from IMU data. All high-value military features are integrated to construct a standardized battlefield state vector. This serves as the core input for subsequent perception optimization.
[0081] S2. Construct a three-layer military domain knowledge constraint model consisting of a battlefield structure geometric knowledge model, a combat tactical rules knowledge model, and a weapon and equipment physical motion consistency knowledge model. The model takes the battlefield state vector as input and outputs battlefield structure geometric constraint terms, combat tactical rules constraint terms, and weapon and equipment physical motion consistency constraint terms. It evaluates the consistency between the geometric structural features in the battlefield state vector and the prior battlefield structure, the consistency between the target state transition in the battlefield state vector and the combat tactical logic, and the consistency between the motion parameters in the battlefield state vector and the physical motion limits of the equipment. In this embodiment, the battlefield structure geometric knowledge model is as follows:
[0082] in, For battlefield structural geometric constraints, For the real-time attitude angle of the combat platform, The desired attitude angle is based on the battlefield prior structure. This represents the deviation value of the battlefield structure symmetry. , These are dynamic weighting coefficients.
[0083] In this embodiment, the combat tactical rule knowledge model is verified by establishing a combat state transition logic matrix. When the target state transition in the battlefield state vector conforms to the preset combat state transition logic, the output constraint term is 0; when the state transition does not conform to the preset combat state transition logic, the output constraint term is a preset penalty coefficient.
[0084] In this embodiment, the knowledge model for the consistency of physical motion of the weapon system is as follows: in, For the physical motion consistency constraint of weapon equipment, , These represent the current and previous moments' movement speeds of the combat platform, respectively. For the acceleration at the current moment, , These are the platform angular velocities at the current and previous moments, respectively. The platform angular acceleration at the current moment, These are the attitude motion constraint weights.
[0085] For example, this embodiment constructs a three-layer military domain knowledge constraint model, where each model has a state vector. The full range of features are jointly evaluated and corrected to form a three-dimensional cross-constraint system of "spatial structure - tactical logic - physical laws".
[0086] (1) Battlefield structural geometric knowledge model (structural model) The core of this model is to constrain all features related to the battlefield spatial structure, ensuring they conform to the inherent geometric priors of the battlefield's geography, architecture, and equipment. The constraint function is:
[0087] in, From The measured attitude angles extracted from the platform's maneuvering characteristics; Based on terrain slope and traffic capacity index Calculated theoretical expected attitude angle; This is the battlefield structure symmetry deviation value, used to determine whether the identified fortification outline conforms to a preset symmetrical structure. This value comprehensively evaluates the following characteristics: From The deviation between the extracted building / tunnel outlines and equipment curvature from the geographical structural features and the prior template; from The spatial rationality of the thermal imaging contours and individual soldier posture contours extracted from the features of combat targets; from Whether the spatial location of the target extracted from the target motion features conforms to the passable area; , The dynamic weighting coefficients are preset according to the scenario in this embodiment. This constraint enables precise correction of battlefield structure perception. The parameters involved in this model are shown in the table below:
[0088] The core of this model is to constrain all features related to the battlefield spatial structure, requiring them to conform to the inherent geometric priors of battlefield geography, architecture, and equipment; this model... The correction logic for all features is as follows:
[0089] First item By matching the platform's attitude with the terrain structure, the attitude and trajectory errors of the platform's maneuvering characteristics are corrected, while the terrain slope and accessibility extraction errors of the geographical structure characteristics are corrected in reverse. Second item By analyzing the deviation between the structural outline and the prior template, errors in the outline and curvature of geographical structural features, the shape and positioning errors of combat target features, and the distance and azimuth jump errors of target motion features are corrected.
[0090] For example, when smoke obscures half of the tank body and causes the lidar point cloud to be missing, the model calls the military equipment prior geometry template to complete the curvature and contour, and at the same time corrects the platform's expected attitude based on terrain slope constraints.
[0091] (2) Knowledge model of combat tactics rules (tactical model) Based on the operational doctrine of combined arms assault operations, establish an operational state transition logic matrix. Defines legal transfer paths between operational states such as search, identification, tracking, locking, attack, and withdrawal. For example: Legitimate transfer: Search → Identify → Track → Lock down → Crack down Illegal transfer: Identify → Strike (skip lock), Track → Evacuate (Strike incomplete) The core of this model is to constrain all features related to combat behavior and target attributes, ensuring they conform to actual combat tactical rules and operational state logic; the constraint function is:
[0092] or
[0093] Among them, the combat state transition logic matrix Clearly define the legal redirection rules for "search → identification → tracking → locking → attack → withdrawal"; penalty coefficient. This ensures that perception results that violate tactical logic are forcibly corrected. The parameters involved in this model are shown in the table below:
[0094] The core of this model is to constrain all features related to combat behavior and target attributes, ensuring they conform to actual combat tactical rules and operational state logic; this model... The correction logic for all features is as follows:
[0095] based on The characteristics of combat targets are used to verify whether the soldier's posture and the target's friend or foe attributes are consistent with the current combat status, and to correct errors in posture extraction and friend or foe identification. based on The target's motion characteristics are verified to determine whether the target's trajectory and speed conform to tactical logic (such as the irregular sudden changes in the relative speed of the formation during an armored group assault), and to correct the jump errors in the target's distance, speed, and azimuth. based on The platform's maneuver characteristics are verified to ensure that the platform's trajectory and maneuver status are in line with the tactical mission (e.g., there will be no acceleration toward the enemy during withdrawal), and the drift errors of the platform's attitude, trajectory, and acceleration are corrected. based on The geographical structure characteristics are used to verify whether the accessibility and assault routes conform to tactical specifications (e.g., armored assault routes will not be planned in areas with steep, impassable slopes), and to correct the extraction errors of terrain slope and accessibility.
[0096] In this embodiment, the way the operational tactical rules knowledge model (tactical model) constrains the state vector comes from the determination of the operational state. This embodiment provides the following example of extracting the operational state from the state vector. Those skilled in the art can also set it according to actual needs. This example is not the only one.
[0097] First, the matrix The standardization definition is as follows: Dimension: 6×6 square matrix (rows = previous state, columns = current state); Status enumeration (index 1-6): 1=Reconnaissance, 2=Identification, 3=Tracking, 4=Locking, 5=Attack, 6=Evacuation; Element rules: (Legal transfer) (Illegal transfer); Core legal transition matrix:
[0098] Furthermore, key constraint rules (supplemented by prior knowledge) Only "forward adjacent jumps" (such as reconnaissance → identification, lock-on → attack) and "same state maintenance" (such as reconnaissance → reconnaissance) are allowed. Cross-state transitions (such as identification → lock) and reverse transitions (such as attack → lock) are both illegal; The evacuation state is a terminal state that can only be maintained and cannot be transitioned to other states.
[0099] Secondly, from Real-time extraction of "combat status feature vectors" based on Four core military characteristics were identified, key indicators strongly correlated with combat status were selected, and a "combat status feature vector" was quantified to generate it. As the direct input for state recognition, the specific extraction rules are as follows:
[0100] The final generated feature vector
[0101]
[0102] Again, based on Real-time identification of current combat status
[0103] A "threshold determination + multi-indicator voting" mechanism is adopted to... The system analyzes the current combat status to ensure accurate and reliable identification results. The specific judgment rules are as follows:
[0104] Exception handling rules If the confidence level of all states The state is determined to be "ambiguous," and the valid state of the previous frame is retained. ; Three consecutive frames of "state ambiguity" trigger a knowledge compensation mechanism to complete the data based on prior templates. Then re-identify.
[0105] Subsequently, the real-time combat status transition logic was extracted. (Transfer path generation) Based on the current frame state Compared to the previous frame state Generate real-time state transition paths As the core object of tactical rule verification: 1. Definition of Transfer Path
[0106] Example: If the previous frame is "Recognition (2)" and the current frame is "Tracking (3)", then .
[0107] 2. Transfer trajectory cache Cache the state transition paths of the last 5 frames It is used to handle state fluctuations in a short period of time and avoid judgment errors caused by single misidentification.
[0108] Finally, based on the prior matrix Complete tactical compliance assessment (core model verification). The operational tactical rules knowledge model combines "transfer path matching + multi-frame consistency verification" with a priori matrix. To achieve real-time tactical compliance determination, the specific process is as follows: 1. Compliance determination of single-frame transfer From the prior matrix Search Corresponding elements : like Preliminary assessment: This is a "compliant transfer," subject to tactical constraints. Take 0 temporarily; like Preliminary assessment: "Illegal transfer," tactical constraints. Temporarily take the penalty coefficient (Recommended 10.0).
[0109] 2. Multi-frame consistency check (suppressing momentary misjudgments) Calculate the percentage of compliant transfers based on the cached 5-frame transfer trajectory:
[0110] like Confirming the "compliant transfer," ultimately ; like The transaction was deemed "suspected illegal transfer" and triggered. Feature re-extraction and state re-identification, if still illegal... (Mild punishment); like Confirmed "illegal transfer", ultimately (Severe penalty), forced optimization process correction eigenvalues.
[0111] 3. Connection logic with the tactical constraint model The judgment result directly affects the tactical constraints. And by globally optimizing the objective function, it corrects the error in reverse. Features that could lead to illegal transitions (such as misjudgment of target distance or insufficient attitude stability) are identified to ensure that the state transition in the next frame returns to compliant logic.
[0112] The overall process of the above steps is illustrated below:
[0113] (3) Knowledge model of the consistency of physical motion of weapons and equipment (physical model) The core of this model is to constrain all kinematic and dynamic characteristics, ensuring they conform to the physical motion limits and continuity of personnel / equipment. The constraint function is:
[0114] in, , The velocity of the moving subject is extracted (platform velocity is extracted from platform maneuver features, and target velocity is extracted from target motion features). The acceleration of the moving subject is calculated (platform acceleration is extracted from platform maneuver features, and target acceleration is calculated based on the velocity at adjacent time points). , For the platform angular velocity, from The attitude change rate calculation in the platform's maneuver characteristics; For the platform's angular acceleration, from The attitude change rate and acceleration fluctuation in the platform's maneuver characteristics are calculated. The attitude and motion constraint weights are set to ensure that the magnitudes of the attitude and linear motion constraints match. The parameters involved in this model are shown in the table below:
[0115] The core of this model is to constrain all kinematic and dynamic characteristics, ensuring they conform to the physical motion limits and continuity laws of personnel / equipment. This model... The correction logic for all features is as follows:
[0116] First item : Motion continuity constraints, speed and trajectory jump errors of the correction platform's maneuvering characteristics, and speed, distance, and azimuth instantaneous shift errors of the target's motion characteristics; Second item Physical limit constraints, acceleration and attitude change rate errors of the correction platform's maneuver characteristics (e.g., maximum acceleration of armored vehicles ≤ 3m / s²), and acceleration errors of the target's motion characteristics (maximum running speed of a single soldier ≤ 12m / s). Third item Physical limits of attitude motion are constrained to correct the rate of attitude change of the platform's maneuvering characteristics from exceeding the limit and attitude jump error, thus eliminating false perception results that violate physical laws, such as "armored vehicles turning 180° instantly and UAVs changing attitude suddenly without power." Extended constraints: Based on the human physical limit of individual soldier posture change rate, correct the individual soldier posture change error of combat target characteristics; based on the physical matching relationship between terrain slope and equipment acceleration, correct the terrain slope extraction error of geographical structure characteristics.
[0117] S3. The observation error term of the multimodal battlefield perception data is weighted and fused with the three types of constraint terms output by the three-layer military domain knowledge constraint model to construct a joint optimization objective function based on the battlefield state vector. For example, Define a sensor observation error term (i.e., the observation error term of multimodal battlefield perception data) to characterize the deviation between the state vector and the original sensor observation data:
[0118] in, Based on the state vector The predicted observations calculated in reverse This is the integrated value of the raw observation data from multiple sensors.
[0119] By weighted and fused with the sensor error term and the three-layer knowledge constraint term, a joint optimization objective function is constructed:
[0120] in, These are the weighting coefficients for battlefield structure, tactical rules, and physical consistency knowledge constraints, which can be adaptively adjusted according to different combat scenarios such as urban warfare, field warfare, and border patrols. Iterative optimization is performed using the Gauss-Newton method, with the state vector updated in each iteration. The process continues until the total error meets the convergence condition or reaches the maximum number of iterations, ultimately outputting battlefield perception results (target localization, situation identification, platform motion status) that conform to military domain knowledge, are highly accurate, and robust.
[0121] S4. Based on the current battlefield situation, by quantifying the confidence level of each knowledge constraint in the current battlefield environment, and using an exponential function as a regulator, dynamically allocate the weight of each constraint in the joint optimization objective function so that the optimization model can adapt to the real-time battlefield environment. In this embodiment, the weight of each constraint in the joint optimization objective function is dynamically assigned, including: Calculate the first Confidence of class knowledge constraints ,in For the first The variance of the error of the constraint term; Update the weight coefficients based on confidence level. ,in As the initial baseline weights, The adjustment factor is constructed using an exponential function, which makes the weights positively correlated with the confidence level.
[0122] An example is an adaptive weight adjustment mechanism based on battlefield situation. When the battlefield environment changes (such as when the enemy releases smoke grenades causing a decrease in visibility), the reliability of each knowledge constraint changes accordingly. This embodiment dynamically adjusts the weights by calculating the confidence level in real time.
[0123] First, calculate the variance of the error for each constraint term. Let the nearest... The structural constraint error sequence at each time step is: ,but:
[0124] Similarly, calculate and .
[0125] Confidence level is defined as:
[0126] When smoke causes missing point cloud data and increases structure recognition error... rise, Decrease. Update the weights according to the following formula:
[0127] Among them, the battlefield adjustment coefficient The initial weights are preset to... , , .
[0128] If the confidence level of the structural constraints in front of the smoke screen is... (If the error is small), then:
[0129] After the smoke cloud dissipated, the point cloud was severely missing, the structure recognition error increased, and the confidence level dropped to [missing information]. ,at this time:
[0130] As can be seen, the real-time weight of structural constraints decreased from 2.476 to 0.911, meaning that the decrease in confidence led to a decrease in weight, achieving a positive correlation adjustment where reliable constraints have greater influence and unreliable constraints have less influence. Meanwhile, the weights of physical and tactical constraints are automatically adjusted according to their respective confidence levels, ensuring that the optimization process is always dominated by the most reliable knowledge available at present.
[0131] S5. Calculate the missing proportion of battlefield perception data. When the missing proportion exceeds the preset threshold, trigger the knowledge compensation mechanism to reconstruct the battlefield state vector and predict and complete the historical state. Then, re-input the compensated battlefield state vector into the joint optimization objective function for optimization. In this embodiment, triggering the knowledge compensation mechanism includes: Calculate the proportion of missing battlefield data ,in For the amount of missing data, This represents the total data volume for the corresponding modality; when When the threshold is exceeded, the missing geometric perception data is reconstructed by calling a standardized structural template through the battlefield structure geometric knowledge model, and then reconstructed using motion prediction formulas. Predict and complete the battlefield state vector at the current moment, where This represents the adjacent sensing time interval. The template library is pre-built based on battlefield geographic information.
[0132] For example, the proportion of data loss when sensors are severely damaged or communication is interrupted. It may exceed the threshold. In this embodiment, the missing percentage is defined as follows:
[0133] in, This represents the number of missing points in the point cloud of the current frame. Theoretically, this should be considered. Preset field scenario thresholds. .
[0134] When the reconnaissance vehicle entered the dense smoke area, large areas of the lidar point cloud were missing. This triggers a knowledge compensation mechanism: Structural template reconstruction: Based on the current location (calculated from the IMU track), the standardized terrain template of the area is pre-stored in the battlefield structural geometric knowledge model to perform geometric reconstruction of the missing point cloud area and supplement the battlefield structural features. Historical state prediction: based on the state vector of the previous time step and speed of movement The current state is completed using a motion prediction formula:
[0135] in The time interval between adjacent perceptions.
[0136] The reconstructed and predicted compensation data are integrated to form a complete battlefield state vector. Re-enter step 3 to optimize the model.
[0137] S6. Iteratively solve the joint optimization objective function until the convergence condition is met, and output the optimized battlefield state vector as the final high-precision battlefield perception result.
[0138] For example, suppose that after 5 iterations, the joint optimization objective function converges, and the optimized battlefield state vector is output. The vector contains: Precise target location information: The enemy tank is located at (850m, 120°), with a speed of 45km / h, and has been verified by tactical rules to conform to the logic of armored group assault formation; Accurate platform motion status: The reconnaissance vehicle's attitude angle deviation is less than 0.1°, the position error is less than 0.5m, and the attitude change rate and angular acceleration both meet the vehicle's physical limits; Reliable battlefield situation: Three high-threat targets were identified, and their spatial location and thermal imaging profiles were verified by the structural model and matched the prior template with a 92% accuracy.
[0139] Will The data is sent to the vehicle-mounted command and control system to generate a battlefield situation map, supporting commanders' decision-making and collaborative operations with unmanned platforms.
[0140] Example 2 Please see Figure 2 A multimodal perception optimization system based on domain knowledge enhancement, based on the method described in one embodiment, the system comprising: The acquisition module is used to collect multimodal battlefield perception data, perform spatiotemporal synchronization and military feature extraction, and construct standardized battlefield state vectors. The constraint module is used to construct a three-layer military domain knowledge constraint model, which includes a battlefield structure geometric knowledge model, a combat tactical rules knowledge model, and a weapon and equipment physical motion consistency knowledge model. The model takes the battlefield state vector as input and outputs battlefield structure geometric constraint terms, combat tactical rules constraint terms, and weapon and equipment physical motion consistency constraint terms. It evaluates the consistency between the geometric structural features in the battlefield state vector and the prior battlefield structure, the consistency between the target state transition in the battlefield state vector and the combat tactical logic, and the consistency between the motion parameters in the battlefield state vector and the physical motion limits of the equipment. The module is used to weight and fuse the observation error terms of multimodal battlefield perception data with the three types of constraint terms output by the three-layer military domain knowledge constraint model to construct a joint optimization objective function based on the battlefield state vector. The weighting module is used to dynamically allocate the weight of each constraint in the joint optimization objective function by quantifying the confidence level of each knowledge constraint in the current battlefield environment based on the current battlefield situation and using an exponential function as a regulator, so that the optimization model can adapt to the real-time battlefield environment. The compensation module is used to calculate the missing proportion of battlefield perception data. When the missing proportion exceeds the preset threshold, the knowledge compensation mechanism is triggered to reconstruct the battlefield state vector and predict and complete the historical state. The compensated battlefield state vector is then re-input into the joint optimization objective function for optimization. The iteration module is used to iteratively solve the joint optimization objective function until the convergence condition is met, and outputs the optimized battlefield state vector as the final high-precision battlefield perception result.
[0141] In this embodiment, in order to better utilize the method described in one of the embodiments, this application proposes a multimodal perception optimization system based on domain knowledge enhancement. Each module corresponds to each step of the above method, and its specific principle has been described above and will not be repeated here.
[0142] Example 3 A multimodal perception optimization device based on domain knowledge enhancement includes: At least one processor and a memory communicatively connected to said at least one processor; The memory stores instructions that can be executed by the at least one processor, which, when executed by the at least one processor, enables the at least one processor to perform the method described in one of the embodiments.
[0143] In this embodiment, to better run and process the method described in one of the embodiments, the above method is stored in a memory, and the stored method is executed using a processor. It should be noted that the principle and effect of each step have been described above and will not be elaborated further here.
[0144] Example 4 A computer-readable storage medium storing a computer program that, when executed by a processor, implements the method described in one of the embodiments.
[0145] In this embodiment, to better operate and use the method described in one of the embodiments, the above method is stored in a computer-readable storage medium, and the above method is implemented using a processor. It should be noted that the principle and effect of each step have been described above and will not be elaborated further here.
[0146] The above description is only a part of the embodiments of the present invention and does not limit the scope of protection of the present invention. Any equivalent device or equivalent process transformation made based on the content of the present invention specification and drawings, or direct or indirect application in other related technical fields, are similarly included within the patent protection scope of the present invention.
Claims
1. A multimodal perception optimization method based on domain knowledge enhancement, characterized in that, The method includes: Collect multimodal battlefield perception data, perform spatiotemporal synchronization and military feature extraction, and construct a standardized battlefield state vector; A three-layer military domain knowledge constraint model is constructed, consisting of a battlefield structure geometric knowledge model, a combat tactical rules knowledge model, and a weapon and equipment physical motion consistency knowledge model. The model takes the battlefield state vector as input and outputs battlefield structure geometric constraint terms, combat tactical rules constraint terms, and weapon and equipment physical motion consistency constraint terms. It evaluates the consistency between the geometric structural features in the battlefield state vector and the prior battlefield structure, the consistency between the target state transition in the battlefield state vector and the combat tactical logic, and the consistency between the motion parameters in the battlefield state vector and the physical motion limits of the equipment. The observation error term of the multimodal battlefield perception data is weighted and fused with the three types of constraint terms output by the three-layer military domain knowledge constraint model to construct a joint optimization objective function based on the battlefield state vector; Based on the current battlefield situation, by quantifying the confidence level of each knowledge constraint in the current battlefield environment, and using an exponential function as a regulator, the weight of each constraint in the joint optimization objective function is dynamically allocated, so that the optimization model can adapt to the real-time battlefield environment. The missing proportion of battlefield perception data is calculated. When the missing proportion exceeds a preset threshold, a knowledge compensation mechanism is triggered to reconstruct the battlefield state vector and predict and complete the historical state. The compensated battlefield state vector is then re-input into the joint optimization objective function for optimization. The joint optimization objective function is solved iteratively until the convergence condition is met, and the optimized battlefield state vector is output as the final high-precision battlefield perception result.
2. The multimodal perception optimization method based on domain knowledge enhancement as described in claim 1, characterized in that, The process of collecting multimodal battlefield perception data, performing spatiotemporal synchronization and military feature extraction, and constructing a standardized battlefield state vector specifically includes: Deploy lidar, visible light cameras, infrared thermal imaging cameras, millimeter-wave radar, and inertial measurement units to collect multimodal battlefield perception data; A composite calibration method combining military-grade timestamp hard matching and anti-interference interpolation soft alignment is used for the multimodal sensing data to uniformly anchor all sensing data to the same operational time node, thereby achieving spatiotemporal synchronization. The terrain slope, accessibility, building and tunnel outlines, and equipment curvature are extracted from the synchronized lidar point cloud as battlefield geographical structure features. Individual soldier combat posture, equipment thermal imaging features, and military target-specific texture features are extracted from visible light or infrared images as combat target features. Target distance, movement speed, and azimuth are extracted from millimeter-wave radar as target motion status features. Attitude change rate, acceleration fluctuation, and maneuver trajectory are extracted from inertial measurement units as maneuver features of combat platforms. All extracted military features are integrated to construct the battlefield state vector.
3. The multimodal perception optimization method based on domain knowledge enhancement as described in claim 1, characterized in that, The battlefield structure geometric knowledge model is as follows: in, For battlefield structural geometric constraints, For the real-time attitude angle of the combat platform, The desired attitude angle is based on the battlefield prior structure. This represents the deviation value of the battlefield structure symmetry. , These are dynamic weighting coefficients.
4. The multimodal perception optimization method based on domain knowledge enhancement as described in claim 1, characterized in that, The combat tactical rule knowledge model is verified by establishing a combat state transition logic matrix. When the target state transition in the battlefield state vector conforms to the preset combat state transition logic, the output constraint term is 0; when the state transition does not conform to the preset combat state transition logic, the output constraint term is a preset penalty coefficient.
5. The multimodal perception optimization method based on domain knowledge enhancement as described in claim 1, characterized in that, The knowledge model for the consistency of physical motion of the weapon system is as follows: in, For the physical motion consistency constraint of weapon equipment, , These represent the current and previous moments' movement speeds of the combat platform, respectively. For the acceleration at the current moment, , These are the angular velocities at the current moment and the previous moment, respectively. Let be the angular acceleration at the current moment. These are the attitude motion constraint weight coefficients.
6. The multimodal perception optimization method based on domain knowledge enhancement as described in claim 1, characterized in that, Dynamically assign weights to each constraint in the joint optimization objective function, including: Calculate the first Confidence of class knowledge constraints ,in, For the first The variance of the error of the constraint term; Update the weight coefficients based on confidence level. ,in, As the initial baseline weights, The adjustment factor is constructed using an exponential function, which makes the weights positively correlated with the confidence level.
7. The multimodal perception optimization method based on domain knowledge enhancement as described in claim 1, characterized in that, Triggering knowledge compensation mechanisms includes: Calculate the proportion of missing battlefield data ,in For the amount of missing data, This represents the total data volume for the corresponding modality; when When the threshold is exceeded, the missing geometric perception data is reconstructed by calling a standardized structural template through the battlefield structure geometric knowledge model, and then reconstructed using motion prediction formulas. For the current battlefield state vector Perform predictive completion, among which The time interval between adjacent perceptions. The state vector from the previous time step. The speed of motion.
8. A multimodal perception optimization system based on domain knowledge enhancement, characterized in that, Based on the method according to any one of claims 1-7, the system comprises: The acquisition module is used to collect multimodal battlefield perception data, perform spatiotemporal synchronization and military feature extraction, and construct standardized battlefield state vectors. The constraint module is used to construct a three-layer military domain knowledge constraint model, which consists of a battlefield structure geometric knowledge model, a combat tactical rules knowledge model, and a weapon and equipment physical motion consistency knowledge model. The model takes the battlefield state vector as input and outputs battlefield structure geometric constraint terms, combat tactical rules constraint terms, and weapon and equipment physical motion consistency constraint terms. It evaluates the consistency between the geometric structural features in the battlefield state vector and the prior battlefield structure, the consistency between the target state transition in the battlefield state vector and the combat tactical logic, and the consistency between the motion parameters in the battlefield state vector and the physical motion limits of the equipment. The module is used to weight and fuse the observation error terms of multimodal battlefield perception data with the three types of constraint terms output by the three-layer military domain knowledge constraint model to construct a joint optimization objective function based on the battlefield state vector. The weighting module is used to dynamically allocate the weight of each constraint in the joint optimization objective function by quantifying the confidence level of each knowledge constraint in the current battlefield environment based on the current battlefield situation and using an exponential function as a regulator, so that the optimization model can adapt to the real-time battlefield environment. The compensation module is used to calculate the missing proportion of battlefield perception data. When the missing proportion exceeds the preset threshold, the knowledge compensation mechanism is triggered to reconstruct the battlefield state vector and predict and complete the historical state. The compensated battlefield state vector is then re-input into the joint optimization objective function for optimization. The iteration module is used to iteratively solve the joint optimization objective function until the convergence condition is met, and outputs the optimized battlefield state vector as the final high-precision battlefield perception result.
9. A multimodal perception optimization device based on domain knowledge enhancement, characterized in that, include: At least one processor and a memory communicatively connected to said at least one processor; The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method as described in any one of claims 1 to 7.
10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the method of any one of claims 1 to 7.