A multi-modal intelligent adjudication wargame playthrough device and system

By using multimodal perception and feature decoupling techniques, the exploration rate decay factor of the reinforcement learning model is converted, which solves the problem of physical trajectory information loss in the wargaming system, improves decision generation efficiency and computational robustness, and enables fast and accurate situational reasoning.

CN122114193BActive Publication Date: 2026-07-31FUJIAN JUNZUAN INTELLIGENT EQUIP CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
FUJIAN JUNZUAN INTELLIGENT EQUIP CO LTD
Filing Date
2026-04-29
Publication Date
2026-07-31

AI Technical Summary

Technical Problem

Existing wargaming systems have technical limitations in the process of converting physical interaction parameters into input states for the computational model. This leads to the loss of continuous action trajectory information of physical entities within the pre-decision window of the move-making process, making it impossible to effectively extract decision deterministic features. Consequently, decision generation efficiency is low and reasoning convergence is slow.

Method used

The multimodal perception module acquires the three-dimensional continuous coordinate sequence of the physical deduction entity, the feature decoupling module calculates the position variance and steady-state retention time, converts them into the exploration rate decay factor of the reinforcement learning model, the policy mapping module limits the search step size boundary, the instruction generation module generates state transition discrimination instructions that satisfy global game consistency, and the task is routed to the corresponding processing module as needed through the heterogeneous computing offloading bus.

Benefits of technology

It realizes the direct conversion of behavioral uncertainties in the physical space interaction process into a computational model, improves the inference convergence speed, reduces computational overhead, and ensures computational robustness and decision generation efficiency under complex adversarial situations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122114193B_ABST
    Figure CN122114193B_ABST
Patent Text Reader

Abstract

This invention relates to the field of wargaming technology and discloses a multimodal intelligent adjudication manual wargaming device and system, including a multimodal perception module, a feature decoupling module, a strategy mapping module, and an instruction generation module. The multimodal perception module acquires the coordinate sequence before the landing point of the wargaming entity; the feature decoupling module determines the intention feature parameters based on the coordinate position variance and retention time; the strategy mapping module converts the intention feature parameters into a reinforcement learning exploration rate decay factor to limit the search boundary; and the instruction generation module adjusts the action sampling distribution to generate discrimination instructions. This invention directly maps interaction features to model search constraints, realizing the transformation of behavioral uncertainty into the search boundary of the computational model, improving the convergence speed under complex game situations, and effectively alleviating the computational bottleneck caused by undirected search in high-dimensional space.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to a multimodal intelligent adjudication manual wargaming device and system, belonging to the field of wargaming technology. Background Technology

[0002] Current wargaming systems, as tools for simulating adversarial situations and generating tactical decisions, face a surge in the complexity of decision-making logic when dealing with multi-domain collaboration and nonlinear evolution tasks. The mainstream technical architecture combines physical chessboard interaction with a back-end digital processing unit. This architecture uses physical entities to provide a tactical perception interface and processes battle damage parameters, resource scheduling, and collaborative game logic through a back-end computing engine, balancing the intuitiveness of simulation operations with the level of automation in data decision-making. However, existing simulation architectures suffer from limitations in the technical mechanism of converting physical interaction parameters into input states for the computational model. In addition to hardware limitations such as physical interaction interfaces and spatial mapping of simulation entities, the control methods and decision-making algorithms are also insufficient in handling the uncertainty of physical interactions. For example, Chinese invention patent application CN121031386B discloses an intelligent wargaming simulation system based on a digital twin fusion model. While it improves simulation accuracy through digital twin modeling and multimodal feature fusion, decision generation relies on discrete state sampling or predetermined historical sequences after entity placement, failing to explore the continuous micro-motion characteristics of physical trajectories within the window before the move decision.

[0003] However, existing inference architectures have limitations at the technical mechanism level in the process of converting physical interaction parameters into input states of the computational model. Traditional processing methods are based on discrete state sampling, extracting the landing point coordinates as the model's feature vector only when the physical piece moves to the target position and lands. This mechanism leads to the loss of continuous action trajectory information of physical entities within the pre-decision window of the move decision, resulting in the inability to extract the implicit decision deterministic features. To address the challenge of low decision generation efficiency, conventional linear improvement paths focus on increasing the scale of backend computing power or introducing high-frequency state acquisition logic. Simply increasing the sampling frequency cannot establish the connection between the variance of physical trajectory micro-movements and the search space of the policy network. Since the computational model lacks prior constraints on the game intention in the physical space, when facing high-dimensional nonlinear adversarial tasks, the model conducts an undirected search in the global action solution space. This leads to convergence lag in the multi-agent policy evolution process and causes inference link delays.

[0004] Therefore, the technical problem to be solved by this invention is how to translate the uncertainty of physical space interaction into the sampling constraint boundary of the policy generation network by constructing a mapping path between the variance of the physical chess piece movement trajectory and the exploration rate decay factor of a specific calculation model, thereby improving the inference convergence speed under complex adversarial situations. Summary of the Invention

[0005] To address the problems in the background art, the technical solution of the present invention is as follows: A multimodal intelligent adjudication manual wargame simulation system, comprising:

[0006] The multimodal perception module is used to acquire the three-dimensional continuous coordinate sequence of the physical simulation entity within a preset sampling time. The three-dimensional continuous coordinate sequence corresponds to the preset hovering spatiotemporal window of the physical simulation entity before contacting the simulation interface. The multimodal perception module eliminates systematic position noise in the three-dimensional continuous coordinate sequence through a multi-sensor fusion mechanism.

[0007] The feature decoupling module is used to calculate the position variance and steady-state retention time of the three-dimensional continuous coordinate sequence in the Cartesian coordinate system, and to determine the intention feature parameters that represent the inference intention and have probabilistic attributes based on the position variance and steady-state retention time.

[0008] The policy mapping module is used to convert the intent feature parameters into the exploration rate decay factor of the reinforcement learning model, and to limit the policy search step size boundary of the reinforcement learning model in the multi-dimensional action space based on the exploration rate decay factor.

[0009] The instruction generation module is used to adjust the action sampling probability distribution of the reinforcement learning model based on the policy search step size boundary, and generate state transition discrimination instructions that correspond to the current inferred situation and satisfy global game consistency.

[0010] Preferably, it also includes a heterogeneous computing offloading bus; the heterogeneous computing offloading bus is used to identify the task attributes of the inference task, and when the task attribute is a linear rule matching task, it routes the inference task to the rule matching module based on the directed acyclic graph; the heterogeneous computing offloading bus is also used to route the inference task to the strategy mapping module when the task attribute is a nonlinear game task.

[0011] Preferably, the feature decoupling module determines the intention feature parameters by the following steps: Step S31, performing first-order difference processing on the three-dimensional continuous coordinate sequence to determine the transient displacement vector set of the physical deduction entity within the preset sampling time; Step S32, calculating the displacement change statistical features of the physical deduction entity based on the transient displacement vector set, and fitting the displacement change statistical features with the preset intention template to determine the intention feature parameters.

[0012] Preferably, the policy mapping module converts the exploration rate decay factor of the reinforcement learning model into the following steps: Step S41, establish a monotonic mapping relationship between the intention feature parameter and the neuron activation threshold; Step S42, when the intention feature parameter monotonically increases, decrease the neuron activation threshold to lock the local policy search path of the reinforcement learning model.

[0013] Preferably, the instruction generation module is used to acquire the global simulation data of the digital twin representing the simulation interface, inject the action sampling probability distribution into the global simulation data of the digital twin, and output an adjudication report containing the situation assessment score and decision weight.

[0014] Preferably, the policy mapping module further includes a prior weight bias unit; the prior weight bias unit is used to map the trajectory oscillation features of the physical deduction entity into the prior probability weight bias of the deep neural network, and to use the prior probability weight bias to perturb the initial feature vector of the reinforcement learning model in order to guide the model to predict the direction.

[0015] Preferably, it also includes a self-supervised evolution module; the self-supervised evolution module is used to extract the inference element set with timestamps and clean the inference element set into reinforcement learning training sample pairs with causal chain labels; the self-supervised evolution module is also used to drive the reinforcement learning model to perform adaptive iteration of model parameters based on the reinforcement learning training sample pairs.

[0016] Preferably, the multimodal perception module is equipped with an edge sensing node array, which includes an infrared matrix sensor and an ultrasonic ranging sensor, used to acquire spatial pose data of the physically deduced entity.

[0017] Preferably, the system is used for multi-domain collaborative tasks, including acquiring the state parameters of combat units in multi-dimensional combat domains, and performing coupled calculations on the state parameters of combat units according to state transition discrimination instructions to solve the global performance index.

[0018] A multimodal intelligent decision-making manual wargame simulation device is provided for running a multimodal intelligent decision-making manual wargame simulation system.

[0019] Compared with the prior art, the beneficial effects of the present invention are:

[0020] 1. In multimodal intelligent adjudication manual wargame simulation, by establishing a mapping path between the geometric space variance of the physical piece movement trajectory and the exploration rate decay factor of the policy network, the behavioral uncertainty in the physical space interaction process is directly transformed into the search boundary of a specific computational model. This enables the reinforcement learning model to dynamically adjust the sampling variance of the action output vector according to the characteristics of the operator's action trajectory, avoiding the model from performing undirected exploration in the high-dimensional action space, thereby improving the inference convergence speed of the computational model when dealing with complex adversarial situations.

[0021] 2. This scheme utilizes edge sensing nodes to capture continuous coordinate sequences within a preset time window before the physical chess pieces are placed, and extracts the trajectory oscillation and hovering features contained in the sequence to reconstruct the state confidence component. The physical residuals, which are regarded as noise in traditional information processing, are transformed into prior probability weight biases of deep learning networks, thereby enhancing the system's perception dimension of implicit adversarial intentions in the context of incomplete information games, and thus reducing the computational overhead in the decision generation process.

[0022] 3. By embedding a heterogeneous computing offloading bus in the underlying architecture, the system routes deterministic rule tasks such as basic maneuver calculations to a lightweight static rule matching engine based on directed acyclic graphs, while routing high-dimensional nonlinear tasks involving multi-agent collaborative game and battle loss probability calculations to the artificial intelligence model layer. This achieves decoupling and on-demand allocation of computational load, constructs a computational robustness barrier for the system under high-frequency concurrent interaction of physical entities and data surge conditions, and avoids inference link deadlock in deep neural networks under extreme conditions. Attached Figure Description

[0023] Figure 1 This is a complete logic diagram of the multimodal perception interaction to decision constraint process of the present invention;

[0024] Figure 2 This is a diagram of the heterogeneous computing offloading bus and its task routing architecture of the present invention.

[0025] The objectives, features, and advantages of this invention will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation

[0026] The technical solutions of the embodiments of this application will be clearly described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. All other embodiments obtained by those skilled in the art based on the embodiments of this application are within the scope of protection of this application.

[0027] A multimodal intelligent adjudication manual wargame simulation system includes:

[0028] The multimodal perception module is used to acquire the three-dimensional continuous coordinate sequence of the physical simulation entity within a preset sampling time. The three-dimensional continuous coordinate sequence corresponds to the preset hovering spatiotemporal window of the physical simulation entity before contacting the simulation interface. The multimodal perception module eliminates systematic position noise in the three-dimensional continuous coordinate sequence through a multi-sensor fusion mechanism.

[0029] The feature decoupling module is used to calculate the position variance and steady-state retention time of the three-dimensional continuous coordinate sequence in the Cartesian coordinate system, and to determine the intention feature parameters that represent the inference intention and have probabilistic attributes based on the position variance and steady-state retention time.

[0030] The policy mapping module is used to convert the intent feature parameters into the exploration rate decay factor of the reinforcement learning model, and to limit the policy search step size boundary of the reinforcement learning model in the multi-dimensional action space based on the exploration rate decay factor.

[0031] The instruction generation module is used to adjust the action sampling probability distribution of the reinforcement learning model based on the policy search step size boundary, and generate state transition discrimination instructions that correspond to the current inferred situation and satisfy global game consistency.

[0032] Preferably, it also includes a heterogeneous computing offloading bus; the heterogeneous computing offloading bus is used to identify the task attributes of the inference task, and when the task attribute is a linear rule matching task, it routes the inference task to the rule matching module based on the directed acyclic graph; the heterogeneous computing offloading bus is also used to route the inference task to the strategy mapping module when the task attribute is a nonlinear game task.

[0033] Preferably, the feature decoupling module determines the intention feature parameters by the following steps: Step S31, performing first-order difference processing on the three-dimensional continuous coordinate sequence to determine the transient displacement vector set of the physical deduction entity within the preset sampling time; Step S32, calculating the displacement change statistical features of the physical deduction entity based on the transient displacement vector set, and fitting the displacement change statistical features with the preset intention template to determine the intention feature parameters.

[0034] Preferably, the policy mapping module converts the exploration rate decay factor of the reinforcement learning model into the following steps: Step S41, establish a monotonic mapping relationship between the intention feature parameter and the neuron activation threshold; Step S42, when the intention feature parameter monotonically increases, decrease the neuron activation threshold to lock the local policy search path of the reinforcement learning model.

[0035] Preferably, the instruction generation module is used to acquire the global simulation data of the digital twin representing the simulation interface, inject the action sampling probability distribution into the global simulation data of the digital twin, and output an adjudication report containing the situation assessment score and decision weight.

[0036] Preferably, the policy mapping module further includes a prior weight bias unit; the prior weight bias unit is used to map the trajectory oscillation features of the physical deduction entity into the prior probability weight bias of the deep neural network, and to use the prior probability weight bias to perturb the initial feature vector of the reinforcement learning model in order to guide the model to predict the direction.

[0037] Preferably, it also includes a self-supervised evolution module; the self-supervised evolution module is used to extract the inference element set with timestamps and clean the inference element set into reinforcement learning training sample pairs with causal chain labels; the self-supervised evolution module is also used to drive the reinforcement learning model to perform adaptive iteration of model parameters based on the reinforcement learning training sample pairs.

[0038] Preferably, the multimodal perception module is equipped with an edge sensing node array, which includes an infrared matrix sensor and an ultrasonic ranging sensor, used to acquire spatial pose data of the physically deduced entity.

[0039] Preferably, the system is used for multi-domain collaborative tasks, including acquiring the state parameters of combat units in multi-dimensional combat domains, and performing coupled calculations on the state parameters of combat units according to state transition discrimination instructions to solve the global performance index.

[0040] A multimodal intelligent decision-making manual wargame simulation device is provided for running a multimodal intelligent decision-making manual wargame simulation system.

[0041] Example 1: In a campaign-level wargame scenario involving multi-domain force coordination and parameter coupling, the multimodal perception module captures the continuous three-dimensional coordinate sequence of the physical simulation entity before it comes into contact with the simulation interface through an edge sensor node array. For physical movement events, the feature decoupling module uses the timestamp of the detected physical simulation entity's landing point coordinates as a reference to retrieve a preset time window backward. The continuous three-dimensional coordinate set within the range is used to calculate the geometric space variance of the three-dimensional coordinate set relative to the landing point coordinates. The landing point coordinates are then converted into a baseline state feature vector. The reciprocal of the geometric space variance is defined as the confidence component of the state. Finally, these are concatenated to form a multi-dimensional state feature vector sequence.

[0042] The strategy mapping module receives a sequence of multi-dimensional state feature vectors. The strategy generation network uses the confidence component as a dynamic exploration adjustment parameter, adjusting the distribution variance hyperparameter in the action sampling function based on the confidence component. When the confidence component exceeds a preset deterministic threshold, the strategy generation network causes the sampling variance of the corresponding node's output action distribution to decay and converge to the extreme action vector to calculate the battle damage parameter and resource decay index. The strategy mapping module maps the physical trajectory variance to the neural network exploration hyperparameters, converting the physical residual into the prior probability weight bias of the multi-agent strategy network. The instruction generation module obtains the representation inference interface. The system generates global simulation data of the digital twin and writes the action sampling probability distribution into the global simulation data of the digital twin. It outputs an adjudication report containing situation assessment scores and decision weights. The decision tensor overwrites the underlying global situation map database and drives the situation rendering synchronization of the front-end digital twin visualization interface. The physical inference space and the logical operation environment form a closed loop. The system maintains millisecond-level operation convergence under the condition of concurrent interaction of physical entities. It converts the behavioral parameters in the physical space into sampling variance hyperparameters inside the reinforcement learning model. The calculation model performs calculations according to the confidence boundary. Physical uncertainty is converted into a reconstructed state of computing power convergence.

[0043] Example 2: In a digital twin wargame simulation platform simulating multi-domain adversarial scenarios, the experimental data originates from a heterogeneous game event sequence containing continuous trajectories in three-dimensional space generated by the simulation platform. The simulation environment sampling rate is set to 100Hz, and Gaussian white noise with a signal-to-noise ratio of 20dB is superimposed on the original coordinate sequence to simulate background disturbances at edge sensing nodes. The core parameter is preset to the sampling duration. The determination logic depends on the trade-off between sampling coverage and processor computational load; specifically, the selection rule is to decrease the value as the signal spectral bandwidth increases. The value of is set in this experiment. The sample group of this invention is set to 150ms, and includes a feature decoupling module and a strategy mapping module; a control group A is set to remove the feature decoupling module and only use discrete landing point coordinate input; and a sample group with a preset sampling duration is also included. In control group B, set at 500ms, the feature decoupling module acquires the original coordinate sequence and eliminates high-frequency disturbances using a multi-sensor fusion mechanism. The calculated position variance is 12.45, and the intention feature parameter with probabilistic attributes is determined. The policy mapping module calculates the exploration rate decay factor based on the intention feature parameter, adjusting the policy search step size of the reinforcement learning model from a globally undirected search state to a locally convergent state. Experimental measurements show that the model convergence step size of control group A is 4580 steps, while the model convergence step size of the sample group of this invention is reduced to 860 steps. Boundary test data of control group B indicate that when the preset sampling duration... After more than 200ms, the computational latency generated by the feature decoupling module surged from 5.2ms to 28.4ms, and the overall system response time broke through the 20ms real-time decision threshold. This proves that mapping the variance of the physical trajectory to the sampling variance hyperparameter inside the reinforcement learning model can constrain the search boundary in the multidimensional action space and realize the reconstruction of the uncertainty of physical behavior towards the convergence of computing power.

[0044] Example 3: In a manual wargaming environment simulating a large-scale joint exercise, when multiple physical simulation entities simultaneously move or switch positions on the sensing simulation interface, the multimodal perception module uses an edge sensing node array to collect a three-dimensional continuous coordinate sequence. This three-dimensional continuous coordinate sequence includes... There are discrete time sampling points, where... The values ​​are positive integers. The spatial resolution of the edge sensor node array is 1 mm, and the sampling frequency is set to 100 Hz. For the acquired raw data, a trajectory compensation procedure based on kinematic state transition is applied. Infrared and ultrasonic sensor signals are prone to spatial occlusion and transient signal loss when the physical operator's limbs intervene. When the signal amplitude received by any edge sensor node drops below 1.5 times the idle noise baseline and lasts for more than 20 milliseconds, the processor determines that spatial occlusion has occurred, extracts the continuous coordinate set within the 50-millisecond window before the signal breakpoint, calculates the extrapolated displacement vector based on the uniformly accelerated spatial motion model, and continuously writes the calculated virtual interpolated coordinates into the breakpoint interval, outputting a three-dimensional continuous coordinate sequence with equidistant time and spatial continuity. Since the 50-millisecond time window is much smaller than the typical delay time of human hand muscle reaction, within this extremely short scale, the motion state of the physically extrapolated entity is mainly dominated by limb inertia, exhibiting strong physical continuity. Therefore, the uniformly accelerated spatial motion model is used to extrapolate the breakpoint. This system can capture the movement trend of a hand the instant the sensor signal is lost. This prediction is not aimed at obtaining an absolutely precise landing point, but rather at maintaining the continuity of the feature vector flow within the logical operation environment. This prevents the reinforcement learning model from falling into inference deadlock due to sudden changes in the input signal. The calculated virtual interpolated coordinates serve as a transitional constraint, translating the short-term motion inertia of physical space into the state conservation of the computing system. This maintains the logical coherence of the twin potential evolution when the physical signal is momentarily obscured. The feature decoupling module, based on the Fitts Law principle in ergonomics—which objectively correlates the uncertainty of target selection in the operational space with the degree of divergence of the trajectory endpoints—determines the quantization processing path for the three-dimensional continuous coordinate sequence. The feature decoupling module calculates and determines the intention feature parameters, extracts the three-dimensional coordinate set of the physically deduced entity within a preset hovering time-space window, calculates the Euclidean distance of each coordinate point in the three-dimensional coordinate set relative to the center position of the final landing point, and determines the position variance as the arithmetic mean of the sum of squared distances. Finally, the duration of physical deduction entities within the position deviation threshold is statistically analyzed and determined as the steady-state retention time.

[0045] The strategy mapping module obtains the location variance. With steady-state retention time, the location variance is converted using the intention transformation operator. Mapping steady-state retention time to the exploration rate decay factor of the reinforcement learning model Combining the exponential decay mathematical model for regulating the convergence space in the control algorithm, the dimensionless coefficients are calculated using the following formula to establish a single closed numerical transformation relationship: , where η represents the dimensionless control variable of the exploration rate decay factor of the output, and its value range is forcibly limited to between 0 and 1; The preset baseline exploration rate is a dimensionless fixed constant, set to 0.1 based on the empirical value of the basic grid pathfinding test. The dimensionless bias constant representing the intention sensitivity coefficient is fixed at 0.5 based on the measured average muscle response lag parameter. The steady-state retention time represents the statistical output. The extracted location variance is represented by the transformation operator, which follows a negative correlation when the location variance... When the steady-state retention time decreases and the steady-state retention time increases, The value of decreases accordingly. In a specific set of computational data, when the position variance... Take 5.2 Furthermore, when the steady-state retention time is 120ms, the strategy mapping module determines the exploration rate decay factor according to the linear mapping rule. The exploration rate decay factor is 0.15. The variance constraint variable, acting as the action sampling function of the policy generation network, transforms the search step size of the reinforcement learning model in the multi-dimensional action space from a global search state to a local convergence state centered on the intent feature parameters. The instruction generation module reads the sampling probability distribution output by the policy generation network, injects the sampling probability distribution into the global simulation data of the digital twin, and outputs state transition discrimination instructions that satisfy global game consistency. The decision tensor synchronously overwrites the underlying global situation map database, completing the synchronous situation rendering of the digital twin visualization interface. The trajectory features of the physically inferred entities are mapped to the search boundary constraints of the neural network through a discretized calculation procedure. The system controls the decision response latency to within 15ms under multi-entity concurrent interaction conditions. The computational model dynamically adjusts the internal policy search parameters according to the state feedback of the physical space, completing the underlying reconstruction of behavior features converging to computing power.

[0046] Example 4: This example combines Figures 1 to 2 A description of a multimodal intelligent adjudication manual wargaming device and system, such as... Figure 1As shown, this invention provides a full-process logical implementation of multimodal perception interaction to decision constraints. First, the system captures the interactive actions generated by physically deduced entities in real time through a multimodal perception module. This module uses a built-in multi-sensor fusion mechanism to eliminate systematic position noise in the physical environment and outputs a set of three-dimensional continuous coordinate sequences corresponding to a preset hovering spatiotemporal window. Next, a feature decoupling module receives this sequence and calculates the position variance and steady-state retention time of the physical entity in space. By performing first-order difference processing on the coordinate sequence, the transient displacement vector is fitted with a preset intention template, thereby decoupling the intention feature parameters with probabilistic attributes. Subsequently, a policy mapping module maps the intention... Graph parameters are transformed into exploration rate decay factors for the reinforcement learning model. During this process, the prior weight bias unit simultaneously extracts trajectory oscillation features and maps them to the prior probability weight bias of the deep neural network, perturbing the initial feature vector of the model to determine the policy search step size boundary. Finally, the instruction generation module adjusts the action sampling probability distribution of the reinforcement learning model based on the boundary and injects global simulation data from the digital twin to generate state transition discrimination instructions that satisfy global game consistency. At the same time, the self-supervised evolution module performs causal chain cleaning on the inference elements with timestamps, driving the model parameters to undergo adaptive iterative evolution, forming a closed-loop logical flow from perception to decision.

[0047] like Figure 2 As shown, this invention provides a heterogeneous computing offloading bus and its task routing architecture. The system includes a heterogeneous computing offloading bus, serving as a core hub connecting multimodal perception, feature decoupling, and downstream computing units. This bus possesses task attribute recognition capabilities. Upon receiving an inference task, it extracts the task's semantic vector and calculates its complexity score. If identified as a linear rule matching task, the bus automatically routes the task flow to a rule matching module based on a directed acyclic graph, using a depth-first search algorithm for rapid decision-making, avoiding redundant computing power consumption. If identified as a nonlinear game task, the bus routes the task flow to a policy mapping module, triggering an artificial intelligence inference path based on deep reinforcement learning. This architecture, through a heterogeneous offloading mechanism, ensures that when facing high-frequency concurrent interactions, the computational load can be decoupled and distributed between the lightweight rule engine and the reconstructed deep learning model, thereby constructing a robust computing power barrier.

[0048] Example 5: In a deployment scenario adapted to edge sensing hardware, regarding the parameter determination of the intent conversion operator in the feature decoupling module, the system establishes a controlled engineering experiment to generate a feature mapping table. This procedure, under a baseline environment shielded from interference, uses mechanical devices to drive a physical deduction entity according to a preset... The standard trajectory generates displacement, where, The coordinates are positive integers. The multimodal perception module synchronously records the three-dimensional continuous coordinate sequence corresponding to the standard trajectory, and the feature decoupling module calculates the corresponding position variance. Finally, the calculation results are aligned with the preset action commands of the mechanical device, and the linear coefficients and bias terms in the intention conversion operator are obtained by least squares fitting, thus determining the exploration rate decay factor. With location variance The proportional relationship between them anchors the weight allocation logic within the operator to the physical range perceived by the hardware.

[0049] When the system is applied to mobile simulation scenarios involving electromagnetic environmental fluctuations, the multimodal sensing module initiates a baseline noise compensation procedure before receiving simulation commands. The sensor array samples at 100Hz for 5 seconds in an idle state without physical interaction. The processor acquires the idle sampling data and calculates the root mean square value of the signal amplitude as a systematic position noise benchmark for the current environment. This benchmark is then injected into the multi-sensor fusion mechanism. When the multimodal sensing module detects a drastic change in the signal amplitude of the infrared matrix sensor (determined as operator intervention), the system automatically reduces the confidence weight of this noise benchmark and calls a pre-stored dynamic interference correction coefficient in real time. This correction coefficient, determined based on prior experimental data of the electromagnetic characteristics of the hand, is used to offset local spatial field strength fluctuations caused by operator intervention. The multi-sensor fusion mechanism dynamically adjusts the observation noise covariance matrix of the Kalman filter operator. Increase in the hand intervention area The value of determines the position variance output by the feature decoupling module. It can effectively filter out random swaying noise generated by limbs, retaining only the physical residuals related to the deduction intention. By dynamically adjusting the covariance matrix parameters of the filtering operator, it can offset environmental background disturbances. At this time, the position variance output by the feature decoupling module is... Automatically subtract environmental noise components to reduce the exploration rate decay factor generated by the policy mapping module. The physical residuals caused by the inferrer's intent are used to maintain the policy convergence state of the reinforcement learning model under different background noise intensities.

[0050] Example 6: In deployment scenarios adapting to edge sensor arrays with different hardware specifications, the multimodal sensing module applies a sensor benchmark calibration procedure. This calibration procedure divides the physical projection interface into sections... A calibration matrix consisting of sampling units, wherein... This represents the number of horizontal grid cells. To determine the vertical grid number, standard calibration entities are sequentially placed at the center nodes of the calibration grid, and the edge sensor node array acquires the induced level signal at the center nodes. The processor is based on the sensed level signal The first derivative of the spatial distribution of the voltage level is calculated to determine the linear gain coefficient of the coordinate mapping operator, and the standard deviation of the voltage level fluctuation under no-load conditions is calculated by the processor. As the initial zero-point bias for calculating the position variance in the feature decoupling module, a spatial non-uniformity compensation table for the inductive inference interface is established by traversing and sampling the global grid, thus determining the linear correspondence between the physical space of the inference interface and the digital coordinate system. In applications configured to handle nonlinear game tasks, the policy mapping module applies an offline optimization parameter-finding procedure using a reinforcement learning model, which sets an exploration rate decay factor. The value range is from 0.1 to 0.9. In the simulation environment, a continuous three-dimensional coordinate sequence of the gradient of the inferred intention from 0% to 100% is generated. The processor records the convergence rate of the action sampling function under different confidence components and determines the slope of the mapping curve of the intention conversion operator. The slope k is used to calibrate the intent sensitivity coefficient. , The slope is a proportionality constant with a physical value range of -0.8 to -0.5. This range was determined through offline simulation tests on game scenarios of varying complexity. The stiffness that determines the transformation of physical space uncertainty into computational search constraints, when When the variance is less than -0.8, a small decrease in physical variance can lead to an excessive shrinkage of the search radius, causing the model to lose its ability to consider the global optimum when faced with complex trajectories with deceptive trajectories, resulting in inference deadlock; while when When the value is greater than -0.5, the physical feedback has insufficient guiding effect on the convergence of computing power, causing the system's response latency to exceed the 20-millisecond real-time threshold under multi-entity concurrent conditions. Therefore, selecting the range of -0.8 to -0.5 can maintain the computational overhead within the system's robustness barrier while ensuring game diversity. In a specific set of deployment parameters, when the processor identifies a confidence component value of 0.85, the policy generation network determines the appropriate response based on the slope of the mapping curve. The sampling variance bias of the action sampling function is calculated to be -1.2, which shrinks the search radius of the reinforcement learning model in the multi-dimensional action space by 45.6% around the predicted intention landing point. The calibrated parameter matrix is ​​stored in the underlying control register. The nonlinear mapping between the physical trajectory variance and the neural network hyperparameters is converted into numerical correspondence logic, and the system generates a mapping from the perceptual input to the model search constraints.

[0051] When processing hybrid inference task flows containing deterministic decision-making logic and uncertain game-theoretic decision-making, the heterogeneous computing offloading bus acquires the task semantic vector, extracts the number of decision factors in the task semantic vector, and calculates the task complexity score corresponding to the number of decision factors. The heterogeneous computing offloading bus performs keyword scanning on the input inference instruction flow, and counts the entity tags (such as troop size), spatial constraints (such as terrain obstacles), and logical operators (such as if...then...) contained therein.The frequency of occurrence of each decision factor is used to construct a multi-dimensional task semantic vector. The number of decision factors corresponds to the sum of the dimensions of the non-zero elements in the semantic vector. The task complexity score is derived from the normalized number of decision factors and is used to quantify the task's dependence on nonlinear computing resources. When the instruction stream is filled with a large number of simple position coordinate comparisons, the decision factors are highly concentrated in the deterministic dimension. Based on this, the system identifies it as a linear rule matching task, thereby achieving accurate routing of the computational load. When the task complexity score is lower than a preset linear threshold, the heterogeneous computing offloading bus identifies the current inference task as a linear rule matching task and routes it accordingly. The system then moves to the rule matching module based on a directed acyclic graph (DAG). This module uses a depth-first search algorithm to traverse the decision path composed of logical nodes and unidirectional connecting edges, outputting a decision conclusion that conforms to the preset tactical logic. When the task complexity score exceeds a preset linear threshold, the heterogeneous computing offloading bus determines the task attribute as a nonlinear game task and routes it to the policy mapping module to call the reinforcement learning model to process the multi-domain game computation. This task attribute-driven heterogeneous offloading mechanism reduces the system's computational overhead when processing basic decision logic. The instruction generation module obtains the sampling probability distribution output by the policy generation network and calculates the relationship between the sampling probability distribution and the numerical... The interaction entropy of historical state sequences in the digital twin global simulation data is used to determine the weight of the impact of action decisions on the consistency of the global game. In the specific calculation process, the instruction generation module treats the sampled probability distribution as a candidate action tensor for the current game situation. By convolving this tensor with the game rule operators pre-stored in the digital twin global simulation data, the feasibility of the action is evaluated across the entire graph. The calculation of interaction entropy is used to measure the semantic offset between the current sampled action and the historical optimal path. If the offset is within a preset consistency threshold, the impact weight is set to a high value, allowing the sampled distribution to directly modify the entity attributes in the digital twin environment. This transformation from local sampling probability to global consistency influence weights essentially establishes a conflict review mechanism at the computing power layer. This ensures that surface-level physical operations, after being solved by the AI ​​model, still conform to the overall evolutionary constraints of the overall deductive logic. When the influence weights meet preset coordination conditions, the instruction generation module translates the sampling probability distribution into state transition judgment instructions and generates an adjudication report containing logical unit addresses and state increment values. The decision tensor synchronously overwrites the underlying global situation map database, driving the front-end digital twin visualization interface to correspondingly render multi-dimensional situations, achieving closed-loop synchronization between the physical interaction space and the logical operation architecture under global game constraints.

[0052] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above, and that the present invention can be implemented in other specific forms without departing from the spirit or essential characteristics of the present invention.

[0053] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention.

Claims

1. A multi-modal intelligent adjudication wargame play system, characterized in that, include: The multimodal perception module is used to acquire the three-dimensional continuous coordinate sequence of the physical simulation entity within a preset sampling time. The three-dimensional continuous coordinate sequence corresponds to the preset hovering spatiotemporal window of the physical simulation entity before contacting the simulation interface. The multimodal perception module eliminates systematic position noise in the three-dimensional continuous coordinate sequence through a multi-sensor fusion mechanism. The feature decoupling module is used to calculate the position variance and steady-state retention time of the three-dimensional continuous coordinate sequence in the Cartesian coordinate system, and to determine the intention feature parameters that represent the inference intention and have probabilistic attributes based on the position variance and steady-state retention time. The quantization processing path for the three-dimensional continuous coordinate sequence is determined. The feature decoupling module calculates and determines the intent feature parameters, extracts the three-dimensional coordinate set of the physically deduced entity within the preset hovering spatiotemporal window, calculates the Euclidean distance of each coordinate point in the three-dimensional coordinate set relative to the center position of the final landing point, and determines the position variance as the arithmetic mean of the sum of squared distances. Finally, the duration of physical deduction entities within the position deviation threshold is statistically analyzed and determined as the steady-state retention time. The policy mapping module converts intent feature parameters into exploration rate decay factors for the reinforcement learning model and limits the policy search step size boundary of the reinforcement learning model in the multi-dimensional action space based on the exploration rate decay factors; the policy mapping module obtains the position variance. With steady-state retention time, the location variance is converted using the intention transformation operator. Mapping steady-state retention time to the exploration rate decay factor of the reinforcement learning model Combining the exponential decay mathematical model for regulating the convergence space in the control algorithm, the dimensionless coefficients are calculated using the following formula to establish a single closed numerical transformation relationship: ,in, The exploration rate decay factor, representing the output, is a dimensionless control variable whose value range is forcibly limited to between 0 and 1. The preset baseline exploration rate is a dimensionless fixed constant, set to 0.1 based on the empirical value of the basic grid pathfinding test. The dimensionless bias constant representing the intention sensitivity coefficient is fixed at 0.5 based on the measured average muscle response lag parameter. The slope of the mapping curve representing the intent conversion operator is used to calibrate the intent sensitivity coefficient. ,and It is a proportionality constant, and its physical value ranges from -0.8 to -0.5; The steady-state retention time represents the statistical output. The variance of the extracted location, the intention transformation operator follows a negative correlation, and the exploration rate decay factor. It serves as a variance constraint variable for the action sampling function of the policy generation network, enabling the search step size of the reinforcement learning model in the multidimensional action space to be transformed from a global search state to a local convergence state centered on the intent feature parameters. The instruction generation module is used to adjust the action sampling probability distribution of the reinforcement learning model based on the policy search step size boundary, and generate state transition discrimination instructions that correspond to the current inferred situation and satisfy global game consistency.

2. A multi-modal intelligent adjudication wargaming play system according to claim 1, wherein, It also includes a heterogeneous computing offloading bus; the heterogeneous computing offloading bus is used to identify the task attributes of the inference task, and when the task attribute is a linear rule matching task, it routes the inference task to the rule matching module based on the directed acyclic graph. The heterogeneous computing offloading bus is also used to route the inference task to the policy mapping module when the task attribute is a nonlinear game task.

3. A multi-modal intelligent adjudication wargaming play system according to claim 1, wherein, The feature decoupling module determines the intent feature parameters by the following steps: Step S31, performing first-order difference processing on the three-dimensional continuous coordinate sequence to determine the transient displacement vector set of the physical deduction entity within the preset sampling time; Step S32, calculating the displacement change statistical features of the physical deduction entity based on the transient displacement vector set, and fitting the displacement change statistical features with the preset intent template to determine the intent feature parameters.

4. A multi-modal intelligent adjudication wargaming play system according to claim 1, wherein, The policy mapping module converts the exploration rate decay factor of the reinforcement learning model into the following steps: Step S41, establish a monotonic mapping relationship between the intention feature parameter and the neuron activation threshold; Step S42, when the intention feature parameter monotonically increases, decrease the neuron activation threshold to lock the local policy search path of the reinforcement learning model.

5. The multimodal intelligent adjudication manual wargame simulation system according to claim 1, characterized in that, The instruction generation module is used to acquire the global simulation data of the digital twin representing the simulation interface, inject the action sampling probability distribution into the global simulation data of the digital twin, and output an adjudication report containing the situation assessment score and decision weight.

6. A multi-modal intelligent adjudication wargaming play system according to claim 1, wherein, The policy mapping module also includes a prior weight bias unit; the prior weight bias unit is used to map the trajectory oscillation features of the physical deduction entity into the prior probability weight bias of the deep neural network, and to use the prior probability weight bias to perturb the initial feature vector of the reinforcement learning model in order to guide the model to predict the direction.

7. A multi-modal intelligent adjudication wargaming play system according to claim 1, wherein, It also includes a self-supervised evolution module; the self-supervised evolution module is used to extract the inference element set with timestamps and clean the inference element set into reinforcement learning training sample pairs with causal chain annotations; the self-supervised evolution module is also used to drive the reinforcement learning model to perform adaptive iteration of model parameters based on the reinforcement learning training sample pairs.

8. A multi-modal intelligent adjudication wargaming play system according to claim 1, wherein, The multimodal perception module is equipped with an edge sensing node array, which includes an infrared matrix sensor and an ultrasonic ranging sensor, used to acquire spatial pose data of physically deduced entities.

9. A multi-modal intelligent adjudication wargaming play system according to claim 1, wherein, The system is used for multi-domain collaborative tasks, including acquiring the state parameters of combat units in multi-dimensional combat domains, and performing coupled calculations on the state parameters of combat units based on state transition discrimination instructions to solve global performance indicators.

10. A multimodal intelligent adjudication manual wargame simulation device, characterized in that, The multimodal intelligent adjudication manual wargaming device is used to run the multimodal intelligent adjudication manual wargaming system described in claim 1.