Security command system based on digital twin electronic sand table

By constructing a digital twin electronic sand table spatial model with access constraints and mean-field multi-agent reinforcement learning, the problem of automated decision-making in complex scenarios of existing security command systems has been solved. Stable and continuous risk situation assessment and collaborative scheduling of multiple security units have been achieved, thereby improving the intelligence level of the security command system.

CN121836263AInactive Publication Date: 2026-04-10SHANGHAI DUOYU INFORMATION TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-01-08
Publication Date
2026-04-10
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing security command systems based on digital twins or electronic sand tables struggle to achieve automated decision-making in complex scenarios. They lack a unified expression of real-time access constraints and dynamic lockdown status, and the computational complexity is high when multiple security units are coordinated and scheduled, making it difficult to form a stable and continuous risk situation assessment. Furthermore, command intentions and automated decision-making models are difficult to coordinate.

Method used

A digital twin electronic sand table spatial model with access constraints is constructed. A continuous risk situation is formed through time segmentation and spatial statistics. A collaborative decision-making process of multiple security units is introduced. Combined with mean field multi-agent reinforcement learning and time consistency judgment mechanism, stable generation and executable output are achieved.

Benefits of technology

It enhances the intelligence and stability of security command, enabling stable and continuous dispatching decisions in complex and dynamic scenarios, and improves the executability of collaborative dispatching of multiple security units and the accurate characterization of risk situations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121836263A_ABST
    Figure CN121836263A_ABST
Patent Text Reader

Abstract

The invention discloses a security command system based on a digital twin electronic sand table, and the system comprises a digital twin electronic sand table spatial modeling module which constructs a spatial model containing a regional structure, nodes, edges and constraints; the security state and event acquisition module is used for acquiring unit state and event information and mapping the unit state and event information; the reachable relation and neighborhood determination module is used for calculating a reachable relation based on constraints and determining an average field neighborhood; the risk situation and average field construction module is used for generating a node risk situation and forming an average field state representation; the command intention modulation module performs command constraint modulation on the average field state; the average field multi-agent reinforcement learning decision module is used for generating a scheduling action based on the state of the security unit and the modulated average field; and the time consistency judgment and action output module is used for judging the stability of the continuous decision period and outputting actions. According to the invention, the stability and execution efficiency of security command are improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of digital twinning and multi-agent reinforcement learning, and particularly relates to a digital twinning electronic sand table security command system. BACKGROUND

[0002] With the acceleration of urbanization and the increasing number of complex scenes such as large parks, transportation hubs and public venues, security command systems are gradually evolving from traditional manual patrol and static monitoring to digital and intelligent direction. In recent years, digital twinning technology and electronic sand table technology have been introduced into the security field. Through spatial modeling and visual presentation of real scenes, command personnel can intuitively understand the structure of the security area, personnel distribution and event situation in the virtual environment. At the same time, some systems begin to try to combine intelligent analysis algorithms to assist in scheduling security forces to improve overall response efficiency and prevention and control capabilities.

[0003] However, the existing security command systems based on digital twinning or electronic sand table still have obvious technical limitations in practical application. On the one hand, most existing systems focus on spatial visualization display, although they can present security areas, camera locations and personnel positions, the spatial model often lacks unified expression of real traffic constraints, access rules and dynamic containment states, resulting in subsequent scheduling analysis still relying on manual experience judgment, making it difficult to directly support automated decision-making in complex scenarios. On the other hand, the existing security systems mostly use rule triggering or simple weight superposition methods to handle security events, lack comprehensive modeling of event time evolution characteristics and spatial correlation, and are difficult to form stable and continuous risk situation assessment results, which may easily lead to response imbalance in multi-point concurrent event or high-frequency alarm scenarios.

[0004] In addition, for the problem of coordinated scheduling of multiple security units, existing technologies usually use fixed patrol strategies, preset rule allocation or single-agent optimization methods, which are difficult to depict the mutual influence relationship between multiple security units. When introducing intelligent decision-making algorithms such as reinforcement learning, decisions are often made directly based on individual states without fully considering the overall state distribution of neighboring security units, resulting in unstable strategies or a sharp increase in computational complexity when the number of security units increases. At the same time, existing systems generally lack a mechanism to integrate human command intentions into the decision-making process in a structured manner, and temporary strategy adjustments by command personnel are mostly achieved through manual intervention, making it difficult to effectively coordinate with automatic decision-making models.

[0005] Therefore, how to provide a digital twinning electronic sand table security command system is a problem that those skilled in the art need to solve. SUMMARY

[0006] An object of the present application is to provide a digital twin electronic sand table-based security command system. The present application constructs a space model containing traffic constraints based on a digital twin electronic sand table, limits security dispatching to real accessibility; forms a continuous risk situation by time segmentation and space statistics of security events, and modulates the average field state together with command intent, introduces a multi-security unit collaborative decision-making process; combines average field multi-agent reinforcement learning and time consistency determination mechanism, realizes stable generation and executable output of security dispatching in complex dynamic scenarios, and effectively improves the intelligent level and stability of security command.

[0007] According to an embodiment of the present application, a digital twin electronic sand table-based security command system comprises: A digital twin electronic sand table space modeling module is configured to construct a digital twin electronic sand table space model containing security area structure, traffic node set, connected edge set and traffic constraint information. A security state and event collection module is configured to collect security unit state information and security event information, and map the security event information to the traffic node set. An accessible relationship and neighborhood determination module is configured to calculate the accessible relationship between security units based on traffic constraint information, and determine a neighborhood security unit set participating in average field calculation. A risk situation and average field construction module is configured to generate risk situation data corresponding to traffic nodes, and construct an average field state representation. A command intent modulation module is configured to modulate the average field state representation based on command intent information. An average field multi-agent reinforcement learning decision module is configured to generate dispatching actions based on security unit state information and modulated average field state representation. A time consistency determination and action output module is configured to determine the consistency of the average field state representation in a continuous decision-making period and output dispatching actions.

[0008] Optionally, the modules are realized through the following methods: S1, a digital twin electronic sand table space model is constructed, which contains security area structure, traffic constraint relationship and position distribution of security units in space. S2, in each decision-making period, the state information of each security unit is obtained, and the corresponding security event information is obtained. S3, according to the traffic constraint relationship in the digital twin electronic sand table space model, the accessible relationship between security units is calculated, and the neighborhood security unit set participating in average field calculation is determined for each security unit based on the accessible relationship. S4, calculate the risk situation weight corresponding to each security unit, and weight the state information of the neighboring security units based on the risk situation weight to form an average field state representation of risk modulation; S5, receive the command intention information input by the command terminal, and convert the command intention information into an average field modulation parameter, limit the set of neighboring security units participating in the average field state representation and the corresponding state weight; S6, for each security unit, combine the state information of the security unit with the corresponding average field state representation to form a reinforcement learning state input, and calculate the corresponding scheduling action based on the improved average field multi-agent reinforcement learning strategy model; S7, based on the change result of the average field state representation in the continuous decision period, perform consistency judgment on the scheduling action, output the scheduling action corresponding to the last decision period when the change result meets the stability condition, and output the scheduling action corresponding to the current decision period when the change result does not meet the stability condition.

[0009] Optionally, the S1 specifically comprises: S11, obtain the basic space data of the security scene, the basic space data including building contour line, floor identification, wall boundary line, entrance position, access control point position, road center line, passable area boundary line and impassable area boundary line; S12, perform space coordinate unification processing on the basic space data to determine the origin of the unified coordinate system, the direction of the coordinate axis and the coordinate unit, and convert the coordinates of each geometric object in the basic space data to the unified coordinate system; S13, generate a passable unit set based on the wall boundary line, the passable area boundary line and the impassable area boundary line, the passable unit set being composed of a plurality of passable grid units, each passable grid unit corresponding to a continuous passable space, and any two passable grid units being non-overlapping in space; S14, perform adjacency relationship judgment on the passable grid units according to the road center line and the entrance position, and record the adjacency relationship between two passable grid units when the shared boundary of the two passable grid units does not coincide with the wall boundary line; S15, perform passable constraint annotation on the recorded adjacency relationship, the passable constraint annotation including passable direction identification, passable time period identification, access control passable rule identification and containment state identification; S16, construct a passable node set and a connected edge set based on the passable grid units and the adjacency relationship, the passable node set including node identification and spatial geometric center coordinates of the passable grid units, and the connected edge set including edge identification and corresponding passable constraint annotation between adjacent passable grid units; S17, obtain initial deployment data of the security unit set, the initial deployment data including security unit identification, security unit type identification, security unit initial position coordinates and security unit available state identification, and map the security unit initial position coordinates to node identification in the set of passage nodes to form a position distribution record of the security unit in the digital twin electronic sand table space model.

[0010] Optionally, the S2 specifically includes: S21, in each decision cycle, collect real-time position coordinates of each security unit, the real-time position coordinates being obtained by a positioning device and expressed in a unified coordinate system of the digital twin electronic sand table space model; S22, in the decision cycle, collect motion state data of each security unit, the motion state data including current moving speed, moving direction identification and static state identification; S23, in the decision cycle, collect task state data of each security unit, the task state data including current execution task type identification, task remaining time length and task occupation state identification; S24, in the decision cycle, collect security and protection event data, the security and protection event data including event identification, event type identification, event occurrence time and event occurrence position coordinates, the event occurrence position coordinates being mapped to passage node identification in the digital twin electronic sand table space model; S25, perform timestamp alignment processing on the real-time position coordinates, motion state data and task state data of the security unit, and perform time merging processing on the security and protection event data within the same decision cycle to form a security unit state information set and a security and protection event information set corresponding to the decision cycle.

[0011] Optionally, the S3 specifically includes: S31, based on the digital twin electronic sand table space model, read a set of passage nodes and a set of connected edges, each connected edge in the set of connected edges being associated with passage direction identification, passage time period identification, access control passage rule identification and containment state identification; S32, for each security unit, determine a starting passage node identification mapped with the current position of the security unit; S33, in the current decision cycle, perform constraint traversal calculation on the set of passage nodes according to the passage direction identification, passage time period identification, access control passage rule identification and containment state identification of the set of connected edges to generate a set of reachable passage nodes satisfying passage constraint conditions; S34, in the set of reachable passage nodes, further perform clipping on passage nodes according to a preset maximum passage level number to form a set of effective reachable passage nodes constrained by a space range; S35, in the set of effective accessible passage nodes, screening the passage node identifier mapped with the security unit, and recording the corresponding security unit identifier to form a candidate neighborhood security unit set corresponding to the security unit; S36, performing deduplication and time consistency screening on the candidate neighborhood security unit set, eliminating the security unit identifier that has not changed position within the continuous decision period to form a neighborhood security unit set participating in the average field calculation, and the neighborhood security unit set does not contain the security unit identifier corresponding to the security unit.

[0012] Optionally, the S4 specifically comprises: S41, in the current decision period, based on the formed security event information set, time segmenting the security event according to the event occurrence time, and counting the corresponding security event occurrence number of each passage node in each time segment in the digital twin electronic sand table space model; S42, according to the event type identifier of the security event, configuring the corresponding event influence coefficient for different event types, and performing weighted accumulation of the event occurrence number and the corresponding event influence coefficient in each time segment to generate the node risk sequence corresponding to each passage node; S43, performing time smoothing processing on the node risk sequence to obtain the node risk value reflecting the risk situation of the current decision period, and writing the node risk value into the digital twin electronic sand table space model as risk situation data; S44, for each security unit, in the range of passage nodes corresponding to the neighborhood security unit set, according to the node risk value of each passage node and the spatial distance from the security unit to the corresponding passage node, calculating the risk situation weight corresponding to the neighborhood security unit; S45, based on the risk situation weight, performing hierarchical weighted statistical processing on the state information of each security unit in the neighborhood security unit set to form an average field state representation corresponding to the security unit, and the average field state representation contains the weighted distribution result of the neighborhood security unit state under different risk levels.

[0013] Optionally, the S5 specifically comprises: S51, in the current decision period, receiving the command intention information input by the command terminal, and the command intention information includes the defense area control level identifier, the security strategy stage identifier and the security unit scheduling priority identifier; S52, according to the defense area control level identifier, dividing the passage nodes in the digital twin electronic sand table space model into several command control levels, and assigning corresponding node modulation coefficients to different command control levels; S53, performing phase mapping processing on the node risk value according to the security policy phase identification, to generate a phase risk modulation sequence corresponding to the current security policy phase; S54, performing joint modulation calculation on the risk situation weight based on the node modulation coefficient and the phase risk modulation sequence, to generate a command intention modulation weight set varying with the decision cycle; S55, within the range of the neighbor security unit set, performing reconstruction processing on the state information of the neighbor security unit according to the command intention modulation weight set, to form an average field state representation containing command control level information and policy phase information.

[0014] Optionally, the S6 specifically includes: S61, for each security unit, obtaining the security unit state information of the security unit within the current decision cycle, and obtaining the average field state representation corresponding to the security unit, while reading the historical state components reserved in at least one historical decision cycle, performing splicing processing on the security unit state information, the average field state representation and the historical state components in accordance with a unified state dimension order, to form an initial joint state vector; S62, on the basis of constructing the initial joint state vector, performing credibility evaluation processing on the average field state representation within the current decision cycle, calculating the change degree of the average field state representation and the average field state representation corresponding to the previous decision cycle in each statistical dimension, to generate a credibility index representing the stability of the average field; S63, according to the credibility index, performing weight redistribution processing on the average field state component and the security unit state component in the initial joint state vector, reducing the weight of the average field state component relative to the security unit state component when the credibility index is lower than a preset threshold, and increasing the weight of the average field state component relative to the security unit state component when the credibility index is not lower than the preset threshold, to form a decision state vector after credibility modulation; S64, inputting the decision state vector after credibility modulation into the average field conditioned policy network corresponding to the security unit, the average field conditioned policy network outputting an action evaluation result for a preset action set, the action evaluation result including an action benefit evaluation component and an action uncertainty evaluation component; S65, according to the action uncertainty evaluation component, performing risk consistency screening on the preset action set, to eliminate actions with action uncertainty higher than a preset uncertainty threshold, to obtain an initial candidate action set, and performing reachability constraint screening in the initial candidate action set in combination with the reachable passing node set, to form a final candidate action set; S66, in the final candidate action set, determine the scheduling action of the security unit based on the joint ranking result of the action benefit evaluation component and the action uncertainty evaluation component, and map the scheduling action to a passing node sequence or node residence identifier in the digital twin electronic sand table space model; S67, store the passing node sequence or node residence identifier corresponding to the scheduling action and the decision cycle identifier in association, and perform update processing on the historical state component according to a preset time decay factor.

[0015] Optionally, the S7 specifically comprises: S71, in the current decision cycle, read the mean field state representation and the mean field state representation corresponding to the previous decision cycle, and perform difference calculation on the two in the corresponding state dimension to obtain a mean field state change amount; S72, compare the mean field state change amount with a preset stability determination threshold to generate a stability determination result corresponding to the current decision cycle; S73, when the stability determination result shows that the mean field state change amount does not exceed the stability determination threshold, select the scheduling action corresponding to the previous decision cycle as the output scheduling action of the current decision cycle; S74, when the stability determination result shows that the mean field state change amount exceeds the stability determination threshold, select the scheduling action calculated according to the mean field multi-agent reinforcement learning in the current decision cycle as the output scheduling action; S75, map the output scheduling action to a passing node sequence or node residence identifier in the digital twin electronic sand table space model, and issue the output scheduling action to the corresponding security unit for execution, while recording the output scheduling action and the decision cycle identifier to form an action output record.

[0016] The beneficial effects of the present application are: The application maps the space structure, the passing rule and the dynamic control state in the real security scene into a computable space constraint model, so that the scheduling decision of the multi-security unit is directly limited by the physical accessibility, and the executable of the scheduling scheme in the complex scene is fundamentally improved; at the situation awareness level, the application constructs a node risk sequence based on the time segmentation statistics of the security event and the event type influence coefficient, and forms continuous and stable risk situation data through time smoothing, introduces the time evolution characteristics and the space distribution characteristics of the risk into the average field construction process at the same time, and realizes the accurate description of the key area and the high-risk area; at the collaborative decision level, the statistical state of the neighborhood security unit is embedded into the multi-agent reinforcement learning strategy generation process in the form of average field, and through the joint modulation mechanism of the risk situation and the command intention, the automatic decision can respond to the objective risk change and reflect the strategy stage and the control level of the command personnel, and realize the collaborative scheduling under the human-machine collaborative constraint; at the same time, the consistency of the average field state change in the continuous decision period is determined, the frequent switching of the scheduling action under the small state fluctuation is effectively inhibited, and the stability and continuity of the security command scheme are significantly improved. In summary, the application has achieved significant improvement in the aspects of space modeling accuracy, situation awareness continuity, multi-security unit collaborative decision-making ability and command stability, and can effectively adapt to the actual application requirements of complex dynamic security scenes. BRIEF DESCRIPTION OF DRAWINGS

[0017] The accompanying drawings are included to provide a further understanding of the application, and constitute a part of the specification, together with the embodiments of the application, to explain the application, and do not constitute a limitation on the application. In the drawings: Fig. 1 A structure schematic diagram of a security command system based on a digital twin electronic sand table is provided for the application; Fig. 2 A work flow chart of a security command system based on a digital twin electronic sand table is provided for the application; Fig. 3 A flowchart of a security command system based on a digital twin electronic sand table is provided for the application. DETAILED DESCRIPTION

[0018] The application will now be further described in detail in conjunction with the drawings. These drawings are all simplified schematic diagrams, and only illustrate the basic structure of the application in a schematic manner, and therefore only show the components related to the application.

[0019] Reference Figs. 1-3 A security command system based on a digital twin electronic sand table, comprising: The digital twin electronic sand table space modeling module is configured to construct a digital twin electronic sand table space model containing a security area structure, a set of access nodes, a set of connected edges, and access constraint information. The security state and event collection module is configured to collect security unit state information and security event information, and map the security event information to the set of access nodes. The reachable relationship and neighborhood determination module is configured to calculate the reachable relationship between security units based on the access constraint information, and determine a set of neighborhood security units participating in average field calculation. The risk situation and average field construction module is configured to generate risk situation data corresponding to the access nodes, and construct an average field state representation. The command intention modulation module is configured to modulate the average field state representation based on command intention information. The average field multi-agent reinforcement learning decision module is configured to generate a scheduling action based on the security unit state information and the modulated average field state representation. The time consistency determination and action output module is configured to determine the consistency of the average field state representation in a continuous decision cycle and output the scheduling action.

[0020] In this embodiment, the modules are implemented through the following methods: S1, a digital twin electronic sand table space model is constructed, which contains a security area structure, an access constraint relationship, and a position distribution of security units in space. S2, in each decision cycle, the state information of each security unit is obtained, and the corresponding security event information is obtained. S3, according to the access constraint relationship in the digital twin electronic sand table space model, the reachable relationship between security units is calculated, and a set of neighborhood security units participating in average field calculation is determined for each security unit based on the reachable relationship. S4, the risk situation weight corresponding to each security unit is calculated, and the state information of the neighborhood security units is weighted based on the risk situation weight, forming a risk-modulated average field state representation. S5, the command intention information input by the command terminal is received, and the command intention information is converted into average field modulation parameters, which limit the set of neighborhood security units participating in the average field state representation and the corresponding state weight. S6, for each security unit, the state information of the security unit and the corresponding average field state representation are combined to form a reinforcement learning state input, and the corresponding scheduling action is calculated based on an improved average field multi-agent reinforcement learning strategy model. S7, based on the change result of the average field state representation in the continuous decision period, performing consistency judgment on the scheduling action, outputting the scheduling action corresponding to the last decision period when the change result meets the stability condition, and outputting the scheduling action corresponding to the current decision period when the change result does not meet the stability condition.

[0021] The present application aims at the problem that the multi-security unit cooperative decision in the digital twin electronic sand table security command scene is easily affected by group state fluctuation, risk situation change and command strategy switching. A decision mechanism with structural innovation is proposed under the average field multi-agent reinforcement learning framework. First, the average field credibility evaluation and dynamic weight modulation mechanism are introduced to quantitatively judge the stability of the average field state representation, and accordingly the influence proportion of the average field information in the decision input is adaptively adjusted to avoid the amplification effect of average field statistical distortion on the scheduling result. Second, the action evaluation result is divided into two components of reward evaluation and uncertainty evaluation in the action generation stage, and the risk consistency screening mechanism is introduced before the reachability constraint to make the scheduling action meet the requirements of risk stability and spatial executability. In addition, the historical scheduling state is managed through the time decay memory mechanism, so that the decision process considers both the recent scheduling trend and the real-time situation change, effectively suppressing the strategy oscillation under multi-decision period.

[0022] In this embodiment, S1 specifically comprises: S11, acquiring basic space data of a security scene, the basic space data including building contour lines, floor identifiers, wall boundary lines, entrance positions, access control point positions, road center lines, passable area boundary lines, and impassable area boundary lines; S12, performing space coordinate unification processing on the basic space data to determine a unified coordinate system origin, coordinate axis direction, and coordinate unit, and converting coordinates of each geometric object in the basic space data to the unified coordinate system; S13, generating a passable unit set based on the wall boundary lines, passable area boundary lines, and impassable area boundary lines, the passable unit set being composed of a plurality of passable grid units, each passable grid unit corresponding to a continuous passable space, and any two passable grid units being spatially non-overlapping; S14, performing adjacency relationship judgment on the passable grid units according to the road center lines and the entrance positions, and recording the adjacency relationship between two passable grid units when they share a boundary and the shared boundary does not coincide with the wall boundary line; S15, performing passable constraint labeling on the recorded adjacency relationship, the passable constraint labeling including passable direction identifiers, passable time period identifiers, access control passable rule identifiers, and containment state identifiers; S16, construct a set of passing nodes and a set of connected edges based on the passing grid unit and the adjacency relationship, the set of passing nodes containing node identifiers and spatial geometric center coordinates of the passing grid unit, and the set of connected edges containing edge identifiers and corresponding passing constraint annotations between adjacent passing grid units; S17, obtain initial deployment data of the set of security units, the initial deployment data containing security unit identifiers, security unit type identifiers, security unit initial position coordinates, and security unit available state identifiers, and map the security unit initial position coordinates to node identifiers in the set of passing nodes to form a position distribution record of the security units in the digital twin electronic sand table space model.

[0023] In the present application, the digital twin electronic sand table space model takes the passing grid unit, the set of passing nodes, and the set of connected edges as a unified spatial expression basis. Through the structured annotation of the passing direction, the access control rule, the time period restriction, and the containment state, the space model has reachability judgment ability and dynamic constraint bearing ability in the construction stage. The space model, after being generated, serves as the only spatial basis for subsequent security unit reachability calculation, neighborhood division, and mean field state statistics, thereby ensuring that the multi-security unit decision-making process is executed under consistent spatial constraint conditions, and realizing stable coupling between the digital twin space structure and the collaborative decision-making algorithm.

[0024] In the present embodiment, the S2 specifically includes: S21, in each decision-making period, collect real-time position coordinates of each security unit, the real-time position coordinates being obtained by a positioning device and represented in a unified coordinate system of the digital twin electronic sand table space model; S22, in the decision-making period, collect motion state data of each security unit, the motion state data containing current moving speed, moving direction identifier, and static state identifier; S23, in the decision-making period, collect task state data of each security unit, the task state data containing current execution task type identifier, task remaining time length, and task occupation state identifier; S24, in the decision-making period, collect security event data, the security event data containing event identifier, event type identifier, event occurrence time, and event occurrence position coordinates, the event occurrence position coordinates being mapped to passing node identifiers in the digital twin electronic sand table space model; S25, perform timestamp alignment processing on the real-time position coordinates, the motion state data, and the task state data of the security unit, and perform time merging processing on the security event data within the same decision-making period to form a set of security unit state information and a set of security event information corresponding to the decision-making period.

[0025] The application collects the position, movement and task state of the security unit in each decision cycle by using a unified time reference, and realizes the alignment of multi-source state data under the same space semantics by taking the passing node in the digital twin electronic sand table space model as a space index; at the same time, the security events are merged into the corresponding decision cycle according to the occurrence time, and the mapping of the event position and the passing node is completed, thereby forming a state information set with time consistency and space consistency.

[0026] In the embodiment, the S3 specifically comprises: S31, reading a passing node set and a connected edge set based on the digital twin electronic sand table space model, each connected edge in the connected edge set being associated with a passing direction identifier, a passing time period identifier, an access control passing rule identifier and an enclosed state identifier; S32, determining, for each security unit, a starting passing node identifier mapped with the current position of the security unit; S33, in the current decision cycle, performing constraint traversal calculation on the passing node set according to the passing direction identifier, the passing time period identifier, the access control passing rule identifier and the enclosed state identifier of the connected edge set, to generate a set of reachable passing nodes meeting the passing constraint condition; S34, in the set of reachable passing nodes, further pruning the passing nodes according to a preset maximum passing level number to form a set of effective reachable passing nodes constrained by a space range; S35, in the set of effective reachable passing nodes, screening the passing node identifiers mapped with the security units and recording the corresponding security unit identifiers to form a set of candidate neighborhood security units corresponding to the security unit; S36, performing deduplication and time consistency screening on the set of candidate neighborhood security units, eliminating the security unit identifiers that have not changed the position in the continuous decision cycle to form a set of neighborhood security units participating in the average field calculation, and the set of neighborhood security units does not contain the security unit identifier corresponding to the security unit.

[0027] In the application, the constraint graph composed of the passing nodes and the connected edges is used as the basic structure in each decision cycle, the reachable relationship is periodically updated in combination with the decision cycle identifier, and the space range pruning rule is introduced synchronously in the constraint traversal process, so that the set of reachable passing nodes is limited within the space scale related to the current decision; at the same time, the time consistency screening is performed on the candidate neighborhood security units to form a stable neighborhood set, so that the subsequent average field state statistics are consistent in both space and time dimensions, thereby ensuring the stable execution of the multi-security unit collaborative decision process.

[0028] In the embodiment, the S4 specifically comprises: S41, in the current decision period, based on the formed security event information set, the security events are time segmented according to the time when the events occur, and the number of security events corresponding to each passing node in each time segment is counted in the digital twin electronic sand table space model; S42, according to the event type identification of the security event, the corresponding event influence coefficient is configured for different event types, and the number of events and the corresponding event influence coefficient are weighted and accumulated in each time segment to generate the node risk sequence corresponding to each passing node; S43, the node risk sequence is executed time smoothing processing, the node risk value reflecting the risk situation of the current decision period is obtained, and the node risk value is written into the digital twin electronic sand table space model as risk situation data; S44, for each security unit, in the passing node range corresponding to the neighbor security unit set, according to the node risk value of each passing node and the spatial distance from the security unit to the corresponding passing node, the risk situation weight corresponding to the neighbor security unit is calculated; S45, based on the risk situation weight, the state information of each security unit in the neighbor security unit set is executed hierarchical weighted statistical processing to form the average field state representation corresponding to the security unit, and the average field state representation includes the weighted distribution result of the neighbor security unit state under different risk levels.

[0029] In the application, the event influence coefficient is configured according to the event type identification of the security event, the event level corresponding to the event type is determined by the pre-established event classification table, or the influence weight set is configured based on the statistical result of the historical security event; The time smoothing processing is executed for the node risk sequence corresponding to the passing node, which is used to weaken the instantaneous influence of the burst event on the risk situation evaluation in a single decision period, so as to form the risk situation data reflecting the overall trend of the continuous decision period; The spatial distance is calculated based on the spatial geometric center coordinates of the passing node in the digital twin electronic sand table space model, the spatial distance reflects the spatial position relationship of the security unit to the corresponding passing node under the constraint of passing, and is used as a spatial constraint factor in the generation of the average field state representation.

[0030] In the embodiment, the S5 specifically includes: S51, in the current decision period, the command intention information input by the command terminal is received, and the command intention information includes the defense area control level identification, the security strategy stage identification and the security unit scheduling priority identification; S52, according to the defense area control level identification, the passing nodes in the digital twin electronic sand table space model are divided into several command control levels, and the corresponding node modulation coefficient is allocated for different command control levels; S53, performing phase mapping processing on the node risk value according to the identified security policy phase, to generate a phase risk modulation sequence corresponding to the current security policy phase; S54, performing joint modulation calculation on the risk situation weight based on the node modulation coefficient and the phase risk modulation sequence, to generate a set of command intention modulation weights varying with decision cycles; S55, within the range of the set of neighborhood security units, performing reconstruction processing on the state information of the neighborhood security units according to the set of command intention modulation weights, to form an average field state representation containing command and control level information and strategy phase information.

[0031] In the present application, the command and control level is generated by the command terminal according to the preset security plan or the defense area level configuration, and is written into the digital twin electronic sand table space model with the passable node as the smallest control unit; the node modulation coefficient is updated with the decision cycle, and the joint modulation with the phase risk modulation sequence is completed under the same time reference, so as to form a modulation weight set that can dynamically change with the security policy phase; the average field state representation keeps the neighborhood structure unchanged in the reconstruction process, and only the weight distribution of the security unit state is updated, so as to ensure that the command intention participates in the subsequent reinforcement learning decision without destroying the multi-agent collaborative structure.

[0032] In the present embodiment, the S6 specifically comprises: S61, for each security unit, obtaining the security unit state information of the security unit within the current decision cycle, and obtaining the average field state representation corresponding to the security unit, while reading the historical state components reserved in at least one historical decision cycle, performing splicing processing on the security unit state information, the average field state representation and the historical state components in accordance with a unified state dimension order, to form an initial joint state vector; S62, on the basis of constructing the initial joint state vector, performing credibility evaluation processing on the average field state representation within the current decision cycle, calculating the change degree of the average field state representation and the average field state representation corresponding to the previous decision cycle in each statistical dimension, to generate a credibility index representing the stability of the average field; S63, according to the credibility index, performing weight redistribution processing on the average field state component and the security unit state component in the initial joint state vector, reducing the weight of the average field state component relative to the security unit state component when the credibility index is lower than a preset threshold, and increasing the weight of the average field state component relative to the security unit state component when the credibility index is not lower than the preset threshold, to form a decision state vector after credibility modulation; S64, input the decision state vector modulated by the credibility into an average field conditioned policy network corresponding to the security unit, the average field conditioned policy network outputs action evaluation results for a preset action set, the action evaluation results include action return evaluation components and action uncertainty evaluation components; S65, perform risk consistency screening on the preset action set according to the action uncertainty evaluation components, eliminate actions with action uncertainty higher than a preset uncertainty threshold, obtain an initial candidate action set, and perform reachability constraint screening on the initial candidate action set in combination with a set of reachable passing nodes to form a final candidate action set; S66, determine a scheduling action of the security unit in the final candidate action set based on a joint ranking result of the action return evaluation components and the action uncertainty evaluation components, and map the scheduling action into a passing node sequence or a node residence identifier in a digital twin electronic sand table space model; S67, store the passing node sequence or the node residence identifier corresponding to the scheduling action in association with the decision cycle identifier, and perform update processing on the historical state components according to a preset time decay factor.

[0033] In the average field multi-agent reinforcement learning decision process of the application, in order to solve the problems of dynamic changes of neighborhood security units, frequent event triggering and stage distortion of group statistical information in the digital twin electronic sand table security scene, the application introduces a decision mechanism combining average field credibility modulation and risk consistency constraints in the decision step. In each decision cycle, the system not only constructs a joint state vector containing the individual state of the security unit, the average field statistical state and the historical state component, but also performs stability analysis on the average field state representation in the current decision cycle and the adjacent historical cycle, and generates a credibility index reflecting the reliability of the average field statistical result. The credibility index is used to dynamically adjust the influence proportion of the average field state component in the decision input, so that the average field information fully participates in the collaborative decision when the neighborhood structure is stable, and is automatically weakened when the neighborhood changes rapidly or the statistics are insufficient, thereby avoiding the amplification effect of average field error in the scheduling decision. At the same time, the application performs structured processing on the output results of the strategy network in the action generation stage, splits the evaluation results of the scheduling action into action revenue evaluation components and action uncertainty evaluation components, and introduces the uncertainty evaluation results into the candidate action screening process. The system preferentially eliminates actions with high uncertainty under the current risk situation and spatial constraint conditions, and then performs reachability constraint verification on the remaining actions, thereby ensuring that the generated scheduling action meets the requirements in terms of risk consistency and spatial executability. In addition, the application manages the historical state by introducing a time decay memory mechanism, so that the decision process can perceive the recent scheduling trend and avoid strategy oscillation, significantly improving the stability and robustness in the multi-security unit continuous scheduling scene.

[0034] In the embodiment, the S7 specifically includes: S71, in the current decision cycle, read the average field state representation and the average field state representation corresponding to the last decision cycle, and perform difference calculation on the two in the corresponding state dimension to obtain the average field state change amount; S72, compare the average field state change amount with a preset stability determination threshold to generate a stability determination result corresponding to the current decision cycle; S73, when the stability determination result shows that the average field state change amount does not exceed the stability determination threshold, select the scheduling action corresponding to the last decision cycle as the output scheduling action of the current decision cycle; S74, when the stability determination result shows that the average field state change amount exceeds the stability determination threshold, select the scheduling action calculated according to the average field multi-agent reinforcement learning in the current decision cycle as the output scheduling action; S75, map the output scheduling action to a passing node sequence or a node residence identifier in the digital twin electronic sand table space model, and issue the output scheduling action to the corresponding security unit for execution while recording the output scheduling action and the decision cycle identifier to form an action output record.

[0035] In the present application, the calculation of the average field state change quantity is based on the numerical difference of each state component in the average field state representation within the adjacent decision cycle, and the alignment processing is completed in the same state dimension sequence to ensure the consistency of the change quantity calculation; the stability determination threshold is used as a configuration parameter associated with the security strategy stage, and different values can be used in different security operation stages to adapt to the risk situation change rate; in the output stage, the scheduling action first completes the corresponding relationship verification with the passing node in the digital twin electronic sand table space model, and then forms the action output record that can be issued, so as to ensure that the action output meets the security command requirements in terms of time stability and space executability.

[0036] Embodiment 1 In order to verify the feasibility and effectiveness of the present application in practical application, the security command system based on the digital twin electronic sand table proposed by the present application is applied to the security command scene of a large urban comprehensive transportation hub. The transportation hub includes a high-speed rail station, a subway interchange area, a commercial corridor and an underground parking area, with a total building area of about 320,000 square meters, a daily passenger flow of more than 250,000 people, a complex security area level, a large number of passing channels, and a dense event type. The traditional security scheduling mode based on manual command and fixed rules is prone to response delay, scheduling conflict and frequent instruction changes during peak hours, and it is difficult to meet the real-time and stability requirements.

[0037] In this scenario, first, based on the architectural design drawings, the access control system configuration table and the field surveying data, a digital twin electronic sand table space model is constructed, and each floor channel, entrance, gate area, closed channel and temporary control area in the hub is mapped to a passing node set, and a connected edge set is formed according to the wall boundary, access control rules and passing direction. Each passing node is bound with a space coordinate and a passing attribute, so that the scheduling path of the security unit is always constrained by the real space accessibility in the subsequent calculation. The security unit includes fixed post security personnel, patrol security personnel and mobile emergency security units, with a total number of 86, and the initial position is mapped to the passing node in real time through the positioning device.

[0038] The system takes 10 seconds as a decision-making period during operation, continuously collects the real-time position, motion state and task state of each security unit in each decision-making period, and accesses the security event data generated by the video analysis system, the access control alarm system and the manual reporting terminal. The security event is uniformly mapped to the passing node in the digital twin space, and is divided into congestion anomaly, personnel retention, intrusion into forbidden area, equipment anomaly and other categories according to the event type. The system performs time segmentation statistics on the security event according to the occurrence time in each decision-making period, and forms a node risk sequence by combining the event type influence coefficient, and generates continuous risk situation data through time smoothing processing, so as to avoid the risk assessment fluctuation caused by single burst event.

[0039] In the multi-security unit cooperative decision-making process, the system calculates the reachable relationship between security units according to the passing constraint relationship in the digital twin space, and determines the neighbor security unit set within the preset passing level range, only the spatially real cooperative security units are included in the mean field calculation. Subsequently, the system constructs the mean field state representation based on the state information of the neighbor security unit, the node risk value and the spatial distance relationship, and introduces the defense area control level and the security strategy stage information input by the command terminal to modulate the mean field weight, so that the key defense area obtains higher cooperative scheduling priority in the high risk stage.

[0040] In the scheduling action generation stage, the system splices the individual state of the security unit and the modulated mean field state representation to form a joint state vector, and inputs the mean field conditioned strategy network to calculate the candidate scheduling action. The generated scheduling action is filtered through the time consistency judgment mechanism before output, when the mean field state change amount in the continuous decision-making period is lower than the stability judgment threshold, the system keeps the scheduling action of the last period unchanged, so as to avoid frequent adjustment of security deployment in the case of high passenger flow but risk change is not significant, and improve the execution stability.

[0041] In order to verify the effectiveness of the method of the application, in the actual deployment test of continuous operation for seven days, the system of the application is compared with the traditional rule command system, and the experimental results are shown in Table 1: Table 1 Comparison of application effect of traffic hub security command system From the comparison results of Table 1, it can be seen that the system of the present application has obvious advantages in many key security command indicators. First, the average event response time is reduced from 48.6 seconds to 31.2 seconds, with a decrease of about 35.8%, indicating that the dispatching mechanism based on risk situation and average field collaborative decision can complete the security force reorganization faster. Second, the number of dispatching instruction changes during peak period is reduced from 19.4 times per hour to 7.1 times, indicating that the time consistency determination effectively suppresses the execution disturbance caused by frequent dispatching. The invalid dispatching ratio of security unit and the path conflict occurrence rate are reduced to 5.3% and 3.9% respectively, reflecting that the traffic constraint modeling and reachable relationship calculation significantly improve the spatial executability of the dispatching scheme. At the same time, the success rate of multi-point concurrent event disposal is improved to 94.6%, and the number of manual intervention by command personnel is greatly reduced, indicating that the present application not only improves the automation level in complex scenarios, but also effectively reduces the command burden, and the overall security command efficiency and stability are significantly improved. The present application can combine digital twin space modeling, risk situation perception, command intention modulation and average field multi-agent reinforcement learning in complex dynamic security scenarios, ensure the spatial executability and decision stability, realize the efficient collaborative dispatching of multiple security units, and has good engineering implementability and significant practical application effect.

[0042] The above describes only the preferred embodiments of the present application, but the protection scope of the present application is not limited thereto, any person skilled in the art can make equivalent replacement or change according to the technical scheme and inventive concept of the present application within the technical range disclosed by the present application, which should be covered within the protection scope of the present application.

Claims

1. A digital twin-based electronic sand table security command system, characterized in that, The method comprises the following steps: A digital twin electronic sand table space modeling module is used to construct a digital twin electronic sand table space model containing security area structures, a set of traffic nodes, a set of connected edges, and traffic constraint information; A security state and event collection module is used to collect security unit state information and security event information, and to map the security event information to the set of traffic nodes; A reachable relationship and neighborhood determination module is used to calculate the reachable relationship between security units based on the traffic constraint information, and to determine a set of neighborhood security units participating in average field calculation; A risk situation and average field construction module is used to generate risk situation data corresponding to the traffic nodes, and to construct an average field state representation; A command intention modulation module is used to modulate the average field state representation based on command intention information; An average field multi-agent reinforcement learning decision module is used to generate scheduling actions based on security unit state information and the modulated average field state representation; A time consistency determination and action output module is used to determine the consistency of the average field state representation in consecutive decision cycles and output scheduling actions.

2. The digital-twin-based electronic sand table security command system according to claim 1, characterized in that, The modules are connected through the following methods: S1, a digital twin electronic sand table space model is constructed, which contains security area structures, traffic constraint relationships, and the location distribution of security units in space; S2, in each decision cycle, the state information of each security unit is obtained, and the corresponding security event information is obtained; S3, according to the traffic constraint relationship in the digital twin electronic sand table space model, the reachable relationship between security units is calculated, and based on the reachable relationship, a set of neighborhood security units participating in average field calculation is determined for each security unit; S4, the risk situation weight corresponding to each security unit is calculated, and the state information of the neighborhood security units is weighted based on the risk situation weight, forming a risk-modulated average field state representation; S5, the command intention information input by the command terminal is received, and the command intention information is converted into average field modulation parameters, which limit the set of neighborhood security units participating in the average field state representation and the corresponding state weight; S6, for each security unit, the state information of the security unit and the corresponding average field state representation are combined to form a reinforcement learning state input, and the corresponding scheduling action is calculated based on an improved average field multi-agent reinforcement learning strategy model; S7, based on the change result of the average field state representation in consecutive decision cycles, the consistency of the scheduling action is determined, and when the change result meets the stability condition, the scheduling action corresponding to the last decision cycle is output, and when the change result does not meet the stability condition, the scheduling action corresponding to the current decision cycle is output.

3. The digital-twin-based electronic sand table security command system according to claim 2, characterized in that, The S1 specifically comprises: S11, the basic space data of the security scene is obtained, which contains building contour lines, floor identifiers, wall boundary lines, entrance positions, access control point positions, road center lines, passable area boundary lines, and impassable area boundary lines; S12, performing spatial coordinate uniform processing on the basic space data to determine a uniform coordinate system origin, coordinate axis direction and coordinate unit, and converting coordinates of each geometric object in the basic space data to the uniform coordinate system; S13, generating a pass unit set based on the wall boundary line, passable area boundary line and impassable area boundary line, the pass unit set being composed of a plurality of pass grid units, each pass grid unit corresponding to a continuous passable space, and any two pass grid units not overlapping in space; S14, performing adjacency relationship judgment on the pass grid units according to the road center line and the entrance and exit positions, and recording the adjacency relationship between two pass grid units when they share a boundary and the shared boundary does not coincide with the wall boundary line; S15, performing pass constraint labeling on the recorded adjacency relationship, the pass constraint labeling including pass direction identification, pass time period identification, access control pass rule identification and containment state identification; S16, constructing a pass node set and a connected edge set based on the pass grid units and the adjacency relationship, the pass node set including node identification and spatial geometric center coordinates of the pass grid units, and the connected edge set including edge identification and corresponding pass constraint labeling between adjacent pass grid units; S17, obtaining initial deployment data of the security unit set, the initial deployment data including security unit identification, security unit type identification, security unit initial position coordinates and security unit available state identification, and mapping the security unit initial position coordinates to node identification in the pass node set to form a position distribution record of the security unit in the digital twin electronic sand table space model.

4. The digital-twin-based electronic sand table security command system according to claim 2, wherein, The S2 specifically includes: S21, in each decision cycle, collecting real-time position coordinates of each security unit, the real-time position coordinates being obtained by a positioning device and represented in a uniform coordinate system of the digital twin electronic sand table space model; S22, in the decision cycle, collecting motion state data of each security unit, the motion state data including current moving speed, moving direction identification and stationary state identification; S23, in the decision cycle, collecting task state data of each security unit, the task state data including current execution task type identification, task remaining time length and task occupation state identification; S24, in the decision cycle, collecting security event data, the security event data including event identification, event type identification, event occurrence time and event occurrence position coordinates, the event occurrence position coordinates being mapped to pass node identification in the digital twin electronic sand table space model; S25, performing timestamp alignment processing on the real-time position coordinates, motion state data and task state data of the security unit, and performing time merging processing on the security event data within the same decision cycle to form a security unit state information set and a security event information set corresponding to the decision cycle.

5. The digital-twin-based electronic sand table security command system according to claim 2, wherein, The S3 specifically includes: S31, reading a set of passing nodes and a set of connected edges based on the digital twin electronic sand table space model, each connected edge in the set of connected edges being associated with a passing direction identifier, a passing time period identifier, an access control passing rule identifier, and an enclosed state identifier; S32, for each security unit, determining a starting passing node identifier mapped with the current position of the security unit; S33, in the current decision-making period, performing constraint traversal calculation on the set of passing nodes according to the passing direction identifier, the passing time period identifier, the access control passing rule identifier, and the enclosed state identifier of the set of connected edges, to generate a set of reachable passing nodes satisfying the passing constraint condition; S34, in the set of reachable passing nodes, further pruning the passing nodes according to a preset maximum passing level number to form an effective set of reachable passing nodes constrained by a space range; S35, in the set of effective reachable passing nodes, screening passing node identifiers mapped with security units and recording corresponding security unit identifiers to form a candidate neighborhood security unit set corresponding to the security unit; S36, performing deduplication and time consistency screening on the candidate neighborhood security unit set to exclude security unit identifiers that have not changed positions in consecutive decision-making periods to form a neighborhood security unit set participating in average field calculation, the neighborhood security unit set not containing the security unit identifier corresponding to the security unit.

6. The digital-twin-based electronic sand table security command system according to claim 2, wherein, The S4 specifically includes: S41, in the current decision-making period, based on the formed security event information set, time segmenting the security events according to the event occurrence time, and in the digital twin electronic sand table space model, counting the number of security events corresponding to each passing node in each time segment; S42, according to the event type identifier of the security event, configuring a corresponding event influence coefficient for different event types, and performing weighted accumulation of the number of events and the corresponding event influence coefficient in each time segment to generate a node risk sequence corresponding to each passing node; S43, performing time smoothing processing on the node risk sequence to obtain a node risk value reflecting the risk situation of the current decision-making period, and writing the node risk value into the digital twin electronic sand table space model as risk situation data; S44, for each security unit, in the range of passing nodes corresponding to the neighborhood security unit set, calculating the risk situation weight corresponding to the neighborhood security unit according to the node risk value of each passing node and the spatial distance from the security unit to the corresponding passing node; S45, based on the risk situation weight, performing hierarchical weighted statistical processing on the state information of each security unit in the neighborhood security unit set to form an average field state representation corresponding to the security unit, the average field state representation including the weighted distribution result of the neighborhood security unit state at different risk levels.

7. The digital-twin-based electronic sand table security command system according to claim 2, wherein, The S5 specifically includes: S51, in the current decision-making period, receiving command intention information input by a command terminal, the command intention information including a defense area control level identifier, a security policy phase identifier, and a security unit dispatch priority identifier; S52, according to the defense area control level identifier, the passing node in the digital twin electronic sand table space model is divided into several command control levels, and the corresponding node modulation coefficient is allocated to different command control levels; S53, according to the security strategy stage identifier, the stage mapping processing is performed on the node risk value, and the stage risk modulation sequence corresponding to the current security strategy stage is generated; S54, based on the node modulation coefficient and the stage risk modulation sequence, the joint modulation calculation is performed on the risk situation weight, and the command intention modulation weight set changing with the decision cycle is generated; S55, in the range of the neighbor security unit set, the state information of the neighbor security unit is reconstructed according to the command intention modulation weight set, and the average field state representation containing command control level information and strategy stage information is formed.

8. The digital-twin-based electronic sand table security command system according to claim 2, wherein, The S6 specifically includes: S61, for each security unit, the security unit state information of the security unit in the current decision cycle is obtained, the average field state representation corresponding to the security unit is obtained, at least one historical state component reserved in at least one historical decision cycle is read, and the security unit state information, the average field state representation and the historical state component are spliced according to the unified state dimension order to form an initial joint state vector; S62, on the basis of constructing the initial joint state vector, the average field state representation is evaluated in the current decision cycle, the change degree of the average field state representation and the average field state representation corresponding to the last decision cycle in each statistical dimension is calculated, and a credibility index representing the stability of the average field is generated; S63, according to the credibility index, the weight of the average field state component and the security unit state component in the initial joint state vector is redistributed, the weight of the average field state component relative to the security unit state component is reduced when the credibility index is lower than a preset threshold, the weight of the average field state component relative to the security unit state component is increased when the credibility index is not lower than the preset threshold, and a decision state vector after credibility modulation is formed; S64, the decision state vector after credibility modulation is input into the average field conditioned strategy network corresponding to the security unit, the average field conditioned strategy network outputs the action evaluation result for the preset action set, and the action evaluation result includes action benefit evaluation component and action uncertainty evaluation component; S65, according to the action uncertainty evaluation component, the risk consistency screening is performed on the preset action set, the action with the action uncertainty higher than the preset uncertainty threshold is removed, the initial candidate action set is obtained, and the reachability constraint screening is performed in the initial candidate action set combined with the reachable passing node set, and the final candidate action set is formed; S66, in the final candidate action set, the scheduling action of the security unit is determined based on the joint sorting result of the action benefit evaluation component and the action uncertainty evaluation component, and the scheduling action is mapped into the passing node sequence or the node residence identifier in the digital twin electronic sand table space model; S67, store the passing node sequence or node residence identification corresponding to the scheduling action in association with the decision cycle identification, and perform update processing on the historical state component according to a preset time decay factor.