Multimodal emergency human-machine interaction intelligent decision method and system
By employing a multimodal emergency human-computer interaction intelligent decision-making method, the problems of information distortion and decision delay caused by single-modal command input are solved. Through a multi-level cognitive state map and a human-computer bidirectional conditional probability inference framework, a consistent decision intention distribution is generated, and the decision granularity is dynamically adjusted, thus achieving efficient and reliable human-computer collaborative decision-making in emergency scenarios.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- ZHONGKESOFTGOLD (BEIJING) SOFTWARE TECHNOLOGY CO LTD
- Filing Date
- 2026-06-18
- Publication Date
- 2026-07-28
AI Technical Summary
In existing emergency human-computer interaction technologies, single-modal command input leads to information distortion or incompleteness, lacks the utilization of multimodal information complementarity, the decision-making system fails to adjust the decision granularity in real time, ignores changes in the operator's cognitive state, resulting in operational errors and decision delays, and poor human-computer collaboration.
By employing a multimodal emergency human-computer interaction intelligent decision-making method, semantic association analysis is performed to extract complementary and redundant information, construct a multi-level cognitive state map, calculate cognitive occupancy, establish a human-computer bidirectional conditional probability inference framework, generate a decision intent distribution vector for consistency verification, dynamically construct multi-granularity decision schemes, and generate decision outputs in conjunction with resource scheduling constraints.
It significantly improves the accuracy and efficiency of human-machine collaborative decision-making in emergency scenarios, adapts to decision-making needs under different cognitive states, reduces decision-making bias, and improves the compliance and reliability of decisions.
Smart Images

Figure CN122471352A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of emergency human-computer interaction technology, and in particular to a multimodal emergency human-computer interaction intelligent decision-making method and system. Background Technology
[0002] In the field of emergency human-computer interaction decision-making, existing conventional practices typically rely on single-modal command input, such as voice commands or gesture recognition, and perform response matching through a pre-set rule base. Such methods simplify the operator's intent into discrete triggering events, ignoring the semantic complementarity between multimodal information. For example, voice commands can be distorted by environmental noise, and gestures can be incomplete due to visual obstruction; however, joint analysis of both can improve the robustness of intent recognition. Current methods for fusing multimodal features mostly employ weighted summation or feature concatenation, lacking the removal of redundant information between modalities and the extraction of complementary information. This results in semantic ambiguity in the fused representation, making it difficult to accurately reflect the operator's true needs in high-pressure emergency scenarios.
[0003] Conventional approaches to decision-making typically employ fixed process templates or rigid knowledge bases, directly mapping predefined response plans to historical event types. This approach fails to consider the changing cognitive states of operators, such as fluctuations in attention allocation and working memory load over time. When the decision-making system outputs a solution, it does not assess the operator's ability to effectively process complex information, thus easily leading to operational errors or decision delays under high-load emergency situations. Existing technologies also fail to collaboratively model the constraints between humans and machines, often optimizing the behavior of both independently, lacking a holistic understanding of intent. Figure 1 The lack of consistent joint verification has led to frequent human-machine conflicts or misunderstandings.
[0004] Other methods attempt to monitor operator cognitive load through physiological signals, but only use this as a post-event evaluation indicator, without applying it to real-time adjustments to decision granularity or solution complexity. Most systems still output operational steps at a fixed granularity, ignoring the fundamental constraint of limited operator cognitive resources. When cognitive load is saturated, providing overly detailed decision-making solutions exacerbates cognitive overload; when cognitive load is low, coarse solutions lack necessary guidance. This static decision output limits the adaptability and efficiency of human-machine collaboration, making it difficult to meet the dual demands of flexibility and safety in emergency scenarios. Summary of the Invention
[0005] This invention provides a multimodal emergency human-computer interaction intelligent decision-making method and system, which can solve the problems in the prior art.
[0006] A first aspect of this invention provides a multimodal emergency human-computer interaction intelligent decision-making method, comprising:
[0007] Semantic association analysis is performed on the modal features in the multimodal feature sequence of human-computer interaction in emergency scenarios to extract complementary and redundant information between modalities and generate a fused semantic representation vector;
[0008] The temporal dependency analysis of the human-computer interaction multimodal feature sequence is performed by constructing a multi-level cognitive state graph. The cognitive occupancy of each node is calculated by graph convolution propagation mechanism and aggregated to form a global cognitive load tensor.
[0009] Based on the fused semantic representation vector and the global cognitive load tensor, a human-machine bidirectional conditional probability inference framework is constructed. This framework performs joint posterior probability estimation of the operator's decision intent and machine-executable constraints, generating a result after human-machine interaction. Figure 1 The intent distribution vector for consistency verification;
[0010] The intent distribution vector is semantically mapped to the standardized handling process to identify the handling process that best matches the current operator's intent and to extract decision nodes and resource scheduling constraints.
[0011] Based on the dynamic weight coefficients of each candidate intent hypothesis in the intent distribution vector and the saturation distribution of the global cognitive load tensor, a multi-granularity decision scheme set is dynamically constructed. The decision granularity of the multi-granularity decision scheme set is inversely related to the occupancy status of the global cognitive load tensor.
[0012] The set of multi-granularity decision schemes is matched with the decision nodes under constraints, and the decision output result is generated by combining the resource scheduling constraints with logical reasoning.
[0013] Semantic association analysis is performed on the modal features in the multimodal feature sequence of human-computer interaction in emergency scenarios to extract complementary and redundant information between modalities and generate a fused semantic representation vector, including:
[0014] A multimodal semantic dependency hypergraph structure is constructed, with each modal feature in the human-computer interaction multimodal feature sequence as a hypergraph node. The semantic collaboration strength between arbitrary modal subsets is calculated through a multimodal combination association measurement mechanism. A set of hyperedges reflecting multimodal combination semantics is established. The number of nodes connected by each hyperedge in the hyperedge set is adaptively determined by the multimodal combination association measurement mechanism based on the semantic collaboration strength.
[0015] Hypergraph decomposition is performed based on the multimodal semantic dependency hypergraph structure. The dominant semantic subspace and redundant semantic subspace are extracted by eigenvalue decomposition. The feature vector set corresponding to the dominant semantic subspace represents the complementary information component that contributes the most to the fusion semantics among the features of each modality. The feature vector set corresponding to the redundant semantic subspace represents the redundant information component that repeatedly expresses the same semantics among the features of each modality.
[0016] The modal features in the human-computer interaction multimodal feature sequence are orthogonally projected onto the dominant semantic subspace to suppress the information components in the redundant semantic subspace, thereby obtaining the decoupled semantic component representations of each modality.
[0017] The semantic components of each modality are represented by cross-modal interaction modeling, and the cross-modal interaction modeling results are reduced in dimensionality through tensor decomposition to generate the fused semantic representation vector.
[0018] Temporal dependency analysis was performed on the multimodal feature sequences of human-computer interaction by constructing a multi-level cognitive state graph. The cognitive occupancy of each node was calculated through a graph convolutional propagation mechanism and aggregated to form a global cognitive load tensor, including:
[0019] The human-computer interaction multimodal feature sequence is hierarchically structured and modeled according to the time dimension to construct a multi-level cognitive state map containing a single-moment cognitive state layer, a continuous time period cognitive state layer, and a cross-time period cognitive state layer, and directed edge connections reflecting temporal dependencies are established within and between each level.
[0020] Based on the topological structure of the multi-level cognitive state graph, the feature representations of nodes at each level are iteratively updated through a multi-level graph convolutional propagation mechanism. During the graph convolutional propagation process, the feature contributions of adjacent nodes are weighted and aggregated according to the temporal dependency strength of directed edges. The node features of the continuous time period cognitive state layer and the cross time period cognitive state layer are propagated and fused to the single time period cognitive state layer through a cross-level feature transfer mechanism.
[0021] Based on the updated node feature representations of each level according to the multi-level graph convolutional propagation mechanism, the cognitive occupancy value corresponding to each node in the multi-level cognitive state graph is calculated. The cognitive occupancy value reflects the degree to which the cognitive state features represented by the corresponding node occupy the cognitive resources of the current operator.
[0022] The cognitive occupancy values of all nodes in the multi-level cognitive state graph are tensorized and organized according to the hierarchical structure and temporal relationship, and aggregated to form the global cognitive load tensor.
[0023] Based on the fused semantic representation vector and the global cognitive load tensor, a human-machine bidirectional conditional probability inference framework is constructed. This framework performs joint posterior probability estimation of the operator's decision intent and machine-executable constraints, generating a result after human-machine interaction. Figure 1 The intent distribution vector for consistency verification includes:
[0024] The fused semantic representation vector is used as an observation variable to input the prior probability distribution of the operator's decision intention, and the global cognitive load tensor is used as a conditional variable to input the conditional probability distribution of the machine-executable constraints. A human-machine bidirectional conditional probability inference framework is constructed. The prior probability distribution of the operator's decision intention represents the probability distribution of the operator's intention space, and the conditional probability distribution of the machine-executable constraints represents the probability distribution of the machine-executable action space under a given cognitive load state.
[0025] Based on the aforementioned human-machine bidirectional conditional probability inference framework, a joint posterior probability estimation is performed on the prior probability distribution of the operator's decision intention and the conditional probability distribution of the machine's executable constraints through a variational Bayesian inference mechanism. This process incorporates human-machine intention into the variational inference process. Figure 1 Consistency constraints, the human-machine interface Figure 1 Consistency constraints require minimizing the difference in probability distribution between the operator's decision intention and the machine-executable constraints;
[0026] Based on the joint posterior probability estimation results output by the variational Bayesian inference mechanism, the posterior probability distribution of the operator's decision intention in the intention space is extracted, and it is verified whether the posterior probability distribution satisfies the human-machine intention... Figure 1 Consistency constraints will be implemented through human-machine interaction. Figure 1 The posterior probability distribution of the consistency check is transformed into the intent distribution vector.
[0027] The intent distribution vector is semantically mapped to the standardized handling process to identify the handling process that best matches the current operator's intent and to extract decision nodes and resource scheduling constraints, including:
[0028] A two-layer semantic alignment network for intent-process is constructed, comprising an intent semantic projection layer and a process semantic projection layer. The intent semantic projection layer converts the intent distribution vector into an intent semantic embedding representation, and the process semantic projection layer converts the standardized processing flow into a process semantic embedding representation. The two-layer semantic alignment network maps the intent semantic embedding representation and the process semantic embedding representation to a unified semantic alignment space through a cross-layer semantic alignment mechanism.
[0029] The intent-process hierarchical conditional probability inference framework is constructed by taking the intent probability distribution in the intent distribution vector as the top-level conditional probability source and establishing a multi-level conditional probability propagation link from the top-level intent probability distribution to the bottom-level processing process.
[0030] In the multi-level conditional probability propagation link, inter-process semantic dependency constraints are introduced to represent the dependency and mutual exclusion relationships between different handling processes. The conditional probability is propagated layer by layer from the top-level intention probability distribution through the variational conditional probability inference mechanism and the inter-process semantic dependency constraints are fused. The handling process with the highest posterior activation probability of the bottom-level handling process is taken as the handling process that best matches the current operator's intention.
[0031] The process structure of the handling procedure that best matches the current operator's intention is analyzed, and the decision node and the resource scheduling constraints are extracted.
[0032] Based on the dynamic weight coefficients of each candidate intent hypothesis in the intent distribution vector and the saturation distribution of the global cognitive load tensor, a multi-granularity decision scheme set is dynamically constructed. The decision granularity of the multi-granularity decision scheme set is inversely adjusted to the occupancy status of the global cognitive load tensor, including:
[0033] The dynamic weight coefficients of each candidate intention hypothesis are extracted from the intention distribution vector, and the saturation distribution is calculated from the global cognitive load tensor.
[0034] An inverse adjustment mapping relationship is established between the granularity of the decision-making scheme and the cognitive load saturation. The granularity level of the decision-making scheme is adaptively determined according to the load occupancy level in the saturation distribution through a granularity feedback adjustment mechanism. The granularity level includes the abstraction level of the decision-making scheme and the expansion depth of the operation steps. When the saturation distribution indicates that the cognitive load occupancy level is increasing, the granularity level is reduced to reduce the expansion depth of the operation steps. When the cognitive load occupancy level is decreasing, the granularity level is increased to increase the expansion depth of the operation steps.
[0035] Based on the dynamic weight coefficients, the importance of each candidate intent hypothesis is ranked. For the importance ranking, corresponding decision schemes are generated according to the granularity level. The generation process of the decision scheme is controlled according to the granularity level through a hierarchical granularity expansion mechanism. The depth of the operation steps of the decision scheme is inversely proportional to the cognitive load occupancy in the saturation distribution. The decision schemes generated for different candidate intent hypotheses are summarized to form the multi-granularity decision scheme set.
[0036] The set of multi-granularity decision schemes is matched with the decision nodes under constraints, and the decision output is generated by combining the resource scheduling constraints with logical reasoning, including:
[0037] Analyze the execution conditions and prerequisite dependencies of each decision scheme in the multi-granularity decision scheme set;
[0038] The constraint propagation mechanism propagates the node type and logical constraint relationship between the decision nodes to each decision scheme in the multi-granularity decision scheme set, establishing a constraint matching relationship between the execution conditions of the decision scheme and the node type of the decision node. The constraint propagation mechanism selects decision schemes that satisfy the logical constraint relationship between the decision nodes based on the execution conditions and the prerequisite dependencies to form a constraint-matched decision scheme set.
[0039] A resource constraint-driven logical reasoning link is constructed. The logical reasoning link receives the constraint matching decision scheme set and the resource scheduling constraints as input. The logical reasoning mechanism verifies the feasibility of each decision scheme in the constraint matching decision scheme set according to the resource scheduling constraints. The logical reasoning mechanism checks the matching degree between the resources required by each decision scheme and the currently available resources based on the resource availability conditions and resource conflict rules in the resource scheduling constraints. The decision schemes that pass the feasibility verification are selected, and the execution order of the selected decision schemes is sorted according to the logical constraint relationship of the decision nodes to obtain the decision output result.
[0040] A second aspect of this invention provides a multimodal emergency human-computer interaction intelligent decision-making system, comprising:
[0041] The semantic fusion unit is used to perform semantic association analysis on the features of each modality in the multimodal feature sequence of human-computer interaction in emergency scenarios, extract complementary and redundant information between modalities, and generate a fused semantic representation vector.
[0042] The cognitive load unit is used to perform temporal dependency analysis on the human-computer interaction multimodal feature sequence by constructing a multi-level cognitive state graph, and to calculate the cognitive occupancy of each node through a graph convolution propagation mechanism and aggregate them to form a global cognitive load tensor.
[0043] The intent estimation unit is used to construct a human-machine bidirectional conditional probability inference framework based on the fused semantic representation vector and the global cognitive load tensor, perform joint posterior probability estimation of the operator's decision intent and machine executable constraints, and generate a human-machine bidirectional conditional probability inference framework. Figure 1 The intent distribution vector for consistency verification;
[0044] The process matching unit is used to perform semantic space mapping between the intent distribution vector and the standardized handling process, identify the handling process that best matches the current operator's intent, and extract decision nodes and resource scheduling constraints.
[0045] The decision scheme unit is used to dynamically construct a multi-granularity decision scheme set based on the dynamic weight coefficients of each candidate intention hypothesis in the intention distribution vector and the saturation distribution of the global cognitive load tensor. The decision granularity of the multi-granularity decision scheme set is inversely adjusted to the occupancy status of the global cognitive load tensor.
[0046] The decision output unit is used to perform constraint matching between the set of multi-granularity decision schemes and the decision node, and generate a decision output result by combining the resource scheduling constraints through logical reasoning.
[0047] A third aspect of the present invention provides an electronic device, comprising:
[0048] processor;
[0049] Memory used to store processor-executable instructions;
[0050] The processor is configured to invoke instructions stored in the memory to execute the aforementioned method.
[0051] A fourth aspect of the present invention provides a computer-readable storage medium having stored thereon computer program instructions that, when executed by a processor, implement the aforementioned method.
[0052] The beneficial effects of this application are as follows:
[0053] Multimodal feature semantic association analysis extracts complementary and redundant information between modalities to generate fused semantic representation vectors, comprehensively capturing multidimensional information features in emergency scenarios. This significantly enhances the accuracy and semantic integrity of perception of complex situations, laying a solid data foundation for human-computer interaction decision-making and effectively reducing decision bias caused by missing information or redundant interference.
[0054] The multi-level cognitive state graph accurately calculates the occupancy of each cognitive node through temporal dependency analysis and graph convolutional propagation, and aggregates them to form a global cognitive load tensor. This enables real-time quantitative perception of the operator's cognitive state, allowing the system to dynamically assess the operator's psychological load, help predict cognitive change trends, and provide a reliable basis for adaptively adjusting interaction strategies. This allows for timely intervention in cases of cognitive overload, preventing a decline in decision-making quality.
[0055] A human-machine bidirectional conditional probability inference framework, built upon a fusion of semantic representation vectors and a global cognitive load tensor, performs joint posterior probability estimation of operator decision-making intentions and machine-executable constraints. This generates a consistency-verified intention distribution vector, effectively eliminating human-machine intention conflicts and information asymmetry. Through semantic space mapping and standardized processing procedures, this intention distribution vector is precisely matched, automatically identifying the most relevant processes and extracting decision nodes and resource scheduling constraints. This ensures that decisions are both compliant with regulations and relevant to actual on-site conditions, significantly improving the accuracy and compliance of human-machine collaborative decision-making.
[0056] Based on the dynamic weight coefficients of each candidate intent hypothesis in the intent distribution vector and the saturation distribution of the global cognitive load tensor, a multi-granularity decision scheme set is dynamically constructed. The decision granularity and cognitive load have an inverse adjustment relationship. When the operator's cognition is overloaded, it automatically simplifies to a coarse-grained scheme to reduce the decision burden. When cognition is sufficient, it provides a fine-grained scheme to optimize decision accuracy. The decision output generated by constraint matching and logical reasoning has high adaptability and executability, significantly improving the efficiency and reliability of human-machine collaborative decision-making in emergency scenarios and fully adapting to the decision-making needs under different cognitive states. Attached Figure Description
[0057] Figure 1 A flowchart illustrating a multimodal emergency human-computer interaction intelligent decision-making method;
[0058] Figure 2 This is a schematic diagram of the human-machine two-way conditional probability inference framework. Detailed Implementation
[0059] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0060] The technical solution of the present invention will be described in detail below with reference to specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments.
[0061] Figure 1 This is a flowchart illustrating the multimodal emergency human-computer interaction intelligent decision-making method according to an embodiment of the present invention.
[0062] Multimodal emergency human-computer interaction intelligent decision-making methods include:
[0063] Semantic association analysis is performed on the modal features in the multimodal feature sequence of human-computer interaction in emergency scenarios to extract complementary and redundant information between modalities and generate a fused semantic representation vector;
[0064] The temporal dependency analysis of the human-computer interaction multimodal feature sequence is performed by constructing a multi-level cognitive state graph. The cognitive occupancy of each node is calculated by graph convolution propagation mechanism and aggregated to form a global cognitive load tensor.
[0065] Based on the fused semantic representation vector and the global cognitive load tensor, a human-machine bidirectional conditional probability inference framework is constructed. This framework performs joint posterior probability estimation of the operator's decision intent and machine-executable constraints, generating a result after human-machine interaction. Figure 1 The intent distribution vector for consistency verification;
[0066] The intent distribution vector is semantically mapped to the standardized handling process to identify the handling process that best matches the current operator's intent and to extract decision nodes and resource scheduling constraints.
[0067] Based on the dynamic weight coefficients of each candidate intent hypothesis in the intent distribution vector and the saturation distribution of the global cognitive load tensor, a multi-granularity decision scheme set is dynamically constructed. The decision granularity of the multi-granularity decision scheme set is inversely related to the occupancy status of the global cognitive load tensor.
[0068] The set of multi-granularity decision schemes is matched with the decision nodes under constraints, and the decision output result is generated by combining the resource scheduling constraints with logical reasoning.
[0069] In one optional implementation, semantic association analysis is performed on the modal features in the multimodal feature sequence of human-computer interaction in emergency scenarios to extract complementary and redundant information between modalities and generate a fused semantic representation vector, including:
[0070] A multimodal semantic dependency hypergraph structure is constructed, with each modal feature in the human-computer interaction multimodal feature sequence as a hypergraph node. The semantic collaboration strength between arbitrary modal subsets is calculated through a multimodal combination association measurement mechanism. A set of hyperedges reflecting multimodal combination semantics is established. The number of nodes connected by each hyperedge in the hyperedge set is adaptively determined by the multimodal combination association measurement mechanism based on the semantic collaboration strength.
[0071] Hypergraph decomposition is performed based on the multimodal semantic dependency hypergraph structure. The dominant semantic subspace and redundant semantic subspace are extracted by eigenvalue decomposition. The feature vector set corresponding to the dominant semantic subspace represents the complementary information component that contributes the most to the fusion semantics among the features of each modality. The feature vector set corresponding to the redundant semantic subspace represents the redundant information component that repeatedly expresses the same semantics among the features of each modality.
[0072] The modal features in the human-computer interaction multimodal feature sequence are orthogonally projected onto the dominant semantic subspace to suppress the information components in the redundant semantic subspace, thereby obtaining the decoupled semantic component representations of each modality.
[0073] The semantic components of each modality are represented by cross-modal interaction modeling, and the cross-modal interaction modeling results are reduced in dimensionality through tensor decomposition to generate the fused semantic representation vector.
[0074] In emergency scenarios, human-computer interaction simultaneously generates feature data from multiple modalities, including voice commands, gestures, facial expressions, eye movements, and physiological signals. These modal features have complex semantic relationships, containing both complementary information and redundant information that repeatedly expresses the same meaning. To effectively extract complementary information between modalities and suppress redundancy, it is necessary to construct a hypergraph structure capable of characterizing high-order combinations of multimodal features.
[0075] When constructing a multimodal semantic dependency hypergraph, each modal feature extracted from the multimodal feature sequence of human-computer interaction is represented as a set of nodes in the hypergraph. ,in This represents the total number of modalities. Unlike ordinary graph structures where each edge connects only two nodes, hyperedges in a hypergraph can connect any number of nodes, thus enabling the modeling of combinatorial semantic relationships between multiple modalities. For any subset of modalities... The semantic co-strength of each modality feature within the subset is calculated using a multimodal combination correlation measurement mechanism. Semantic co-strength This reflects the synergistic gain generated when multiple modalities jointly express a certain semantic meaning. A higher value indicates stronger semantic complementarity and more prominent combined expressive power within that modal subset. When the preset collaboration strength threshold is exceeded, the hyperedge corresponding to the modality subset is included in the hyperedge set. The number of nodes connected by the hyperedge is determined by The size of the hyperedge is adaptively determined—modal subsets with high cooperation strength form higher-order hyperedges, while subsets with lower cooperation strength but still meeting the threshold requirements form lower-order hyperedges. This adaptive hyperedge generation mechanism avoids the oversimplification of multimodal relationships by fixed-order hypergraph structures, while preventing computational redundancy caused by introducing a large number of low-correlation hyperedges.
[0076] Based on the constructed multimodal semantic dependency hypergraph structure, hypergraph decomposition is performed to extract semantic subspaces. The hypergraph's association matrix... This describes the membership relationships between each node and each hyperedge, where The total number of hyperedges. This is determined by the Laplacian matrix of the hypergraph. Spectral domain analysis of the hypergraph structure. It is constructed by combining the node degree matrix, the hyperedge degree matrix, and the hyperedge weight matrix. Eigenvalue decomposition yields a set of feature vectors arranged in ascending order of eigenvalues. Feature vectors with smaller eigenvalues correspond to low-frequency semantic components in the hypergraph that exhibit gradual changes. These components, consistent across multiple modal nodes, represent complementary information that contributes most to the fused semantics among the features of each modality, constituting the dominant semantic subspace. Eigenvectors with larger eigenvalues correspond to high-frequency components in the hypergraph that exhibit dramatic changes. These components often reflect redundant information formed by the repeated expression of the same semantics among different modalities, constituting a redundant semantic subspace. The dimensional division between the dominant semantic subspace and the redundant semantic subspace is based on a preset threshold for the cumulative contribution rate of feature values. Confirmed, current The sum of these eigenvalues accounts for a proportion of the sum of all eigenvalues. At that time, before The subspace spanned by the 1 feature vectors is the dominant semantic subspace, while the remaining feature vectors span the redundant semantic subspace.
[0077] The feature vectors of each modality in the multimodal feature sequence of human-computer interaction are orthogonally projected onto the dominant semantic subspace. Let the first... The original feature vectors of each modality are The orthogonal basis matrix of the dominant semantic subspace is Then the first The decoupling semantic components of each modality are characterized as follows ,in This is the low-dimensional representation after projection. Mathematically, orthogonal projection is equivalent to setting the components of the original feature vector in the redundant semantic subspace direction to zero, retaining only the projected components in the dominant semantic subspace direction, thus effectively suppressing redundant information. After orthogonal projection transformation, the semantic components of each modality are represented... Already unified Alignment in the semantic subspace provides a semantically consistent feature foundation for subsequent cross-modal interaction modeling.
[0078] In the cross-modal interaction modeling stage, the semantic components after decoupling from each modality are represented. We perform pairwise and higher-order interaction computations to capture co-activation patterns between different modalities within the dominant semantic subspace. Cross-modal interaction can be achieved through a bilinear attention mechanism, for any two modalities... and Its interactive features are ,in This is a learnable cross-modal interaction weight matrix. The interaction features of all modal pairs are organized into a third-order tensor. This tensor fully records the interaction information of each modality pair in the dominant semantic subspace. However, when the number of modalities... When the dimensions of a third-order tensor are large, the dimensionality increases significantly, leading to excessive computational overhead when used directly. Therefore, a tensor decomposition mechanism is used to reduce the dimensionality of the cross-modal interaction modeling results. Tucker decomposition or CP decomposition is employed to decompose the tensor. A low-rank approximation is performed, decomposing the core tensor into a product of several low-rank factor matrices and the core tensor. This significantly compresses the representation dimension while preserving key cross-modal interaction information. After decomposition, the vectorized representation of the core tensor is extracted and concatenated with the weighted convergence result of the decoupled semantic component representations from each modality. This concatenation is then linearly transformed and mapped to a fixed-dimensional semantic space, ultimately generating a fused semantic representation vector. This vector contains complementary information components of each modality in the dominant semantic subspace, captures the collaborative patterns of multimodal joint semantics through cross-modal interaction modeling, and effectively removes redundant information expressed repeatedly between modalities through the suppression operation of redundant semantic subspace, providing high-quality multimodal fusion feature input for subsequent cognitive state graph analysis and decision intention inference.
[0079] In real-world emergency scenarios, multimodal feature sequences collected at different times may exhibit partial modality loss. For example, in noisy environments, the low signal-to-noise ratio of speech modalities can lead to feature extraction failure. In such cases, the hypergraph structure automatically degenerates into a sub-hypergraph containing only valid modality nodes, and the hyperedge set is dynamically adjusted accordingly. The extraction of the dominant semantic subspace is still performed on the sub-hypergraph composed of valid modalities, ensuring the robustness of the fused semantic representation vector under modality loss conditions. The calculations of orthogonal projection transformation and tensor decomposition can be completed independently on the valid modality subset without the need to complete or fill in missing modalities, thereby reducing the impact of modality loss on the final fusion result and improving the applicability of the overall method in complex emergency environments.
[0080] In one optional implementation, temporal dependency analysis is performed on the human-computer interaction multimodal feature sequence by constructing a multi-level cognitive state graph, and the cognitive occupancy of each node is calculated and aggregated to form a global cognitive load tensor through a graph convolution propagation mechanism, including:
[0081] The human-computer interaction multimodal feature sequence is hierarchically structured and modeled according to the time dimension to construct a multi-level cognitive state map containing a single-moment cognitive state layer, a continuous time period cognitive state layer, and a cross-time period cognitive state layer, and directed edge connections reflecting temporal dependencies are established within and between each level.
[0082] Based on the topological structure of the multi-level cognitive state graph, the feature representations of nodes at each level are iteratively updated through a multi-level graph convolutional propagation mechanism. During the graph convolutional propagation process, the feature contributions of adjacent nodes are weighted and aggregated according to the temporal dependency strength of directed edges. The node features of the continuous time period cognitive state layer and the cross time period cognitive state layer are propagated and fused to the single time period cognitive state layer through a cross-level feature transfer mechanism.
[0083] Based on the updated node feature representations of each level according to the multi-level graph convolutional propagation mechanism, the cognitive occupancy value corresponding to each node in the multi-level cognitive state graph is calculated. The cognitive occupancy value reflects the degree to which the cognitive state features represented by the corresponding node occupy the cognitive resources of the current operator.
[0084] The cognitive occupancy values of all nodes in the multi-level cognitive state graph are tensorized and organized according to the hierarchical structure and temporal relationship, and aggregated to form the global cognitive load tensor.
[0085] The multimodal feature sequences of human-computer interaction are hierarchically structured and modeled according to the time dimension. The feature data on the original timeline are divided into granular segments, with the sampling frame as the smallest unit. The operator's state features at a single moment (including eye-tracking fixation, gestures, and speech fragments) are organized into nodes of the single-moment cognitive state layer; aggregated state features within a time window consisting of several consecutive frames are organized into nodes of the continuous-time-span cognitive state layer; and features spanning multiple time windows and reflecting cognitive evolution trends over a longer time span are organized into nodes of the cross-time-span cognitive state layer. These three levels together constitute a multi-level cognitive state graph, denoted as _____. The set of nodes in the single-moment cognitive state layer is as follows: The set of nodes in the cognitive state layer over a continuous time period is The set of nodes in the cross-time cognitive state layer is All three conditions are met. Within each level, directed edges reflecting temporal dependencies are established between adjacent moments or adjacent time windows. The direction of the directed edges follows the chronological order, i.e., from earlier moment nodes to later moment nodes. Between levels, directed edges are established from top to bottom between a cognitive state layer node in a continuous time period and the cognitive state layer nodes it contains in a single moment, and similarly, directed edges are established from top to bottom between a cognitive state layer node in a cross-time period and the cognitive state layer nodes in a continuous time period it encompasses, thus forming a temporal dependency propagation path between levels.
[0086] After constructing the topological structure of the multi-level cognitive state graph, the feature representations of each level node are iteratively updated using a multi-level graph convolutional propagation mechanism. Let the node... In the The features in the subgraph convolution iteration are represented as follows: Its initial value It is obtained by linear mapping of the original multimodal features within the corresponding time or time period. In each graph convolution iteration, the nodes... The feature update rule is to update all its incoming neighbor nodes. After weighted aggregation of features, a new feature representation is obtained through nonlinear activation transformation. (Directed edge) The time-dependent strength weight is denoted as The weight is determined by the time interval between the two nodes connected by the edge and the correlation between their corresponding cognitive states; the shorter the interval and the higher the correlation, the greater the weight. In the The features after the nth iteration are represented as follows: ,in For graph convolution, the weight matrix can be learned. For bias vectors, It is a non-linear activation function. For nodes The set of incoming neighbor nodes. The total number of iterations is denoted as . After completing all iterations, the final feature representation of each node is obtained. .
[0087] The cross-level feature transfer mechanism is executed synchronously during graph convolution propagation. Nodes in the cognitive state layer across consecutive time periods. In each iteration, the features not only receive aggregated information from neighboring nodes at the same level, but also receive top-down propagation information from the corresponding parent node in the cross-time cognitive state layer; single-time cognitive state layer nodes The features simultaneously receive propagation and fusion information from the corresponding parent nodes of the cognitive state layers in consecutive time periods. The top-down cross-level propagation weights are denoted as... This mechanism controls the proportion of influence of high-level node features on the updates of low-level node features. It ensures that long-term cognitive evolution trends effectively affect the feature representations of short-term cognitive states, enabling nodes at a single-time cognitive state layer to carry both local temporal information and global temporal context information after iterative updates.
[0088] After obtaining the updated node feature representations through the multi-level graph convolutional propagation mechanism, the cognitive occupancy value for each node is calculated. Cognitive occupancy reflects the degree to which the cognitive state characteristics represented by that node consume the limited cognitive resources of the current operator; it is a core quantitative indicator for measuring the operator's cognitive load level at a given time or time period. For nodes... Its cognitive occupancy By using the final node features The cognitive occupancy evaluation function is obtained by inputting the input, which is implemented by a two-layer fully connected network. The output is a scalar value with a normalized range. The interval, in which This indicates that the state corresponding to this node consumes almost no cognitive resources. This indicates that cognitive resources are in a state of complete saturation. In emergency scenarios, when operators process multiple information sources simultaneously, the cognitive occupancy values of multiple nodes may approach saturation at the same time. At this point, the global cognitive load tensor will exhibit a high occupancy distribution characteristic, providing a basis for the dynamic adjustment of subsequent decision granularity.
[0089] The cognitive occupancy values of all nodes in the multi-level cognitive state graph are tensorized and organized according to the hierarchical structure and temporal relationship, and aggregated to form a global cognitive load tensor. . Specifically, This is a third-order tensor, with three dimensions corresponding to the hierarchy, time, and feature dimensions, respectively. The size of the hierarchy dimension is... This corresponds to a single-moment cognitive state layer, a continuous-time cognitive state layer, and a cross-time cognitive state layer; the size of the time dimension is the maximum number of nodes in each level, with insufficient positions padded with zeros; the size of the feature dimension is... This is a concatenated representation of the final feature vector of each node and a scalar representation of cognitive load. During tensor organization, nodes within the same level are arranged in temporal order to ensure that their relative positions in the time dimension are consistent with the original timeline. Cross-level correspondences are maintained through index alignment in the level dimension, allowing cognitive state information from different levels within the same time span to form corresponding relationships in the tensor structure. The final global cognitive load tensor... The system fully encodes the multi-granularity cognitive load distribution of operators in the current emergency scenario, from a single moment to multiple time periods, providing a structured cognitive state input for the subsequent construction of a human-machine bidirectional conditional probability inference framework based on the fusion of semantic representation vectors and global cognitive load tensors.
[0090] In real-world emergency scenarios, operators' cognitive states often exhibit non-stationary characteristics as the complexity of the task dynamically changes. For example, in the initial stages of an emergency, the influx of a large amount of information causes a rapid increase in the cognitive occupancy of nodes in the cognitive state layer across different time periods. However, as the handling process stabilizes, the occupancy of nodes in the cognitive state layer at a single moment gradually decreases. The hierarchical structure of a multi-level cognitive state graph can effectively capture this dynamic change in cognition across time scales. The graph convolutional propagation mechanism utilizes temporal-dependent strength weights. Adaptively adjust the information transmission ratio between nodes at different time distances, thereby increasing the global cognitive load tensor. It can accurately reflect the cognitive resource allocation status of operators at different time granularities, providing a high-precision cognitive load perception basis for intelligent decision-making systems.
[0091] In one optional implementation, a human-machine bidirectional conditional probability inference framework is constructed based on the fused semantic representation vector and the global cognitive load tensor. This framework performs joint posterior probability estimation of the operator's decision intent and machine-executable constraints, generating a result after human-machine interaction. Figure 1 The intent distribution vector for consistency verification includes:
[0092] The fused semantic representation vector is used as an observation variable to input the prior probability distribution of the operator's decision intention, and the global cognitive load tensor is used as a conditional variable to input the conditional probability distribution of the machine executable constraints. A human-machine bidirectional conditional probability inference framework is constructed. The prior probability distribution of the operator's decision intention represents the probability distribution of the operator's intention space, and the conditional probability distribution of the machine executable constraints represents the probability distribution of the machine's executable action space under a given cognitive load state.
[0093] Based on the aforementioned human-machine bidirectional conditional probability inference framework, a joint posterior probability estimation is performed on the prior probability distribution of the operator's decision intention and the conditional probability distribution of the machine's executable constraints through a variational Bayesian inference mechanism. This introduces human-machine intention into the variational inference process. Figure 1 Consistency constraints, the human-machine interface Figure 1 Consistency constraints require minimizing the difference in probability distribution between the operator's decision intention and the machine-executable constraints;
[0094] Based on the joint posterior probability estimation results output by the variational Bayesian inference mechanism, the posterior probability distribution of the operator's decision intention in the intention space is extracted, and it is verified whether the posterior probability distribution satisfies the human-machine intention... Figure 1 Consistency constraints will be implemented through human-machine interaction. Figure 1 The posterior probability distribution of the consistency check is transformed into the intent distribution vector.
[0095] like Figure 2 As shown, the method includes:
[0096] Fusion semantic representation vector Carrying comprehensive semantic information from multimodal perception channels after decoupling through a hypergraph structure and cross-modal interaction, this information is introduced as an observed variable into the prior probability distribution modeling process of the operator's decision-making intent. Specifically, with As input, the prior probability distribution of the operator's decision intention is obtained by mapping it to the latent variable space of the intention through a parameterized probabilistic coding network. ,in As a latent variable representing intent, this distribution characterizes the probability density of various decision-making intentions held by operators under the current multimodal observation conditions in the intent space. Simultaneously, the global cognitive load tensor... As a condition variable, the conditional probability distribution of machine-executable constraints Modeling is performed, among which This represents the executable actions of the machine under a given cognitive load. This conditional probability distribution reflects the probability allocation of the legally executable action space under different cognitive load saturation levels—when... When the cognitive load is high, the constraints on machine-executable actions tighten, and the conditional probability of high-risk actions is suppressed; when the cognitive load is low, the space of executable actions expands accordingly. Combining the probability distributions in these two directions forms a human-machine two-way conditional probability inference framework. This framework simultaneously captures the uncertainty of human intention and the uncertainty of machine execution constraints, providing a complete probabilistic basis for subsequent joint posterior estimation.
[0097] After the human-machine bidirectional conditional probability inference framework is constructed, a variational Bayesian inference mechanism is used to jointly estimate the prior probability distribution of the operator's decision intention and the conditional probability distribution of the machine's executable constraints. The core idea of variational Bayesian inference is to introduce a processable variational distribution. To approximate the true joint posterior distribution By minimizing the KL divergence between the two To optimize the variational distribution parameters. Variational distribution Decomposition is performed using the mean-field assumption, i.e. ,in and These are the parameter sets for the strain distribution. The objective is optimized using the evidence lower bound (ELBO) of variational inference. and Iterative updates are performed to gradually approximate the true posterior distribution.
[0098] In the process of variational reasoning, human-machine interaction is introduced. Figure 1 Consistency constraints are applied to ensure that the probability distribution difference between human intent and machine constraints is minimized. This constraint is added to the optimization objective of variational inference as a regularization term, specifically by adding an additional term for human-machine intent to the ELBO loss. Figure 1 Sexual punishment items ,in The consistency constraint strength coefficient, The Jensen-Shannon divergence measures the symmetry difference between the distribution of operator decision intent and the distribution of machine-executable constraints. Compared to KL divergence, Jensen-Shannon divergence is symmetric and bounded, providing a more stable measure of the difference between the two distributions and avoiding numerical divergence when the support sets of the distributions do not overlap. By iterating through gradient descent on the above joint optimization objective, the desired result is obtained, satisfying the human-machine intention... Figure 1 Variational posterior distribution parameters of consistency constraints and .
[0099] Optimal variational posterior parameters based on variational Bayesian inference output Extract the posterior probability distribution of the operator's decision intention in the intention space. Human-machine interaction was used to analyze the posterior probability distribution. Figure 1 Consistency verification, the verification method is calculation Is it below the preset consistency threshold? If the divergence value satisfies Then it is considered that the current posterior probability distribution has passed the human-machine intention. Figure 1 Consistency verification indicates that the operator's decision-making intent and the machine-executable constraints are sufficiently consistent at the probability distribution level, and there is no fundamental execution conflict; if this condition is not met, then the consistency constraint strength coefficient in the variational inference process is adjusted. Adaptively increase the size and re-execute variational inference iterations until the verification passes.
[0100] Through human-machine intention Figure 1 Posterior probability distribution of consistency verification It is then transformed into a discretized intention distribution vector. ,in This represents the total number of predefined intent categories. The transformation process is achieved by performing an integral approximation or Monte Carlo sampling estimation on the posterior probabilities of each discrete intent category in the latent variable space of the intent, obtaining the posterior probability value corresponding to each candidate intent hypothesis, and then normalizing the probability values of all candidate intents to ensure... It satisfies the probabilistic simplex constraint, meaning that all components are non-negative and their sum is 1. The Middle Each component Indicates the first The posterior probability of candidate intentions under the current multimodal observation and cognitive load conditions is a vector that fully characterizes the probability distribution of the operator's decision intentions in the entire intention space, providing a probabilistic basis at the intention level for the dynamic construction of subsequent multi-granularity decision schemes.
[0101] Intent distribution vector The quality of human-machine interaction directly impacts the reliability of subsequent decision-making outputs. In real-world emergency scenarios, operator behavior signals are often subject to ambiguity due to noise interference or cognitive overload. Variational Bayesian inference frameworks, by explicitly modeling uncertainty, can retain necessary probability diffusion information in the intent distribution vector, avoiding overly deterministic hard decisions regarding ambiguous intents. Simultaneously, human-machine interaction... Figure 1 The introduction of consistency constraints ensures that the operational intent represented by the intent distribution vector is feasible at the machine execution level, fundamentally avoiding the output of intents that the machine cannot execute to the subsequent decision-making process with a high probability, thereby improving the safety and reliability of the entire emergency human-machine interaction decision-making chain.
[0102] In one optional implementation, the intent distribution vector is semantically mapped to a standardized handling process to identify the handling process that best matches the current operator's intent and to extract decision nodes and resource scheduling constraints, including:
[0103] A two-layer semantic alignment network for intent-process is constructed, comprising an intent semantic projection layer and a process semantic projection layer. The intent semantic projection layer converts the intent distribution vector into an intent semantic embedding representation, and the process semantic projection layer converts the standardized processing flow into a process semantic embedding representation. The two-layer semantic alignment network maps the intent semantic embedding representation and the process semantic embedding representation to a unified semantic alignment space through a cross-layer semantic alignment mechanism.
[0104] The intent-process hierarchical conditional probability inference framework is constructed by taking the intent probability distribution in the intent distribution vector as the top-level conditional probability source and establishing a multi-level conditional probability propagation link from the top-level intent probability distribution to the bottom-level processing process.
[0105] In the multi-level conditional probability propagation link, inter-process semantic dependency constraints are introduced to represent the dependency and mutual exclusion relationships between different handling processes. The conditional probability is propagated layer by layer from the top-level intention probability distribution through the variational conditional probability inference mechanism and the inter-process semantic dependency constraints are fused. The handling process with the highest posterior activation probability of the bottom-level handling process is taken as the handling process that best matches the current operator's intention.
[0106] The process structure of the handling procedure that best matches the current operator's intention is analyzed, and the decision node and the resource scheduling constraints are extracted.
[0107] A semantic space mapping is performed between the intent distribution vector and the standardized processing flow to construct an intent-flow two-layer semantic alignment network. This network contains two parallel projection layers: an intent semantic projection layer and a flow semantic projection layer. The intent semantic projection layer receives the intent distribution vector. It is mapped into a high-dimensional intention semantic embedding representation through multi-layer nonlinear transformation. This embedding represents the comprehensive semantic information encoded in the semantic space of the operator's current decision-making intent, with a dimension of [missing information]. The process semantic projection layer targets a pre-stored standardized processing procedure library and performs semantic projection on each processing procedure. ( , To extract the structured text description and operation step sequence of the total number of processes, a process semantic encoder is used to generate the corresponding process semantic embedding representation. The dimensions are the same This ensures that the two types of embedding representations are in comparable semantic spaces of the same dimension.
[0108] The core of the cross-layer semantic alignment mechanism lies in unifying the mapping of intent semantic embedding representation and process semantic embedding representation to a semantic alignment space, making semantically similar intents and processing flows closer together in this space. Specifically, this is achieved through a learnable alignment projection matrix. Perform linear transformations on the two types of embedding representations to obtain the aligned intent representation. With process representation ,in and These are the alignment projection matrices for the intent side and the process side, respectively. In the semantic alignment space, the cosine semantic similarity between the intent embedding and each processing process embedding is calculated. ,Right now
[0109]
[0110] This serves as the prior input for subsequent conditional probability inference.
[0111] In the intent-process hierarchical conditional probability inference framework, the intent distribution vector is... Posterior probability values of each candidate intent category ( This serves as the top-level conditional probability source. The framework divides the entire inference process into multiple levels: the top level corresponds to the abstract intent category of the operator, the middle level corresponds to the functional category of the handling procedure, and the bottom level corresponds to specific standardized handling procedure instances. Conditional probability propagation from the top to the middle levels is achieved through the intent-functional category conditional probability matrix. Implementation, in which The matrix elements represent the total number of functional categories in the processing workflow. Indicates the first Activation of the first class under the condition of intent The conditional probability of a functional process. The propagation of conditional probability from the intermediate layer to the bottom layer is achieved through the functional category-specific process conditional probability matrix. Implementation, matrix elements Indicates the first Activate the first function under the functional process conditions The conditional probability of each specific handling procedure. Through chain-like conditional probability propagation, the underlying... The initial activation probability of each processing step under a given top-level intent distribution. It can be represented as .
[0112] Different handling processes may have dependencies (i.e., the activation of one process depends on the completion of another) or mutual exclusions (i.e., two processes cannot be activated simultaneously in the same emergency scenario). These relationships are encoded into a process dependency matrix. , of which elements A positive value indicates a process. process A dependency exists; a negative value indicates mutual exclusion, while zero indicates no constraint. In the variational conditional probability inference mechanism, a variational posterior distribution of the process activation probability is introduced. ,in Given a set of variational parameters, optimization is achieved by maximizing a variational lower bound that includes penalties for semantic dependencies between processes. The optimization objective is to optimize the initial activation probability. Based on this, a dependency constraint correction term is added, so that both dependency constraints and mutual exclusion constraints are included in the final posterior activation probability. The calculation process. After convergence of variational inference, the posterior activation probability... The highest-level handling procedure is the one identified as the one that best matches the current operator's intention. ,Right now .
[0113] Determine the most suitable treatment procedure Subsequently, the process structure is analyzed to extract decision nodes and resource scheduling constraints. The standardized handling process is stored in the form of a directed acyclic graph (DAG), where each node corresponds to a specific operation step or decision point. Node attributes include operation type label, execution condition predicate, required resource type set, and time constraint parameters. Decision node identification is based on node out-degree judgment: nodes with an out-degree greater than 1 indicate that there are multiple subsequent execution paths for this step, belonging to decision nodes requiring path selection by operators or automated decision-making systems. All nodes with an out-degree greater than 1 are extracted and constitute a set of decision nodes. Resource scheduling constraints are determined by traversing all nodes in the directed acyclic graph of the processing flow, summarizing the resource types, quantity upper and lower bounds, scheduling priorities, and resource mutual exclusion constraints required by each node, forming a structured set of resource scheduling constraints. The set stores each constraint in the form of a triple, which includes the constraint type, the constraint subject resource identifier, and the constraint quantization parameter.
[0114] Extracted set of decision nodes With resource scheduling constraint set This will serve as input for subsequent multi-granularity decision-making scheme generation and constraint matching stages, ensuring that the final generated decision output is consistent with the operational specifications of the standardized handling process, while also meeting resource availability constraints and temporal logic constraints in emergency scenarios. The entire semantic space mapping and process matching process fully utilizes the multi-category intent probability information carried by the intent distribution vector. Through hierarchical conditional probability propagation and inter-process constraint fusion, it achieves a precise mapping from abstract operational intents to specific executable handling processes.
[0115] In one optional implementation, a multi-granularity decision scheme set is dynamically constructed based on the dynamic weight coefficients of each candidate intent hypothesis in the intent distribution vector and the saturation distribution of the global cognitive load tensor. The decision granularity of the multi-granularity decision scheme set is inversely adjusted to the occupancy status of the global cognitive load tensor, including:
[0116] The dynamic weight coefficients of each candidate intention hypothesis are extracted from the intention distribution vector, and the saturation distribution is calculated from the global cognitive load tensor.
[0117] An inverse adjustment mapping relationship is established between the granularity of the decision-making scheme and the cognitive load saturation. The granularity level of the decision-making scheme is adaptively determined according to the load occupancy level in the saturation distribution through a granularity feedback adjustment mechanism. The granularity level includes the abstraction level of the decision-making scheme and the expansion depth of the operation steps. When the saturation distribution indicates that the cognitive load occupancy level is increasing, the granularity level is reduced to reduce the expansion depth of the operation steps. When the cognitive load occupancy level is decreasing, the granularity level is increased to increase the expansion depth of the operation steps.
[0118] Based on the dynamic weight coefficients, the importance of each candidate intent hypothesis is ranked. For the importance ranking, corresponding decision schemes are generated according to the granularity level. The generation process of the decision scheme is controlled according to the granularity level through a hierarchical granularity expansion mechanism. The depth of the operation steps of the decision scheme is inversely proportional to the cognitive load occupancy in the saturation distribution. The decision schemes generated for different candidate intent hypotheses are summarized to form the multi-granularity decision scheme set.
[0119] From the intention distribution vector The dynamic weight coefficients of each candidate intent hypothesis are extracted, and the posterior probability value of each candidate intent is used. ( The weights are used as the importance weights of the candidate intent hypothesis in the current emergency scenario. Since the evolution of the emergency scenario is time-varying, the posterior probability distributions of each candidate intent hypothesis will shift with the update of the operator's behavioral signals at different times. Therefore, these posterior probability values are treated as dynamic weight coefficients, rather than static fixed values. Meanwhile, from the global cognitive load tensor... The saturation distribution is calculated. Specifically, the saturation distribution is calculated. Along feature dimension The direction is normalized to obtain the load occupancy ratio of each cognitive state node at the current time, and the cognitive occupancy value of each node is denoted as . The global saturation index is obtained by taking a weighted average of the occupancy of all nodes. This is used to characterize the degree of overall cognitive resource utilization. When Exceeding the preset high load threshold When, it is considered that the operator is in a state of cognitive overload; when Below the low load threshold At that time, it was believed that operators had sufficient cognitive margin to process more complex decision-making information.
[0120] Establish an inverse adjustment mapping relationship between decision-making granularity and cognitive load saturation, and define granularity levels. ,in This corresponds to the coarsest granularity (i.e., a highly abstract decision-making scheme with very few operational steps). This corresponds to the finest granularity (i.e., decision-making schemes with fully expanded steps and rich operational details). The total number of predefined granularity levels. The granularity feedback adjustment mechanism achieves adaptive granularity determination through the following mapping logic: [The normalized global saturation index is then used for...] Mapping to granular level index The mapping relationship satisfies Follow The granularity level decreases as cognitive load increases; that is, the higher the cognitive load, the lower the granularity level selected. Specifically, this inverse adjustment can be achieved using piecewise linear mapping or a monotonically decreasing function, ensuring that when... Approaching 1 o'clock tend ,when Approaching 0 tend .
[0121] Particle size level Simultaneously control two dimensions of the decision-making scheme: the level of abstraction and the depth of operational steps. The level of abstraction determines whether the decision-making scheme is presented as a goal-level description (e.g., "start backup power") or an action-level description (e.g., "sequentially execute circuit breaker reset, bus charging, and load switching"). The depth of operational steps determines the number of sub-steps and levels under each decision node. In situations with high cognitive load ( In situations where the cognitive load is low, the granularity level is reduced so that the decision-making scheme only presents the key operational objectives and the minimum necessary steps, thereby reducing the information processing burden on operators; In this context, by increasing the granularity level and fully unfolding the hierarchical structure of the operation steps, detailed execution guidance can be provided to operators. This reverse adjustment relationship ensures that the human-computer interface maintains an appropriate information density under different cognitive states.
[0122] Determining the granularity level Next, the importance of each candidate intent hypothesis is ranked based on dynamic weighting coefficients. The posterior probability value of each candidate intent hypothesis is then used as the ranking factor. Based on the sorting criteria, arrange them in descending order to obtain the importance ranking sequence. ,in The candidate intent hypothesis with the highest posterior probability has the highest priority. After sorting, for each candidate intent hypothesis... ( According to the current granularity level Generate corresponding decision-making schemes. The hierarchical granularity expansion mechanism plays a core control role in the generation process: for granularity levels... (Coarse-grained) Only generates a top-level decision objective description, without expanding sub-steps; for intermediate-grained levels, according to... Corresponding unfolding depth parameters The decision tree is expanded recursively, and each decision node can be expanded at most once. Layer sub-nodes; for granularity level (Fine-grained) Fully expand all reachable operation steps, including resource scheduling details and operation timing constraints.
[0123] Depth of Operational Steps for Decision-Making Plan It is inversely proportional to the degree of cognitive load occupancy in the saturation distribution, and can be quantified in the following way: Values With normalized saturation The complement of the product, i.e. ,in This is the preset maximum unfolding depth. When... When it approaches 1, When the value approaches 0, the decision-making process degenerates into a pure objective-level description; when When it approaches 0, Approaching The decision-making scheme reaches its maximum depth of development. In actual engineering implementation, To obtain an integer value, the above calculation result can be rounded down to the nearest integer.
[0124] For candidate intent hypotheses that rank high in importance, their dynamic weighting coefficients should also be considered when generating decision-making options. The decision-making process differentiates the content of the proposed solutions: Decision-making solutions corresponding to candidate intent assumptions with higher weights are allocated more operational steps at the same granularity level, while decision-making solutions corresponding to candidate intent assumptions with lower weights have their non-critical branches appropriately simplified while maintaining consistency in granularity. This differentiation ensures that, given limited cognitive resources, the most probable intent assumption receives the most comprehensive decision support.
[0125] The decision schemes generated based on different candidate intent assumptions are aggregated to form a multi-granularity decision scheme set. Each element in the set corresponds to a decision scheme based on a candidate intent hypothesis, with the schemes organized by granularity level. The hierarchical structure of organizational operation steps. The overall granularity is determined by the global saturation index at the current moment. The decision is dynamically made and updated in real time as the emergency scenario evolves and the operator's cognitive state changes. When the emergency scenario enters a new phase or the cognitive load changes significantly, the calculation is recalculated. And trigger the granularity feedback adjustment mechanism to... The depth of each option is readjusted to ensure that the set of multi-granularity decision options always adapts to the operator's real-time cognitive state, thereby minimizing the operator's cognitive burden while ensuring the integrity of the decision.
[0126] In one optional implementation, the set of multi-granularity decision schemes is constrained and matched with the decision node, and the decision output result is generated by logical reasoning combined with the resource scheduling constraints, including:
[0127] Analyze the execution conditions and prerequisite dependencies of each decision scheme in the multi-granularity decision scheme set;
[0128] The constraint propagation mechanism propagates the node type and logical constraint relationship between the decision nodes to each decision scheme in the multi-granularity decision scheme set, establishing a constraint matching relationship between the execution conditions of the decision scheme and the node type of the decision node. The constraint propagation mechanism selects decision schemes that satisfy the logical constraint relationship between the decision nodes based on the execution conditions and the prerequisite dependencies to form a constraint-matched decision scheme set.
[0129] A resource constraint-driven logical reasoning link is constructed. The logical reasoning link receives the constraint matching decision scheme set and the resource scheduling constraints as input. The logical reasoning mechanism verifies the feasibility of each decision scheme in the constraint matching decision scheme set according to the resource scheduling constraints. The logical reasoning mechanism checks the matching degree between the resources required by each decision scheme and the currently available resources based on the resource availability conditions and resource conflict rules in the resource scheduling constraints. The decision schemes that pass the feasibility verification are selected, and the execution order of the selected decision schemes is sorted according to the logical constraint relationship of the decision nodes to obtain the decision output result.
[0130] After constructing the set of multi-granularity decision-making schemes, it is necessary to perform constraint matching with the decision nodes extracted from the standardized handling process, and combine resource scheduling constraints to generate the final executable decision output through logical reasoning. This process begins with the set of multi-granularity decision-making schemes. The structured analysis of each decision scheme involves extracting its execution conditions and prerequisite dependencies. Execution conditions describe the state prerequisites that must be met for the scheme to proceed, such as specific equipment being ready, specific information being confirmed, or specific operations being completed. Prerequisite dependencies characterize the sequential constraints between this scheme and other schemes; that is, the execution of one scheme depends on the completion of another. A complete analysis of execution conditions and prerequisite dependencies forms the basis for subsequent constraint propagation and logical reasoning. The analysis results are organized in the form of a structured dependency graph, where nodes represent decision schemes, directed edges represent prerequisite dependencies, and the node attributes include predicate expressions for the execution conditions.
[0131] The core function of the constraint propagation mechanism is to set up decision nodes. The node type information and logical constraint relationships between decision nodes are propagated to each decision scheme in the multi-granularity decision scheme set, thereby establishing a correspondence between execution conditions and decision node types at the decision scheme level. The node type of a decision node reflects its semantic role in the handling process, such as trigger node, confirmation node, execution node, or termination node. Different types of nodes correspond to different logical constraint rules. The constraint propagation process adopts an iterative approach: for each decision scheme, the decision node type involved in its execution conditions is retrieved, and the logical constraint rules associated with that node type are introduced into the constraint set of that scheme; then, it is propagated upstream along the directed edges of the preceding dependencies, simultaneously incorporating the logical constraints of the upstream schemes. The propagation terminates when the constraint sets of all schemes no longer change in two consecutive iterations, i.e., a fixed point is reached. After constraint propagation, the consistency of the execution conditions of each decision scheme with the logical constraint rules of the matched decision nodes is checked: if the predicate expression in the execution conditions is compatible with the constraint rules of the corresponding node type, then the scheme passes the constraint matching; otherwise, it is removed from the candidate set. Finally, the decision schemes that pass the constraint matching are retained, forming the constraint-matched decision scheme set. .
[0132] In the set of constraint-matching decision solutions Once determined, a resource-constrained, logical reasoning chain is constructed to verify the resource feasibility of each decision option in the set. The logical reasoning chain is based on... With resource scheduling constraint set As two inputs. Resource scheduling constraint set. It includes two types of rules: resource availability conditions and resource conflict rules. Resource availability conditions describe the actual availability status of various resources (including human resources, equipment resources, communication resources, etc.) at the current moment, expressed as key-value pairs of resource identifiers and available quantity or availability status flags; resource conflict rules describe the resource competition situations that may occur when multiple decision schemes are executed simultaneously, such as the same device cannot be occupied by two schemes at the same time, and the same communication channel can only be allocated to one task at the same time period.
[0133] Logical reasoning mechanism Each decision-making option in the process undergoes a feasibility verification process. For each option... Required resource set traversal For each type of resource, check whether its demand does not exceed [the specified limit]. The available amount of the corresponding resources. Assume the scheme... For resource types The demand is The current available quantity is The resource availability verification condition is: For all All conditions must be met. If any type of resource does not meet this condition, then the solution... If the feasibility verification fails, the solution is removed from the candidate set. For solutions that pass the resource availability verification, a resource conflict check is further performed: all verified solutions are paired, and the resource conflict rules are used to determine if there are any incompatible solution pairs. If a conflict exists, the higher-priority solution is retained and the lower-priority solution is removed based on its priority score. The priority score of a solution comprehensively considers its position in the intent distribution vector. The posterior probability value of the corresponding intention hypothesis The final priority score is obtained by weighting and summing the consistency score of the scheme during the constraint matching phase, with the weight determined by the preset parameters of the emergency scenario.
[0134] After feasibility verification and conflict resolution, a subset of decision-making schemes that passed the verification were obtained. .right The plan is based on decision nodes The execution order is determined by the logical constraints between decision nodes. The execution order is sorted using a topological sorting algorithm: a directed acyclic graph is constructed based on the logical constraints between decision nodes, where each node corresponds to a specific logical constraint. In the algorithm, directed edges represent the execution order constraints between decision schemes. Topological sorting starts from the node with an in-degree of zero (i.e., the scheme with no prerequisites) and proceeds layer by layer according to dependencies to generate a linear execution sequence that satisfies all logical constraints. If a cycle is detected during topological sorting, it indicates a contradiction in the constraints, triggering an exception handling process: the lowest priority scheme in the cycle is located and removed, and topological sorting is re-executed until a cycle-free execution sequence is generated.
[0135] After topological sorting, the sorted sequence of decision schemes is integrated with the corresponding resource allocation schemes and execution timing information, and encapsulated into a structured decision output. The decision output contains three levels of information: an execution scheme sequence (each decision scheme and its operational steps arranged in execution order), a resource allocation list (the specific resources used by each scheme and their allocation time periods), and a logical constraint summary (pre-dependent relationships and conflict resolution explanations between schemes). The decision output is output in a standardized format for subsequent human-computer interaction interface rendering and operator confirmation. In scenarios with high operator cognitive load, the decision output is automatically simplified to the core operational steps of high-priority schemes to reduce information density and ensure that operators can quickly understand and execute decision recommendations with limited cognitive resources.
[0136] A second aspect of this invention provides a multimodal emergency human-computer interaction intelligent decision-making system, comprising:
[0137] The semantic fusion unit is used to perform semantic association analysis on the features of each modality in the multimodal feature sequence of human-computer interaction in emergency scenarios, extract complementary and redundant information between modalities, and generate a fused semantic representation vector.
[0138] The cognitive load unit is used to perform temporal dependency analysis on the human-computer interaction multimodal feature sequence by constructing a multi-level cognitive state graph, and to calculate the cognitive occupancy of each node through a graph convolution propagation mechanism and aggregate them to form a global cognitive load tensor.
[0139] The intent estimation unit is used to construct a human-machine bidirectional conditional probability inference framework based on the fused semantic representation vector and the global cognitive load tensor, perform joint posterior probability estimation of the operator's decision intent and machine executable constraints, and generate a human-machine bidirectional conditional probability inference framework. Figure 1 The intent distribution vector for consistency verification;
[0140] The process matching unit is used to perform semantic space mapping between the intent distribution vector and the standardized handling process, identify the handling process that best matches the current operator's intent, and extract decision nodes and resource scheduling constraints.
[0141] The decision scheme unit is used to dynamically construct a multi-granularity decision scheme set based on the dynamic weight coefficients of each candidate intention hypothesis in the intention distribution vector and the saturation distribution of the global cognitive load tensor. The decision granularity of the multi-granularity decision scheme set is inversely adjusted to the occupancy status of the global cognitive load tensor.
[0142] The decision output unit is used to perform constraint matching between the set of multi-granularity decision schemes and the decision node, and generate a decision output result by combining the resource scheduling constraints through logical reasoning.
[0143] A third aspect of the present invention provides an electronic device, comprising:
[0144] processor;
[0145] Memory used to store processor-executable instructions;
[0146] The processor is configured to invoke instructions stored in the memory to execute the aforementioned method.
[0147] A fourth aspect of the present invention provides a computer-readable storage medium having stored thereon computer program instructions that, when executed by a processor, implement the aforementioned method.
[0148] This invention can be a method, apparatus, system, and / or computer program product. The computer program product may include a computer-readable storage medium having computer-readable program instructions loaded thereon for performing various aspects of the invention.
[0149] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A multimodal emergency human-computer interaction intelligent decision-making method, characterized in that, include: Semantic association analysis is performed on the modal features in the multimodal feature sequence of human-computer interaction in emergency scenarios to extract complementary and redundant information between modalities and generate a fused semantic representation vector; The temporal dependency analysis of the human-computer interaction multimodal feature sequence is performed by constructing a multi-level cognitive state graph. The cognitive occupancy of each node is calculated by graph convolution propagation mechanism and aggregated to form a global cognitive load tensor. Based on the fused semantic representation vector and the global cognitive load tensor, a human-machine bidirectional conditional probability inference framework is constructed. Joint posterior probability estimation is performed on the operator's decision intention and the machine's executable constraints to generate an intention distribution vector that has passed the human-machine intention consistency verification. The intent distribution vector is semantically mapped to the standardized handling process to identify the handling process that best matches the current operator's intent and to extract decision nodes and resource scheduling constraints. Based on the dynamic weight coefficients of each candidate intent hypothesis in the intent distribution vector and the saturation distribution of the global cognitive load tensor, a multi-granularity decision scheme set is dynamically constructed. The decision granularity of the multi-granularity decision scheme set is inversely related to the occupancy status of the global cognitive load tensor. The set of multi-granularity decision schemes is matched with the decision nodes under constraints, and the decision output result is generated by combining the resource scheduling constraints with logical reasoning.
2. The method according to claim 1, characterized in that, Semantic association analysis is performed on the modal features in the multimodal feature sequence of human-computer interaction in emergency scenarios to extract complementary and redundant information between modalities and generate a fused semantic representation vector, including: A multimodal semantic dependency hypergraph structure is constructed, with each modal feature in the human-computer interaction multimodal feature sequence as a hypergraph node. The semantic collaboration strength between arbitrary modal subsets is calculated through a multimodal combination association measurement mechanism. A set of hyperedges reflecting multimodal combination semantics is established. The number of nodes connected by each hyperedge in the hyperedge set is adaptively determined by the multimodal combination association measurement mechanism based on the semantic collaboration strength. Hypergraph decomposition is performed based on the multimodal semantic dependency hypergraph structure. The dominant semantic subspace and redundant semantic subspace are extracted by eigenvalue decomposition. The feature vector set corresponding to the dominant semantic subspace represents the complementary information component that contributes the most to the fusion semantics among the features of each modality. The feature vector set corresponding to the redundant semantic subspace represents the redundant information component that repeatedly expresses the same semantics among the features of each modality. The modal features in the human-computer interaction multimodal feature sequence are orthogonally projected onto the dominant semantic subspace to suppress the information components in the redundant semantic subspace, thereby obtaining the decoupled semantic component representations of each modality. The semantic components of each modality are represented by cross-modal interaction modeling, and the cross-modal interaction modeling results are reduced in dimensionality through tensor decomposition to generate the fused semantic representation vector.
3. The method according to claim 1, characterized in that, Temporal dependency analysis was performed on the multimodal feature sequences of human-computer interaction by constructing a multi-level cognitive state graph. The cognitive occupancy of each node was calculated through a graph convolutional propagation mechanism and aggregated to form a global cognitive load tensor, including: The human-computer interaction multimodal feature sequence is hierarchically structured and modeled according to the time dimension to construct a multi-level cognitive state map containing a single-moment cognitive state layer, a continuous time period cognitive state layer, and a cross-time period cognitive state layer, and directed edge connections reflecting temporal dependencies are established within and between each level. Based on the topological structure of the multi-level cognitive state graph, the feature representations of nodes at each level are iteratively updated through a multi-level graph convolutional propagation mechanism. During the graph convolutional propagation process, the feature contributions of adjacent nodes are weighted and aggregated according to the temporal dependency strength of directed edges. The node features of the continuous time period cognitive state layer and the cross time period cognitive state layer are propagated and fused to the single time period cognitive state layer through a cross-level feature transfer mechanism. Based on the updated node feature representations of each level according to the multi-level graph convolutional propagation mechanism, the cognitive occupancy value corresponding to each node in the multi-level cognitive state graph is calculated. The cognitive occupancy value reflects the degree to which the cognitive state features represented by the corresponding node occupy the cognitive resources of the current operator. The cognitive occupancy values of all nodes in the multi-level cognitive state graph are tensorized and organized according to the hierarchical structure and temporal relationship, and aggregated to form the global cognitive load tensor.
4. The method according to claim 1, characterized in that, Based on the fused semantic representation vector and the global cognitive load tensor, a human-machine bidirectional conditional probability inference framework is constructed. This framework performs joint posterior probability estimation of the operator's decision intent and machine executable constraints, generating an intent distribution vector that has passed human-machine intent consistency verification. The fused semantic representation vector is used as an observation variable to input the prior probability distribution of the operator's decision intention, and the global cognitive load tensor is used as a conditional variable to input the conditional probability distribution of the machine-executable constraints. A human-machine bidirectional conditional probability inference framework is constructed. The prior probability distribution of the operator's decision intention represents the probability distribution of the operator's intention space, and the conditional probability distribution of the machine-executable constraints represents the probability distribution of the machine-executable action space under a given cognitive load state. Based on the aforementioned human-machine bidirectional conditional probability inference framework, a joint posterior probability estimation is performed on the prior probability distribution of the operator's decision intention and the conditional probability distribution of the machine's executable constraints through a variational Bayesian inference mechanism. A human-machine intention consistency constraint is introduced during the variational inference process, which requires minimizing the difference in probability distribution between the operator's decision intention and the machine's executable constraints. Based on the joint posterior probability estimation results output by the variational Bayesian inference mechanism, the posterior probability distribution of the operator's decision intention in the intention space is extracted, and it is verified whether the posterior probability distribution satisfies the human-machine intention consistency constraint. The posterior probability distribution that passes the human-machine intention consistency verification is then transformed into the intention distribution vector.
5. The method according to claim 1, characterized in that, The intent distribution vector is semantically mapped to the standardized handling process to identify the handling process that best matches the current operator's intent and to extract decision nodes and resource scheduling constraints, including: A two-layer semantic alignment network for intent-process is constructed, comprising an intent semantic projection layer and a process semantic projection layer. The intent semantic projection layer converts the intent distribution vector into an intent semantic embedding representation, and the process semantic projection layer converts the standardized processing flow into a process semantic embedding representation. The two-layer semantic alignment network maps the intent semantic embedding representation and the process semantic embedding representation to a unified semantic alignment space through a cross-layer semantic alignment mechanism. The intent-process hierarchical conditional probability inference framework is constructed by taking the intent probability distribution in the intent distribution vector as the top-level conditional probability source and establishing a multi-level conditional probability propagation link from the top-level intent probability distribution to the bottom-level processing process. In the multi-level conditional probability propagation link, inter-process semantic dependency constraints are introduced to represent the dependency and mutual exclusion relationships between different handling processes. The conditional probability is propagated layer by layer from the top-level intention probability distribution through the variational conditional probability inference mechanism and the inter-process semantic dependency constraints are fused. The handling process with the highest posterior activation probability of the bottom-level handling process is taken as the handling process that best matches the current operator's intention. The process structure of the handling procedure that best matches the current operator's intention is analyzed, and the decision node and the resource scheduling constraints are extracted.
6. The method according to claim 1, characterized in that, Based on the dynamic weight coefficients of each candidate intent hypothesis in the intent distribution vector and the saturation distribution of the global cognitive load tensor, a multi-granularity decision scheme set is dynamically constructed. The decision granularity of the multi-granularity decision scheme set is inversely adjusted to the occupancy status of the global cognitive load tensor, including: The dynamic weight coefficients of each candidate intention hypothesis are extracted from the intention distribution vector, and the saturation distribution is calculated from the global cognitive load tensor. An inverse adjustment mapping relationship is established between the granularity of the decision-making scheme and the cognitive load saturation. The granularity level of the decision-making scheme is adaptively determined according to the load occupancy level in the saturation distribution through a granularity feedback adjustment mechanism. The granularity level includes the abstraction level of the decision-making scheme and the expansion depth of the operation steps. When the saturation distribution indicates that the cognitive load occupancy level is increasing, the granularity level is reduced to reduce the expansion depth of the operation steps. When the cognitive load occupancy level is decreasing, the granularity level is increased to increase the expansion depth of the operation steps. Based on the dynamic weight coefficients, the importance of each candidate intent hypothesis is ranked. For the importance ranking, corresponding decision schemes are generated according to the granularity level. The generation process of the decision scheme is controlled according to the granularity level through a hierarchical granularity expansion mechanism. The depth of the operation steps of the decision scheme is inversely proportional to the cognitive load occupancy in the saturation distribution. The decision schemes generated for different candidate intent hypotheses are summarized to form the multi-granularity decision scheme set.
7. The method according to claim 1, characterized in that, The set of multi-granularity decision schemes is matched with the decision nodes under constraints, and the decision output is generated by combining the resource scheduling constraints with logical reasoning, including: Analyze the execution conditions and prerequisite dependencies of each decision scheme in the multi-granularity decision scheme set; The constraint propagation mechanism propagates the node type and logical constraint relationship between the decision nodes to each decision scheme in the multi-granularity decision scheme set, establishing a constraint matching relationship between the execution conditions of the decision scheme and the node type of the decision node. The constraint propagation mechanism selects decision schemes that satisfy the logical constraint relationship between the decision nodes based on the execution conditions and the prerequisite dependencies to form a constraint-matched decision scheme set. A resource constraint-driven logical reasoning link is constructed. The logical reasoning link receives the constraint matching decision scheme set and the resource scheduling constraints as input. The logical reasoning mechanism verifies the feasibility of each decision scheme in the constraint matching decision scheme set according to the resource scheduling constraints. The logical reasoning mechanism checks the matching degree between the resources required by each decision scheme and the currently available resources based on the resource availability conditions and resource conflict rules in the resource scheduling constraints. The decision schemes that pass the feasibility verification are selected, and the execution order of the selected decision schemes is sorted according to the logical constraint relationship of the decision nodes to obtain the decision output result.
8. A multimodal emergency human-computer interaction intelligent decision-making system, used to implement the method as described in any one of claims 1-7, characterized in that, include: The semantic fusion unit is used to perform semantic association analysis on the features of each modality in the multimodal feature sequence of human-computer interaction in emergency scenarios, extract complementary and redundant information between modalities, and generate a fused semantic representation vector. The cognitive load unit is used to perform temporal dependency analysis on the human-computer interaction multimodal feature sequence by constructing a multi-level cognitive state graph, and to calculate the cognitive occupancy of each node through a graph convolution propagation mechanism and aggregate them to form a global cognitive load tensor. The intent estimation unit is used to construct a human-machine bidirectional conditional probability inference framework based on the fused semantic representation vector and the global cognitive load tensor, perform joint posterior probability estimation of the operator's decision intent and machine executable constraints, and generate an intent distribution vector that has passed the human-machine intent consistency verification. The process matching unit is used to perform semantic space mapping between the intent distribution vector and the standardized handling process, identify the handling process that best matches the current operator's intent, and extract decision nodes and resource scheduling constraints. The decision scheme unit is used to dynamically construct a multi-granularity decision scheme set based on the dynamic weight coefficients of each candidate intention hypothesis in the intention distribution vector and the saturation distribution of the global cognitive load tensor. The decision granularity of the multi-granularity decision scheme set is inversely adjusted to the occupancy status of the global cognitive load tensor. The decision output unit is used to perform constraint matching between the set of multi-granularity decision schemes and the decision node, and generate a decision output result by combining the resource scheduling constraints through logical reasoning.
9. An electronic device, characterized in that, include: processor; Memory used to store processor-executable instructions; The processor is configured to invoke instructions stored in the memory to execute the method according to any one of claims 1 to 7.
10. A computer-readable storage medium having computer program instructions stored thereon, characterized in that, When the computer program instructions are executed by the processor, they implement the method described in any one of claims 1 to 7.