Dynamic interaction collaborative decision-making method and device, equipment and storage medium

By real-time preprocessing of multimodal data and construction of dynamic interactive flow graphs, combined with attention fusion and policy network inference, the problem of insufficient adaptability of traditional VLA models in dynamic interactive tasks is solved, and efficient and reliable collaborative decision-making is achieved.

CN120975228APending Publication Date: 2025-11-18PING AN TECH (SHENZHEN) CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202511063308.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-30
Publication Date
2025-11-18

AI Technical Summary

Technical Problem

Traditional VLA models cannot capture the dynamic changes of multimodal data in real time in dynamic interactive tasks, resulting in insufficient adaptability and reliability of collaborative decision-making, making it difficult to meet the needs of efficiency and multi-factor collaboration in fields such as finance and healthcare.

Method used

By receiving and preprocessing real-time multimodal data, a dynamic interactive flow graph is constructed, and deep fusion features are generated based on attention fusion. The policy network is used to infer and output collaborative decision-making actions, and the policy network is updated based on the execution results.

Benefits of technology

It improves the ability to capture multimodal data changes in dynamic interactions, enables timely feedback of collaborative decision-making actions, meets the needs of finance, healthcare and other fields for efficient decision-making and multi-factor collaboration, and enhances the adaptability and reliability of collaborative decision-making.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120975228A_ABST
    Figure CN120975228A_ABST
Patent Text Reader

Abstract

The invention relates to the fields of artificial intelligence, financial science and technology and digital medical treatment, discloses a collaborative decision-making method, device and equipment for dynamic interaction and a storage medium, and can be applied to collaborative decision-making processing in a dynamic interaction scene in the fields of finance and medical treatment. The method comprises the following steps: receiving multi-modal data collected in real time and preprocessing the multi-modal data to obtain multi-modal interaction features; constructing a dynamic interaction flow graph based on the multi-modal interaction features; utilizing the dynamic interaction flow graph to guide the multi-modal interaction features to generate deep fusion features based on attention fusion; and inputting the deep fusion feature and the dynamic interaction flow graph into a strategy network for reasoning, and outputting a collaborative decision action. According to the method, the dynamic interaction flow graph is constructed by adopting the multi-modal interaction characteristics which are acquired and preprocessed in real time, and the characteristics are fused based on the dynamic interaction flow graph, so that the capturing capability of multi-modal data changes is improved, dynamic extraction and fusion of the characteristics are realized, and the requirements of financial and medical fields on decision-making high efficiency and multi-element collaboration are met.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the fields of artificial intelligence, financial technology and digital medicine, and in particular to a dynamic interaction collaborative decision-making method, device and equipment and storage medium. BACKGROUND

[0002] In the collaborative decision-making of dynamic interaction tasks, the visual-language-action (VLA) model is mainly relied on to realize the integration of multi-modal data and decision output. Among them, the VLA model is a technical model capable of processing visual, language and action multi-modal data, and realizing multi-modal information interaction and collaborative processing. By integrating data of different modalities, the understanding, analysis and decision-making of complex scenes are realized.

[0003] With the development of the current VLA model, the collaborative decision-making technology of multi-modal data interaction tasks has been widely penetrated into the financial, medical and other industries, but it still faces many technical bottlenecks in the application process. In particular, the traditional VLA model mostly adopts a static feature extraction and fusion method, which cannot track and understand the dynamic changes of multi-modal data in the interaction process in real time. For example, in the intelligent financial advisor service scenario in the financial field, when the service robot processes customer asset data, customer voice consultation and its own mobile posture, the traditional model is difficult to capture the fluctuations of the data, the immediate adjustment of the consultation and the dynamic changes of the posture in real time, resulting in delayed advisor recommendations, which may reduce customer service experience and investment efficiency; for example, in the remote multi-specialist collaborative surgery in the medical field, when the system integrates real-time action data of surgical instruments, high-definition surgical field pictures and voice guidance between experts, the traditional model is difficult to capture the rapid adjustment of the instrument action, the real-time change of the picture and the emergency change of the voice instruction in real time, which may lead to collaboration failure and threaten patient safety; for example, in the multi-robot collaborative drug delivery task in the medical field, when multiple robots collaborate to deliver drugs, the traditional model is difficult to capture the action coordination between robots, posture adjustment, and dynamic changes of language instructions of the operator in real time, resulting in low collaboration efficiency and even task failure.

[0004] Therefore, the adaptability and reliability of the traditional VLA model in dynamic interaction tasks are significantly insufficient, and it is difficult to meet the demand for decision-making efficiency and multi-element collaboration in actual application scenarios. SUMMARY

[0005] The present application provides a dynamic interaction collaborative decision-making method, device, equipment and storage medium to solve the technical problem that the existing VLA model collaborative decision-making cannot meet the demand for decision-making efficiency and collaboration.

[0006] In a first aspect, a dynamic interaction collaborative decision-making method is provided, comprising:

[0007] receive real-time collected multi-modal data and pre-process the multi-modal data to obtain multi-modal interaction features;

[0008] construct a dynamic interaction flow graph based on the multi-modal interaction features;

[0009] generate deep fusion features by fusing the multi-modal interaction features based on attention using the dynamic interaction flow graph as guidance;

[0010] input the deep fusion features and the dynamic interaction flow graph into a policy network for inference, and output a collaborative decision action;

[0011] collect an execution result of the collaborative decision action, and update the policy network based on the execution result.

[0012] In a second aspect, a collaborative decision device for dynamic interaction is provided, comprising:

[0013] a receiving module configured to receive real-time collected multi-modal data and pre-process the multi-modal data to obtain multi-modal interaction features;

[0014] a constructing module configured to construct a dynamic interaction flow graph based on the multi-modal interaction features;

[0015] a fusing module configured to generate deep fusion features by fusing the multi-modal interaction features based on attention using the dynamic interaction flow graph as guidance;

[0016] a decision module configured to input the deep fusion features and the dynamic interaction flow graph into a policy network for inference, and output a collaborative decision action;

[0017] an updating module configured to collect an execution result of the collaborative decision action, and update the policy network based on the execution result.

[0018] In a third aspect, a computer device is provided, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the collaborative decision method for dynamic interaction when executing the computer program.

[0019] In a fourth aspect, a computer readable storage medium is provided, which stores a computer program, wherein the computer program is executable on a processor to implement the steps of the collaborative decision method for dynamic interaction.

[0020] Compared with the prior art, the application has the beneficial effects that: the application adopts the multi-modal interactive features after real-time collection and preprocessing to construct a dynamic interactive flow graph, and fuses the multi-modal interactive features based on the dynamic interactive flow graph, thereby improving the capture ability of the multi-modal data changes in the dynamic interaction, and realizing dynamic extraction and fusion of the features, so that the collaborative decision action can be timely fed back, and the needs of the fields such as finance and medical treatment for decision efficiency and multi-element collaboration are met; through strategy network reasoning and result feedback updating, the adaptability and reliability of the collaborative decision are improved, and the accuracy of the collaborative decision is optimized.

[0021] The above description is only a summary of the technical scheme of the application, in order to more clearly understand the technical means of the application, the application can be implemented according to the content of the specification, and in order to make the above and other purposes, characteristics and advantages of the application more obvious and easy to understand, the following preferred embodiments are described in detail as follows. BRIEF DESCRIPTION OF DRAWINGS

[0022] Figure 1 is an application environment schematic diagram of the collaborative decision method of the dynamic interaction in an embodiment of the application;

[0023] Figure 2 is a flow schematic diagram of the collaborative decision method of the dynamic interaction in an embodiment of the application;

[0024] Figure 3 is Figure 2 is a specific implementation flow schematic diagram of step S10 in the application;

[0025] Figure 4 is Figure 2 is a specific implementation flow schematic diagram of step S20 in the application;

[0026] Figure 5 is Figure 2 is another specific implementation flow schematic diagram of step S30 in the application;

[0027] Figure 6 is Figure 2 is a specific implementation flow schematic diagram of step S40 in the application;

[0028] Figure 7 is Figure 2 is a specific implementation flow schematic diagram of step S50 in the application;

[0029] Figure 8 is a structure schematic diagram of the collaborative decision device of the dynamic interaction in an embodiment of the application;

[0030] Figure 9 is a structure schematic diagram of the computer device in an embodiment of the application;

[0031] Figure 10is another structural schematic view of the computer device in an embodiment of the present application. DETAILED DESCRIPTION

[0032] In order to make the objectives, technical solutions and advantages of the present application clearer, further detailed description will be made to the present application in combination with the accompanying drawings and specific embodiments. The technical solutions in the embodiments of the present application will be described clearly and completely in combination with the accompanying drawings of the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative work fall within the scope of protection of the present application.

[0033] It should be understood that, when used in the specification and the appended claims, the terms “comprise” and “include” indicate the presence of the described features, integers, steps, operations, elements, and / or components, but do not exclude one or more other features, integers, steps, operations, elements, components, and / or groups thereof.

[0034] It should also be understood that the terms used in the present application specification are only for the purpose of describing specific embodiments and are not intended to limit the present application. As used in the present application specification and the appended claims, the singular forms “a”, “an” and “the” are intended to include the plural forms, unless the context clearly indicates otherwise.

[0035] It should be further understood that the term “and / or” used in the present application specification and the appended claims means any combination of one or more of the associated listed items and all possible combinations, and includes these combinations.

[0036] Please refer to Figure 1 and Figure 2 , Figure 1An application scenario schematic diagram of the dynamic interactive collaborative decision-making method provided by the embodiment of the present application is shown in FIG. 1. In the application scenario, an execution terminal communicates with a server through a network. The execution terminal is configured to collect visual information, instruction voice and action data collected in real time in a dynamic interactive scenario, and can also receive a collaborative decision-making action returned by the server to drive an interactive subject such as a robot to execute the collaborative decision-making action. The server can receive the multi-modal data collected in real time through the network and pre-process the multi-modal data to obtain multi-modal interaction features. The server can also construct a dynamic interaction flow graph based on the multi-modal interaction features. The server can guide the multi-modal interaction features to be fused based on attention by using the dynamic interaction flow graph, and generate deep fusion features. The server can input the deep fusion features and the dynamic interaction flow graph into a policy network for inference, and output a collaborative decision-making action. The server can collect an execution result of the collaborative decision-making action, and update the policy network based on the execution result. The execution terminal is a device integrated with detection and control, including but not limited to a robot controller, a smart robot body and an AGV control system. The server can be implemented by an independent server or a server cluster composed of multiple servers. The present application will be described in detail through specific embodiments.

[0037] Referring to FIG. 1, Figure 2 Figure 2 An application scenario schematic diagram of the dynamic interactive collaborative decision-making method provided by the embodiment of the present application is shown in FIG. 1. In the application scenario, an execution terminal communicates with a server through a network. The execution terminal is configured to collect visual information, instruction voice and action data collected in real time in a dynamic interactive scenario, and can also receive a collaborative decision-making action returned by the server to drive an interactive subject such as a robot to execute the collaborative decision-making action. The server can receive the multi-modal data collected in real time through the network and pre-process the multi-modal data to obtain multi-modal interaction features. The server can also construct a dynamic interaction flow graph based on the multi-modal interaction features. The server can guide the multi-modal interaction features to be fused based on attention by using the dynamic interaction flow graph, and generate deep fusion features. The server can input the deep fusion features and the dynamic interaction flow graph into a policy network for inference, and output a collaborative decision-making action. The server can collect an execution result of the collaborative decision-making action, and update the policy network based on the execution result. The execution terminal is a device integrated with detection and control, including but not limited to a robot controller, a smart robot body and an AGV control system. The server can be implemented by an independent server or a server cluster composed of multiple servers. The present application will be described in detail through specific embodiments.

[0038] S10: receiving multi-modal data collected in real time and pre-processing the multi-modal data to obtain multi-modal interaction features.

[0039] Step S10 is an input link of the whole process. By pre-processing the multi-modal data collected in real time, the problem of incomplete single modal information is solved. The pre-processed multi-modal data retains the key features of the original information, and provides high-quality input for the construction of the dynamic interaction flow graph.

[0040] The pre-processing is a process of cleaning, converting and standardizing the original collected data, which is used to eliminate noise interference, unify data format, and ensure the consistency and availability of the multi-modal data. The multi-modal data refers to data types from different information channels. In the embodiment, the multi-modal data includes visual information, instruction voice and action data, which are used to comprehensively capture various information in the dynamic interactive scenario, provide complete data basis for subsequent pre-processing and feature extraction, and provide input basis for the VLA model, so as to realize accurate collaborative decision-making.

[0041] ​In a specific implementation, for example, in the intelligent financial advisor service scenario, the position, posture and motion trajectory information of the service robot is obtained through visual detection, and these information of the service robot is the visual information, the instruction voice issued by the customer or the service personnel to the service robot is the instruction voice, and the guiding gesture and other actions of the service robot are the action data. After preprocessing, the visual information, the instruction voice and the action data provide basic data support for the interactive decision system of the intelligent advisor, and the interactive decision system can accurately judge the customer intention, the service personnel demand and the current state of the service robot, and then support the intelligent advisor system to plan a more practical action strategy for the service robot, so as to realize more efficient and personalized financial service guidance. For another example, in the remote surgery guidance scenario in the medical field, the picture of the surgery site obtained through video acquisition is the visual information, the voice guidance content given by the expert for the surgery process is the instruction voice, and the posture and action speed of the surgical instrument in the surgery are recorded as the action data. After preprocessing, these information provide basic data support for the remote decision system, and then support the expert to give remote accurate guidance.

[0042] In some embodiments of the present application, as shown in Figure 3 A specific preprocessing scheme is provided, in S10, that is, the real-time collected multi-modal data is received and preprocessed to obtain multi-modal interaction features, specifically including the following steps S11-S15.

[0043] S11: receiving visual information, instruction voice and action data collected in real time in a dynamic interaction scenario.

[0044] In this embodiment, the dynamic interaction scenario refers to an environment in which real-time information exchange and behavior interaction exist, such as machine cooperation and human-computer interaction. It can be understood that the multi-modal data of the present embodiment includes the above-mentioned visual information, instruction voice and action data, and the real-time collection of multi-modal data relies on the execution terminal, therefore, the execution terminal integrates the visual, voice and action collection functions. Preferably, the execution terminal can use multiple high-definition cameras and depth cameras to collect visual information of the dynamic interaction scenario from multiple angles in real time; through a microphone array, instruction voice in the interaction process is collected in real time; with the help of inertial sensors, joint angle sensors and other devices, action data of the robot is collected in real time.

[0045] In specific implementation, the execution terminal can be a robot controller, which can detect the state of each robot in the robot system, centrally process all sensor signals of the robot system, run a bottom-layer control algorithm, and drive each motor actuator to execute; the execution terminal can also be an intelligent robot body, which is equipped with a control system and can also collect interactive information such as user voice and gesture action by using a camera module carried by itself, and can also feed back information required by the user according to the user's interactive information; the execution terminal can also be an AGV (Automated Guided Vehicle) control system, which can control an automatic driving unmanned carrier vehicle and has functions such as environment perception, positioning and navigation, motion control, communication management, and safety monitoring.

[0046] S12: detecting an interactive subject from the visual information by using a target detection algorithm, and extracting a dynamically changing visual feature vector of the interactive subject by using a visual feature extraction network.

[0047] Step S12 helps to realize the conversion of visual information from original pixels to structured features, solves the problem that traditional static features cannot reflect the motion trend by capturing dynamic changing features, and ensures the accurate association of the visual feature vector and the interactive subject through the linkage of target detection and feature extraction, thereby improving the pertinence of the visual feature vector.

[0048] It can be understood that the target detection algorithm is a technology for identifying specific objects in images / videos to locate key interactive subjects such as people and objects from complex visual information. In this embodiment, a YOLOv9 target detection algorithm is used for detection of the interactive subject. The YOLOv9 target detection algorithm solves the information loss problem in deep network training through programmable gradient information and lightweight architecture design, significantly optimizes the balance between speed and accuracy, and enables accurate capture of interactive subjects in dynamic interactive scenes.

[0049] It can be understood that the visual feature extraction network is a kind of deep neural network, which is used to convert the visual information such as image or video frame of the detected interactive subject into high-dimensional vector representation, and encode the dynamic changing posture, motion and other key features. In this embodiment, a Transformer visual feature extraction network such as Swin TransformerV2 is used to extract a visual feature vector F_V\in\mathbb{R}^{m\times d_V} containing appearance, action details and other information, where m is the number of features and d_V is the dimension of the visual feature, so as to improve the expression richness and accuracy of the visual feature by modeling the visual features of different scales, and provide high-quality visual basis for subsequent construction of dynamic interactive flow graph and fusion of multi-modal interactive features.

[0050] Preferably, while the visual feature vector is extracted, the visual features of adjacent frames are subjected to time difference calculation to capture the dynamic change information of the interactive subject. By comparing the differences in visual features of adjacent frames, the motion trend, speed change, posture adjustment and other dynamic information of the interactive subject can be directly captured. For example, in a multi-robot cooperative handling task, the moving direction, speed and relative motion state of the robot hand with other objects can be perceived in real time, making up for the deficiency of static features in describing dynamic processes.

[0051] S13: Obtain the language feature vector of the instruction voice through voice recognition technology and a language large model.

[0052] For step S13, by converting the instruction voice into semantic features and combining voice recognition and language large model coding, the key information and interactive intent in the instruction voice can be accurately captured to generate a language feature vector that retains semantic connotation and is easy to fuse with other modal features, thereby improving the cross-modal understanding ability of the VLA model.

[0053] In specific implementation, the instruction voice can be converted into text using voice recognition technology, and the text can be subjected to word segmentation, part-of-speech tagging and sentiment analysis to identify instruction keywords, interactive intent and sentiment tendency. Then, a pre-trained language model such as ChatGLM is used to encode the text to obtain the language feature vector F_L\in\mathbb{R}^{n\times d_L}, n being the number of language features and d_L being the dimension of language features. In addition, a language interaction history record is established to record the semantic association information of the preceding and subsequent instruction voices. The language interaction history record can clearly indicate the cause-and-effect relationship between the preceding and subsequent instruction voices, such as the preceding instruction being the cause and the subsequent instruction being the effect. From the perspective of time sequence and logical relationship, the language interaction history record can sort out the time sequence and internal logical chain of the instruction voices to ensure coherent understanding of the interactive intent and facilitate accurate anticipation of subsequent actions by the VLA model.

[0054] S14: Extract the time sequence dynamic characteristics of the action data through a time convolution network to obtain an action feature vector.

[0055] For step S14, the problem that traditional static features cannot reflect the time sequence changes of actions is solved. Through the long-time dependency modeling capability of the time convolution network, the rhythm, intensity and trend of the action can be accurately captured, and the generated action feature vector contains both instantaneous state and dynamic trend, providing a key basis for understanding the intent of the interactive behavior and a core support for reasoning for collaborative decision-making actions.

[0056] Among them, the temporal convolutional network (TCN) is a convolutional neural network good at processing time series data, which captures long-time dependence through dilated convolution. In the present embodiment, it is used to extract the time series dynamic characteristics of the action data. The time series dynamic characteristics are the change rules of the action in the time dimension, such as speed, acceleration, and action interval, which are used to reflect the dynamic properties of the interactive behavior.

[0057] In specific implementation, the collected action data can be first filtered and normalized to remove noise interference; then the time series dynamic characteristics of the action are extracted through the TCN, and thus the action feature vector F_A\in\mathbb{R}^{p\timesd_A} is obtained, where p is the number of action features, and d_A is the dimension of the action features. At the same time, the dynamic parameters such as the speed and acceleration of the action are calculated to enrich the action feature representation.

[0058] S15: Align and integrate the visual feature vector, the language feature vector, and the action feature vector to generate a multi-modal interactive feature.

[0059] Among them, the feature alignment refers to matching the features of different modalities in the time dimension, such as a certain visual frame corresponding to the speech and action at the same time. In the present embodiment, it is used to ensure the spatio-temporal consistency of the cross-modal features. Step S15 solves the cross-modal data asynchronization problem by using accurate time alignment, and retains the unique information of each modality while establishing the association by using feature integration, thereby providing a unified feature input for subsequent dynamic interactive flow modeling, i.e., the construction of the dynamic interactive flow graph, and improving the overall understanding ability of the VLA model to complex interactive scenarios.

[0060] In specific implementation, the dynamic time warping (DTW) algorithm can be used to allow a time deviation of ±100ms for matching, and the feature timestamp cross-validation is used to ensure the alignment accuracy; the feature splicing method is used in the integration process to obtain the multi-modal interactive feature F_{pre}=[F_V;F_L;F_A], wherein the visual feature vector can be grouped according to the target, the language feature vector can be segmented according to the sentence, and the action feature vector can be arranged according to the time step.

[0061] S20: Construct a dynamic interactive flow graph based on the multi-modal interactive feature.

[0062] It can be understood that in the present embodiment, the server carries the VLA model, i.e., the visual-language-action model, which is used for the construction of the dynamic interactive flow graph, the deep fusion of features, the reasoning of collaborative decision actions, and the reasoning optimization. The dynamic interactive flow graph is the product of the dynamic interactive flow modeling in the VLA model, and provides guidance for the interactive events and relationships for the deep fusion of features.

[0063] For step S20, converting the multi-modal interaction features into a structured graph representation solves the problem that traditional sequence models are difficult to capture complex causal relationships, and the dynamic interaction flow graph structure can be updated in real time following the multi-modal interaction features collected and pre-processed in real time, ensuring that the VLA model can track the continuous changes of the interaction scene and improve the depth of understanding of the dynamic scene.

[0064] In some embodiments of the present application, as shown in Figure 4 A specific dynamic interaction graph construction scheme is provided, S20, that is, a dynamic interaction flow graph is constructed based on the multi-modal interaction features, specifically including steps S21-S23.

[0065] S21: Using pre-defined interaction event types and interaction event detection algorithms, interaction events are identified from the multi-modal interaction features, and the interaction events are taken as nodes.

[0066] In this embodiment, the interaction event types are pre-defined typical interaction behavior categories, such as "object handover", "instruction transmission", "action coordination", etc., and the interaction event types are the classification criteria for interaction event identification. The interaction event detection algorithm is a hybrid algorithm combining rules and machine learning, which is used to accurately identify interaction events from multi-modal interaction features F_{pre}. Rules are used to handle explicit logic, and machine learning is used to judge ambiguous scenarios. For example, when two robot hands are detected to approach, and the action data shows a grasping and releasing action, and there is a related instruction in the language data, it is determined as an "object handover" event; for example, in the application scenario of an intelligent audit counter for insurance claims, the interaction event types of "material submission", "question raising", "signature confirmation", etc. are defined, and the interaction event is identified as "signature confirmation" from the features of the monitoring video uploaded by the customer, the customer's consultation voice and signature action, so as to perform the next step of feedback, such as feedback of the customer's claim progress and prompt of the customer's related matters needing attention, reducing manual intervention operation; for example, in the remote surgery application scenario, the interaction event types of "instrument transfer", "hemostasis operation", "suture", etc. are pre-defined, and the current instrument is identified to be performing a gauze winding action from the features of the surgery video, the doctor's instructions and the instrument movement, so that the current interaction event type is determined as "hemostasis operation", which is beneficial to realize the automatic recording and standardized evaluation of the surgery process.

[0067] S22: Obtain the temporal relationship, causal relationship and logical relationship between the nodes through a graph neural network, and take the temporal relationship, causal relationship and logical relationship as edges.

[0068] The graph neural network (GNN) is a neural network specially processing graph structure data, aggregates the features of nodes and their neighbors through a message passing mechanism, and is also used to mine deep relationships between node interaction events from multi-modal interaction features, so as to improve the accuracy of relationship identification. The temporal relationship refers to the sequence of interaction events in the time dimension; the causal relationship refers to the driving and driven relationship between events, such as "instruction transmission" leading to "action execution"; the logical relationship refers to the internal relationship of interaction events formed according to domain rules, such as "grabbing" usually followed by "moving". The temporal relationship, the causal relationship and the logical relationship jointly constitute the attributes of the edges of the dynamic interaction flow graph in the embodiment, and completely depict the association mode of the interaction events.

[0069] The following is a specific implementation example of step S22: a 2-layer GraphSAGE framework is used as the core structure of the GNN, 64-dimensional hidden features are set in each layer, and the relationship between nodes is processed through a "sampling-aggregation-updating" process, wherein the temporal relationship is determined by the difference between event timestamps, such as a time interval less than 2s being marked as a direct temporal edge, the causal relationship is determined by the conditional probability calculation of multi-modal features, such as "the probability of action feature change after the appearance of instruction feature > 0.85" being determined as a causal edge, and the logical relationship is matched based on a preset rule library, such as "object handover" needing to be completed before "grabbing", and the edge weights are assigned according to the strength of the relationship, for example, the weight of the causal relationship is 0.6, the weight of the temporal relationship is 0.2, and the weight of the logical relationship is 0.2, which can also be dynamically adjusted according to the interaction scene.

[0070] S23: updating the feature representation of the node through a message passing mechanism and the temporal relationship, the causal relationship and the logical relationship, and constructing a dynamic interaction flow graph according to the node and the edge.

[0071] The message passing mechanism is a process in which nodes in the GNN update their own features through neighbor node features, and in the embodiment, the message passing mechanism is used to allow the node features to fuse the information of the neighbor nodes, so as to enhance the expression of the context of the interaction events. The dynamic interaction flow graph is a dynamic graph structure composed of updated nodes and edges, which is a structured representation of the dynamic interaction scene in the embodiment, and can be updated in real time through reward feedback and a message passing mechanism, so as to efficiently reflect the changes of the dynamic interaction scene.

[0072] In specific implementation, the dynamic interaction flow graph is recorded as G_{IF}=(V,E), where V is a set of interaction event nodes, and E is an edge set. The feature representation of the node is updated through a message passing mechanism, and in the embodiment, the node feature update formula is defined as:

[0073] h_i^{(l+1)}=\sigma\left(\sum_{j\inN(i)}W^{(l)}h_j^{(l)}+b^{(l)}\right),

[0074] wherein h_i^{(l)} represents the feature of node i in the l-th layer of the dynamic interaction flow graph, N(i) is the neighbor node set of node i, W^{(l)} and b^{(l)} are learnable weights and biases, and sigma is an activation function. With the continuous generation of new interaction events, i.e., nodes, the updating of node features enables the dynamic interaction flow graph to be updated in real time following the changes in multi-modal interaction data.

[0075] S30: generating a deep fusion feature by guiding the multi-modal interaction features to be fused based on attention using the dynamic interaction flow graph.

[0076] wherein the attention fusion is an attention mechanism-based feature fusion method that highlights key information by calculating attention weights between features. In this embodiment, the multi-modal feature fusion is guided by the dynamic interaction flow graph, so that it focuses more on information related to the current interaction event. The deep fusion feature is a high-dimensional feature vector that integrates visual, language, action features and dynamic interaction flow graph information. In this embodiment, it is used as the core input of the strategy network and can comprehensively reflect the multi-modal data and relationship structure of the interaction scene.

[0077] For step S30, the attention mechanism is more accurately focused on features strongly related to the current interaction event by the guidance of the dynamic interaction flow graph. The deep fusion feature contains both the details of the multi-modal interaction features and the interaction event relationship, solving the problem of ignoring interaction relationships in traditional static feature fusion. It provides more comprehensive input for subsequent decision-making and enhances the adaptability of the VLA model to dynamic scenes.

[0078] In some embodiments of the present application, as shown in Figure 5 a specific feature fusion scheme is provided. In S30, i.e., the multi-modal interaction features are fused based on attention using the dynamic interaction flow graph to generate a deep fusion feature, which specifically includes steps S31-S34.

[0079] S31: calculating attention weights of the node features and the multi-modal interaction features.

[0080] wherein the attention weight is a numerical value that measures the correlation between two multi-modal interaction features, and the numerical value is 0 to 1. The higher the weight, the stronger the correlation. In this embodiment, it is used to quantify the degree of association between the dynamic interaction flow graph node and the multi-modal interaction feature, providing a priority basis for feature fusion.

[0081] S32: generating attention-enhanced features by weighted aggregation of the multi-modal interaction features based on the attention weights.

[0082] wherein the weighted aggregation is a process of weighted summation of the multi-modal interaction features according to the attention weights, and in the present embodiment, the weighted aggregation gives higher weights to the features strongly related to the interaction event, and enhances the proportion of the features in the aggregation result.

[0083] S33: fusing the attention-enhanced features through a multi-layer perception mechanism.

[0084] wherein the multi-layer perception mechanism (MLP) is a feature processing module composed of a multi-layer fully connected neural network, and realizes feature fusion through nonlinear transformation, and in the present embodiment, is used to deeply integrate the visual, language and action features after attention enhancement, and to mine the cross-modal correlation. The fusion refers to converting the features of different modalities into a unified representation, while retaining the unique information of each modality and establishing a correlation. The multi-layer perception mechanism (MLP) has nonlinear capability, and can break through the limitation of simple concatenation, and realize organic fusion of multi-modal features.

[0085] S34: outputting the fused attention-enhanced features as deep fusion features by using residual connection and layer normalization technology.

[0086] wherein the residual connection is a connection mode of directly adding the input attention-enhanced features to the network output, and in the present embodiment, is used to alleviate the gradient vanishing problem of deep network, and retain the original feature information. The layer normalization is a normalization processing of the output features of the neural network layer, and in the present embodiment, is used to stabilize the training process, accelerate the convergence of the fusion model, and enhance the generalization ability of the attention-enhanced features.

[0087] For steps S31-S34, by quantifying the correlation strength of the nodes and the multi-modal interaction features, the subsequent feature fusion can preferentially focus on the features strongly related to the current interaction event, filter irrelevant information interference, and improve the signal-to-noise ratio and pertinence of the fused features; by weighted aggregation, the attention weights are converted into actual feature enhancement, so that the details related to the current interaction event in the multi-modal interaction features, such as instruction keywords and key actions, are strengthened, and irrelevant information is weakened, providing more focused and effective intermediate features for subsequent deep fusion, and improving the quality of cross-modal fusion; through the nonlinear transformation capability of the multi-layer perception mechanism, the multi-modal interaction features after attention enhancement are improved from physical concatenation to semantic fusion, the deep correlation between cross-modalities is mined, such as the correlation between the “turn left” instruction and the corresponding limb action, and more semantically consistent fused features are generated; the residual connection ensures that the key information in the attention-enhanced features is not annihilated by the deep network; and the layer normalization solves the instability problem of fusion caused by the distribution difference of multi-modal interaction features, so that the deep fusion features contain not only cross-modal semantic correlation, but also core information of the original features, and have more stable distribution characteristics, providing high-quality input for the strategy network.

[0088] The following is a specific implementation example of steps S31-S34: combining the node features of the dynamic interaction flow graph G_{IF} with the preprocessed multi-modal interaction features F_{pre} to design an interaction flow guided multi-head attention mechanism. Taking visual features as an example, the attention weight between each visual feature and the interaction event node is calculated to highlight the visual information related to the current interaction event. The attention weight calculation formula is defined as:

[0089] e_{ij}=\text{LeakyReLU}(a^T[W_1f_{V_i}\parallelW_2h_{j}]),

[0090] where f_{V_i} is an element of the visual feature vector, h_{j} is the feature vector of interaction event node j, a^T is a learnable attention vector, W_1 and W_2 are weight matrices, and \parallel represents vector concatenation. After normalizing the attention weight, the visual features are weighted and aggregated to obtain attention-enhanced visual features. Similarly, the above processing is performed on language and action features to obtain attention-enhanced features. Then, the attention-enhanced features are concatenated and fused through a multi-layer perception mechanism. At the same time, residual connection and layer normalization techniques are introduced to enhance the expressiveness of the features and the training stability of the model, resulting in deep fusion features F_{fusion}.

[0091] S40: input the deep fusion features and the dynamic interaction flow graph into the policy network for inference and output the collaborative decision action.

[0092] The policy network is a decision-making model based on a deep neural network. In this embodiment, feature fusion is the link that the VLA model processes the mutual relationship of multi-modal interaction features, and the deep fusion features output by it are input into the policy network together with the dynamic interaction flow graph. The policy network is the core part of the collaborative decision reinforcement learning in the VLA model, which is used to output the collaborative decision action, and its parameters are continuously optimized through the online model update mechanism. The pre-processing, dynamic communication graph construction, feature fusion, and collaborative decision formation form a complete processing and decision-making link that supports each other. The collaborative decision action is the next action that the VLA model directly guides the interactive subject in the dynamic interaction scenario to perform. By inputting the deep fusion features and the dynamic interaction flow graph, the policy network can consider both the specific features and the overall logic of the interaction scenario, and the output collaborative decision action not only conforms to the real-time state but also follows the interaction rules, solving the one-sidedness problem caused by the traditional decision-making model relying on a single feature and improving the accuracy and coordination of decision-making.

[0093] In specific implementation, a collaborative decision-making policy network based on a deep neural network, \pi_{\theta}(a|s), can be constructed in advance, which takes the deep fusion feature F_{fusion} and the encoded representation of the dynamic interaction flow graph G_{IF} as the input state s, and outputs the probability distribution of the collaborative decision-making action a. Preferably, the policy network adopts a multi-layer fully connected neural network structure combined with a Transformer module, which can enhance the processing capability of long sequence interaction information and significantly improve the continuity and adaptability of time series decision-making in dynamic interaction scenarios.

[0094] In some embodiments of the present application, as shown in Figure 6 a specific decision-making scheme is provided, S40, that is, the deep fusion feature and the dynamic interaction flow graph are input into the policy network for inference, and the collaborative decision-making action is output, specifically including the following steps S41-S43.

[0095] S41: input the encoded state of the deep fusion feature and the dynamic interaction flow graph into the policy network.

[0096] S42: output the probability distribution of the collaborative decision-making action through the policy network.

[0097] S43: output the collaborative decision-making action according to the probability distribution.

[0098] For steps S41-S43, the encoded state of the dynamic interaction flow graph refers to a low-dimensional vector obtained by compressing and encoding the dynamic interaction flow graph through a graph neural network, which contains the overall information of event nodes and relationship edges, providing a structured relationship basis for the dynamic interaction scenario; the probability distribution refers to a set of probability values assigned by the policy network to all possible collaborative decision-making actions, and the sum of the probabilities is 1, which reflects the rationality of each action in this embodiment, and quantifies the priority of each action through the probability distribution.

[0099] In specific implementation, a greedy selection and constraint checking mechanism can be used to determine the output result, preferentially selecting the action with the highest probability in the probability distribution, and checking whether the action meets the safety constraints, such as whether the robot action will cause a collision, whether the surgical action is beyond the safe range, etc. If the constraints are not met, the action with the second highest probability is selected and rechecked until an action that meets the constraints is selected as the collaborative decision-making action.

[0100] For example, in the collaboration of intelligent financial consultant robots, the collaborative decision-making actions of "robot A showing products" and "robot B recording demands" are selected according to the probability distribution to receive customers and improve service efficiency; for example, in the guidance of intelligent robot rehabilitation training, the collaborative decision-making action of "adjusting the angle of the instrument" is selected according to the probability distribution to ensure that the exhibition action of the intelligent robot meets the needs of the patient's rehabilitation stage and improves the training effect.

[0101] S50: collecting an execution result of the cooperative decision action, and updating the policy network based on the execution result.

[0102] For step S50, by constructing a closed loop of "decision-making - execution - feedback - optimization", the policy network is optimized in reverse, solving the problem of decision-making deviation of the policy network caused by environmental changes in a dynamic interaction scenario, so that the VLA model can continuously adapt to scenario changes and improve long-term decision-making performance.

[0103] In some embodiments of the present application, as shown in Figure 7 The specific updating scheme S50, i.e. collecting an execution result of the cooperative decision action, and updating the policy network based on the execution result, specifically includes the following steps S51-S55.

[0104] S51: collecting an execution result of the cooperative decision action.

[0105] S52: obtaining a reward feedback of the execution result according to a reward function.

[0106] S53: generating a state-action-reward sequence according to an input of the policy network, the cooperative decision action, and the reward feedback.

[0107] For steps S51-S53, the execution result is the actual effect data generated after the interactive agent executes the cooperative decision action, including task completion degree, interaction efficiency, action safety, etc., which is used as raw data to evaluate the quality of decision-making in this embodiment. The reward function is a function that quantifies the execution result into a reward value, which is used to evaluate the pros and cons of the cooperative decision action and provide a learning signal for the policy network. Positive reward indicates good decision-making, and negative reward indicates poor decision-making. The reward feedback is a numerical value calculated by the reward function, which is used to guide the model to learn towards a better decision-making direction. In this embodiment, the state-action-reward sequence refers to a set of triplets arranged in chronological order, which can be used as training samples for policy network updating. The dispersed input, action, and reward information is organized into structured time-series training samples, preserving the dynamic correlation of the interaction process, and providing high-quality data for parameter optimization of the policy network through reinforcement learning algorithms, ensuring that the learning process can capture the time-series dependency.

[0108] In specific implementation, a reward function R for dynamic multi-modal interaction scenarios is designed, considering factors such as task completion degree, interaction efficiency, and action safety. For example, in a robot cooperative handling task, a positive reward is given for successfully handling an object, and a negative reward is given for collision or task timeout. The model interacts with the environment according to the cooperative decision action output by the policy network, obtains the reward signal r fed back by the environment, and records the state-action-reward sequence as (s_t, a_t, r_t).

[0109] S54: calculating a loss function of the policy network according to the state-action-reward sequence by using a proximal policy optimization algorithm.

[0110] S55: updating the policy network based on the loss function.

[0111] The proximal policy optimization (PPO) algorithm is a commonly used reinforcement learning algorithm, which guarantees training stability by limiting the policy update amplitude, i.e., a clipping parameter; the loss function is a function for measuring the difference between the predicted value of the policy network and the target value, which is constructed based on the proximal policy optimization algorithm, and preferably, the loss function formula of the policy network is:

[0112] L(\theta)=\mathbb{E}_{s_t,a_t\sim\pi_{\theta_{old}}}[\min(r_t(\theta)

[0113] \hat{A}_t,\text{clip}(r_t(\theta),1-\epsilon,1+\epsilon)\hat{A}_t)],

[0114] wherein r_t(\theta) is an importance sampling ratio, \hat{A}_t is an advantage function estimate, and \epsilon is a clipping parameter; the parameters \theta of the policy network are updated by a back propagation algorithm, so that the policy network can gradually learn the optimal collaborative decision-making strategy.

[0115] In some embodiments of the present application, after S50, i.e., collecting the execution results of the collaborative decision-making actions, and updating the parameters of the policy network based on the execution results, the following steps are further included:

[0116] monitoring whether the execution results and the multi-modal data change in real time;

[0117] If the execution results and the multi-modal data change, the execution results are used as training samples to update the parameters of the policy network online.

[0118] The online updating mode can adjust the parameters in real time during the running of the model without interrupting the service, so that the model can quickly adapt to the dynamically changing scene. The execution results and the corresponding multi-modal data, decision-making actions and other information are used for incremental training of the policy network, real-time monitoring and incremental updating, so that the policy network can quickly respond to changes in the scene.

[0119] In specific implementation, not only the parameters of the strategy network are updated, but also the construction of the dynamic interaction flow graph and the multi-modal interaction feature fusion module are optimized, the detection accuracy of the model on the interaction event, the feature fusion mode and the like are adjusted, and through continuous online updating and optimization, the model can quickly adapt to the changes of the dynamic multi-modal interaction scene, and the accuracy and efficiency of the collaborative decision are continuously improved.

[0120] It can be seen that in the above scheme, the multi-modal interaction features after real-time collection and preprocessing are used to construct the dynamic interaction flow graph, and the multi-modal interaction features are fused based on the dynamic interaction flow graph, which improves the capture ability of the multi-modal data changes in the dynamic interaction, and realizes dynamic extraction and fusion of the features, so that the collaborative decision action can be fed back in time, and the needs of the fields such as finance and medical treatment for decision efficiency and multi-element collaboration are met; through the strategy network reasoning and result feedback updating, the adaptability and reliability of the collaborative decision are improved, and the accuracy of the collaborative decision is optimized.

[0121] It should be understood that the size of the serial number of each step in the above embodiment does not mean the order of execution, and the execution order of each process should be determined according to its function and inherent logic, and should not constitute any limitation on the implementation process of the embodiment of the present application.

[0122] In an embodiment, the present application provides a collaborative decision device 100 for dynamic interaction, which corresponds to the collaborative decision method for dynamic interaction in the above embodiment. As shown in the figure, the collaborative decision device 100 for dynamic interaction comprises a receiving module 101, a construction module 102, a fusion module 103, a decision module 104 and an updating module 105. The detailed description of each functional module is as follows: Figure 8

[0123] The receiving module 101 is used for receiving real-time collected multi-modal data and performing preprocessing to obtain multi-modal interaction features.

[0124] The construction module 102 is used for constructing a dynamic interaction flow graph based on the multi-modal interaction features.

[0125] The fusion module 103 is used for guiding the multi-modal interaction features to be fused based on attention using the dynamic interaction flow graph, to generate deep fusion features.

[0126] The decision module 104 is used for inputting the deep fusion features and the dynamic interaction flow graph into a strategy network for reasoning, and outputting a collaborative decision action.

[0127] The updating module 105 is used for collecting an execution result of the collaborative decision action, and updating the strategy network based on the execution result.

[0128] In an embodiment, the receiving module 101 is specifically used for: ​

[0129] receiving visual information, instruction speech and action data collected in real time in a dynamic interaction scenario;

[0130] detecting an interaction subject from the visual information by using a target detection algorithm, and extracting a dynamically changing visual feature vector of the interaction subject by using a visual feature extraction network;

[0131] obtaining a language feature vector of the instruction speech by using a speech recognition technology and a language large model;

[0132] extracting a time sequence dynamic characteristic of the action data by using a time convolution network to obtain an action feature vector;

[0133] aligning and integrating the visual feature vector, the language feature vector and the action feature vector to generate a multi-modal interaction feature.

[0134] In an embodiment, the construction module 102 is specifically configured to:

[0135] recognizing an interaction event from the multi-modal interaction feature by using a predefined interaction event type and an interaction event detection algorithm, and taking the interaction event as a node;

[0136] obtaining a time sequence relationship, a causal relationship and a logical relationship between the nodes by using a graph neural network, and taking the time sequence relationship, the causal relationship and the logical relationship as an edge;

[0137] updating a feature representation of the node by using a message passing mechanism and the time sequence relationship, the causal relationship and the logical relationship, and constructing a dynamic interaction flow graph according to the node and the edge.

[0138] In an embodiment, the fusion module 103 is specifically configured to:

[0139] calculating an attention weight of the feature of the node and the multi-modal interaction feature;

[0140] weighting and aggregating the multi-modal interaction feature based on the attention weight to generate an attention enhanced feature;

[0141] fusing the attention enhanced feature by using a multi-layer perception mechanism;

[0142] outputting the fused attention enhanced feature as a deep fusion feature by using a residual connection and a layer normalization technology.

[0143] In an embodiment, the decision module 104 is specifically configured to:

[0144] inputting the deep fusion feature and an encoded state of the dynamic interaction flow graph into a policy network;

[0145] output, by the policy network, a probability distribution of a cooperative decision action;

[0146] output, according to the probability distribution, the cooperative decision action.

[0147] In an embodiment, the updating module 105 is specifically configured to:

[0148] collect an execution result of the cooperative decision action;

[0149] obtain, according to a reward function, a reward feedback of the execution result;

[0150] generate, according to an input of the policy network, the cooperative decision action and the reward feedback, a state-action-reward sequence;

[0151] calculate, according to the state-action-reward sequence, a loss function of the policy network by using a proximal policy optimization algorithm;

[0152] update the policy network based on the loss function.

[0153] The specific limitations of the cooperative decision apparatus 100 for dynamic interaction can refer to the limitations of the cooperative decision method for dynamic interaction in the above, and will not be repeated here. Each module in the above cooperative decision apparatus 100 for dynamic interaction can be realized by software, hardware and combinations thereof, in whole or in part. The above modules can be embedded in or independent of the processor in the computer device in hardware form, or can be stored in the memory in the computer device in software form, so as to be called and executed by the processor to perform the operations corresponding to each module.

[0154] In one embodiment, a computer device 200 is provided, which can be a server, and its internal structure diagram can be as shown in Figure 9 The computer device 200 includes a processor 220, a memory and a network interface 250 connected through a system bus 210. The processor 220 of the computer device is configured to provide computing and control capabilities. The memory of the computer device 200 includes a non-volatile and / or volatile storage medium, an internal memory 240. The non-volatile storage medium 230 stores an operating system 231, a computer program 232 and a database 233. The internal memory 240 provides an environment for the operating system and the computer program in the non-volatile storage medium 230 to run. The network interface 250 of the computer device 200 is configured to communicate with an external execution terminal through a network connection. The computer program is executed by the processor 220 to implement the functions or steps of the cooperative decision method for dynamic interaction server. That is, the processor 220 implements the following steps when executing the computer program:

[0155] receive real-time collected multi-modal data and pre-process the multi-modal data to obtain multi-modal interaction features;

[0156] construct a dynamic interaction flow graph based on the multi-modal interaction features;

[0157] fuse the multi-modal interaction features based on attention using the dynamic interaction flow graph to generate deep fusion features;

[0158] input the deep fusion features and the dynamic interaction flow graph into a policy network for inference and output a collaborative decision action;

[0159] collect an execution result of the collaborative decision action and update the policy network based on the execution result.

[0160] In one embodiment, a computer device 300, which can be an execution terminal, has an internal structure as shown in Figure 10 The computer device includes a processor 320, a memory, a network interface 350, a display screen 370 and an input device 360 connected through a system bus 310. The processor 320 of the computer device is configured to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium 330 and an internal memory 340. The non-volatile storage medium 330 stores an operating system 331 and a computer program 332. The internal memory provides an environment for the operating system 331 and the computer program 332 in the non-volatile storage medium 330 to run. The network interface 350 of the computer device 300 is configured to communicate with an external server through a network connection. The computer program is executed by the processor 320 to implement a function or a step of an execution terminal side of a dynamic interaction collaborative decision method. That is, when the processor 320 executes the computer program 332, the following steps are implemented:

[0161] receive real-time collected multi-modal data and pre-process the multi-modal data to obtain multi-modal interaction features;

[0162] construct a dynamic interaction flow graph based on the multi-modal interaction features;

[0163] fuse the multi-modal interaction features based on attention using the dynamic interaction flow graph to generate deep fusion features;

[0164] input the deep fusion features and the dynamic interaction flow graph into a policy network for inference and output a collaborative decision action;

[0165] collect an execution result of the collaborative decision action and update the policy network based on the execution result.

[0166] In one embodiment, a computer readable storage medium is provided, having stored thereon a computer program, which when executed by a processor implements the following steps:

[0167] receiving real-time collected multi-modal data and pre-processing to obtain multi-modal interaction features;

[0168] constructing a dynamic interaction flow graph based on the multi-modal interaction features;

[0169] using the dynamic interaction flow graph to guide the multi-modal interaction features to be fused based on attention, to generate deep fusion features;

[0170] inputting the deep fusion features and the dynamic interaction flow graph into a policy network for inference, and outputting a collaborative decision action;

[0171] collecting the execution result of the collaborative decision action, and updating the policy network based on the execution result.

[0172] It should be noted that the functions or steps that the above computer readable storage medium or computer device can implement can be referred to the related description in the foregoing method embodiments, and will not be described here to avoid repetition.

[0173] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer readable storage medium. When the computer program is executed, it can include the processes of the above-mentioned embodiments. Any reference to memory, storage, database or other medium used in the embodiments provided by the present application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. As an illustration but not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0174] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the above-mentioned division of each functional unit and module is exemplified, and in actual application, the above-mentioned functions can be completed by different functional units and modules according to needs, that is, the internal structure of the device is divided into different functional units or modules to complete all or part of the functions described above.

[0175] The above-described embodiments are only used to illustrate the technical solutions of the present application, rather than limit them; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that the technical solutions recorded in the foregoing embodiments can be modified, or some technical features can be replaced by equivalents; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application, and should be included in the protection scope of the present application.

Claims

1. A method of dynamic interactive collaborative decision making, characterized by, The method comprises the following steps: receiving real-time collected multi-modal data and preprocessing to obtain multi-modal interaction features; constructing a dynamic interaction flow graph based on the multi-modal interaction features; using the dynamic interaction flow graph to guide the multi-modal interaction features to be fused based on attention, and generating deep fusion features; inputting the deep fusion features and the dynamic interaction flow graph into a policy network for reasoning, and outputting collaborative decision actions; collecting the execution results of the collaborative decision actions, and updating the policy network based on the execution results.

2. The method of dynamic interactive collaborative decision making of claim 1, wherein, The method comprises the following steps: receiving real-time collected multi-modal data and preprocessing to obtain multi-modal interaction features; receiving real-time collected visual information, instruction speech and action data in a dynamic interaction scene; detecting the interaction subject from the visual information by using a target detection algorithm, and extracting the dynamic change visual feature vector of the interaction subject by using a visual feature extraction network; obtaining the language feature vector of the instruction speech by using speech recognition technology and a language large model; extracting the time sequence dynamic characteristics of the action data by using a time convolution network to obtain an action feature vector; 3. The method of dynamically interactive, collaborative decision-making of claim 2, wherein, aligning and integrating the visual feature vector, the language feature vector and the action feature vector to generate multi-modal interaction features. The method comprises the following steps: using a pre-defined interaction event type and an interaction event detection algorithm to identify interaction events from the multi-modal interaction features, and taking the interaction events as nodes; obtaining the time sequence relationship, causal relationship and logical relationship between the nodes by using a graph neural network, and taking the time sequence relationship, causal relationship and logical relationship as edges; 4. The method of dynamically interactive, collaborative decision-making of claim 3, wherein, updating the feature representation of the nodes by using a message passing mechanism and the time sequence relationship, causal relationship and logical relationship, and constructing a dynamic interaction flow graph according to the nodes and the edges. The method comprises the following steps: calculating the attention weight of the node features and the multi-modal interaction features; weighting and aggregating the multi-modal interaction features based on the attention weight to generate attention enhanced features; fusing the attention enhanced features by using a multi-layer perception mechanism; 5. The method for dynamic interactive and collaborative decision making according to claim 4, wherein, outputting the fused attention enhanced features as deep fusion features by using residual connection and layer normalization technology. The method comprises the following steps: inputting the deep fusion features and the encoded state of the dynamic interaction flow graph into the policy network; outputting the probability distribution of the collaborative decision actions by using the policy network; 6. The method of dynamically interactive, collaborative decision-making of claim 5, wherein, outputting the collaborative decision actions according to the probability distribution. The method comprises the following steps: collecting the execution results of the collaborative decision actions; obtaining the reward feedback of the execution results according to a reward function; generating a state-action-reward sequence according to the input of the policy network, the collaborative decision actions and the reward feedback; calculating the loss function of the policy network according to the state-action-reward sequence by using a proximal policy optimization algorithm; updating the policy network based on the loss function.

7. The method of dynamically interactive, collaborative decision-making of claim 6, wherein, The method further comprises: monitoring the execution result and the multi-modal data in real time; if the execution result and the multi-modal data change, updating the parameters of the policy network online with the execution result as a training sample.

8. A dynamic interactive co-decision device, characterized in that, The method comprises: receiving real-time collected multi-modal data and pre-processing the multi-modal data to obtain multi-modal interaction features; constructing a dynamic interaction flow graph based on the multi-modal interaction features; fusing the multi-modal interaction features based on attention using the dynamic interaction flow graph to generate deep fusion features; inputting the deep fusion features and the dynamic interaction flow graph into a policy network for inference and outputting a collaborative decision action; updating the policy network based on the execution result of the collaborative decision action.

9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, The processor executes the computer program to implement the steps of the collaborative decision method of the dynamic interaction according to any one of claims 1 to 7.

10. A computer-readable storage medium storing a computer program, the computer program comprising instructions that, when executed by a computer, cause the computer to perform the method of any one of claims 1 to 9. The computer program is executed by the processor to implement the steps of the collaborative decision method of the dynamic interaction according to any one of claims 1 to 7.

Citation Information

Cited By

  • Collaborative decision determination method and device for equipment, equipment and medium

    CN121997267A

  • A multi-scene multi-modal interaction behavior prediction method and system based on a timing diagram

    CN122490425A