Human-machine cooperation assembly cognitive reasoning method oriented to space-time dynamic evolution

By constructing a spatiotemporal hypergraph of knowledge for human-machine collaborative assembly and using a stacked graph neural network with multi-event Hawkes processes, the problem of representing time-varying higher-order associations in personalized product assembly is solved, realizing an efficient human-machine collaborative assembly strategy and improving the accuracy and efficiency of the assembly process.

CN121094136APending Publication Date: 2025-12-09DONGHUA UNIV
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202511264140.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-05
Publication Date
2025-12-09

AI Technical Summary

Technical Problem

Existing technologies cannot effectively characterize the time-varying higher-order correlations of dynamic assembly during the assembly process of personalized products, resulting in inaccurate task allocation in human-machine collaboration and affecting the accuracy of assembly strategies.

Method used

We construct a knowledge spatiotemporal hypergraph for human-machine collaborative assembly, generate spatial hyperedges, temporal hyperedges, and task hyperedges through visual feature extraction, and combine a stacked graph neural network of multi-event Hawkes processes to learn high-order task associations between assembly nodes, capture time-varying features, and optimize network performance through a loss function to achieve cognitive reasoning for the dynamic assembly process.

Benefits of technology

It effectively characterizes time-varying unpaired relationships in the dynamic assembly process, improves the efficiency and accuracy of human-machine collaborative assembly, reduces redundant information interference, and reveals the dynamic triggering mechanism of time-varying tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121094136A_ABST
    Figure CN121094136A_ABST
Patent Text Reader

Abstract

The invention relates to a time-space dynamic evolution-oriented man-machine cooperation assembly cognitive inference method, which comprises the following steps of: extracting visual features in an assembly scene, generating a scene graph, constructing a time hyperedge, a space hyperedge and a task hyperedge, and fusing the three types of hyperedges to form a hyperedge set; time-varying non-pairwise relationships among human operators, robots, assembly operations and various assembly component nodes are represented, a hyperedge incidence matrix is defined, a man-machine cooperation assembly knowledge space-time hypergraph is constructed, and hyperedge representation among assembly components is realized; designing a stacked graph neural network with a self-excitation characteristic based on a multi-event hokes process, learning high-order task association among assembly nodes, and updating the change of the high-order task association along with time to realize space-time hypergraph representation; and modeling a self-excitation process among different sub-tasks, and capturing individual features and collective association to realize man-machine cooperation assembly. The problem that man-machine cooperation cognition in a time-varying task is difficult to infer is solved, and man-machine cooperation assembly efficiency and initiative are improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to a knowledge graph reasoning and human-computer collaboration technology, in particular to a human-computer collaboration assembly cognitive reasoning method for space-time dynamic evolution. BACKGROUND

[0002] At present, large-scale personalized production has gradually become a popular production mode. Personalized products are usually considered configurable products, which means that they can be customized according to customer requirements and options. These products allow multiple variants to be generated through different parameter and configuration combinations, resulting in multiple assembly paths. This requires highly flexible manufacturing capabilities and relies on the adaptability of operators to introduce dynamic characteristics into the assembly process. In order to cope with this challenge, human-robot collaboration has attracted attention from both industry and academia. Most researches do not endow human-robot collaboration systems with self-organizing ability of dynamic assembly task allocation, thereby affecting the active collaboration between human operators and robots. In order to study the bidirectional autonomous cognitive ability of human operators and robots, the main focus is on transforming environmental perception cues into graph structures, and focusing on pairwise relationships between operators, robots and components in discrete snapshots.

[0003] However, personalized product assembly involves a large number of diversified components, and these components do not have a fixed one-to-one relationship with humans and machines, but exhibit time-evolving many-to-many non-paired relationships. Therefore, the current reasoning method focusing on paired relationships in discrete snapshots cannot generate accurate assembly strategies. SUMMARY

[0004] In view of the problem that the method focusing on the paired relationship of the assembly components in the discrete snapshot cannot represent the time-varying high-order correlation of the dynamic assembly, resulting in inaccurate task allocation and affecting the human-robot collaboration cognitive strategy, the present application provides a human-robot collaboration assembly cognitive reasoning method for space-time dynamic evolution.

[0005] The technical scheme of the present application is as follows:

[0006] A human-robot collaboration assembly cognitive reasoning method for space-time dynamic evolution, comprising the following steps:

[0007] Step S100: By extracting the visual features in the assembly scene, a scene graph is generated, time hyper-edges, space hyper-edges and task hyper-edges are constructed, and the three types of hyper-edges are fused to form a hyper-edge set, representing the time-varying non-paired relationships between human operators, robots, assembly operations and various assembly component nodes, defining a hyper-edge association matrix, constructing a human-robot collaboration assembly knowledge space-time hypergraph, and realizing the hyper-edge representation between assembly components;

[0008] Step S200: design a stacked graph neural network with self-activation characteristics based on a multi-event Hawkes process, learn the high-order task correlation between each assembly node, and update its change over time to realize space-time hypergraph representation; model the self-activation process between different sub-tasks, capture individual features and collective correlation at the same time, reveal the evolution mechanism of the time-varying assembly task, decouple the operations that need to be completed by the human and the machine respectively, and realize efficient human-machine collaborative assembly.

[0009] Further, it specifically includes the following steps:

[0010] Step S100: construct a human-machine collaborative assembly knowledge space-time hypergraph to represent the time-varying non-paired relationship in the dynamic assembly process; perceive the assembly space through a depth camera, dynamically detect visual clues using a graph R-CNN, and convert them into a graph structure to realize the structured processing of scene features; define three types of hyperedges, spatial hyperedges, temporal hyperedges, and task hyperedges, and fuse them to form a hyperedge set for operators, robots, assembly operations, and assembly component nodes in the assembly scene: spatial hyperedges represent high-order correlation based on the Euclidean distance between nodes; temporal hyperedges represent time sequence dependency based on task timestamps and sliding time windows; task hyperedges represent collaborative dependency based on semantic information; finally, the features of the three types of hyperedges, spatial proximity, time dependency, and task correlation, are spliced and fused to generate a hyperedge feature vector, and the association between hyperedges and nodes is established through a hyperedge association matrix to complete the construction of the space-time hypergraph and realize comprehensive modeling of the multi-node time-varying non-paired relationship in the dynamic assembly process;

[0011] Step S200: design a stacked graph neural network with self-activation characteristics based on a multi-event Hawkes process to learn the high-order task correlation between nodes and capture the time-varying task evolution mechanism; the multi-event Hawkes process estimates the probability of future events by defining a conditional intensity function to model the space-time evolution law between multiple dynamic events; the stacked graph neural network with self-activation characteristics includes node embedding, hyperedge embedding, time fusion, and graph pooling modules, which integrate historical information through recursive aggregation and mapping of multi-layer neighbor features, and design a time representation network to realize feature transmission under time sequence dependency by fusing the Hawkes process, capture the time sensitivity of historical neighbor messages, and stack features layer by layer;

[0012] Step S300: optimize cognitive reasoning performance by decoupling individual features and collective correlation features; the individual features convert event prior experience into unique parameters through a learnable conversion model, dynamically adjust network parameters with a feature linear modulation model to ensure that prior experience quickly adapts to event uniqueness; the collective correlation features represent event features by connecting space-time hypergraph nodes, explore time sequence propagation rules using a graph neural network, and generate a new event predictor to capture systematic dynamic interaction features;

[0013] Step S400: defining loss functions to optimize network performance, including negative log-likelihood loss of individual feature capture optimization and smooth L1 loss of collective association feature capture optimization, ensuring consistency of hyperedge and node relationship transmission through joint optimization of the two types of losses, ultimately realizing dynamic cognitive reasoning of time-varying tasks in human-robot collaborative assembly process, decoupling human-robot operation and improving collaboration efficiency.

[0014] Further, in step S100:

[0015] The human-robot collaborative assembly knowledge spatio-temporal hypergraph generates time-varying non-paired relationships of each node by structuring scene features, representing dynamic assembly process;

[0016] The scene feature structuring perceives the assembly space through a depth camera, dynamically detects visual cues using a graph R-CNN over time, and converts the visual cues into a graph structure with assembly relationships;

[0017] The time-varying non-paired relationship representation is through the hyperedge of each node in the human-robot collaborative assembly knowledge spatio-temporal hypergraph, representing the assembly relationship involving multiple entities, modeling the time-varying non-paired relationship in dynamic assembly; at different time points, the human-robot collaborative assembly knowledge spatio-temporal hypergraph dynamically represents the operator, robot, assembly operation, assembly component and its time-varying non-paired relationship, relating the dynamic assembly states at different stages;

[0018] The time-varying non-paired relationship represents the many-to-many high-order relationship between multiple nodes under time-varying conditions, which is different from the one-to-one paired relationship, and has the advantages of complex relationship representation and improvement of redundant information interference.

[0019] Further, in step S100:

[0020] The spatial hyperedge, i.e. the geographical proximity between nodes, represents the complex high-order association between nodes in the same space; if the Euclidean distance between nodes is below a specified threshold, they belong to the same spatial hyperedge; the formula is defined as follows:

[0021] ε space ={e|distance(v i ,v j )<θ}

[0022] Where ε spdce represents the set of spatial hyperedges describing the geographical proximity between nodes; e represents a hyperedge containing spatially related nodes; distance represents the Euclidean distance between nodes v i and v jthe Euclidean distance between two nodes, and θ is a predefined distance threshold for determining whether the nodes belong to the same spatial hyperedge; if the spatial distance between two nodes is less than the threshold θ, they are considered to belong to the same hyperedge;

[0023] The time hyperedge represents the temporal dependency relationship between nodes in time-varying tasks; based on continuous task timestamps, multiple nodes with time dependency are grouped into the same hyperedge; a time window is set and the time hyperedge is constructed using task instructions, defined as follows:

[0024] ε time = {e'|time(v i ) < time(v j ), gap≤Δt}

[0025] where ε time represents a set of time hyperedges describing the temporal dependency relationship between nodes, e' represents a hyperedge containing time-dependent nodes, time(v i ) represents the timestamp of node v i occurrence, gap represents the time difference between two nodes, and Δt is the value of the time window; if the time difference between nodes is less than the time window value Δt, these nodes exhibit temporal dependency and are considered to belong to the same time hyperedge; if the time interval between these consecutive actions is less than the specified time window Δt, the nodes involved in these actions are classified as belonging to the same time hyperedge;

[0026] The task hyperedge represents the collaboration relationship between nodes in human-machine collaborative tasks; these hyperedges rely on language cues, task plans and process information to extract task semantics to build task hyperedges between multiple nodes, defined as follows:

[0027]

[0028] where ε task represents a set of task hyperedges describing the task dependency relationship between nodes; e'' represents a task hyperedge representing the relationship between nodes within the same human-machine collaborative task; v i and v j represent nodes involved in the same human-machine collaborative task; represent that node v i and node v j belong to the same human-machine collaborative task; these nodes are extracted from language cues and their relationship is defined as a task hyperedge;

[0029] The hyperedge set represents the fusion and collection of spatial hyperedges, time hyperedges and task hyperedges; this set comprehensively considers spatial, temporal and task features to establish relationships between nodes in human-machine collaboration; defined as follows:

[0030] H ε =concat(space features ,time features ,task features )

[0031] where H ε represents the hyperedge feature vector, space features represents the hyperedge space proximity feature, time features represents the hyperedge time dependency feature, and task features represents the hyperedge task related feature; concat represents the connection operation of combining multi-modal features into a vector; in general, the hyperedge features include task type, time dependency, priority information and tool use semantics;

[0032] The hyperedge association matrix represents the relationship between all hyperedges and nodes, and realizes the construction of the space-time hypergraph.

[0033] Further, the hyperedge association matrix takes the operator, robot, assembly operation, assembly component nodes, and space hyperedge, time hyperedge and task hyperedge association as an example, converts the relationship between the hyperedge and the node into a hyperedge connection matrix composed of rows and columns, and at the same time, defines the matrix value to determine whether there is a relationship between the hyperedge and the node, and the formula is defined as follows:

[0034] v={v operator ,v robot ,v operation ,v component}

[0035] ε={e Space ,e time ,e task}

[0036] e Space ={v operator ,v robot ,v operation}

[0037] e time ={v operator ,v operation ,v component}

[0038] e task ={v robot ,v operation ,v component}

[0039]

[0040] where v represents the node set, v operator ,vrobot operation component correspond to operator, robot, assembly operation, part respectively, ε represents hyperedge set, e space time task correspond to spatial hyperedge, temporal hyperedge and task hyperedge respectively, A ε represents hyperedge connection matrix describing the relationship between hyperedge and node, where each row represents a node and each column represents a hyperedge.

[0041] Further, in step S200:

[0042] The multi-event Hox process can determine whether multiple different nodes are associated at a certain time, estimate the future relationship between multiple nodes by defining the condition strength, i.e., defining the condition strength function of the triplets l, m, n forming a hyperedge at time t, generating the probability of occurrence of a certain event, modeling the evolution mechanism between multiple dynamic events in the time-varying task, and revealing the spatiotemporal evolution law of the event. The formula of the condition strength function is defined as follows:

[0043]

[0044]

[0045] Here, μ l,m,n (t) represents the basic rate of event nodes l, m and n forming a link at time t; S m′,n′ (t′)k(t-t′) represents the historical neighbor influence of event l at time t′ on the current event at time t; S l′,n′ (t′)k(t-t′) represents the historical neighbor influence of event m at time t′ on the current event at time t; S l′,m′ (t′)k(t-t′) represents the historical neighbor influence of event n at time t′ on the current event at time t; (l,m′,n′,t′)∈H l (t) represents the historical event of node l before time t; (l′,m,n′,t′)∈H m (t) represents the historical event of node m before time t; (l′,m′,n,t′)∈H n (t) represents the historical event of node n before time t; H l (t), H m (t) and H n ​​​​(t) respectively represent the set of historical events of nodes l, m and n before time t; in addition, m' and n' are regarded as the historical neighbors of l; similarly, other combination pairs can be defined; k(t-t') represents the influence function, t' represents a certain historical time, and k represents an exponential decay kernel, and k decays over time;

[0046] The conditional intensity describes the expected number of events occurring in a unit time interval at future time t influenced by past events; if the sequence of past event occurrence times is represented as {t1, t2, …, tn}, then the conditional intensity λ(t) at time t can be represented as: r

[0047]

[0048] Here, λ(t) represents the basic intensity function, which represents the event occurrence rate without any past event influence, and α(t-t i ) is the influence function, which represents the influence of past events on future event occurrence.

[0049] Further, in step S200:

[0050] The stacked graph neural network with self-activation characteristics receives, aggregates and maps the features of the multi-layer neighbor nodes through the node embedding, hyper-edge embedding, time fusion and graph pooling module, integrates the historical node features, and at the same time, designs a time representation network of the graph neural network, and integrates with the Hawkes process to realize the feature transmission between nodes under time sequence dependence; the formula is defined as follows:

[0051]

[0052] Here, is defined as the time representation of the event node l located in layer p, i.e., the event at time t, d p represents the dimension of the embedding vector; respectively represent the feature representation of the historical neighbors m' and n' located in layer p-1 at historical time t', represent the aggregated representation of the historical neighbors m' and n' in layer p-1 at historical time t', and σ represents an activation function, and and are all learnable weight matrices; at the same time, the embedding of the historical neighbors of is mapped to capture the time sensitivity of past events; in addition, within the kernel function k s (t-t'), the time decay effect is represented using SoftMax; the time representation of the node is described by aggregating the node's own information and the historical neighbor message from the previous layer; the layer-by-layer feature stacking is used for GNN. ​

[0053] Further, in step S200:

[0054] The individual characteristics are the non-pairwise relationships of nodes in the decoupled spatio-temporal hypergraph, which ensure the uniqueness of events, connect knowledge through encoding general relationships, generate prior experience with adaptive time-varying characteristics, build a learnable conversion model, make the prior experience quickly adapt to the uniqueness of each event, fit the transfer function for each event, and capture the individual characteristics of each event; the prior experience ex is converted into an event (l, m, n, t) with specific parameters δ l,m,n,t , as shown below:

[0055]

[0056] Wherein, is defined as the event-specific time representation of nodes l, m and n, and con() in it refers to the connection operation; based on The event prior ex is transformed into the unique parameters δ of the event through the learnable conversion model ψ l,m,n,t , which is parameterized by η;

[0057] The learnable conversion model is linearly modulated by features FiLM, which is a simple feature affine transformation on the intermediate layer features of the neural network, dynamically adjusts the network parameters, scales and moves the prior experience, and ensures that the prior experience can quickly adapt to the uniqueness of each event; scaling and moving are achieved through the scaling factor scf C and the moving factor shf (l,m,n,t) expressed by the fully connected layer FCL (l,m,n) , as follows:

[0058]

[0059] Where ω scf and ω shf are the learnable weight matrices of FCL C , B scf and B shf are the bias vectors of FCL C ;

[0060] The transfer function is an important part of generating conditional strength, which is realized by using the softplus function.

[0061] Further, in step S200:

[0062] Decouple the non-pairwise relationships of nodes in the spatio-temporal hypergraph, capture all event collective correlation characteristics, and ensure the systematicness and dynamic interaction characteristics of events;

[0063] The collective association feature decouples the non-paired relationships between nodes in the spatiotemporal hypergraph, ensuring the systematic and dynamic interaction characteristics of events. By connecting the nodes in the spatiotemporal hypergraph, it characterizes the event features associated with the nodes. Over time, a graph neural network is used to explore the temporal propagation characteristics of event features, fusing the spatial features and temporal evolution of events to capture the collective association features of all events. The propagation and trends of all relevant nodes in the assembly process—operator, robot, operation, and component nodes—are essentially the same concept. For ease of discussion, a new event is defined here as... This represents the new events that occur at time t for nodes l, m, and n; the following uses a fully connected layer FCL. C A predictor for new events of nodes was built:

[0064]

[0065] in, δ is defined as the event-specific time representation of nodes l, m, and n. I Indicates a fully connected layer (FCL) C Required parameters.

[0066] Furthermore, in step S200:

[0067] Define a loss function, optimize network performance, ensure the consistency of relational propagation of all hyperedges and nodes in the spatiotemporal hypergraph, model the evolution of time-varying unpaired relations in the spatiotemporal hypergraph, and realize human-computer collaborative cognitive strategy reasoning.

[0068] The loss function includes simultaneous optimization of individual event features and collective correlation features. Individual feature capture optimization is defined by negative log-likelihood loss, which can provide appropriate conditional strength for the occurrence or non-occurrence of events. Collective correlation feature capture optimization is defined by smooth L1 loss, which increases error penalty and achieves better network performance.

[0069] The individual feature capture optimization assumes that the event (l,m,n,t)∈O has already occurred, and its loss L n (l,m,n,t) is as follows:

[0070]

[0071] Where, λ l,m,n (t) and λ l,m,q (t) denotes the conditional strength of the formation of hyperedges for the ternary pairs (l,m,n) and (l,m,q) at time t, respectively, and q represents the conditional strength based on the distribution d. n The nodes of the sampled negative samples, the distribution defines the events. Instances that do not occur; meanwhile, Q ne This represents the number of negative samples for each event that occurred;

[0072] The collective correlation feature capture optimization, loss L C (l, m, n, t) as follows:

[0073]

[0074] Wherein, N represents the predicted node propagation value, N l,m,n N represents the true value of the node.

[0075] The beneficial effects of the present application are:

[0076] 1. For the redundant information brought by the pair relationship, a time hypergraph is constructed to represent the complex non-pair relationship between the time-varying tasks in human-machine collaboration.

[0077] 2. For the dynamic and time-varying nature of human-machine collaboration tasks, the self-excitation process of Hawkes process is introduced into the stacked GNN architecture to capture individual and collective correlation features of assembly tasks at the same time, so as to reveal the dynamic triggering mechanism between time-varying tasks. BRIEF DESCRIPTION OF DRAWINGS

[0078] Figure 1 A spatiotemporal dynamic evolution-oriented human-machine collaboration assembly cognitive reasoning method is provided.

[0079] Figure 2 A spatiotemporal dynamic evolution-oriented human-machine collaboration assembly cognitive reasoning method is provided.

[0080] Figure 3 A spatiotemporal dynamic evolution-oriented human-machine collaboration assembly cognitive reasoning method is provided. DETAILED DESCRIPTION

[0081] The present application will be described in detail below in conjunction with the drawings and specific embodiments. The present embodiment is implemented on the basis of the technical solution of the present application, and gives a detailed implementation and specific operation process, but the protection scope of the present application is not limited to the following examples.

[0082] As Figure 1 shown, the embodiment of the present application provides a spatiotemporal dynamic evolution-oriented human-machine collaboration assembly cognitive reasoning method, which comprises:

[0083] Step S100: as Figure 1 , 2As shown, by extracting visual features in the assembly scene, a scene graph is generated, time hyper-edges, space hyper-edges and task hyper-edges are constructed, and the three types of hyper-edges are fused to form a hyper-edge set, representing the time-varying non-paired relationship between human operators, robots, assembly operations and various assembly component nodes, defining a hyper-edge association matrix, constructing a human-robot collaborative assembly knowledge spatio-temporal hypergraph, and realizing the hyper-edge representation between assembly components.

[0084] In step S100, the human-robot collaborative assembly knowledge spatio-temporal hypergraph is generated by structuring the scene features, generating the time-varying non-paired relationship of each node, and representing the dynamic assembly process.

[0085] The scene feature structuring perceives the assembly space through a depth camera, dynamically detects visual clues over time using a graph R-CNN, and converts the visual clues into a graph structure with assembly relationships.

[0086] The time-varying non-paired relationship representation is realized through the hyper-edges of each node in the spatio-temporal hypergraph (i.e., the abbreviation of human-robot collaborative assembly knowledge spatio-temporal hypergraph), which represents the assembly relationship involving multiple entities, effectively modeling the time-varying non-paired relationship in dynamic assembly. At different time points, the human-robot collaborative assembly knowledge spatio-temporal hypergraph dynamically represents the operators, robots, assembly operations, assembly components and their time-varying non-paired relationships, and associates the dynamic assembly states at different stages.

[0087] Among them, the "dynamic assembly state" is the interaction relationship and task execution situation between man, machine and object at a certain time point in the assembly process.

[0088] And "associating the dynamic assembly states at different stages" is to string these time sequence states through the model to form a complete dynamic evolution representation of the assembly task.

[0089] The time-varying non-paired relationship represents the many-to-many high-order relationship between multiple nodes under the condition of time continuous change, which is different from the one-to-one paired relationship, and has the advantages of complex relationship representation and improvement of redundant information interference.

[0090] Note: The two names are different, one is a description of complex relationships, and the other is how to learn and represent such complex relationships.

[0091] The space hyper-edge, i.e., the geographical proximity between nodes, represents the complex high-order association between nodes in the same space. If the Euclidean distance between nodes is below a specified threshold, they belong to the same space hyper-edge, for example, nuts and bolts related to robots indicate that these components may need corresponding robot operations. The formula is defined as follows:

[0092] ε space = {e | distance(vi ,v j )<θ}

[0093] where ε space represents the spatial hyperedge set that describes the geographical proximity relationship between nodes. e represents the hyperedge that contains spatially related nodes. distance represents the Euclidean distance between nodes v i and v j in the assembly process, and θ is a predefined distance threshold for determining whether the nodes belong to the same spatial hyperedge. If the spatial distance between two nodes is less than the threshold θ, they are considered to belong to the same hyperedge. For example, if the wrench and the nut are very close to each other and both are close to the robot, and the distance between them is less than the threshold, it can be concluded that these nodes belong to the same hyperedge, indicating that the robot can use the wrench to tighten the nut.

[0094] The temporal hyperedge represents the temporal dependency relationship between nodes in time-varying tasks. Based on continuous task timestamps, multiple nodes with temporal dependencies are grouped into the same hyperedge. Set the time window and use the task instructions to construct the temporal hyperedge, defined as follows:

[0095] ε time = {e'|time(v i ) < time(v j ), gap≤Δt}

[0096] where ε time represents the temporal hyperedge set that describes the temporal dependency relationship between nodes, e' represents the hyperedge that contains temporally related nodes, time(v i ) represents the timestamp of the node v i occurrence, and gap represents the time difference between two nodes. Δt is the value of the time window. If the time difference between nodes is less than the time window value Δt, these nodes exhibit temporal dependency and are considered to belong to the same temporal hyperedge. For example, the operator places the nut in the correct position, the robot moves the wrench to the nut's position, and then rotates the wrench to tighten the nut. If the time interval between these consecutive actions is very short (less than the specified time window Δt), the nodes involved in these actions are classified as belonging to the same temporal hyperedge.

[0097] The task hyperedge represents the collaboration relationship between nodes in human-robot collaboration tasks. These hyperedges mainly rely on language cues, task plans, and process information to extract task semantics (e.g., robot transports the reducer cover, operator tightens the nut) to build task hyperedges between multiple nodes, defined as follows:

[0098]

[0099] where ε task represents the set of task hyper-edges describing the task dependency relationship between nodes. e” represents the task hyper-edge representing the relationship between nodes within the same human-robot collaborative task. i and v j represent the nodes involved in the same human-robot collaborative task. represents the node v i and the node v j belong to the same human-robot collaborative task. Typically, these nodes are extracted from linguistic cues, and their relationship is defined as a task hyper-edge. For example, in the task description “rotate the wrench to tighten the bolt and nut”, the wrench, nut, and bolt are considered to belong to the same task hyper-edge.

[0100] The hyper-edge set represents the fusion and collection of spatial hyper-edges, temporal hyper-edges, and task hyper-edges. This set comprehensively considers spatial, temporal, and task characteristics to establish relationships between nodes in human-robot collaboration. It provides a foundation for the best cognitive strategy of the robot, and the formula is defined as follows:

[0101] H ε = concat(space features ,time features ,task features )

[0102] where H ε represents the feature vector of the hyper-edge, space features represents the spatial proximity feature of the hyper-edge, time features represents the temporal dependency feature of the hyper-edge, and task features represents the task-related feature of the hyper-edge. Concat represents the concatenation operation of combining multi-modal features into a vector. Overall, the hyper-edge features include task type, temporal dependency, priority information, and tool usage semantics.

[0103] The hyper-edge association matrix represents the relationship between all hyper-edges and nodes, realizing the construction of a spatio-temporal hypergraph. Taking the operator, robot, assembly operation, assembly component nodes, and spatial hyper-edges, temporal hyper-edges, and task hyper-edges as examples (where the node and hyper-edge association is only a general example), the relationship between hyper-edges and nodes is converted into a hyper-edge connection matrix composed of rows and columns, and at the same time, the matrix value is defined to determine whether there is a relationship between hyper-edges and nodes, and the formula is defined as follows:

[0104] v = {v operator ,v robot ,v operation ,v component}

[0105] ε = {e space ,e time ,etask}

[0106] e space ={v operator ,v robot ,v operation}

[0107] e time ={v operator ,v operation ,v component}

[0108] e task ={v robot ,v operation ,v component}

[0109]

[0110] where v represents a node set, v operator ,v robot ,v operation ,v component correspond to operators, robots, assembly operations, and parts, respectively, ε represents a hyperedge set, e space ,e time ,e task correspond to spatial hyperedges, temporal hyperedges, and task hyperedges, respectively, A ε represents a hyperedge connection matrix describing the relationship between hyperedges and nodes, where each row represents a node and each column represents a hyperedge. For example, the first row represents node v operator , which belongs to hyperedge e space and hyperedge e time , so the first row will be [1, 1, 0]. If all nodes in a column belong to the same hyperedge, i.e., v∈e s , then A ε will be 1. (For example, e space ={v operator ,v robot ,v operation}, which generally refers to hyperedges between different nodes.)

[0111] Step S200: As Figure 3 shown, a stacked graph neural network with self-activation characteristics based on a multi-event Hoax process is designed to learn high-order task correlations between assembly nodes and update their changes over time, realizing spatio-temporal hypergraph representation; modeling the self-activation process between different sub-tasks while capturing individual features and collective correlations, revealing the evolution mechanism of time-varying assembly tasks, decoupling the operations that need to be completed by humans and machines respectively, and realizing efficient human-machine collaborative assembly.

[0112] The spatio-temporal hypergraph representation is a unified hypergraph representation of the complex relationships between the human and the machine in the assembly process, which is used to support cognitive reasoning and strategy generation.

[0113] The spatio-temporal hypergraph is a knowledge graph representing the human-machine collaboration relationship;

[0114] The spatio-temporal hypergraph representation is a learning and representation of the knowledge graph.

[0115] The multi-event Hawkes process can determine whether multiple different nodes are associated at a certain time, estimate the future relationship between multiple nodes by defining the condition strength (a function of the condition strength of the triplets l, m, n forming a hyperedge at time t), generate the probability of an event occurring, model the evolution mechanism between multiple dynamic events in the time-varying task, and reveal the event spatio-temporal evolution law. The formula of the condition strength function is defined as follows:

[0116]

[0117]

[0118] Here, μ l,m,n (t) represents the basic rate at which event nodes l, m, and n form a link at time t, which is independent of the historical assembly events. S m′,n′ (t′)k(t-t′) represents the historical neighbor influence of event l at time t′ on the historical neighbor of the current event at time t; S l′,n′ (t′)k(t-t′) represents the historical neighbor influence of event m at time t′ on the historical neighbor of the current event at time t; S l′,m′ (t′)k(t-t′) represents the historical neighbor influence of event n at time t′ on the historical neighbor of the current event at time t. (l, m′, n′, t′) ∈ H l (t) represents the historical event of node l before time t; (l′, m, n′, t′) ∈ H m (t) represents the historical event of node m before time t; (l′, m′, n, t′) ∈ H n (t) represents the historical event of node n before time t. H l (t), H m (t) and H n (t) represent the historical event set of nodes l, m, and n before time t, respectively. In addition, m′ and n′ are considered as the historical neighbors of l. Similarly, combinations such as l′ and m′, l′ and n′, etc. can be defined. k(t-t′) represents the influence function, t′ represents a certain historical time, and k represents an exponential decay kernel, which decays over time.

[0119] The conditional strength, a key concept in Hawkes processes, represents the probability of an event occurring at a future moment given historical information. It describes the expected number of future events occurring within a unit time interval influenced by historical events. If we represent the sequence of past event occurrence times as {t1, t2, ..., t...} r Then the conditional strength λ(t) at time t can be expressed as:

[0120]

[0121] Here, λ(t) represents the fundamental intensity function, indicating the event occurrence rate in the absence of any past events, while α(t) represents the event occurrence rate. i α(tt) is an influence function, representing the impact of past events on the occurrence of future events. Typically, α(tt) i The conditional strength is non-negative and decreases over time, indicating that the influence of past events on future events is diminishing. Essentially, conditional strength describes the rate at which a given past event influences the occurrence of a future event, a key concept in Hawkes processes.

[0122] The self-excited stacked graph neural network, through node embedding, hyperedge embedding, temporal fusion, and graph pooling modules, recursively receives, aggregates, and maps features from multiple layers (multiple layers refer to multiple vertically connected feature transformation layers in the stacked graph neural network; each layer performs nonlinear mapping and aggregation of input features through a weight matrix) of neighboring nodes, integrating historical node features. Simultaneously, a temporal representation network of the graph neural network is designed and fused with a Hawkes process to achieve feature transfer between nodes under temporal dependence.

[0123] The formula is defined as follows:

[0124]

[0125] here, d is defined as the temporal representation of an event node l (i.e., the node representation of an event) located in layer p at time t. p Indicates the dimension of the embedding vector. Let m′ and n′, located in layer p-1, represent the feature representations of historical neighbors m′ and n′ at historical time t′, respectively. Let represent the aggregate representation of historical neighbors m′ and n′ in layer p-1 at historical time t′, and σ represent an activation function. and Both are learnable weight matrices. Simultaneously, utilizing... Mapping the embeddings of its historical neighbors captures the temporal sensitivity of past events. Furthermore, in the kernel function k... s(t-t') is represented using SoftMax to capture the time decay effect. To effectively capture the time sensitivity between the base strength and the historical neighbor information, the temporal representation of a node must be described by aggregating the node's own information and the historical neighbor messages from the previous layer. Moreover, due to the complexity of the temporal hypergraph, the representational capacity of the model needs to be enhanced, thus a layer-wise feature stacking is used for GNN. More specifically, for the initialization of the node messages in the first layer of the network, the input node features can play an important role.

[0126] The individual characteristics, mainly decoupling the non-paired relationship of nodes in the spatio-temporal hypergraph, ensure the uniqueness of the event, generate prior experience with adaptive time-varying characteristics by encoding general relationship connection knowledge, build a learnable conversion model, make the prior experience quickly adapt to the uniqueness of each event, fit the transfer function for each event, and capture the individual characteristics of each event. The prior experience ex is converted into the event-specific parameter δ l,m,n,t (l, m, n, t), which is specifically as follows:

[0127]

[0128] Wherein, is defined as the event-specific temporal representation of nodes l, m and n, and con() in it refers to the connection operation. Based on The event prior ex is transformed into the unique parameters δ l,m,n,t of the event through the learnable conversion model ψ, and the process is parameterized by η.

[0129] The learnable conversion model dynamically adjusts the network parameters through feature linear modulation (FiLM), that is, a simple feature affine transformation on the intermediate layer features of the neural network, scales and moves the prior experience, and ensures that the prior experience can quickly adapt to the uniqueness of each event. Scaling and moving are achieved through the scaling factor scf C and the moving factor shf (l,m,n,t) expressed by the fully connected layer FCL (l,m,n) , which is specifically as follows:

[0130]

[0131] Where ω scf and ω shf are the learnable weight matrices of FCL C , B scf and B shf are the bias vectors of FCL C .

[0132] The transfer function is an important part of generating conditional strength, which is usually implemented using the softplus function.

[0133] Decouple the non-pairwise relationship of nodes in the space-time hypergraph, capture the collective correlation characteristics of all events, and ensure the systematicness and dynamic interaction characteristics of events;

[0134] The collective correlation characteristics mainly decouple the non-pairwise relationship of nodes in the space-time hypergraph, ensure the systematicness and dynamic interaction characteristics of events, represent the event characteristics associated with the nodes by connecting the nodes in the space-time hypergraph, explore the time sequence propagation characteristics of the event characteristics through the graph neural network as time goes by, and fuse the spatial characteristics and time evolution of the events to realize the capture of the collective correlation characteristics of all events. In the above assembly, the propagation and trend of all related nodes, operators, robots, operations, and component nodes are essentially a concept. For ease of discussion, here the new event is defined as , which represents a new event of nodes l, m, and n occurring at time t. Hereinafter, a fully connected layer FCL is used C A predictor of the new event of the node is constructed:

[0135]

[0136] wherein, is defined as the event-specific time representation of nodes l, m, and n, and δ I represents a fully connected layer FCL C required parameters.

[0137] A loss function is defined to optimize network performance and ensure the consistency of relationship transmission of all hyperedges and nodes on the space-time hypergraph, model the evolution of time-varying non-pairwise relationships in the space-time hypergraph, and realize the reasoning of human-robot collaborative cognitive strategy.

[0138] The loss function includes simultaneous optimization of individual event feature capture and collective correlation feature capture. The individual feature capture optimization is defined by a negative log likelihood, which can provide appropriate conditional intensity for the occurrence or non-occurrence of an event. The collective correlation feature capture optimization is defined by a smooth L1 loss, which increases the error penalty to achieve superior network performance.

[0139] The individual feature capture optimization assumes that the event (l, m, n, t) ∈ O has occurred, and its loss L n (l, m, n, t) is as follows:

[0140]

[0141] wherein, λ l,m,n (t) and λ l,m,q (t) represent the conditional intensity of the triplets (l, m, n) and (l, m, q) forming a hyperedge at time t, respectively, and q represents a node sampled based on a distribution d n of negative samples, and the distribution defines the event Examples that do not occur. ne represents the number of negative samples for each occurrence event.

[0142] The collective correlation feature captures optimization, and the loss L C (l, m, n, t) as follows:

[0143]

[0144] wherein, represent the predicted node propagation value, N l,m,n represent the true value of the node.

[0145] In summary, for the problem of human-robot collaborative cognitive reasoning in time-varying tasks, the application proposes a human-robot collaborative assembly cognitive reasoning method for spatio-temporal dynamic evolution, constructs a spatio-temporal hypergraph to represent time-varying non-paired relationships, and reduces the redundant interference of one-to-one paired relationships; design a stack graph neural network with a hawks process to learn the event correlation representation in the spatio-temporal hypergraph, reveal the individual characteristics and collective correlation of subtasks, realize robot cognitive reasoning, and improve the efficiency and initiative of human-robot collaborative assembly, which has important theoretical significance and practical value.

[0146] The above-described embodiments only express one embodiment of the application, which is described in detail and specifically, but it cannot be understood as a limitation on the scope of the patent. It should be noted that for ordinary skilled in the art, without departing from the concept of the application, a number of modifications and improvements can be made, which are within the scope of protection of the application. Therefore, the scope of protection of the patent of the application should be subject to the appended claims.

Claims

1. A human-machine collaborative assembly cognitive reasoning method oriented to spatio-temporal dynamic evolution, characterized in that, Comprise the following steps: Step S100: generate a scene graph by extracting visual features in the assembly scene, construct time hyperedge, space hyperedge and task hyperedge, and fuse the three types of hyperedges to form a hyperedge set, representing the time-varying non-paired relationship between human operators, robots, assembly operations and various assembly component nodes, define the hyperedge association matrix, construct the human-robot collaborative assembly knowledge space-time hypergraph, and realize the hyperedge representation between assembly components; Step S200: design a stacked graph neural network with self-activation characteristics based on a multi-event Hawkes process, learn the high-order task association between nodes, and update its change over time, realize the space-time hypergraph representation; model the self-activation process between different subtasks, capture individual features and collective association at the same time, reveal the evolution mechanism of time-varying assembly tasks, decouple the operations that need to be completed by humans and machines respectively, and realize efficient human-robot collaborative assembly.

2. The method of claim 1, wherein, Specifically, the following steps are included: Step S100: Construct a human-robot collaborative assembly knowledge space-time hypergraph to represent the time-varying non-paired relationship in the dynamic assembly process; perceive the assembly space through a depth camera, dynamically detect visual clues using a graph R-CNN and convert them into a graph structure to realize the structured processing of scene features; define three types of hyperedges, space hyperedge, time hyperedge and task hyperedge, for operators, robots, assembly operations and assembly component nodes in the assembly scene, and fuse them to form a hyperedge set: the space hyperedge is based on the Euclidean distance between nodes, representing the high-order association of geographical proximity; The time hyperedge is based on the task timestamp and the sliding time window, representing the time sequence dependency; The task hyperedge is based on semantic information, representing the collaborative dependency; finally, the features of the three types of hyperedges, spatial proximity, time dependency and task correlation, are spliced and fused to generate a hyperedge feature vector, and the association between hyperedges and nodes is established through a hyperedge association matrix to complete the construction of the space-time hypergraph, realizing the comprehensive modeling of the multi-node time-varying non-paired relationship in the dynamic assembly process; Step S200: design a stacked graph neural network with self-activation characteristics based on a multi-event Hawkes process, learn the high-order task association between nodes and capture the time-varying task evolution mechanism; the multi-event Hawkes process estimates the probability of future events by defining a conditional intensity function, modeling the spatio-temporal evolution law between multiple dynamic events; the stacked graph neural network with self-activation characteristics includes node embedding, hyperedge embedding, time fusion, and graph pooling modules, which integrate historical information through recursive aggregation and mapping of multi-layer neighbor features, and design a time representation network to realize feature transmission under time sequence dependency, capture the time sensitivity of historical neighbor messages, and stack features layer by layer; Step S300: decouple individual features and collective association features to optimize cognitive reasoning performance; the individual features convert event prior experience into unique parameters through a learnable conversion model, and dynamically adjust network parameters combined with a feature linear modulation model to ensure that prior experience quickly adapts to event uniqueness; the collective association features represent event features by connecting space-time hypergraph nodes, explore the time sequence propagation law using a graph neural network, and generate a new event predictor to capture systematic dynamic interaction features; Step S400: Define loss function to optimize network performance, including individual feature capture optimized negative log-likelihood loss and collective association feature capture optimized smooth L1 loss, ensure consistency of hyperedge and node relationship transmission through joint optimization of two types of loss, finally realize dynamic cognitive reasoning of time-varying task in human-machine collaborative assembly process, decouple human-machine operation and improve collaboration efficiency.

3. The method of claim 1, wherein, In step S100: The human-machine collaborative assembly knowledge spatio-temporal hypergraph represents the time-varying non-paired relationship of each node through scene feature structuring, representing the dynamic assembly process; The scene feature structuring perceives the assembly space through a depth camera, dynamically detects visual clues using R-CNN, and converts visual clues into graph structures with assembly relationships; The time-varying non-paired relationship representation is through the hyperedge of each node in the human-machine collaborative assembly knowledge spatio-temporal hypergraph, representing the assembly relationship involving multiple entities, modeling the time-varying non-paired relationship in dynamic assembly; At different time points, the human-machine collaborative assembly knowledge spatio-temporal hypergraph dynamically represents the operator, robot, assembly operation, assembly component and their time-varying non-paired relationship, relating the dynamic assembly status at different stages; The time-varying non-paired relationship represents the many-to-many high-order relationship between multiple nodes under time continuity, which is different from one-to-one paired relationship, and has the advantages of complex relationship representation and improvement of redundant information interference.

4. The method of claim 1, wherein, In step S100: The spatial hyperedge, i.e. the geographical proximity between nodes, represents the complex high-order association between nodes in the same space; if the Euclidean distance between nodes is below a specified threshold, they belong to the same spatial hyperedge; the formula is defined as follows: e space = {e | distance(u i ,v j ) < θ} where ε space represents a spatial hyperedge set describing the geographical proximity relationship between nodes; e represents a hyperedge containing spatially related nodes; distance represents the Euclidean distance between each node υ i and v j during the assembly process, and θ is a predefined distance threshold for determining whether the nodes belong to the same spatial hyperedge; if the spatial distance between two nodes is less than the threshold θ, they are considered to belong to the same hyperedge; The time hyperedge represents the time sequence dependency relationship between nodes in time-varying tasks; Based on continuous task timestamps, multiple nodes with time dependency are grouped into the same hyperedge; Set a time window and use task instructions to construct a time hyperedge, the formula is defined as follows: e time = {e' | time(u i ) < time(v j ), gap < At} where ε time represents a time hyperedge set describing the time-dependent relationship between nodes, e’ represents a hyperedge containing time-dependent nodes, time(υ i ) represents the timestamp of the node v i ’s appearance, gap represents the time difference between two nodes, and Δt is the value of the time window; if the time difference between nodes is less than the time window value Δt, these nodes exhibit time dependence and are considered to belong to the same time hyperedge; if the time interval between these consecutive actions is less than the specified time window Δt, the nodes involved in these actions are classified as belonging to the same time hyperedge; The task hyperedge represents the collaboration relationship between nodes in human-machine collaborative tasks; these hyperedges rely on language clues, task plans and process information to extract task semantics to build task hyperedges between multiple nodes, the formula is defined as follows: where ε task represents the set of task hyper-edges describing the task dependency relationships between nodes; e" represents the task hyper-edges representing the relationships between nodes within the same human-machine collaborative task; i and v j represent the nodes involved in the same human-machine collaborative task; represents the node v i and the node v j belong to the same human-machine collaborative task; these nodes are extracted from linguistic cues and their relationships are defined as task hyper-edges; The hyperedge set represents the fusion and set of spatial hyperedges, time hyperedges and task hyperedges; this set comprehensively considers spatial, temporal and task features to establish relationships between nodes in human-machine collaboration; the formula is defined as follows: H ε = concat(space features , time features , task features ) where H ε represents hyperedge space proximity features, space features represents hyperedge time dependent features, time features represents hyperedge task dependent features; task features represents hyperedge task dependent features; concat represents the connection operation of combining multi-modal features into a vector; in general, the hyperedge feature includes task type, time dependency, priority information and tool usage semantics; The hyperedge association matrix represents the relationship between all hyperedges and nodes, realizing spatio-temporal hypergraph construction.

5. The method of claim 4, wherein, The hyperedge association matrix, taking the operator, robot, assembly operation, assembly component nodes and spatial hyperedge, time hyperedge and task hyperedge association as an example, converts the relationship between hyperedges and nodes into a hyperedge connection matrix composed of rows and columns, at the same time, defines the matrix value to determine whether there is a relationship between hyperedges and nodes, the formula is defined as follows: v = {v operator , v robot , v operation , v component} ε = {e space , e time , e task} e space = {u operator , u robot , u operation} e time = {u operator , u operation , u component} e task = {u robot , u operation , u component ) where v represents the set of nodes, v operator , v robot , v operation , v component correspond to operators, robots, assembly operations, parts, respectively, e space , e time , e task correspond to spatial hyper-edges, temporal hyper-edges and task hyper-edges, respectively, A ε represents the hyper-edge incidence matrix describing the relationship between hyper-edges and nodes, where each row represents a node and each column represents a hyper-edge.

6. The method of claim 1, wherein, In step S200: The multi-event Hawkes process can determine whether multiple different nodes are associated at a certain time, estimate the future relationship between multiple nodes by defining the condition strength, i.e., defining the condition strength function of the triplets l, m, n forming a hyperedge at time t, generate the probability of occurrence of a certain event, model the evolution mechanism between multiple dynamic events in the time-varying task, and reveal the spatiotemporal evolution law of the event. The formula definition of the condition strength function is as follows: Here, μ l,m,n (t) represents the basic rate at which event nodes l, m and n form links at time t; S m′,n′ (t') k(t - t') represents the historical neighbor influence of event l's historical neighbors m' and n' at time t' on the current event's historical neighbors at time t; S l′,n′ (t') k(t - t') represents the historical neighbor influence of event m's historical neighbors l' and n' at time t' on the current event's historical neighbors at time t; S l′,m′ (t') k(t - t') represents the historical neighbor influence of event n's historical neighbors l' and m' at time t' on the current event's historical neighbors at time t; (l, m', n', t') e Hl(t) represents the historical events of node l prior to time t; (l', m, n', t') e H m (t) represents the historical events of node m prior to time t; (l', m', n, t') e H n (t) represents the historical events of node n prior to time t; H l (t), H m (t) and H n (t) respectively represent the historical event sets of nodes l, m and n prior to time t; Furthermore, m' and n' are considered as historical neighbors of l; Similarly, other combinations can be defined; k(t - t') represents the influence function, t' represents a certain historical time, k represents an exponential decay kernel, k decays over time; The conditional intensity describes the expected number of events occurring in a future time unit, given the history of events; if the ordered list of past event times is denoted as {t1, t2,..., t r}, then the conditional intensity λ(t) at time t can be expressed as: Here, λ(t) represents the base intensity function, which denotes the event occurrence rate without any past event influence, while α(t - t i ) is the influence function, which denotes the past event influence on future event occurrence.

7. The method of claim 1, wherein, In step S200: The stacked graph neural network with self-excitation characteristics receives, aggregates, and maps the features of multiple layers of neighbor nodes through the node embedding, hyperedge embedding, time fusion, and graph pooling modules, integrates the historical node features, and at the same time, designs a time representation network of the graph neural network, which is combined with the Hawkes process to realize the feature transmission between nodes under time sequence dependence; the formula definition is as follows: Here, The node of an event l, i.e., an event, defined as located in layer p, represents a temporal representation at time t, d p denotes the dimension of the embedding vector; denotes the feature representation of the historical neighbors m' and n' located in layer p-1 at historical time t', respectively, denotes the aggregated representation of the historical neighbors m' and n' in layer p-1 at historical time t', σ denotes an activation function, and and are learnable weight matrices; meanwhile, the embedding of its historical neighbors are mapped using to capture the temporal sensitivity of past events; further, within the kernel function k s (t-t'), a SoftMax is used to represent the temporal decay effect; the temporal representation of a node is described by aggregating its own information and the messages from the historical neighbors from the previous layer; a layer-wise feature stacking is used for GNNs.

8. The method of claim 1, wherein, In step S200: The individual feature is the non-paired relationship of the node in the decoupled spatiotemporal hypergraph, ensures the uniqueness of the event, generates prior experience with adaptive time-varying characteristics by encoding general relationship connection knowledge, builds a learnable conversion model to make the prior experience quickly adapt to the uniqueness of each event, fits the transmission function for each event, and captures the individual feature of each event. The prior experience ex is converted into an event (l,m,n,t) with specific parameters δ l,m,n,t as follows: where, The event-specific time representation defined as nodes l, m and n, with con() referring to the concatenation operation; based on The event prior ex is transformed into the event's uniqueness parameter δ by a learnable transformation model ψ l,m,n,t The process is parameterized by η; The learnable conversion model dynamically adjusts network parameters by feature linear modulation FiLM, that is, performing a simple feature affine transformation on the intermediate layer features of the neural network, scaling and moving the prior experience, and ensuring that the prior experience can quickly adapt to the uniqueness of each event; scaling and moving are performed by a fully connected layer FCL C Scaling factor scf of expression (l ,m,n,t) And moving factor shf (l,m,n) The implementation is as follows: where ω scf and ω shf are the learnable weight matrices of the FCL C , B scf and B shf are the bias vectors of the FCL C ; The transmission function is an important part of generating the condition strength, which is realized by using the softplus function.

9. The method of claim 1, wherein, In step S200: The non-paired relationship of the node in the decoupled spatiotemporal hypergraph captures the collective correlation features of all events, and ensures the systematicness and dynamic interaction features of the events; The collective correlation feature is the non-paired relationship of the node in the decoupled spatiotemporal hypergraph, which ensures the systematicness and dynamic interaction features of the events, represents the event features associated with the node by connecting the nodes in the spatiotemporal hypergraph, explores the time sequence propagation characteristics of the event features through the graph neural network as time elapses, and fuses the spatial features and time evolution of the event to realize the collective correlation feature capture of all events; All related nodes in the above assembly, operators, robots, operations, and component nodes are essentially a concept; For ease of discussion, here we define a new event as denotes a new event at time t for nodes l, m, and n; below, we use a fully connected layer FCL C A predictor of node new events is constructed: wherein, defined as event-specific time representation of nodes l, m and n, δ I denotes a fully connected layer FCL C parameters required.

10. The method of claim 1, wherein, In step S200: Define the loss function to optimize the network performance and ensure the consistency of the relationship transmission of all hyperedges and nodes on the spatiotemporal hypergraph, model the evolution of the time-varying non-paired relationship in the spatiotemporal hypergraph, and realize the reasoning of the human-machine collaborative cognitive strategy; The loss function includes optimizing the individual feature and collective correlation feature capture of the event, the individual feature capture optimization is defined by the negative log-likelihood loss, which can provide appropriate condition strength for the occurrence or non-occurrence of the event, and the collective correlation feature capture optimization is defined by the smooth L1 loss, which increases the error penalty to achieve superior network performance. The individual feature capture optimization, assuming that an event (l,m,n,t) e O has occurred, its loss L I (l,m,n,t) as follows: where λ l,m,n (t) and λ l,m,q (t) represent the conditional intensity of the superedge of ternary (l, m, n) and (l, m, q) at time t, respectively, q represents the distribution d n based on the sampled negative samples, which defines the instances of the event not occurring; meanwhile, N ne represents the number of negative samples for each occurring event; The collective correlation feature captures optimization, whose loss L C (l, m, n, t) as follows: wherein, N represents a predicted node propagation value, l,m,n Y represents a true value of a node.

Citation Information

Patent Citations

  • Personnel operation intention recognition method for man-machine cooperative assembly

    CN114445741A