Method and system for occasion cognition based on structured representation and causal reasoning
By adopting the method of structured representation and causal reasoning in robots, constructing a hypergraph of scene image sequences and performing causal reasoning, the problem of difficulty in effectively representing and reasoning complex scenes in existing technologies is solved, more accurate and reasonable scene cognition is achieved, and the adaptive behavior of robots in different scenes is supported.
Patent Information
- Application Number
- CN202510732694.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-03
- Publication Date
- 2025-09-19
AI Technical Summary
Existing technologies find it difficult to effectively represent and reason about the formation process of complex situations. Especially in robotic application scenarios, the requirements for situational cognition tasks are gradually evolving from single-factor recognition to multi-factor cognitive understanding, but research is still in its infancy.
A method based on structured representation and causal reasoning is adopted. The robot collects the scene image sequence, and the scene association element recognition algorithm is used to construct a hypergraph. The activity sequence causal reasoning model and the scene sequence causal reasoning model are combined to realize the structured representation and credible reasoning of the scene elements.
It improves the accuracy and rationality of situation cognition results, reduces the complexity of situation representation, can effectively identify and reason about multi-factor information in complex situations, and supports robots' reasonable judgment and behavioral adaptation in different situations.
Smart Images

Figure CN120673361A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of situation cognition of robots, and in particular to a situation cognition method and system based on structured representation and causal reasoning. Background Art
[0002] Occasion is a concept that integrates multiple factors, including time, location, participants, purpose of the event, and sociocultural context. It significantly influences and constrains people's behavior, language, and psychology. The complexity of occasions makes it difficult to characterize the formation of specific occasions using simple definitions. While technologies for identifying single elements such as location, people, and objects have been widely developed, the formation of an occasion is not only related to the characteristics of these individual elements but also closely integrated with time, relationships between people, cultural symbols, and social context. This complex coupled causal relationship makes it difficult to achieve limited, structured representation and reasoning.
[0003] However, the demand for situational awareness tasks in robotics applications is becoming increasingly important. Situational awareness encompasses more than just the recognition of the objective environment and location; it encompasses the understanding and comprehension of multiple elements, including specific times, locations, participants, activity objectives, rules, and functions. For example, meetings require quietness, gatherings require dynamic obstacle avoidance, workplaces require restricted access, and service environments require attention to the status of individuals. Situational awareness tasks are gradually evolving from identifying single elements of a scene (such as location and human behavior) to understanding multiple elements. However, research on using symbolic language to appropriately represent situations and on how to achieve cognitive reasoning based on key information about situations is still in its infancy. Summary of the Invention
[0004] Purpose of the invention: The purpose of the present invention is to provide a situation cognition method and system based on structured representation and causal reasoning, which deconstructs the situation element cognition through structured representation, solves the representation problem in the situation cognition process, and improves the accuracy and rationality of the situation cognition results through causal relationship analysis.
[0005] Technical Solution: To achieve the above objectives, the situation recognition method based on structured representation and causal reasoning described in the present invention includes the following steps:
[0006] S1: image sequence I of the scene where the robot collects data;
[0007] S2: Using the occasion-related element recognition algorithm to extract the occasion elements in the image sequence I to construct the hypergraph G;
[0008] S3: Use the activity sequence causal reasoning model to infer the activity sequence A under the trustworthy evaluation of the hypergraph G;
[0009] S4: Use the occasion sequence causal reasoning model to infer the occasion cognition sequence O under the credible evaluation of the hypergraph G and the activity sequence A, and obtain the occasion cognition result of the occasion in which the robot is located.
[0010] Preferably, the image sequence I is a collection of multiple frames of images captured by the robot from a first perspective during movement.
[0011] Preferably, the occasion elements include people, objects, places, relationship attributes, action attributes, cultural symbols, and social background.
[0012] Preferably, the method of extracting the occasion elements from the image sequence I using the occasion-related element recognition algorithm to construct the hypergraph G is:
[0013] S201, each frame of the image sequence I is sent to the visual encoder for encoding, and a visual feature vector is output; the text prompt is sent to the semantic decoder for feature extraction, and a semantic feature vector is output;
[0014] S202, sending the visual feature vector and the semantic feature vector into the multimodal attention block for information fusion, and outputting the fused multimodal feature vector;
[0015] S203. Send the multimodal feature vector to the scene element extraction module, output the identified scene elements and features through the multi-layer perceptron and classifier, and construct a hypergraph of the image sequence I by identifying the scene elements and features in each frame of the image.
[0016] Preferably, the hypergraph structured representation is:
[0017] G=(V,E R ,E B ,φ P ,φ U ,φ L ,φ C )
[0018] Among them, V = P∪U∪L∪C is the node set of the hypergraph, E R is the set of hyperedges, E B is the set of behavior hyperedges; P is the set of person nodes; U is the set of object nodes; L is the set of location nodes; C is the set of other nodes;
[0019] Define the feature mapping function of the node: character features: is the character feature space; object features: is the object feature space; place feature: is the place feature space; other features: For other feature spaces.
[0020] Preferably, the hypergraph G includes a relationship hyperedge e r ∈E R and behavioral hyperedge e b ∈E B , each hyperedge is a multi-tuple representation of the relationship and behavior, and the multi-tuple structure is expressed as:
[0021]
[0022] in, is the subset of nodes in the node set V that are related to the relationship hyperedge or behavior hyperedge; is the relationship type in the relationship hyperedge; is the behavior type in the behavior hyperedge, l is the location node where the relationship or behavior occurs, and t is the time when the relationship or behavior occurs.
[0023] Preferably, the method of using the activity sequence causal reasoning model to infer the activity sequence A under the trustworthy evaluation of the hypergraph G is:
[0024] The activity sequence causal reasoning model includes a hypergraph encoding module, a cross-order message passing module and a trusted causal inference module, wherein the hypergraph encoding module encodes the node features in the hypergraph into feature vectors and constructs an initial node feature matrix X(x i ∈X), and at the same time, the hyperedge feature encoding in the hypergraph is converted into the initial feature vector containing the high-order relationship of the nodes and the initial hyperedge feature matrix E(e j ∈E);
[0025] The cross-order message passing module, for each hyperedge j∈(E B ∪E R ), use the attention mechanism to aggregate the node features contained in the hyperedge to obtain the updated feature e of the hyperedge j j ′:
[0026]
[0027] Where E′ is the updated hyperedge feature matrix, α i,j is the attention coefficient of node i to hyperedge j;
[0028] The cross-order message passing module uses the node feature anti-aggregation mechanism to aggregate the representations of all hyperedges belonging to each node i, and obtains the updated feature x′ of node i i :
[0029]
[0030] Where X′ is the updated node feature matrix, β i,j is the attention coefficient of hyperedge j to node i;
[0031] The credible causal inference module includes a structural causal model and a confidence evaluation function, wherein the structural causal model concatenates the node feature matrix X′ and the hyperedge feature matrix E′ into a comprehensive feature vector z activity = [mean(X′)||mean(E′)], inferring the predefined activity category A based on the comprehensive feature vector candidate The probability distribution of :
[0032] P(A candidate )=softmax(linear(z))
[0033] Where, mean(·) represents the mean of each row vector of the matrix, linear(·) represents the linear fully connected layer operation, and softmax(·) represents the normalized activation function operation;
[0034] Trust evaluation function is based on activity category A candidate The probability distribution of each activity is used to evaluate the credibility value, and the activity a(a∈A candidate ) as the reasoning result of activity cognition, construct activity sequence A.
[0035] 9. Preferably, the method of using the occasion sequence causal reasoning model to infer the occasion cognition sequence O under the trustworthy evaluation of the hypergraph G and the activity sequence A is:
[0036] The occasion sequence causal reasoning model includes a hypergraph encoding module, an activity sequence encoding module, a cross-order message passing module and a causal inference module, wherein the hypergraph encoding module encodes the node features in the hypergraph into feature vectors and constructs an initial node feature matrix X(x i ∈X), and at the same time, the hyperedge encoding in the hypergraph is converted into the initial feature vector containing the high-order relationship of the nodes and the initial hyperedge feature matrix E(e j ∈E).
[0037] The activity sequence encoding module encodes the activities of the activity sequence A and the confidence corresponding to the activities into an activity feature matrix:
[0038] encoder(A)=Embedding(a)⊙MLP(c a )
[0039] Where Embedding(·) represents the encoding operation of the learnable embedding matrix, MLP(·) represents the encoding mapping operation of the multilayer perceptron, and ⊙ represents element-by-element multiplication.
[0040] The cross-order message transmission module realizes bidirectional information flow based on the hyperedge feature aggregation and node feature de-aggregation mechanism, and obtains the updated node feature matrix X′ and the hyperedge feature matrix E′;
[0041] The causal inference module includes a structural causal model and a confidence evaluation function, wherein the structural causal model concatenates the node feature matrix X′, the hyperedge feature matrix E′, and the activity feature matrix encoder (A) into a comprehensive feature vector z occasion = [mean(X′)||mean(E)||mean(encoder(A))], inferring the predefined occasion category O based on the comprehensive feature vector candidate The probability distribution of :
[0042] P(O candidate )=softmax(linear(z occasion ))
[0043] Where, mean(·) represents the mean of each row vector of the matrix, linear(·) represents the linear fully connected layer operation, and softmax(·) represents the normalized activation function operation;
[0044] Credibility evaluation function is based on the occasion category O candidate The probability distribution of each occasion is used to evaluate the credibility value, and the occasions o(o∈O candidate ) as the reasoning result of occasion cognition, construct the living occasion cognition sequence O.
[0045] The present invention provides a situation recognition system based on structured representation and causal reasoning, comprising:
[0046] Data acquisition module: The robot collects image sequences I of the scene;
[0047] Hypergraph construction module: uses the occasion-related element recognition algorithm to extract the occasion elements in the image sequence I to construct the hypergraph G;
[0048] Activity sequence reasoning module: uses the activity sequence causal reasoning model to infer the activity sequence A under the trustworthy evaluation of the hypergraph G;
[0049] Situation cognition reasoning module: Use the situation sequence causal reasoning model to infer the situation cognition sequence O under the credible evaluation of the hypergraph G and the activity sequence A, and obtain the situation cognition result of the situation where the robot is located.
[0050] Beneficial effects: 1. In the perceptual reasoning process of occasion cognition, a hypergraph-based structured representation method for occasion-related elements is designed, hypergraph nodes are constructed to represent occasion elements, and multi-tuple hyperedges are constructed to structuredly represent relationships and behavioral attributes. Hyperedges are used to infer credible activity attributes, and hypergraphs are used to achieve credible occasion cognition, solving the problem that traditional binary representation methods cannot accurately and intuitively represent multi-node features such as relationship attributes and behavioral attributes, reducing the complexity of occasion representation, and improving the logic of occasion cognitive reasoning; 2. A hypergraph neural network architecture considering cross-order message passing and causal relationship evaluation is designed, and the interference of occasion-irrelevant variables is reduced through cross-order message passing models. The confidence score is used to evaluate the credibility of occasion cognitive reasoning, thereby improving the reliability and rationality of occasion cognitive results; 3. Based on the occasion sequence of occasion cognitive reasoning results, it is possible to make reasonable judgments on the occasion of the robot's environment, judge the important information contained in the scene, reduce human-machine conflicts, and integrate people into the occasion. BRIEF DESCRIPTION OF THE DRAWINGS
[0051] Figure 1 It is a flow chart of the present invention;
[0052] Figure 2 This is a structural diagram of the associated element extraction network in this case;
[0053] Figure 3 This is an example graph for constructing a hypergraph. DETAILED DESCRIPTION
[0054] The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0055] like Figure 1 As shown, the specific steps of the occasion recognition method of the present invention are as follows:
[0056] S1: The robot collects an image sequence I of the scene based on its visual sensor, where the image sequence I is a collection of multiple frames of images captured by the robot from the first perspective during its movement in the scene, including but not limited to RGB images, depth images, infrared images, and other types.
[0057] The occasions where the robot is used include but are not limited to: daily activities at home, work, entertainment, and social occasions.
[0058] S2: Use the occasion-related element recognition algorithm to extract the occasion elements in the image sequence I to construct the hypergraph G.
[0059] Among them, occasion elements include but are not limited to: people, objects, places, relationship attributes, action attributes, cultural symbols, social background, etc.
[0060] Character characteristics include, but are not limited to, clothing characteristics (type, color), hair characteristics (length, color), expression, and posture (standing, sitting).
[0061] Objects include but are not limited to: tables, chairs, beds, cups, televisions, beds, clothes...; the characteristics of objects include but are not limited to: state (on, off, running, empty), color.
[0062] Places include but are not limited to: bedrooms, living rooms, kitchens, classrooms, lawns, offices...; characteristics of places include but are not limited to: purpose, space (crowded, spacious).
[0063] Relationship attributes include but are not limited to: object-to-object relationships (contains, adjacent to, above, below), person-to-person relationships (adjacent to, attached to, above, below, holds), and person-to-person relationships (adjacent to, facing, back-to-back).
[0064] Action attributes include but are not limited to: grabbing, pushing, throwing, washing, running, eating, and drinking.
[0065] Cultural symbols refer to some expressive features in an image that can express cultural characteristics, such as the red theme color of a Chinese wedding.
[0066] Social background refers to some expressive features in the image that can express social characteristics, such as the place names contained in street signs in neighborhoods and places, and the text descriptions of the design time and event themes in conference venues.
[0067] The context-related element recognition algorithm uses a transformer network architecture and can extract context-related elements using a multimodal attention mechanism based on prompt text and image sequences. Specifically, it comprises a visual encoder and semantic decoder, a multimodal attention module, and a context-related element extraction module. The text prompt is a natural language description of the image processing task, and it provides prompts for the identification of elements such as people, objects, places, relationships, and behaviors in the image that the context-related element extraction task focuses on.
[0068] The visual encoder consists of stacked convolutional blocks (CNN), residual blocks (ResNet), and self-attention blocks, which perform multi-level and multi-scale visual feature extraction on image sequences. The stacked convolutional blocks are used to extract multi-scale features of the image; the residual blocks are used to alleviate the vanishing gradient problem, allowing the network to learn deeper features; and the self-attention blocks are used to enhance the value of the global structure and feature relationships of the image, and output visual feature vectors to the semantic decoder.
[0069] The semantic decoder includes an embedding block, a positional encoding block, a multi-head self-attention block, a feedforward network, a multi-head cross-attention block, and a feedforward network to extract semantic features from the prompt text. The embedding block converts the text into a word vector representation, capturing semantic information; the positional encoding adds position information to the word vector; the multi-head self-attention block performs calculations within the semantic features to mine semantic information associations, providing rich semantic features for the multi-head cross-attention block; the multi-head cross-attention block performs cross-attention calculations, "querying" image information through the semantic representation of the text, achieving a fusion of image and semantic information; and the feedforward network performs nonlinear transformations on the features output by the previous processing layer.
[0070] The multimodal attention module includes a visual self-attention block, a semantic self-attention block, and a cross-attention block. The visual self-attention block performs self-attention on the visual feature vector to highlight key information in the visual features; the semantic self-attention block performs self-attention on the semantic feature vector to enhance the expression of semantic information; and the cross-attention block performs cross-attention calculations on the outputs of the visual and semantic self-attention blocks to achieve a deep fusion of visual and semantic information.
[0071] The occasion factor extraction module includes a multi-layer perceptron (MLP) and a classifier, which identifies factors related to the "occasion" indicated by the prompt text from the deeply integrated and well-defined features. The MLP consists of a stack of fully connected layers, batch normalization layers, and activation function layers to map features to the appropriate dimensions. The classifier classifies the features and outputs different probabilities of occasion factors.
[0072] like Figure 2 As shown in FIG, the process of extracting occasion elements from the image sequence I using the occasion-related element recognition algorithm is as follows:
[0073] S201. Send each frame of the image sequence I to the visual encoder for encoding and output a visual feature vector; send the text prompt to the semantic decoder for feature extraction and output a semantic feature vector.
[0074] S202: Send the visual feature vector and the semantic feature vector to the multimodal attention block to calculate self-attention and cross-attention, fuse the visual information and semantic information, and output the fused multimodal feature vector.
[0075] S203: Send the fused multimodal feature vector to the occasion factor extraction module, and output the recognized occasion factors and features through the multi-layer perceptron and classifier.
[0076] By identifying the scene elements and features in each frame of the image, a hypergraph of the image sequence I is constructed. The mathematical expression of the hypergraph structured representation is:
[0077] G=(V,E R , E B ,φ P ,φ U ,φ L ,φ C )
[0078] Among them, V = P∪U∪L∪C is the node set of the hypergraph, E R is the set of hyperedges, E B is the set of behavior hyperedges; P is the set of character nodes, p i ∈P; U is the object node set, u i ∈U; L is the set of place nodes, l k ∈L; C is a set of other nodes (including cultural symbols, social background, etc.), c m ∈C.
[0079] Each node represents an occasion element and has characteristic attributes, which define the node's feature mapping function: Character characteristics: ( is the character feature space); object features: Location Features: Other features:
[0080] Hyperedge e of a relation in a hypergraph r ∈E R and behavioral hyperedge e b ∈E B ,relationship hyperedges are used to describe the relationship attributes between nodes, and behavior hyperedges are used to describe the behavior attributes between nodes. Each hyperedge is in the form of a tuple to fully represent the relationship and behavior. The tuple structure is expressed as:
[0081]
[0082] in, is the subset of nodes in the node set V that are related to the relationship hyperedge or behavior hyperedge; The type of relationship in the relationship hyperedge; The behavior type in the behavior hyperedge, l: the location node where the relationship or behavior occurs, t: the time when the relationship or behavior occurs.
[0083] S3: Use the activity sequence causal reasoning model to infer the activity sequence A under the corresponding occasion credibility evaluation in the hypergraph G;
[0084] The activity sequence causal reasoning model is a hypergraph neural network architecture that includes a hypergraph encoding module, a cross-order message passing module, and a trusted causal inference module.
[0085] The hypergraph encoding module is used to encode the node features in the hypergraph into feature vectors and construct the initial node feature matrix X(x i ∈X), and convert the multi-tuple hyperedge encoding into the initial feature vector containing the high-order relationship between nodes and construct the initial hyperedge feature matrix E(e j ∈E), this module uses the hyperedge-node association matrix to determine the affiliation between nodes and hyperedges.
[0086] For example: Let the hyperedge-node association matrix of the hypergraph be H, where H i,j =1 means node i belongs to hyperedge j, otherwise H i,j =0, which can be used to represent high-order relationships between nodes during hyperedge encoding.
[0087] The cross-level message transmission module is based on the hyperedge feature aggregation and node feature anti-aggregation mechanism to achieve bidirectional information flow. The hyperedge feature aggregation mechanism is: for each hyperedge j∈(E B ∪E R ), use the attention mechanism to aggregate the node features contained in the hyperedge, and thus obtain the updated feature representation e of the hyperedge j j ';
[0088]
[0089] Where E is the updated hyperedge feature matrix, a i,j is the attention coefficient of node i to hyperedge j:
[0090]
[0091] Where W is the learnable weight matrix, λ is the learnable attention vector, and x i is the initial feature of node feature i, y j is the initial feature of hyperedge j, and || represents the vector concatenation operation.
[0092] The node feature anti-aggregation mechanism aggregates the representations of all hyperedges belonging to each node i to update the node feature x′ i :
[0093]
[0094] Where X′ is the updated node feature matrix, β i,j is the attention coefficient of hyperedge j to node i;
[0095]
[0096] Among them, W′ is the learnable weight matrix and λ′ is the learnable attention vector.
[0097] The credible causal inference module infers the activity sequence A based on the structural causal model (SCM) and the confidence credible evaluation function.
[0098] The structural causal model (SCM) concatenates the node feature matrix X′ and the hyperedge feature matrix E into a comprehensive feature vector z = [mean(X′)||mean(E)], and infers the potential activity category A through the fully connected layer of the structural equation neural network. candidate Probability distribution of (predefined activity categories):
[0099] P(A candidate )=softmax(linear(z))
[0100] Where mean(·) means finding the mean of each row vector of the matrix, linear(·) means the linear fully connected layer operation, and softmax(·) means the normalized activation function operation.
[0101] Credibility evaluation function for activity category A candidate The object in the evaluation credibility value, for example, for A candidate An activity a in , its confidence score c a It can be defined as:
[0102] c a =P(A candidate =a)×conf(a)
[0103] Where conf(a) is the confidence of the neural network output of the trust evaluation function, which is obtained by machine learning. Only activities a(a∈A candidate ) as the reasoning result of activity cognition, construct activity sequence A.
[0104] S4: Taking the hypergraph G constructed in S2 and the activity cognition results in the activity sequence A obtained in S3 as input, the robot infers the occasion cognition sequence O under credible evaluation based on the occasion sequence causal reasoning model.
[0105] The occasion sequence causal reasoning model takes a hypergraph G and an activity sequence A as input and outputs an occasion cognition sequence O under credible evaluation. Its architecture includes a hypergraph neural network architecture consisting of a hypergraph encoding module, an activity sequence encoding module, a cross-order message passing module, and a causal inference module.
[0106] The hypergraph encoding module is used to encode the node features in the hypergraph into feature vectors and construct the initial node feature matrix X(x i∈X), and convert the multi-tuple hyperedge encoding into the initial feature vector containing the high-order relationship between nodes and construct the initial hyperedge feature matrix E(e j ∈E).
[0107] The activity sequence encoding module encodes the activities and confidences of the activity sequence into an activity feature matrix:
[0108] encoder(A)=Embedding(a)⊙MLP(ca)
[0109] Where Embedding(·) represents the encoding operation of the learnable embedding matrix, MLP(·) represents the encoding mapping operation of the multilayer perceptron, and ⊙ represents element-by-element multiplication.
[0110] The cross-order message passing module, similar to S3, implements bidirectional information flow based on hyperedge feature aggregation and node feature deaggregation. The updated node feature matrix is X′, and the hyperedge feature matrix is E′.
[0111] The credible causal inference module is based on the structural causal model (SCM) and the confidence evaluation function to infer the occasion sequence O. The structural causal model (SCM) concatenates the input data node feature matrix X′, the hyperedge feature matrix E′, and the activity feature matrix encoder (A) to obtain the comprehensive feature vector: occasion = [mean(X′)||mean(E)||mean(encoder(A))], inferring the predefined occasion category O based on the comprehensive feature vector candidate The probability distribution of :
[0112] P(O candidate )=softmax(linear(z occasion ))
[0113] Where, mean(·) represents the mean of each row vector of the matrix, linear(·) represents the linear fully connected layer operation, and softmax(·) represents the normalized activation function operation;
[0114] Credibility evaluation function is based on the occasion category O candidate The probability distribution of each occasion is used to evaluate the credibility value, and the occasions with credibility scores higher than the threshold τ′ are retained (o∈0 candidate ) as the reasoning result of occasion cognition, construct the living occasion cognition sequence O.
[0115] like Figure 3 As shown in Figure 2, taking a typical family dinner as an example, the process of inferring the occasion sequence is as follows:
[0116] S1. The robot collects image sequence I in real time in the scene, with a total of 4 frames of images {T1, T2, T3, T4}.
[0117] S2. The robot feeds the image sequence I and the text prompt into the occasion element extraction algorithm to extract occasion-related elements.
[0118] The text prompt is designed as follows: {Please identify instances of elements such as people, objects, and places in the image, detect the relationship features and action features between these elements, and classify them according to predefined categories.}
[0119] The scene element extraction results in the {T1, T2, T3, T4} images include kid (Black hair, gray top, standing), drinking (personl, drink bootles), etc., which contain elements such as person instances, object instances, place instances, relationship attributes, and action attributes.
[0120] Based on the results of the occasion factor extraction, a hypergraph is constructed according to the structure of the hypergraph:
[0121] The nodes in the hypergraph include: <kid> , <person1> , <person2>,<drink bottles>, <dishes>,<dining table>, <room>Nodes. <kid> , <person1> , <person2>Belongs to the character node subset;<drink bottles> , <dishes>,<dining table> Belongs to the object node subset; <room>Belongs to the venue node subset;
[0122] The hyperedge set contains four hyperedges, among which:
[0123] {( <kid> , <person1> , <person2> ) <surrounding>(<dining table> )<room><T1,T2,T3,T4>} belongs to the relational hyperedge;
[0124] {( <personl> , <person2>)<clinking glasses><drink bottles> <room><T1,T2>} belongs to behavioral hyperedge;
[0125] { <person2><talking with> <kid> <room><T2,T3>} belongs to behavioral hyperedge;
[0126] { <person1> <drinking><drink bottles> <room><T3,T4>} belongs to behavioral hyperedge.
[0127] S3 takes the hypergraph constructed in S2 as input and feeds it into the causal reasoning model to infer the activity cognitive reasoning results under the credible evaluation in the scenario, the activity sequence:
[0128] {Activities: {{Name: Toasting}, {Confidence: 0.85}, {{Name: Have a mealtogether}, {Confidence: 0.9}}}
[0129] Furthermore, according to S4, the hypergraph constructed by S2 and the activity sequence obtained by reasoning in S3 are fed into the causal reasoning model to infer the cognitive reasoning results of the occasion under the credible evaluation in the occasion. The occasion sequence is:
[0130] {Occasions: {{Name: Family gathering}, {Confidence: 0.85}, {{Name: Familyreunion dinner}, {Confidence: 0.85}}}
[0131] This paper defines the robot's situational cognition process as the dynamic perception, reasoning, and cognitive process of the situation when the robot is performing tasks in a human-like social setting. To address the difficulty of representing situations, this paper employs information processing methods from cognitive psychology to deconstruct the situational cognition process into a structured cognitive formation process based on individual characteristics, situational elements, activity attributes, and situation type. This process is then structured using a hypergraph, and cognitive reasoning is performed in stages using causal analysis, reducing the complexity of situational cognition.
[0132] The situation sequence based on situational cognitive reasoning can enable reasonable judgment of the situation in the robot's environment and determine the important information contained in the scene. For example, people gathered around a dining table, enjoying a meal, indicating a family gathering. This provides reasonable adaptive guidance for the robot's subsequent tasks in the environment. For example, in robot navigation tasks, it can assist in setting appropriate movement speeds, restricted areas, and safe distances. In addition, in robot service tasks in conference settings, reasonable situational cognitive results can help the robot reasonably avoid noisy behaviors that disrupt the atmosphere of thinking and discussion in the crowd, provide more gentle and comfortable service, reduce human-machine conflicts, and integrate into the human environment.< / room> < / drinking> < / person1> < / room> < / kid> < / room> < / personl> < / surrounding> < / person2> < / person1> < / kid> < / room> < / dishes> < / person1> < / kid> < / room> < / dishes> < / person1> < / kid>
Claims
1. A situation recognition method based on structured representation and causal reasoning, characterized by: include: S1: image sequence I of the scene where the robot collects data; S2: Using the occasion-related element recognition algorithm to extract the occasion elements in the image sequence I to construct the hypergraph G; S3: Use the activity sequence causal reasoning model to infer the activity sequence A under the trustworthy evaluation of the hypergraph G; S4: Use the occasion sequence causal reasoning model to infer the occasion cognition sequence O under the credible evaluation of the hypergraph G and the activity sequence A, and obtain the occasion cognition result of the occasion in which the robot is located.
2. The situation recognition method according to claim 1, characterized in that: The image sequence I is a collection of multiple frames of images captured by the robot from the first perspective during its movement.
3. The situation recognition method according to claim 1, characterized in that: The occasion elements include people, objects, places, relationship attributes, action attributes, cultural symbols, and social background.
4. The situation recognition method according to claim 1, characterized in that: The method of extracting occasion elements from the image sequence I using the occasion-related element recognition algorithm to construct the hypergraph G is as follows: S201, sending each frame of the image sequence I to a visual encoder for encoding, and outputting a visual feature vector; Send the text prompt to the semantic decoder for feature extraction and output the semantic feature vector; S202, sending the visual feature vector and the semantic feature vector into the multimodal attention block for information fusion, and outputting the fused multimodal feature vector; S203. Send the multimodal feature vector to the scene element extraction module, output the identified scene elements and features through the multi-layer perceptron and classifier, and construct a hypergraph of the image sequence I by identifying the scene elements and features in each frame of the image.
5. The situation recognition method according to claim 4, characterized in that: The hypergraph structure is characterized as follows: G=(V,E R ,E B ,f P ,f U ,f L ,f C ) Among them, V = P∪U∪L∪C is the node set of the hypergraph, E R is the set of hyperedges, E B is the set of behavior hyperedges; P is the set of person nodes; U is the set of object nodes; L is the set of location nodes; C is the set of other nodes; Define the feature mapping function of the node: character features: is the character feature space; object features: is the object feature space; place feature: is the place feature space; other features: For other feature spaces.
6. The situation recognition method according to claim 5, characterized in that: The hypergraph G includes the relationship hyperedge e r ∈E R and behavioral hyperedge e b ∈E B , each hyperedge is a multi-tuple representation of the relationship and behavior, and the multi-tuple structure is expressed as: in, is the subset of nodes in the node set V that are related to the relationship hyperedge or behavior hyperedge; is the relationship type in the relationship hyperedge; is the behavior type in the behavior hyperedge, l is the location node where the relationship or behavior occurs, and t is the time when the relationship or behavior occurs.
7. The situation recognition method according to claim 1, characterized in that: The method of using the activity sequence causal reasoning model to infer the activity sequence A under the trustworthy evaluation of the hypergraph G is: The activity sequence causal reasoning model includes a hypergraph encoding module, a cross-order message passing module and a trusted causal inference module, wherein the hypergraph encoding module encodes the node features in the hypergraph into feature vectors and constructs an initial node feature matrix X(x i ∈X), and at the same time, the hyperedge feature encoding in the hypergraph is converted into the initial feature vector containing the high-order relationship of the nodes and the initial hyperedge feature matrix E(e j ∈E); The cross-order message passing module, for each hyperedge j∈(E B ∪E R ), use the attention mechanism to aggregate the node features contained in the hyperedge to obtain the updated feature e of the hyperedge j j ′: Where E′ is the updated hyperedge feature matrix, α i,j is the attention coefficient of node i to hyperedge j; The cross-order message passing module uses the node feature anti-aggregation mechanism to aggregate the representations of all hyperedges belonging to each node i, and obtains the updated feature x′ of node i i : Where X′ is the updated node feature matrix, β i,j is the attention coefficient of hyperedge j to node i; The credible causal inference module includes a structural causal model and a confidence evaluation function, wherein the structural causal model concatenates the node feature matrix X′ and the hyperedge feature matrix E′ into a comprehensive feature vector z activity = [mean(X′)||mean(E′)], inferring the predefined activity category A based on the comprehensive feature vector candidate The probability distribution of : P(A candidate )=soft max(linear(z)) Where, mean(·) represents the mean of each row vector of the matrix, linear(·) represents the linear fully connected layer operation, and softmax(·) represents the normalized activation function operation; Trust evaluation function is based on activity category A candidate The probability distribution of each activity is used to evaluate the credibility value, and the activity a(a∈A candidate ) as the reasoning result of activity cognition, construct activity sequence A.
8. The situation recognition method according to claim 1, characterized in that: The method of using the occasion sequence causal reasoning model to infer the occasion cognitive sequence O under the credible evaluation of the hypergraph G and the activity sequence A is: The occasion sequence causal reasoning model includes a hypergraph encoding module, an activity sequence encoding module, a cross-order message passing module and a causal inference module, wherein the hypergraph encoding module encodes the node features in the hypergraph into feature vectors and constructs an initial node feature matrix X(x i ∈X), and at the same time, the hyperedge encoding in the hypergraph is converted into the initial feature vector containing the high-order relationship of the nodes and the initial hyperedge feature matrix E(e j ∈E). The activity sequence encoding module encodes the activities of the activity sequence A and the confidence corresponding to the activities into an activity feature matrix: encoder(A)=Embedding(a)⊙MLP(c a ) Where Embedding(·) represents the encoding operation of the learnable embedding matrix, MLP(·) represents the encoding mapping operation of the multilayer perceptron, and ⊙ represents element-by-element multiplication. The cross-order message transmission module realizes bidirectional information flow based on the hyperedge feature aggregation and node feature de-aggregation mechanism, and obtains the updated node feature matrix X′ and the hyperedge feature matrix E′; The causal inference module includes a structural causal model and a confidence evaluation function, wherein the structural causal model concatenates the node feature matrix X′, the hyperedge feature matrix E′, and the activity feature matrix encoder (A) into a comprehensive feature vector z occasion = [mean(X′)||mean(E)||mean(encoder(A))], inferring the predefined occasion category O based on the comprehensive feature vector candidate The probability distribution of : P(O candidate )=softmax(linear(z occasion )) Where, mean(·) represents the mean of each row vector of the matrix, linear(·) represents the linear fully connected layer operation, and softmax(·) represents the normalized activation function operation; Credibility evaluation function is based on the occasion category O candidate The probability distribution of each occasion is used to evaluate the credibility value, and the occasions o(o∈O candidate ) as the reasoning result of occasion cognition, construct the living occasion cognition sequence O.
9. A situation recognition system based on structured representation and causal reasoning, characterized by: include: Data acquisition module: The robot collects image sequences I of the scene; Hypergraph construction module: uses the occasion-related element recognition algorithm to extract the occasion elements in the image sequence I to construct the hypergraph G; Activity sequence reasoning module: uses the activity sequence causal reasoning model to infer the activity sequence A under the trustworthy evaluation of the hypergraph G; Situation cognition reasoning module: Use the situation sequence causal reasoning model to infer the situation cognition sequence O under the credible evaluation of the hypergraph G and the activity sequence A, and obtain the situation cognition result of the situation where the robot is located.