Body intelligence layered memory perception method, system and electronic device

By employing a hierarchical memory perception method, the problem of poor environmental adaptability caused by the single-layer architecture in embodied intelligence systems is solved. This method achieves decoupling between millisecond-level response and task-level decision-making, thereby improving the system's dynamic environmental adaptability and communication continuity.

CN120724133BActive Publication Date: 2025-11-07NANJING UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511212077.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-28
Publication Date
2025-11-07
Estimated Expiration
2045-08-28

AI Technical Summary

Technical Problem

In existing embodied intelligence systems, the single-layer perception and decision-making architecture leads to competition and conflict in computing resources between millisecond-level response and task-level decision-making, resulting in poor environmental adaptability. This is particularly evident in industrial inspection, underwater exploration, and aerial monitoring, where issues such as equipment failure response delays and cross-task knowledge transfer interruptions occur.

Method used

A hierarchical memory perception method is adopted. Environmental data is collected through the perception processing module. Feature extraction and sensor scheduling are performed in the short-term perception layer, dynamic environmental map is constructed in the medium-term context layer, and event memory chain is constructed and cross-task causal transfer is performed in the long-term cognition layer to generate target instructions. The communication interaction module is used to transmit instructions and receive feedback data, thereby decoupling response and decision-making.

Benefits of technology

It improves adaptability to dynamic environments, reduces energy consumption of redundant data, enhances adaptive capabilities, supports accurate environmental situation estimation and task planning, and ensures adaptive switching of communication links and continuity of commands.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120724133B_ABST
    Figure CN120724133B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of embodied intelligence, and provides an embodied intelligent layered memory perception method, system and electronic equipment. The method collects environment data through a perception processing module, performs feature extraction and sensor scheduling at a short-time perception layer, outputs weighted fusion features, transmits the weighted fusion features to a cognitive decision module through a communication interaction module, the cognitive decision module receives the weighted fusion features, constructs a dynamic environment graph at a medium-time context layer, constructs an event memory chain and performs cross-task causal migration at a long-time cognitive layer, and outputs a long-term cognitive feature strategy. A target instruction is generated based on the weighted fusion features, the dynamic environment graph and the long-term cognitive feature through a control module, and the target instruction is transmitted to a target mechanism through the communication interaction module. The method decouples response and decision through a layered architecture to improve dynamic environment adaptability.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of embodied intelligence, and in particular to an embodied intelligence layered memory perception method and system and electronic device. BACKGROUND

[0002] Embodied intelligence is a technical system that realizes environmental adaptation, autonomous decision-making, and complex task execution, and can be applied to industrial detection equipment, ocean exploration platforms, and aerial monitoring systems. In the industrial field, equipment detection needs to be performed in dangerous environments. Ocean exploration relies on underwater autonomous platforms to complete resource exploration. Aerial monitoring requires continuous tracking of airspace targets. All three fields have cross-scale collaboration needs ranging from millisecond-level to task-level decision-making.

[0003] For example, a single-layer perception and decision-making architecture is used to implement multi-modal data processing and task control. An industrial detection system synchronously processes visual sensor data and equipment maintenance strategies through a centralized control module. An underwater submersible integrates sonar point clouds and historical exploration knowledge in a single computing unit. A smart office system couples real-time document parsing and project long-term planning based on a static rule engine. A control hub enhances processing power by increasing processor clock speed or adding parallel computing units.

[0004] The single-layer architecture leads to a competitive conflict between millisecond-level responses and task-level decisions in terms of computing resources. In the industrial field, the response delay to equipment failures disrupts cross-task knowledge transfer in underwater exploration. The mutual interference between aerial target tracking and airspace planning results in poor environmental adaptability. SUMMARY

[0005] The present application provides an embodied intelligence layered memory perception method and system and electronic device to solve the problem of poor environmental adaptability.

[0006] In a first aspect, the present application provides an embodied intelligence layered memory perception method, comprising:

[0007] Collecting environmental data through a perception processing module, performing feature extraction and sensor scheduling in a short-time perception layer to output weighted fusion features;

[0008] Transmitting the weighted fusion features to a cognitive decision-making module through a communication interaction module;

[0009] Receiving the weighted fusion features through a cognitive decision-making module, constructing a dynamic environment map in a medium-time context layer, and constructing an event memory chain and performing cross-task causal transfer in a long-time cognitive layer to output long-term cognitive features;

[0010] Generating target instructions based on the weighted fusion features, dynamic environment map, and long-term cognitive features through a control module;

[0011] transmit the target instruction to the target mechanism through the communication interaction module, and receive feedback data sent by the target mechanism to the long-term cognitive layer.

[0012] The hierarchical architecture decouples millisecond-level response from task-level decision, improving dynamic environmental adaptability.

[0013] In some feasible embodiments, the feature extraction and sensor scheduling performed at the short-term perception layer include:

[0014] Preprocessing is performed on the environmental data to generate preprocessed data, the environmental data being data collected by a multi-modal sensor array;

[0015] Feature extraction is performed on the preprocessed data by a convolutional neural network to output local dynamic features;

[0016] The local dynamic features are aggregated based on a time decay function to generate short-term dynamic information, wherein the decay rate of the time decay function is adjusted based on the medium-term context layer;

[0017] Sensor activation weights are calculated, and weighted fusion is performed on the short-term dynamic information to output the weighted fusion features.

[0018] The sensor activation weights are dynamically adjusted to reduce redundant data energy consumption and enhance the adaptive ability of short-term features to physical disturbances.

[0019] In some feasible embodiments, the construction of the dynamic environment graph at the medium-term context layer includes:

[0020] Spatial clustering is performed on the weighted fusion features to generate a node set;

[0021] The state vectors of the nodes in the node set are iteratively updated by a graph neural network;

[0022] Symbolic physical rules are injected into the graph construction system to calculate environmental disturbance coefficients;

[0023] According to the state vectors and environmental disturbance coefficients, the graph edge weights in the graph construction system are adjusted to generate a dynamic environment graph.

[0024] The graph neural network and the symbolic physical rules are fused to construct a dynamic environment graph, improving the robustness of environmental state representation.

[0025] In some feasible embodiments, the construction of the event memory chain at the long-term cognitive layer includes:

[0026] The dynamic environment graph is parsed to output entity relationships and pre-stored causal graphs;

[0027] generate a time-stamped potential event representation based on the entity relationship and a pre-stored causal graph;

[0028] link the potential event representation in chronological order to construct an event memory chain storing event and timing mapping relationship.

[0029] Organize the event memory chain to store the spatiotemporal causal relationship, and provide a structured knowledge base for cross-task migration.

[0030] In some feasible embodiments, the cross-task causal migration performed by the long-term cognitive layer includes:

[0031] Calculate the intra-task event causal relationship matrix through event co-occurrence probability and conditional probability;

[0032] Calculate the semantic similarity between the source task event and the target task event in the event memory chain to generate a target matching event;

[0033] Determine the dependence coefficient of the source task event on the target event according to the causal relationship matrix;

[0034] Based on the dependence coefficient and the target matching event, weight the feature embedding of the source task event and output the migration representation of the target event;

[0035] Output the long-term cognitive feature based on the migration representation of the target event.

[0036] Cross-task causal migration is achieved through semantic similarity and event contribution weight, which improves the generalization performance of the target task.

[0037] In some feasible embodiments, the generation of the target instruction includes:

[0038] Integrate the weighted fusion features of the short-term perception layer, the dynamic environment graph of the medium-term context layer, and the long-term cognitive feature of the long-term cognitive layer;

[0039] Correlate the weighted fusion features and the dynamic environment graph through a cross-layer attention mechanism to generate spatiotemporal correlation features;

[0040] Fuse the long-term cognitive feature and the migration representation through a long-term memory activation function to generate knowledge-enhanced features;

[0041] Decode the fusion result of the spatiotemporal correlation features and the knowledge-enhanced features through a decoder to output a unified perception state vector;

[0042] Generate the target instruction based on the unified perception state vector.

[0043] Integrate multi-layer features and decode them into a unified perception state to support accurate environment situation estimation and task planning.

[0044] In some possible embodiments, the transmitting the target instruction to the target mechanism by the communication interaction module comprises:

[0045] The transmission channel is established through a water acoustic communication link, a laser communication link, and a near-field optical communication link.

[0046] The water acoustic communication link, the laser communication link, and the near-field optical communication link are switched according to a quality signal-to-noise ratio of the communication link.

[0047] In the case of interruption of the communication link, the target instruction is stored by using a cache management unit.

[0048] In the case of resuming communication, the target instruction is retransmitted to the target mechanism.

[0049] The communication link is adaptively switched, and the cache management is used to guarantee the continuity of instruction transmission in a weak network.

[0050] In a second aspect, the present application provides a somatic intelligent layered memory perception system, comprising:

[0051] The perception processing module is configured to collect environment data, perform feature extraction and sensor scheduling at a short-time perception layer, and output weighted fusion features.

[0052] The cognitive decision module is configured to receive the weighted fusion features, construct a dynamic environment graph at a medium-time context layer, and construct an event memory chain and perform cross-task causal migration at a long-time cognitive layer, and output long-term cognitive features.

[0053] The control module is configured to generate a target instruction based on the weighted fusion features, the dynamic environment graph, and the long-term cognitive features.

[0054] The communication interaction module is configured to transmit the weighted fusion features to the cognitive decision module, transmit the target instruction to a target mechanism, and receive feedback data sent by the target mechanism to the long-time cognitive layer.

[0055] In some possible embodiments, the system further comprises an energy management module configured to:

[0056] Obtain power consumption of a sensor array and the cognitive decision module.

[0057] Based on the power consumption, allocate power supply priorities of the sensor array and the cognitive decision module according to a task execution stage.

[0058] The power supply priorities are dynamically allocated based on the task stage, and the overall energy efficiency of the system is optimized.

[0059] In a third aspect, the present application provides an electronic device, comprising: a processor, and a memory connected with the processor in communication;

[0060] The memory stores computer-executable instructions.

[0061] The processor executes the computer-executable instructions stored in the memory to implement the embodied intelligent hierarchical memory perception method.

[0062] From the above technical solutions, the present application provides an embodied intelligent hierarchical memory perception method, system and electronic device. The method comprises: collecting environmental data by a perception processing module, performing feature extraction and sensor scheduling at a short-time perception layer to output weighted fusion features; transmitting the weighted fusion features to a cognitive decision module by a communication interaction module; receiving the weighted fusion features by the cognitive decision module, constructing a dynamic environment map at a medium-time context layer, and constructing an event memory chain and performing cross-task causal migration at a long-time cognitive layer to output long-term cognitive features; generating target instructions based on the weighted fusion features, the dynamic environment map and the long-term cognitive features by a control module; transmitting the target instructions to a target mechanism by the communication interaction module, and receiving feedback data sent by the target mechanism to the long-time cognitive layer. The method decouples response and decision through a hierarchical architecture to improve dynamic environment adaptability. BRIEF DESCRIPTION OF DRAWINGS

[0063] In order to more clearly illustrate the technical solutions of the present application, the drawings needed in the embodiments will be briefly introduced below. Obviously, other drawings can also be obtained by those skilled in the art without creative labor on the basis of these drawings.

[0064] Figure 1 The flowchart of the embodied intelligent hierarchical memory perception method provided by the embodiments of the present application;

[0065] Figure 2 The flowchart of the short-time perception layer work provided by the embodiments of the present application;

[0066] Figure 3 The flowchart of the medium-time context layer work provided by the embodiments of the present application;

[0067] Figure 4 The flowchart of constructing an event memory chain provided by the embodiments of the present application;

[0068] Figure 5 The flowchart of performing cross-task causal migration provided by the embodiments of the present application;

[0069] Figure 6 The flowchart of generating target instructions provided by the embodiments of the present application;

[0070] Figure 7 A hierarchical architecture mapping schematic diagram provided for Embodiment 2 of the present application;

[0071] Figure 8 An event chain construction flowchart provided for Embodiment 2 of the present application;

[0072] Figure 9 A process schematic diagram for analyzing causal relationships in a historical event chain provided for Embodiment 2 of the present application;

[0073] Figure 10 A migration historical experience flowchart provided for Embodiment 2 of the present application;

[0074] Figure 11 A hierarchical architecture mapping schematic diagram provided for Embodiment 3 of the present application;

[0075] Figure 12 A schematic diagram of a body-aware intelligent hierarchical memory perception system structure provided for the present application. DETAILED DESCRIPTION

[0076] The embodiments will be described in detail below with reference to the drawings. When the following description refers to the drawings, the same numbers in different drawings represent the same or similar elements unless otherwise indicated. The embodiments described in the following examples do not represent all the embodiments consistent with the present application.

[0077] As shown in Figure 1 , some embodiments of the present application provide a body-aware intelligent hierarchical memory perception method, comprising the following steps:

[0078] S100: Collecting environment data through a perception processing module, performing feature extraction and sensor scheduling in a short-time perception layer to output weighted fusion features.

[0079] The perception processing module is used to collect environment data. In some implementations, the environment data includes sonar point cloud, optical image, water flow vector, and other multi-modal sensor information. In another implementation, the environment data includes task data, time node, and other data. That is, the method provided in the present embodiment can be applied to different environments and different environment data can be obtained based on different application environments.

[0080] The short-time perception layer is used to perform feature extraction and sensor scheduling. Feature extraction refers to the process of separating key information patterns from environment data, and sensor scheduling refers to a mechanism for dynamically configuring sensor working states according to task requirements.

[0081] After performing feature extraction and sensor scheduling, the weighted fusion features are generated by fusing the multi-modal data after weight distribution, and the standardized environment observation values, i.e., perception data, are output.

[0082] The perception processing module collects environmental data through a sensor array, performs feature extraction in a short-time perception layer, converts raw data into machine-processable digital signals, synchronously performs sensor scheduling, adjusts the collection frequency of each sensor according to the complexity of the environment, and outputs weighted fusion features and standardized perception data.

[0083] S200: transmitting the weighted fusion features to the cognitive decision module through the communication interaction module.

[0084] The cognitive decision module is used for receiving the weighted fusion features and performing layered decision-making, and the communication interaction module establishes a transmission channel. For example, the weighted fusion features are sent through a water acoustic communication link, when the quality of the water acoustic channel is detected to decrease, the communication link is automatically switched to a laser communication link or a near-field optical communication link, and if the communication link is interrupted, a buffer queue is enabled to temporarily store data, and the transmission is completed after the link is restored.

[0085] In some embodiments, the communication interaction module establishes a transmission channel through a water acoustic communication link, a laser communication link, and a near-field optical communication link. The water acoustic communication link is a communication channel that uses sound waves to transmit data in water medium and is suitable for medium and long distance underwater information interaction. The laser communication link is a wireless optical communication channel that carries information through a laser beam and has high bandwidth and low delay characteristics. The near-field optical communication link is a communication channel based on short-distance visible light or infrared light and is suitable for near-distance anti-interference transmission. The transmission channel is a redundant communication path set composed of the water acoustic communication link, the laser communication link, and the near-field optical communication link.

[0086] According to the quality signal-to-noise ratio of the communication link, the water acoustic communication link, the laser communication link, and the near-field optical communication link are switched. After establishing the transmission channel composed of the water acoustic communication link, the laser communication link, and the near-field optical communication link, the quality signal-to-noise ratio of each link is monitored in real time. The quality signal-to-noise ratio is an index for quantifying the stability of the communication link and is calculated by the ratio of the received signal strength to the noise power.

[0087] The quality signal-to-noise ratio is dynamically calculated by the ratio of the received signal power to the background noise power. For example, when the quality signal-to-noise ratio of the water acoustic communication link is lower than a first threshold, a link switching judgment process is started.

[0088] In the case of communication link interruption, the target instruction is stored by using a buffer management unit. The buffer management unit is a temporary register module for storing instruction data to be transmitted, including a first-in-first-out queue and an error checking mechanism. In the case of restoring communication, the target instruction is retransmitted to the target mechanism.

[0089] For example, the transmission channel is switched according to the link quality state. If the quality signal-to-noise ratio of the laser communication link is higher than the second threshold, the transmission channel is switched to the laser communication link. If the laser communication link is unavailable and the quality signal-to-noise ratio of the near-field optical communication link meets the third threshold, the near-field optical communication link is switched.

[0090] When the quality signal-to-noise ratio of all communication links is lower than the minimum working threshold, it is determined that the communication link is interrupted, and the cache management unit stores the target instruction to be sent. The target instruction is stored in a first-in-first-out queue in chronological order, and the queue uses a circular storage structure. At the same time, a parity check code is added to ensure data integrity.

[0091] The communication link state recovery signal is continuously monitored. When the quality signal-to-noise ratio of any link rises to the working threshold, the cache instruction retransmission mechanism is triggered. The target instruction is extracted from the queue in the storage order, transmitted to the target mechanism through the recovered communication link, and the corresponding instruction data in the queue is cleared and the cache pointer is updated after the transmission is completed.

[0092] Taking the underwater mechanical arm grasping task as an example, the control module generates a target instruction of a clamping force of 25N. Due to ocean current noise, the quality signal-to-noise ratio of the underwater acoustic communication link decreases below the threshold, and the transmission channel is automatically switched to the laser communication link to complete the instruction transmission. When the laser link is interrupted due to suspended matter shielding, the instruction is stored in the cache management unit, and the instruction is retransmitted to the mechanical arm actuator from the queue after the near-field optical communication link is recovered.

[0093] The communication link quality monitoring can realize real-time sensing of channel state changes, avoid instruction loss due to link degradation, switch the transmission channel, ensure optimal communication path selection under different environmental conditions, and provide instruction temporary storage capability by the cache management unit to eliminate the execution window period during communication interruption.

[0094] S300: receiving the weighted fusion feature through the cognitive decision module, constructing a dynamic environment graph at a medium-term situational layer, and constructing an event memory chain and performing cross-task causal migration at a long-term cognitive layer to output a long-term cognitive feature.

[0095] The medium-term situational layer is used to construct a dynamic environment graph. The dynamic environment graph is a topological structure describing the correlation relationship of environmental elements, and the dynamic environment graph represents the spatial distribution relationship of the environmental elements.

[0096] The long-term cognitive layer is used to construct an event memory chain. The event memory chain is structured data recording the temporal relationship of events, recording the time sequence of historical events, and performing cross-task causal migration refers to a mechanism for migrating historical task knowledge to a new task. The long-term cognitive feature is a global fusion feature that fuses the event migration representation and the event memory chain.

[0097] S400: generating target instructions based on the weighted fusion features, dynamic environment map, and long-term cognitive features by the control module.

[0098] In some embodiments, the target instructions are control signals for driving the execution mechanism, including action type, direction parameter, intensity threshold, and other elements. In other embodiments, the target instructions are generating target files, sending actions, etc. That is, the target instructions are dynamically adjusted based on the application environment.

[0099] For example, the weighted fusion features, dynamic environment map, and long-term cognitive features provide action planning framework, real-time environment parameters, etc., combined with generating executable target instructions, which include a set of action instructions, such as parameterized commands such as "heading angle 35°±2°, mechanical arm gripping force 25N±3N".

[0100] S500: transmitting the target instructions to the target mechanism through the communication interaction module, and receiving the feedback data sent by the target mechanism to the long-term cognitive layer.

[0101] The target mechanism is a physical execution terminal that receives the target instructions. Like the target instructions, the target mechanism is determined based on the type of the target instructions. If the target instructions are control signals for driving the execution mechanism, the target mechanism can include power devices such as rotating thrusters and multi-joint mechanical arms. The feedback data is the state monitoring information returned after the target mechanism executes the instructions.

[0102] The communication interaction module transmits the target instructions to the target mechanism. After the target mechanism executes the instructions, it returns the feedback data through the communication link. The feedback data includes execution result code and environment state change amount, which is returned to the long-term cognitive layer for updating the event memory chain.

[0103] For example, in the task of underwater topographic mapping, the perception processing module collects sonar data, the short-term perception layer outputs the weighted fusion features of the seabed topography, the communication interaction module transmits the features through the laser link, the cognitive decision module constructs the topographic map and migrates the historical mapping strategy, and outputs the path planning strategy data; the control module generates obstacle avoidance navigation instructions, and the thruster returns the actual track deviation feedback data to the long-term cognitive layer after execution.

[0104] This application illustrates the application scenarios of underwater unmanned underwater vehicle (Unmanned Underwater Vehicle, UUV) platform, intelligent ship captain system, and office system.

[0105] Embodiment 1

[0106] On the UUV platform, it is necessary to realize multi-modal perception, situational reasoning, autonomous navigation, and task execution in a complex dynamic environment.

[0107] For example, Figure 2In some embodiments, the feature extraction and sensor scheduling are performed in a short-time perception layer, including:

[0108] S201: performing preprocessing on the environmental data to generate preprocessed data.

[0109] The environmental data is data collected by a multi-modal sensor array, including acoustic reflection signals, optical image frames, fluid dynamics vectors, and other physical quantities. The multi-modal sensor array is a hardware collection of heterogeneous sensors such as sonar, camera, flow meter, etc.

[0110] The short-time perception layer runs on a millisecond time scale and is mainly used for fast acquisition, standardization and dynamic feature extraction of multi-modal raw data streams of underwater environment. The multi-modal input at time t is set as follows:

[0111] ;

[0112] wherein, represents the sonar point cloud, represents the optical image frame, represents the water flow vector field, is an image with H pixels in height, W pixels in width, and C channels.

[0113] The preprocessing is normalization preprocessing, including mean normalization and variance scaling operation, to enhance the numerical stability of multi-modal data, see the following formula:

[0114] ;

[0115] wherein, is the historical mean of each modal data, is the variance of each modal data.

[0116] S202: performing feature extraction on the preprocessed data by a convolutional neural network, wherein the convolutional neural network is a deep model composed of convolutional layers and activation functions, for extracting local features of the data.

[0117] The local dynamic features are output by the following formula:

[0118] ;

[0119] wherein, is a convolutional parameter.

[0120] S203: aggregating the local dynamic features based on a time decay function to generate short-term dynamic information by the following formula:

[0121] ;

[0122] ;

[0123] where t is the time, is a weight coefficient, is a physical disturbance coefficient from the mid-time context layer.

[0124] Weight function By introducing top-down adjustment is made, so that short-time feature aggregation can dynamically adapt to complex environmental changes.

[0125] In addition, the short-time perception layer introduces a sensor adaptive scheduling mechanism. According to the current feature state and task demand, the activation weight of each sensor is dynamically calculated by the following formula. The sensor activation weight is a probability vector quantifying the importance of the sensor, which is mapped to the [0, 1] interval through the Sigmoid function:

[0126] ;

[0127] where, is a Sigmoid function, represents the switching probability of the i-th sensor at time t, is a weight matrix, is a bias vector, and M is the number of sensors. The scheduling weight is used to calculate the weighted perception feature by the following formula:

[0128] ;

[0129] where represents element-wise multiplication, realizing the closed-loop coupling of sensor behavior and feature perception.

[0130] S204: Recalculate the sensor activation weight, perform weighted fusion on the short-term dynamic information, and output the weighted fusion feature. Finally, the short-time perception layer outputs the weighted fusion feature by the following formula:

[0131] ;

[0132] where, is a short-time perception layer mapping function.

[0133] The weighted fusion feature is used as the input of the mid-time context layer context state modeling.

[0134] Exemplarily, the sonar array collects point cloud data, the camera acquires front image, the flow meter records water flow direction, the preprocessing stage unifies three modal data scales, the convolutional neural network identifies the convex object profile in the point cloud and the texture edge in the image, the time decay function fuses the features in the last milliseconds, the decay rate increases when the water flow suddenly changes, the activation weight calculation shows that the camera is the most important, the sonar is the second, and the flow meter collection is turned off. The weighted fusion feature outputs the obstacle position and size parameters.

[0135] The medium-time context layer operates at a time scale of seconds to minutes, aiming to build a dynamic map and fuse physical rules for context reasoning. In some embodiments, a dynamic environment map is built at the medium-time context layer, such as Figure 3 As shown, the method comprises the following steps:

[0136] S301: According to the short-time layer feature, i.e. the weighted fusion feature The underwater area is clustered to form a node set V(t), and the node set is a topological unit group generated by spatial clustering, each node representing an environmental sub-region:

[0137] ;

[0138] wherein, represents a κ-cluster-based embedding clustering function.

[0139] S302: The state vector of the node in the node set is iteratively updated by the graph neural network.

[0140] The graph neural network is a deep learning model for processing graph structure data, which updates the state by aggregating neighbor node information. The node state vector is a feature vector describing the node attributes, including location, type, risk value, and other environmental element information. Iterative update means gradually optimizing the node state through multiple rounds of information transmission:

[0141] ;

[0142] wherein, is the neighbor set of node , σ is the Sigmoid mapping function, W is the weight, b is the bias, and l, t are variables.

[0143] To realize neural-symbol coupling, S303: a symbolic physical rule engine is introduced to calculate the environmental disturbance coefficient. Symbolic physical rules are domain knowledge expressed in mathematical formulas or logical statements, such as fluid mechanics equations and obstacle motion laws. The physical disturbance coefficient is adjusted by the following formula:

[0144] ;

[0145] wherein, Δk depends on the environmental disturbance factor.

[0146] The disturbance coefficient acts on the edge weight calculation described by the following formula to realize the structure adaptation of the dynamic environment graph:

[0147]

[0148] wherein, is an edge weight calculation function, which dynamically calculates the connection strength between two nodes according to their states and the current environmental physical parameters.

[0149] S304: Adjust the graph edge weight value in the graph construction system according to the state vector and environmental disturbance coefficient dynamics to generate a dynamic environment graph.

[0150] wherein, the graph edge weight value represents the interaction strength between nodes, such as obstacle relevance, water flow influence factor, etc.

[0151] Finally, the context encoding is output by the following formula, which is a vector representation of the dynamic environment graph for the long-term cognitive layer:

[0152]

[0153] wherein, is the hth head graph attention module, is a physical rule constraint, is a physical regularization strength.

[0154] As shown in Figure 4 , in some embodiments, the output dynamic environment graph is input to the long-term cognitive layer, and the event memory chain is constructed in the long-term cognitive layer, including:

[0155] S401: Analyzing the dynamic environment graph to output entity relationships and pre-stored causal graphs.

[0156] Analyzing the dynamic environment graph refers to the calculation process of extracting the correlation between nodes from the graph, and the entity relationship is a structured data describing the interaction attributes of environmental elements, such as obstacle position relevance, water flow influence strength, etc.

[0157] S402: Generating a time-stamped potential event representation based on the entity relationship and the pre-stored causal graph.

[0158] The pre-stored causal graph is a rule base that stores prior knowledge in the field, including event type definition, causal relationship probability matrix and constraint conditions. The potential event representation describes the low-dimensional feature vector of the environmental event, including event type, spatial position, intensity level, etc. semantic information, time index is a logical label marking the time of event occurrence, supporting absolute timestamp or relative time sequence encoding.

[0159] ​​The long-time cognitive layer operates at the hour-to-task level time scale, and mainly undertakes event memory chain construction, causal graph learning and cross-task knowledge transfer. First, combined with context encoding and a preset causal graph , a latent event representation is generated by the following formula :

[0160] ;

[0161] wherein, is a graph embedding extracted by a graph neural network, explicitly fusing node entity information, edge weight dynamics and symbolic physical rules, is a long-term task memory chain, is a time index of event occurrence, W is a weight, b is a bias, the superscript (K) is a task number, K=1 is a source task, K=2 is a target task, and the subscript i is the number of the event in the task.

[0162] S403: Linking the latent event representation in time index order to construct an event memory chain that stores the time sequence mapping relationship of events.

[0163] The event memory chain is a sequence of latent event representations organized in time index order, storing the time sequence mapping relationship between events, and the mapping process can capture the time-space dependency and its modulation effect on the target event.

[0164] To solve the cross-task knowledge reuse problem, as shown in Figure 5 , in some embodiments, the long-time cognitive layer performs cross-task causal transfer, including:

[0165] S501: Counting the event co-occurrence frequency in the source task event memory chain, and calculating the intra-task event causal relationship matrix.

[0166] First, based on the event co-occurrence probability and the conditional probability, the joint probability is calculated, first to determine the intra-task dependency and the influence of the event on , to build the intra-source task event causal relationship matrix, is the i-th event in the source task (task 1).

[0167] S502: Then calculate the semantic similarity between the source task events and the target task events in the event memory chain to generate target matching events.

[0168] To achieve task generalization, a cross-task memory transfer operator is introduced, and the event in the target task that is most similar to the source task event is found by the following formula, ​For the jth event in the target task (task 2):

[0169] ;

[0170] S503: Determine the dependence coefficient of the source task event on the target event according to the causal relationship matrix;

[0171] ;

[0172] wherein, is the event dependence coefficient, indicating which source task events have the greatest impact on the target event .

[0173] S504: Based on the dependence coefficient and the target matching event, weight the feature embedding of the source task event and output the migration representation of the target event.

[0174] The external event embedding obtained by migration is:

[0175] ;

[0176] If is larger, it means should be migrated to the target task more.

[0177] The long-term cognitive layer finally outputs the long-term cognitive feature through the following formula:

[0178] ;

[0179] wherein, is a memory network for extracting key causal patterns, is the attention weight of the migrated event.

[0180] The global fusion module realizes the integration of features across time scales, and the global memory integration function is defined as follows:

[0181] ;

[0182] wherein, , is a cross-layer attention mechanism, is a long-term memory activation function.

[0183] The final output unified perception state, as shown in Figure 6 , in some embodiments, generates a target instruction, including the following steps:

[0184] S601: Integrate the weighted fusion features of the short-term perception layer, the dynamic environment map of the medium-term situation layer, and the long-term cognitive features of the long-term cognitive layer.

[0185] S602: Correlate the weighted fusion features with the dynamic environment graph through a cross-layer attention mechanism to generate spatio-temporal correlation features;

[0186] S603: Fuse the long-term cognitive features and the transfer representation through a long-term memory activation function to generate knowledge-enhanced features;

[0187] S604: Decode the fusion result of the spatio-temporal correlation features and the knowledge-enhanced features through a decoder to output a unified perception state vector;

[0188] S605: Generate a target instruction based on the unified perception state vector.

[0189] In the embodiment, by The decoder outputs a unified perception state, and the decoder is a pre-trained large model that is fine-tuned:

[0190] ;

[0191] Through the obtained perception state result, in combination with different scene applications, functions such as environment situation estimation, task planning and abnormal warning are supported, and a multi-scale perception and cognition closed loop is constructed.

[0192] The weighted fusion features of the short-time perception layer are feature vectors output by millisecond-level time scale processing, contain real-time dynamic information of the environment, the context encoding of the medium-time context layer is a dynamic environment graph vector representation constructed at a time scale of seconds to minutes, represents the correlation relationship of environment elements, and the long-term cognitive features of the long-time cognitive layer are knowledge-enhanced features generated at a task-level time scale, which fuse historical event chains and transfer event representations.

[0193] The weighted fusion features provide the instantaneous state of the environment, the context encoding describes the correlation of environment elements, and the long-term cognitive features contribute historical experience knowledge. After the feature dimensions are aligned, they are input into a fusion module. The weighted fusion features and the context encoding are associated through a cross-layer attention mechanism. The cross-layer attention mechanism calculates the spatio-temporal dependence between features through multi-head attention. The weighted fusion features are used as query vectors, and the context encoding is used as key-value vectors. An attention weight matrix is calculated to identify the spatio-temporal correlation strength between each element in the weighted fusion features and the context encoding. The context encoding is aggregated to generate spatio-temporal correlation features. This feature retains the details of real-time data while embedding the environmental context.

[0194] The long-term cognitive features and the transfer representation are fused through a long-term memory activation function. The activation function adopts a gated recurrent unit structure. The long-term cognitive features are input as the main channel, and the transfer representation is input as the auxiliary channel. An update gate signal is generated to control the injection proportion of the transfer representation. The knowledge-enhanced features are output. When the confidence of the transfer representation is lower than a threshold, the gate signal automatically attenuates the weight of the auxiliary channel.

[0195] The decoder parses the spatio-temporal correlation feature and the knowledge enhanced feature, the decoder is fine-tuned based on the pre-trained Transformer architecture, the input layer receives the spliced fusion feature, the global dependence between features is extracted through multiple layers of self-attention mechanism, a unified perception state vector is output, and a target instruction is generated based on the vector. The instruction parameters include action type, spatial coordinates, intensity threshold and other executable elements.

[0196] For example, the perception processing module collects seabed environment data through a sonar array, an optical camera and a CTD sensor, performs standardization preprocessing on the sonar point cloud, image frames and temperature-salinity-depth data, eliminates dimensional differences, extracts local dynamic features through a three-dimensional convolutional neural network, identifies reef contours and biological activities, aggregates features in the last millisecond based on a time decay function, automatically increases the decay rate when there is a sudden change in water flow, calculates sensor activation weights, and when the optical visibility is low, the camera is turned off and the sonar weight is increased. The weighted fusion features are output for obstacle avoidance decision-making.

[0197] Spatial clustering is performed on the acoustic features, the detection area is divided into grid nodes, and the graph neural network iteratively updates the node state, fuses the sonar echo intensity and water flow data, injects the symbolic physical rule "flow rate > 2 m / s increases the risk of corrosion", and the environmental disturbance coefficient increases with the increase of turbulence intensity, reducing the edge weight between nodes in high flow rate areas. A dynamic environmental map identifying high-risk corrosion areas is constructed.

[0198] The seabed topographic map is further parsed to extract the "hydrothermal vent and mineral deposition" entity relationship. The pre-stored causal map "vent activity, mineral enrichment probability 0.9" rule is matched to generate a potential representation of hydrothermal activity events. The vent activation, mineral deposition, and sensor reading anomaly event chain are linked by timestamp, and stored as an exploration task memory chain.

[0199] The semantic similarity (0.88) of the hydrothermal sampling event in the South China Sea exploration task and the cold spring sampling event in the East China Sea task is calculated. The co-occurrence probability of the source task is counted, the hydrothermal activity and mineral enrichment probability is 0.92, the dependence coefficient is 0.90, the mineral analysis strategy feature is weighted and migrated, and the cold spring area sampling migration representation is output.

[0200] The sonar real-time weighted fusion feature, seabed topography context encoding and historical exploration long-term cognitive feature are integrated, the current pose is associated with the mineral location through cross-layer attention, the spatio-temporal correlation feature containing the optimal path is output, the similar mineral area collection parameters are injected through the long-term memory activation function, the knowledge enhanced feature is generated, and the decoder outputs the operation instruction "propulsion direction 120°±5°, mechanical arm gripping force 35N±3N".

[0201] Embodiment 2

[0202] The short-term perception layer receives the sea state information of the marine environment received by the navigation sensor, and multi-modal information such as video and equipment information from the equipment monitor and camera in the ship-borne equipment, performs a preprocessing, and extracts local dynamic features through convolution operation. In the scenario of the present embodiment, the environmental data can include sea state data, video and equipment information from the equipment monitor and camera in the ship-borne equipment. The execution process of the short-term perception layer can be seen in steps S201-S204, which will not be repeated here.

[0203] As shown in Figure 7 , the short-term perception layer receives through the navigation sensor, the equipment detector, and the video analysis system. The navigation sensor receives the data of the ship GPS and / or gyroscope, the sea state data, the vibration and / or temperature data of the ship-borne facilities detected by the equipment detector, and the camera data of the ship-borne facilities.

[0204] The target position is calculated through the navigation sensor in the short-term perception layer, the equipment detector outputs the equipment information, and the video analysis system outputs the video stream information.

[0205] The medium-term context layer receives the target position, the equipment information, and the equipment information, constructs the cruising posture graph, the equipment detection model, and the anomaly detection engine, converts the physical environment into a weighted node network, and checks the collision avoidance rules in combination with the neural symbol engine. The execution process of the medium-term context layer can be seen in steps S301-S304, which will not be repeated here.

[0206] The long-term cognitive layer establishes an event memory chain through the historical database and the cruising knowledge base, and finds the optimal cruising route and maintenance strategy through causal graph learning. The execution process of the long-term cognitive layer can be seen in steps S401-S403 and S501-S504, which will not be repeated here.

[0207] Specifically, the sea state received by the navigation sensor is normalized and preprocessed with multi-modal information X(t) such as sea state, video, and equipment, i.e. sea state, video, and equipment.

[0208] The preprocessed data performs convolution operation to extract local dynamic features The data is captured for a short period of dynamic information to generate :

[0209] ;

[0210] ;

[0211] That is, the features f1(τ) within the time window of t-Δt from the current time t are weighted and summed to obtain the features fused with short-term dynamic information The input is The local dynamic feature (extracted by convolutional neural network, etc.) at time point τ, where τ belongs to the time window [t-Δt, t].

[0212] The weight function is an exponential decay function, including a time decay factor and an environmental disturbance coefficient, where the time decay factor is where (t-τ) is the time difference between the current time t and the historical time τ, and α is the decay coefficient, which controls the rate of time decay. The larger the time difference (t-τ), the smaller the weight, i.e., the smaller the weight of the more distant data.

[0213] The environmental disturbance coefficient is a coefficient from the middle context layer, representing the intensity of the current environmental disturbance (such as sudden strong water flow, obstacles, etc.).

[0214] Multiply the feature f1(τ) at each time point τ in the time window [t-Δt, t] by the corresponding weight w(τ), and then perform weighted summation to obtain the fused feature .

[0215] In the short-term perception layer, a sensor adaptive scheduling mechanism is adopted. Since multiple modal sensors work simultaneously, it may cause high energy consumption, information redundancy, and excessive computational load. Therefore, non-critical sensors need to be dynamically turned off, and the current task requirements (such as navigation sensors in the exploration stage, equipment detection instruments, and video analysis systems in the detection stage) as well as environmental changes (prefer navigation sensors in strong water flow) and feature importance (start high-precision sensors when detecting anomalies) need to be considered.

[0216] The sensor activation probability is:

[0217] ;

[0218] Based on the current extracted environmental features , the activation probability of each sensor is calculated through a neural network. When the feature shows strong water flow disturbance, the weight of the navigation sensor increases, and when the feature shows obstacles in front, the weights of the sonar and camera increase.

[0219] Feature weighted fusion:

[0220] ;

[0221] where is the element-wise multiplication (Hadamard product), is the sensor activation weight vector, is the original feature vector (grouped by sensors).

[0222] In the middle-time emotional layer, a dynamic map of the environment is constructed, analogous to spatial memory (experience analysis). Specifically, according to the output of the first layer The cruise area is clustered to form a node set :

[0223] ;

[0224] Then, a node state update is performed through a graph neural network:

[0225] ;

[0226] In the middle-time emotional layer, a dynamic graph is used, that is, the state of each node is dynamically updated according to the state of its neighbor nodes, forming a dynamic graph. This dynamic graph can better represent the changes in the underwater environment.

[0227] Neighbor node state aggregation: the new state of a node is obtained by weighted summation of the states of its neighbor nodes, the weight matrix is used to adjust the contribution of neighbor node states, and the bias vector is used to adjust the initial value of the node state. The activation function σ introduces nonlinearity, enabling the model to learn complex patterns. The weight matrix and the bias vector can be dynamically adjusted through training to adapt to different environmental conditions.

[0228] The symbolic physical rule engine is introduced to enhance the model's understanding and reasoning ability for the physical environment. The symbolic physical rule engine generates symbolic rules based on known physical rules (such as water flow, temperature change, etc.). These rules can be simple logical expressions, such as "if the water flow velocity increases, the connection strength between nodes increases."

[0229] The physical disturbance coefficient is added, which will be dynamically adjusted according to the physical changes of the environment, and the physical disturbance coefficient is calculated by the following formula:

[0230] ;

[0231] The physical disturbance is used for edge weight calculation in the graph neural network:

[0232] ;

[0233] According to the states of two nodes and the current environmental physical parameters, the connection strength between them is dynamically calculated, and the final output is the scenario encoding. This is used by the next layer, the long-time cognitive layer.

[0234] The long-term cognitive layer is used to process cross-task knowledge transfer and long-term memory construction. Specifically, first, the construction of the event memory chain is performed, which is used to record and organize key events encountered by the system in different tasks. The long-term cognitive layer first generates potential event representations in combination with the features and causal graph of the medium-term context layer. These event representations are organized according to the time sequence and causal relationship to form an event memory chain CT.

[0235] As shown in Figure 8 , when receiving a cruise task instruction, the long-term cognitive layer starts the event chain construction process. The input of the medium-term layer is the navigation sensor, and the detection stage requires the scene encoding of the equipment detection instrument and the video analysis system.

[0236] The purpose of causal graph learning is to understand the causal relationship between events. By calculating the conditional probability, the influence degree of one event on another event can be determined. This understanding of causal relationship helps to utilize existing knowledge for reasoning and decision-making when facing new tasks.

[0237] As shown in Figure 9 , by analyzing the causal relationship in the historical event chain, when a vibration anomaly is detected, an automatic warning is given. The warning information can be: "vibration anomaly (0.15mm / s) may cause power drop (probability 87%), suggest checking lubrication system".

[0238] The purpose of cross-task knowledge transfer is to transfer the knowledge learned in the source task to the new target task. The similarity between the event in the source task and the event in the target task is calculated by the following formula, and combined with the causal influence coefficient, the experience in the source task can be applied to the new task, which can quickly adapt to the new task and make efficient decisions based on existing experience.

[0239] ;

[0240] wherein, is the event representation of the target task, is the event in the source task, is the event in the target task, is the causal influence coefficient, is the similarity function.

[0241] As shown in Figure 10 , when performing a new cruise task, historical experience is transferred. For example, the current task is to perform a cruise task transfer, the task is a South China economic cruise speed strategy, and the target task is an East China cruise speed optimization. The effect is a 12% reduction in fuel consumption.

[0242] The device inspection migration task is a generator bearing fault handling scheme, and the target task is a diesel engine set preventive inspection. This event memory and causal migration-based architecture enables the intelligent captain system to have experience accumulation-reasoning decision-making capabilities, enabling a leap from passive response to active prevention.

[0243] In the global fusion layer, the features of the short-term perception layer, the medium-term context layer, and the long-term cognitive layer are fused across time scales to generate the final decision output. Specifically, the features of different time scales are integrated and a unified perception state is output to support system decision-making and task execution.

[0244] Specifically, the global fusion layer receives features from the short-term perception layer , the medium-term context layer and the long-term cognitive layer , and integrates them into a global feature .

[0245] Through cross-time scale fusion, the global fusion layer can consider information at different time scales to generate a comprehensive perception state. This fusion mechanism can make more accurate and stable decisions in dynamic environments. The global fusion layer outputs a unified perception state through a decoder, which can be used to support environmental situation estimation, task planning, and anomaly warning functions. The decoder is usually a fine-tuned pre-trained large model that maps the global feature to the final perception state, and finally outputs navigation instructions, device inspection information, or anomaly warnings.

[0246] For example, the navigation sensor collects sea wave data, and the device monitor obtains engine vibration signals. In the preprocessing stage, the vibration frequency and wave height dimensions are unified. The convolutional neural network extracts abnormal pulse features in the vibration spectrum, the time decay function fuses short-term features, the decay rate increases under stormy weather, the activation weight calculates the priority of the navigation sensor, and when a device anomaly is detected, the high-precision vibration sensor is started. The weighted fusion features are output to the fault diagnosis module.

[0247] Navigation feature clustering generates channel nodes, and graph neural networks update node state representations of buoy positions. Inject the navigation rule "visibility <100 meters requires deceleration", the environmental disturbance coefficient increases with fog concentration, adjust the edge weight value between nodes for navigation risk, and output a dynamic channel graph with storm influence weight.

[0248] Analyze the device state graph to extract the entity relationship between bearing wear and vibration intensification. Match the causal graph "wear >0.2mm, power drop probability 0.85" rule to generate a bearing abnormal event representation. Link the wear occurrence, vibration overrun, and power drop event chain in time sequence to construct a device fault memory chain.

[0249] Match the South China Sea speed strategy event with the East China Sea task speed optimization event semantic similarity (0.85). Statistical source task condition probability: economic speed, oil consumption reduction probability 0.89, dependence coefficient 0.87, migrate historical speed curve characteristics, generate East China Sea route migration representation.

[0250] Fusion of current heading weighted fusion features, channel state situation encoding and historical navigation long-term cognitive features, cross-layer attention association of ship position and buoy position, generation of collision avoidance space-time features, memory activation function migration of historical storm response strategy, enhancement of decision features. The decoder outputs the navigation instruction "speed down to 12 knots, right rudder 15°".

[0251] Embodiment 3

[0252] As Figure 11 shown, for the field of intelligent office system, a hierarchical memory perception architecture is used to realize intelligent management. The short-term perception layer analyzes the user uploaded planning document in real time, extracts key information such as task description, time node and responsible person through NLP technology, and monitors user operation behavior and system data changes. The execution process of the short-term perception layer can be seen in the steps of S201-S204, which will not be repeated here.

[0253] The medium-term situation layer constructs a dynamic project status map, uses a graph neural network to model task dependency relationships, and combines a risk propagation model to predict delay probabilities, such as automatically calculating the impact weight on the test phase when the design task is delayed. The execution process of the medium-term situation layer can be seen in the steps of S301-S304, which will not be repeated here.

[0254] The long-term cognitive layer retrieves the historical project knowledge base, and through cross-project transfer learning, adapts the task structure and risk strategy of similar projects to the new project, such as reusing the demand change handling scheme of A project to B project. The execution process of the long-term cognitive layer can be seen in the steps of S401-S403 and S501-S504, which will not be repeated here.

[0255] The global fusion layer integrates the outputs of the three layers to generate decisions, which are verified by the rule engine in the large model hub before triggering the execution layer to automatically generate a plan, send a responsible person notification, and update the system status, forming a perception, analysis, decision, and execution closed loop.

[0256] For example, the user terminal collects document text, operation behavior logs, and system performance data. The text data is normalized by word vector, and the operation log is converted into a time series vector. The convolutional neural network extracts local features of key phrases in the document. The time decay function focuses on recent operation events, and the decay rate increases when the system is lagging. The activation weight dynamically allocates resources, and the voice collection weight is increased during important meetings. The output weighted fusion features are provided for agenda analysis.

[0257] Project document feature clustering generation task node, graph neural network modeling node dependency relationship, inject "design delay causes test delay" project rule, environmental disturbance coefficient increases with the number of requirement changes, adjust the influence edge weight between task nodes, construct a project state graph with delay risk weight.

[0258] Parse the project graph and extract "requirement change, development delay" entity relationship. Match the causal graph "change frequency > 3, delay probability 0.75" rule to generate a requirement change event representation. Link the change submission, design delay, and test compression event chain according to time index to form a project management memory chain.

[0259] Calculate the similarity between A project requirement freeze event and B project requirement change event (0.78). Calculate the source task joint probability: requirement change, test compression probability 0.82, dependency coefficient 0.80, migrate requirement control strategy features, and output B project migration representation.

[0260] Integrate document keyword weighted fusion features, project dependency context encoding, and historical case long-term cognitive features, cross-layer attention links current tasks with milestone nodes, output risk spatio-temporal features, memory activation function injects similar project review rules, generates knowledge enhanced features, and decoder outputs design review start, delay warning notification management instructions.

[0261] Based on the above embodiment intelligent hierarchical memory perception method, as shown in Figure 12 The embodiment also provides an embodiment intelligent hierarchical memory perception system, which comprises:

[0262] The perception processing module is configured to collect environmental data, perform feature extraction and sensor scheduling in the short-time perception layer, and output weighted fusion features;

[0263] The cognitive decision module is configured to receive the weighted fusion features, construct a dynamic environment graph in the medium-time context layer, and construct an event memory chain and perform cross-task causal migration in the long-time cognitive layer, and output long-term cognitive features;

[0264] The control module is configured to generate target instructions based on the weighted fusion features, dynamic environment graph, and long-term cognitive features;

[0265] The communication interaction module is configured to transmit the weighted fusion features to the cognitive decision module, and transmit the target instructions to the target mechanism, and receive feedback data sent by the target mechanism to the long-time cognitive layer.

[0266] In some embodiments, the system further comprises an energy management module, which is configured to:

[0267] Obtain the power consumption of the sensor array and the cognitive decision module;

[0268] based on the power consumption, assigning power supply priority of the sensor array and the cognitive decision module according to a task execution stage.

[0269] The energy management module is a control system for coordinating power supply, including a high energy density power supply unit, an energy recovery mechanism, and a power supply scheduling control unit. The power supply scheduling control unit is an integrated circuit for executing a dynamic power consumption strategy, and distributes current paths according to load requirements.

[0270] A current monitoring circuit is deployed in the power supply scheduling control unit to collect the working current of each probe of the sensor array and the processor load current of each module, calculate the total power consumption per unit time, and generate an energy consumption evaluation report in combination with the remaining power of the power supply unit.

[0271] The current task execution stage is identified, the environment perception stage is mainly for sensor data collection, the decision calculation stage is mainly for processor inference, the instruction execution stage is mainly for power mechanism operation, and the standby maintenance stage turns off non-core components. The stage identifier code is read through the system state register to determine the type of the dominant hardware component.

[0272] The application also provides a computer program product, including a computer program which, when executed by a processor, implements the above-mentioned embodiment of the embodied intelligent hierarchical memory perception method.

[0273] The computer readable storage medium described above can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disk. The readable storage medium can be any available medium that can be accessed by a general-purpose or special-purpose computer.

[0274] An exemplary readable storage medium is coupled to the processor, so that the processor can read information from the readable storage medium, and can write information to the readable storage medium. Of course, the readable storage medium can also be an integral part of the processor. The processor and the readable storage medium can be located in an application specific integrated circuit (ASIC). Of course, the processor and the readable storage medium can also exist as discrete components in the embodied intelligent data processing device.

[0275] Those skilled in the art can understand that all or part of the steps of the above-mentioned method embodiments can be completed by program instruction related hardware. The foregoing program can be stored in a computer readable storage medium. The program executes to perform the steps of the above-mentioned method embodiments; and the foregoing storage medium includes various storage media that can store program codes, such as ROM, RAM, magnetic disk or optical disk.

[0276] The similar parts among the embodiments provided in the application can be referred to each other, the specific embodiments provided above are only a few examples under the general concept of the application, and do not constitute the limitation of the protection scope of the application. For those skilled in the art, any other embodiments extended according to the application scheme without creative labor shall fall within the protection scope of the application.

Claims

1. A body-aware hierarchical memory perception method, characterized in that, The method comprises: collecting environment data through a perception processing module, performing feature extraction and sensor scheduling in a short-time perception layer to output weighted fusion features; the feature extraction and sensor scheduling in the short-time perception layer to output the weighted fusion features comprises: performing preprocessing on the environment data to generate preprocessed data, the environment data being data collected through a multi-modal sensor array; performing feature extraction on the preprocessed data through a convolutional neural network to output local dynamic features; aggregating the local dynamic features based on a time decay function to generate short-term dynamic information, wherein the decay rate of the time decay function is adjusted based on the medium-time context layer; calculating sensor activation weights, performing weighted fusion on the short-term dynamic information, and outputting the weighted fusion features; transmitting the weighted fusion features to a cognitive decision module through a communication interaction module; receiving the weighted fusion features through the cognitive decision module, constructing a dynamic environment graph in a medium-time context layer, and constructing an event memory chain and performing cross-task causal migration in a long-time cognitive layer to output long-term cognitive features; the construction of the dynamic environment graph in the medium-time context layer comprises: performing spatial clustering on the weighted fusion features to generate a node set; iteratively updating the state vector of the nodes in the node set through a graph neural network; injecting symbolic physical rules into a graph construction system to calculate an environmental disturbance coefficient; adjusting the graph edge weight value in the graph construction system according to the state vector and the environmental disturbance coefficient to generate a dynamic environment graph; the construction of the event memory chain in the long-time cognitive layer comprises: analyzing the dynamic environment graph to output entity relationships and a pre-stored causal graph; generating a time-stamped potential event representation based on the entity relationships and the pre-stored causal graph; linking the potential event representations in chronological order to construct an event memory chain that stores event and timing mapping relationships; the cross-task causal migration in the long-time cognitive layer comprises: calculating an intra-task event causal relationship matrix through event co-occurrence probability and conditional probability; calculating the semantic similarity between source task events and target task events in the event memory chain to generate target matching events; determining the dependence coefficient of the source task events on the target events according to the causal relationship matrix; weighting and migrating the feature embedding of the source task events based on the dependence coefficient and the target matching events to output the migration representation of the target events; outputting long-term cognitive features based on the migration representation of the target events; generating target instructions based on the weighted fusion features, the dynamic environment graph, and the long-term cognitive features through a control module; transmitting the target instructions to a target mechanism through a communication interaction module and receiving feedback data sent by the target mechanism to the long-time cognitive layer.

2. The embodied intelligent hierarchical memory perception method of claim 1, wherein, the generation of the target instructions comprises: integrating the weighted fusion features of the short-time perception layer, the dynamic environment graph of the medium-time context layer, and the long-term cognitive features of the long-time cognitive layer; associating the weighted fusion features and the dynamic environment graph through a cross-layer attention mechanism to generate spatio-temporal correlation features; The long-term memory activation function is fused with the long-term cognitive features to generate knowledge-enhanced features; The fusion result of the spatio-temporal correlation features and the knowledge-enhanced features is decoded by a decoder to output a unified perception state vector; Based on the unified perception state vector, a target instruction is generated.

3. The embodied intelligent hierarchical memory perception method of claim 1, wherein, The target instruction is transmitted to a target mechanism through a communication interaction module, including: A transmission channel is established through a water acoustic communication link, a laser communication link, and a near-field optical communication link; According to the quality signal-to-noise ratio of the communication link, the water acoustic communication link, the laser communication link, and the near-field optical communication link are switched; In the case of communication link interruption, the target instruction is stored by a cache management unit; In the case of resuming communication, the target instruction is retransmitted to the target mechanism.

4. A body-aware hierarchical memory perception system, comprising: It includes: The perception processing module is configured to collect environmental data, perform feature extraction and sensor scheduling at the short-time perception layer to output weighted fusion features; The short-time perception layer performs feature extraction and sensor scheduling to output weighted fusion features, including: The environmental data is preprocessed to generate preprocessed data, and the environmental data is collected by a multi-modal sensor array; The preprocessed data is subjected to feature extraction by a convolutional neural network to output local dynamic features; Based on a time decay function, the local dynamic features are aggregated to generate short-term dynamic information, wherein the decay rate of the time decay function is adjusted based on the medium-time context layer; The sensor activation weight is calculated, and the short-term dynamic information is subjected to weighted fusion to output the weighted fusion features; The cognitive decision module is configured to receive the weighted fusion features, construct a dynamic environment graph at the medium-time context layer, and construct an event memory chain and perform cross-task causal migration at the long-time cognitive layer to output long-term cognitive features; The dynamic environment graph is constructed at the medium-time context layer, including: The weighted fusion features are subjected to spatial clustering to generate a node set; The state vector of the nodes in the node set is iteratively updated by a graph neural network; Symbolic physical rules are injected into the graph construction system to calculate environmental disturbance coefficients; According to the state vector and environmental disturbance coefficients, the graph edge weight value in the graph construction system is adjusted to generate a dynamic environment graph; The event memory chain is constructed at the long-time cognitive layer, including: The dynamic environment graph is parsed to output entity relationships and pre-stored causal graphs; Based on the entity relationships and pre-stored causal graphs, a time-stamped potential event representation is generated; The potential event representations are linked in chronological order to construct an event memory chain that stores event and timing mapping relationships; Cross-task causal migration is performed at the long-time cognitive layer, including: An intra-task event causal relationship matrix is calculated by event co-occurrence probability and conditional probability; The semantic similarity between source task events and target task events in the event memory chain is calculated to generate a target matching event; According to the causal relationship matrix, the dependence coefficient of the source task event on the target event is determined; based on the dependency coefficient and the target matching event, weighting the feature embedding of the source task event to output a migration representation of the target event; based on the migration representation of the target event, output a long-term cognitive feature; a control module configured to generate a target instruction based on the weighted fusion feature, a dynamic environment map, and the long-term cognitive feature; a communication interaction module configured to transmit the weighted fusion feature to the cognitive decision module, and transmit the target instruction to a target mechanism, and receive feedback data sent by the target mechanism to the long-term cognitive layer.

5. The embodied intelligent hierarchical memory perception system of claim 4, wherein, Further comprising an energy management module configured to: acquire power consumption of the sensor array and the cognitive decision module; based on the power consumption, according to a task execution stage, allocate power supply priority of the sensor array and the cognitive decision module.

6. An electronic device, comprising: comprising: a processor, and a memory in communication connection with the processor; the memory stores computer execution instructions; the processor executes the computer execution instructions stored in the memory to realize the embodied intelligent hierarchical memory perception method according to any one of claims 1 to 3.

Citation Information

Patent Citations

  • Sensitive active control method and system for body intelligence

    CN120406171A

  • Multi-mode perception and interaction method and device in personal environment

    CN120492561A