An Incremental Cognitive System and Method for Multidimensional Ontology of Intelligent Agents

By using an intelligent agent multidimensional ontology incremental cognitive system, multimodal data processing and neural network training are employed to solve the problem that intelligent agents have difficulty representing complex geometric configurations and multidimensional physical properties in open environments. This enables intelligent agents to accurately perceive multidimensional features of their ontology and achieve safe and efficient whole-body motion control.

CN119807712BActive Publication Date: 2025-10-31TONGJI UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411976287.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-31
Publication Date
2025-10-31
Estimated Expiration
2044-12-31

AI Technical Summary

Technical Problem

Existing technologies are unable to effectively characterize the complex geometric configurations and multidimensional physical properties of intelligent agents, making it difficult for intelligent agents to achieve safe and efficient whole-body motion control in open environments.

Method used

An incremental cognitive system based on multidimensional ontology is adopted, which includes an ontology motion perception module, a cognitive evaluation module, a memory pool module, a knowledge mining module, and a cognitive cultivation module. Through multimodal data processing and neural network training, a cognitive model of the agent's multidimensional ontology characteristics is established.

Benefits of technology

It enables intelligent agents to accurately perceive and represent multi-dimensional features of the body, effectively plan the physical interaction between the body at any position and the outside world, and improve the safety and effectiveness of motion control.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119807712B_ABST
    Figure CN119807712B_ABST
Patent Text Reader

Abstract

This invention relates to an incremental ontology-based cognitive system and method for intelligent agents with multi-dimensional ontology. The invention includes an ontology motion perception module, an ontology cognitive evaluation module, a memory pool module, an ontology knowledge mining module, and an ontology cognitive training module. The ontology motion perception module is used to obtain multimodal data; the ontology cognitive evaluation module is used to obtain ontology cognitive quantization error using the multimodal data; the memory pool module is used for data storage; the ontology knowledge mining module is used to establish an ontology knowledge base using multimodal data mining; and the ontology cognitive training module is used to construct an ontology cognitive model and train the ontology cognitive model under the guidance of the ontology cognitive quantization error and the ontology knowledge base. Compared with existing technologies, this invention has advantages such as multi-dimensional feature representation and perception of the ontology, improved accuracy of cognitive data, and reduced overall error.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intelligent agent cognitive technology, and in particular to an intelligent agent multi-dimensional ontology incremental cognitive system and method. Background Technology

[0002] Embodied agents, possessing perception, reasoning, and movement capabilities based on embodied interaction, have the potential for autonomous and stable movement in open, unstructured environments, much like humans. These challenging scenarios demand higher levels of adaptability and intelligence from the agent. The agent needs to handle real-time physical interactions between its arbitrary position and the external environment during autonomous movement to avoid catastrophic movement failures in operational scenarios. Currently, to effectively represent the agent's topological structure and meet operational and planning requirements, the relationships between moving parts are represented by DH (Denavit-Hartenberg) matrices after being abstracted into virtual links. By constructing DH matrix chains (virtual moving links and joints), a description of the agent's positional and motion spaces can be quickly built. However, agent ontology representation methods based on abstract links struggle to represent the agent's complex geometric configuration and multidimensional physical characteristics. Characterizing and detecting the embodied interactions between arbitrary body positions and the external environment and applying them to the agent's motion control remains a significant challenge.

[0003] Currently, the main solution to the problem of safe interaction management for intelligent agents is based on sensing devices (such as electronic skin, joint torque sensors, and motor current feedback) to detect interactive collisions. Appropriate compliant control methods (impedance control, admittance control, etc.) are designed to control the contact force within a safe threshold range to avoid catastrophic motion failures. However, this compliant motion control method based on virtual levers lacks consistency with task planning, making it difficult to coordinate the safety of the agent's actions with the effectiveness of task execution. Summary of the Invention

[0004] The purpose of this invention is to overcome the shortcomings of the prior art and provide an intelligent agent multi-dimensional ontology incremental cognition system and method.

[0005] The objective of this invention can be achieved through the following technical solutions:

[0006] According to one aspect of the present invention, a multi-dimensional ontology incremental cognitive system for intelligent agents is provided, characterized in that the system includes an ontology motion perception module, an ontology cognitive evaluation module, a memory pool module, an ontology knowledge mining module, and an ontology cognitive cultivation module;

[0007] The system comprises the following modules: an ontology motion perception module for collecting and processing ontology perception data to obtain multimodal data; an ontology cognition evaluation module for using multimodal data to invert ontology activities in digital space and comparing the inversion results with ontology cognition results to obtain ontology cognition quantification error; a memory pool module for data storage; an ontology knowledge mining module for using multimodal data to mine the intrinsic relationship between agent perception modalities and ontology representation dimensions to establish an ontology knowledge base as the basis for data extraction and training of the ontology cognition cultivation module; and an ontology cognition cultivation module for constructing an ontology cognition model and training the ontology cognition model under the guidance of ontology cognition quantification error and ontology knowledge base to map multimodal data to a high-dimensional space to enable the agent to cognize the multidimensional characteristics of the ontology.

[0008] As a preferred technical solution, the proprioception module includes a color depth (RGB-D) visual perception unit, an embodied interaction perception unit, and a proprioception unit. The RGB-D visual perception unit is used to collect two-dimensional images and depth information of the agent's movement process from its own perspective and transmit them to the visual perception queue. The embodied interaction perception unit is used to collect four modes: temperature, vibration, proximity, and pressure and transmit them to the depth interaction perception queue. The proprioception unit is used to collect the angular displacement, angular velocity, and joint torque of the joints during the agent's movement and transmit them to the proprioception queue.

[0009] As a preferred technical solution, the proprioception module also includes a signal uniform calibration unit. This unit is used to use a time-series consistency calibration algorithm to extract time samples at corresponding moments from the visual perception queue, the embodied interaction perception queue, and the proprioception queue, respectively, using the embodied interaction perception signal as the time reference. The sampled and aligned data are then grouped by action to construct visual / tactile-proprioception pairs.

[0010] As a preferred technical solution, the ontology cognition evaluation module includes an ontology action inversion unit, a cognitive bias unit, and a data interaction unit. The ontology action inversion unit is used to reproduce the agent's real-world limb movements in digital space, converting multimodal information into action-state tensor pairs. The cognitive bias unit uses the inverted actions as a reference for the real actions to calculate the ontology cognition quantization error. The data interaction unit is used to extract multimodal data and ontology cognition result data for ontology cognition bias quantization and to send the ontology cognition quantization error.

[0011] As a preferred technical solution, the memory pool module includes a short-term memory unit and a long-term memory unit; the short-term memory unit is used to store multimodal data acquired by the agent in the current time sequence, providing a quantitative data source for the multidimensional ontology cognition evaluation module; the long-term memory unit is used to store multimodal data exceeding the preset tolerance of ontological perception bias, providing training data for the training of the ontology cognition model.

[0012] As a preferred technical solution, the ontology cognition training module includes a forward encoding unit, a backward encoding unit, a backbone network unit, a forward decoding unit, and a backward decoding unit. The forward encoding unit includes a visual encoding network built on a ResNet50 pre-trained deep residual network, an ontology perception encoding network built on an embedding and multilayer perceptron (MLP) network structure, and a tactile encoder network built on MLP network technology. The backward encoding unit includes an intelligent agent ontology material point set encoding network built on an MLP network. The backbone network includes a multilayer deep neural network based on a Transformer structure, used to map multimodal information to a high-dimensional space to implicitly represent the multidimensional motion features of the intelligent agent ontology. The forward decoder includes a Transformer and MLP combined network, used to output the spatial state of the intelligent agent ontology material point set. The backward decoder includes an MLP network, used to output predicted actions in joint space.

[0013] According to another aspect of the present invention, an agent-based multi-dimensional ontology incremental cognitive system is provided, characterized in that the method is applied to the agent-based multi-dimensional ontology incremental cognitive system as described above, wherein the method utilizes a memory pool module for data storage, and the method includes the following steps:

[0014] S1. Collect and process the body's perception data using the body motion perception module to obtain multimodal data;

[0015] S2. Using the ontology cognition training module, the neural network model is initially trained based on multimodal data to establish an ontology cognition model, and the ontology cognition model is used to process multimodal data to obtain ontology cognition results.

[0016] S3. Using the ontology cognition assessment module, ontology activities are inverted in the digital space based on multimodal data, and the inversion results are compared with ontology cognition results to obtain ontology cognition quantification error.

[0017] S4. Use the ontology knowledge mining module to process multimodal data and obtain an ontology knowledge base;

[0018] S5. Quantify the error of ontology cognition to clarify the target training ontology cognition dimension, clarify the target training perception modality from the knowledge base, train the ontology cognition dimension and the target training perception modality based on the target, and use the ontology cognition cultivation module to train the ontology cognition model based on multimodal data.

[0019] S6. The trained ontology cognition model processes multimodal data to output ontology cognition results and returns to execute step S4 to realize multidimensional ontology incremental cognition of the agent.

[0020] As a preferred technical solution, the process of obtaining the ontology cognition quantification error in S3 is as follows: First, the motion state of the agent in the ontology cognition model is discretized into a finite number of material point sets. Then, visual / tactile-proprioceptive sensations are extracted from the memory pool module as input to predict the spatial state of these material point sets, so as to realize the reproduction of physical reality's limb movements in digital space. Finally, the error of the ontology cognition result is quantified by the mean square error method as the ontology cognition quantification error.

[0021] As a preferred technical solution, the specific process of using the ontology knowledge mining module in S4 to process multimodal data and obtain the knowledge base is as follows: First, the perceptual modality and ontology representation are materialized to obtain entity nodes. Then, the multimodal data is differentiated to obtain multimodal data with gradient characteristics. Then, the multimodal data with gradient characteristics is fused into the entity nodes to obtain entity data. Based on Bayes' theorem, the ontology knowledge base is constructed using the entity data.

[0022] As a preferred technical solution, the training process for training the ontology cognition model in S5 uses mean square error (MSE) as the loss function.

[0023] Compared with the prior art, the present invention has the following beneficial effects:

[0024] 1. In this invention, an incremental ontological cognitive model is established and trained through an intelligent agent multi-dimensional ontological incremental cognitive system composed of an ontological motion perception module, an ontological cognition evaluation module, a memory pool module, an ontological knowledge mining module, and an ontological cognition cultivation module. This ontological cognitive model endows the intelligent agent with the ability to model multi-dimensional motion features of the whole body. Based on this ability, the intelligent agent cognitive model can represent and perceive any position of the body, enabling the intelligent agent to accurately perceive and represent multi-dimensional ontological features, thereby effectively planning the physical interaction between any position of the body and the outside world.

[0025] 2. In this invention, the ontology cognition evaluation module includes an ontology action inversion unit, a cognitive bias unit, and a data interaction unit. The agent inverts its ontology activities in a digital space and compares the inversion results with the ontology cognition results to obtain the ontology cognition quantification error. By observing its own movement in real time and comparing it with the ontology cognition network results, the ontology cognition quantification error is obtained, providing a continuous source of data for updating the ontology cognition network and discovering new knowledge. By using the ontology cognition training module to conduct targeted training on data with large cognitive biases, the ontology cognition network can be continuously improved and enriched during the agent's movement, enhancing its accuracy.

[0026] 3. In this invention, the intrinsic relationship between the agent's perceptual modality and the ontology representation dimension is explored through the ontology knowledge mining module. Multimodal data is used to mine the intrinsic relationship between the agent's perceptual modality and the ontology representation dimension, so as to establish an ontology knowledge base as the basis for data extraction and training of the ontology cognition training module. This enables the ontology cognition network to selectively extract corresponding perceptual modality data from the memory pool system to train the ontology cognition dimension with large errors, thereby reducing the overall cognitive data error. Attached Figure Description

[0027] Figure 1 This is a block diagram of the multi-dimensional ontology incremental cognitive system of the present invention;

[0028] Figure 2 This is a flowchart of the multimodal ontological motion sensing system of the present invention.

[0029] Figure 3 This is a flowchart of the multi-dimensional ontology cognition assessment process of the present invention;

[0030] Figure 4 This is a flowchart of the offline / online ontology cognition training process of the present invention; Detailed Implementation

[0031] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.

[0032] Embodied agents, possessing perception, reasoning, and movement capabilities based on embodied interaction, have the potential for autonomous and stable movement in open, unstructured environments, much like humans. These challenging scenarios demand higher levels of adaptability and intelligence from the agent. The agent needs to handle real-time physical interactions between its arbitrary position and the external environment during autonomous movement to avoid catastrophic movement failures in operational scenarios. Currently, to effectively represent the agent's topological structure to meet operational and planning needs, the relationships between moving parts are represented by DH matrices after being abstracted into virtual links. By constructing DH matrix chains (virtual moving links and joints), descriptions of the agent's positional and motion spaces can be quickly built. However, agent ontology representation methods based on abstract links struggle to represent the agent's complex geometric configuration and multidimensional physical characteristics. Characterizing and detecting the embodied interactions between the body at any position and the external environment and applying them to the agent's motion control remains a significant challenge.

[0033] Currently, the main solution to the problem of safe interaction management for intelligent agents is based on sensing devices (such as electronic skin, joint torque sensors, and motor current feedback) to detect interactive collisions. Appropriate compliant control methods (impedance control, admittance control, etc.) are designed to control the contact force within a safe threshold range to avoid catastrophic motion failures. However, this compliant motion control method based on virtual levers lacks consistency with task planning, making it difficult to coordinate the safety of the agent's actions with the effectiveness of task execution.

[0034] In recent years, with the improvement of computing power, machine learning methods, especially deep learning, have shown outstanding advantages in simulating multi-factor nonlinear strongly coupled problems due to their powerful nonlinear fitting capabilities. Deep neural networks implicitly represent the multi-dimensional geometric and physical features of an intelligent agent's ontology in a high-dimensional space, enabling a more detailed and direct digital representation of the agent's ontology, thus providing the possibility for achieving full-body motion control and management. Furthermore, the intelligent agent continuously observes and understands its body movements from its own perspective during motion, eliminating reliance on prior knowledge in motion control and allowing for timely adaptation to changes in the body's activity diagram (component damage, actuator replacement, etc.). Unlike ontology representation methods based on abstract symbols, the intelligent agent, based on physical reality, continuously recognizes the multi-dimensional motion features of its body from its own perspective to build an intuitive ontology model, which will significantly alleviate the challenges of full-body interaction planning for intelligent agents.

[0035] Example 1

[0036] In this embodiment, an intelligent agent multi-dimensional ontology incremental cognitive system is applied, the system structure of which is as follows: Figure 1As shown, the system includes an ontology motion perception module, an ontology cognition evaluation module, a memory pool module, an ontology knowledge mining module, and an ontology cognition training module. The ontology motion perception module collects and processes ontology perception data to obtain multimodal data. The ontology cognition evaluation module uses multimodal data to invert ontology activities in the digital space and compares the inversion results with ontology cognition results to obtain ontology cognition quantification error. The memory pool module stores data. The ontology knowledge mining module uses multimodal data to mine the intrinsic relationship between the agent's perception modalities and ontology representation dimensions, establishing an ontology knowledge base as the basis for data extraction and training in the ontology cognition training module. The ontology cognition training module constructs an ontology cognition model and trains it under the guidance of the ontology cognition quantification error and the ontology knowledge base, mapping multimodal data to a high-dimensional space to enable the agent to cognize the multidimensional characteristics of the ontology.

[0037] In this embodiment, by establishing a bidirectional ontology cognition mapping between joint space and motion space, the intelligent agent can predict the form of the ontology in motion space based on joint movements. At the same time, based on the target space state of the ontology, the intelligent agent can generate joint movements that conform to the intent, thereby enabling the management of embodied interactions at any position of the body.

[0038] Furthermore, by temporally monitoring and evaluating the deviation between ontological action facts and ontological cognitive networks, a continuous stream of multimodal ontological perception data can be obtained for ontological cognition, thus providing temporal increments.

[0039] Simultaneously, by mining high-level ontology knowledge, the ontology cognitive network can selectively extract data from relevant modalities for specialized training, improving training efficiency and reducing workload. The agent's ontology understanding matures and becomes more comprehensive through continuous ontology activities.

[0040] In this embodiment, the ontology motion perception module is used to collect and process perception data, and distinguish the agent from the external environment through ontology recognition; the ontology cognition evaluation module is used to reproduce the ontology's motion in digital space based on observed multimodal sensing signals, and quantify the error between the ontology's cognized actions and the actual actions; the memory pool system module is used to collect the robot's motion morphology facts and ontology motion cognition result data, providing a data source for the training of the agent's ontology cognition model and the multidimensional ontology cognition evaluation module; the ontology cognition cultivation module is used to implicitly construct a neural network model that maps from joint space to motion space, mapping multimodal data to a high-dimensional space to enable the agent to cognize the multidimensional characteristics of the ontology; the ontology knowledge mining module is used to mine the intrinsic relationship between the agent's perception modalities and ontology representation dimensions, providing a basis for the extraction of training data and model training for the ontology cognition cultivation module.

[0041] In this embodiment, the proprioception module collects visual images, proprioceptive sensations, and tactile data in real time based on the action task and constructs image / tactile-proprioceptive pairs. After the action task is completed, the collected data is encapsulated and sent to the memory pool module for limb motion data storage.

[0042] The multi-dimensional ontology cognition assessment module extracts multimodal perception data and ontology action cognition result data from the memory pool system for quantitative analysis of ontology cognition bias, and discretizes the agent to represent it as having n The tensor of a set of matter points, followed by the error obtained from quantization. It is used as the basis for data storage in the transmission memory pool system.

[0043] In this embodiment, the memory pool system receives multimodal ontological motion perception data collected by the multimodal ontological motion perception system and ontological motion results pre-rendered by the ontological cognition training module, and then stores both in the short-term memory module for use by the multidimensional ontological cognition evaluation module for data retrieval; simultaneously, it receives the quantification results of ontological cognition deviation from the multidimensional ontological cognition evaluation module, and when the error... Exceeding the tolerance of ontology cognition When the time is right, the memory pool system will migrate the corresponding multimodal ontology action data from the short-term memory module to the long-term memory module; otherwise, it will delete the data from the short-term memory module.

[0044] In this embodiment, the ontology cognition training module is a bidirectional ontology cognition network based on a transformer network architecture, connecting joint space to action space. The module receives action sequences or predicted ontology actions or action commands from the planner's target state output. This output is transmitted to a memory pool system for short-term memory storage. Simultaneously, the predicted spatial state is sent to the planner, and the predicted action commands are sent to the agent's controller. During the training phase, based on the ontology cognition's deviation in a specific dimension and the modality-dimensional mapping learned by the ontology knowledge mining module 150, matching modality data is extracted from the memory pool system for online training.

[0045] In this embodiment, the ontology knowledge mining module obtains training data from the memory pool system and constructs a modality-relationship-dimension basic knowledge graph unit based on a bottom-up approach. Subsequently, the obtained high-level ontology knowledge is transmitted to the ontology cognition cultivation module for online training according to the request sending mode.

[0046] In this embodiment, the multimodal body motion sensing system is as follows: Figure 2 As shown, the specific process is as follows:

[0047] The RGB-D visual perception unit collects two-dimensional images and depth information of the agent's own motion process from its own perspective. The collected signals are arranged into a sequence according to time order and sent to the host computer's visual perception queue through the communication link using the right-in-right-out principle.

[0048] The embodied interaction sensing unit collects signals including four modes: temperature, vibration, proximity, and pressure. This unit is a distributed tactile sensing array. After the tactile sensing unit encapsulates the data, it can transmit signals with other units. Multiple units form a tactile sensing group. During the movement of the unit robot, multimodal signals of the agent interacting with the outside world are collected according to the system's specific sampling frequency. Each tactile sensing group packages the signals and sends them to the embodied interaction sensing queue of the host computer according to the right-in, right-out principle.

[0049] The proprioception unit collects the angular displacement, angular velocity and joint torque of the joints during the movement of the intelligent agent to complete the proprioception signal observation. This unit consists of joint encoders and joint torque sensors distributed in each joint of the intelligent agent. After the collected signals are encapsulated into a unified data format, they are transmitted to the host computer's proprioception queue through the communication link using the right-in-right-out principle.

[0050] The signal uniform calibration unit uses a time consistency calibration algorithm, taking the embodied interaction perception signal as the time reference, to extract time samples at corresponding moments from the visual perception queue, the embodied interaction perception queue and the proprioception queue respectively. The aligned data is grouped by action to construct visual / tactile-proprioception pairs, forming multimodal limb action data packets, which are then sent to the memory pool system.

[0051] In this embodiment, the specific implementation of the multi-dimensional ontology cognition assessment module is as follows: Figure 3 As shown, the specific steps are as follows:

[0052] The body motion inversion unit is a deep neural network for robot body motion recognition. This network discretizes the motion state of the intelligent agent into a finite set of material points, extracts visual / tactile proprioception from the memory pool system as input to predict the spatial state of these material point sets, and realizes the reproduction of physical reality limb movements in digital space.

[0053] The cognitive bias assessment unit takes the spatial state of the set of material points obtained from the limb movement inversion module as the motion fact, and uses the mean squared error method to quantify the error in predicting the cognitive results of proprioception.

[0054] The data interaction unit establishes a communication link between the multi-dimensional ontology cognition evaluation module and the memory pool system based on the communication method of sending requests. After the agent completes the action task, the data interaction module requests the visual / tactile-proprio perception pair and ontology cognition result action data from the memory pool system, and sends the received action data to the ontology action inversion module. At the same time, when the cognition deviation module obtains the quantification error, it sends the quantification result to the memory pool system.

[0055] In this embodiment, the memory pool system is implemented as follows:

[0056] The short-term memory unit stores and forwards multimodal data generated by actions performed in a short period of time, receives limb activity data generated by the multimodal propriokinetic motion perception system and the propriokinetic cognition training module, and sends the stored limb activity data according to the request of the multidimensional propriokinetic cognition assessment module.

[0057] The long-term memory module dynamically stores limb activity data that exceeds the tolerance of proprioceptive bias, providing training data for the proprioceptive cognition model. By receiving the proprioceptive cognition quantification error and limb activity labels from the cognitive bias module, the long-term memory module retrieves limb activity data from the short-term memory module. When the error exceeds the tolerance range, the data is stored in this module; otherwise, the data is deleted.

[0058] In this embodiment, the specific implementation of the ontology cognition training module is as follows: Figure 4 As shown, the specific process is as follows:

[0059] First, the model is built. The entire network structure consists of five parts: forward encoder, inverse encoder, backbone network, forward decoder, and inverse decoder. The forward encoder includes: a visual encoder module based on a ResNet50 pre-trained network; a proprioceptive encoder module based on embedding and MLP network structures; and a tactile encoder module based on MLP network technology. The inverse encoder module is an agent ontology point set encoding module built from an MLP network. The backbone network is a multi-layer deep neural network based on a Transformer structure, which can map multimodal information to a high-dimensional space to implicitly represent the multi-dimensional motion features of the agent ontology. The forward decoder combines a Transformer structure and an MLP structure to output the spatial state of the agent ontology point set. The inverse decoder, built from an MLP, outputs the predicted actions in joint space.

[0060] Then, the model is trained offline using the pre-collected multimodal dataset. The dataset is divided into a training set (80%) and a test set (20%), and training is performed based on the MSE loss function.

[0061] Next, the model is deployed, and the trained model is migrated to the intelligent agent computing system.

[0062] Finally, the model enters the online training phase. After receiving the model training task from the operator, the agent's ontology cognition model retrieves the ontology cognition quantification error from the memory pool system to determine the ontology cognition dimension to be trained. Using a request-response communication mode, it obtains the perceptual modality to be trained from the ontology knowledge mining module and retrieves the corresponding modality's training data from the memory pool system for supervised training. The corresponding online training task is completed by evaluating the loss result after model training.

[0063] In this embodiment, the ontology knowledge mining module is implemented as follows:

[0064] First, the ontology elements are materialized, and the perceptual modality and ontology representation dimension are materialized into independent entity nodes respectively;

[0065] Then, structured data is constructed, and the time-series data is differentiated to obtain multimodal ontology motion data with gradient characteristics. The data is then stored based on entity nodes.

[0066] Finally, ontology knowledge reasoning is performed. Using entity data with gradient features, and based on Bayes' theorem, a high-level ontology knowledge base of entity-weight-entity is constructed.

[0067] In summary, this system establishes and trains an incremental ontological cognitive model through a multi-dimensional ontological incremental cognitive system for intelligent agents, consisting of an ontological motion perception module, an ontological cognition evaluation module, a memory pool module, an ontological knowledge mining module, and an ontological cognition cultivation module. This ontological cognitive model endows the intelligent agent with the ability to model multi-dimensional motion features of the whole body. Based on this ability, the intelligent agent cognitive model can represent and perceive any position of the body, enabling the intelligent agent to accurately perceive and represent multi-dimensional ontological features, thereby effectively planning the physical interaction between any position of the body and the external environment.

[0068] Example 2

[0069] In this embodiment, a multi-dimensional ontology incremental cognition method for intelligent agents is applied. This method is applied to a system containing an ontology motion perception module, an ontology cognition evaluation module, a memory pool module, an ontology knowledge mining module, and an ontology cognition cultivation module. The method utilizes the memory pool module for data storage and includes the following steps:

[0070] S1. Collect and process the body's perception data using the body motion perception module to obtain multimodal data;

[0071] S2. Using the ontology cognition training module, the neural network model is initially trained based on multimodal data to establish an ontology cognition model, and the ontology cognition model is used to process multimodal data to obtain ontology cognition results.

[0072] S3. Using the ontology cognition assessment module, ontology activities are inverted in the digital space based on multimodal data, and the inversion results are compared with ontology cognition results to obtain ontology cognition quantification error.

[0073] S4. Use the ontology knowledge mining module to process multimodal data and obtain an ontology knowledge base;

[0074] S5. Quantify the error of ontology cognition to clarify the target training ontology cognition dimension, clarify the target training perception modality from the knowledge base, train the ontology cognition dimension and the target training perception modality based on the target, and use the ontology cognition cultivation module to train the ontology cognition model based on multimodal data.

[0075] S6. The trained ontology cognition model processes multimodal data to output ontology cognition results and returns to execute step S4 to realize multidimensional ontology incremental cognition of the agent.

[0076] In this method, the process of obtaining the ontology cognition quantification error in S3 is as follows: First, the motion state of the agent in the ontology cognition model is discretized into a finite number of material point sets. Then, visual / tactile proprioception is extracted from the memory pool module as input to predict the spatial state of these material point sets, so as to realize the reproduction of physical reality's limb movements in digital space. Finally, the error of the ontology cognition result is quantified by the mean square error method as the ontology cognition quantification error.

[0077] In this method, the specific process of using the ontology knowledge mining module in S4 to process multimodal data and obtain the knowledge base is as follows: First, the perceptual modality and ontology representation are materialized to obtain entity nodes. Then, the multimodal data is differentiated to obtain multimodal data with gradient characteristics. Then, the multimodal data with gradient characteristics is fused into the entity nodes to obtain entity data. Based on Bayes' theorem, the ontology knowledge base is constructed using the entity data.

[0078] In this method, the training process of the ontology cognition model in S5 uses MSE as the loss function.

[0079] In this embodiment, a proprioception module is used to collect and process sensory data to distinguish the intelligent agent from the external environment. The proprioception module includes multiple types of sensing units and devices, including: an RGB-D visual sensing unit for geometric and image perception of the embodied intelligent agent, enabling the distinction between the agent's motion and the external environment; an embodied interaction sensing unit for collecting object signals at the interactive interface and observing the external physical interaction stimuli received by the intelligent agent during its movement; a proprioception unit for observing the position, velocity, and torque signals in the joint space of the intelligent agent during its movement, completing the detection of proprio power unit commands and states; and a signal uniform calibration unit for performing time-uniform registration of the collected multimodal proprio observation signals to unify the sampling frequency and start and end times of the multimodal sensing signals.

[0080] In this embodiment, the ontology cognition assessment module utilizes observed multimodal sensor signals to reproduce the ontology's motion in digital space, thereby quantifying the error between the ontology cognition result and the actual action. This multidimensional ontology cognition assessment module includes several sub-modules, including: an ontology action inversion module: reproducing physical limb movements in digital space and converting multimodal information into action-state tensor pairs; a cognitive bias slice module: using the inverted action as a reference for the real action to calculate the agent's cognitive error regarding the ontology cognition result; and a data interaction module: extracting multimodal perception data and ontology cognition result data for ontology cognitive bias quantification and sending the calculation results.

[0081] In this embodiment, a memory pool module is used to store the action form facts of the intelligent agent, providing a data source for the training of the agent's ontology cognition model and the multi-dimensional ontology cognition evaluation. The memory pool system includes two subsystems: a short-term memory module, which stores the temporal multimodal data of the robot's past actions in a short period of time, providing a quantitative data source for the multi-dimensional ontology cognition evaluation module; and a long-term memory module, which stores limb activity data that exceeds the tolerance of proprioceptive bias, providing training data for the ontology cognition model training.

[0082] In this embodiment, an ontology cognition training module implicitly constructs a neural network model that maps from joint space to action space for the physical entity of the intelligent agent. This maps multimodal data to a high-dimensional space, enabling the intelligent agent to recognize the multidimensional characteristics of the ontology. The ontology cognition training module includes both offline and online modes. The offline mode is for batch training of large amounts of perceptual data, mainly targeting the implicit representation training stage of appearance geometry. The online mode targets the implicit representation training stage of the multimodal high-dimensional space of the physical entity and makes real-time adjustments to key representation points.

[0083] In this embodiment, the ontology knowledge mining module is used to explore the intrinsic relationship between the agent's perceptual modality and ontology representation dimension, which serves as the basis for data extraction and training of the ontology cognition cultivation module.

[0084] In this embodiment, the multimodal body motion sensing system is as follows: Figure 2 As shown, the specific process is as follows:

[0085] The RGB-D visual perception unit collects two-dimensional images and depth information of the agent's movement from its own perspective. The collected signals are arranged in chronological order and sent to the host computer's visual perception queue via a right-in, right-out communication link. The embodied interaction perception unit collects signals including four modalities: temperature, vibration, proximity, and pressure. This unit is a distributed tactile perception array. The tactile perception units encapsulate data for signal transmission with other units, and multiple units form a tactile perception group. During the robot's movement, multimodal signals from the agent's interaction with the outside world are collected at a specific sampling frequency. Each tactile perception group packages the signals and sends them to the host computer's embodied interaction perception unit using a right-in, right-out principle. In the queue, the proprioception unit collects the angular displacement, angular velocity, and joint torque of the joints during the movement of the intelligent agent to complete the proprioception signal observation. This unit consists of joint encoders and joint torque sensors distributed in each joint of the intelligent agent. After encapsulating the collected signals into a unified data format, they are transmitted to the host computer's proprioception queue through the communication link using the right-in, right-out principle. The signal uniform calibration unit uses a time consistency calibration algorithm, with the embodied interaction perception signal as the time reference, to extract the corresponding time samples from the visual perception queue, embodied interaction perception queue, and proprioception queue respectively. The aligned data is grouped by action to construct visual / tactile-proprioception pairs, forming multimodal limb action data packets, which are then sent to the memory pool system.

[0086] In this embodiment, the specific implementation of the multi-dimensional ontology cognition assessment module is as follows: Figure 3 As shown, the specific steps are as follows:

[0087] The ontology motion inversion unit is a deep neural network used for robot ontology motion recognition. This network discretizes the agent's motion state into a finite set of material points, extracts visual / tactile proprioceptive sensations from the memory pool system as input, and predicts the spatial state of these material point sets, thus reproducing physical limb movements in digital space. The cognitive bias evaluation unit takes the spatial state of the material point sets obtained from the ontology motion inversion module as the motion fact, and quantifies the error in predicting the ontology motion cognitive result using the mean square error method. The data interaction unit establishes a communication link between the multi-dimensional ontology cognition evaluation module and the memory pool system based on the communication method of sending requests. After the agent completes the action task, the data interaction module requests the visual / tactile-proprio perception pair and ontology cognition result action data from the memory pool system, and sends the received action data to the ontology action inversion module. At the same time, when the cognition deviation module obtains the quantification error, it sends the quantification result to the memory pool system.

[0088] In this embodiment, the memory pool system is implemented as follows:

[0089] The short-term memory unit stores and forwards multimodal data generated by actions performed within a short period of time. It receives limb activity data generated by the multimodal proprioception system and the proprioception training module, and sends the stored limb activity data according to the request of the multidimensional proprioception assessment module. The long-term memory module dynamically stores limb activity data that exceeds the tolerance of proprioceptive bias, providing training data for proprioception model training. By receiving the proprioception quantification error and limb activity labels obtained from the cognitive bias module, the long-term memory module retrieves limb activity data from the short-term memory module. When the error exceeds the tolerance range, the data is stored in this module; otherwise, the data is deleted.

[0090] In this embodiment, the specific implementation of the ontology cognition training module is as follows: Figure 4 As shown, the specific process is as follows:

[0091] First, the model is built. The entire network structure consists of five parts: forward encoder, inverse encoder, backbone network, forward decoder, and inverse decoder. The forward encoder includes a visual encoder module based on a ResNet50 pre-trained network, a proprioceptive encoder module based on embedding and MLP network structures, and a tactile encoder module based on MLP network technology. The inverse encoder module is an agent ontology point set encoding module built from an MLP network. The backbone network is a multi-layer deep neural network based on a Transformer structure, which can map multimodal information to a high-dimensional space to implicitly represent the multi-dimensional motion features of the agent ontology. The forward decoder combines a Transformer structure and an MLP structure to output the spatial state of the agent ontology point set. The inverse decoder, built from an MLP, outputs the predicted actions in joint space. Then, the model is trained offline using pre-collected multimodal datasets. The model is divided into a training set (80%) and a test set (20%), and trained using the MSE loss function. The model is then deployed to the agent computing system. Finally, in the online training phase, the agent ontology cognitive model receives the training task from the operator, retrieves the ontology cognitive quantification error from the memory pool to determine the ontology cognitive dimension to be trained, and uses a request-response communication mode to obtain the perceptual modality to be trained from the ontology knowledge mining module. It then retrieves the corresponding modality's training data from the memory pool for supervised training, and completes the corresponding online training task by evaluating the loss result after model training.

[0092] In this embodiment, the ontology knowledge mining module is implemented as follows:

[0093] First, ontology elements are materialized, with the perceptual modality and ontology representation dimension materialized as independent entity nodes. Then, structured data is constructed by differentiating time-series data to obtain multimodal ontology motion data with gradient characteristics, and the data is stored based on entity nodes. Finally, ontology knowledge reasoning is performed, using entity data with gradient features and based on Bayes' theorem, to construct a high-level ontology knowledge base of entity-weight-entity.

[0094] In summary, this method enables agents to accurately represent and perceive the multi-dimensional features of an ontology. By establishing a bidirectional ontology cognition mapping between joint space and motion space, the agent can accurately predict the ontology's form in motion space based on joint movements. Simultaneously, based on the ontology's target space state, the agent can generate joint movements that conform to the intent, thereby managing embodied interactions at any position of the body.

[0095] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in the present invention, and these modifications or substitutions should all be covered within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.

Claims

1. A multi-dimensional ontology incremental cognitive system for intelligent agents, characterized in that, The system includes an ontology motion perception module, an ontology cognition assessment module, a memory pool module, an ontology knowledge mining module, and an ontology cognition cultivation module; The body motion sensing module is used to collect and process the body's sensing data to obtain multimodal data; The ontology cognition assessment module is used to invert ontology activities in the digital space using multimodal data, and compare the inversion results with ontology cognition results to obtain ontology cognition quantification error; The memory pool module is used for data storage; The ontology knowledge mining module is used to mine the intrinsic relationship between the agent's perceptual modality and ontology representation dimensions using multimodal data, so as to establish an ontology knowledge base as the basis for data extraction and training of the ontology cognition cultivation module; the ontology representation dimensions include the agent's geometric configuration and physical characteristics. The ontology cognition training module is used to construct an ontology cognition model and train the ontology cognition model under the guidance of ontology cognition quantification error and ontology knowledge base, so as to map multimodal data to a high-dimensional space to enable the agent to recognize the multidimensional characteristics of the ontology. The ontology cognition training module includes a forward encoding unit, a backward encoding unit, a backbone network unit, a forward decoding unit, and a backward decoding unit. The forward encoding unit includes a visual encoding network built on a ResNet50 pre-trained network, an ontology perception encoding network built on an embedding and MLP network structure, and a tactile encoder network built on MLP network technology. The backward encoding unit includes an agent ontology material point set encoding network built on an MLP network. The backbone network includes a multi-layer deep neural network based on a Transformer structure, used to map multimodal information to a high-dimensional space to implicitly represent the multi-dimensional motion features of the agent ontology. The forward decoder includes a Transformer and MLP combined network, used to output the spatial state of the agent ontology material point set. The backward decoder includes an MLP network, used to output predicted actions in joint space.

2. The multi-dimensional ontology incremental cognitive system for intelligent agents according to claim 1, characterized in that, The aforementioned proprioception module includes an RGB-D visual perception unit, an embodied interaction perception unit, and a proprioception unit. The RGB-D visual perception unit is used to collect two-dimensional images and depth information of the proprioception process observed by the agent from its own perspective and transmit them to the visual perception queue. The embodied interaction perception unit is used to collect four modes: temperature, vibration, proximity, and pressure and transmit them to the depth interaction perception queue. The proprioceptive unit is used to collect the angular displacement, angular velocity and joint torque of the joints during the movement of the intelligent agent and transmit them to the proprioceptive queue.

3. The intelligent agent multi-dimensional ontology incremental cognitive system according to claim 2, characterized in that, The proprioception module also includes a signal uniform calibration unit. This unit is used to use a time-series consistency calibration algorithm to extract time samples at corresponding moments from the visual perception queue, the embodied interaction perception queue and the proprioception queue, respectively, using the embodied interaction perception signal as the time reference. The aligned data is then grouped by action to construct visual / tactile-proprioception pairs.

4. The multi-dimensional ontology incremental cognitive system for intelligent agents according to claim 1, characterized in that, The ontology cognition assessment module includes an ontology action inversion unit, a cognitive bias unit, and a data interaction unit. The ontology action inversion unit reproduces the agent's real-world limb movements in digital space, converting multimodal information into action-state tensor pairs. The cognitive bias unit uses the inverted actions as a reference for the real actions to calculate the ontology cognition quantization error. The data interaction unit extracts multimodal data and ontology cognition result data for ontology cognition bias quantization and sends the ontology cognition quantization error.

5. The multi-dimensional ontology incremental cognitive system for intelligent agents according to claim 1, characterized in that, The memory pool module includes a short-term memory unit and a long-term memory unit. The short-term memory unit is used to store multimodal data acquired by the agent in the current time sequence, providing a quantitative data source for the multidimensional ontology cognition evaluation module. The long-term memory unit is used to store multimodal data that exceeds the preset tolerance of ontological perception bias, providing training data for the training of the ontology cognition model.

6. A multi-dimensional ontology incremental cognition method for intelligent agents, characterized in that, The method is applied to an agent-based multi-dimensional ontology incremental cognitive system as described in any one of claims 1-5. The method utilizes a memory pool module for data storage and includes the following steps: S1. Collect and process the body's perception data using the body motion perception module to obtain multimodal data; S2. Using the ontology cognition training module, the neural network model is initially trained based on multimodal data to establish an ontology cognition model, and the ontology cognition model is used to process multimodal data to obtain ontology cognition results. S3. Using the ontology cognition assessment module, ontology activities are inverted in the digital space based on multimodal data, and the inversion results are compared with ontology cognition results to obtain ontology cognition quantification error. S4. Use the ontology knowledge mining module to process multimodal data and obtain an ontology knowledge base; S5. Quantify the error of ontology cognition to clarify the target training ontology cognition dimension, clarify the target training perception modality from the knowledge base, train the ontology cognition dimension and the target training perception modality based on the target, and use the ontology cognition cultivation module to train the ontology cognition model based on multimodal data. S6. The trained ontology cognition model processes multimodal data to output ontology cognition results and returns to execute step S4 to realize multidimensional ontology incremental cognition of the agent.

7. The multi-dimensional ontology incremental cognition method for intelligent agents according to claim 6, characterized in that, The process of obtaining the ontology cognition quantification error in S3 is as follows: First, the motion state of the agent in the ontology cognition model is discretized into a finite number of material point sets. Then, visual / tactile proprioception is extracted from the memory pool module as input to predict the spatial state of these material point sets, so as to realize the reproduction of physical reality's limb movements in digital space. Finally, the error of the ontology cognition result is quantified by the mean square error method as the ontology cognition quantification error.

8. The multi-dimensional ontology incremental cognition method for intelligent agents according to claim 6, characterized in that, The specific process of using the ontology knowledge mining module in S4 to process multimodal data and obtain the knowledge base is as follows: First, the perceptual modality and ontology representation are materialized to obtain entity nodes. Then, the multimodal data is differentiated to obtain multimodal data with gradient characteristics. Then, the multimodal data with gradient characteristics is fused into the entity nodes to obtain entity data. Based on Bayes' theorem, the ontology knowledge base is constructed using the entity data.

9. The multi-dimensional ontology incremental cognition method for intelligent agents according to claim 6, characterized in that, The training process for training the ontology cognitive model in S5 uses MSE as the loss function.

Citation Information

Patent Citations

  • Intelligent agent control method and device, equipment and storage medium

    CN117518907A

  • Multi-task cognitive brain-inspired modeling method

    WO2024103345A1