General-purpose intelligent agent and control method therefor

By designing an active general intelligent agent and utilizing transparent and interpretable abstract thinking and a data-driven model, the fragility and uncontrollability of existing AI technologies are solved, enabling autonomous decision-making and widespread application of general artificial intelligence.

WO2025214260A1PCT designated stage Publication Date: 2025-10-16CHENGDU YUANJI TONGZHI TECHNOLOGY CO LTD

Patent Information

Application Number
PCT/CN2025/087245
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-04-10
Filing Date
2025-04-03
Publication Date
2025-10-16

AI Technical Summary

Technical Problem

Existing AI technologies suffer from fragility, uncontrollability, lack of interpretability, and high resource consumption, and their application scope is narrow, making it difficult to achieve general artificial intelligence.

Method used

Design an active general intelligent agent, which includes an input module, a consciousness module, a self-awareness module, a subconscious module and an information exchange module. It perceives and recognizes through transparent and explainable abstract thinking forms, uses data-driven models to realize subconscious functions, and combines a universal language for information exchange and output.

Benefits of technology

It realizes a transparent and controllable intelligent agent that makes autonomous decisions, learns independently, and continuously evolves. It can understand the physical world and create the cultural world, reduces the demand for data and computing resources for end-to-end processing, and has broad versatility.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025087245_16102025_PF_FP_ABST
    Figure CN2025087245_16102025_PF_FP_ABST
Patent Text Reader

Abstract

Disclosed in the present invention is a general-purpose intelligent agent. The general-purpose intelligent agent comprises: an input module, which is configured to acquire a preprocessed input signal; a consciousness module, which is configured to perform perception and cognitive operations in a transparent and interpretable abstract thinking form; a self-awareness module, which is configured for perception and cognition of an intelligent agent body; a subconsciousness module, which is configured to use a data-driven model to implement perception and cognitive functions at a subconscious level; an information exchange module, which is configured to perform mutual exchange between perception, cognition and other information generated by the consciousness module, the self-awareness module and the subconsciousness module, thereby enabling integration and utilization; and an output module, which is configured to output a processing result on demand. In a machine intelligent agent of the present invention, a machine uses a general-purpose language as an underlying language for thinking and interaction, a thinking process and result are completely transparent, and thus autonomous decision-making, autonomous learning and continuous evolution are realized under fully controllable conditions. Further disclosed is a method for constructing a general-purpose intelligent agent and a general-purpose language. Compared with the prior art in which a natural language is used, the machine intelligent agent based on the general-purpose language has wide versatility in application.
Need to check novelty before this filing date? Find Prior Art

Description

A general agent and a control method thereof TECHNICAL FIELD

[0001] The present application belongs to the field of artificial intelligence, and particularly relates to a general agent and a control method thereof. BACKGROUND

[0002] Since its inception, the research of artificial intelligence (AI) has roughly three schools: symbolism, connectionism and behaviorism. Symbolism mainly uses a large number of if-else and other symbolic logic reasoning to realize expert systems and knowledge representation, which has the advantages of clear logical rules and easy interpretability, but its limitations are that it is disconnected from the physical world and it is difficult to cover everything, falling into an endless process of manually adding symbolic rules; Connectionism advocates realizing AI by simulating the interconnection and weight between neurons, simplifying the causal logic between things into data correlation, which has good end-to-end processing information capability, but its disadvantage is that the training of the network requires a large amount of time and computing resources, and lacks interpretability; Behaviorism represented by reinforcement learning lets an agent constantly take different actions and interacts with the environment to obtain different rewards to learn appropriate strategies through continuous trial and error, which has the advantage of being able to process real-time environmental information, but it can only be used for specific agents trained in specific environments, and the application range is narrow.

[0003] Symbolism, connectionism and behaviorism show different aspects of intelligence from different angles, and each has been applied in the corresponding field. Although their implementation paths are different, they all imply a common assumption: intelligence is hidden in data and rules, and as long as there are as many symbolic rules and multi-modal data as possible for the machine to explore, it can approach artificial general intelligence (AGI). Therefore, these technologies can be collectively referred to as passive AI, which develops machines that exhibit intelligent behavior (which can be called "intelligent illusion") on the basis of the vast amount of data and specific rule representations generated by human intelligence activities, through probability statistical analysis and manual checking and filling. SUMMARY

[0004] In view of the problems in the prior art, the present application provides an active general agent and a control method thereof.

[0005] To solve the above technical problems, the present application is realized by the following way:

[0006] A general agent, comprising:

[0007] 1) an input module for obtaining pre-processed input signals;

[0008] 2) a consciousness module for perception and cognition operations in a transparent and interpretable abstract thinking form;

[0009] 3) a self-awareness module for perception and cognition of the agent itself;

[0010] 4) a subconscious module for implementing perception and cognition functions under the subconscious using a data-driven model;

[0011] 5) an information exchange module for exchanging and fusing the perception and cognition information generated by the consciousness module, the self-awareness module, and the subconscious module;

[0012] 6) an output module for outputting the processing results as needed;

[0013] The input module is connected to the output module through an information processing module, which includes two relatively independent groups of modules, one group including the consciousness module and the self-awareness module, and the other group being the subconscious module. The information processing module processes the input signals obtained by the input module, and the processing results are sent to the output module. The information exchange module is responsible for information exchange and fusion between the two relatively independent groups of modules.

[0014] Further, the signal preprocessing of the input module specifically includes:

[0015] The signals include received general language statements, natural language statements, codes, agent memory bank data, positioning tracking, vision, hearing, touch, smell, taste, and other types of modal signals received by the body sensing sensor and other external sensors; the preprocessing includes various common signal preprocessing methods such as sampling, changing image resolution, signal enhancement, denoising, and filtering, as well as various common methods for converting signals from signal space to feature space through linear or nonlinear transformation for better feature dimensionality reduction or clustering in the transformed domain.

[0016] Further, the consciousness module includes:

[0017] 21) a feeling sub-module for converting the input signals obtained by the input module into clues through bottom-level intuitive processing;

[0018] 22) a perception sub-module for arousing relevant concepts and perception logic using the clues, and constructing objects and scenes in a new perspective and posture in the construction space;

[0019] 23) a cognition sub-module for arousing relevant logic using the constructed object and scene conditions and task parts, generating judgments, reasoning, or planning in the construction space to obtain rational cognition;

[0020] 24) Output modal conversion sub-module, convert the general language form of thinking into natural language text, audio, image, video and other modal signals as needed;

[0021] 25) Long-term memory sub-module, store the long-term memory data of the agent, including concept library, logic library, ontology model library, etc.

[0022] As a preferred mode, the processing flow of the sensing sub-module includes:

[0023] 211) Call the input signal obtained by the input module, and obtain the attention preset parameters, including the focus point, the attention type and the value range, the minimum attention granularity, and other intuitive process-related calculation parameters;

[0024] 212) According to the attention preset parameters, use the region segmentation algorithm to select the plane or space region of interest in the signal digital space;

[0025] 213) Structure the selected region, convert it into a meta-point set composed of basic elements of meta-points, and store it as a clue in the construction space.

[0026] Further, the meta-point in step 213) is the basic constituent element of the machine thinking space, and the meta-point set is represented as follows:

[0027]

[0028] Where, P i represents the physical attribute, E i represents the extended attribute, V i represents the value attribute, R i represents the reference attribute, O i represents other attributes; In the clue extraction stage, the meta-point adopts the skeleton line extraction or the detection and extraction method of various key points and key lines in image processing such as corner point, centroid and edge, or adopts various methods such as point cloud estimation trained by neural network, Gaussian distribution estimation for combination and simplification, and converts the extraction result into structured data; The meta-point set is converted into a corresponding clue graph, and stored in the construction space as a visual auxiliary image of the corresponding clue.

[0029] As a preferred mode, the processing flow of the perception sub-module includes:

[0030] 221) Call the clue obtained by the sensing sub-module;

[0031] 222) Use the clue to arouse related concepts in the concept library, and store them in the construction space;

[0032] 223) With the aroused concept as a template and the relevant clues as materials, the object is constructed in the construction space with a new perspective and posture, the existence and state of the object are perceived, and the relevant perceptual logic is aroused to construct the environment, conditions, and tasks, etc.

[0033] 224) The sensory sub-module is called again to find new clues to fill and update, support and enrich the constructed object and scene;

[0034] 225) The similarities and differences between the object, scene and the relevant concept in the concept library are compared, and a new concept is formed on the basis of the existing concept to expand the concept library;

[0035] 226) Steps 221) to 225) are performed multiple rounds to observe and analyze the scene and object more carefully.

[0036] As a preferred mode, the specific process in step 222) is to find one or more concepts in the concept library, which are the same or similar to the clues obtained in step 221) in one or more aspects of the data structure, and call the relevant concept from the concept library to the construction space;

[0037] The specific process of step 223) is that after arousing the concept, the concept is filled with all or part of the clue meta-point according to the matching relationship between the concept and the clue, the attribute of each meta-point in the concept is updated, and the object is constructed in the space-time coordinate system with a new perspective and posture, so that the existence and state of the object are perceived, and the relevant perceptual logic is aroused to construct the environment, conditions, and tasks, etc.

[0038] The new clues in step 224) refer to the clues found in the sensory sub-module, which pay attention to the clues or meta-points in the clues that are not used in the arousal process, or the clues found by changing the attention parameters and returning to the original input signal;

[0039] The specific process of step 225) includes: under the circumstances of the object and scene obtained in step 224), judging the similarities and differences between them and the relevant concepts in the concept library, and updating the new concept on the basis of the existing concept to form a new concept with simplified and highlighted difference features, and storing it in the concept library as a high-level concept category;

[0040] The specific process of step 226) includes: based on the clues, concepts, objects and scenes obtained in the construction space, multiple rounds of steps 221) to 225) are performed as the attention parameters change, the object and scene are constantly updated, the existing object and scene are observed in detail, and new objects and scenes are generated with the newly aroused concepts and logic after arousing more similar concepts and perceptual logic.

[0041] Further, the cognitive sub-module processing flow comprises:

[0042] 231) calling the perception sub-module to obtain objects and scenes;

[0043] 232) deepening the construction of conditions and task scenes, while arousing relevant logic, including directly arousing relevant logic in the logic library and indirectly extracting relevant logic from specific scene concepts involving judgment, reasoning or planning in the concept library, and storing in the construction space;

[0044] 233) generating judgment, reasoning or planning containing specific content in the construction space with the aroused logic as a template and relevant objects and scenes as materials, to obtain preliminary rational cognition;

[0045] 234) finding new objects or new scenes to fill and update, support and enrich the constructed judgment, reasoning or planning.

[0046] 235) forming judgment, reasoning or planning as a specific scene, forming new concepts on the basis of existing concepts to expand the concept library; at the same time, finding rules in judgment, reasoning or planning, and updating new logic on the basis of existing logic to expand the logic library;

[0047] 236) Steps 231) to 235) are performed multiple times to analyze and consider the task and generated judgment, reasoning or planning more carefully.

[0048] As a preferred mode, the specific process of step 233) is: after arousing the logic, according to the matching relationship between the logic and the objects and scenes, the meta-points in the logic are filled by the corresponding object and scene meta-point set, and the parameter values such as the extension attribute values in the logic meta-points are calculated, and the attributes of each meta-point in the logic are updated. In this way, judgment, reasoning or planning containing specific content is generated in the construction space to obtain preliminary rational cognition;

[0049] The new objects or new scenes in step 234) refer to the objects and scenes found, focusing on the objects and scenes that are not used in the logic arousing process, or calling the perception sub-module to return to the original input signal to change the attention parameters and find new clues, objects and scenes, and the meta-points in the judgment, reasoning or planning meta-point set still have no content filling. The relevant concepts are aroused by the concept library to generate new objects and scenes as filling content.

[0050] The flow of the output modality conversion sub-module specifically includes: according to the clue, object, scene, judgment, reasoning or planning content of the meta-point set or graph form generated by each sub-module in the construction space, and all or selected part of the content is directly output as a general language sentence; and according to the need, the general language form content is converted into natural language text, audio, image, video and other modal signals.

[0051] Further, the self-awareness module includes:

[0052] 31) A proprioceptive sub-module for receiving proprioceptive signals, including navigation positioning, time perception, vision, hearing, touch, smell, taste, inertia, balance and other types of proprioceptive signals on the body of the agent and each component;

[0053] 32) A proprioceptive perception sub-module for estimating model parameters using proprioceptive signals and proprioceptive models, and combining objects and scenes in the construction space to calculate the position and posture of the body of the agent including the trunk and each component relative to other objects in the scene, and correcting the proprioceptive model and updating the proprioceptive model library when the proprioceptive model is not applicable;

[0054] 33) A proprioceptive cognitive sub-module for filling the proprioceptive body as a content or element into the awakened logic if the proprioceptive body is needed in the judgment, reasoning or planning process, and participating in the judgment, reasoning or planning process;

[0055] 34) A proprioceptive action control sub-module for outputting the proprioceptive control signal as the action sequence, hardware control program and other information generated by the planning result when the proprioceptive action is needed;

[0056] 35) A proprioceptive model library for storing the structural parameter model of the trunk and each component of the agent in mathematics, physics, graph, motion and interaction.

[0057] As a preferred mode, the subconscious module includes:

[0058] 41) A subconscious feeling sub-module for implementing the feeling function under the subconscious using a data-driven model;

[0059] 42) A subconscious perception sub-module for implementing the perception function under the subconscious using a data-driven model;

[0060] 43) A subconscious cognitive sub-module for implementing the cognitive function under the subconscious using a data-driven model;

[0061] 44) A subconscious output modality conversion sub-module for implementing the multi-modal signal output function under the subconscious using a data-driven model.

[0062] As a preferred mode, the sub-module 41) processing flow specifically includes: using the trained model parameters, obtaining the model output signal according to the model input signal; the model input signal is the input signal obtained by the input module, and the model output signal is a general language sentence matched with the clue element point set or pattern generated by the feeling sub-module in the consciousness module;

[0063] The sub-module 42) processing flow specifically includes: using the trained model parameters, obtaining the model output signal according to the model input signal; the model input signal is respectively the input signal obtained by the input module or the general language sentence matched with the clue element point set or pattern generated by the feeling sub-module in the consciousness module, and the model output signal includes the general language sentence matched with the object, scene element point set or pattern generated by the perception sub-module in the consciousness module;

[0064] The sub-module 43) processing flow specifically includes: using the trained model parameters, obtaining the model output signal according to the model input signal; the model input signal is respectively the input signal obtained by the input module or the general language sentence matched with the clue element point set or pattern generated by the feeling sub-module in the consciousness module or the general language sentence matched with the object, scene element point set or pattern generated by the perception sub-module in the consciousness module, and the model output signal includes the general language sentence matched with the cognitive element point set or pattern generated by the cognitive sub-module in the consciousness module;

[0065] The sub-module 44) processing flow specifically includes: using the trained model parameters, obtaining the model output signal according to the model input signal; the model input signal is respectively the input signal obtained by the input module or the general language sentence matched with the clue element point set or pattern generated by the feeling sub-module in the consciousness module or the general language sentence matched with the object, scene element point set or pattern generated by the perception sub-module in the consciousness module or the general language sentence matched with the cognitive element point set or pattern generated by the cognitive sub-module in the consciousness module, and the model output signal includes the multi-modal data matched with the output signal of the output module.

[0066] As a preferred mode, the information exchange module processing flow includes:

[0067] 51) Add the input signal obtained by the input module and the feeling, perception, cognition, multi-modal output data generated by the feeling, perception, cognition, output modal conversion sub-modules in the consciousness module into the training data set for the sub-modules of the subconscious module to train the model;

[0068] 52) The feeling, perception, cognition, output modalities generated by each sub-module in the subconscious module are converted into feeling, perception, cognition, multi-modal output data, which are provided as reference data to each sub-module of the conscious module for perception, analysis, correction by each sub-module of the conscious module, and fusion with the data generated by the conscious module itself to obtain output results.

[0069] As a preferred mode, the output module processing flow includes:

[0070] According to the interaction and action requirements, the general language statements, natural language statements, codes, agent memory bank data, ontology control signals and other image, video, audio and other modal signals generated by each module are output.

[0071] In a second aspect, the present application also provides a general intelligent agent control method, which includes:

[0072] S1, memory preset and intelligent agent pre-training;

[0073] S2, value-driven task generation;

[0074] S3, task-oriented autonomous decision-making;

[0075] S4, task execution and self-evolution;

[0076] S5, multiple rounds of steps S2-S4 are performed to actively interact with the environment and the user.

[0077] As a preferred mode of the above method, the step S1 specifically includes: pre-setting the concept library, logic library and ontology model library in the long-term memory sub-module of the conscious module, and setting attribute values according to requirements; collecting model input data and expected model output data for the network model contained in each sub-module of the subconscious module, training the model, and updating model parameters;

[0078] The step S2 specifically includes: task analysis of the relevant task scenarios recalled in the memory bank according to the observed objects and scenes, and selection of appropriate tasks through value judgment;

[0079] The step S3 specifically includes: autonomous generation of judgment, reasoning and planning for the selected task, and decomposition into specific steps and action sequences such as movement and operation for task solving;

[0080] The step S4 specifically includes: the intelligent agent outputs the ontology control signal to perform actions or directly completes the cognitive task in the constructed space, and realizes self-evolution by adaptively adjusting the memory bank of the conscious module and the model parameters of each module of the subconscious module through observation of the scene construction and action control error during task execution;

[0081] The step S5 specifically comprises: performing the step S2-S4 for multiple rounds, the agent actively interacts with the environment and the user, continuously generates tasks and autonomously decides to execute, the agent thinking is transparent and controllable in the whole process, and the agent adjusts the attributes through learning.

[0082] Compared with the prior art, the application has the beneficial effects:

[0083] The active AI of the application uses abstract meta-point to uniformly represent and interpret various modal signals, and autonomously constructs a world model, thereby having the ability to understand the physical world and create a cultural world. The active AI is a general artificial intelligence system designed to learn and expand advanced concepts and cognitive logic by focusing on and analyzing limited data outside its own cognitive range; under the active AI framework, symbolism, connectionism and behaviorism play an efficient role in different parts of the general intelligent system, wherein the subconscious module as a "fast system" splits the end-to-end neural network "black box" processing into multi-level dimension reduction processing such as sensation, perception and cognition under the guidance of the "slow system" conscious module, greatly solves the problems of data-driven model vulnerability and uncontrollability through abstract form concept and logic representation, and greatly reduces the demand for massive data, model size and computing power in an end-to-end manner.

[0084] In the machine agent, the machine uses a general language as the underlying thinking and interaction language, and the thinking process and result are completely transparent, so that autonomous decision-making, autonomous learning and continuous evolution can be realized in the case of alignment with human values and complete controllability. The general language is embedded in the abstract concept interaction language of the underlying physical and logical representation, the method for constructing the general agent is also the method for constructing the general language, and compared with the prior art of natural language, the machine agent based on the general language has wide generality in application. BRIEF DESCRIPTION OF DRAWINGS

[0085] In order to more clearly illustrate the technical solutions in the embodiments of the application, the following will briefly introduce the drawings needed to be used in the embodiment description. Obviously, the drawings in the following description are only some embodiments of the application, and for those skilled in the art, other drawings can also be obtained without creative labor based on these drawings.

[0086] Fig. 1 is a schematic diagram of the system composition of the general agent of the application;

[0087] Fig. 2 is a schematic diagram of the input module and the output module of the general agent of the application;

[0088] Fig. 3 is a schematic diagram of the conscious module and the self-conscious module of the general agent of the application;

[0089] Fig. 4 is a schematic diagram of the subconscious module of the general agent of the application;

[0090] Fig. 5 is a schematic diagram of the information exchange module of the general intelligent agent of the present application;

[0091] Fig. 6 is a schematic diagram of the input picture of embodiment 2 of the present application;

[0092] Fig. 7 is a schematic diagram of the extracted meta-points in the input signal of embodiment 2 of the present application;

[0093] Fig. 8 is a schematic diagram of the conversion of the meta-point set of embodiment 2 of the present application;

[0094] Fig. 9 is a schematic diagram of the circular meta-point set and concept pattern of the concept library of embodiment 2 of the present application;

[0095] Fig. 10 is a schematic diagram of the decomposed scene of the object constructed by embodiment 2 of the present application;

[0096] Fig. 11 is a schematic diagram of the objects perceived by the machine and the interpretation of the given conditions and tasks of embodiment 2 of the present application;

[0097] Fig. 12 is a schematic diagram of the machine cognition process of embodiment 2 of the present application;

[0098] Fig. 13 is a schematic diagram of the completion of the task by analogy by the machine of embodiment 2 of the present application;

[0099] Fig. 14 is a schematic diagram of the image captured by the camera of embodiment 3 of the present application;

[0100] Fig. 15 is a schematic diagram of the three-dimensional scene generated in the constructed space of embodiment 3 of the present application;

[0101] Fig. 16 is a schematic diagram of the conversion of the three-dimensional scene to a two-dimensional form of embodiment 3 of the present application;

[0102] Fig. 17 is a schematic diagram of the sending of the command by the user in the general language of embodiment 3 of the present application;

[0103] Fig. 18 is a schematic diagram of the logical analysis of the physical environment by the robot of embodiment 3 of the present application;

[0104] Fig. 19 is a schematic diagram of the various movement modes in the logical library of the robot of embodiment 3 of the present application;

[0105] Fig. 20 is a schematic diagram of the prediction of the behavior process and the result of the command by the user by the robot of embodiment 3 of the present application.

[0106] The specific embodiments of the present application have been shown through the above-described figures, and will be described in more detail hereinafter. These figures and the written description are not intended to limit the scope of the inventive concept in any way, but to illustrate the inventive concept to those skilled in the art by referring to specific embodiments. DETAILED DESCRIPTION

[0107] The exemplary embodiments will be described in detail herein below with reference to a few examples. The implementations described in these examples do not represent all implementations consistent with the present application. Instead, they are merely examples of apparatuses and methods consistent with some aspects of the present application as detailed in the appended claims.

[0108] The terms "first", "second", "third", etc. (if any), or "step 1", "step 2", "step 3", etc. (if any) in the description and claims of this application and the above drawings, if any, are used for distinguishing between similar elements and not necessarily for describing a sequential or chronological order. It is to be understood that the use of these terms herein is to distinguish one feature from another and not necessarily describe a sequential or chronological order. It is to be understood that the data so used in this description and claims is merely provisional and should not be unduly construed in support of or in derogation of the patenting of the subject application. Furthermore, the terms "comprising", "having", "including", and the like, as used in the specification are used in their broadest and most general sense and are intended to be construed as open-ended terms, meaning that the listed steps, features, components, elements or the like can be supplemented after the fact by additional steps, features, components, elements, or the like. Additionally, it is understood that the use of "including", "comprising", "having" and the like, are not meant to be limiting and do not exclude the presence of non-recited, additional steps, features, components, elements, or the like.

[0109] Embodiment 1

[0110] The general intelligent agent composition and working mode of the present application will be described in detail below in combination with the drawings and specific embodiments.

[0111] As shown in FIG. 1, a general intelligent agent includes the following modules:

[0112] 1) an input module for obtaining pre-processed input signals;

[0113] 2) a consciousness module for perception and cognition operation in a transparent and interpretable abstract thinking form;

[0114] 3) a self-consciousness module for perception and cognition of the intelligent agent itself;

[0115] 4) a subconsciousness module for realizing perception and cognition functions under subconsciousness using a data-driven model;

[0116] 5) an information exchange module for exchanging and fusing the information generated by the consciousness module, the self-consciousness module and the subconsciousness module;

[0117] 6) an output module for outputting the processing results as needed;

[0118] The input module is connected with the output module through an information processing module, the information processing module includes two relatively independent groups of modules, one group includes consciousness module and self-consciousness module (for perceiving the outside world and the body respectively), and the other group is subconsciousness module, the information processing module processes input signals obtained by the input module, and the processing result is sent to the output module, and the information exchange module is responsible for information exchange and fusion between the two relatively independent groups of modules.

[0119] As shown in FIG. 2, the input module specifically includes receiving signals and pre-processing the signals, the signals including received general language sentences, natural language sentences (voice or text), codes, agent memory database data, positioning tracking, vision, hearing, touch, smell, taste and various types of modal signals received by other external sensors of the body sensing sensor; the pre-processing includes various common signal pre-processing methods such as sampling, changing image resolution, signal enhancement, denoising and filtering, and various common methods for converting signals from signal space to feature space through linear or nonlinear transformation for some modal signals, in the transformation domain in order to better perform feature dimension reduction or clustering.

[0120] The general language is an abstract concept interaction language embedded with bottom layer physical and logical representation, as a bottom layer human-computer interaction language, has the advantages of maximum expression flexibility, accurate and sceneized intention expression, transparent safety, complete interpretability and the like. The patent application number: 202410201258.1, patent name "human-computer interaction method, device and equipment and storage medium" provides a construction method of general language at the perception level. The present application extends the general language to the cognitive level, which can more completely express abstract perception and cognitive information.

[0121] The output module specifically includes: according to the interaction and action requirements, outputting the general language sentences, natural language sentences (voice or text), codes, agent memory database data, ontology control signals and other image, video, audio and various types of modal signals generated by the consciousness module, self-consciousness module and subconsciousness module.

[0122] As shown in FIG. 3, the thick lines and dark boxes correspond to the self-consciousness module, and the other lines and boxes correspond to the consciousness module, the consciousness module includes:

[0123] 21) Sensory sub-module, the input signals obtained by the input module are converted into clues through bottom layer intuitive processing;

[0124] 22) Perception sub-module, the clues are used to arouse related concepts and perception logic, and the objects and scenes are constructed in the construction space with new perspectives and postures;

[0125] 23) cognitive sub-module, using the constructed object and scene condition part and task part, arousing relevant logic, generating judgment, reasoning or planning in the construction space, obtaining rational cognition;

[0126] 24) output modal conversion sub-module, converting the general language form of thinking into natural language text, audio, image, video and other modal signals as needed;

[0127] 25) long-term memory sub-module, storing the long-term memory data of the agent pre-set or learned, including concept library, logic library, ontology model library, etc.

[0128] The sensory sub-module, the perceptual sub-module, the cognitive sub-module and the output modal conversion sub-module are sequentially connected, wherein the concepts and logic data used in each module are searched and accessed in the long-term memory sub-module.

[0129] The self-awareness module includes:

[0130] 31) proprioceptive sub-module, for receiving proprioceptive signals, including navigation positioning, time perception, vision, hearing, touch, smell, taste, inertia, balance and other types of proprioceptive signals on the agent trunk and each component;

[0131] 32) proprioceptive perception sub-module, using proprioceptive signals and proprioceptive model to estimate model parameters, and combining objects and scenes in the construction space, calculating the position and posture of the agent body including the trunk and each component relative to other objects in the scene, and correcting and updating the proprioceptive model when the proprioceptive model is not applicable;

[0132] 33) proprioceptive cognitive sub-module, if the body needs to participate in the judgment, reasoning or planning process, the body is filled into the aroused logic as a content or element, and participates in the judgment, reasoning or planning process;

[0133] 34) proprioceptive action control sub-module, when the body needs to act, the planning results generate action sequences, hardware control programs and other information as body control signals;

[0134] 35) proprioceptive model library, storing the structural parameter model of the agent trunk and each component in mathematics, physics, graphics, motion, interaction, etc. The model library can be stored in the long-term memory sub-module.

[0135] In FIG. 3, the nodes P0, P1, P2, P3, P4 are located at different positions in the data processing flow of the consciousness module, and the data types at the nodes are respectively: P0: multi-modal data in the input signal obtained by the input module; P1: general language statements representing the clue element point set or graph generated by the sensation sub-module; P2: general language statements representing the object, scene element point set or graph generated by the perception sub-module; P3: general language statements representing the cognitive element point set or graph generated by the cognitive sub-module; P4: multi-modal data in the output signal output by the output module.

[0136] As shown in FIG. 4, the subconscious module includes:

[0137] 41) a subconscious sensation sub-module, which implements the sensation function under the subconscious using a data-driven model;

[0138] 42) a subconscious perception sub-module, which implements the perception function under the subconscious using a data-driven model;

[0139] 43) a subconscious cognitive sub-module, which implements the cognitive function under the subconscious using a data-driven model;

[0140] 44) a subconscious output modality conversion sub-module, which implements the multi-modal signal output function under the subconscious using a data-driven model.

[0141] In FIG. 3, the nodes P0, P1, P2, P3, P4 represent the different stage processing processes of the signal in the consciousness module from input to output, and in the subconscious module shown in FIG. 4, there are also five representative nodes P0', P1', P2', P3', P4', and the data types at the nodes are respectively: P0': input signal obtained by the input module; P1': general language statements adapted to the clue element point set or graph generated by the sensation sub-module of the consciousness module; P2': general language statements adapted to the object, scene element point set or graph generated by the perception sub-module of the consciousness module; P3': general language statements adapted to the cognitive element point set or graph generated by the cognitive sub-module of the consciousness module; P4': multi-modal data adapted to the output signal of the output module.

[0142] A type of input data and a type of expected output data can train a set of neural networks, so that 5 nodes are combined in pairs, and there are 10 groups of neural network models of P0'> P1', P0'> P2', …, P3'> P4', etc. According to the type of output signal, they are classified into the subconscious feeling submodule, the subconscious perception submodule, the subconscious cognition submodule and the subconscious output mode conversion submodule. According to the demand and effect, part of the model is selected, and the input or output involves multi-modal data. A single multi-modal model can be trained, or a set of sub-models can be trained for each type of modal data. The training method adopts the pre-training method of batch data, and the model parameters can also be updated in real time with new training data. P0'> P4' represents a multi-modal input to a multi-modal output neural network model, and can also be directly connected to and called existing various single-mode or multi-modal large models.

[0143] As shown in FIG. 5, P0, …, P4 and P0', …, P4' correspond to the representative nodes in the consciousness module of FIG. 3 and the subconscious module of FIG. 4 respectively. The information exchange module 5) processing flow specifically includes:

[0144] 51) The input signal obtained by the input module, and the feeling, perception, cognition, and multi-modal output data generated by the feeling, perception, cognition, and output mode conversion submodules in the consciousness module are added to the training data set for the submodules of the subconscious module to train the model;

[0145] 52) The feeling, perception, cognition, and multi-modal output data generated by the feeling, perception, cognition, and output mode conversion submodules in the subconscious module are provided as reference data to the submodules of the consciousness module for the submodules of the consciousness module to perceive, analyze, correct, and integrate with the data generated by the consciousness module itself, so as to obtain more accurate and smooth output results.

[0146] On the basis of the composition of the agent system, the application provides an agent control method, which comprises:

[0147] S1, memory presetting and agent pre-training;

[0148] S2, value-driven task generation;

[0149] S3, task-oriented autonomous decision-making;

[0150] S4, task execution and self-evolution;

[0151] S5, steps S2-S4 are performed for multiple rounds, and the agent actively interacts with the environment and the user.

[0152] The intelligent agent selects a task by value judgment, autonomously generates judgment, reasoning, planning and specific steps, outputs ontology control signals for action execution or directly completes cognitive tasks in the construction space. The perception and cognition results in the construction space are generated in a predictive manner, and the subsequent accuracy of the prediction results can become the agent's spontaneous reward and punishment signals. Therefore, the agent's observation, decision-making, action, and reward in conscious interaction with the environment and users become a closed loop. During task execution, the alignment of human-machine values and self-evolution are achieved by adjusting the memory library of the consciousness module and the model parameters in the subconscious module.

[0153] Using the above control method, the intelligent agent can include but is not limited to the following eight working modes: autonomous exploration mode, active learning mode, passive learning mode, interactive teaching mode, task-driven autonomous decision-making mode, goal-driven autonomous decision-making mode, value-driven autonomous decision-making mode, and dialogue and multi-modal translation mode. The modules and sub-modules of the intelligent agent are only a logical division of functions. Physically, these modules can be integrated or independent. Some modules can be ignored or not executed during actual implementation, such as the subconscious module when the consciousness module is working.

[0154] Embodiment 2

[0155] Based on Embodiment 1, the object perception, scene understanding, and task reasoning of the present application are described in detail below in conjunction with the accompanying drawings and Embodiment 2.

[0156] The processing flow of the sensing sub-module includes:

[0157] 211) Call the input signal obtained by the input module and obtain the attention preset parameter, which includes the attention point, the attention type and the value range of the feeling, the minimum attention granularity, and other calculation parameters related to attention in the intuitive process; in this embodiment, the input is the picture shown in Figure 6, which is a classification reasoning question adapted from question 91 in a set of visual puzzles - the Bondage problem. The attention point in the attention preset parameter covers the entire image, and the feeling parameter is the gray value with a minimum attention granularity of a circle with a radius of 4.

[0158] 212) According to the attention preset parameter, use the region segmentation algorithm to select the plane or spatial region that is paid attention to in the signal digital space; the region segmentation algorithm includes traditional image segmentation methods, image segmentation methods using deep learning, etc.; the region is marked with a binary mask, the edge is extracted in the form of a closed curve or surface, and the region is represented. The plane or spatial region includes the region range and the corresponding feeling parameter, and the feeling parameter is the feeling type and its value range according to which the region is divided; in this embodiment, each black region is marked with a binary mask by binary quantization, as the plane region that is paid attention to.

[0159] 213) The selected region is structured and represented as a set of meta-points composed of basic elements, and stored as a clue in the construction space. Meta-points are the basic elements of the machine thinking space, which can represent clues, concepts, logic, objects, etc. in the thinking space in the form of abstract meta-point sets. The meta-point set is represented as follows:

[0160]

[0161] where P represents physical properties including location properties, range properties, and sensory properties. The location property is the representative location of the meta-point in the space-time coordinate system, the range property is the extension range size of the meta-point in the space-time coordinate system, and the sensory property is the sensory type and its value range to which the meta-point belongs; E represents extended properties including connection properties and dynamic properties. The connection property is a static property that indicates the connection relationship between meta-points, and obtains the abstract perception of the topological relationship between the components of an object. The dynamic property describes the extension trend of the meta-point in the form of a straight line or a curve in space and the change trend over time, and obtains the abstract feature perception of the straight line or curve in the space-time coordinate system of the object; V represents value view properties including facts, costs, regulations, values, aesthetics, preferences, desires, beliefs, intentions, etc.; R represents reference properties, which represent the target region pointed to by the meta-point in the clue, concept, and object meta-point, and can refer to object or scene meta-point sets in the logic meta-point; O represents other properties. i i i i i

[0162] In the clue extraction stage, the meta-point uses skeleton line extraction or corner point, centroid, edge detection and extraction methods in image processing, or uses point cloud estimation, Gaussian distribution estimation, etc. trained by neural network to combine and simplify, and converts the extraction result into structured data; the meta-point set is converted into a corresponding clue pattern, which is stored as a visual auxiliary image of the corresponding clue in the construction space. The pattern is a graph (which can be represented as G(M)) converted from the meta-point set, which can explicitly depict various extended properties in the meta-point set, such as the topological relationship represented by the connection property in the form of nodes and edges of the graph, and the extension and change trend in space and time represented by the dynamic property in the form of straight lines and curves. The pattern is similar to the abstract representation of simple sketches, schematics, etc. that people often use, and has high understandability and good affinity for users.

[0163] ​​​​​In this embodiment, the skeleton line algorithm and the corner point detection method are combined to extract the meta-point. According to the attention preset parameter, the minimum inscribed circle radius is set to 4, and the obtained meta-points are shown in FIG. 7. Because the number of objects contained in the input picture is large, FIG. 7 only shows the perception result of one of them (the pattern in the second row and the second column of FIG. 6), which contains two types of meta-points: static meta-points (meta-points with empty dynamic attributes) marked with "×", and dynamic meta-points (meta-points with non-empty dynamic attributes) marked with solid circles.

[0164] As an example of the meta-point, the following table gives the data structure of two meta-points in FIG. 7, giving the position, range, connection, and dynamic attribute values. The position attribute value is the meta-point coordinate, the range attribute value is the meta-point radius, the connection attribute value is the connected other meta-point, and the dynamic attribute value has multiple forms. Here, the extension trend of the meta-point is expressed as [direction, bias, curvature radius] by simulating the way the meta-point travels, where the bias has three values: -1 represents left deflection along the travel direction, 0 represents straight travel, and 1 represents right deflection.

[0165] Meta-point A Meta-point D Position attribute (48, 34) (107, 103) Range attribute 151 Connection attribute {B, C} {E, F, G} Dynamic attribute {[ -39 o , -1, 91], [50.6 o , 0]}

[0166] The line clue pattern G(M) converted from the meta-point set is shown in FIG. 8, which indicates the preliminary perception result of the input signal by the machine after the bottom-level intuitive processing. Various straight lines, arcs, and three-lipped irregular objects are perceived.

[0167] The processing flow of the perception sub-module includes:

[0168] 221) Call the clues obtained by the perception sub-module;

[0169] 222) Use the clues to arouse related concepts in the concept library and store them in the construction space. Concepts are homogeneous with clues and are also structured data structures composed of meta-points as basic elements, which can also be converted into concept patterns. In this embodiment, the concept library stores some basic geometric concept patterns obtained through presetting or learning. According to the various clues (meta-point sets) found in step 221), the related concepts are found and called out from the concept library. The concept library contains various concepts such as straight lines, angles, quadrilaterals, arcs, and circles. For example, the dynamic attributes of meta-point A and meta-point B in FIG. 7 have the same curvature radius, and the "circle" concept is aroused as shown in FIG. 9. For example, meta-point C in FIG. 7 has two straight line extension directions, and there are four meta-points including meta-point C associated by two straight lines. According to the principle of feature matching, the "quadrilateral" concept is aroused.

[0170] 223) With the awakened concept as a template and relevant clues as materials, the object is constructed in the construction space with a new perspective and posture, the existence and state of the object are perceived, and the relevant perceptual logic is awakened to build the environment, conditions, and tasks, etc. related scene; it specifically includes:

[0171] After awakening the concept, according to the matching relationship between the concept and the clue, the meta-points in the concept are filled with the corresponding clue meta-points, and the position, range, feeling, extension, etc. of each meta-point in the concept are updated. In this way, the object is constructed in the space-time coordinate system with a new perspective and posture, and the relevant perceptual logic is awakened to build the environment, conditions, and tasks, etc. related scene, and the object and scene are stored in the construction space. Through the construction of specific objects by clues and concepts, the machine perceives the existence and state of the object, and can also recognize the object, its components, and their names. The logical meta-point is a pure rational meta-point, which is filled by the meta-point set of the object or scene referred to in the reference attribute when applied; the extension attribute in the logical meta-point contains the connection relationship and connection strength with other meta-points, which is used to represent different types and strengths of logical relationships. As a kind of logic, perceptual logic is the perception rule of the time, physical and geometric relationship between objects or components of objects acting on the object or component. Perceptual logic can solve the construction of relatively simple and direct environment, conditions and task scenes at the perceptual level. The construction of scenes at a higher cognitive level can be carried out in subsequent cognitive sub-modules.

[0172] Among them, the filling in the concept awakening and object construction process refers to the meta-points in the concept and clues at key points, lines, and graphs, and a matching relationship is established in the similarity calculation, so that the meta-point attributes in the clues can be used instead of the meta-point attributes in the concept. The construction of foreground and background or multiple objects can form a scene, which is a static scene in a plane or space, or a dynamic continuous or sequential slice scene in space-time; the construction is working in a predictive way, and the machine needs to construct and explain the surrounding scene at that time, and construct and predict the subsequent state of the object and scene through relevant concepts and perceptual logic. Scene explanation and understanding are carried out in a predictive way. If the subsequent development deviates from the previous prediction result, it means that the scene construction does not completely match the reality, and the concept and logic need to be awakened and updated, and the construction of the object and scene needs to be updated.

[0173] In this embodiment, the meta-points in the concept and the meta-points in the clue establish a corresponding relationship in the arousal process, and the meta-point attributes in the circular and quadrilateral concepts aroused in step 222) can be replaced by the meta-point attributes in the corresponding clues in step 221), so that new objects such as circles, quadrilaterals, etc. are constructed in the construction space with new perspectives and postures, and the perceptual logic of component decomposition and missing of the objects is aroused according to the mutual characteristics between the object meta-point sets, as shown in FIG. 10. The logical meta-points in the component decomposition are filled by the corresponding object meta-point sets, respectively, so that the constructed scene is: the initial complete circular object is decomposed into two parts, one part contains three quadrilaterals, and the other part is the remaining three irregular pieces. These preliminary perceived objects and scenes are often a direct approximation to the actual situation. In FIG. 10, the circle, quadrilateral, and three irregular pieces are all constructed objects and scenes that can be stored as materials in the construction space. Because the concepts in the concept library, such as the circle concept shown in FIG. 9, also have the text names of some components, the machine can also give the text names of the objects and components constructed in FIG. 10.

[0174] 224) Call the sensory sub-module to find new clues to fill and update, support and enrich the constructed objects and scenes, which specifically include: after the object is constructed, if some components of the object are missing due to the mismatch between the meta-points of the clue and the concept, or the object needs more detailed attention, the machine needs to find new clues to fill and update the object. In the found clues, pay attention to the clues that are not used in the arousal process or the meta-points in the clues that are not used, or return to the original input signal to change the attention parameters and find new clues, so as to support and enrich the constructed objects and scenes. In this embodiment, the feature of the quadrilateral object corresponding to the meta-point set in FIG. 10 is missing a side, and after returning to the original signal and changing the attention parameters, no related clues are found, which indirectly supports the assumption of object component decomposition and missing in step 223) of scene construction.

[0175] 225) Comparing the similarities and differences between the object, scene and the relevant concepts in the concept library, forming new concepts on the basis of existing concepts, expanding the concept library, including: in the case of the object and scene obtained in step 224), judging the similarities and differences between the object and scene and the relevant concepts in the concept library, and updating the object, scene and corresponding clues in a simplified and highlighted manner to form new concepts on the basis of existing concepts, and storing the new concepts in the concept library as high-level concept categories; the user can directly correct and modify the concepts in the concept library, or view the newly generated concepts in the construction space and modify them in subsequent steps through human-computer interaction; the user can also add symbolic names, typical feature descriptions, and derivative relationships with existing concepts in the concept data structure generated by the machine. In this embodiment, the machine adds the three irregular pieces shown in Figure 8 to the concept library, and the user decides whether to keep and add a textual description.

[0176] 226) Specifically includes: based on the original input signal and the user's interaction command, or the clues, concepts and objects obtained in the construction space, multiple rounds of steps 221) to 225) are performed as the attention parameters change, the object and scene are updated, the existing object and scene are observed in detail, and new objects and new scenes are generated after similar concepts and perceptual logic are aroused, and according to the user's needs perceived from the input signal, the relevant objects and scenes in the construction space are classified into the condition category and the task category respectively.

[0177] Wherein the changes of the attention parameters include the changes of the calculation parameters such as the shift of the focus point, the focusing or generalization, the type of feeling and the attention granularity, and the attention parameters are changed by the machine independently during the construction of the object and scene, or can be changed by the user in the interaction.

[0178] In this embodiment, the machine perception process is exemplified by the pattern in the second row and the second column of the input picture, and the objects perceived by the machine in the input picture are shown in Figure 11, and the relevant objects are classified into the condition category (including condition 1 group and condition 2 group) and the task category; because there are obvious letters and question mark prompts in Figure 6, the "selection" logic and the "filling" logic are directly aroused and converted as the specific tasks understood, and the interpretation of more complex conditions and tasks can also be performed in the subsequent cognitive sub-module. In practice, according to the attention, the concept library, the logic library and the materials in the construction space, the machine constructs the scene and the object in a predictive manner, and the prediction or interpretation can be flexible and diverse, as long as it conforms to the machine's own cognition and can understand the input signal in its own way. The interpreted objects and scenes can be used as alternatives to achieve the user's specific purpose in the interaction with the user.

[0179] The cognitive sub-module processing flow includes:

[0180] 231) Call the perception submodule to obtain objects and scenes;

[0181] 232) Deepen the conditions and task scene construction, while arousing relevant logic, including directly arousing in the logic library, and indirectly extracting relevant logic from specific scene concepts involving judgment, reasoning or planning in the concept library, and storing in the construction space;

[0182] 233) With the aroused logic as a template, and the relevant objects and scenes as materials, after arousing the logic, according to the matching relationship between the logic and the objects and scenes, the meta-points in the logic are filled by the corresponding object and scene meta-point set, and the connection strength and other parameters in the logic meta-point are calculated, and the attributes of each meta-point in the logic are updated, in this way, the judgment, reasoning or planning containing specific content is generated in the construction space, and the preliminary rational cognition is obtained.

[0183] In this embodiment, as shown in FIG. 11, after perceiving each object and given conditions and tasks, according to the matching relationship between the logic and the objects and scenes, the "induction", "similarity", "analogy" and other logics are successively aroused in the logic library. As shown in FIG. 12, first, use the "induction" logic 1201 to process each object in the condition category, according to the processing flow of "forming hypothesis, verifying hypothesis", the attribute values in each object meta-point set are respectively extracted and statistically compared according to the category, quantity, numerical value and the like, so as to extract the common features, that is, "three" in condition 1 group and "four" in condition 2 group shown in FIG. 12; use the "similarity" logic 1202 to calculate the similarity of the extracted common features with the two objects in the task category respectively; finally, as shown in FIG. 13, arouse the "analogy" logic, compare and replace the parameters and contents of the "selection" task in FIG. 11 and the "similarity" logic in FIG. 12, so as to obtain the final task answer (FIG. 13 only gives the result in condition 1 group, and condition 2 group can be completed in the same way).

[0184] Step 234) Find new objects or new scenes to fill and update, support and enrich the constructed judgment, reasoning or planning, in the objects and scenes found in step 231), pay attention to the objects and scenes that are not used in the logic arousing process, or call the perception submodule to return to the original input signal to change the attention parameters and find new clues, objects and scenes; the meta-points in the judgment, reasoning or planning meta-point set which still have no content filling can arouse related concepts through the concept library to generate new objects and scenes as filling content.

[0185] 235) Forming judgment, reasoning or planning can be used as a specific scene, forming new concepts on the basis of existing concepts, expanding the concept library; at the same time, looking for rules in the judgment, reasoning or planning, updating and forming new logic on the basis of existing logic, expanding the logic library.

[0186] 236) Repeat steps 231) through 235) multiple times to conduct a more detailed analysis and consideration of the task and the resulting judgments, reasoning, or planning. In this embodiment, the reasoning results can be examined and verified from multiple perspectives. As shown in Figure 12, after the inferred objects are classified into corresponding conditional categories through the fill logic 1203, the expanded group members are further summarized to verify the results for errors.

[0187] It should be noted that the graphical forms of objects, scenes, reasoning, and other content generated in the construction space according to the previous steps, as shown in Figures 8-13, use abstract graphics similar to commonly used stick figures, schematic diagrams, and flow charts. Users can view these directly in the machine's construction space, or the machine can directly output all or selected portions of these contents as statements in a universal language through the output module 6). Users can also send their own requirements to the machine through the input module in the form of stick figures, sketches, schematic diagrams, flow charts, etc., which the machine can also understand after generating the corresponding objects and scenes in the construction space through the previous steps. The perceptual and cognitive content generated in the construction space clearly demonstrates the machine's thinking process and results, essentially forming the machine's underlying language, serving as the universal language for machine thinking, program control, and interaction with users.

[0188] Example 3

[0189] Based on Example 1, the present invention is described in detail in terms of three-dimensional scene understanding, human-computer interaction, and autonomous decision-making in conjunction with the accompanying drawings and Example 3.

[0190] The robot is equipped with binocular cameras, a robotic arm, and wheels, enabling it to move around the room, observe the environment, and grasp objects. The input module collects signals from the binocular cameras. Figure 14 shows an image captured by one of the cameras. After clue-finding and concept-recalling in the awareness module, a 3D scene graph (Figure 15) is generated in the 3D construction space. The robot has only learned the concepts of simple objects and can identify room structures such as walls, ceilings, and floors, as well as objects within the room such as tables and spheres. However, the robot does not recognize other objects, such as outlet panels, closets, fire safety signs on closets, and hand sanitizer. Figure 16 shows the scene graph converted from the 3D scene to a 2D representation, as seen in Figure 14. The clue-finding process utilizes a 3D point extraction method. This involves first extracting a point on a plane from two monocular images, and then using a binocular vision space position estimation algorithm to determine the spatial coordinates of the point.

[0191] The user views the scene construction result of the robot in real time through the output module, and sends instructions or commands to the robot in a common language based on the result. For example, if the user needs the robot to move the ball in the center of the table to the floor position, as shown in FIG. 17, an arrow is added to FIG. 16 and sent to the robot. The arrow here represents the movement or operation of the object represented by the meta-point in the space-time coordinate system, that is, the spatial and temporal process of the object moving from the starting meta-point position of the arrow to the ending meta-point position.

[0192] After receiving the command, the robot generates a task scene in the construction space and performs physical and operation analysis on the relevant objects that need to be operated. As shown in FIG. 18, according to the gravity and balance logic injected in the logic library, the ball to be operated does not fall due to gravity because of the support of the table. As shown in FIG. 19, the movement logic in the logic library has two cases: unobstructed movement (from position A to position B) and obstructed movement (from position A to position B through the obstruction C), according to the position and range attributes in the meta-point set of the table, the ball movement task in FIG. 17 belongs to obstructed movement, so the task is refined into three behavior options in the "obstructed movement" logic, corresponding to three cases of breaking the obstruction with a tool, moving the obstruction away, and bypassing the obstruction. The robot calculates the value attribute values for the three cases respectively, extracts and scores from the value attribute of the meta-point set of the same or similar scene in the logic and concept library, and selects the behavior with the highest score as the candidate. In this embodiment, because the user specifies that the robot cannot damage the items in the room, and predicts that breaking or moving the table may require more effort and cost, and operating the table as a support may have risk cost factors, the robot selects the option of moving the ball by bypassing the obstruction after comprehensive evaluation. Therefore, the robot generates the movement process and result of the object in a predictive manner in the construction space, and feeds back to the user in the form of a dynamic diagram or a static diagram as shown in FIG. 20 in the interactive interface, so that the user can check whether the robot's understanding of the command and the prepared behavior operation are appropriate, and can supplement the command in time if necessary.

[0193] After the robot accurately understands the user's intention, the robot converts the action sequence generated in the construction space into a hardware control program for operation in the self-awareness module, moves to the side of the table through the wheeled device and moves the ball to the specified position using the mechanical arm. During the operation, the robot perceives the changes of the objects and the environment in real time and compares them with the predicted scene result, and if the task fails, the robot re-formulates the action sequence and performs the operation under unsupervised or user supervision.

[0194] Using the intelligent system and control method provided by the present application, the robot can understand the physical environment in an unknown environment, accurately understand the user's intention in a new task, and autonomously plan and decide to complete the user's instruction in a safe operation manner.

[0195] The various embodiments described in this specification are presented for the purpose of illustration and description. Each of the embodiments highlights a different aspect of the disclosure. However, it should be understood that the embodiments are not mutually exclusive and can be combined.

[0196] In several embodiments provided in the present application, it should be understood that the disclosed system, device and method can be implemented in other ways. For example, the above-described device embodiments are merely illustrative, and the division of units is merely a logical function division. In actual implementation, another division manner can be adopted, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the units shown or discussed can be indirect coupling or communication connection through some interfaces, devices or units, and can be electrical, mechanical or other forms. In addition, each functional unit in each embodiment of the present application can be integrated into a processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be realized in the form of hardware or in the form of a software functional unit.

[0197] Other embodiments of the present disclosure will be apparent to those skilled in the art from consideration of the specification and practice of the application disclosed herein. The present application is intended to cover any variations, uses or adaptations of the application following the general principles thereof and including such departures from the present disclosure as come within known use or custom in the art to which the present application pertains. It is to be understood that the application is not limited to the exact details shown and described herein, and that work within the scope of the application can vary under many conditions.

Claims

1. A general intelligent agent, characterized by: include: 1) Input module, used to obtain pre-processed input signals; 2) Awareness module, used for perception and cognitive operations in a transparent and explainable abstract form of thinking; 3) Self-awareness module, which is the perception and cognition of the intelligent body; 4) Subconscious module, which uses data-driven models to realize subconscious perception and cognitive functions; 5) Information exchange module, used to exchange perception, cognition and other information generated by the consciousness module, self-awareness module and subconsciousness module for integration and utilization; 6) Output module, outputs the processing results as needed; The input module is connected to the output module through the information processing module. The information processing module includes two relatively independent groups of modules, one group includes a consciousness module and a self-awareness module, and the other group is a subconscious module. The information processing module processes the input signal obtained by the input module and sends the processing result to the output module. The information exchange module is responsible for the information exchange and integrated utilization between the two relatively independent groups of modules.

2. A general intelligent agent according to claim 1, characterized in that: The awareness module includes: 21) The sensory submodule converts the input signal obtained by calling the input module into clues through the underlying intuitive processing; 22) The perceptual submodule uses cues to evoke relevant concepts and perceptual logic, constructing objects and scenes from new perspectives and postures in the constructed space; 23) The cognitive submodule uses the constructed objects, scene conditions, and tasks to evoke relevant logic, generate judgments, reasoning, or planning within the constructed space, and obtain rational cognition; 24) Output modality conversion submodule, which converts the universal language form of thinking into natural language text, audio, image, and video modality signals as needed; 25) Long-term memory submodule, which stores the long-term memory data preset or learned by the intelligent agent, including the concept library, logic library, and ontology model library.

3. A general intelligent agent according to claim 2, characterized in that: The processing flow of the sensory submodule includes: 211) Calling the input module to obtain an input signal and acquire attention preset parameters, wherein the attention preset parameters include the focus point, the type of feeling of attention and its value range, and the calculation parameters related to attention in the minimum attention fine-grained intuitive process; 212) Based on the preset attention parameters, a region segmentation algorithm is used to select the plane or spatial region of interest in the signal digital space; 213) The selected area is structured and converted into a set of meta-points composed of meta-points as basic elements, and stored as clues in the construction space; the meta-points are the basic components of the machine thinking space, and the meta-point set is represented as follows: ; Among them, P i Represents physical properties, E i Indicates extended attributes, V i Represents the value attribute, R i Indicates the reference attribute, O i Represent other attributes; in the clue extraction stage, the meta-points are simplified by using skeleton line extraction or detection and extraction methods of various key points and key lines in image processing such as corner points, centroids, edges, etc., or point cloud estimation and Gaussian distribution estimation methods trained by neural networks are used, and the extraction results are converted into structured data; the meta-point set is converted into the corresponding clue pattern and stored in the construction space as a visual subsidiary image of the corresponding clue.

4. A general intelligent agent according to claim 2, characterized in that: The processing flow of the perception submodule includes: 221) Call the clues obtained by the sensory submodule; 222) Use clues to recall related concepts in the concept library and store them in the construction space; 223) Using the evoked concepts as templates and relevant clues as materials, construct objects from new perspectives and postures in the constructed space, perceive the existence and status of the objects, and at the same time evoke relevant perceptual logic to construct relevant scenes such as environment, conditions and tasks; 224) The sensory submodule is called again to search for new clues to fill in and update, support and enrich the constructed objects and scenes; 225) Compare the similarities and differences between objects and scenarios and related concepts in the concept library, form new concepts based on existing concepts, and expand the concept library; 226) Steps 221) to 225) are repeated multiple times to observe and analyze the scene and objects in more detail.

5. A general intelligent agent according to claim 4, characterized in that: The specific process of step 222) is: searching for one or some concepts in the concept library that are identical or similar to the clues obtained in step 221) in one or some aspects of the data structure, and transferring the relevant concepts from the concept library to the construction space; The specific process of step 223) is as follows: after the concept is evoked, based on the matching relationship between the concept and the clue, the element points in the concept are fully or partially filled with the corresponding clue element points, and the attributes of each element point in the concept are updated. In this way, the object is constructed in a new perspective and posture in the space-time coordinate system, thereby perceiving the existence and state of the object, and at the same time, the relevant perceptual logic is evoked to construct the environment, conditions and task-related scenes; In step 224), the new clue refers to a clue found by the sensory submodule that is not used in the attention arousal process or an unused element in the clue, or a clue that is found again by changing the attention parameters in the original input signal; The step 225) specifically includes: determining the similarities and differences between the objects and scenes obtained in step 224) and the related concepts in the concept library, and updating the objects, scenes and corresponding clues based on the existing concepts in a manner of simplifying and highlighting the differences to form new concepts, and storing them in the concept library as high-level concept categories; The step 226) specifically includes: based on the clues, concepts, objects and scenes obtained in the construction space, as the attention parameters change, performing steps 221) to 225) multiple times, continuously updating objects and scenes, carefully observing the basis of existing objects and scenes, and generating new objects and new scenes with the newly evoked concepts and logic after evoking more similar concepts and perceptual logic.

6. A general intelligent agent according to claim 2, characterized in that: The cognitive submodule processing flow includes: 231) Call the perception submodule to obtain objects and scenes; 232) Deepen the construction of conditional and task scenarios, and simultaneously evoke relevant logic, including those directly evoked in the logic library and those indirectly extracted from the concept library involving specific scenario concepts involving judgment, reasoning or planning, and store them in the construction space; 233) Using the evoked logic as a template and related objects and scenes as materials, one generates judgments, reasoning, or planning with specific content within the constructed space, thus acquiring preliminary rational cognition; 234) Seeking new objects or new situations to fill in, update, support, and enrich constructed judgments, reasoning, or plans; 235) Forming judgments, reasoning, or planning as a specific scenario, forming new concepts based on existing concepts, and expanding the concept library; at the same time, finding patterns in judgments, reasoning, or planning, and updating and forming new logic based on existing logic, and expanding the logic library; 236) Steps 231) to 235) are performed in multiple rounds to analyze and consider the task and the generated judgment, reasoning, or planning in more detail.

7. A general intelligent agent according to claim 6, characterized in that: The specific process of step 233) is as follows: after the logic is invoked, based on the matching relationship between the logic and the object and scene, the elements in the logic are filled with the corresponding object and scene element sets, and the extended attribute value parameters in the logic element are calculated and the attributes of each element in the logic are updated. In this way, judgments, reasoning or planning containing specific content are generated in the construction space, and preliminary rational cognition is obtained; The new objects or new scenes in step 234) refer to the objects and scenes found that are not utilized by the attention logic arousal process, or the new clues, objects and scenes that are re-searched by calling the perception submodule back to the original input signal to change the attention parameters, and there are still no content-filled points in the judgment, reasoning or planning point set. New objects and scenes are generated as filling content by arousing related concepts through the concept library.

8. A general intelligent agent according to claim 1, characterized in that: The self-awareness module includes: 31) Proprioception submodule, used to receive proprioceptive signals, including navigation positioning, time perception, vision, hearing, touch, smell, taste, inertia, and balance signals from the intelligent body's trunk and various components; 32) The proprioception submodule uses proprioception signals and the proprioception model to estimate model parameters. It also calculates the position and posture of the intelligent body, including the trunk and various components, relative to other objects in the scene, based on the objects and scenes in the constructed space. It also modifies the body model and updates the body model library when the body model is not applicable. 33) Ontology cognition submodule: If ontology participation is required in the judgment, reasoning or planning process, the ontology is filled into the evoked logic as a content or element, and participates in the judgment, reasoning or planning process; 34) The main body action control submodule generates action sequences and hardware control program information from the planning results as main body control signal outputs when the main body needs to take action; 35) Ontology model library, which stores the structural parameter models of the intelligent body backbone and its components in terms of mathematics, physics, graphics, motion, and interaction.

9. A general intelligent agent according to claim 1, characterized in that: The subconscious module includes: 41) Subconscious sensory submodule, which uses a data-driven model to realize subconscious sensory functions; 42) Subconscious perception submodule, which uses a data-driven model to realize subconscious perception functions; 43) Subconscious cognition submodule, which uses data-driven models to realize subconscious cognitive functions; 44) Subconscious output modality conversion submodule, which uses a data-driven model to realize subconscious multimodal signal output function; The processing flow of the submodule 41) specifically includes: using the trained model parameters and obtaining a model output signal according to the model input signal; the model input signal is the input signal obtained by the input module, and the model output signal is a universal language sentence adapted to the clue element set or graph generated by the sensory submodule in the consciousness module; The processing flow of the submodule 42) specifically includes: using the trained model parameters and obtaining a model output signal based on the model input signal; the model input signal is respectively an input signal obtained by the input module or a general language sentence adapted to the clue element set or graph generated by the sensory submodule of the consciousness module; the model output signal includes a general language sentence adapted to the object, scene element set or graph generated by the perception submodule of the consciousness module; The processing flow of the submodule 43) specifically includes: using the trained model parameters and obtaining a model output signal according to the model input signal; the model input signal is respectively an input signal obtained by the input module or a general language sentence adapted to the clue element set or graph generated by the sensory submodule in the consciousness module, or a general language sentence adapted to the object, scene element set or graph generated by the perception submodule in the consciousness module; the model output signal includes a general language sentence adapted to the cognitive element set or graph generated by the cognitive submodule in the consciousness module; The submodule 44) processing flow specifically includes: using the trained model parameters to obtain the model output signal according to the model input signal; the model input signal is respectively the input signal obtained by the input module or the general language sentence adapted to the clue element point set or graph generated by the sensory submodule in the consciousness module or the general language sentence adapted to the object, scene element point set or graph generated by the perception submodule in the consciousness module or the general language sentence adapted to the cognitive element point set or graph generated by the cognitive submodule in the consciousness module, and the model output signal includes multimodal data adapted to the output signal of the output module.

10. A general agent control method according to claim 1, characterized in that: include: S1, memory pre-setting and agent pre-training; S2, value-driven task generation; S3, task-oriented autonomous decision-making; S4, task execution and self-evolution; S5, steps S2 to S4 are repeated multiple times, actively interacting with the environment and the user; The step S1 specifically includes: pre-setting the concept library, logic library, and ontology model library in the long-term memory submodule of the consciousness module, and setting attribute values ​​according to requirements; for the network model contained in each submodule of the subconscious module, collecting model input data and expected model output data, training the model, and updating model parameters; The step S2 specifically includes: performing task analysis on the goals, tasks, instructions arranged by the user, or on the relevant task scenarios evoked in the memory bank based on the observed objects and scenarios, and selecting an appropriate task through value judgment; Said step S3 specifically includes: autonomously generating judgment, reasoning, and planning for the selected task, breaking it down into specific steps and movement and operation action sequences for task solving; The step S4 specifically includes: the intelligent agent outputting a body control signal to perform an action or directly completing a cognitive task in the construction space, and during task execution, by observing scene construction and action control errors, adaptively adjusting the model parameters of the consciousness module memory bank and the subconscious module to achieve self-evolution; The step S5 specifically includes: performing steps S2 to S4 in multiple rounds, wherein the intelligent agent actively interacts with the environment and the user, continuously generates tasks and makes autonomous decisions to execute them. During the entire process, the intelligent agent's thinking is transparent and controllable, and the intelligent agent adjusts its attributes through learning.

Citation Information

Patent Citations

  • Method for realizing humanoid universal artificial intelligence

    CN112215346A

  • Intelligent agent control method and device, computer equipment and storage medium

    CN117391181A

  • Reinforcement learning navigation method based on target driving

    CN117475279A

  • Universal intelligent agent and control method thereof

    CN118261192A

  • New function of recognition module

    JP2019185701A

Cited By

  • Multi-agent cooperation method and device for electromagnetic spectrum monitoring and analysis

    CN121256512A

  • Graph structure-based intelligent decision-making method, apparatus and device for body, and medium

    CN121456805A

  • Photoelectric interconnection cooperative training system and method for multi-mode intelligent agent network

    CN121509269A

  • An AI-native real-time operating system and its control method

    CN122412058A