Method and device for interaction, equipment, storage medium and program product

By constructing a memory database and multimodal memory agent based on identification information, the problem of insufficient information preservation and reasoning capabilities of intelligent systems in long-term interactions is solved, and the high adaptability and coherence of intelligent systems in cross-time periods and multi-round interactions are achieved.

CN120848723APending Publication Date: 2025-10-28BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510933784.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-07
Publication Date
2025-10-28

Smart Images

  • Figure CN120848723A_ABST
    Figure CN120848723A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a method, device and equipment for interaction, a storage medium and a program product. The method comprises the steps of obtaining identification information related to interaction input in response to the received interaction input for a target object; determining target memory information from a memory database based on the identification information, the memory database being constructed based on historical interaction performed with the target object; and determining a response to the interactive input based on the target memory information and the interactive input.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The exemplary embodiments disclosed herein generally relate to the field of computers, and particularly to methods, apparatuses, electronic devices, computer-readable storage media, and computer program products for interaction. Background Technology

[0002] With the rapid development of machine learning and multimodal perception technologies, various intelligent systems are widely used in scenarios such as mobile robots, intelligent dialogue, and smart homes. These systems typically need to continuously receive multimodal input information from the environment and generate responses or complete specific tasks based on the perceived information. In practical applications, intelligent systems often operate in a long-term, continuous interactive state. Therefore, how to enable intelligent systems to have the ability to interact for extended periods in dynamic environments has become a noteworthy issue. Summary of the Invention

[0003] In a first aspect of this disclosure, an interaction method is provided. The method includes: in response to receiving an interactive input for a target object, acquiring identification information related to the interactive input; determining target memory information from a memory database based on the identification information, the memory database being constructed based on historical interactions performed with the target object; and determining a response to the interactive input based on the target memory information and the interactive input.

[0004] In a second aspect of this disclosure, a task execution method is provided. The method includes: in response to receiving target input for a movable target object, acquiring identification information related to the target input; determining target memory information from a memory database based on the identification information; determining a target task for the target object based on the target memory information and the target input; and controlling the target object to move towards a target location corresponding to the target task to execute the target task.

[0005] In a third aspect of this disclosure, an apparatus for interaction is provided. The apparatus includes: an identification information acquisition module configured to acquire identification information related to the interaction input in response to receiving an interaction input for a target object; a memory information determination module configured to determine target memory information from a memory database based on the identification information, the memory database being constructed based on historical interactions performed with the target object; and a response module configured to determine a response to the interaction input based on the target memory information and the interaction input.

[0006] In a fourth aspect of this disclosure, an apparatus for task execution is provided. The apparatus includes: an identification information acquisition module configured to acquire and emit identification information related to the target input in response to receiving target input for a movable target object; a memory information determination module configured to determine target memory information from a memory database based on the identification information; a task determination module configured to determine a target task for the target object based on the target memory information and the target input; and an execution module configured to control the target object to move towards a target location corresponding to the target task to execute the target task.

[0007] In a fifth aspect of this disclosure, an electronic device is provided. The device includes at least one processor; and at least one memory coupled to the at least one processor and storing instructions for execution by the at least one processor. When executed by the at least one processing unit, the instructions cause the electronic device to perform the methods of the first or second aspect.

[0008] In a sixth aspect of this disclosure, a computer-readable storage medium is provided. The medium stores computer instructions that, when executed by a processor, implement the method of the first or second aspect.

[0009] In a seventh aspect of this disclosure, a computer program product is provided. The product includes a computer program, which, when executed by a processor, implements the method according to a first or second aspect of this disclosure.

[0010] It should be understood that the content described in this content section is not intended to limit the key or essential features of the embodiments of this disclosure, nor is it intended to restrict the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description

[0011] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. In the drawings, the same or similar reference numerals denote the same or similar elements, wherein:

[0012] Figure 1 A schematic diagram of an example environment in which embodiments of the present disclosure can be implemented is shown;

[0013] Figure 2 A schematic diagram illustrating an example process of interaction according to some embodiments of the present disclosure is shown;

[0014] Figure 3 A schematic diagram of a structured data object according to some embodiments of the present disclosure is shown;

[0015] Figure 4A flowchart of a method for interaction according to some embodiments of the present disclosure is shown;

[0016] Figure 5 A flowchart of a method for task execution according to some embodiments of the present disclosure is shown;

[0017] Figure 6 A schematic structural block diagram of an interactive device according to some embodiments of the present disclosure is shown;

[0018] Figure 7 A schematic structural block diagram of an apparatus for task execution according to some embodiments of the present disclosure is shown; and

[0019] Figure 8 A block diagram of an electronic device in which one or more embodiments of the present disclosure may be implemented is shown. Detailed Implementation

[0020] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.

[0021] In the description of embodiments of this disclosure, the term "comprising" and similar terms should be understood as open-ended inclusion, i.e., "including but not limited to". The term "based on" should be understood as "at least partially based on". The term "one embodiment" or "the embodiment" should be understood as "at least one embodiment". The term "some embodiments" should be understood as "at least some embodiments". Other explicit and implicit definitions may also be included below.

[0022] In this document, unless explicitly stated otherwise, performing a step in response to A does not mean that the step is performed immediately after A, but may include one or more intermediate steps.

[0023] It is understood that the data involved in this technical solution (including but not limited to the data itself, the acquisition, use, storage or deletion of the data) shall comply with the requirements of relevant laws, regulations and related provisions.

[0024] It is understood that before using the technical solutions disclosed in the various embodiments of this disclosure, relevant users should be informed of the type, scope of use, and usage scenarios of the information involved in this disclosure through appropriate means in accordance with relevant laws and regulations, and authorization should be obtained from the relevant users. Among them, relevant users may include any type of rights holder, such as individuals, enterprises, and groups.

[0025] For example, in response to receiving an active request from a user, a prompt message is sent to the relevant user to clearly inform the user that the requested operation will require obtaining and using the user's information, thereby enabling the relevant user to choose whether to provide information to the software or hardware such as the electronic device, application, server, or storage medium that performs the operation of the technical solution disclosed herein based on the prompt message.

[0026] As an optional but non-restrictive implementation, in response to a user's active request, a prompt message can be sent to the user, such as a pop-up window, where the prompt message can be presented in text format. Furthermore, the pop-up window can also include a selection control allowing the user to choose "agree" or "disagree" to provide information to the electronic device.

[0027] It is understood that the above notification and user authorization process are merely illustrative and do not constitute a limitation on the implementation of this disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of this disclosure.

[0028] As used in this paper, the term "model" refers to a model that learns the relationship between inputs and outputs from training data, enabling it to generate corresponding outputs for a given input after training. Model generation can be based on machine learning techniques. Deep learning is a machine learning algorithm that processes inputs and provides corresponding outputs using multiple layers of processing units. A neural network model is an example of a deep learning-based model. In this paper, "model" may also be referred to as a "machine learning model," "learning model," "machine learning network," or "learning network," and these terms are used interchangeably.

[0029] Figure 1 A schematic diagram of an example environment 100 in which embodiments of the present disclosure can be implemented is shown. In this example environment 100, an application 120 is installed on a terminal device 110. A user 140 can interact with the application 120 via the terminal device 110 and / or an attached device of the terminal device 110.

[0030] In some embodiments, application 120 can be downloaded and installed on terminal device 110. In some embodiments, application 120 can also be accessed in other ways, such as through a web page. Figure 1 In environment 100, in response to application 120 being launched, terminal device 110 can display the interface 150 of application 120.

[0031] In some embodiments, terminal device 110 may be any type of mobile terminal, fixed terminal, or portable terminal, including mobile phones, desktop computers, laptop computers, notebook computers, netbook computers, tablet computers, media computers, multimedia tablets, personal communication system (PCS) devices, personal navigation devices, personal digital assistants (PDAs), audio / video players, digital cameras / camcorders, positioning devices, television receivers, radio receivers, e-book devices, gaming devices, or any combination thereof, including accessories and peripherals of these devices or any combination thereof. In some embodiments, terminal device 110 may also support any type of user-facing interface (such as "wearable" circuitry).

[0032] In some embodiments, terminal device 110 can communicate with server 130 to provide services to application 120. Terminal device 110 can be any type of mobile terminal, fixed terminal, or portable terminal, including mobile phones, desktop computers, laptop computers, notebook computers, netbook computers, tablet computers, media computers, multimedia tablets, personal communication system (PCS) devices, personal navigation devices, personal digital assistants (PDAs), audio / video players, digital cameras / camcorders, television receivers, radio receivers, e-book devices, gaming devices, or any combination thereof, including accessories and peripherals of these devices or any combination thereof. In some embodiments, terminal device 110 can also support any type of user-facing interface (such as "wearable" circuitry). Application 120 can be various types of computing systems / servers capable of providing computing power, including but not limited to mainframes, edge computing nodes, computing devices in cloud environments, etc.

[0033] In some embodiments, server 130 or functions coordinated by server 130 in terminal device 110 (e.g., application 120) can implement a dialogue system to provide a dialogue between user 140 and a digital assistant. For example, terminal device 110 can acquire user input, such as voice commands, text questions, or other forms of interactive content, and provide the user input to server 130. Server 130 can then provide a generated response to the provided user input to terminal device 110 for presentation to the user. This dialogue interaction can be implemented in the form of a digital assistant.

[0034] It should be understood that Figure 1 The environment 100 shown is merely exemplary and is not intended to be limiting. In some embodiments, server 130 may interact with user 140 without requiring the application 120 or interface 150 in terminal device 110.

[0035] In some embodiments, terminal device 110 may include or be included within a movable object deployed in environment 100 (e.g., a robot). In some embodiments, terminal device 110 may be a control device for the movable object (e.g., a robot). The control device may be deployed within the movable object or separately from it. For example, terminal device 110 may be any type of microcontroller, programmable logic controller, industrial computer, single-board computer, field-programmable gate array, digital signal processor, multi-core processor, or any combination thereof, including accessories and peripherals of these devices or any combination thereof. In some embodiments, terminal device 110 may control the movable object based on control commands it generates.

[0036] In some embodiments, the operation of the movable object can be controlled via server 130. Server 130 can receive relevant data for various tasks and remotely control the operation of the movable object. Terminal device 110 can collect terminal device 110 data and communicate with server 130 to send data to server 130. Server 130 determines control instructions for terminal device 110 based on the acquired data and sends the control instructions to terminal device 110. Terminal device 110 controls the movable object based on the received control instructions.

[0037] It should be understood that the structure and function of the various elements in environment 100 are described for illustrative purposes only and do not imply any limitation on the scope of this disclosure.

[0038] As briefly mentioned earlier, with the continuous development of artificial intelligence and multimodal perception technologies, various intelligent systems with perception, understanding, and decision-making capabilities are widely used in various scenarios such as mobile robots, intelligent dialogue, and smart homes. In these scenarios, intelligent systems often need to operate continuously, interacting with interactive objects and the environment for extended periods. Based on the perceived information or received instructions, intelligent systems may also need to perform multiple functions such as dialogue responses, instruction responses, and task execution.

[0039] In practical applications, intelligent systems often face multimodal inputs, including video, audio, and text information, and may also include complex inputs such as scene perception information and sensor status. Furthermore, many tasks that intelligent systems need to perform rely on a grasp of long-term historical information to make accurate judgments. Intelligent systems should have the ability to understand the needs and habits of different users. For example, warehouse robots should remember obstacle areas in their paths. Dialogue systems should continuously understand the identity of the current interaction partner and the context of the interaction.

[0040] Traditional intelligent systems often process information only within the current or short timeframes, lacking the ability to systematically retain and mobilize long-term interaction histories. To extend the agent's perception timeframe, some solutions attempt to reduce input size by compressing redundant modal information or to improve processing length limits by expanding the model's context window. Additionally, some solutions introduce the approach of segmenting long-term multimodal content into several fragments and using pre-trained models to generate text descriptions for each fragment, thus constructing memories in discrete text form. However, these solutions still have significant limitations in practical applications. For example, compression mechanisms may result in the loss of semantic information, affecting subsequent understanding. Model window expansion methods struggle to support memory needs spanning days or even lifetimes. Fragment-based description solutions, lacking a unified global perspective, are prone to memory fragmentation or disjointed events.

[0041] Furthermore, intelligent systems tend to use memory information passively during task execution. Typically, they only retrieve information through simple keyword recall or static retrieval, lacking the ability to comprehensively reason about the content stored in memory. Therefore, a solution is needed for intelligent systems to handle interaction or task execution. On one hand, this solution should enable intelligent systems to reliably and structurally store information across modalities and time periods, constructing a dynamically updated and lifelong expandable long-term memory system. On the other hand, it should enable intelligent systems to combine long-term memory for reasoning and adaptive behavioral decision-making during task execution, thus enabling them to handle more complex and realistic intelligent tasks.

[0042] In view of this, according to embodiments of the present disclosure, an improved interaction scheme is provided. According to this scheme, firstly, in response to receiving interactive input for a target object, identification information related to the interactive input is obtained. Further, based on the identification information, target memory information is determined from a memory database, which is constructed based on historical interactions performed with the target object. Still further, based on the target memory information and the interactive input, a response to the interactive input is determined.

[0043] Therefore, by retrieving relevant memory information from the long-term memory database based on identification information and performing reasoning, more context-relevant and continuous response content can be generated. This approach not only improves the adaptability and coherence of intelligent systems in cross-time and multi-round interactions, but also enhances their depth of understanding of user intent and the intelligence of task responses.

[0044] The following description will continue with reference to the accompanying drawings, which will provide some exemplary embodiments of this disclosure. Figure 2A schematic diagram of an interactive process 200 according to some embodiments of the present disclosure is shown. Part or all of process 200 may be implemented by server 130 or by other devices, such as remote devices (terminal devices or service devices) with computing capabilities. Hereinafter, for ease of discussion, example embodiments will be described primarily with respect to server 130. It should be understood that the actions described with respect to server 130 may be performed at least partially by other entities, such as application 120, or by server 130 in coordination with other terminal devices or servers.

[0045] like Figure 2 As shown, server 130 can construct memory database 230. Memory database 230 can be used to store memory information related to multiple modalities associated with one or more interactive entities. Memory information can refer to data content generated and stored by the intelligent system during operation, based on external input or feedback during task execution. This information can originate from various modalities, including but not limited to video frames, audio clips, text descriptions, image features, or sensor data.

[0046] Because intelligent systems face complex scenarios in practical applications, such as long-term operation, multi-turn interaction, and multimodal input, relying on short-term memory information for task response or behavioral decision-making often fails to meet requirements such as context consistency and intelligent reasoning. To address this, server 130 can store memory information in memory database 230 and construct a long-term memory system that can be dynamically updated and continuously expanded throughout its lifecycle.

[0047] Continue to refer to Figure 2 The server 130 can continuously receive historical inputs 204 from multiple modalities. The historical inputs 204 can originate from video frames, audio clips, text dialogues, voice dialogues collected during multi-round interactions, or from sensor status information during task execution. Furthermore, the server 130 can determine the modality of the historical inputs 204. A modality includes at least one of multiple predetermined modalities. The server 130 can parse and classify the received historical inputs 204 to determine the modality involved, such as image modality, voice modality, text modality, etc.

[0048] Continue to refer to Figure 2Server 130 can utilize memory agent 210 to determine memory information 205 of the modality of historical input based on historical input 204. Memory agent 210 has the ability to stream multimodal input data. Therefore, memory agent 210 can process historical input 204 in real-time or near real-time, such as video data, audio signals, text interaction records, sensor data, or other perceptual modal inputs. After receiving these inputs, memory agent 210 can generate structured memory information 205 based on preset strategies and intelligent judgment mechanisms.

[0049] In some embodiments, the memory database 230 is constructed using multiple functional blocks corresponding to multiple modalities. These functional blocks can be viewed as processing tools designed or configured for different input modalities (such as visual, auditory, text, etc.) to support operations such as modality perception, feature extraction, information encoding, and semantic construction. Each functional block can process the data content of its corresponding modality and output corresponding information for unified management and storage by the server 130. For example, the multiple functional blocks may include an image processing module, an audio recognition module, a retrieval module, an anomaly detection module, and so on.

[0050] In some embodiments, the memory information stored in the memory database 230 may correspond to identification information. Identification information refers to information that can be used to identify an object, a task target, a device, or each area (e.g., the area where a robot performs a task). Identification information ensures that the system can accurately identify and distinguish memory information corresponding to different interactive objects or tasks, thereby enabling correct memory storage, retrieval, and processing. For example, identification information may be information indicating a destination in the environment, or identification of items involved in the input. As another example, identification information may be information used to mark a specific task or event.

[0051] In some embodiments, server 130 may utilize memory agent 210 to determine at least one piece of historical memory information related to the identification information from memory database 230. Server 130 may determine memory information 205 based on at least one piece of historical memory information by utilizing memory agent to call function blocks corresponding to the modalities of historical inputs from multiple function blocks.

[0052] In the process of generating memory information, the memory agent 210 can rely not only on the received input but also access historical memory information stored in the memory database 230 to assist in deciding "which content is worth remembering." For example, the memory agent 210 uses an image processing module or an audio recognition module to determine whether the current interactive input is similar to the identification information in the memory information in the memory database 230. In this case, the memory agent 210 can utilize historical information related to that identification information. In this way, the semantic consistency of the newly generated memory information 205 can be enhanced.

[0053] Alternatively or additionally, the memory agent 210 can autonomously decide whether to record a certain input, or in what modality and granularity to encode and store it, based on a memory strategy. For example, for input content that is highly repetitive or contains no new information, the memory agent 210 can choose to skip memorization.

[0054] In some embodiments, memory information 205 may indicate at least one of descriptive information about historical input 204 or summary information determined based on historical input 204. These two types of content together constitute the core foundation of an agent's long-term memory.

[0055] The descriptive information of historical input 204 can be understood as "episodic memory," which refers to the record of events experienced by an intelligent system (e.g., a conversational agent or a mobile robot) at a specific time and place. Episodic memory emphasizes the reconstructive description of specific events. Episodic memory can typically manifest as a multimodal detailed expression of objects, behaviors, and environments involved in input data (such as a video clip, a voice interaction, or a motion trajectory). For example, after receiving an interactive video, the memory agent 210 can generate a natural language description explaining which characters appeared in the video, what actions were performed, and what environment they were in. As another example, in an embodied intelligence scenario, a warehouse robot collects sensor data, video recordings, and navigation trajectories while inspecting shelves. The memory agent 210 can generate corresponding episodic memory, recording when and where the robot performed what task, and what the result was. This descriptive information can include not only behavioral information but also historical memory information stored in the memory database 230 (e.g., interaction objects, locations, behavioral models, scene attributes, etc.), making the generated memory more structured, readable, and semantically consistent.

[0056] The summarized information determined based on historical input 204 can also be understood as "semantic memory." Semantic memory refers to a higher-level understanding of the environment, interactive objects, concepts, or events formed through the analysis and summarization of historical input data (especially episodic memory). Semantic memory may not correspond to a specific moment or a single event, but can reflect a generalizable knowledge representation accumulated by an intelligent system over a long period of operation. For example, the memory agent 210 can summarize from multiple interactions that "a meeting is held every Monday morning."

[0057] As an example, in embodied intelligence scenarios, warehouse robots can use data acquired over long periods of task execution to summarize information such as which types of goods are often placed in specific areas, or which passageways are prone to congestion during certain time periods, thus forming semantic memories. For example, "Area C is often used for temporary storage of fragile items" or "Aisles in Area B experience heavy traffic during peak hours." This semantic knowledge may not originate from a single specific task, but rather be formed through comparison, induction, and reasoning of multiple episode memories.

[0058] In some embodiments, the memory database 230 may include structured data objects. The structured objects include multiple memory nodes, each corresponding to a different piece of memory information. As an example, Figure 3 A schematic diagram of a structured data object 300 according to some embodiments of the present disclosure is shown.

[0059] like Figure 3 As shown, the memory database 230 may include structured data objects 300 for organizing multimodal memory information in a unified and scalable form. The structured data objects 300 may contain multiple memory nodes, each corresponding to a specific piece of memory information. Each memory node can represent specific memory information within a particular modality. The memory information corresponding to each memory node may not only exist in natural language form but may also include multimodal information such as images, audio clips, and semantic tags. For example, an image node may represent a keyframe or image. An audio node may represent an audio clip. A text node may represent an event description, a semantic summary, or a user attribute tag.

[0060] Continue to refer to Figure 3 Server 130 can construct multiple associated nodes for the same identification information. These memory nodes are connected by edges to represent the relationships between nodes. For example, image node 1, which represents image features related to user A, can be associated with audio node 1, which represents user A's voice commands, and text node 1, which represents user A's task requirements, thereby forming a cross-modal, multi-dimensional memory graph in the structured data object.

[0061] In some embodiments, after determining new memory information 205, server 130 can determine the association between memory information 205 and memory information of multiple modalities stored in memory database 230. Further, server 130 can store memory information 205 in memory database 230 according to the association. After continuously determining the memory information 205 corresponding to the received input 204 using memory agent 210, server 130 can analyze the association between memory information 205 and the multimodal memory information already stored in memory database 203. For example, the association may include, but is not limited to, matching and linking based on dimensions such as object identification, event sequence, and semantic content similarity. After determining the association between memory information 205 and the multimodal memory information already stored in memory database 203, server 130 can store or merge memory information 205 into memory database 230 according to its association with existing memory information.

[0062] For example, if the association indicates that memory information 205 has high similarity or semantic coherence with a memory node in the database, server 130 can update the memory information corresponding to that memory node. If no relevant node is detected, server 130 can create a new memory node and add its connection edges with other nodes in database 230, thereby continuously enriching and expanding the structure and coverage of the long-term memory graph. Through this process, server 130 can not only achieve long-term retention of multimodal historical inputs, but also realize the structured organization and dynamic evolution of semantic relationships between information. In this way, the ability of intelligent agents to understand cross-modal information, perform multi-turn reasoning, and make task decisions can be further improved.

[0063] In some embodiments, in response to an association indicating the existence of at least one piece of memory information related to memory information 205 in the memory database, server 130 may update at least one memory node and at least one edge of the memory node corresponding to the at least one piece of memory information in the structured data object based on the memory information 205 and the association. As an example, the update operation may include, but is not limited to, replacing or supplementing the content of the memory information represented by the memory node, for example, updating the image feature vector; strengthening the edge weights between nodes to indicate that the association is re-validated in a new context; or adding associated edges to introduce more contextual interpretation.

[0064] Alternatively or additionally, in response to the absence of memory information related to memory information 205 in the association indication memory database 230, server 130 may add at least one memory node corresponding to memory information 205 to the structured data object. For example, in structured data object 300, image node 2, audio node 2, and text node 2 may be added. These may correspond to image, audio, and text information of another object in a new interaction. Server 130 may create initial edge relationships for these new nodes. Alternatively or additionally, server 130 may dynamically complete the connection relationships of these memory nodes in the graph structure based on continuously received multimodal input data.

[0065] In some embodiments, server 130 can identify and correct conflicting modal associations. For example, in structured data object 300, image node 1, representing task-related image information of identifier A, and audio node 1, representing task-related audio information of identifier A, are connected by an edge, meaning that the image feature and the audio feature correspond to the same identifier information. However, in subsequent interactions, a new potential pairing may be detected between image node 1 and audio node 2. In this case, server 130 can record the conflicting relationship and perform posterior correction through statistical co-occurrence frequency, confidence assessment, or voting mechanisms, ultimately deciding which relationship to retain or update.

[0066] In the embodiments of this disclosure, a long-term memory mechanism with scalability and self-evolution is achieved by storing multimodal memory information in the form of structured data objects and supporting continuous updates and error correction of memory content based on new inputs. On the one hand, this enables the agent to orderly integrate information from different modalities and spanning different time periods to form a complete and rich memory system. On the other hand, by identifying and correcting conflicting relationships posteriorly, the memory structure can be dynamically adjusted to ensure the accuracy and consistency of the memory content when facing practical problems such as modality recognition errors and contextual biases.

[0067] Return to reference Figure 2While determining memory information 205 based on historical input 204 and storing it in the memory database 230, server 130 can simultaneously receive input indicating real-time interaction or task requirements. For example, while server 130 uses memory agent 210 to analyze and process historical input 204 and generate episodes or semantic memories, it does not interrupt its responsiveness to system interactions. This parallel mechanism allows server 130 to continuously process real-time input from users or the environment, such as voice commands, action instructions, or state triggers, while continuously building and updating long-term memories. Furthermore, server 130 can recall memory information related to the current input from the memory information stored in memory database 230 and generate context-consistent response results or decision behaviors accordingly.

[0068] In some embodiments, server 130 may, in response to receiving interactive input targeting a target object, acquire identification information associated with the interactive input. Each interactive input may correspond to one piece of identification information. The identification information may be used to identify information such as the target object or target item in the interaction. Through the identification information, server 130 can more accurately organize, store, and retrieve memory information, thereby providing a personalized response for each interaction.

[0069] In some embodiments, the target object may be an intelligent dialogue system such as a conversational agent or virtual assistant, or a smart device such as a smart speaker, smart glasses, or smartwatch. The target object may rely on personalized responses or information services based on user input and identity.

[0070] In some embodiments, based on the aforementioned long-term memory system, embodied intelligence (or embodied agent) can possess long-term memory capabilities. This long-term memory capability enables embodied intelligence not only to process current user input and perceived environmental data, but also to combine long-term memory information such as historical interaction experience, task execution records, and user needs to complete more complex and intelligent behavioral decisions and task execution in the actual physical environment.

[0071] Embodied intelligence refers to a physical entity intelligent agent possessing the capabilities of perception, cognition, decision-making, and execution. Embodied intelligent agents typically possess a mobile and perceptible physical form and can include warehouse robots, inspection robots, service robots, delivery robots, unmanned cleaning equipment, and intelligent security terminals. These intelligent agents can not only receive task instructions from users but also perform tasks such as path planning, obstacle avoidance, and target interaction in real-world environments. Utilizing long-term memory mechanisms, embodied intelligent agents can remember historical states related to users and the environment, thereby achieving more efficient and safer intelligent control in complex tasks.

[0072] In some embodiments, the target object can be an embodied intelligent agent, such as a warehouse robot, inspection robot, or service robot. In response to receiving target input for the target object, the server 130 can determine identification information related to the target input. Through the identification information, the server 130 can more accurately obtain memory information related to the target input.

[0073] In some embodiments, the target input corresponds to a first modality among multiple modalities. Server 130 may utilize a device associated with the target object to acquire feature information related to the target input. This feature information may correspond to a second modality among multiple modalities, which differs from the first modality. Server 130 may determine identification information based on the feature information. For example, after a user issues a task command in voice or text (which is an example of the first modality), server 130 may determine the corresponding identification information based on the target object's operating environment or task characteristics.

[0074] In some embodiments, server 130 can determine target memory information from memory database 230 based on identification information. Memory database 230 is constructed based on historical interactions performed with the target object. Based on the identification information, server 130 can identify context or task information related to the target input. Server 130 can then retrieve and determine target memory information related to the current task execution or interaction context from memory database 230.

[0075] As an example, when a user interacts with a virtual assistant through a smart speaker, server 130 can determine identification information related to the target input 201 as the user utters voice input 201, in order to identify the user's identity. Furthermore, server 130 can retrieve relevant memory information from memory database 230 to better understand the user's needs and generate a more user-friendly target response 202.

[0076] As another example, upon receiving the target input 201, the server 130 can receive environmental identification information (e.g., current location, task target location, or current environmental factors) obtained through sensing devices (e.g., head-mounted camera, array microphone, sensors, etc.) attached to the target object (e.g., the robot body). Furthermore, the server 130 can obtain information such as the operator's past task records, path habits, and work preferences to optimize current path planning, work behavior, or collaborative scheduling strategies.

[0077] In some embodiments, the memory database 230 includes memory information of multiple modalities, such as images, audio, text, etc. Each modal of memory information can construct corresponding memory nodes and associated edges around a specific entity in the memory database 230. The server 130 can generate at least one query based on the category of the identification information. The at least one query indicates retrieving at least one piece of memory information corresponding to the category of the identification information from the memory information of multiple modalities. The server 130 can generate at least one query based on the category of the identification information (e.g., audio, image, etc.) after receiving the target input 201. This query is used to retrieve at least one piece of memory information corresponding to that category from the memory information of multiple modalities. The server 130 can retrieve at least one piece of memory information corresponding to the category of the identification information from the memory information of multiple modalities based on at least one query.

[0078] As an example, when receiving a user's voice input "Get me a cup of coffee," server 130 can first extract relevant audio features based on the voice segment and identify the associated identifier as coffee. Then, server 130 can construct a first query based on this identifier. For example, the first query could be: finding memory nodes with an audio modality related to the current device and the identifier "coffee," determining corresponding memory information from the memory database, such as the location of the coffee machine, and returning information from other associated memory nodes, such as different types of coffee pairings. Server 130 can determine target memory information based on at least one piece of memory information.

[0079] In some examples, server 130 can generate more specific queries based on the acquired memory information. For example, a second query could be: coffee + coffee machine B, which attempts to retrieve memory information related to "coffee machine B" and "coffee" from long-term memory. This query might hit memory information such as: an episodic memory: "Yesterday at 10 o'clock, coffee machine B made black coffee," or a semantic memory: "The preferred beverage of coffee machine B is black coffee." Server 130 can further perform contextual reasoning and semantic integration based on the retrieval results through memory agent 220. For example, server 130 determines that the current user's desired coffee is black coffee, and thus determines the target response decision: "Make black coffee."

[0080] In the embodiments of this disclosure, the query constructed by server 130 can not only be directly generated based on identification information, but also undergo multi-hop semantic expansion according to the task context. In this way, layer-by-layer memory tracing of cross-modal semantic associations can be achieved. In each hop of retrieval, server 130 can dynamically evaluate the relevance and confidence of the hit results and flexibly adjust the query strategy until a complete semantic input is formed to support the generation of task behavior.

[0081] In some embodiments, server 130 can determine the response to the interactive input based on target memory information and interactive input. After identifying identification information related to the interactive input, server 130 can recall target memory information related to that identification information from memory database 230. Target memory information includes, for example, past interaction content, demand preference records, the most recent unfinished request, common expressions, etc. Subsequently, the executing agent 220 can semantically fuse the target memory information with the current interactive input, and determine the response result that best meets the user's needs through context matching, intent understanding, and multi-turn tracing mechanisms. The executing agent 220 can be used to receive the current interactive input or target task request, and combine it with the target memory information recalled in memory database 230 to perform operations such as task reasoning, intent understanding, behavior decision-making, and control execution. In some examples, the executing agent 220 and the memory agent 210 can be the same agent, or two agent modules working collaboratively under the same agent system architecture.

[0082] In some embodiments, server 130 can determine a target task for a target object based on target memory information and target input. Server 130 can control the target object to move to the target location corresponding to the target task in order to perform the target task. For embodied intelligence such as warehouse robots, server 130 can extract relevant target memory information from memory database 230 based on the received target input. Target memory information includes, for example, information such as historical task execution results, task execution priority areas, and task execution habitual paths.

[0083] In some embodiments, the target location is determined based on at least one piece of memory information related to the identification information in the memory database 230. The server 130 can determine the target location that the target object needs to go to based on at least one piece of memory information related to the identification information in the memory database 230. That is, the target location in the task is not necessarily given by the user explicitly inputting specific coordinates or area information, but can be inferred by combining the user's semantic expression and existing memory information. For example, the user inputs "Please put this package on shelf A" via voice command. Although the user input does not provide a specific area number or location coordinates, the server 130 can identify the identification information as shelf A and retrieve location information (such as floor, area, etc.) related to shelf A from the memory database 230, thereby generating specific navigation coordinates or route planning instructions.

[0084] Server 130 can combine the target input with relevant memory information and, through the execution of agent 220, make a comprehensive judgment to determine the optimal target task. The target task can indicate the target location or behavioral path of the target object. Furthermore, server 130 can generate control commands for the target object, scheduling the target object to move to the target location and perform navigation, handling, and other operations as planned. For example, server 130 can schedule a warehouse robot to transport items to a designated warehouse location. During task execution, the warehouse robot can refer to past handling trajectories, congestion area records, and historical obstacle perception data based on memory information, thereby avoiding areas prone to congestion or equipment failure, improving path safety and task completion efficiency.

[0085] In some embodiments, server 130 can perform a reflection process based on feedback information during the interaction process or task execution process, thereby further improving the accuracy and adaptability of long-term memory. For example, server 130 can evaluate and adjust relevant memory information in memory database 230 based on feedback information such as user evaluation, task result status, and anomaly detection results, in order to support the updating, correction, or supplementation of existing long-term memories.

[0086] In some embodiments, server 130 may determine first memory information for a target object based on feedback information 203 to the target response 202. The target response 202 may include system-generated response operations, such as dialogue responses, device control, etc. Alternatively or additionally, the target response 202 may include a target task performed by an embodied intelligent agent, such as a transport path, inspection task, etc.

[0087] In some embodiments, feedback information 203 may include information such as evaluation of the result of the target response 202 and description of deviations. Feedback information 203 may be provided by an interactive object, such as user-initiated evaluation or error correction input. Alternatively or additionally, feedback information 203 may include state feedback from the environment in which the target object is located in response to the target response 202. Feedback information 203 may be automatically collected and determined by the device. For example, the robot may detect abnormal events such as obstacles, collisions, or inability to complete instructions during path execution, or the state changes recorded by environmental sensors may not match the expected results of the target task.

[0088] In some embodiments, server 130 can analyze, based on feedback information 203, whether there are semantic understanding errors, user demand deviations, or unreasonable task paths in the current task or response process, and then determine the corresponding memory information. Based on the first memory information, server 130 can update at least a portion of the memory information related to the identification information in memory database 230. Server 130 can correct, supplement, or reconstruct the memory information related to the identification information in memory database 230 based on the first memory information. This may include updating existing memory nodes, adding new memory nodes, adjusting the weights or confidence levels of edges, or merging redundant information. Through this reflective mechanism, it is possible to achieve autonomous perception and self-repair of errors, omissions, or outdated information in previous memories. In this way, a more reliable and dynamically evolving long-term memory system can be constructed, thereby supporting higher-quality reasoning and task decision-making.

[0089] In some embodiments, server 130 may determine whether there is memory information in memory database 230 that conflicts with the first memory information and is related to the identification information. If it is determined that there is second memory information in memory database 230 that conflicts with the first memory information and is related to the identification information, server 130 may update the second memory information based at least on the corresponding confidence levels determined in historical interactions between the first memory information and the second memory information.

[0090] A conflict with the first memory information and its relation to the identification information refers to a situation where, upon recognizing new memory information, it is detected that this new information has an inconsistency or contradiction at the semantic, factual, or logical level with existing memory information in the memory database 230 that is associated with the same identification information (e.g., second memory information). For example, the first and second memory information represent different values ​​or descriptions for the same attribute or context. Another example is the discrepancy between the descriptions of an object, task, location, or item by the first and second memory information. In such cases, the server 130 can update the second memory information based on the confidence levels formed by the first and second memory information in historical interactions. The confidence level quantifies the reliability of a memory information and can be comprehensively evaluated based on one or more factors, such as the number of times the memory information is mentioned or triggered in historical interactions, the weight of the memory information, and user feedback annotations.

[0091] As an example, if the confidence level of the second memory information is less than that of the newly determined first memory information, meaning the second memory information is judged to be incorrect, the server 130 can directly overwrite or correct the value of the second memory information with the first memory information. Alternatively or additionally, if the confidence level of the second memory information is greater than or equal to the confidence level of the newly determined first memory information, meaning the second memory information is still considered reliable, the server 130 can retain the second memory information, only adjusting its corresponding confidence level, assessment weight, or conflict flag, thereby providing a confidence level reference for subsequent decision-making reasoning.

[0092] In some embodiments, if it is determined that there is no memory information in memory database 230 that is related to the identification information and conflicts with the first memory information, server 130 may store the first memory information in memory database 230 in association with the identification information. If server 130 determines that there is no memory information in memory database 230 that is related to the current identification information and conflicts with the first memory information, it indicates that the first memory information is new memory information or newly added cognition. Server 130 may store the first memory information in memory database 230 in association with the interaction identification information for subsequent interaction or task reasoning invocation. For example, server 130 may use the memory node corresponding to the first memory information as a new node in structured data object 300 and establish connection edges with existing multimodal nodes to enrich the long-term memory of the corresponding identification information.

[0093] In some embodiments, server 130 can perform reinforcement learning training on the intelligent system (e.g., an intelligent system including memory agent 210 and executive agent 220) based on environmental feedback information. In this way, its reasoning ability, long-term memory retrieval efficiency, and task-related memory construction ability can be continuously improved.

[0094] As an example, server 130 supports interactive training within a reinforcement learning paradigm, enabling the agent to continuously adjust its inference path, response strategy, or multi-round contextual judgment mechanism based on environmental feedback. For instance, when a user expresses "dissatisfaction" or "denial" regarding a system-generated response, server 130 can treat this as negative feedback and use it for training. In embodied agent scenarios, such as when a robot collides or fails a task after following a certain path, server 130 can determine that the current inference is flawed. Server 130 can use these feedback signals to adjust the parameters of the inference model or the way it uses memory, thereby enhancing its judgment ability in similar scenarios in the future.

[0095] As another example, since retrieval of memory information depends on queries generated by the agent, server 130 can utilize reinforcement learning mechanisms to optimize the strategy for generating memory information queries. For instance, server 130 can train by sampling multiple query generation paths and evaluating the diversity and coverage of memory information retrieved by each generation path. If a generation path retrieves only redundant or highly similar content, causing task failure, its associated query generation method will be given a lower reward. Conversely, if multiple queries corresponding to a generation path hit different memory fragments and support high-quality decisions, server 130 can provide positive feedback to the strategy. Furthermore, server 130 can introduce a retrieval cost mechanism, setting a "penalty" value for each retrieval operation to guide the agent to generate fewer, more useful queries, thereby improving information retrieval efficiency.

[0096] In summary, according to the embodiments of this disclosure, by introducing a memory mechanism oriented towards multimodal input, a solution for interaction and task execution with adaptive memory generation, intelligent retrieval, and dynamic updating capabilities is provided. In this way, related multimodal memories can be retrieved from a memory database based on identification information, and context-consistent responses or task decision results can be generated. Furthermore, by supporting the continuous accumulation and conflict correction of memories, and by reflecting and updating based on feedback information, a long-term memory system that can evolve over time and self-repair can be constructed. In this way, an intelligent system with a closed loop of perception, cognition, and action driven by long-term memory can be realized, widely applicable to various scenarios such as intelligent dialogue and embodied intelligence.

[0097] Figure 4 A flowchart of a method 400 for interaction according to some embodiments of the present disclosure is shown. Method 400 can be implemented in any device. For example, method 400 can be implemented at server 130. Reference is made below. Figure 1 Description method 400.

[0098] In box 410, server 130 responds to receiving interactive input for the target object by obtaining identification information related to the interactive input.

[0099] In box 420, server 130 determines target memory information from a memory database based on identification information. The memory database is constructed based on historical interactions performed with the target object.

[0100] In box 430, server 130 determines the response to the interactive input based on the target memory information and the interactive input.

[0101] In some embodiments, the memory database includes memory information of multiple modalities, and determining target memory information related to identification information from the memory database includes: generating at least one query based on the category of the identification information, the at least one query indicating to query at least one piece of memory information corresponding to the category from the memory information of multiple modalities; querying at least one piece of memory information corresponding to the category of the identification information from the memory information of multiple modalities based on the at least one query; and determining target memory information based on the at least one piece of memory information.

[0102] In some embodiments, method 400 further includes: determining first memory information for a target object based on feedback information for a response; and updating at least a portion of the memory information in the memory database related to the identification information based on the first memory information.

[0103] In some embodiments, updating at least a portion of the memory information in the memory database that is related to the identification information includes: determining whether there is memory information in the memory database that conflicts with the first memory information and is related to the identification information; in response to determining that there is second memory information in the memory database that conflicts with the first memory information and is related to the identification information, updating the second memory information at least based on the corresponding confidence levels determined in historical interactions between the first memory information and the second memory information; and in response to the absence of memory information in the memory database that is related to the identification information and conflicts with the first memory information, storing the first memory information in association with the identification information in the memory database.

[0104] In some embodiments, method 400 further includes: determining a modality of historical input associated with identification information, the modality including at least one of a plurality of predetermined modalities; using a memory agent to determine third memory information of the modality of the historical input based on the historical input; determining the association relationship between the third memory information and memory information of a plurality of modalities in a memory database; and storing the third memory information in the memory database according to the association relationship.

[0105] In some embodiments, the memory database is constructed using multiple functional blocks corresponding to multiple modalities, and determining the third memory information includes: using a memory agent to determine at least one piece of historical memory information related to the identification information from the memory database; and based on the at least one piece of historical memory information, by using the memory agent to invoke the functional block corresponding to the modality of the historical input among the multiple functional blocks to determine the third memory information.

[0106] In some embodiments, the third memory information indicates at least one of the following: descriptive information about historical input, or summary information determined based on historical input.

[0107] In some embodiments, the memory database includes a structured data object, the structured object including a plurality of memory nodes corresponding to multiple memory information, and storing third memory information in the memory database includes: in response to an association indicating that at least one memory information related to the third memory information exists in the memory database, updating at least one memory node and the edge of at least one memory node in the structured data object corresponding to the at least one memory information based on the third memory information and the association; and in response to an association indicating that no memory information related to the third memory information exists in the memory database, adding at least one memory node corresponding to the third memory information in the structured data object.

[0108] Figure 5 A flowchart of a method 500 for task execution according to some embodiments of the present disclosure is shown. Method 500 can be implemented in any device. For example, method 500 can be implemented at server 130. Reference is made below. Figure 1 Description method 500.

[0109] In box 510, server 130, in response to receiving target input for a movable target object, obtains identification information associated with the target input.

[0110] In box 520, server 130 determines the target memory information from the memory database based on the identification information.

[0111] In box 530, server 130 determines the target task for the target object based on target memory information and target input.

[0112] In box 540, server 130 controls the target object to move to the target location corresponding to the target task in order to execute the target task.

[0113] In some embodiments, the memory database includes memory information of multiple modalities, and determining target memory information related to identification information from the memory database includes: generating at least one query based on the category of the identification information, the at least one query indicating to query at least one piece of memory information corresponding to the category from the memory information of multiple modalities; querying at least one piece of memory information corresponding to the category of the identification information from the memory information of multiple modalities based on the at least one query; and determining target memory information based on the at least one piece of memory information.

[0114] In some embodiments, method 500 further includes: determining first memory information for a target object based on feedback information for a target task; and updating at least a portion of the memory information related to the identification information in the memory database based on the first memory information.

[0115] In some embodiments, method 500 further includes: determining a modality of historical input associated with identification information, the modality including at least one of a plurality of predetermined modalities; using a memory agent, determining second memory information corresponding to the modality of the historical input; determining the association between the second memory information and memory information of a plurality of modalities in a memory database; and storing the second memory information in the memory database according to the association.

[0116] In some embodiments, the memory database is constructed using multiple functional blocks corresponding to multiple modalities, and determining the second memory information includes: using a memory agent to determine at least one piece of historical memory information related to the identification information from the memory database; and based on the at least one piece of historical memory information, by using the memory agent to invoke the functional block among the multiple functional blocks corresponding to the modality of the historical input to determine the second memory information.

[0117] In some embodiments, the memory database includes a structured data object, the structured object including a plurality of memory nodes corresponding to multiple pieces of memory information, and storing the second memory information in the memory database includes: in response to an association indicating that at least one piece of memory information related to the second memory information exists in the memory database, updating at least one memory node and the edge of at least one memory node corresponding to the at least one piece of memory information in the structured data object based on the second memory information and the association; and in response to an association indicating that no memory information related to the second memory information exists in the memory database, adding at least one memory node corresponding to the second memory information in the structured data object.

[0118] In some embodiments, the target location is determined based on at least one piece of memory information related to the identification information in the memory database.

[0119] In some embodiments, the target input corresponds to a first mode among a plurality of modes, and obtaining identification information includes: acquiring feature information related to the target input using a device associated with the target object, the feature information corresponding to a second mode among a plurality of modes, the second mode being different from the first mode; and determining identification information based on the feature information.

[0120] Embodiments of this disclosure also provide corresponding apparatus for implementing the above methods or processes. Figure 6 A schematic structural block diagram of an interactive device 600 according to some embodiments of the present disclosure is shown. The device 600 may be implemented in or included in server 130, for example. Various modules / components in the device 600 may be implemented by hardware, software, firmware, or any combination thereof.

[0121] like Figure 6As shown, the device 600 includes an identification information acquisition module 610, configured to acquire identification information related to the interactive input in response to receiving an interactive input for a target object; a memory information determination module 620, configured to determine target memory information from a memory database based on the identification information, the memory database being constructed based on historical interactions performed with the target object; and a response module 630, configured to determine a response to the interactive input based on the target memory information and the interactive input.

[0122] In some embodiments, the memory database includes memory information of multiple modalities, and the memory information determination module 620 is further configured to: generate at least one query based on the category of the identification information, wherein the at least one query indicates that at least one piece of memory information corresponding to the category is retrieved from the memory information of multiple modalities; retrieve at least one piece of memory information corresponding to the category of the identification information from the memory information of multiple modalities based on the at least one query; and determine target memory information based on the at least one piece of memory information.

[0123] In some embodiments, the apparatus 600 further includes an update module configured to: determine first memory information for a target object based on feedback information for a response; and update at least a portion of the memory information in the memory database associated with the identification information based on the first memory information.

[0124] In some embodiments, the update module is further configured to: determine whether there is memory information in the memory database that conflicts with the first memory information and is related to the identification information; in response to determining that there is second memory information in the memory database that conflicts with the first memory information and is related to the identification information, update the second memory information at least based on the corresponding confidence levels determined in historical interactions between the first memory information and the second memory information; and in response to the absence of memory information in the memory database that is related to the identification information and conflicts with the first memory information, store the first memory information in association with the identification information in the memory database.

[0125] In some embodiments, the apparatus 600 further includes a memory module configured to: determine a modality of historical input associated with identification information, the modality including at least one of a plurality of predetermined modalities; determine third memory information of the modality of the historical input based on the historical input using a memory agent; determine the association relationship between the third memory information and memory information of a plurality of modalities in a memory database; and store the third memory information in the memory database according to the association relationship.

[0126] In some embodiments, the memory database is constructed using multiple functional blocks corresponding to multiple modalities, and the memory module is further configured to: use a memory agent to determine at least one piece of historical memory information related to the identification information from the memory database; and based on the at least one piece of historical memory information, use the memory agent to call the functional block corresponding to the modality of the historical input among the multiple functional blocks to determine third memory information.

[0127] In some embodiments, the third memory information indicates at least one of the following: descriptive information about historical input, or summary information determined based on historical input.

[0128] In some embodiments, the memory database includes a structured data object, the structured object including multiple memory nodes corresponding to multiple memory information, and the memory module is further configured to: in response to an association indicating that at least one memory information related to the third memory information exists in the memory database, update at least one memory node and the edge of the at least one memory node in the structured data object corresponding to the at least one memory information based on the third memory information and the association; and in response to an association indicating that no memory information related to the third memory information exists in the memory database, add at least one memory node corresponding to the third memory information in the structured data object.

[0129] The units and / or modules included in device 600 can be implemented in various ways, including software, hardware, firmware, or any combination thereof. In some embodiments, one or more units and / or modules can be implemented using software and / or firmware, such as machine-executable instructions stored on a storage medium. In addition to or as an alternative to machine-executable instructions, some or all of the units and / or modules in device 600 can be implemented at least partially by one or more hardware logic components. By way of example and not limitation, exemplary types of hardware logic components that can be used include field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), complex programmable logic devices (CPLDs), and so on.

[0130] Figure 7 A schematic structural block diagram of an apparatus 700 for task execution according to some embodiments of the present disclosure is shown. The apparatus 700 may be implemented in or included in server 130, for example. Various modules / components in the apparatus 700 may be implemented by hardware, software, firmware, or any combination thereof.

[0131] like Figure 7As shown, the device 700 includes an identification information acquisition module 710, configured to acquire identification information related to the target input in response to receiving identification information for a target input of a movable target object; a memory information determination module 720, configured to determine target memory information of the identification information from a memory database based on the identification information; a task determination module 730, configured to determine a target task for the target object based on the target memory information and the target input; and an execution module 740, configured to control the target object to move to a target position corresponding to the target task in order to execute the target task.

[0132] In some embodiments, the memory database includes memory information of multiple modalities, and the memory information determination module 720 is further configured to generate at least one query based on the category of the identification information, wherein the at least one query indicates to query at least one piece of memory information corresponding to the category from the memory information of multiple modalities; based on the at least one query, query at least one piece of memory information corresponding to the category of the identification information from the memory information of multiple modalities; and based on the at least one piece of memory information, determine target memory information.

[0133] In some embodiments, the apparatus 700 further includes an update module configured to: determine first memory information for a target object based on feedback information for a target task; and update at least a portion of the memory information in the memory database related to the identification information based on the first memory information.

[0134] In some embodiments, the apparatus 700 further includes a memory module configured to: determine a modality of historical input associated with identification information, the modality including at least one of a plurality of predetermined modalities; utilize a memory agent to determine second memory information corresponding to the modality of the historical input; determine the association between the second memory information and memory information of a plurality of modalities in a memory database; and store the second memory information in the memory database according to the association.

[0135] In some embodiments, the memory database is constructed using multiple functional blocks corresponding to multiple modalities, and the memory module is further configured to: use a memory agent to determine at least one piece of historical memory information related to the identification information from the memory database; and based on the at least one piece of historical memory information, use the memory agent to call the functional block among the multiple functional blocks corresponding to the modality of the historical input to determine second memory information.

[0136] In some embodiments, the memory database includes a structured data object, the structured object including a plurality of memory nodes corresponding to multiple memory information, and the memory module is further configured to: in response to an association indicating that at least one memory information related to the second memory information exists in the memory database, update at least one memory node and the edge of at least one memory node in the structured data object corresponding to the at least one memory information based on the second memory information and the association; and in response to an association indicating that no memory information related to the second memory information exists in the memory database, add at least one memory node corresponding to the second memory information in the structured data object.

[0137] In some embodiments, the target location is determined based on at least one piece of memory information related to the identification information in the memory database.

[0138] In some embodiments, the target input corresponds to a first mode among a plurality of modes, and the feature information acquisition module 710 is further configured to: acquire feature information related to the target input using a device associated with the target object, the feature information corresponding to a second mode among a plurality of modes, the second mode being different from the first mode; and determine identification information based on the feature information.

[0139] It should be understood that one or more steps in the above methods can be performed by suitable electronic devices or combinations of electronic devices. Such electronic devices or combinations of electronic devices may include, for example, […]. Figure 1 Server 130 in the middle.

[0140] Figure 8 A block diagram of an electronic device 800 in which one or more embodiments of the present disclosure may be implemented is shown. It should be understood that... Figure 8 The electronic device 800 shown is merely exemplary and should not be construed as limiting the functionality and scope of the embodiments described herein.

[0141] like Figure 8 As shown, electronic device 800 is in the form of a general-purpose electronic device. Components of electronic device 800 may include, but are not limited to, one or more processors or processor 810, memory 820, storage device 830, one or more communication units 840, one or more input devices 850, and one or more output devices 860. Processor 810 may be a physical or virtual processor and is capable of performing various processes according to programs stored in memory 820. In a multiprocessor system, multiple processing units execute computer-executable instructions in parallel to improve the parallel processing capability of electronic device 800.

[0142] Electronic device 800 typically includes multiple computer storage media. Such media can be any available media accessible to electronic device 800, including but not limited to volatile and non-volatile media, removable and non-removable media. Memory 820 can be volatile memory (e.g., registers, cache, random access memory (RAM)), non-volatile memory (e.g., read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory), or some combination thereof. Storage device 830 can be removable or non-removable media and can include machine-readable media, such as flash drives, disks, or any other media capable of storing information and / or data and accessible within electronic device 800.

[0143] Electronic device 800 may further include additional removable / non-removable, volatile / non-volatile storage media. Although not explicitly stated... Figure 8 As shown, disk drives for reading from or writing to removable, non-volatile disks (e.g., "floppy disks") and optical disk drives for reading from or writing to removable, non-volatile optical disks can be provided. In these cases, each drive can be connected to a bus (not shown) via one or more data media interfaces. Memory 820 may include computer program product 825 having one or more program modules configured to perform various methods or actions of various embodiments of this disclosure.

[0144] The communication unit 840 enables communication with other electronic devices via a communication medium. Additionally, the functionality of the components of the electronic device 800 can be implemented using a single computing cluster or multiple computing machines capable of communicating via communication connections. Therefore, the electronic device 800 can operate in a networked environment using logical connections to one or more other servers, network personal computers (PCs), or another network node.

[0145] Input device 850 can be one or more input devices, such as a mouse, keyboard, trackball, etc. Output device 860 can be one or more output devices, such as a monitor, speaker, printer, etc. Electronic device 800 can also communicate with one or more external devices (not shown) via communication unit 840 as needed. These external devices include storage devices, display devices, etc., and can communicate with one or more devices that enable user interaction with electronic device 800, or with any device that enables electronic device 800 to communicate with one or more other electronic devices (e.g., network card, modem, etc.). Such communication can be performed via input / output (I / O) interface (not shown).

[0146] According to an exemplary implementation of this disclosure, a computer-readable storage medium is provided that stores computer-executable instructions thereon, wherein the computer-executable instructions are executed by a processor to implement the methods described above. According to an exemplary implementation of this disclosure, a computer program product is also provided, which is tangibly stored on a non-transitory computer-readable medium and includes computer-executable instructions, which are executed by a processor to implement the methods described above.

[0147] Various aspects of this disclosure are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatuses, devices, and computer program products implemented according to this disclosure. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.

[0148] These computer-readable program instructions can be provided to the processing unit of a general-purpose computer, special-purpose computer, or other programmable model deployment apparatus to produce a machine such that, when executed by the processing unit of the computer or other programmable model deployment apparatus, they create means for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium that causes a computer, programmable model deployment apparatus, and / or other device to operate in a particular manner; thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.

[0149] Computer-readable program instructions can be loaded onto a computer, other programmable model deployment apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable model deployment apparatus, or other device to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable model deployment apparatus, or other device to perform the functions / actions specified in one or more boxes of a flowchart and / or block diagram.

[0150] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of an instruction, which contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.

[0151] Various implementations of this disclosure have been described above. These descriptions are exemplary and not exhaustive, nor are they limited to the disclosed implementations. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described implementations. The terminology used herein is chosen to best explain the principles, practical applications, or improvements to technology in the market, or to enable others skilled in the art to understand the various implementations disclosed herein.

Claims

1. An interaction method, comprising: In response to receiving interactive input for a target object, obtain identification information related to the interactive input; Based on the identification information, target memory information is determined from a memory database, which is constructed based on historical interactions performed with the target object; and Based on the target memory information and the interactive input, a response to the interactive input is determined.

2. The method according to claim 1, wherein the memory database includes memory information of multiple modalities, and determining the target memory information from the memory database includes: Based on the category of the identification information, at least one query is generated, the at least one query indicating that at least one piece of memory information corresponding to the category is retrieved from the memory information of the plurality of modalities; Based on the at least one query, at least one piece of memory information corresponding to the category of the identification information is retrieved from the memory information of the multiple modalities; as well as The target memory information is determined based on the at least one piece of memory information.

3. The method according to claim 1, further comprising: Based on the feedback information in response to the response, first memory information for the target object is determined; as well as Based on the first memory information, at least a portion of the memory information in the memory database related to the identification information is updated.

4. The method of claim 3, wherein updating at least a portion of the memory information in the memory database associated with the identification information comprises: Determine whether there is memory information in the memory database that conflicts with the first memory information and is related to the identification information; In response to determining that there is second memory information in the memory database that conflicts with the first memory information and is related to the identification information, the second memory information is updated at least based on the corresponding confidence levels determined in historical interactions between the first memory information and the second memory information; as well as In response to the absence of memory information in the memory database that is related to the identification information and conflicts with the first memory information, the first memory information is stored in the memory database in association with the identification information.

5. The method according to claim 1, further comprising: Determine the modality of historical inputs associated with the identification information, the modality including at least one of a plurality of predetermined modalities; Using a memory agent, the third memory information of the modality of the historical input is determined based on the historical input; Determine the association between the third memory information and memory information of multiple modalities in the memory database; as well as According to the aforementioned association, the third memory information is stored in the memory database.

6. The method of claim 5, wherein the memory database is constructed using a plurality of functional blocks corresponding to the plurality of modalities, and wherein determining the third memory information includes: Using the memory agent, at least one piece of historical memory information related to the identification information is determined from the memory database; as well as Based on the at least one piece of historical memory information, the third memory information is determined by using the memory agent to call the function block among the plurality of function blocks that corresponds to the modality of the historical input.

7. The method of claim 5, wherein the third memory information indicates at least one of the following: Description information of the historical input, or Summary information determined based on the historical input.

8. The method according to claim 5, wherein the memory database includes a structured data object, the structured object including a plurality of memory nodes respectively corresponding to multiple memory information, and storing the third memory information into the memory database includes: In response to the association indicating that at least one piece of memory information related to the third memory information exists in the memory database, based on the third memory information and the association, at least one memory node in the structured data object corresponding to the at least one piece of memory information and the edge of the at least one memory node are updated; as well as In response to the association indicating that there is no memory information related to the third memory information in the memory database, at least one memory node corresponding to the third memory information is added to the structured data object.

9. A task execution method, comprising: In response to receiving a target input for a movable target object, acquire identification information related to the target input; Based on the identification information, the target memory information is determined from the memory database; Based on the target memory information and the target input, a target task is determined for the target object; as well as Control the target object to move to the target location corresponding to the target task in order to execute the target task.

10. The method of claim 9, wherein the memory database comprises memory information of multiple modalities, and determining the target memory information from the memory database comprises: Based on the category of the identification information, at least one query is generated, the at least one query indicating that at least one piece of memory information corresponding to the category is retrieved from the memory information of the plurality of modalities; Based on the at least one query, at least one piece of memory information corresponding to the category of the identification information is retrieved from the memory information of the multiple modalities; as well as The target memory information is determined based on the at least one piece of memory information.

11. The method of claim 9, further comprising: Based on the feedback information for the target task, first memory information for the target object is determined; as well as Based on the first memory information, at least a portion of the memory information in the memory database related to the identification information is updated.

12. The method according to claim 9, further comprising: Determine the modality of historical inputs associated with the identification information, the modality including at least one of a plurality of predetermined modalities; Using a memory agent, based on second memory information that determines the historical input corresponds to the modality; Determine the association between the second memory information and memory information of multiple modalities in the memory database; as well as According to the aforementioned association, the second memory information is stored in the memory database.

13. The method of claim 12, wherein the memory database is constructed using a plurality of functional blocks corresponding to the plurality of modalities, and wherein determining the second memory information includes: Using the memory agent, at least one piece of historical memory information related to the identification information is determined from the memory database; as well as Based on the at least one piece of historical memory information, the second memory information is determined by using the memory agent to call the function block among the plurality of function blocks that corresponds to the modality of the historical input.

14. The method of claim 12, wherein the memory database comprises a structured data object, the structured object comprising a plurality of memory nodes respectively corresponding to multiple pieces of memory information, and storing the second memory information into the memory database comprises: In response to the association indicating that at least one piece of memory information related to the second memory information exists in the memory database, based on the second memory information and the association, at least one memory node in the structured data object corresponding to the at least one piece of memory information and the edge of the at least one memory node are updated; as well as In response to the association indicating that there is no memory information related to the second memory information in the memory database, at least one memory node corresponding to the second memory information is added to the structured data object.

15. The method of claim 9, wherein the target location is determined based on at least one piece of memory information in the memory database that is associated with the identification information.

16. The method of claim 9, wherein the target input corresponds to a first mode among a plurality of modes, and obtaining the identification information comprises: The feature information related to the target input is obtained using a device associated with the target object, the feature information corresponding to a second mode among the plurality of modes, the second mode being different from the first mode; as well as The identification information is determined based on the feature information.

17. A device for interaction, comprising: The identification information acquisition module is configured to acquire identification information related to the interactive input in response to receiving interactive input for a target object; The memory information determination module is configured to determine target memory information from a memory database based on the identification information, the memory database being constructed based on historical interactions performed with the target object; as well as The response module is configured to determine a response to the interactive input based on the target memory information and the interactive input.

18. An apparatus for performing a task, comprising: The identification information acquisition module is configured to acquire identification information related to the target input in response to receiving a target input for a movable target object; The memory information determination module is configured to determine target memory information from the memory database based on the identification information; The task determination module is configured to determine the target task for the target object based on the target memory information and the target input; as well as The execution module is configured to control the target object to move to the target location corresponding to the target task in order to execute the target task.

19. An electronic device comprising: At least one processor; as well as At least one memory coupled to the at least one processor and storing instructions for execution by the at least one processor, the instructions causing the electronic device to perform the method according to any one of claims 1 to 8 or claims 9 to 16 when executed by the at least one processor.

20. A computer-readable storage medium having stored thereon computer-executable instructions, which are executable by a processor to implement the method according to any one of claims 1 to 8 or claims 9 to 16.

21. A computer program product comprising computer-executable instructions, wherein the computer-executable instructions, when executed by a processor, implement the method according to any one of claims 1 to 8 or claims 9 to 16.

Citation Information

Cited By

  • Memory data processing method based on intelligent agent and related device

    CN121052280A

  • Intelligent agent interaction method and device, equipment, medium and product

    CN121809534A