Task processing method, device and equipment

By acquiring user input information and generating multi-level target instructions based on multimodal data, combined with the user's personalized information, the limitations of intelligent agents in acquiring user data are overcome, resulting in more accurate understanding of user intent and improved task execution efficiency.

CN121326518APending Publication Date: 2026-01-13LENOVO (BEIJING) LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511434615.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-30
Publication Date
2026-01-13

AI Technical Summary

Technical Problem

In existing technologies, intelligent agents acquire user data by relying on in-application data or user-initiated input, which limits the value of data application and places a heavy cognitive burden on users, making it difficult to effectively utilize user data to process user tasks.

Method used

By acquiring user input information, multi-level target instructions are generated based on multimodal data. Combined with the user's personalized information, multi-level target instructions are generated, including short-term memory, long-term memory, preferences and traits, and instructions are generated layer by layer to execute user tasks.

Benefits of technology

It enables a more accurate understanding of user intent, improves the efficiency and personalization of user tasks, breaks down the isolation between devices and applications, and expands the scope of data application.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121326518A_ABST
    Figure CN121326518A_ABST
Patent Text Reader

Abstract

The invention discloses a task processing method, device and equipment. The method comprises the steps of obtaining input information of a user; generating a multi-level target instruction based on the input information and the personalized information of the user; wherein the personalized information represents user feature information generated based on multi-modal data; the multi-level target instructions represent instructions used for executing user tasks corresponding to the input information, and the target instructions of different instruction levels correspond to different pieces of user feature information; and executing the multi-level target instruction to obtain a processing result of the user task corresponding to the input information.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present disclosure relates to, but is not limited to, the technical field of computer, and particularly relates to a task processing method, device and equipment. BACKGROUND

[0002] An agent refers to an entity capable of perceiving an environment, making decisions and taking actions based on the perceived information, so as to achieve a specific goal or task. In order to make the agent more intelligent and closer to user needs, it is necessary to help the agent understand the user better and obtain more user personal data.

[0003] Currently, the related art mainly relies on in-application data to obtain user personal data. However, the isolated application ecosystem limits the application value of user data. In addition, user data can also be obtained through user-initiated input, but this will bring a great cognitive burden to the user.

[0004] Therefore, how to obtain user data and effectively utilize the user data to process user tasks has become a problem to be solved. SUMMARY

[0005] Therefore, the present disclosure at least provides a task processing method, device and equipment.

[0006] The technical solution of the present disclosure is implemented as follows:

[0007] In one aspect, the present disclosure provides a task processing method, which comprises:

[0008] obtaining input information of a user;

[0009] generating a multi-level target instruction based on the input information and personalized information of the user; wherein the personalized information represents user feature information generated based on multi-modal data; the multi-level target instruction represents an instruction for executing a user task corresponding to the input information, and target instructions of different instruction levels correspond to different user feature information;

[0010] executing the multi-level target instruction to obtain a processing result of the user task corresponding to the input information.

[0011] In some embodiments, the user feature information includes at least two types of at least one short-term memory, at least one long-term memory, at least one preference and at least one trait of the user; the abstraction degree and / or generality of the short-term memory, the long-term memory, the preference and the trait increase in turn;

[0012] The generating of the multi-level target instruction based on the input information and the personalized information of the user comprises:

[0013] determine, based on the input information, at least two of the target short-term memory, the target long-term memory, the target preference, and the target trait from the user characteristic information;

[0014] generate, based on the different instruction levels and the at least two of the target short-term memory, the target long-term memory, the target preference, and the target trait, the multi-level target instruction; wherein the different instruction levels are based on instruction planning on the input information.

[0015] In some embodiments, the target instructions of the different instruction levels correspond to different levels of abstraction and / or decomposability;

[0016] generate, based on the different instruction levels and the at least two of the target short-term memory, the target long-term memory, the target preference, and the target trait, the multi-level target instruction, comprising:

[0017] For each instruction level, determine at least one type of target information from the at least two of the target short-term memory, the target long-term memory, the target preference, and the target trait based on the corresponding level of abstraction and / or decomposability; wherein the higher the level of abstraction and / or decomposability of the target instruction of the instruction level, the higher the level of abstraction and / or generality of the corresponding target information.

[0018] generate the target instruction corresponding to the instruction level based on the input information and the at least one type of target information.

[0019] In some embodiments, the multi-modal data represents data from a first perspective of the user; wherein,

[0020] The multi-modal data includes cross-device data collected through at least one device associated with the user;

[0021] and / or,

[0022] The multi-modal data includes cross-device data received through a plurality of devices associated with the user;

[0023] and / or,

[0024] The multi-modal data can be used to process a cross-application user task initiated by the user.

[0025] In some embodiments, the method further comprises:

[0026] generate at least one short-term memory from the multi-modal data of the user;

[0027] generate at least one long-term memory from the correlation information between the at least one short-term memory;

[0028] generate at least one preference from the at least one long-term memory;

[0029] Generate at least one trait based on at least one preference.

[0030] In some implementations, the preferences have weight information; the method further includes:

[0031] Update at least one preference; where,

[0032] If the similarity between an existing second preference and a newly generated first preference in at least one preference is greater than a similarity threshold, the weight of the second preference is increased.

[0033] If the similarity between the newly generated first preference and any existing preference among at least one preference is less than or equal to a similarity threshold, the first preference is stored in the at least one preference, and an initial weight is assigned to the first preference.

[0034] In some implementations, the method further includes increasing the weight of the target preference.

[0035] In some implementations, the method further includes:

[0036] Based on the application scenario information, context information, and processing results corresponding to the input information, a summary information is generated.

[0037] The summary information is stored in association with the target preferences.

[0038] On the other hand, this disclosure also provides a task processing apparatus, including:

[0039] The acquisition module is used to acquire user input information;

[0040] The generation module is used to generate multi-level target instructions based on input information and user personalized information. The personalized information represents user feature information generated based on multimodal data. The multi-level target instructions represent instructions for executing user tasks corresponding to the input information, and different instruction levels correspond to different user feature information.

[0041] The execution module is used to execute multi-level target instructions and obtain the processing results of the user tasks corresponding to the input information.

[0042] In another aspect, this disclosure also provides an electronic device including a memory and at least one processor; wherein the memory stores a computer program that can run on the at least one processor, and the at least one processor executes the program to implement the steps in any of the above embodiments.

[0043] It should be understood that the above general description and the following detailed description are merely exemplary and explanatory, and are not intended to limit the technical solutions of this disclosure. Attached Figure Description

[0044] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this disclosure and, together with the specification, serve to illustrate the technical solutions of this disclosure.

[0045] Figure 1 A schematic diagram illustrating the implementation flow of a task processing method provided in this disclosure;

[0046] Figure 2 This is a schematic diagram of the storage structure of user personalized information in one embodiment of the present disclosure;

[0047] Figure 3 This is a schematic diagram of the process for generating personalized user information in one embodiment of the present disclosure;

[0048] Figure 4 This is a schematic diagram of a task processing flow in one embodiment of the present disclosure;

[0049] Figure 5 This is a schematic diagram of the composition structure of a task processing device provided in this disclosure;

[0050] Figure 6 This is a schematic diagram of the hardware entity of an electronic device provided in this disclosure. Detailed Implementation

[0051] To make the objectives, technical solutions, and advantages of this disclosure clearer, the technical solutions of this disclosure are further described in detail below with reference to the accompanying drawings and embodiments. The described embodiments should not be regarded as limitations on this disclosure. All other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this disclosure.

[0052] In the following description, references are made to “some embodiments,” which describe a subset of all possible embodiments. However, it is understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.

[0053] The terms “first / second / third” are used merely to distinguish similar objects and do not represent a specific ordering of objects. It is understood that “first / second / third” may be interchanged in a specific order or sequence where permitted, so that the embodiments of this disclosure described herein can be implemented in an order other than that illustrated or described herein.

[0054] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure pertains. The terminology used herein is for descriptive purposes only and is not intended to limit the scope of this disclosure.

[0055] This disclosure provides a task processing method that can be executed by an electronic device. The electronic device can be various types of terminals such as laptops, tablets, desktop computers, set-top boxes, and mobile devices (e.g., mobile phones, portable music players, personal digital assistants, dedicated messaging devices, portable gaming devices), or it can be implemented as a server. The server can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms.

[0056] The technical solutions provided in this disclosure will now be clearly and completely described in conjunction with the accompanying drawings.

[0057] Figure 1 This is a schematic diagram illustrating the implementation flow of a task processing method provided in this disclosure, such as... Figure 1 As shown, the method includes the following steps S11 to S13:

[0058] Step S11: Obtain user input information.

[0059] Here, user input information refers to information related to the user used to determine the user's task.

[0060] In some implementations, the input information can be of any data type, such as text, image, voice, or multimodal information.

[0061] In some implementations, the input information can be information obtained in any way. For example, the input information can be a query actively entered by the user through a designated interactive window. Alternatively, the input information can be information determined by actively detecting the user's interaction content during the user interaction process.

[0062] Thus, after obtaining user input, the input can be analyzed to determine the corresponding user task. In some implementations, the input can be fed into a designated model to generate the corresponding user task. In other implementations, the corresponding user task can be manually determined based on the input.

[0063] Step S12: Based on the input information and the user's personalized information, generate multi-level target instructions; wherein, the personalized information represents user feature information generated based on multimodal data; the multimodal data represents data from the user's first-person perspective; the multi-level target instructions represent instructions for executing the user task corresponding to the input information, and different instruction levels correspond to different user feature information.

[0064] Here, multimodal data refers to a data set containing two or more different types or modalities (e.g., text, image, audio, video, sensor data, etc.). Data of different modalities express information through different perceptual methods (e.g., visual, auditory, tactile, etc.) and collaboratively or relatedly describe the same phenomenon or object. Therefore, multimodal data has a richer information expression capability.

[0065] Personalized characteristics refer to the unique traits that distinguish a user from others. These include a user's physical appearance, personality traits, personal experiences (e.g., life, work, or educational background), hobbies, and so on. When performing user tasks, these personalized characteristics allow for a more accurate understanding of the user's intent and relevant background information in the input, leading to better decision-making and more satisfactory results.

[0066] Personalized information representation is based on user feature information generated from multimodal data. That is, personalized features are user feature information obtained after extracting, analyzing, and summarizing information from the multimodal data. In some implementations, this personalized information can be determined by any suitable method. In some implementations, semantic understanding can be performed on the multimodal data to determine personalized information; for example, semantic understanding can be performed on the user's natural language interactions with others, or on images or scenes that the user is interested in, to determine personalized information. In some implementations, attribute recognition can be performed on the multimodal data to determine personalized information; for example, attribute recognition can be performed on the products or scenery that the user is interested in to determine the user's personal preferences.

[0067] Thus, because multimodal data has richer information expression forms and stronger information expression capabilities, it can comprehensively and truthfully describe users' personal lives and experiences. Therefore, the user feature information generated based on this multimodal data, that is, the user's personalized information, can also more comprehensively and accurately describe the user's true characteristics.

[0068] After obtaining the user's input information and determining the corresponding user task, the system further generates multi-level target instructions for executing the user task based on the input information and the user task, combined with the user's personalized information.

[0069] Here, multi-level target instructions refer to task execution instructions generated according to a specified instruction level. In some implementations, the specified instruction level can be determined according to the task corresponding to the instruction. For example, when a user task is divided into multiple subtasks, at least one instruction corresponding to each subtask can be determined as an instruction of the same level, so that the number of subtasks corresponds to the number of instruction levels. In some implementations, the specified instruction level can be determined according to the level of abstraction of the instruction. For example, instruction levels include high-level instruction, medium-level instruction, and low-level instruction, with the level of abstraction decreasing sequentially.

[0070] Different instruction levels correspond to different user characteristic information. That is, when determining the target instructions for different instruction levels, the target instructions for each instruction level are generated based on the user's input information, the user task, and the user characteristic information corresponding to each instruction level. In some implementations, different instruction levels correspond to different types of user characteristic information. For example, high-level and mid-level target instructions correspond to relatively stable personalized information (e.g., the user's physical characteristics, personality traits, personal experiences (e.g., life, work, or educational background characteristics), hobbies, etc.), while low-level target instructions correspond to the user's long-term memory and current experiences.

[0071] Step S13: Execute the multi-level target instruction to obtain the processing result of the user task corresponding to the input information.

[0072] Here, after determining the multi-level target instructions, the multi-level target instructions are executed sequentially to obtain the processing results of the user task corresponding to the input information.

[0073] In some implementations, when the multi-level target instructions correspond to different subtasks of a user task, the multi-level target instructions are executed according to the execution order of the multiple subtask pairs. In some implementations, when the multi-level target instructions represent target instructions with different levels of abstraction, the target instruction with the lowest level of abstraction is executed, and the success of the execution of the target instruction with a higher level of abstraction is determined by a step-by-step verification method.

[0074] In some implementations, the execution result can be evaluated for the target instruction at each instruction level to verify the accuracy of the instruction execution result.

[0075] In the embodiments provided in this disclosure, firstly, multi-level target instructions are generated based on user input information and user personalization information. The personalization information represents user feature information generated from multimodal data, which represents data from the user's first-person perspective. The multi-level target instructions represent instructions used to execute the user task corresponding to the input information, and different instruction levels correspond to different user feature information. Then, the multi-level target instructions are executed to obtain the processing result of the user task corresponding to the input information. In this way, combining different user feature information to generate target instructions at different instruction levels can improve the personalization of the multi-level target instructions, thereby making the execution result more consistent with the user's true intentions.

[0076] Furthermore, the above embodiments generate personalized user information based on multimodal data collected from the user's first-person perspective, which can obtain more comprehensive and accurate user characteristic information, thereby helping to improve the accuracy of understanding user intentions.

[0077] In some implementations, the user characteristic information includes at least two of the following: at least one short-term memory, at least one long-term memory, at least one preference, and at least one trait; the level of abstraction and / or generalization of the short-term memory, the long-term memory, the preference, and the trait increases sequentially.

[0078] Here, short-term memory refers to user characteristic information with a low level of abstraction and / or generalization. In some implementations, short-term memory can be environmental information collected from the user's first-person perspective, such as multimodal information like image information, voice information, text information, and / or video information. In some implementations, short-term memory can be information obtained after preliminary understanding or analysis of current information collected from the user's first-person perspective, such as semantic information or attribute information obtained after semantic understanding of image information or attribute recognition of objects within it.

[0079] Long-term memory refers to user characteristic information that is more abstract and / or generalized than short-term memory. In some implementations, long-term memory is a summary and integration of short-term memory, reflecting a user's experiences and knowledge accumulation in a specific scenario. In some implementations, long-term memory can include semantic memory, episodic memory, and procedural memory; where semantic memory refers to memory about concepts, things, and relationships between things; episodic memory refers to memory related to a specific time or place; and procedural memory refers to memory related to behavioral patterns and habits.

[0080] Preferences refer to user characteristic information that has a higher level of abstraction and / or generalization compared to short-term and long-term memory. In some implementations, preferences are stable tendencies formed by users in different contexts. In some implementations, user preferences may include content preferences, reasoning preferences, and action preferences, etc.; where, content preferences refer to the content that users tend to focus on or are interested in, such as liking to read news or history; reasoning preferences refer to how users tend to evaluate, compare, and make decisions, such as being accustomed to logical analysis; action preferences refer to how users tend to take actions or perform tasks, such as preferring to operate manually.

[0081] A trait refers to user characteristic information with the highest level of abstraction and / or generalization. In some implementations, a trait is a stable personality characteristic exhibited by a user across different scenarios, such as price sensitivity, early riser, or analytical behavior. In some implementations, a trait can be a stable user characteristic generated by aggregating and summarizing at least one user preference and / or at least one long-term memory.

[0082] As described above, user characteristic information (i.e., user personalized information) is constructed layer by layer from concrete to abstract or from concrete to general. Simultaneously, during decision-making, user characteristics with higher levels of abstraction and / or generalization will progressively influence characteristics with lower levels of abstraction and / or generalization. Therefore, in some implementations, such as... Figure 2 As shown, different personalized information about users is stored in layers, from bottom to top: short-term memory 210, long-term memory 220, preferences 230, and traits 240. While each layer of information is constructed from the bottom up, in specific tasks or decisions, the information from each layer may have a top-down impact.

[0083] Here, user feature information includes at least two of the following: at least one short-term memory, at least one long-term memory, at least one preference, and at least one trait. For example, it may include one short-term memory and at least one long-term memory, or at least one long-term memory, at least one preference, and at least one trait. In this way, target instructions at different instruction levels can be generated based on different types of user feature information.

[0084] Thus, the generation of multi-level target instructions based on the input information and the user's personalized information, i.e., step S12 above, can be implemented as the following steps S121 to S122:

[0085] Step S121: Based on the input information, determine at least two categories from the user feature information: target short-term memory, target long-term memory, target preferences, and target traits.

[0086] Here, in some implementations, user feature information is retrieved based on user input information to obtain at least two categories among target short-term memory, target long-term memory, target preference, and target trait, whose relevance to the input information is greater than a specified threshold. For example, if the user feature information includes at least one short-term memory and at least one long-term memory, target short-term memory and / or target long-term memory with a relevance to the input information greater than a specified threshold can be identified. As another example, if the user feature information includes at least one long-term memory, at least one preference, and at least one trait, target long-term memory, target preference, and / or target trait with a relevance to the input information greater than a specified threshold can be identified.

[0087] In some implementations, user feature information is retrieved based on user input information and its corresponding user task, thereby obtaining at least two of the following categories: target short-term memory, target long-term memory, target preference, and target trait, which are more relevant to the input information than a specified threshold.

[0088] Step S122: Based on different instruction levels and at least two of the following: target short-term memory, target long-term memory, target preference, and target features, generate the multi-level target instruction; wherein, the different instruction levels are obtained based on instruction planning of the input information.

[0089] Here, different instruction levels are obtained based on the instruction planning process for the input information.

[0090] In some implementations, during instruction planning, different instruction levels can be determined based on the complexity of the user task corresponding to the input information. For example, for simple tasks with relatively fixed execution strategies, fewer instruction levels can be designed, such as only high-level and low-level instructions; while for complex tasks whose execution strategies need to be dynamically adjusted according to the execution environment, more instruction levels can be designed, such as high-level, mid-level, and low-level instructions.

[0091] In some implementations, during instruction planning, different instruction levels can be determined based on the cognitive level and execution accuracy of the agent executing the instructions. For example, for rule-driven agents with fixed logic, fewer instruction levels can be designed, such as only high-level and low-level instructions; while for highly autonomous agents, more instruction levels can be designed, such as high-level, mid-level, and low-level instructions.

[0092] Thus, after determining different instruction levels, multi-level target instructions are generated based on these different instruction levels and at least two of the determined target short-term memory, target long-term memory, target preferences, and target features. Simultaneously, different instruction levels correspond to different user feature information. For example, if the instruction levels include high-level and low-level instruction levels, and the determined at least two types of user feature information are target long-term memory and target preferences, then high-level target instructions can be generated based on target long-term memory and target preferences, and low-level target instructions can be generated based on target long-term memory. As another example, if the instruction levels include high-level, mid-level, and low-level instruction levels, and the determined at least two types of user feature information are target short-term memory, target long-term memory, and target preferences, then high-level target instructions can be generated based on target long-term memory and target preferences, mid-level target instructions can be generated based on target long-term memory, and low-level target instructions can be generated based on target long-term memory and target short-term memory.

[0093] In the embodiments provided in this application, user feature information includes at least two of the following categories: short-term memory, long-term memory, preferences, and traits. This aligns with human cognitive mechanisms, making the multi-level target instructions generated based on user feature information more consistent with users' cognitive mechanisms and reflect their personalized characteristics, thus simulating the user's real decision-making process. On the other hand, by planning different instruction levels based on the input information and then combining them with different user feature information to generate multi-level target instructions, the design of multi-level target instructions can be matched with the user tasks corresponding to the input information, thereby improving the execution efficiency of user tasks.

[0094] In some implementations, the target instructions at different instruction levels correspond to different levels of abstraction and / or decomposability.

[0095] Here, the level of abstraction refers to the general level of the instruction's description of the user's task, or the granularity and generalization of the task description by different levels of target instructions. For example, when target instructions include high-level, mid-level, and low-level instructions, high-level instructions have a higher level of abstraction and are used to describe the guiding plan for completing the task. They are concise and lack detail, such as "Find suitable restaurants for the user to choose at noon." Mid-level instructions are used to describe the step-by-step plan for completing the guidance and have a certain number and continuity, such as "Instruction 1: Query the current geographical location," "Instruction 2: Find interfaces and tools that can recommend restaurants based on geographical location," and "Instruction 3: Generate an overview of each restaurant." Low-level instructions have a lower level of abstraction and are used to describe the detailed plan for completing each step. They have detailed and complete implementation steps, such as "Instruction 1: Find the recently used restaurant application and obtain the corresponding application interface," "Instruction 2: Generate personalized parameters for calling the restaurant list application interface," and "Instruction 3: Call the user's short-term personalized memory and find a list of recently visited restaurants."

[0096] Decomposability refers to the ability to convert between higher-level and lower-level instructions in a multi-level instruction set. For example, the decomposability of higher-level and mid-level instructions decreases sequentially compared to lower-level instructions that contain execution parameters and can be directly executed. In some implementations, the level of abstraction of an instruction corresponds to its decomposability; that is, the higher the level of abstraction, the higher the decomposability.

[0097] Thus, the generation of the multi-level target instructions based on different instruction levels and at least two of the following: target short-term memory, target long-term memory, target preferences, and target traits, i.e., the above step S122, can be implemented as the following steps S1221 to S1222:

[0098] Step S1221: For each instruction level, based on the corresponding level of abstraction and / or decomposability, determine at least one type of target information from the target short-term memory, the target long-term memory, the target preference, and the target trait; wherein, the higher the level of abstraction and / or decomposability of the target instruction at the instruction level, the higher the level of abstraction and / or generalization of the corresponding target information.

[0099] Here, after determining at least two categories of target short-term memory, target long-term memory, target preference, and target trait based on user input information, at least one target information is determined from these at least two categories based on the level of abstraction and / or decomposability of each instruction level.

[0100] At the same time, when determining at least one target information, the higher the degree of abstraction and / or the higher the degree of decomposability of the target instruction, the higher the degree of abstraction and / or generalization of the target information determined for it.

[0101] In some implementations, when the number of instruction levels is less than or equal to the number of user feature types, a one-to-one correspondence or a one-to-many relationship can be established between instruction levels and user feature information types based on the level of abstraction and / or decomposability of the target instructions at different instruction levels, and the level of abstraction and / or generality of the user feature information. In a one-to-many relationship, the highest level of abstraction and / or highest generality of the user feature information corresponding to an instruction level with a lower level of abstraction and / or higher decomposability is higher than the lowest level of abstraction and / or lowest generality. For example, when multi-level target instructions include high-level and low-level target instructions with decreasing levels of abstraction and / or decomposability, and user feature information includes target short-term memory, target long-term memory, and target preferences, based on the one-to-one correspondence, target preferences and target long-term memory can be identified as the target information corresponding to the high-level and low-level instruction levels, respectively; based on the one-to-many correspondence, target preferences and target long-term memory can be identified as the target information corresponding to the high-level instruction level, and target long-term memory and target short-term memory can be identified as the target information corresponding to the low-level instruction level.

[0102] In some implementations, when the number of instruction levels exceeds the number of user feature types, a many-to-one or many-to-many correspondence can be established between instruction levels and user feature information types. For example, when multi-level target instructions include high-level, mid-level, and low-level target instructions with decreasing levels of abstraction and / or decomposability, and user feature information includes target short-term memory and target long-term memory, based on the many-to-one correspondence, target long-term memory can be identified as target information corresponding to high-level and mid-level instruction levels, and target short-term memory can be identified as target information corresponding to low-level instructions; based on the many-to-many correspondence, both target long-term memory and target short-term memory can be identified as target information corresponding to high-level, mid-level, and low-level instruction levels.

[0103] In some implementations, more types of user feature information can be determined for instruction levels with lower levels of abstraction and / or decomposability. For example, in a multi-level target instruction system comprising high-level, mid-level, and low-level target instructions with decreasing levels of abstraction and / or decomposability, and where user feature information includes target short-term memory, target long-term memory, target preferences, and target characteristics, target short-term memory, target long-term memory, target preferences, and target characteristics can all be considered as target information corresponding to the low-level instruction level, while target long-term memory and target preferences can be considered as target information corresponding to the mid-level instruction level, and target preferences and target characteristics can be considered as target information corresponding to the high-level instruction level.

[0104] Step S1222: Based on the input information and the at least one type of target information, generate the target instruction corresponding to the instruction level.

[0105] Here, after determining at least one type of target information corresponding to each instruction level, a target instruction corresponding to each instruction level is generated based on the input information, the user task corresponding to the input information, and at least one type of target information.

[0106] For example, when the target information corresponding to the low-level instruction is the target long-term memory and the target short-term memory, the low-level target instruction is generated based on the target long-term memory and the target short-term memory, the user's input information, and the user's task.

[0107] In the embodiments provided in this disclosure, at least one type of target information is determined for each instruction level based on the degree of abstraction and / or decomposability of the target instruction and the degree of abstraction and / or generalization of the user feature information. This can achieve the effect of dynamically selecting the most suitable user feature information according to the needs of different instruction levels, thereby improving the rationality, interpretability and personalization of the instructions.

[0108] In some implementations, the multimodal data represents data from the user's first-person perspective; that is, multimodal data is data acquired by simulating or directly recording the user's visual, auditory, or tactile sensory experiences. In other implementations, multimodal data is data collected through a user's first-person perspective device, such as video and audio data collected through augmented reality (AR) glasses, action cameras, or smartphones worn by the user. Thus, because the multimodal data is from the user's first-person perspective, it can more comprehensively and realistically reflect the user's experiences in real life, recreating the user's dynamic decision-making process in complex scenarios.

[0109] In some implementations, the multimodal data includes cross-device data collected through at least one device associated with the user.

[0110] At least one device associated with the user can refer to a device worn or carried by the user. For example, AR glasses, smartwatches, smart earphones worn by the user; or smartphones, action cameras, etc. carried by the user. Furthermore, each of these devices is equipped with at least one sensor capable of collecting multimodal information, such as a camera and microphone in AR glasses and smartphones, a positioning device in smartwatches, a microphone in smart earphones, and a camera in action cameras.

[0111] In this way, the camera installed in the aforementioned device can be used to collect image or video data from the user's first-person perspective to simulate the user's vision; the microphone installed in the aforementioned device can be used to collect audio data from the user's first-person perspective to simulate the user's hearing.

[0112] Furthermore, since at least one device can collect any information from the surrounding environment from the user's first-person perspective, the multimodal data collected by these multiple devices is cross-device, cross-application, cross-ecosystem, and / or cross-real and virtual world data. For example, when using the camera and microphone in AR glasses to collect information about the surrounding environment, the camera can collect environmental information about the user's real world, as well as electronic data displayed or played on the user's smartphone, smartwatch, laptop, etc., and visual and audio data displayed and output by different applications running on the user's electronic devices.

[0113] In addition, when multimodal data is collected using multiple devices associated with a user, the multimodal data collected from these multiple devices can be integrated and summarized to obtain cross-device data.

[0114] In some implementations, the multimodal data includes cross-device data received from multiple devices associated with the user.

[0115] Here, cross-device data received by multiple devices related to the user can refer to cross-device data obtained by summarizing or integrating the data passively received by each device during its operation.

[0116] In some implementations, the data passively received during the operation of each device may include at least one of the following: user input information generated during user interaction with the application, application feedback information, and / or information received by the application in a non-user interaction state (e.g., emails received by an email application, communication messages received by a messaging application, etc.).

[0117] In some implementations, the multimodal data can be used to process cross-application user tasks initiated by the user.

[0118] Here, cross-application user tasks can refer to multiple user tasks initiated by a user through different applications or a single user task. For example, multiple user tasks initiated by a user through different agents running on the same device, multiple user tasks initiated by a user through different applications on different devices, and related tasks initiated by a user through multiple devices.

[0119] In this way, when it is determined that the user task was initiated by the current user, the user's input information can be understood based on the multimodal data, and the user feature information corresponding to the multimodal data can be used to assist the decision-making process of the user task.

[0120] In the embodiments provided in this application, by collecting multimodal data across devices or using multiple devices to receive multimodal data, the isolation between different devices, different applications, and different ecosystems can be broken down, thereby improving the diversity and comprehensiveness of the collected multimodal data. In addition, applying the multimodal data to user-initiated cross-application user tasks can expand the application scope of the multimodal data, thereby increasing its data value.

[0121] In some implementations, the method further includes steps S14 to S17:

[0122] Step S14: Generate at least one short-term memory based on the user's multimodal data.

[0123] Here, after acquiring the user's multimodal data, semantic analysis or attribute recognition is performed on the multimodal data to generate at least one short-term memory. In some implementations, this at least one short-term memory may be a structured, timestamped short-term event description obtained after preliminary understanding or analysis of the multimodal data, such as descriptive information of multimodal data collected within 1 minute or 2 minutes.

[0124] In some implementations, the multimodal data can be input into a first model to facilitate the generation of at least one corresponding short-term memory. In some implementations, the first model can implement any type of model, such as a large language model, a visual language model, a video model, a speech recognition model, etc.

[0125] Step S15: Generate at least one long-term memory based on the correlation information between the at least one short-term memory.

[0126] Here, the correlation information between at least one short-term memory can refer to any type of correlation information, such as logical correlation, temporal correlation, object correlation, etc.; among them, logical correlation refers to whether the at least one short-term memory belongs to the same event with a logical relationship; temporal correlation refers to whether the at least one short-term memory belongs to the same time period divided by a specified duration; object correlation refers to whether the at least one short-term memory is associated with the same object.

[0127] In this way, by integrating and summarizing at least one relevant short-term memory, the corresponding long-term memory can be obtained.

[0128] In some implementations, long-term memories can be generated in any manner. In some implementations, corresponding long-term memories can be generated based on the logical correlation between at least one short-term memory. For example, when a change in event boundary is detected in short-term memory, a long-term memory corresponding to the previous event is generated based on at least one short-term memory corresponding to the previous event. In some implementations, at least one long-term memory can be generated based on the temporal correlation between at least one short-term memory. For example, multiple short-term memories within a certain time period are summarized and integrated to generate a long-term memory corresponding to that time period, where the time period is a time period divided according to a specified duration, which can be 1 hour, 2 hours, or other durations.

[0129] Step S16: Generate at least one preference based on the at least one long-term memory.

[0130] Here, at least one long-term memory is integrated and summarized to determine the user's choice tendency in different scenarios, thereby generating user preferences in different scenarios.

[0131] In some implementations, user preference information can be generated based on at least one long-term memory and multiple short-term memories of the user.

[0132] In some implementations, a large language model can be used to integrate and summarize at least one long-term memory of a user, thereby generating at least one preference.

[0133] Step S17: Generate at least one trait based on the at least one preference.

[0134] Here, at least one user preference is aggregated to summarize the user's stable characteristics across scenarios, that is, at least one trait of the user.

[0135] Below, in conjunction with Figure 3 The process of generating personalized user information in one embodiment of this disclosure will be described below; wherein, the multimodal data related to the user is video data. The embodiment will now be described in conjunction with steps S301 to S309:

[0136] Step S301: The acquired video data 310 is segmented to obtain multiple video segments 320, such as segment 321, segment 322 and segment 323; then, step S302 is executed.

[0137] Here, video data 310 can be a video clip taken by the user some time later, such as a video clip taken by the user between 8:00 AM and 9:00 AM.

[0138] In some implementations, the video data 310 can be segmented in the following four ways, with the priority of these four methods decreasing in that order:

[0139] The first method is to use the video transition position as the cutting point when a video transition occurs; where a video transition refers to a situation where more than 80% of the pixels in the video frame change.

[0140] The second method is to use the end point of a continuous audio segment in the video as the segmentation point;

[0141] The third method involves dividing the segment at fixed time intervals; these fixed time intervals can be 1 minute, 2 minutes, etc.

[0142] The fourth method is to split the data according to the maximum fragment length; the maximum fragment length can be, for example, 5 minutes or 6 minutes.

[0143] Step S302: Each video segment after segmentation and the preset short-term memory extraction prompts are input into a multimodal model to generate corresponding short-term memories, such as short-term memory 331, short-term memory 332, short-term memory 333, short-term memory 334 and short-term memory 335; then, step S303 is executed.

[0144] Here, the multimodal model can output structured descriptive information for each short-term memory 330.

[0145] Step S303: Stream-based determination of the correlation between multiple short-term memories; treat multiple short-term memories with a correlation exceeding the correlation threshold as the same event; then, execute step S304.

[0146] For example, short-term memory 331, short-term memory 332, and short-term memory 333 are identified as event 341, and short-term memory 334 and short-term memory 335 are identified as event 342.

[0147] In some implementations, different short-term memories can be input into a large language model to determine whether the different short-term memories belong to the same event.

[0148] In some implementations, firstly, structured descriptions of multiple short-term memories are input into an embedding model to calculate semantic similarity scores between different short-term memories. Then, if the semantic similarity score is greater than the upper similarity limit, the corresponding different short-term memories are determined to belong to the same event; if the semantic similarity score is less than the lower similarity limit, the corresponding different short-term memories are determined not to belong to the same event. If the semantic similarity score is greater than the lower similarity limit but less than the upper similarity limit, the corresponding different short-term memories are input into a large language model to determine whether the different short-term memories belong to the same event.

[0149] In this way, multiple short-term memories can be divided into multiple events using the above method.

[0150] Step S304: Generate corresponding long-term memories based on multiple short-term memories belonging to the same event; generate at least one preference based on at least one long-term memory; generate at least one trait based on at least one long-term memory and at least one preference; then, execute step S305.

[0151] In some implementations, multiple short-term memories belonging to the same event are input into a large language model to generate corresponding long-term memories 350. For example, long-term memory 351 is generated based on event 341, and long-term memory 352 is generated based on event 342. In some implementations, long-term memory 350 includes identification information of the corresponding multiple short-term memories 330 for tracking the corresponding short-term memories.

[0152] In some implementations, at least one long-term memory is input into a large language model to integrate and summarize the at least one long-term memory to obtain at least one preference. For example, long-term memory 361 is generated based on long-term memory 351 and long-term memory 352; long-term memory 362 is generated based on long-term memory 352 and long-term memory 353.

[0153] In some implementations, at least one preference is input into a large language model to integrate and summarize the at least one preference, thereby obtaining at least one trait. For example, traits 371 and 372 are generated based on preferences 361 and 362.

[0154] In the embodiments provided in this disclosure, by simulating the cognitive mechanism of the human brain, multi-modal data with the same perspective as the user's first-person perspective are extracted in multiple layers to construct the user's multi-layered personalized traits, so as to better assist in processing user-initiated user tasks.

[0155] In some implementations, the preference has weighting information.

[0156] Here, weight information is used to indicate the importance of the corresponding preference. For example, the higher the weight, the easier it is for the corresponding preference to be retrieved or used in the user's decision-making process; or, the higher the weight, the greater the influence of the corresponding preference on the decision outcome.

[0157] Thus, the method further includes the following step S16:

[0158] Step S16, update the at least one preference; wherein,

[0159] If the similarity between an existing second preference and a newly generated first preference in at least one of the preferences is greater than a similarity threshold, the weight of the second preference is increased.

[0160] If the similarity between the newly generated first preference and any existing preference among the at least one preferences is less than or equal to the similarity threshold, the first preference is stored in the at least one preference, and an initial weight is assigned to the first preference.

[0161] Here, the first preference refers to the new preference generated based on the new multimodal data of the user and at least one long-term memory that has been generated when extracting traits based on the new multimodal data of the user.

[0162] When updating a first preference to at least one stored or generated preference, firstly, the similarity between the first preference and the at least one preference is determined using a large language model or embedding model; then, the update method for the at least one preference is determined based on the determined similarity.

[0163] At least one second preference already exists, referring to a preference whose similarity to the first preference is greater than a similarity threshold. When a second preference exists, it represents a preference that the user maintains over a long period. To avoid redundant storage of highly similar preference information, the at least one preference is updated by increasing its weight. In some implementations, while increasing the weight of the second preference, the semantic information of the second preference is updated using a large language model based on the semantic information of the first preference, so that the updated semantics of the second preference can more accurately represent the semantic information of the first preference.

[0164] Meanwhile, if the similarity between the newly generated first preference and any of the existing preferences is less than or equal to the similarity threshold, the first preference is characterized as a newly generated preference by the user. Therefore, the first preference is stored in the at least one stored preference and an initial weight is assigned to the first preference.

[0165] In some implementations, the initial weight of each preference can be any value, for example, 0.8.

[0166] In some implementations, the weight of the second preference can be increased by a specified percentage, for example, by increasing the current weight of the second preference by 5%.

[0167] In some implementations, the at least one preference includes a third preference. Thus, after continuously performing feature extraction based on the user's multimodal data, if the third preference is not updated, its weight is reduced, for example, by decreasing the current weight of the third preference by 5%. The specified number of times can be any preset number, such as 100 or 200 times.

[0168] In some implementations, the method further includes the step S17:

[0169] Step S17 increases the weight of the target preference.

[0170] Here, target preference refers to the preference used in step S122 above to generate multi-level target instructions. It is clear that target preference refers to the preference retrieved and referenced during the execution of user tasks.

[0171] In this way, the target preference is used in the user task decision-making process, representing that the target preference is a preference with high usage or high importance. Therefore, the weight of the target preference is increased so that the target preference is retrieved first in subsequent task processing.

[0172] In some implementations, the method further includes the following steps S18 to S19:

[0173] Step S18: Generate summary information based on the application scenario information, context information, and processing result corresponding to the input information;

[0174] Step S19: The summary information is associated with and stored with the target preference.

[0175] Here, application scenario information refers to the specific environment and context when the user inputs the information, such as dining, work, or entertainment scenarios.

[0176] Contextual information refers to any information related to the user task corresponding to the input information. Examples include the task objective, the environmental information perceived during task execution, and user feedback during and after task execution.

[0177] In this way, the application scenario information, context information, and processing results corresponding to the input information are summarized to extract key information and generate corresponding summary information.

[0178] The generated summary information is stored in association with the target preference, so that the summary information can be retrieved simultaneously when the target preference is retrieved.

[0179] In the embodiments provided in this disclosure, by generating application scenario information, context information, and summary information corresponding to the processing results of the user input information, and storing the summary information in association with the target preference, when the target preference is retrieved again and applied to the processing of the user task, the application scenario, task objective, execution result, and user feedback in the summary information can be referenced to make task decisions, thereby helping to improve the accuracy of the next task decision.

[0180] Below, in conjunction with Figure 4 The following describes a task processing flow in one embodiment provided in this disclosure; wherein at least one intelligent agent is used to implement the task processing method. The embodiment will now be described in conjunction with steps S401 to S409:

[0181] Step S401: The task-generating agent 420 receives user input information 410; then, step S402 is executed.

[0182] Here, the user's input information 410 is "book a fun museum".

[0183] In step S402, the task generation agent 420 calls the large language model to generate initial multi-level instructions 430; then, step S403 is executed.

[0184] In some implementations, the large language model performs semantic understanding and enhancement on the input information 410 to obtain enhanced input information. For example, after enhancing the above input information, the result is: "Analyze the user's travel information and plans, help the user find a museum that meets the 'fun' criteria, and provide ticket booking information."

[0185] In this way, the large language model generates initial multi-level instructions 430 based on the enhanced input information, namely, initial high-level instructions, initial intermediate-level instructions, and initial low-level instructions.

[0186] For example, initial advanced instructions include: "Select museum type", "When and where"; initial intermediate instructions include: "Find commonly used apps for museums", "Plan geographical location", "Read from where to form calendar"; initial low-level instructions include: "Open app to get museum list", "Read this week's calendar and check", "Find items and free time in calendar".

[0187] In step S403, the task generating agent 420 retrieves the user's personalized information based on the input information to obtain the target personalized information 440; then, step S404 is executed.

[0188] Here, the retrieved user's targeted personalization information can include targeted short-term memory, targeted long-term memory, targeted preferences, and targeted traits related to the input information. For example, targeted traits include "habit of waking up early" and "price sensitivity"; targeted preferences include "history," "geography," "money," and "preference for using the first and second apps"; targeted long-term memory includes "using the second app," etc.

[0189] Step S404: Based on the target personalized information 440, the task generation agent 420 generates the initial multi-level instructions 430 corresponding to the updated multi-level instructions 450; then, step S405 is executed.

[0190] Here, firstly, the enhanced input information, target traits, target preferences, and initial advanced instructions are input into the large language model. Combined with prompt word templates for generating advanced instructions, the large language model generates updated advanced instructions. For example, based on the search results, if the user's target preferences are "history, geography, and currency" and their target trait is "habitual early riser," then the updated advanced instruction would be: "Help the user find museums related to 'history, geography, and currency,' and since the user is 'habitual early riser,' they can book tickets for when the museum opens."

[0191] Then, the updated advanced instructions, enhanced input information, target preferences, and initial intermediate instructions are input into the large language model. Combined with the prompt word template for generating intermediate instructions, the large language model generates updated intermediate instructions. For example, based on the search results, if the user's target preference is to use "first application and second application," then the updated intermediate instructions would be: "Read the user's calendar and itinerary information; find the user's geographical location for the corresponding time period; use 'first application or second application' to find a list of museums; list the museums' opening hours."

[0192] Next, the updated intermediate-level instructions, enhanced input information, target traits, target preferences, target long-term memory, and initial low-level instructions are input into the large language model. Combined with the prompt word template for generating low-level instructions, the large language model generates updated low-level instructions. For example, based on the search results, if the target trait is "price-sensitive" and the target long-term memory is "recently used second application," the generated updated low-level instructions would be: "Call the calendar and itinerary information interfaces to group user travel information by time; retrieve destinations from the travel information; call the 'second application's' attraction list interface, with parameters including 'type: museum'; sorting: because the user is 'price-sensitive,' sort by price in ascending order; obtain the ticket subscription link."

[0193] In step S405, the task execution agent 460 executes the low-level instructions sequentially according to the updated order of the low-level instructions, and obtains the instruction execution result 470; then, step S406 is executed.

[0194] In step S406, the task execution agent 460 sends the instruction execution result 470 and the low-level instruction to the verification agent 480 for evaluation; the verification agent 480 determines whether the instruction execution result 470 meets the requirements of the low-level instruction; if it does, then step S407 is executed; if it does not, then steps S405 to S406 are executed again.

[0195] Here, in some implementations, the task execution agent 460 executes multiple low-level instructions one by one, and after each low-level instruction is executed, the execution result of the low-level instruction is handed over to the verification agent 480 for verification.

[0196] In some implementations, after the task execution agent 460 has completed the execution of all the multiple low-level instructions, the complete execution result is handed over to the verification agent 470 for verification.

[0197] In step S407, after the low-level instructions are executed, the task execution agent 460 merges and rewrites the low-level instructions and their execution results, and submits them to the verification agent 480 for evaluation to determine whether the execution results of the low-level instructions meet the requirements of the corresponding intermediate instructions. If they do, step S408 is executed; if they do not, the large language model is called to optimize the low-level instructions based on the intermediate instructions, and steps S405 to S407 are re-executed based on the optimized low-level instructions.

[0198] In step S408, the task execution agent 460 merges and rewrites the intermediate instructions and their execution results, and submits them to the verification agent 480 for evaluation to determine whether the execution results of the intermediate instructions meet the requirements of the corresponding high-level instructions. If they do, step S409 is executed. If they do not, the large language model is called to optimize the intermediate instructions based on the high-level instructions, optimize the low-level instructions based on the optimized intermediate instructions, and re-execute steps S405 to S408 based on the optimized intermediate and low-level instructions.

[0199] Step S409: After all advanced instructions have been executed, the updated execution result 470 is obtained; the updated execution result 470 is output to the user 490.

[0200] Here, detailed information about the museums found and links to ticket subscriptions will be provided to users in Markdown format.

[0201] As can be seen from the above embodiments, during the execution of user tasks, by calling the user's short-term memory, long-term memory, preferences, and traits constructed layer by layer, multi-level personalized information is invoked. This allows for precise matching of user needs during task execution, providing users with a personalized service experience, and thus improving task execution accuracy and user satisfaction. On the other hand, the user's personalized information is constructed by simulating the cognitive mechanism of the human brain and using multimodal data collected from the user's first-person perspective. This can more comprehensively, realistically, and across devices, applications, and ecosystems reflect the user's personal characteristics. Therefore, understanding the user's intent based on this personalized information and generating multi-level instructions can further improve the accuracy of understanding the user's intent, increase the personalization of multi-level instructions, and make the execution results more in line with the user's true intentions.

[0202] Based on the foregoing embodiments, this disclosure provides a task processing device, which includes the included units and the modules included in each unit, which can be implemented by a processor in a computer device; of course, it can also be implemented by specific logic circuits; in the implementation process, the processor can be a central processing unit (CPU), a microprocessor unit (MPU), a digital signal processor (DSP), or a field programmable gate array (FPGA), etc.

[0203] Figure 5 This is a schematic diagram of the composition structure of a task processing device provided in this disclosure, such as... Figure 5 As shown, the task processing device 500 includes: an acquisition module 510, a generation module 520, and an execution module 530, wherein:

[0204] A task processing device, comprising:

[0205] The acquisition module 510 is used to acquire user input information;

[0206] The generation module 520 is used to generate multi-level target instructions based on the input information and the user's personalized information; wherein, the personalized information represents user feature information generated based on multimodal data; the multi-level target instructions represent instructions for executing the user task corresponding to the input information, and the target instructions at different instruction levels correspond to different user feature information;

[0207] The execution module 530 is used to execute the multi-level target instructions to obtain the processing result of the user task corresponding to the input information.

[0208] In some embodiments, the user feature information includes at least two of the following: at least one short-term memory, at least one long-term memory, at least one preference, and at least one trait; the level of abstraction and / or generalization of the short-term memory, the long-term memory, the preference, and the trait increases sequentially.

[0209] The generation module 520 is used for:

[0210] Based on the input information, at least two categories from the user feature information are determined: target short-term memory, target long-term memory, target preferences, and target traits.

[0211] Based on different instruction levels and at least two of the following: target short-term memory, target long-term memory, target preferences, and target features, the multi-level target instructions are generated; wherein the different instruction levels are obtained based on instruction planning of the input information.

[0212] In some embodiments, the target instructions at different instruction levels correspond to different levels of abstraction and / or decomposability.

[0213] The generation module 520 is used for:

[0214] For each instruction level, based on the corresponding level of abstraction and / or decomposability, at least one type of target information is determined from at least two of the target short-term memory, the target long-term memory, the target preference, and the target trait; wherein, the higher the level of abstraction and / or the higher the level of decomposability of the target instruction at the instruction level, the higher the level of abstraction and / or the higher the generality of the corresponding target information.

[0215] Based on the input information and the at least one type of target information, a target instruction corresponding to the instruction level is generated.

[0216] In some embodiments, the multimodal data represents data from the user's first-person perspective;

[0217] The multimodal data includes cross-device data collected through at least one device associated with the user;

[0218] And / or,

[0219] The multimodal data includes cross-device data received from multiple devices associated with the user;

[0220] And / or,

[0221] The multimodal data can be used to process cross-application user tasks initiated by the user.

[0222] In some embodiments, the device 500 further includes a feature generation module;

[0223] The feature generation module is used for:

[0224] Based on the user's multimodal data, generate at least one short-term memory;

[0225] Based on the correlation information between the at least one short-term memory, generate at least one long-term memory;

[0226] Based on the at least one long-term memory, generate at least one preference;

[0227] At least one trait is generated based on the at least one preference.

[0228] In some embodiments, the preference has weight information; the feature generation module is further configured to update the at least one preference; wherein,

[0229] If the similarity between an existing second preference and a newly generated first preference in at least one of the preferences is greater than a similarity threshold, the weight of the second preference is increased.

[0230] If the similarity between the newly generated first preference and any existing preference among the at least one preferences is less than or equal to the similarity threshold, the first preference is stored in the at least one preference, and an initial weight is assigned to the first preference.

[0231] In some embodiments, the feature generation module is further configured to increase the weight of the target preference.

[0232] In some embodiments, the feature generation module is further configured to:

[0233] Based on the application scenario information, context information, and processing results corresponding to the input information, a summary information is generated.

[0234] The summary information is stored in association with the target preference.

[0235] The descriptions of the apparatus embodiments above are similar to those of the method embodiments above, and have similar beneficial effects. In some embodiments, the functions or modules included in the apparatus provided in this disclosure can be used to perform the methods described in the method embodiments above. For technical details not disclosed in the apparatus embodiments of this disclosure, please refer to the descriptions of the method embodiments of this disclosure for understanding.

[0236] If the technical solution disclosed herein involves personal information, the product using this technical solution has clearly informed the user of the personal information processing rules and obtained the user's voluntary consent before processing the personal information. If the technical solution disclosed herein involves sensitive personal information, the product using this technical solution has obtained the user's separate consent before processing the sensitive personal information, and also meets the requirement of "express consent". For example, at personal information collection devices such as cameras, clear and prominent signs are set up to inform users that they have entered the scope of personal information collection and that personal information will be collected. If an individual voluntarily enters the collection scope, it is deemed that they have agreed to the collection of their personal information; or on the personal information processing device, with clear signs / information informing users of the personal information processing rules, authorization is obtained from the individual through pop-up information or by asking the individual to upload their personal information; wherein, the personal information processing rules may include information such as the personal information processor, the purpose of personal information processing, the processing method, and the types of personal information processed.

[0237] It should be noted that, in the embodiments of this disclosure, if the above-described task processing method is implemented as a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the embodiments of this disclosure, or the part that contributes to related technologies, can be embodied in the form of a software product. This software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the methods described in the various embodiments of this disclosure. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), magnetic disks, or optical disks. Thus, the embodiments of this disclosure are not limited to any specific hardware, software, or firmware, or any combination of hardware, software, and firmware.

[0238] This disclosure provides a computer device including a memory and a processor. The memory stores a computer program that can run on the processor. When the processor executes the program, it implements some or all of the steps in the above-described method.

[0239] This disclosure provides a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements some or all of the steps in the above-described method. The computer-readable storage medium may be transient or non-transient.

[0240] This disclosure provides a computer program including computer-readable code, wherein when the computer-readable code is executed in a computer device, a processor in the computer device performs some or all of the steps in the above-described method.

[0241] This disclosure provides a computer program product, which includes a non-transitory computer-readable storage medium storing a computer program. When the computer program is read and executed by a computer, it implements some or all of the steps in the above-described method. This computer program product can be implemented specifically through hardware, software, or a combination thereof. In some embodiments, the computer program product is specifically embodied as a computer storage medium; in other embodiments, the computer program product is specifically embodied as a software product, such as a software development kit (SDK), etc.

[0242] It should be noted that the descriptions of the various embodiments above tend to emphasize the differences between them, while their similarities or commonalities can be referenced interchangeably. The descriptions of the above embodiments of the device, storage medium, computer program, and computer program product are similar to the descriptions of the above method embodiments and have similar beneficial effects. For technical details not disclosed in the embodiments of the device, storage medium, computer program, and computer program product of this disclosure, please refer to the descriptions of the method embodiments of this disclosure for understanding.

[0243] It should be noted that, Figure 6 This is a schematic diagram of a hardware entity of the electronic device in this disclosure, such as... Figure 6 As shown, the hardware entity of the electronic device 600 includes: a processor 601, a communication interface 602, and a memory 603, wherein:

[0244] The processor 601 typically controls the overall operation of the electronic device 600.

[0245] Communication interface 602 enables electronic devices to communicate with other terminals or servers via a network.

[0246] The memory 603 is configured to store instructions and applications executable by the processor 601, and can also cache data to be processed or already processed (e.g., image data, audio data, voice communication data, and video communication data) in the processor 601 and various modules in the electronic device 600. It can be implemented using flash memory or random access memory (RAM). Data transfer between the processor 601, the communication interface 602, and the memory 603 can be performed via bus 604.

[0247] It should be understood that the phrase "an embodiment" or "one embodiment" throughout the specification means that a specific feature, structure, or characteristic related to the embodiment is included in at least one embodiment of this disclosure. Therefore, "in one embodiment" or "one embodiment" appearing throughout the specification does not necessarily refer to the same embodiment. Furthermore, these specific features, structures, or characteristics can be combined in any suitable manner in one or more embodiments. It should be understood that in the various embodiments of this disclosure, the sequence numbers of the above steps / processes do not imply a sequential order of execution; the execution order of each step / process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this disclosure. The sequence numbers of the above embodiments of this disclosure are merely descriptive and do not represent the superiority or inferiority of the embodiments.

[0248] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.

[0249] In the several embodiments provided in this disclosure, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of units is only a logical functional division, and in actual implementation, there may be other division methods, such as: multiple units or components may be combined, or integrated into another system, or some features may be ignored or not executed. In addition, the coupling, direct coupling, or communication connection between the various components shown or discussed may be through some interfaces, and the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms.

[0250] The units described above as separate components may or may not be physically separate. The components shown as units may or may not be physical units. They may be located in one place or distributed across multiple network units. Some or all of the units may be selected to achieve the purpose of this embodiment according to actual needs.

[0251] In addition, each functional unit in the various embodiments of this disclosure can be integrated into one processing unit, or each unit can be a separate unit, or two or more units can be integrated into one unit; the integrated unit can be implemented in hardware or in the form of hardware plus software functional units.

[0252] Those skilled in the art will understand that all or part of the steps of the above method embodiments can be implemented by hardware related to program instructions. The aforementioned program can be stored in a computer-readable storage medium. When the program is executed, it performs the steps of the above method embodiments. The aforementioned storage medium includes various media that can store program code, such as mobile storage devices, read-only memory (ROM), magnetic disks, or optical disks.

[0253] Alternatively, if the integrated units described above are implemented as software functional modules and sold or used as independent products, they can also be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this disclosure, or the part that contributes to related technologies, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the methods described in the various embodiments of this disclosure. The aforementioned storage medium includes various media capable of storing program code, such as mobile storage devices, ROM, magnetic disks, or optical disks.

[0254] The above description is merely an embodiment of this disclosure, but the scope of protection of this disclosure is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this disclosure should be included within the scope of protection of this disclosure.

Claims

1. A task processing method, comprising: Obtain user input information; Based on the input information and the user's personalized information, multi-level target instructions are generated; wherein, the personalized information represents user feature information generated based on multimodal data; the multi-level target instructions represent instructions for executing the user task corresponding to the input information, and target instructions at different instruction levels correspond to different user feature information; Execute the multi-level target instructions to obtain the processing result of the user task corresponding to the input information.

2. The method according to claim 1, wherein the user feature information includes at least two of the following: at least one short-term memory, at least one long-term memory, at least one preference, and at least one trait; the level of abstraction and / or generalization of the short-term memory, the long-term memory, the preference, and the trait increases sequentially. The generation of multi-level target instructions based on the input information and the user's personalized information includes: Based on the input information, at least two categories from the user feature information are determined: target short-term memory, target long-term memory, target preferences, and target traits. Based on different instruction levels and at least two of the following: target short-term memory, target long-term memory, target preferences, and target features, the multi-level target instructions are generated; wherein, the different instruction levels are obtained based on instruction planning of the input information.

3. The method according to claim 2, wherein the target instructions at different instruction levels correspond to different levels of abstraction and / or decomposability; The generation of the multi-level target instructions based on different instruction levels and at least two of the following: target short-term memory, target long-term memory, target preferences, and target traits, includes: For each instruction level, based on the corresponding level of abstraction and / or decomposability, at least one type of target information is determined from at least two of the target short-term memory, the target long-term memory, the target preferences, and the target traits; Based on the input information and the at least one type of target information, generate the target instruction corresponding to the instruction level; The higher the level of abstraction and / or decomposability of the target instruction at the instruction level, the higher the level of abstraction and / or generalization of the corresponding target information.

4. The method according to any one of claims 1 to 3, wherein the multimodal data representation perspective is data from the user's first-person perspective; wherein, The multimodal data includes cross-device data collected through at least one device associated with the user; And / or, The multimodal data includes cross-device data received from multiple devices associated with the user; And / or, The multimodal data can be used to process cross-application user tasks initiated by the user.

5. The method according to claim 2, further comprising: Based on the user's multimodal data, generate at least one short-term memory; Based on the correlation information between the at least one short-term memory, generate at least one long-term memory; Based on the at least one long-term memory, generate at least one preference; At least one trait is generated based on the at least one preference.

6. The method according to claim 5, wherein the preference has weight information, and the method further comprises: Update at least one preference; wherein, If the similarity between an existing second preference and a newly generated first preference in at least one of the preferences is greater than a similarity threshold, the weight of the second preference is increased. If the similarity between the newly generated first preference and any existing preference among the at least one preferences is less than or equal to the similarity threshold, the first preference is stored in the at least one preference, and an initial weight is assigned to the first preference.

7. The method according to claim 6, further comprising: Increase the weight of the target preference.

8. The method according to claim 2, further comprising: Based on the application scenario information, context information, and processing results corresponding to the input information, a summary information is generated. The summary information is stored in association with the target preference.

9. A task processing apparatus, comprising: The acquisition module is used to acquire user input information; The generation module is used to generate multi-level target instructions based on the input information and the user's personalized information; wherein, the personalized information represents user feature information generated based on multimodal data; the multi-level target instructions represent instructions for executing the user task corresponding to the input information, and the target instructions at different instruction levels correspond to different user feature information; The execution module is used to execute the multi-level target instructions and obtain the processing result of the user task corresponding to the input information.

10. An electronic device comprising a memory and at least one processor; wherein, The memory stores a computer program that can run on the at least one processor, which, when executing the program, implements the steps of the method according to any one of claims 1 to 8.