Task processing method and device, equipment, storage medium and program product

By identifying multiple candidate processing results and evaluating their quality scores during task processing, the problem of insufficient accuracy and flexibility in traditional task processing methods is solved, and more efficient human-computer interaction is achieved.

CN121833097APending Publication Date: 2026-04-10BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
BEIJING ZITIAO NETWORK TECH CO LTD
Filing Date
2024-10-08
Publication Date
2026-04-10

AI Technical Summary

Technical Problem

Traditional task processing methods rely on user input for high accuracy but lack flexibility, failing to provide multiple processing results for users to choose from, resulting in inaccurate and inflexible processing outcomes.

Method used

Based on the context information of the user request, multiple task processing modes are determined, multiple candidate processing results are generated, and the optimal processing result is selected and a response is provided through quality score evaluation.

Benefits of technology

It improves the accuracy and flexibility of task processing, and can provide multiple processing results to choose from according to user needs, thus enhancing the effectiveness of human-computer interaction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121833097A_ABST
    Figure CN121833097A_ABST
Patent Text Reader

Abstract

The invention provides a task processing method and device, equipment, a storage medium and a program product. The method comprises the following steps: based on context information associated with a user request of a target user, determining a group of tasks matched with a user demand indicated by the user request; for each task in the group of tasks, according to a plurality of task processing modes, determining a plurality of candidate processing results of the task based on the context information, determining respective quality scores of the plurality of candidate processing results of the task, and determining the quality scores of the plurality of candidate processing results of the task at least based on the respective quality scores of the plurality of candidate processing results; determining at least one target processing result for the task from the plurality of candidate processing results; and providing a response to the user request at least based on the at least one target processing result of the group of tasks. Therefore, the processing accuracy and effectiveness of the task expected to be executed by the user in the man-machine interaction process can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] Example embodiments of the present disclosure generally relate to the field of computers, and in particular, to a method, an apparatus, an electronic device, a computer-readable storage medium and a computer program product for task processing. BACKGROUND

[0002] With the development of information technology, various terminal devices can provide various services to people in work and life, etc. For example, an application providing services can be deployed in a terminal device. The terminal device or the application can provide a task processing function to a user to assist the user in using the terminal device or the application. The terminal device can receive a user request from the user, determine a task based on the user request, and determine a processing result of the task. SUMMARY

[0003] In a first aspect of the present disclosure, a method for task processing is provided. The method comprises: determining, based on context information associated with a user request of a target user, a set of tasks matching a user demand indicated by the user request, the context information comprising at least the user request; for each task in the set of tasks, determining, according to a plurality of task processing modes, a plurality of candidate processing results of the task based on the context information respectively, different task processing modes in the plurality of task processing modes defining different task processing strategies, determining a quality score of each of the plurality of candidate processing results of the task, and determining at least one target processing result for the task from the plurality of candidate processing results based on at least the quality score of each of the plurality of candidate processing results; and providing a response to the user request based on at least the at least one target processing result of each of the set of tasks.

[0004] In a second aspect of the present disclosure, an apparatus for task processing is provided. The apparatus comprises: a task determination module configured to determine, based on context information associated with a user request of a target user, a set of tasks matching a user demand indicated by the user request, the context information comprising at least the user request; a result determination module configured to, for each task in the set of tasks, determine, according to a plurality of task processing modes, a plurality of candidate processing results of the task based on the context information respectively, different task processing modes in the plurality of task processing modes defining different task processing strategies, determine a quality score of each of the plurality of candidate processing results of the task, and determine at least one target processing result for the task from the plurality of candidate processing results based on at least the quality score of each of the plurality of candidate processing results; and a response providing module configured to provide a response to the user request based on at least the at least one target processing result of each of the set of tasks.

[0005] In a third aspect of the disclosure, an electronic device is provided. The electronic device includes at least one processing unit; and at least one memory coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, the instructions when executed by the at least one processing unit cause the electronic device to perform the method according to the first aspect of the disclosure.

[0006] In a fourth aspect of the disclosure, a computer-readable storage medium is provided, having stored thereon a computer program which, when executed by a processor, causes the processor to perform the method according to the first aspect of the disclosure.

[0007] In a fifth aspect of the disclosure, a computer program product is provided. The computer program product is tangibly stored in a computer storage medium and includes computer- executable instructions that, when executed by a device, cause the device to perform the method of the first aspect.

[0008] It should be understood that nothing in the Summary is to be construed as a limitation of any embodiment of the disclosure to a particular set of features or functions described in the Summary. Other features, aspects, and advantages of the disclosure will become apparent from the following detailed description, taken in conjunction with the accompanying drawings. BRIEF DESCRIPTION OF DRAWINGS

[0009] The above and other features, aspects, and advantages of embodiments of the disclosure will become more apparent from the following detailed description taken in conjunction with the accompanying drawings. In the drawings, like or similar elements are referred to with like or similar reference numerals, in which:

[0010] Figure 1 A schematic diagram showing an example environment in which embodiments of the disclosure can be implemented is shown;

[0011] Figure 2 An example for task processing according to some embodiments of the disclosure is shown;

[0012] Figure 3 A flowchart of a method for task processing according to some embodiments of the disclosure is shown;

[0013] Figure 4 An exemplary block diagram of an apparatus for task processing according to some embodiments of the disclosure is shown; and

[0014] Figure 5 A block diagram of an electronic device that can implement one or more embodiments of the disclosure is shown. DETAILED DESCRIPTION

[0015] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.

[0016] In the description of embodiments of this disclosure, the term "comprising" and similar terms should be understood as open-ended inclusion, i.e., "including but not limited to". The term "based on" should be understood as "at least partially based on". The term "one embodiment" or "the embodiment" should be understood as "at least one embodiment". The term "some embodiments" should be understood as "at least some embodiments". Other explicit and implicit definitions may also be included below.

[0017] In this document, unless explicitly stated otherwise, performing a step in response to A does not mean that the step is performed immediately after A, but may include one or more intermediate steps.

[0018] It is understood that the data involved in this technical solution (including but not limited to the data itself, the acquisition, use, storage or deletion of the data) shall comply with the requirements of relevant laws, regulations and related provisions.

[0019] It is understood that before using the technical solutions disclosed in the various embodiments of this disclosure, users should be informed of the types, scope of use, and usage scenarios of the personal information involved in this disclosure through appropriate means in accordance with relevant laws and regulations, and user authorization should be obtained.

[0020] For example, in response to receiving a user's active request, a prompt message is sent to the user to clearly inform the user that the requested operation will require the acquisition and use of the user's personal information, thereby enabling the user to choose whether to provide personal information to the software or hardware such as electronic devices, applications, servers or storage media that perform the operation of the technical solution disclosed herein, based on the prompt message.

[0021] As an optional but non-restrictive implementation, in response to a user's active request, a prompt message can be sent to the user, such as a pop-up window, where the prompt message can be presented in text format. Furthermore, the pop-up window can also include a selection control allowing the user to choose "agree" or "disagree" to provide personal information to the electronic device.

[0022] It is understood that the above notification and user authorization process are merely illustrative and do not constitute a limitation on the implementation of this disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of this disclosure.

[0023] As used in this paper, the term "model" refers to a model that learns the relationship between inputs and outputs from training data, enabling it to generate corresponding outputs for a given input after training. Model generation can be based on machine learning techniques. Deep learning is a machine learning algorithm that processes inputs and provides corresponding outputs using multiple layers of processing units. A neural network model is an example of a deep learning-based model. In this paper, "model" may also be referred to as a "machine learning model," "learning model," "machine learning network," or "learning network," and these terms are used interchangeably.

[0024] A neural network is a machine learning network based on deep learning. A neural network processes input and provides a corresponding output, typically consisting of an input layer, an output layer, and one or more hidden layers between the input and output layers. Neural networks used in deep learning applications often include many hidden layers, thus increasing the network's depth. The layers of a neural network are connected sequentially, so that the output of the previous layer is provided as the input to the next layer. The input layer receives the input to the neural network, while the output layer's output serves as the final output. Each layer of a neural network includes one or more nodes (also called processing nodes or neurons), each node processing the input from the layer above.

[0025] Machine learning typically comprises three phases: training, testing, and application (also known as inference). In the training phase, a given model is trained using a large amount of training data, iteratively updating parameter values ​​until the model can consistently generate inferences that meet the expected goals from the training data. Through training, the model can be considered to have learned the relationship between inputs and outputs (also known as an input-output mapping) from the training data. The parameter values ​​of the trained model are determined. In the testing phase, test inputs are applied to the trained model to test whether it can provide the correct output, thus determining the model's performance. The testing phase can sometimes be integrated into the training phase. In the application or inference phase, the trained model can be used to process actual model inputs based on the trained parameter values ​​to determine the corresponding model output.

[0026] Figure 1A schematic diagram of an example environment 100 in which embodiments of the present disclosure can be implemented is shown. In this example environment 100, an application 112 and a digital assistant 114 are installed on a client device 110. A user 140 can interact with the application 112 via the client device 110 and / or an attached device of the client device 110. In some implementations, the application 112 may be authorized to capture speech via an audio capture device (e.g., a microphone) of the client device 110, capture images via an image capture device (e.g., a camera) of the client device 110, and so on.

[0027] In some embodiments, application 112 and digital assistant 114 may be downloaded and installed on client device 110. In some embodiments, application 112 and digital assistant 114 may also be accessed in other ways, such as through a web page.

[0028] In embodiments of this disclosure, application 112 can be any suitable application with task processing capabilities, which may include, but is not limited to, one or more of the following: chat application components (also known as instant messaging application components), browser application components, planning application components, document application components, audio / video conferencing application components, email application components, task application components, calendar application components, goal and key results (OKR) application components, etc. It is understood that, although Figure 1 The image shows a single application component, but in reality, multiple application components can be installed on the client device 110. In some embodiments, application 112 may include a multi-functional collaboration platform, such as an office collaboration platform (also known as an office suite), which can provide integration of various types of business components to facilitate people's office work, communication, and other activities. In a multi-functional collaboration platform, people can launch different business components as needed to complete corresponding information processing, sharing, communication, etc.

[0029] In some embodiments, the digital assistant 114 may be provided by a separate application business component, or it may be integrated into an application 112 capable of providing content entities. The application business component providing the client interface for the digital assistant may correspond to a single-function application business component or a multi-functional collaboration platform, such as an office suite or other collaboration platform capable of integrating multiple components. It is understood that, similar to application business components, although... Figure 1 The image shows a single digital assistant, but there can actually be multiple digital assistants.

[0030] In some embodiments, the digital assistant 114 supports the use of plugins. Each plugin can provide one or more functions of the application. Such plugins include, but are not limited to, one or more of the following: search plugin, contact plugin, messaging plugin, document plugin, form plugin, email plugin, calendar plugin, schedule plugin, task plugin, etc.

[0031] Digital assistant 114 is a user's intelligent assistant, possessing intelligent dialogue and information processing capabilities. In embodiments of this disclosure, digital assistant 114 is used to interact with user 140 to assist user 140 in using terminal devices or applications. In some embodiments, multiple interaction modes between user 140 and digital assistant 114 can be provided, and users can flexibly switch between these modes. When a certain interaction mode is triggered, a corresponding interaction area is presented to facilitate interaction between user 140 and digital assistant 114. The interaction methods between user 140 and digital assistant 114 differ under different interaction modes, thus flexibly adapting to the interaction needs of different application scenarios.

[0032] In environment 100, in response to the launch of application 112, client device 110 may present an interface 150 of application 112 and / or digital assistant 114. Interface 150 may, for example, include an interactive interface for application 112 and digital assistant 114. In some embodiments, interface 150 may present an interaction window between user 140 and digital assistant 114. In the interaction window, user 140 can converse with digital assistant 114 by inputting natural language, images, audio files, video files, web page files, etc., to instruct the digital assistant to assist in completing various tasks.

[0033] The interaction window between the digital assistant 114 and the user 140 may include a session window, such as a session window in an instant messaging application or an instant messaging module of a specific application. In the session window, the interaction between the digital assistant 114 and the user 140 may be presented in the form of session messages. Alternatively or additionally, the interaction window between the digital assistant 114 and the user 140 may also include other types of windows, such as a floating window, in which the user 140 can trigger the digital assistant 114 to perform corresponding operations by entering commands, selecting shortcuts, etc.

[0034] In some embodiments, the digital assistant 114 may support a conversation window interaction mode, also known as conversation mode. In this interaction mode, a conversation window is presented between the user 140 and the digital assistant 114, where the user 140 and the digital assistant 114 interact through conversation messages. In conversation mode, the digital assistant 114 can perform tasks based on the conversation messages in the conversation window. In the interaction window, the user 140 inputs interaction messages, and the digital assistant 114 responds to the user's input by providing a reply message. A conversation window with the digital assistant 114 can be opened by selecting the digital assistant 114. The conversation window may include interface elements for information interaction, such as input boxes, message lists, message bubbles, etc.

[0035] In some embodiments, a communication connection is established between the client device 110 and the server device 120. The communication connection can be established via wired or wireless means. The communication connection may include, but is not limited to, Bluetooth, mobile network, Universal Serial Bus (USB), and Wireless Fidelity (WiFi) connections, etc., and the embodiments of this disclosure are not limited in this respect. In the embodiments of this disclosure, the client device 110 and the server device 120 can perform signaling interaction through their communication connection to provide services to the application 112 and / or the digital assistant 114.

[0036] like Figure 1 As shown, server device 120 can invoke one or more machine learning models, such as machine learning model 130-1, machine learning model 130-2, ..., machine learning model 130-M, where M is any suitable positive integer. For ease of description, the one or more machine learning models are collectively referred to as machine learning model 130 below, to support the task processing functions of application 112 based on the output of machine learning model 130. Machine learning model 130 can be deployed on server device 120 or on other devices.

[0037] Machine learning model 130 can be based on any suitable model architecture, including but not limited to Transformer models, convolutional neural networks (CNNs), recurrent neural networks (RNNs), deep neural networks (DNNs), and so on. In some embodiments, machine learning model 130 can be based on a language model (LM). A language model, by learning from a large corpus, is capable of question answering. Machine learning model 130 can also be based on other suitable models. It should be noted that if machine learning model 130 includes multiple machine learning models, these multiple machine learning models can have different uses and functions, and this disclosure does not limit them.

[0038] Client device 110 can be any type of mobile terminal, fixed terminal, or portable terminal, including mobile phones, desktop computers, laptop computers, notebook computers, netbook computers, tablet computers, media computers, multimedia tablets, personal communication system (PCS) devices, personal navigation devices, personal digital assistants (PDAs), audio / video players, digital cameras / camcorders, positioning devices, television receivers, radio receivers, e-book devices, gaming devices, or any combination thereof, including accessories and peripherals of these devices or any combination thereof. In some embodiments, client device 110 may also support any type of user-facing interface (such as "wearable" circuitry).

[0039] The server-side device 120 can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks, and big data and artificial intelligence platforms. The server-side device 120 may include, for example, computing systems / servers such as mainframes, edge computing nodes, and computing devices in a cloud environment, etc.

[0040] It should be understood that the structure and function of the various elements in environment 100 are described for illustrative purposes only and do not imply any limitation on the scope of this disclosure.

[0041] In human-computer interaction, users can input requests via voice or text. By analyzing the user's needs corresponding to the received requests, the task to be performed is determined and the task processing is provided. Traditionally, the task input from the user is received, and the processing result is determined based on predetermined rules, methods, or algorithms, and then provided to the user. On the one hand, this task processing method relies on the accuracy of the user's input task; if the user's input task is inaccurate, the final processing result provided to the user will also be inaccurate. On the other hand, this task processing method has poor flexibility, usually only determining a single processing result and not providing multiple processing results for the user to choose from.

[0042] In view of the above, according to embodiments of the present disclosure, an improved task processing scheme is provided. According to the scheme of the embodiments of the present disclosure, based on context information associated with a user request of a target user, a set of tasks matching the user needs indicated by the user request is determined, the context information including at least the user request. For each task in the set of tasks, according to multiple task processing modes, multiple candidate processing results for the task are determined based on the context information. Different task processing modes in the multiple task processing modes define different task processing strategies. A quality score is determined for each of the multiple candidate processing results of the task. Based at least on the quality scores of each of the multiple candidate processing results, at least one target processing result for the task is determined from the multiple candidate processing results. Based at least on at least one target processing result for each of the set of tasks, a response to the user request is provided.

[0043] In this way, a set of tasks matching user needs can be identified. For each task in the set, multiple candidate processing results can be determined, and from these multiple candidate processing results, at least one target processing result with a higher quality score for the task can be selected. A response to the user's request can be provided based on at least one target processing result for each task in the set. This helps improve the accuracy and effectiveness of processing tasks expected by the user during human-computer interaction.

[0044] The following description will continue with reference to the accompanying drawings, which will provide some exemplary embodiments of this disclosure.

[0045] Figure 2 Example 200 for task processing according to some embodiments of the present disclosure is shown. Reference will be made to... Figure 1 Let's use environment 100 to describe example 200.

[0046] Example 200 includes an architecture 205 for task processing. For ease of description, it is illustrated by assuming that architecture 205 is implemented on server device 120. It should be noted that if architecture 205 is implemented on client device 110, some operations described with reference to client device 110 may require the assistance of server device 120. It should also be noted that the operations performed by client device 110 may specifically be performed by relevant applications installed on client device 110. As shown in the figure, architecture 205 involves a task planning agent 220, a task processing agent 230, a decision evaluation agent 240, a reflection agent 250, and a presentation agent 260.

[0047] In some embodiments, client device 110 may receive user requests from a target user. For example, client device 110 receives user requests during an interaction between the target user (e.g., user 140) and a digital assistant. For instance, client device 110 may receive user input from the target user within the interface of application 112 and / or digital assistant 114, and identify the received user input as a user request. The user request may, for example, be presented in the interface as a session message from the target user.

[0048] Client device 110 can, in response to receiving a user request, provide the user request to server device 120. Task planning agent 220 in server device 120 can, in response to receiving a user request, determine a set of tasks that match the user needs indicated by the user request based on context information associated with the user request of the target user.

[0049] Task planning agent 220 can determine user needs based on user requests in any suitable manner. For example, task planning agent 220 can perform semantic analysis on user requests to determine user needs. Alternatively, task planning agent 220 can determine user needs based on predefined rules. Furthermore, task planning agent 220 can also use a trained machine learning model to determine user needs based on user requests. Contextual information includes at least the user request.

[0050] In some embodiments, the task planning agent 220 may also obtain reference information matching the user request from an information table 210 associated with the target user. In this case, the context information includes the user request and the reference information. This information table 210 may be created or updated by the server device 120 based at least on the information collected by the information collection component 201. The information collection component 201 may, for example, be deployed at the client device 110. The client device 110 may utilize the information collection component 201 to obtain information associated with the target user. The information collection component 201 may include any suitable components, including but not limited to a positioning component for obtaining location information, a camera for acquiring images / videos, a microphone for acquiring audio, etc. It is understood that in the embodiments of this disclosure, the acquired data, the data processing and storage methods, etc., should all obtain prior authorization from the user and other rights holders associated with the user, and should comply with the relevant laws and regulations and the agreements between rights holders.

[0051] In some embodiments, the client device 110 may periodically use the information collection component 201 to collect information only when authorization is obtained. For example, the client device 110 may provide an authorization request interface to the user before the first information collection, and receive the user's permission configuration via this interface. The authorization request interface may include, for example, permission prompt information and at least one operation control. The at least one operation control may include an authorization confirmation control. The permission prompt information may include, for example, the text "Allow periodic collection of video, audio, and other information?", and the client device 110 may determine that authorization from the user has been obtained in response to detecting a trigger on the authorization confirmation control.

[0052] In some embodiments, information that does not require active user input can be identified as implicit information 204 associated with the target user. Implicit information 204 may include, for example, audio, images, video, and location information of the target user periodically collected by the information collection component 201, as well as audio, video, images, and weather information of the target user's environment. In some embodiments, implicit information 204 may also include historical records of the target user associated with tasks, such as the target user's historical requests and historical responses to historical user requests. In some embodiments, implicit information 204 may also include device information of the client device (i.e., client device 110) corresponding to the target user, including but not limited to real-time battery level, device model, system version, and system memory.

[0053] In some embodiments, client device 110 may also receive user requests. User requests can be requests in any suitable form (e.g., text, audio, etc.). For example, client device 110 may present an interactive interface with a target user. The interactive interface may provide input boxes, at least one option, etc. Client device 110 may receive user requests from the target user via input boxes. Client device 110 may also determine that a user request has been received in response to receiving a selection operation by the target user on at least one option (e.g., the text corresponding to the option may be determined as the user request). Alternatively or additionally, in some embodiments, client device 110 may also acquire audio from the target user via a speaker and determine this audio as the user request. It is understood that if the user request is a non-text type, it may be converted to a text type (e.g., audio may be converted to text). The acquired user request may be determined as display information 202 for the target user.

[0054] Client device 110 can send acquired information (including display information 202 and / or implicit information 204) to server device 120 via a communication connection with server device 120. Server device 120 can create or update information table 210 associated with the target user based on the acquired information. In some embodiments, server device 120 can analyze the acquired information to determine multiple information items associated with the target user's status information and / or the target user's physical environment. Server device 120 can analyze the information in any suitable manner; for example, server device 120 can analyze the information using a trained machine learning model. It is understood that this is merely an example, and information table 210 may actually include any suitable content. Server device 120 can store the analysis results (i.e., multiple information items) in information table 210. For example, information table 210 may include weather-related information items, such as sunny weather. The information stored in information table 210 may be referred to as reference information.

[0055] The task planning agent 220 can retrieve reference information matching the user request from the information table 210 in any appropriate manner. For example, the task planning agent 220 can directly search the information table 210 based on the user request to retrieve reference information matching the user request. Alternatively, the task planning agent 220 can retrieve reference information matching the user request from the information table 210 based on predetermined matching rules.

[0056] Regarding the specific method of determining a set of tasks that match the user's needs indicated by the user request based on contextual information, in some embodiments, the task planning agent 220 may utilize a trained machine learning model to determine a set of tasks. Specifically, the task planning agent 220 uses a prompt word template for determining a set of tasks to generate prompt word input for determining a set of tasks based on the user's needs and contextual information. This prompt word template may include pre-configured prompt word information that instructs the machine learning model to determine a set of tasks based on the contextual information and the user's needs. This prompt word template may also include one or more of the following: configuration information for the machine learning model, specification of at least one function that the machine learning model can invoke, definition of the workflow for at least one function, at least one example (each example may indicate a sample user need and a set of sample tasks corresponding to the sample need), and requirements for the model output of the machine learning model. Referring to Table 1, Table 1 shows an example of a prompt word template for determining a set of tasks:

[0057] Table 1

[0058]

[0059]

[0060] As shown in Table 1, the prompt word template used to determine a set of tasks includes prompt word information such as setting information, function 1, function 2, workflow, and examples (including example input and example task list). The five tasks included in the example task list constitute a set of sample tasks. The sample task planning agent 220 can obtain prompt word input for the machine learning model by filling the "Input" field at the bottom of the prompt word template shown in Table 1 with context information associated with the current user request. The sample task planning agent 220 can provide this prompt word input to the trained machine learning model and obtain the model output of the machine learning model, which indicates a set of tasks that match the user's needs. The model output can be, for example, a task list, which includes a set of tasks that match the user's needs.

[0061] A set of tasks is provided to task processing agent 230. Task processing agent 230 can determine multiple candidate processing results for each task in the set of tasks based on context information, according to multiple task processing modes (e.g., task processing mode 232-1, task processing mode 232-2, ..., task processing mode 232-N, where N is any suitable positive integer). Different task processing modes define different task processing strategies. For example, the multiple task processing modes may include six task processing modes: result-based mode, trial-and-error mode (also known as exploration mode), general strategy mode, experience-based mode, conformity mode, and attribute-based mode, each defining a different task processing strategy.

[0062] Task processing agent 230 can process a group of tasks together and determine multiple candidate processing results for each task in the group at once. Alternatively or additionally, task processing agent 230 can also process each task in the group sequentially, determining multiple candidate processing results for only one task at a time. For ease of description, the following example illustrates determining multiple candidate processing results for only one task at a time. It can be understood that the number of multiple candidate processing results for each task is related to the number of multiple task processing modes. In some embodiments, for each task, task processing agent 230 can determine only one processing result for that task according to a task processing mode. Of course, in some examples, task processing agent 230 can also determine multiple processing results for a task according to a task processing mode.

[0063] Regarding the specific method of determining multiple candidate processing results for each task according to multiple task processing modes, in some embodiments, for each task processing mode among multiple task processing modes, the task processing agent 230 can use the first prompt word template corresponding to the task processing mode to generate a first prompt word input for the first machine learning model based on context information and task.

[0064] Similarly, the first prompt word template may include pre-configured prompt word information, which instructs the first machine learning model to process the input task based on the corresponding task processing strategy. The prompt word information included in the first prompt word template may also include one or more of the following: configuration information for the first machine learning model, specification of at least one function that the first machine learning model can invoke, definition of the workflow for at least one function, at least one example (each example indicating a sample task and the task processing result based on the corresponding task processing strategy), and requirements for the model output of the first machine learning model.

[0065] In some embodiments, the prompt word information in the first prompt word template corresponding to each of the multiple task processing modes is different. For example, the prompt word information in the first prompt word template corresponding to task processing mode A includes setting information A for the first machine learning model, and the prompt word information in the first prompt word template corresponding to task processing mode B includes setting information B for the first machine learning model, wherein setting information A is different from setting information B.

[0066] Taking six task processing modes—result-based, trial-and-error, general strategy, experience-based, conformity, and attribute-based—as examples, Tables 2 to 7 output the first prompt word templates corresponding to these six task processing modes:

[0067] Table 2

[0068]

[0069]

[0070] As shown in Table 2, the prompt word template corresponding to the result pattern includes prompt word information such as setting information, function 1, function 2, workflow, and examples (including example input, example solution analysis, and example results and suggestions). Example input includes example tasks, and example results and suggestions show example processing results for the example tasks. Task processing agent 230 can obtain the first prompt word input for the first machine learning model based on the result pattern by filling the context information and task into the "Input" field at the bottom of the first prompt word template shown in Table 2. Task processing agent 230 can provide the first prompt word input to the first machine learning model and obtain the model output of the first machine learning model, which indicates the candidate processing results corresponding to the result pattern. The model output can be, for example, a result and suggestion.

[0071] Table 3

[0072]

[0073]

[0074]

[0075] As shown in Table 3, the prompt word template corresponding to the trial-and-error mode includes prompt word information such as setting information, functions 1-3, workflow, and examples (including example input, example solution analysis, and example results and suggestions). Example input includes example tasks, and example results and suggestions show example processing results for the example tasks. The task processing agent 230 can obtain the first prompt word input for the first machine learning model in the trial-and-error mode by filling the context information and task into the "Input" field at the bottom of the first prompt word template shown in Table 3. The task processing agent 230 can provide the first prompt word input to the first machine learning model and obtain the model output of the first machine learning model, which indicates the candidate processing results corresponding to the trial-and-error mode. The model output can, for example, be a result and suggestion.

[0076] Table 4

[0077]

[0078]

[0079] As shown in Table 4, the first prompt word template corresponding to the general strategy pattern includes prompt word information such as setting information, functions 1-3, workflow, and examples (including example input, example solution analysis, and example results and suggestions). Example input includes example tasks, and example results and suggestions show example processing results for the example tasks. The task processing agent 230 can obtain the first prompt word input for the first machine learning model under the general strategy pattern by filling the context information and task into the "Input" field at the bottom of the first prompt word template shown in Table 4. The task processing agent 230 can provide the first prompt word input to the first machine learning model and obtain the model output of the first machine learning model, which indicates the candidate processing results corresponding to the general strategy pattern. The model output can be, for example, a result and suggestion.

[0080] Table 5

[0081]

[0082]

[0083] As shown in Table 5, the prompt word template corresponding to the experience mode includes prompt word information such as setting information, functions 1-5, workflow, and examples (including example input, example solution analysis, and example results and suggestions). Example input includes example tasks, and example results and suggestions show example processing results for the example tasks. The task processing agent 230 can obtain the first prompt word input for the first machine learning model based on the experience mode by filling the context information and task into the "Input" field at the bottom of the first prompt word template shown in Table 5. The task processing agent 230 can provide the first prompt word input to the first machine learning model and obtain the model output of the first machine learning model, which indicates the candidate processing results corresponding to the experience mode. The model output can be, for example, a result and suggestion.

[0084] Table 6

[0085]

[0086]

[0087]

[0088] As shown in Table 6, the first prompt word template corresponding to the conformity mode includes prompt word information such as setting information, functions 1-4, workflow, and examples (including example input, example solution analysis, and example results and suggestions). Example input includes example tasks, and example results and suggestions show example processing results for the example tasks. The task processing agent 230 can obtain the first prompt word input for the first machine learning model in the conformity mode by filling the context information and task into the "Input" field at the bottom of the first prompt word template shown in Table 6. The task processing agent 230 can provide the first prompt word input to the first machine learning model and obtain the model output of the first machine learning model, which indicates the candidate processing results corresponding to the conformity mode. The model output can be, for example, a result and suggestion.

[0089] Table 7

[0090]

[0091]

[0092] As shown in Table 7, the prompt word template corresponding to the attribute pattern includes prompt word information such as setting information, functions 1-3, workflow, and examples (including example input, example solution analysis, and example results and suggestions). Example input includes example tasks, and example results and suggestions show example processing results for the example tasks. Task processing agent 230 can obtain the first prompt word input for the first machine learning model based on the attribute pattern by filling the context information and task into the "Input" field at the bottom of the first prompt word template shown in Table 7. Task processing agent 230 can provide the first prompt word input to the first machine learning model and obtain the model output of the first machine learning model, which indicates the candidate processing results corresponding to the attribute pattern. The model output can be, for example, a result and suggestion.

[0093] It is understood that Tables 1 and 2 through 7 above only output examples of prompt word templates (or first prompt word templates). For the sake of brevity, some content of the prompt word information in each table of this disclosure has been omitted. It is understood that locations A, B, C, etc., in each table can be any location, and locations with the same identifier in different tables can also be different locations (for example, location A in Table 2 and location A in Table 3 can be different locations). It is understood that XXX in each table represents any appropriate text or numerical value.

[0094] It is understood that in the embodiments of this disclosure, the machine learning models that receive different prompt word inputs can be the same model (e.g., all can be LM) or different models (e.g., models with different functions), and this disclosure does not limit this.

[0095] Task processing agent 230 can obtain multiple candidate processing results for a task by providing multiple first prompt words corresponding to the aforementioned multiple task processing modes to the first prompt word model. Task processing agent 230 can provide multiple candidate processing results corresponding to each of the multiple tasks to decision evaluation agent 240. For each task, decision evaluation agent 240 can determine the quality score of each of the multiple candidate processing results for the task.

[0096] Regarding the specific method for determining the quality scores of each of the multiple candidate processing results, in some embodiments, for each candidate processing result, the decision evaluation agent 240 can determine multiple evaluation scores of the candidate processing result from multiple evaluation dimensions, and determine the weights corresponding to each of the multiple evaluation dimensions based on user needs. The decision evaluation agent 240 can determine the evaluation scores in any suitable manner. For example, the decision evaluation agent 240 can determine the evaluation scores based on predetermined rules or algorithms. As another example, the decision evaluation agent 240 can also use a second machine learning model to determine the evaluation scores. Specifically, the decision evaluation agent 240 can use a second prompt word template for determining the evaluation scores to generate a second prompt word input for the second machine learning model based on the multiple candidate processing results.

[0097] This second prompt word template includes at least pre-configured prompt word information, which instructs the second machine learning model to score multiple candidate processing results from multiple evaluation dimensions. Similarly, the prompt word information included in the second prompt word template may also include one or more of the following: configuration information for the second machine learning model, specification of at least one function that the second machine learning model can invoke, definition of the workflow for at least one function, at least one example (each example may indicate the sample processing result and the sample evaluation score corresponding to the sample processing result), and requirements for the model output of the second machine learning model.

[0098] The decision evaluation agent 240 can obtain the model output of the second machine learning model by providing the second prompt word as input. This model output can indicate multiple evaluation scores for the candidate processing result across multiple evaluation dimensions. The weight corresponding to each evaluation dimension is positively correlated with the target user's emphasis on that evaluation dimension. That is, the more the target user values ​​that evaluation dimension, the greater its weight. Multiple evaluation dimensions may include, for example, the consistency between the candidate processing result and the user's needs, the time required to implement the candidate processing result, and the resources required to implement the candidate processing result.

[0099] The decision evaluation agent 240 can then determine the quality score of the candidate processing result based on multiple evaluation scores and the weights corresponding to each of the multiple evaluation dimensions. For example, if the evaluation dimension of the consistency between the candidate processing result and the user's needs corresponds to evaluation score A and weight A, the evaluation dimension of the time required to implement the candidate processing result corresponds to evaluation score B and weight B, and the evaluation dimension of the resources required to implement the candidate processing result corresponds to evaluation score C and weight C, then the quality score of the candidate processing result A is evaluation score A * weight A + evaluation score B * weight B + evaluation score C * weight C.

[0100] In some embodiments, the decision evaluation agent 240 may also directly utilize a machine learning model to determine the quality scores of each of the multiple candidate processing results. For example, the decision evaluation agent 240 may directly utilize a third prompt word template for determining quality scores to generate a third prompt word input for a third machine learning model based at least on the multiple candidate processing results. The third prompt word template includes at least pre-configured prompt word information that instructs the third machine learning model to score the multiple candidate processing results separately. The decision evaluation agent 240 can obtain the model output of the third machine learning model by providing the third prompt word input, the model output indicating the quality scores of each of the multiple candidate processing results.

[0101] The decision evaluation agent 240 can provide quality scores of each of the multiple candidate processing results to the reflection agent 250. The reflection agent 250 can determine at least one target processing result for the task from the multiple candidate processing results, based at least on the quality scores of each candidate processing result. In some embodiments, the reflection agent 250 can determine at least one target processing result for the task from the multiple candidate processing results based on the quality scores of each candidate processing result, according to predetermined rules. Alternatively or additionally, in some embodiments, the reflection agent 250 can also obtain quality screening conditions. The reflection agent 250 can determine at least one target processing result by determining whether the quality scores of the candidate processing results meet the quality screening conditions. In some embodiments, the quality screening conditions can also be determined by the reflection agent 250 based on context information, which is not limited in this disclosure.

[0102] Quality screening criteria can indicate that a score threshold is exceeded. Reflection agent 250 can, in response to at least one quality score among multiple candidate processing results exceeding the score threshold, determine at least one candidate processing result corresponding to that at least one quality score as at least one target processing result for the task. Quality screening criteria can also indicate that a score ranking position is satisfied. In some embodiments, in addition to determining the individual quality scores of the multiple candidate processing results, decision evaluation agent 240 can also rank the multiple candidate processing results based on their individual quality scores (e.g., in descending order). Decision evaluation agent 240 can also provide the ranking results to reflection agent 250.

[0103] Of course, in some embodiments, the decision evaluation agent 240 may only determine the quality scores of each of the multiple candidate processing results, and the reflection agent 250 may, in response to receiving the quality scores of each of the multiple candidate processing results, sort the multiple candidate processing results based on their respective quality scores (e.g., in descending order). The reflection agent 250 may, in response to at least one quality score among the multiple candidate processing results reaching a ranking position in the ranking results, determine at least one candidate processing result corresponding to this at least one quality score as at least one target processing result for the task. For example, if the ranking position is 2, the reflection agent 250 may determine the two candidate processing results corresponding to the two quality scores ranking first and second in the ranking results as two target processing results for the task. The quality screening condition may also simultaneously indicate exceeding a score threshold and satisfying the ranking position. It is understood that the quality screening condition may also indicate other appropriate content, which is not limited in this disclosure.

[0104] In some embodiments, the reflection agent 250 may, in response to the fact that the quality scores of multiple candidate processing results do not meet the quality screening criteria, feed back the tasks corresponding to these multiple candidate processing results to the task planning agent 220, so as to instruct the task planning agent 220 to redetermine the task to be executed based on context information. The task processing agent 230 may execute the redetermined task according to multiple task processing modes to obtain multiple candidate processing results corresponding to the task.

[0105] In some embodiments, the reflexive agent 250 may also directly utilize a machine learning model to determine at least one target processing result for a task from multiple candidate processing results. For example, the reflexive agent 250 may directly utilize a fourth cue word template for determining at least one target processing result for a task from multiple candidate processing results, at least based on the quality scores of each of the multiple candidate processing results, to generate a fourth cue word input for a fourth machine learning model. The fourth cue word template includes at least pre-configured cue word information that instructs the fourth machine learning model to determine at least one target processing result for a task from multiple candidate processing results. The reflexive agent 250 can obtain the model output of the fourth machine learning model by providing the fourth cue word input, the model output indicating at least one target processing result.

[0106] In some embodiments, if the quality scores of multiple candidate processing results corresponding to at least one task in a set of tasks do not meet the quality screening criteria (i.e., the at least one task needs to be re-determined), the rethinking agent 250 may store other tasks besides the at least one task in the set of tasks (i.e., save tasks that do not need to be re-determined), and in response to at least one processing result among the multiple candidate processing results corresponding to the at least one re-determined task meeting the quality screening criteria (i.e., the at least one re-determined task does not need to be re-determined), provide the at least one re-determined task and the other previously stored tasks to the presentation agent 260. Alternatively or additionally, in some embodiments, the rethinking agent 250 may first provide the tasks that do not need to be re-determined to the presentation agent 260, and then subsequently provide the at least one re-determined task to the presentation agent 260.

[0107] The presentation agent 260 can provide a response to a user request based at least on at least one target processing result of each of a set of tasks. In some embodiments, the presentation agent 260 can provide a response to a target user by presenting a response delivery interface. Specifically, the presentation agent 260 can determine the presentation style for providing the response based at least on at least one target processing result. The presentation agent 260 can also determine the presentation style for providing the response, for example, based on contextual information and at least one target processing result.

[0108] In some embodiments, the presentation agent 260 may determine the presentation style based on predetermined rules. Alternatively or additionally, in some embodiments, the presentation agent 260 may utilize a machine learning model to determine the presentation style. The presentation style indicates a set of interface elements used to display at least one target processing result. The set of interface elements may include any suitable interface elements such as text, images, and icons, and may include static or dynamic interface elements. The presentation agent 260 may provide a response to a target user by presenting a set of interface elements in a response-providing interface.

[0109] In some embodiments, the presentation agent 260 may provide a response to a target user by playing audio. The presentation agent 260 may determine the audio, for example, based on contextual information and at least one target processing result, and provide a response to the target user by playing the audio. It is understood that the presentation agent 260 may provide a response in any suitable manner, such as vibration, flashing lights, etc., and this disclosure is not limited thereto.

[0110] In some embodiments, server device 120 may provide a response and the manner in which the response is provided to client device 110, instructing client device 110 to provide the response in the manner specified in the response. Client device 110 may also receive interactions from the target user regarding the response. For example, if a task corresponds to two target processing results, the response may include the two processing results for that task. These two target processing results may, for example, be presented as two options on the response providing interface. Client device 110 may receive a selection operation from the target user for either of these two options. Alternatively or additionally, in some embodiments, client device 110 may also receive interactive operations from the target user such as comments, likes, ratings, etc., regarding the response. Client device 110 may, for example, feed back the received interactions with the response to server device 120.

[0111] Server-side device 120 may update information table 210 based, for example, on user requests, user needs, a set of tasks matching the user needs, at least one target processing result for each of the tasks, responses to user requests, and interactions with those responses. Server-side device 120 may update information table 210 in any suitable manner; for example, it may use a trained machine learning model. Thus, by repeatedly executing the aforementioned task processing and updating information table 210, the matching degree between information table 210 and the target user can be increased. This facilitates more accurate responses to subsequent user requests from the target user.

[0112] In summary, according to the embodiments of this disclosure, a set of tasks matching user needs can be determined. For each task in the set of tasks, multiple candidate processing results can be determined, and from these multiple candidate processing results, at least one target processing result with a higher quality score for the task can be determined. A response to a user request can be provided based at least on at least one target processing result for each of the tasks in the set.

[0113] Figure 3 A flowchart of a task processing method 300 according to some embodiments of the present disclosure is shown. Method 300 can be implemented at server device 120 or client device 110. Reference will be made to... Figure 1 The environment 100 is described in the method 300 implemented on the server device 120.

[0114] In box 310, server device 120 determines a set of tasks that match the user needs indicated by the user request based on context information associated with the user request of the target user, the context information including at least the user request.

[0115] In box 320, server device 120 determines multiple candidate processing results for each task in a set of tasks based on context information according to multiple task processing modes. Different task processing modes in the multiple task processing modes define different task processing strategies, determine the quality scores of each of the multiple candidate processing results of the task, and determine at least one target processing result for the task from the multiple candidate processing results based on the quality scores of each of the multiple candidate processing results.

[0116] In box 330, server device 120 provides a response to a user request based on at least one target processing result of each of a set of tasks.

[0117] In some embodiments, method 300 further includes: in response to receiving a user request, obtaining reference information matching the user request from an information table associated with a target user, the information table including multiple information items associated with at least one of the following: the target user's status information, the physical environment in which the user is located, wherein the context information includes user input and reference information.

[0118] In some embodiments, determining multiple candidate processing results for a task based on context information according to multiple task processing modes includes: for each task processing mode, generating a first prompt word input for a first machine learning model based on context information and the task using a first prompt word template corresponding to the task processing mode, wherein the first prompt word template includes at least pre-configured prompt word information, the prompt word information instructing the first machine learning model to process the input task based on the corresponding task processing strategy; and obtaining the model output of the first machine learning model by providing the first prompt word input to the first machine learning model, wherein the model output indicates the candidate processing result corresponding to the task processing mode.

[0119] In some embodiments, the prompt information includes at least one of the following: configuration information for the first machine learning model, specification of at least one function that the first machine learning model can call, definition of the workflow for at least one function, at least one example, each example indicating a sample task and the task processing result of processing the sample task based on the corresponding task processing strategy, and requirements for the model output of the first machine learning model.

[0120] In some embodiments, the prompt word information in the first prompt word template corresponding to each of the multiple task processing modes is different.

[0121] In some embodiments, determining the quality score of each of the multiple candidate processing results of a task includes: for each of the multiple candidate processing results, determining multiple evaluation scores of the candidate processing result from multiple evaluation dimensions, the multiple evaluation dimensions including: consistency between the candidate processing result and user requirements, time required to implement the candidate processing result, and resources required to implement the candidate processing result; determining the weights corresponding to each of the multiple evaluation dimensions based on user requirements; and determining the quality score of the candidate processing result based on the multiple evaluation scores of the candidate processing result and the weights corresponding to each of the multiple evaluation dimensions.

[0122] In some embodiments, determining multiple evaluation scores for each of a plurality of candidate processing results from multiple evaluation dimensions includes: generating a second prompt word input for a second machine learning model based on the plurality of candidate processing results using a second prompt word template for determining the evaluation scores, the second prompt word template including at least pre-configured prompt word information, the prompt word information instructing the second machine learning model to score the plurality of candidate processing results from multiple evaluation dimensions respectively; and obtaining the model output of the second machine learning model by providing the second prompt word input to the second machine learning model, the model output indicating the evaluation scores of the plurality of candidate processing results respectively in multiple evaluation dimensions.

[0123] In some embodiments, determining at least one target processing result for a task from multiple candidate processing results based at least on their respective quality scores includes: if at least one quality score among the multiple candidate processing results satisfies a quality screening condition, determining at least one candidate processing result corresponding to at least one quality score as at least one target processing result for the task, wherein the quality screening condition indicates that the score threshold is exceeded or the score ranking position is met.

[0124] In some embodiments, method 300 further includes: if the quality scores of multiple candidate processing results do not meet the quality screening conditions, redetermining the task to be executed based on context information; and executing the redetermined task according to multiple task processing modes to obtain multiple candidate processing results corresponding to the task.

[0125] In some embodiments, providing a response to a user request includes: determining a presentation style for providing the response based at least on at least one target processing result, the presentation style indicating a set of interface elements for displaying at least one target processing result; and presenting the set of interface elements to provide the response.

[0126] Embodiments of this disclosure also provide corresponding apparatus for implementing the above methods or processes. Figure 4 An exemplary structural block diagram of an apparatus 400 for task processing according to some embodiments of the present disclosure is shown. The apparatus 400 may be implemented as or included in a server device 120 or a client device 110. Various modules / components in the apparatus 400 may be implemented by hardware, software, firmware, or any combination thereof.

[0127] like Figure 4 As shown, the apparatus 400 includes a task determination module 410, configured to determine a set of tasks matching the user needs indicated by the user request based on context information associated with the user request of the target user, wherein the context information includes at least the user request. The apparatus 400 also includes a result determination module 420, configured to, for each task in the set of tasks, determine multiple candidate processing results for the task based on context information according to multiple task processing modes, wherein different task processing modes define different task processing strategies, determine a quality score for each of the multiple candidate processing results, and determine at least one target processing result for the task from the multiple candidate processing results based at least on the quality scores of each of the multiple candidate processing results. The apparatus 400 also includes a response providing module 430, configured to provide a response to the user request based at least on at least one target processing result for each of the set of tasks.

[0128] In some embodiments, the apparatus 400 further includes a reference information acquisition module configured to, in response to receiving a user request, acquire reference information matching the user request from an information table associated with a target user, the information table including multiple information items associated with at least one of the following: the target user's status information, the physical environment in which the user is located, wherein the context information includes user input and reference information.

[0129] In some embodiments, the result determination module 420 is further configured to: for each of the multiple task processing modes, generate a first prompt word input for the first machine learning model based on context information and the task using a first prompt word template corresponding to the task processing mode, wherein the first prompt word template includes at least pre-configured prompt word information, the prompt word information instructing the first machine learning model to process the input task based on the corresponding task processing strategy; and obtain the model output of the first machine learning model by providing the first prompt word input to the first machine learning model, wherein the model output indicates the candidate processing result corresponding to the task processing mode.

[0130] In some embodiments, the prompt information includes at least one of the following: configuration information for the first machine learning model, specification of at least one function that the first machine learning model can call, definition of the workflow for at least one function, at least one example, each example indicating a sample task and the task processing result of processing the sample task based on the corresponding task processing strategy, and requirements for the model output of the first machine learning model.

[0131] In some embodiments, the prompt word information in the first prompt word template corresponding to each of the multiple task processing modes is different.

[0132] In some embodiments, the result determination module 420 is further configured to: for each of the multiple candidate processing results, determine multiple evaluation scores of the candidate processing result from multiple evaluation dimensions, the multiple evaluation dimensions including: consistency between the candidate processing result and user requirements, time required to implement the candidate processing result, and resources required to implement the candidate processing result; determine the weights corresponding to each of the multiple evaluation dimensions based on user requirements; and determine the quality score of the candidate processing result based on the multiple evaluation scores of the candidate processing result and the weights corresponding to each of the multiple evaluation dimensions.

[0133] In some embodiments, the result determination module 420 is further configured to: generate a second prompt word input for a second machine learning model based on multiple candidate processing results using a second prompt word template for determining evaluation scores, the second prompt word template including at least pre-configured prompt word information, the prompt word information instructing the second machine learning model to score the multiple candidate processing results from multiple evaluation dimensions respectively; and obtain the model output of the second machine learning model by providing the second prompt word input to the second machine learning model, the model output indicating the evaluation scores of the multiple candidate processing results in multiple evaluation dimensions respectively.

[0134] In some embodiments, the result determination module 420 is further configured to: if at least one quality score among the quality scores of the plurality of candidate processing results satisfies the quality screening condition, determine at least one candidate processing result corresponding to at least one quality score as at least one target processing result for the task, wherein the quality screening condition indicates that the score threshold is exceeded or the score ranking position is met.

[0135] In some embodiments, the apparatus 400 further includes: a task re-determination module, configured to re-determine the task to be executed based on context information if the quality scores of the multiple candidate processing results do not meet the quality screening conditions; and a task re-execution module, configured to execute the re-determined task according to multiple task processing modes to obtain multiple candidate processing results corresponding to the task.

[0136] In some embodiments, the response providing module 430 is further configured to: determine a presentation style for providing a response based at least on at least one target processing result, the presentation style indicating a set of interface elements for displaying at least one target processing result; and present a set of interface elements to provide a response.

[0137] The units and / or modules included in device 400 can be implemented in various ways, including software, hardware, firmware, or any combination thereof. In some embodiments, one or more units and / or modules can be implemented using software and / or firmware, such as machine-executable instructions stored on a storage medium. In addition to or as an alternative to machine-executable instructions, some or all of the units and / or modules in device 400 can be implemented at least partially by one or more hardware logic components. By way of example and not limitation, exemplary types of hardware logic components that can be used include field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-chips (SoCs), complex programmable logic devices (CPLDs), and so on.

[0138] It should be understood that one or more steps in the above methods can be performed by suitable electronic devices or combinations of electronic devices. Such electronic devices or combinations of electronic devices may include, for example, […]. Figure 1 The server device 120 or the client device 110.

[0139] Figure 5 A block diagram of an electronic device 500 in which one or more embodiments of the present disclosure may be implemented is shown. It should be understood that... Figure 5 The electronic device 500 shown is merely exemplary and should not be construed as limiting the functionality and scope of the embodiments described herein. Figure 5 The electronic device 500 shown can be used to achieve Figure 1 The server device 120 or the client device 110, or Figure 4 Device 400.

[0140] like Figure 5 As shown, electronic device 500 is in the form of a general-purpose electronic device. Components of electronic device 500 may include, but are not limited to, one or more processors or processing units 510, memory 520, storage device 530, one or more communication units 540, one or more input devices 550, and one or more output devices 560. Processing unit 510 may be a physical or virtual processor and is capable of performing various processes according to programs stored in memory 520. In a multiprocessor system, multiple processing units execute computer-executable instructions in parallel to improve the parallel processing capability of electronic device 500.

[0141] Electronic device 500 typically includes multiple computer storage media. Such media can be any accessible media that is accessible to electronic device 500, including but not limited to volatile and non-volatile media, removable and non-removable media. Memory 520 can be volatile memory (e.g., registers, cache, random access memory (RAM)), non-volatile memory (e.g., read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory), or some combination thereof. Storage device 530 can be removable or non-removable media and can include machine-readable media, such as flash drives, disks, or any other media that can be used to store information and / or data and can be accessed within electronic device 500.

[0142] Electronic device 500 may further include additional removable / non-removable, volatile / non-volatile storage media. Although not explicitly stated... Figure 5As shown, disk drives for reading from or writing to removable, non-volatile disks (e.g., "floppy disks") and optical disk drives for reading from or writing to removable, non-volatile optical disks can be provided. In these cases, each drive can be connected to a bus (not shown) via one or more data media interfaces. Memory 520 may include computer program product 525 having one or more program modules configured to perform various methods or actions of various embodiments of this disclosure.

[0143] Communication unit 540 enables communication with other electronic devices via a communication medium. Additionally, the functionality of components of electronic device 500 can be implemented using a single computing cluster or multiple computing machines capable of communicating via communication connections. Therefore, electronic device 500 can operate in a networked environment using logical connections to one or more other servers, network personal computers (PCs), or another network node.

[0144] Input device 550 can be one or more input devices, such as a mouse, keyboard, trackball, etc. Output device 560 can be one or more output devices, such as a monitor, speaker, printer, etc. Electronic device 500 can also communicate with one or more external devices (not shown) via communication unit 540 as needed. These external devices include storage devices, display devices, etc., and can communicate with one or more devices that enable user interaction with electronic device 500, or with any device that enables electronic device 500 to communicate with one or more other electronic devices (e.g., network card, modem, etc.). Such communication can be performed via input / output (I / O) interface (not shown).

[0145] According to an exemplary implementation of this disclosure, a computer-readable storage medium is provided that stores computer-executable instructions thereon, wherein the computer-executable instructions are executed by a processor to implement the methods described above. According to an exemplary implementation of this disclosure, a computer program product is also provided, which is tangibly stored on a non-transitory computer-readable medium and includes computer-executable instructions, which are executed by a processor to implement the methods described above.

[0146] Various aspects of this disclosure are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatuses, devices, and computer program products implemented according to this disclosure. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.

[0147] These computer-readable program instructions can be provided to a processing unit of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that, when executed by the processing unit of the computer or other programmable data processing apparatus, they create means for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium that causes a computer, programmable data processing apparatus, and / or other device to operate in a particular manner. Thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.

[0148] Computer-readable program instructions can be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions that execute on the computer, other programmable data processing apparatus, or other device to perform the functions / actions specified in one or more boxes of a flowchart and / or block diagram.

[0149] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of an instruction, which contains one or more executable instructions for implementing the specified logical function. In some, as newer, implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.

[0150] Various implementations of this disclosure have been described above. These descriptions are exemplary and not exhaustive, nor are they limited to the disclosed implementations. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described implementations. The terminology used herein is chosen to best explain the principles, practical applications, or improvements to technology in the market, or to enable others skilled in the art to understand the various implementations disclosed herein.

Claims

1. A method for task processing, comprising: determining, based on context information associated with a user request of a target user, a set of tasks that match a user demand indicated by the user request, the context information comprising at least the user request; for each task in the set of tasks, determining, based on the context information, a plurality of candidate processing results of the task respectively according to a plurality of task processing modes, different task processing modes in the plurality of task processing modes defining different task processing strategies, determining a quality score of each of the plurality of candidate processing results of the task, determining, based on at least the quality score of each of the plurality of candidate processing results, at least one target processing result for the task from the plurality of candidate processing results; and providing a response to the user request based on at least the at least one target processing result of each of the set of tasks. 2.The method of claim 1, further comprising: in response to receiving the user request, obtaining reference information matching the user request from an information table associated with the target user, the information table comprising a plurality of information items associated with at least one of: state information of the target user, a physical environment in which the target user is located, wherein the context information comprises the user request and the reference information. 3.The method of claim 1, wherein determining, based on the context information, a plurality of candidate processing results of the task according to a plurality of task processing modes comprises: for each task processing mode in the plurality of task processing modes, generating, based on the context information and the task, a first prompt input for a first machine learning model using a first prompt template corresponding to the task processing mode, the first prompt template comprising at least preconfigured prompt information indicating that the first machine learning model processes an input task based on a corresponding task processing strategy, and obtaining a model output of the first machine learning model by providing the first prompt input to the first machine learning model, the model output indicating a candidate processing result corresponding to the task processing mode. 4.The method of claim 3, wherein the prompt information comprises at least one of: setting information of the first machine learning model, a designation of at least one function that the first machine learning model can invoke, a definition of a workflow of the at least one function, at least one example, each example indicating a sample task and a task processing result of the sample task based on the corresponding task processing strategy, a requirement on a model output of the first machine learning model. 5.The method of claim 3, wherein the prompt information in the first prompt template corresponding to each of the plurality of task processing modes is different. 6.The method of claim 1, wherein determining a quality score of each of the plurality of candidate processing results of the task comprises: for each candidate processing result in the plurality of candidate processing results, ​ determine a plurality of evaluation scores of the candidate processing result from a plurality of evaluation dimensions, the plurality of evaluation dimensions comprising: consistency between the candidate processing result and the user demand, time required for the candidate processing result to be implemented, resources consumed for the candidate processing result to be implemented, determine respective weights of the plurality of evaluation dimensions corresponding to the user demand, and determine a quality score of the candidate processing result based on the plurality of evaluation scores of the candidate processing result and the respective weights of the plurality of evaluation dimensions.

7. The method of claim 6, wherein determining, for each candidate processing result of the plurality of candidate processing results, a plurality of evaluation scores of the candidate processing result from a plurality of evaluation dimensions comprises: generating, based on the plurality of candidate processing results, a second prompt input for a second machine learning model using a second prompt word template for determining evaluation scores, the second prompt word template comprising at least preconfigured prompt word information indicating that the second machine learning model scores the plurality of candidate processing results from the plurality of evaluation dimensions respectively; and obtaining a model output of the second machine learning model by providing the second prompt input to the second machine learning model, the model output indicating evaluation scores of the plurality of candidate processing results in the plurality of evaluation dimensions respectively.

8. The method of claim 1, wherein determining at least one target processing result for the task from the plurality of candidate processing results based on at least quality scores of the plurality of candidate processing results respectively comprises: if at least one quality score of the quality scores of the plurality of candidate processing results satisfies a quality screening condition, determining at least one candidate processing result corresponding to the at least one quality score as at least one target processing result for the task, the quality screening condition indicating exceeding a score threshold or satisfying a score ranking position.

9. The method of claim 8, further comprising: if none of the quality scores of the plurality of candidate processing results satisfies the quality screening condition, re-determining a task to be executed based on the context information; and executing the re-determined task in the plurality of task processing modes to obtain a plurality of candidate processing results corresponding to the task.

10. The method of claim 1, wherein providing a response to the user request comprises: determining, based on at least one target processing result, a presentation style for providing the response, the presentation style indicating a set of interface elements for presenting the at least one target processing result; and presenting the set of interface elements to provide the response.

11. An apparatus for task processing, comprising: a task determination module configured to determine, based on context information associated with a user request of a target user, a set of tasks matching a user demand indicated by the user request, the context information comprising at least the user request; ​ ​ a result determination module configured to, for each task in the set of tasks, determine, based on the context information, a plurality of candidate processing results for the task according to a plurality of task processing modes, different ones of the plurality of task processing modes defining different task processing strategies, determine a quality score for each of the plurality of candidate processing results for the task, and determine, based on at least the quality score for each of the plurality of candidate processing results, at least one target processing result for the task from the plurality of candidate processing results; and a response providing module configured to provide a response to the user request based on at least the at least one target processing result for each of the set of tasks.

12. An electronic device, comprising: at least one processing unit; and at least one memory coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, the instructions which, when executed by the at least one processing unit, cause the electronic device to perform the method according to any one of claims 1-10.

13. A computer-readable storage medium having stored thereon a computer program, the computer program being executable by a processor to implement the method according to any one of claims 1-10.

14. A computer program product comprising a computer program, wherein the computer program, when executed by a processor, implements the method according to any one of claims 1-10. ​