Information processing method and device based on large model, equipment and medium
By using large models of different sizes to process instruction intention identification and execution sequence sorting in user requests, the problem that traditional methods cannot accurately identify user needs when facing compound requests is solved, and accurate understanding and efficient execution of user needs are achieved.
Patent Information
- Application Number
- CN202510138475.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-07
- Publication Date
- 2025-05-30
AI Technical Summary
Traditional user instruction intention recognition methods cannot accurately identify and realize user needs when facing multitasking compound requests.
By sorting instruction intent recognition and execution sequences by large models of different sizes, the first large model is used to identify user requests, the second large model is used to determine the execution sequence of instruction intent, and the interface corresponding to multiple instruction intents is called sequentially based on the execution sequence.
Effectively split and process compound instructions in user requests, realize accurate understanding and efficient execution of user needs, significantly reduce the dependence on manual analysis and rule arrangement, and improve the stability and accuracy of task execution.
Smart Images

Figure CN120067222A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of artificial intelligence technologies, and in particular to technologies such as machine learning, deep learning, and large models. Specifically, it relates to a method for information processing based on a large model, an apparatus for information processing based on a large model, an electronic device, a computer-readable storage medium, and a computer program product. Background Art
[0002] Artificial intelligence is a discipline that studies how to make a computer simulate certain human thinking processes and intelligent behaviors (such as learning, reasoning, thinking, planning, etc.). It has both hardware-level technologies and software-level technologies. Artificial intelligence hardware technologies generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, and big data processing; artificial intelligence software technologies mainly include several major directions such as natural language processing technology, computer vision technology, speech recognition technology, and machine learning / deep learning, big data processing technology, and knowledge graph technology.
[0003] In recent years, the natural language understanding technology in the field of artificial intelligence has continuously made breakthroughs, and intent recognition has gradually become an important research direction. By parsing a user request (Query), an intent recognition model can accurately capture the user's needs or operation purposes, providing efficient support for downstream applications.
[0004] The methods described in this section are not necessarily methods that have been previously envisioned or adopted. Unless otherwise specified, no method described in this section should be considered prior art solely because it is included in this section. Similarly, unless otherwise specified, the problems mentioned in this section should not be considered to have been recognized in any prior art. Summary of the Invention
[0005] The present disclosure provides a method for information processing based on a large model, an apparatus for information processing based on a large model, an electronic device, a computer-readable storage medium, and a computer program product.
[0006] According to one aspect of the present disclosure, there is provided a method for information processing based on a large model. The method includes: performing intent recognition on a user request by using a first large model to obtain a plurality of instruction intents; determining an execution sequence of the plurality of instruction intents by using a second large model so that the execution sequence satisfies the dependency relationships among the plurality of instruction intents, wherein the number of parameters of the second large model is greater than that of the first large model; and sequentially invoking interfaces corresponding to the plurality of instruction intents based on the execution sequence.
[0007] According to another aspect of the present disclosure, there is provided an information processing apparatus based on a large model. The apparatus includes: a first recognition unit configured to recognize keyword entries in a target document in response to detecting a first browsing operation of a user on the target document; a first determination unit configured to determine relevant content of the keyword entries, where the relevant content at least includes context information of the keyword entries in the target document; a first invocation unit configured to invoke the large model to generate an entry explanation for the keyword entries based on the relevant content; and a first display unit configured to display the entry explanation in response to detecting a first interaction operation of the user on the keyword entries and in response to determining that the large model has completed generating the entry explanation.
[0008] According to another aspect of the present disclosure, there is provided an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and these instructions are executed by the at least one processor to enable the at least one processor to execute the above method.
[0009] According to another aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause a computer to execute the above method.
[0010] According to another aspect of the present disclosure, there is provided a computer program product including a computer program, wherein the computer program implements the above method when executed by a processor.
[0011] According to one or more embodiments of the present disclosure, by separately completing instruction intention recognition and parsing and execution sequence sorting by large models of different scales, the embodiments of the present disclosure can effectively split and process composite instructions in a user request, and achieve accurate understanding and efficient execution of user requirements.
[0012] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] The drawings exemplarily illustrate embodiments and form a part of the specification, and are used together with the written description of the specification to explain the exemplary embodiments of the embodiments. The illustrated embodiments are for illustrative purposes only and do not limit the scope of the claims. In all the drawings, the same reference numerals refer to similar but not necessarily identical elements.
[0014] Figure 1 A schematic diagram showing an exemplary system in which the various methods described herein can be implemented according to an embodiment of the present disclosure;
[0015] Figure 2 The figure shows a flowchart of a large model-based information processing method according to an embodiment of the present disclosure;
[0016] Figures 3A - 3C The figure shows a schematic diagram of a user operating on an interaction interface according to an exemplary embodiment of the present disclosure;
[0017] Figure 4 The figure shows a flowchart of a process of sequentially invoking interfaces corresponding to multiple instruction intents based on an execution sequence according to an exemplary embodiment of the present disclosure;
[0018] Figure 5 The figure shows a flowchart of a large model-based information processing method according to an embodiment of the present disclosure;
[0019] Figure 6 The figure shows a structural block diagram of a large model-based information processing apparatus according to an embodiment of the present disclosure; and
[0020] Figure 7 The figure shows a structural block diagram of an exemplary electronic device capable of implementing the embodiments of the present disclosure. Detailed implementation manners
[0021] The following makes an explanation of exemplary embodiments of the present disclosure with reference to the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding, which should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.
[0022] In the present disclosure, unless otherwise specified, the terms "first", "second", etc. are used to describe various elements and are not intended to limit the positional relationship, timing relationship, or importance relationship of these elements. Such terms are only used to distinguish one element from another. In some examples, the first element and the second element may refer to the same instance of the element, and in certain cases, based on the context description, they may also refer to different instances.
[0023] In the descriptions of various examples in the present disclosure, the terms used are only for the purpose of describing specific examples and are not intended to be limiting. Unless the context clearly indicates otherwise, if the number of elements is not specifically limited, the element may be one or more. In addition, the term "and / or" used in the present disclosure covers any one of the listed items and all possible combinations.
[0024] In the related art, traditional user instruction intent recognition methods cannot accurately recognize and implement user requirements when facing multi-task composite requests.
[0025] To solve the above problems, by separately completing the instruction intention recognition and parsing and the execution sequence sorting with large models of different scales, the embodiments of the present disclosure can effectively split and process the composite instructions in the user request, realizing the accurate understanding and efficient execution of the user's needs.
[0026] The embodiments of the present disclosure will be described in detail below with reference to the accompanying drawings.
[0027] Figure 1 FIG. shows a schematic diagram of an exemplary system 100 in which the various methods and apparatuses described herein can be implemented according to an embodiment of the present disclosure. Referring Figure 1 to, the system 100 includes one or more client devices 101, 102, 103, 104, 105, and 106, a server 120, and one or more communication networks 110 that couple the one or more client devices to the server 120. The client devices 101, 102, 103, 104, 105, and 106 can be configured to execute one or more applications.
[0028] In an embodiment of the present disclosure, the server 120 can run one or more services or software applications that enable the execution of the methods of the present disclosure.
[0029] In certain embodiments, the server 120 can also provide other services or software applications, and these services or software applications can include non-virtual environments and virtual environments. In certain embodiments, these services can be provided as web-based services or cloud services, for example, provided to users of the client devices 101, 102, 103, 104, 105, and / or 106 under a software as a service (SaaS) model.
[0030] In Figure 1 the configuration shown, the server 120 can include one or more components that implement the functions performed by the server 120. These components can include software components, hardware components, or a combination thereof that can be executed by one or more processors. Users operating the client devices 101, 102, 103, 104, 105, and / or 106 can sequentially utilize one or more client applications to interact with the server 120 to utilize the services provided by these components. It should be understood that various different system configurations are possible, which can be different from the system 100. Therefore, Figure 1 is an example of a system for implementing the various methods described herein and is not intended to be limiting.
[0031] Users can use client devices 101, 102, 103, 104, 105, and / or 106 for human-computer interaction. The client devices can provide an interface that enables the users of the client devices to interact with the client devices. The client devices can also output information to the users via this interface. Although Figure 1 only six client devices are depicted, those skilled in the art will be able to understand that the present disclosure can support any number of client devices.
[0032] Client devices 101, 102, 103, 104, 105, and / or 106 can include various types of computer devices, such as portable handheld devices, general-purpose computers (such as personal computers and laptop computers), workstation computers, wearable devices, smart screen devices, self-service terminal devices, service robots, gaming systems, thin clients, various messaging devices, sensors, or other sensing devices, etc. These computer devices can run various types and versions of software applications and operating systems, such as MICROSOFT Windows, APPLE iOS, UNIX-like operating systems, Linux, or Linux-like operating systems (such as GOOGLE Chrome OS); or include various mobile operating systems, such as MICROSOFT WindowsMobile OS, iOS, Windows Phone, Android. Portable handheld devices can include cellular phones, smartphones, tablets, personal digital assistants (PDAs), etc. Wearable devices can include head-mounted displays (such as smart glasses) and other devices. Gaming systems can include various handheld gaming devices, Internet-enabled gaming devices, etc. The client devices are capable of executing various different applications, such as various Internet-related applications, communication applications (such as email applications), short message service (SMS) applications, and can use various communication protocols.
[0033] Network 110 can be any type of network well-known to those skilled in the art, which can support data communication using any one of a variety of available protocols (including but not limited to TCP / IP, SNA, IPX, etc.). By way of example only, one or more networks 110 can be a local area network (LAN), an Ethernet-based network, token ring, wide area network (WAN), the Internet, virtual network, virtual private network (VPN), intranet, extranet, blockchain network, public switched telephone network (PSTN), infrared network, wireless network (such as Bluetooth, WIFI), and / or any combination of these and / or other networks.
[0034] Server 120 may include one or more general-purpose computers, dedicated server computers (such as PC (Personal Computer) servers, UNIX servers, midrange servers), blade servers, mainframes, server clusters, or any other suitable arrangement and / or combination. Server 120 may include one or more virtual machines running a virtual operating system, or other computing architectures involving virtualization (such as one or more flexible pools of logical storage devices that can be virtualized to maintain virtual storage devices of the server). In various embodiments, server 120 may run one or more services or software applications that provide the functions described below.
[0035] The computing units in server 120 may run one or more operating systems including any of the above operating systems as well as any commercially available server operating systems. Server 120 may also run any one of a variety of additional server applications and / or middleware applications, including HTTP servers, FTP servers, CGI servers, JAVA servers, database servers, and the like.
[0036] In some embodiments, server 120 may include one or more applications to analyze and combine data feeds and / or event updates received from users of client devices 101, 102, 103, 104, 105, and / or 106. Server 120 may also include one or more applications to display data feeds and / or real-time events via one or more display devices of client devices 101, 102, 103, 104, 105, and / or 106.
[0037] In some embodiments, server 120 may be a server of a distributed system, or a server incorporating a blockchain. Server 120 may also be a cloud server, or an intelligent cloud computing server or intelligent cloud host with artificial intelligence technology. A cloud server is a host product in the cloud computing service system, which solves the defects of difficult management and weak business scalability existing in traditional physical hosts and virtual private server (VPS, Virtual Private Server) services.
[0038] System 100 may also include one or more databases 130. In certain embodiments, these databases can be used to store data and other information. For example, one or more of databases 130 can be used to store information such as audio files and video files. Databases 130 can reside in various locations. For example, the database used by server 120 can be local to server 120, or can be remote from server 120 and can communicate with server 120 via a network-based or dedicated connection. Databases 130 can be of different types. In certain embodiments, the database used by server 120 can be, for example, a relational database. One or more of these databases can store, update, and retrieve data to and from the database in response to commands.
[0039] In certain embodiments, one or more of databases 130 can also be used by an application to store application data. The database used by the application can be a different type of database, such as a key-value store, an object store, or a conventional store supported by a file system.
[0040] Figure 1 System 100 can be configured and operated in various ways to enable the application of the various methods and apparatuses described in this disclosure.
[0041] According to one aspect of the present disclosure, an information processing method based on a large model is provided. As Figure 2 shown, the method 200 includes: step S201, using a first large model to perform intent recognition on a user request to obtain multiple instruction intents; step S202, using a second large model to determine an execution sequence of the multiple instruction intents so that the execution sequence satisfies the dependency relationships between the multiple instruction intents, wherein the number of parameters of the second large model is greater than that of the first large model; and step S203, based on the execution sequence, sequentially invoking interfaces corresponding to the multiple instruction intents.
[0042] Thus, by separately completing instruction intent recognition and execution sequence sorting with large models of different scales, the embodiments of the present disclosure can effectively split and process composite instructions in a user request, achieving accurate understanding and efficient execution of user requirements.
[0043] The method of the present disclosure can be used for product forms based on dialogue interaction (e.g., chatbot). In addition, the method of the present disclosure can be used to perform various tasks such as understanding, editing, continuing writing, and generating on documents, files, texts, images, or various different types of objects.
[0044] Before step S201, a user request can be obtained. In some embodiments, the user request can be a text-based request (Query) input by the user, or a request triggered by the user's operations such as clicking in the interaction interface, which is not limited herein. The user request can be a multi-task composite request, that is, it implies multiple tasks to be completed. These tasks may be parallel or may have a dependency relationship. By performing intent recognition on the multi-task composite request, these tasks, that is, multiple instruction intents, can be disassembled therefrom.
[0045] In an exemplary embodiment, Figures 3A - 3C FIG. shows a schematic diagram of a user operating in an interaction interface according to an exemplary embodiment of the present disclosure. As Figure 3A shown, the user can perform a click operation on the file 310 uploaded by the user in the canvas 300 (interaction interface), and then can select from among a plurality of candidate instructions 320 and 330 that are pop-up displayed. After selecting an instruction, secondary editing can also be performed, as Figure 3B shown. The candidate instructions can include a marking instruction 320 and a generation instruction 330. The marking instruction will not be immediately recognized and executed after the user's operation, but will be overall recognized only when the subsequent generation instruction is triggered. As Figure 3C shown, the generation instruction can further include a generation type and a search and Q&A type, that is, the user expresses the type of content to be generated through a click operation or a request (Query) input, or directly triggers the call of a large model such as AI search and Q&A.
[0046] At this time, the backend will trigger the recognition and parameter extraction of the current generative instruction, and call different downstream interfaces according to the recognition result. At the same time, it will synchronously trigger the intent recognition and execution of the marking instructions for all the circled elements on the canvas. For different elements, their instruction intents are concurrently recognized, and for a single element, all its intents are recognized and returned and executed in sequence. Finally, the caller obtains the instruction execution results of each element from the intent side, and finally passes them all to the operator of the downstream policy side for acceptance.
[0047] The large model (or deep learning large model) described in the present disclosure can be a large language model. The deep learning large model has an end-to-end characteristic and can directly generate reply data based on the user's input data without relying on functional components or other inputs outside the deep learning large model. In other words, the deep learning large model itself has a generation function. A large language model generally refers to a deep learning large model with billions or even hundreds of billions of parameters, which are usually trained on large-scale text data or other modal data. Large language models can be used for various natural language processing tasks, such as text generation, language translation, and question-and-answer systems.
[0048] Large deep learning models can, for example, adopt an N-layer Transformer network structure with an encoder and a decoder, or a Unified pre-trained Language Model (UniLM) network structure. It can be understood that large deep learning models can also be other neural network models based on the Transformer network structure, which are not limited herein. The inputs and outputs of large deep learning models are both composed of tokens. Each token can correspond to a single character, character, word, or special symbol. Large deep learning models can be trained using pre-training tasks and generation tasks to possess the above-mentioned generation capabilities.
[0049] In step S201, a first large model is used for intent recognition to parse multiple instruction intents from a user request.
[0050] In some embodiments, the multiple instruction intents can include two categories: generation instruction intents and marking instruction intents. Generation instruction intents can further include one or more of document generation, paper generation, image interpretation, search recommendation, summarization, question answering, or other instruction intents for generation; marking instruction intents can further include one or more of reference content, reference style, reference logical structure, typesetting, expansion, rewriting, style change, or other instruction intents for marking or editing.
[0051] In step S202, a second large model is used to sort the multiple instruction intents obtained in step S201 to obtain an execution sequence of the multiple instruction intents.
[0052] In some embodiments, the second large model can synthesize the dependencies and business logics between the instruction intents to produce a "list of instruction execution orders", that is, an execution sequence. In addition to including the execution order of the multiple instruction intents, this execution sequence also includes the execution results of other instruction intents required when each instruction intent is executed.
[0053] According to some embodiments, the first large model can be a lightweight large language model, and the second large model can be a large-scale large language model. To balance the model effect and inference performance, the present disclosure uses a lightweight first large model with a smaller parameter scale to perform a simpler intent recognition task. To ensure that the obtained execution sequence needs to satisfy the dependencies between the multiple instruction intents, the present disclosure uses a large-scale second large model with a larger parameter scale and stronger inference ability to perform the sorting task of the multiple instruction intents.
[0054] In some embodiments, the first large model may adopt the Qianfan ERNIE-Lite-8k large language model, and the second large model may adopt the Qianfan ERNIE-3.5-8k large language model. It can be understood that the first large model and the second large model may adopt other models that can meet the above task requirements, which are not limited herein.
[0055] In step S203, after obtaining the execution sequences of multiple instruction intents, the interfaces (such as application programming interface API) corresponding to these instruction intents can be determined, and these interfaces can be called in sequence based on the determined execution sequences to satisfy the user request.
[0056] According to some embodiments, each of the multiple instruction intents may include an intent identifier and execution parameters. While identifying the intent in the user request, the first large model can also identify the specific execution parameters of these intents. In an exemplary embodiment, after the first large model identifies a user request, an instruction intent "PPT generation" can be obtained, and three parameters "theme", "number of pages", and "language" can be parsed.
[0057] Figure 4 FIG. shows a flowchart of a process 400 for sequentially calling interfaces corresponding to multiple instruction intents based on an execution sequence according to an exemplary embodiment of the present disclosure. The process 400 can be used to implement the above step S203. As Figure 4 shown, the process 400 may include: step S401, for the target instruction intent to be executed currently, based on the intent identifier of the target instruction intent, determine the target interface to be called; step S402, generate target instruction code based on the execution parameters of the target instruction intent; and step S403, call the target interface based on the target instruction code.
[0058] Thus, by decomposing the user's vague multi-task composite request based on natural language into clear instruction intents and corresponding parameters, the system can correctly call the interfaces to process each sub-task item by item and execute accurately. This method not only significantly reduces the dependence on manual parsing and rule arrangement in complex interaction scenarios, but also greatly improves the stability and accuracy of task execution. With the process of automatic decomposition and ordered execution, the system can quickly respond to and complete the diverse needs of users, enhancing efficiency and flexibility at the same time.
[0059] In some embodiments, each instruction intent may have a corresponding intent identifier, and a mapping relationship can be established between the intent identifier and the interface, so that the corresponding interface can be quickly found after the intent is recognized to achieve the corresponding target instruction intent.
[0060] In step S401, after the execution parameters are parsed, corresponding instruction execution codes can be generated. Since each instruction intent has a corresponding API interface, the system can map the execution parameters to the fields required by the API interface to form an executable call script, that is, an instruction code.
[0061] According to some embodiments, multiple instruction intents can adopt a JSON structure, and the JSON structure can include fields corresponding to intent identifiers and fields corresponding to execution parameters. The advantages of adopting the JSON structure are as follows: the format is general, the scalability is strong, and the integration and transmission on multiple terminals are simple; the hierarchical key-value pairs enable accurate and efficient parsing, reducing custom logic; it has good readability, facilitating debugging and maintenance; it cooperates more smoothly with the front-end and back-end docking, greatly improving the development efficiency and system compatibility.
[0062] In some embodiments, the intent identifier can be the ID or name of the instruction intent, or the ID or name of the interface corresponding to the instruction intent. The field corresponding to the execution parameter can use an object as the data format, which includes multiple sets of "name / value" pairs, that is, multiple specific execution parameters.
[0063] In some embodiments, the information processing method may further include: after determining the execution sequence, performing JSON correction on the JSON structures corresponding to multiple instruction intents.
[0064] Thus, by verifying and correcting the JSON structure corresponding to each instruction intent after determining the execution sequence, the correctness and consistency of the data structure in the downstream execution stage can be ensured, avoiding execution exceptions or errors caused by problems such as missing fields and type mismatches. This method helps to unify the data parsing process of each link, thereby improving the overall processing efficiency and stability, and further ensuring the accurate execution of instructions and the reliable output of results.
[0065] In some embodiments, in addition to the JSON structure, the instruction intent can also adopt other structured data formats, which are not limited herein.
[0066] In some embodiments, when the first large model obtains the initially parsed instruction intent and execution parameters, if it is detected that some key parameters are missing, have abnormal formats, or need to be further refined, the second large model can be used to complete parameter supplementation, that is, to supplement or correct these parameters through context association and deeper language understanding.
[0067] According to some embodiments, the information processing method may further include: using a second large model to complete the execution parameters of multiple instruction intents based on a user request. Step S402, generating a target instruction code based on the execution parameters of the target instruction intent may include: in response to determining that the second large model has completed the execution parameters of the target instruction intent, generating a target instruction code based on the completed execution parameters.
[0068] Thus, through the above method, even a first large model with a relatively small parameter scale can be used to initially parse the execution parameters of the instruction intent, and a second large model with a larger parameter scale is used to complete the execution parameters, thereby making full use of the performance advantages of the first large model and the more powerful natural language understanding ability of the second large model, and achieving efficient and accurate acquisition of the execution parameters.
[0069] According to some embodiments, the first large model may be trained using samples corresponding to multiple preset instruction intents. As Figure 5 shown, the information processing method 500 may further include: Step S502, obtaining a first prompt text corresponding to the new instruction intent; Step S503, using the second large model to identify the new instruction intent in the user request based on the first prompt text; and Step S504, in response to determining that the second large model has identified the new instruction intent in the user request, adding the new instruction intent to the multiple instruction intents. Steps S501, S505, and S506 in method 500 may respectively refer to steps S201 - S203 in method 200, which will not be elaborated here.
[0070] Thus, by training the first large model using samples corresponding to multiple preset instruction intents, the first large model can be enabled to have the ability to identify and parse these preset instruction intents. And if there are new intents temporarily added due to business (which the first large model has not seen before), the second large model can be used to identify and parse the new instruction intents through the above method, so that reasoning can still be carried out and a reasonable order can be given, without the need to train the model again, realizing rapid support for pluggable addition and subtraction of instruction intents.
[0071] In some embodiments, the first prompt text may include the name of the new instruction intent and at least one corresponding execution parameter, and may also include content such as the meaning explanation of the new instruction intent.
[0072] It can be understood that in the present disclosure, different steps using the same large model can be implemented by a single call to the large model. For example, Step S503, Step S504 (Step S202), and the above step of using the second large model to complete the execution parameters of multiple instruction intents can be implemented based on the combined text of multiple prompt texts in a single call.
[0073] According to some embodiments, each of the multiple preset instruction intents may include at least one preset execution parameter, and the number of samples corresponding to each of the multiple instruction intents may be positively correlated with the number of preset execution parameters included in each of the multiple preset instruction intents.
[0074] Thus, by making the number of parameters included in each preset instruction intent positively correlated with the corresponding number of samples, diverse parameter combination scenarios can be fully covered, enhancing the generalization and robustness of the first large model under different instruction intent requirements, and further improving the parsing accuracy and execution success rate of complex requests. This design effectively reduces invalid samples or parameter duplication, making model training more targeted.
[0075] In some embodiments, for each type of intent, samples with different numbers are constructed according to the number of parameters, and sample diversity is ensured, that is, different permutations and combinations of parameters, so that the model can more robustly recognize different parameters of the intent. The first large model can be trained for N rounds by fine-tuning with full-parameter instructions. In an exemplary embodiment, 5 rounds of training can be performed.
[0076] According to another aspect of the present disclosure, an information processing device based on a large model is provided. As Figure 6 shown, the device 600 includes: a first recognition unit 610, a determination unit 620, and a call unit 630. The first recognition unit 610 is configured to use the first large model to perform intent recognition on a user request to obtain multiple instruction intents; the determination unit 620 is configured to use the second large model to determine an execution sequence of the multiple instruction intents so that the execution sequence satisfies the dependency relationship between the multiple instruction intents, and the number of parameters of the second large model is greater than that of the first large model; and the call unit 630 is configured to sequentially call interfaces corresponding to the multiple instruction intents based on the execution sequence.
[0077] The operations of the above-mentioned first recognition unit 610, determination unit 620, and call unit 630 may respectively correspond to the operations of step S201, step S202, and step S203 as Figure 2 shown. Therefore, the details of each aspect will not be elaborated here.
[0078] In some embodiments, each of the multiple instruction intents includes an intent identifier and an execution parameter. The call unit 630 may include (not shown in the figure): a determination subunit, a generation subunit, and a call subunit. The determination subunit is configured to, for the target instruction intent to be executed currently, determine the target interface to be called based on the intent identifier of the target instruction intent; the generation subunit is configured to generate target instruction code based on the execution parameter of the target instruction intent; and the call subunit is configured to call the target interface based on the target instruction code.
[0079] The operations of the above-mentioned determination subunit, generation subunit, and invocation subunit can respectively correspond to the operations of step S401, step S402, and step S403 as shown below. Therefore, the details of each aspect will not be elaborated here. Figure 4 shown.
[0080] In some embodiments, the apparatus 600 may further include (not shown in the figure): a completion unit. The completion unit is configured to use the second large model to complete the execution parameters of multiple instruction intents based on a user request, and the generation subunit is configured to generate the target instruction code based on the completed execution parameters in response to determining that the second large model has completed the execution parameters of the target instruction intent.
[0081] In some embodiments, the first large model is trained using samples corresponding to multiple preset instruction intents respectively, and the apparatus 600 may further include (not shown in the figure): an acquisition unit, a second recognition unit, and a supplementation unit. The acquisition unit is configured to acquire a first prompt text corresponding to the new instruction intent; the second recognition unit is configured to use the second large model to recognize the new instruction intent in the user request based on the first prompt text; and the supplementation unit is configured to supplement the new instruction intent into the multiple instruction intents in response to determining that the second large model has recognized the new instruction intent in the user request.
[0082] The operations of the above-mentioned acquisition unit, second recognition unit, and supplementation unit can respectively correspond to the operations of step S501, step S502, and step S503 as shown below. Therefore, the details of each aspect will not be elaborated here. Figure 5 shown.
[0083] In some embodiments, each of the multiple preset instruction intents may include at least one preset execution parameter, and the number of samples corresponding to each of the multiple instruction intents may be positively correlated with the number of preset execution parameters included in each of the multiple preset instruction intents.
[0084] In some embodiments, the multiple instruction intents may adopt a json structure, and the json structure may include a field corresponding to the intent identifier and a field corresponding to the execution parameter. The apparatus 600 may further include (not shown in the figure): a correction unit. The correction unit is configured to perform json correction on the json structures corresponding to the multiple instruction intents after determining the execution sequence.
[0085] In some embodiments, the first large model may be a lightweight large language model, and the second large model may be a large-scale large language model.
[0086] In the technical solution of the present disclosure, the collection, storage, use, processing, transmission, provision, and disclosure of the user's personal information and other processing comply with the provisions of relevant laws and regulations and do not violate public order and good customs.
[0087] According to an embodiment of the present disclosure, an electronic device, a readable storage medium, and a computer program product are also provided.
[0088] Reference Figure 7 , the structural block diagram of the electronic device 700 that can be used as the server or client of the present disclosure will now be described. It is an example of a hardware device that can be applied to various aspects of the present disclosure. The electronic device is intended to represent various forms of digital electronic computer devices, such as, laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, personal digital processing, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are only examples and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0089] As Figure 7 shown, the electronic device 700 includes a computing unit 701, which can execute various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 702 or the computer program loaded from the storage unit 708 into the random access memory (RAM) 703. In the RAM 703, various programs and data required for the operation of the electronic device 700 can also be stored. The computing unit 701, the ROM 702, and the RAM 703 are connected to each other through a bus 704. The input / output (I / O) interface 705 is also connected to the bus 704.
[0090] Multiple components in the electronic device 700 are connected to the I / O interface 705, including: an input unit 706, an output unit 707, a storage unit 708, and a communication unit 709. The input unit 706 can be any type of device capable of inputting information into the electronic device 700. The input unit 706 can receive input digital or character information, and generate key signal inputs related to the user settings and / or function controls of the electronic device, and can include but are not limited to a mouse, a keyboard, a touch screen, a trackpad, a trackball, a joystick, a microphone, and / or a remote control. The output unit 707 can be any type of device capable of presenting information, and can include but are not limited to a display, a speaker, a video / audio output terminal, a vibrator, and / or a printer. The storage unit 708 can include but are not limited to magnetic disks and optical discs. The communication unit 709 allows the electronic device 700 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks, and can include but are not limited to a modem, a network card, an infrared communication device, a wireless communication transceiver, and / or a chipset, such as a Bluetooth device, an 802.11 device, a WiFi device, a WiMax device, a cellular communication device, and / or the like.
[0091] The computing unit 701 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 701 include but are not limited to a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 701 executes the various methods, processes, and / or operations described above. For example, in some embodiments, these methods, processes, and / or operations can be implemented as a computer software program that is tangibly contained in a machine-readable medium, such as the storage unit 708. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 700 via the ROM 702 and / or the communication unit 709. When the computer program is loaded into the RAM 703 and executed by the computing unit 701, one or more steps of the methods, processes, and / or operations described above can be executed. Alternatively, in other embodiments, the computing unit 701 can be configured to execute these methods, processes, and / or operations in any other suitable manner (e.g., by means of firmware).
[0092] The various embodiments of the systems and techniques described above in this specification can be implemented in digital electronic circuitry, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems-on-chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be a special-purpose or general-purpose programmable processor that receives data and instructions from, and transmits data and instructions to, a storage system, at least one input device, and at least one output device.
[0093] The program code for implementing the methods of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general purpose computer, special purpose computer, or other programmable data processing device, such that the program codes, when executed by the processor or controller, cause the functions / operations specified in the flowchart and / or block diagram to be implemented. The program code can be executed entirely on the machine, partly on the machine, as a stand-alone software package partly on the machine and partly on a remote machine, or entirely on the remote machine or server.
[0094] In the context of this disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0095] To provide for interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can also be used to provide for interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0096] The systems and techniques described herein can be implemented in a computing system that includes backend components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes frontend components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system that includes any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: local area network (LAN), wide area network (WAN), the Internet, and blockchain networks.
[0097] A computer system can include a client and a server. The client and the server are generally remote from each other and typically interact through a communication network. The client - server relationship is created by computer programs running on respective computers and having a client - server relationship with each other. The server can be a cloud server, a server of a distributed system, or a server incorporating blockchain.
[0098] It should be understood that the various forms of the processes shown above can be used, steps reordered, added, or deleted. For example, the steps recited in this disclosure can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and this is not limited herein.
[0099] Although embodiments or examples of the present disclosure have been described with reference to the accompanying drawings, it should be understood that the above methods, systems, and devices are merely exemplary embodiments or examples, and the scope of the present invention is not limited by these embodiments or examples, but is only defined by the authorized claims and their equivalent scope. Various elements in the embodiments or examples may be omitted or replaced by their equivalent elements. In addition, the steps may be executed in an order different from that described in the present disclosure. Further, the various elements in the embodiments or examples may be combined in various ways. Importantly, with the evolution of technology, many of the elements described herein may be replaced by equivalent elements that emerge after the present disclosure.
Claims
1. An information processing method based on a large model, comprising: Using the first model to identify the intent of the user request to obtain multiple instruction intents; Determining the execution sequence of the plurality of instruction intentions by using a second large model so that the execution sequence satisfies the dependency relationship between the plurality of instruction intentions, wherein the parameter amount of the second large model is greater than that of the first large model; as well as Based on the execution sequence, the interfaces corresponding to the multiple instruction intentions are called in sequence.
2. The method according to claim 1, wherein: Each of the plurality of instruction intents includes an intent identifier and an execution parameter. The calling of the interfaces corresponding to the plurality of instruction intentions in sequence based on the execution sequence includes: For the target instruction intent that needs to be executed currently, based on the intent identifier of the target instruction intent, determine the target interface that needs to be called; generating target instruction code based on the execution parameters of the target instruction intention; and Based on the target instruction code, the target interface is called.
3. The method according to claim 2, further comprising: Using the second large model, completing the execution parameters of the plurality of instruction intentions based on the user request, Wherein, the generating the target instruction code based on the execution parameter of the target instruction intention comprises: In response to determining that the second large model completes the execution parameters of the target instruction intent, the target instruction code is generated based on the completed execution parameters.
4. The method according to claim 2, wherein: The first model is obtained by training with samples corresponding to a plurality of preset instruction intentions. Wherein, the method further comprises: Obtaining the first prompt text corresponding to the newly added instruction intention; using the second large model to identify the new instruction intention in the user request based on the first prompt text; and In response to determining that the second large model recognizes the newly added instruction intent in the user request, the newly added instruction intent is added to the multiple instruction intents.
5. The method according to claim 4, wherein: Each of the multiple preset instruction intentions includes at least one preset execution parameter, and the number of samples corresponding to each of the multiple preset instruction intentions is positively correlated with the number of preset execution parameters included in each of the multiple preset instruction intentions.
6. The method according to claim 2, wherein: The multiple instruction intentions adopt a json structure, and the json structure includes a field corresponding to the intention identifier and a field corresponding to the execution parameter, Wherein, the method further comprises: After determining the execution sequence, JSON correction is performed on the JSON structures corresponding to each of the multiple instruction intentions.
7. The method according to any one of claims 1 to 6, wherein: The first large model is a lightweight large language model, and the second large model is a large-scale large language model.
8. An information processing device based on a large model, comprising: A first recognition unit is configured to use a first large model to perform intent recognition on a user request to obtain a plurality of instruction intents; a determining unit configured to determine an execution sequence of the plurality of instruction intentions by using a second large model so that the execution sequence satisfies a dependency relationship between the plurality of instruction intentions, wherein the parameter amount of the second large model is greater than that of the first large model; as well as The calling unit is configured to call the interfaces corresponding to the multiple instruction intentions in sequence based on the execution sequence.
9. The device according to claim 8, wherein: Each of the plurality of instruction intents includes an intent identifier and an execution parameter. Wherein, the calling unit includes: A determination subunit is configured to determine a target interface to be called based on an intent identifier of a target instruction intent that needs to be executed currently; a generating subunit configured to generate a target instruction code based on an execution parameter of the target instruction intention; and The calling subunit is configured to call the target interface based on the target instruction code.
10. The apparatus according to claim 9, further comprising: a completion unit, configured to use the second large model to complete the execution parameters of the plurality of instruction intentions based on the user request, The generating subunit is configured to generate the target instruction code based on the completed execution parameters in response to determining that the second large model has completed the execution parameters of the target instruction intent.
11. The device according to claim 10, wherein: The first model is obtained by training with samples corresponding to a plurality of preset instruction intentions. Wherein, the device further comprises: An acquisition unit, configured to acquire a first prompt text corresponding to the newly added instruction intention; a second recognition unit configured to recognize the newly added instruction intention in the user request based on the first prompt text by using the second large model; and The supplementation unit is configured to supplement the newly added instruction intent into the multiple instruction intents in response to determining that the second large model recognizes the newly added instruction intent in the user request.
12. The device according to claim 11, wherein Each of the multiple preset instruction intentions includes at least one preset execution parameter, and the number of samples corresponding to each of the multiple preset instruction intentions is positively correlated with the number of preset execution parameters included in each of the multiple preset instruction intentions.
13. The device according to claim 9, wherein: The multiple instruction intentions adopt a json structure, and the json structure includes a field corresponding to the intention identifier and a field corresponding to the execution parameter, Wherein, the device further comprises: The correction unit is configured to perform JSON correction on the JSON structures corresponding to each of the multiple instruction intentions after determining the execution sequence.
14. The device according to any one of claims 8 to 13, wherein: The first large model is a lightweight large language model, and the second large model is a large-scale large language model.
15. An electronic device, characterized in that: The electronic device comprises: at least one processor; and a memory communicatively coupled to the at least one processor; wherein The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 7.
16. A non-transitory computer-readable storage medium storing computer instructions, characterized in that: The computer instructions are used to cause the computer to execute the method according to any one of claims 1-7.
17. A computer program product comprising a computer program, wherein: When the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.