Model training method and device, model training application method and device and server
By introducing an intention disassembly model in question-and-answer processing, the question-and-answer sentences are cascaded and decomposed and derivative intent statements are generated, the problem of low accuracy in question-and-answer processing in the prior art is solved, and more efficient and accurate question-and-answer processing is achieved.
Patent Information
- Application Number
- CN202411867444.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-18
- Publication Date
- 2025-05-16
Smart Images

Figure CN120012854A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence technology, and in particular to a model training and its use method, device and server. Background Art
[0002] With the continuous development of deep learning technology, large language models (LLM) came into being. Text semantic understanding and natural language task processing through large language models have gradually become popular. They can be applied to different technical fields such as voice interaction control of display devices and image recognition, affecting users' lives.
[0003] However, when using a large language model to process question-answering sentences, the model usually focuses on the text of the question-answering sentence itself, rather than the overall processing process. As a result, when the large language model is used to process complex question-answering tasks, the processing results often fail to meet expectations and have poor accuracy. Summary of the invention
[0004] The present application provides a model training and its use method, device and server to solve the technical problem of poor accuracy of question and answer processing results in the prior art.
[0005] In a first aspect, some embodiments provide a model training method, comprising:
[0006] Acquire at least one first question-and-answer statement; the first question-and-answer statement is used to instruct question-and-answer processing to be performed based on at least two single application tools;
[0007] For each first question and answer statement, call the intention disassembly model to cascade disassemble the first question and answer statement until there is no composite intention statement in the disassembly result; the composite intention statement cannot be processed for question and answer based on a single application tool;
[0008] Generate a derived intent statement according to the decomposition result corresponding to each first question-answer statement and the decomposition relationship between different intent statements in the corresponding decomposition result;
[0009] According to the derived intention statements, updating the decomposition results of each of the first question and answer statements;
[0010] The intention disassembly model is trained using the updated disassembly results as training samples.
[0011] In some of the above-mentioned embodiments, the first question and answer statement is cascadedly decomposed by introducing an intent decomposition model, so that there are no more compound intent statements in the decomposition results that cannot be processed for question and answer based on a single application tool. This realizes the decomposition of the complex first question and answer statement into simple question and answer statements, ensures the atomicity of the question and answer processing process, avoids only referring to the corresponding semantics of the first question and answer statement itself, and ignores the task characteristics of the first question and answer statement, thereby improving the accuracy of the question and answer execution results. In addition, based on the disassembly results corresponding to each first question and answer sentence and the disassembly relationship between different intent sentences in the corresponding disassembly results, derived intent sentences are generated, thereby realizing the merging of some sentences. Subsequently, based on the generated derived intent sentences, the disassembly results of the first question and answer sentences are updated, and the updated disassembly results are used as training samples to train the intent disassembly model, so that the trained intent disassembly model can not only disassemble a single intent sentence for question and answer processing based on a single application tool, but also has the ability to disassemble derived intent sentences, avoiding the model from mechanically disassembling a single intent sentence, in order to affect the correlation between single intent sentences, thereby affecting the model's question and answer efficiency and the accuracy of the question and answer results, and further improving the accuracy of the disassembly results and the question and answer efficiency of the intent disassembly model.
[0012] In a second aspect, some embodiments further provide a model using method, including:
[0013] Acquire a second question-and-answer statement; the second question-and-answer statement is used to instruct to perform question-and-answer processing based on at least two single application tools;
[0014] Calling the intention disassembly model to perform cascade disassembly on the second question-answering task until there are no more compound intention sentences in the disassembly result; the intention disassembly model is trained based on the method provided by some embodiments of the first aspect;
[0015] When a single intent statement is disassembled, an application tool matching the single intent statement is run in the display device to perform question-and-answer processing on the corresponding single intent statement; the single intent statement can be processed with questions and answers based on a single application tool.
[0016] In some of the above-mentioned embodiments, by introducing a trained intent disassembly model, the second question and answer statement is cascaded and disassembled until there is no longer a complex intent statement in the disassembly result, and when a single intent statement that can be processed for question and answer based on a single application tool is disassembled, an application tool matching the single intent statement in the display device is run to perform question and answer processing on the corresponding single intent statement, thereby disassembling the complex second question and answer statement into a single intent statement that allows independent execution to perform atomic calls to the application tool, thereby improving the accuracy of question and answer processing.
[0017] In a third aspect, some embodiments further provide a model training device, including:
[0018] A first acquisition module is used to acquire at least one first question and answer sentence; the first question and answer sentence is used to instruct question and answer processing based on at least two single application tools;
[0019] A first calling module is used to call the intention disassembly model for each first question and answer sentence, so as to cascade disassemble the first question and answer sentence until there is no composite intention sentence in the disassembly result; the composite intention sentence cannot be processed for question and answer based on a single application tool;
[0020] A first generation module is used to generate a derived intent statement according to the disassembly result corresponding to each first question and answer statement and the disassembly relationship between different intent statements in the corresponding disassembly result;
[0021] A first updating module, configured to update the decomposition results of each of the first question-and-answer statements according to the derived intention statements;
[0022] The model training module is used to train the intention disassembly model using the updated disassembly results as training samples.
[0023] In some of the above embodiments, the first question and answer statement is cascadedly decomposed by introducing an intention decomposition model, until there are no more compound intention statements in the decomposition results that cannot be processed based on a single application tool, thereby realizing the decomposition of the complex first question and answer statement into simple question and answer statements, ensuring the atomicity of the question and answer processing process, and avoiding only referring to the semantics corresponding to the first question and answer statement itself while ignoring the task characteristics of the first question and answer statement, thereby improving the accuracy of the question and answer execution results. In addition, by introducing a first generation module, a derived intention statement is generated according to the decomposition results corresponding to each first question and answer statement, and the decomposition relationship between different intention statements in the corresponding decomposition results, thereby realizing the merging of some statements. Correspondingly, the disassembly result of the first question and answer statement is subsequently updated by the first updating module based on the generated derived intent statement, and the intent disassembly model is trained by the model training module using the updated disassembly result as a training sample, so that the trained intent disassembly model can not only disassemble a single intent statement for question and answer processing based on a single application tool, but also has the ability to disassemble derived intent statements, avoiding the model from mechanically disassembling a single intent statement, in order to affect the correlation between single intent statements, thereby affecting the model's question and answer efficiency and the accuracy of the question and answer results, and further improving the accuracy of the disassembly results and the question and answer efficiency of the intent disassembly model.
[0024] In a fourth aspect, some embodiments further provide a model using device, characterized in that it includes:
[0025] A second acquisition module is used to acquire a second question and answer statement; the second question and answer statement is used to instruct to perform question and answer processing based on at least two single application tools;
[0026] A second calling module is used to call the intention disassembly model to perform cascade disassembly on the second question-answering task until there is no longer a composite intention statement in the disassembly result; the intention disassembly model is trained based on the device provided in some embodiments of the third aspect;
[0027] The first running module is used to run the application tool matching the single intent statement in the display device when the single intent statement is disassembled, so as to perform question and answer processing on the corresponding single intent statement; the single intent statement can be processed with questions and answers based on a single application tool.
[0028] In some of the above embodiments, by introducing a trained intent disassembly model, the second question and answer statement is cascaded and disassembled until there is no longer a complex intent statement in the disassembly result, and when a single intent statement that can be processed for question and answer based on a single application tool is disassembled by the first running module, the application tool matching the single intent statement in the display device is run to perform question and answer processing on the corresponding single intent statement, thereby disassembling the complex second question and answer statement into a single intent statement that allows independent execution to perform atomic calls to the application tool, thereby improving the accuracy of question and answer processing.
[0029] In a fifth aspect, some embodiments further provide a server, including:
[0030] A communication device configured to be communicatively connected with the display device;
[0031] and at least one processor connected to the communication device and configured to perform the following steps:
[0032] Acquire at least one first question-and-answer statement; the first question-and-answer statement is used to instruct question-and-answer processing to be performed based on at least two single application tools;
[0033] For each first question and answer statement, call the intention disassembly model to cascade disassemble the first question and answer statement until there is no composite intention statement in the disassembly result; the composite intention statement cannot be processed for question and answer based on a single application tool;
[0034] Generate a derived intent statement according to the decomposition result corresponding to each first question-answer statement and the decomposition relationship between different intent statements in the corresponding decomposition result;
[0035] According to the derived intention statements, updating the decomposition results of each of the first question and answer statements;
[0036] The intention disassembly model is trained using the updated disassembly results as training samples.
[0037] In some of the above-mentioned embodiments, the first question and answer statement is cascadedly decomposed by introducing an intent decomposition model, so that there are no more compound intent statements in the decomposition results that cannot be processed for question and answer based on a single application tool. This realizes the decomposition of the complex first question and answer statement into simple question and answer statements, ensures the atomicity of the question and answer processing process, avoids only referring to the corresponding semantics of the first question and answer statement itself, and ignores the task characteristics of the first question and answer statement, thereby improving the accuracy of the question and answer execution results. In addition, based on the disassembly results corresponding to each first question and answer sentence and the disassembly relationship between different intent sentences in the corresponding disassembly results, derived intent sentences are generated, thereby realizing the merging of some sentences. Subsequently, based on the generated derived intent sentences, the disassembly results of the first question and answer sentences are updated, and the updated disassembly results are used as training samples to train the intent disassembly model, so that the trained intent disassembly model can not only disassemble a single intent sentence for question and answer processing based on a single application tool, but also has the ability to disassemble derived intent sentences, avoiding the model from mechanically disassembling a single intent sentence, in order to affect the correlation between single intent sentences, thereby affecting the model's question and answer efficiency and the accuracy of the question and answer results, and further improving the accuracy of the disassembly results and the question and answer efficiency of the intent disassembly model.
[0038] In a sixth aspect, some embodiments further provide a server, including:
[0039] A communication device configured to be communicatively connected with the display device;
[0040] and at least one processor connected to the communication device and configured to perform the following steps:
[0041] Acquire a second question-and-answer statement; the second question-and-answer statement is used to instruct to perform question-and-answer processing based on at least two single application tools;
[0042] Calling the intention disassembly model to perform cascade disassembly on the second question-answering task until there are no more compound intention sentences in the disassembly result; the intention disassembly model is trained based on the method provided by some embodiments of the first aspect;
[0043] When a single intent statement is disassembled, an application tool matching the single intent statement is run in the display device to perform question-and-answer processing on the corresponding single intent statement; the single intent statement can be processed with questions and answers based on a single application tool.
[0044] In some of the above-mentioned embodiments, by introducing a trained intent disassembly model, the second question and answer statement is cascaded and disassembled until there is no longer a complex intent statement in the disassembly result, and when a single intent statement that can be processed for question and answer based on a single application tool is disassembled, an application tool matching the single intent statement in the display device is run to perform question and answer processing on the corresponding single intent statement, thereby disassembling the complex second question and answer statement into a single intent statement that allows independent execution to perform atomic calls to the application tool, thereby improving the accuracy of question and answer processing. BRIEF DESCRIPTION OF THE DRAWINGS
[0045] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0046] Figure 1 A schematic diagram of an operation scenario between a display device and a control device provided in some embodiments;
[0047] Figure 2 A schematic diagram of the hardware configuration of a display device provided in some embodiments;
[0048] Figure 3 A schematic diagram of the hardware configuration of a control device provided in some embodiments;
[0049] Figure 4 A schematic diagram of software configuration of a display device provided in some embodiments;
[0050] Figure 5 A flowchart of a model training method provided in some embodiments;
[0051] Fig. 6A A flowchart of steps for generating a derived intent statement provided for some embodiments;
[0052] Figure 6B A schematic diagram of an intention disassembly tree provided for some embodiments;
[0053] Figure 6C A schematic diagram of an intention disassembly tree provided for some embodiments;
[0054] Figure 7 A flowchart of a method for using a model provided in some other embodiments;
[0055] Fig. 8A A flowchart of a question-and-answer processing method provided for other embodiments;
[0056] Figure 8B A flowchart of a question-and-answer processing method provided for other embodiments;
[0057] Figure 8C A schematic diagram of a processing procedure of an intention disassembly model provided in some embodiments;
[0058] Fig.8D A role distribution interaction diagram under the UATRA architecture provided for some embodiments;
[0059] Fig. 8E A schematic diagram of a cascade disassembly process provided for some embodiments;
[0060] Fig. 9 is a structural block diagram of a model training device in some embodiments;
[0061] Fig.10 A structural block diagram of a model using device in some embodiments;
[0062] Fig.11 FIG. 4 is a diagram showing the internal structure of a computer device in one embodiment. DETAILED DESCRIPTION
[0063] The following embodiments are described in detail, and examples thereof are shown in the accompanying drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementations described in the following embodiments do not represent all implementations consistent with the present application. They are only examples of systems and methods consistent with some aspects of the present application as detailed in the claims.
[0064] It should be noted that the brief description of terms in this application is only for the convenience of understanding the embodiments described below, and is not intended to limit the embodiments of this application. Unless otherwise specified, these terms should be understood according to their common and usual meanings.
[0065] The terms "first", "second", "third", etc. in the specification and claims of this application and the above drawings are used to distinguish similar or similar objects or entities, and do not necessarily mean to limit a specific order or sequence, unless otherwise noted. It should be understood that the terms used in this way can be interchangeable under appropriate circumstances.
[0066] The terms "comprises," "comprising," and "having," and any variations thereof, are intended to cover but not exclude inclusion, for example, a product or device comprising a list of components is not necessarily limited to all the components expressly listed but may include other components not expressly listed or inherent to such product or device.
[0067] The term "module" refers to any known or later developed hardware, software, firmware, artificial intelligence, fuzzy logic, or combination of hardware and / or software code that is capable of performing the functions associated with that element.
[0068] In the embodiment of the present application, the display device 200 generally refers to a device with image display and data processing capabilities. For example, the display device 200 includes but is not limited to a smart TV, a mobile terminal, a computer, a monitor, an advertising screen, a wearable device, a virtual reality device, an augmented reality device, etc.
[0069] Figure 1 This is a schematic diagram of an operation scenario between a display device and a control device provided in some embodiments of the present application. Figure 1 As shown in FIG. 1 , the user can operate the display device 200 through touch operation, the mobile terminal 300 and the control device 100. For example, the control device 100 can be a remote controller, a stylus pen, a handle, etc.
[0070] The mobile terminal 300 can be used as a control device for performing human-computer interaction between the user and the display device 200. The mobile terminal 300 can also be used as a communication device for establishing a communication connection with the display device 200 and performing data interaction. In some embodiments, the mobile terminal 300 can install software applications with the display device 200, and achieve connection and communication through a network communication protocol to achieve the purpose of one-to-one control operation and data communication. The audio and video content displayed on the mobile terminal 300 can also be transmitted to the display device 200 to achieve a synchronous display function.
[0071] like Figure 1 As also shown in FIG. 1 , the display device 200 also communicates data with the server 400 through various communication methods. The display device 200 may be allowed to communicate and connect through a local area network (LAN), a wireless local area network (WLAN) and other networks.
[0072] The display device 200 can provide a broadcast receiving television function, and can also additionally provide an intelligent network television function with a computer support function, including but not limited to network television, smart television, Internet Protocol Television (IPTV), etc.
[0073] Figure 2 Some embodiments of the present application provide Figure 1 2 is a block diagram of the hardware configuration of the display device 200.
[0074] In some embodiments, the display device 200 may include at least one of a tuner 210 , a communication device 220 , a detector 230 , a device interface 240 , a controller 250 , a display 260 , an audio output device 270 , a memory, a power supply, and a user input interface 280 .
[0075] In some embodiments, the detector 230 is used to collect signals from the external environment or external interactions. For example, the detector 230 includes a light receiver, a sensor for collecting ambient light intensity; or, the detector 230 includes an image collector, such as a camera, which can be used to collect external environment scenes, user attributes, or user interaction gestures; or, the detector 230 includes a sound collector, such as a microphone, which is used to receive external sounds, such as collecting voice data corresponding to the user's question-and-answer task.
[0076] In some embodiments, the display 260 includes a display function component for presenting a picture, and a driving component for driving an image display. The display 260 is used to receive an image signal output from the controller 250 for display. For example, the display 260 can be used to display video content, image content, and components of a menu control interface and a user control UI interface (User Interface). In an optional implementation, the display 260 can display question-and-answer tasks and corresponding question-and-answer replies.
[0077] In some embodiments, the communication device 220 is a component for communicating with an external device or server 400 according to various communication protocol types. The display device 200 may be provided with a plurality of communication devices 220 according to different supported communication modes. For example, when the display device 200 supports wireless network communication, the display device 200 may be provided with a communication device 220 including a WiFi (Wireless Fidelity) function. When the display device 200 supports Bluetooth connection communication, the display device 200 needs to be provided with a communication device 220 including a Bluetooth function.
[0078] The communication device 220 can enable the display device 200 to communicate with the external device or server 400 by wireless or wired connection. Among them, the wired connection can connect the display device 200 to the external device through components such as data cables and interfaces. The wireless connection can connect the display device 200 to the external device through wireless signals or wireless networks. The display device 200 can establish a connection relationship with the external device directly, or indirectly establish a connection relationship through a gateway, a router, a connection device, etc. In some optional embodiments, the display device 200 can call the intent disassembly model with the server through the communication device 220.
[0079] In some embodiments, the controller 250 may include at least one of a central processing unit, a video processor, an audio processor, a graphics processor, and a power processor, and a first interface to an nth interface for input / output. The controller 250 controls the operation of the display device and responds to the user's operation through various software control programs stored in the memory. The controller 250 controls the overall operation of the display device 200.
[0080] In some embodiments, the controller 250 and the tuner-demodulator 210 may be located in different separate devices, that is, the tuner-demodulator 210 may also be located in an external device of the main device where the controller 250 is located, such as an external set-top box.
[0081] In some embodiments, the user may input a user command through a graphical user interface (GUI) displayed on the display 260 , and the user input interface receives the user input command through the graphical user interface (GUI).
[0082] In some embodiments, the audio output device 270 may be a local speaker of the display device 200, or may be an external audio output device of the display device 200. In particular, for the external audio output device of the display device 200, the display device 200 may also be provided with an external audio output terminal, and the audio output device may be connected to the display device 200 through the external audio output terminal to output the sound of the display device 200.
[0083] In some embodiments, the user input interface 280 may be used to receive question and answer statements etc. inputted by the user.
[0084] Figure 3 Some embodiments of the present application provide Figure 1 The hardware configuration diagram of the control device in the figure. Figure 3 As shown, the control device 100 may include: a controller 110 , a communication interface 130 , a user input / output interface 140 , a power supply 180 , and a memory 190 .
[0085] The control device 100 is configured to control the display device 200 , and can receive user input operation instructions, and convert the operation instructions into instructions that the display device 200 can recognize and respond to, playing the role of an interactive intermediary between the user and the display device 200 .
[0086] In some embodiments, the control device 100 may be a smart device, for example, the control device 100 may be installed with various applications for controlling the display device 200 according to user needs.
[0087] In some embodiments, Figure 1As shown, the mobile terminal 300 or other intelligent electronic devices can play a similar function as the control device 100 after installing the application for controlling the display device 200 .
[0088] The controller 110 includes a processor 112, a RAM (Random Access Memory) 113, a ROM (Read-Only Memory) 114, a communication interface 130, and a communication bus. The controller 110 is used to control the operation and operation of the control device 100, as well as the communication and cooperation between the internal components and the external and internal data processing functions.
[0089] The communication interface 130 implements communication of control signals and data signals with the display device 200 under the control of the controller 110. The communication interface 130 may include at least one of other near field communication modules such as a WiFi chip 131, a Bluetooth module 132, and an NFC (Near Field Communication) module 133.
[0090] The user input / output interface 140 includes at least one of other input interfaces such as a microphone 141, a touch panel 142, a sensor 143, a button 144, etc. In some embodiments, the user input interface may be configured to receive a question and answer statement, etc.
[0091] In some embodiments, the control device 100 includes at least one of a communication interface 130 and an input / output interface 140. The control device 100 is configured with a communication interface 130, such as a WiFi, Bluetooth, NFC or other module, and can encode the user input command through the WiFi protocol, Bluetooth protocol, or NFC protocol and send it to the display device 200.
[0092] The memory 190 is used to store various operating programs, data and applications for driving and controlling the control device 100 under the control of the controller. The memory 190 can store various control signal instructions input by the user.
[0093] The power supply 180 is used to provide operating power support for each component of the control device 100 under the control of the controller.
[0094] In order to perform user interaction, in some embodiments, the display device 200 may run an operating system. The operating system is a computer program for managing and controlling hardware resources and software resources in the display device 200. The operating system may provide a user interface (control the display device), allow the user to interact with the display device 200, and support the running of various application programs.
[0095] It should be noted that the operating system can be a native operating system based on a specific operating platform, a third-party operating system deeply customized based on a specific operating platform, or an independent operating system specially developed for the display device.
[0096] The operating system can be divided into different modules or layers according to the functions implemented, such as Figure 4 As shown, Figure 4 Some embodiments of the present application provide Figure 1 Schematic diagram of software configuration in the display device. In some embodiments, the system of the display device 200 can be divided into three layers, namely, application layer, middleware layer and hardware layer from top to bottom.
[0097] The application layer mainly includes commonly used applications on TV and application frameworks. Common applications are mainly applications developed based on browsers, such as HTML5 APPs (HyperText Markup Language 5 Applications, applications based on web technologies such as the fifth edition of Hypertext Markup Language); and native applications;
[0098] The Application Framework is a complete program model that has all the basic functions required by standard application software, such as file access, data exchange, etc., as well as the user interfaces of these functions (toolbars, status bars, menus, dialog boxes).
[0099] Native apps can support online or offline, message push or local resource access.
[0100] The middleware layer includes various TV protocols, multimedia protocols, system components and other middleware. The middleware can use the basic services (functions) provided by the system software to connect various parts of the application system or different applications on the network, and can achieve the purpose of resource sharing and function sharing.
[0101] The hardware layer mainly includes the hardware abstraction layer (HAL) interface, hardware and drivers. The HAL interface is a unified interface for all TV chips to connect, and the specific logic is implemented by each chip. Drivers mainly include: audio driver, display driver, Bluetooth driver, camera driver, WIFI driver, USB driver, HDMI driver, sensor driver (such as fingerprint sensor, temperature sensor, pressure sensor, etc.), and power driver.
[0102] It should be noted that the above example is only a simple division of the operating system functions and does not constitute a limitation on the specific operating system form of the display device 200 in the embodiment of the present application. Depending on factors such as the function of the display device and the type of operating system, the number of levels and specific level types contained in the operating system may be expressed in other forms.
[0103] In the prior art, the intent decomposition model generated by the large language model, machine learning model or deep learning model usually focuses on the text of the first question and answer sentence itself, rather than on the task processing, when processing the first question and answer sentence. Therefore, when the question and answer intent corresponding to the first question and answer sentence is relatively complex, the question and answer reply content will be unsatisfactory, thus affecting the accuracy of the question and answer results.
[0104] In order to overcome the above problems, in some optional embodiments, see Figure 5 The model training method shown is applied to the display device 200 or other devices connected to the display device 200 in communication, such as a server, and may include the following steps:
[0105] S510, obtaining at least one first question and answer statement; the first question and answer statement is used to instruct question and answer processing based on at least two single application tools.
[0106] The first question-and-answer statement can be understood as a question-and-answer statement dedicated to participating in the training of the intent decomposition model. The first question-and-answer statement can be a voice question-and-answer statement or a text question-and-answer statement, etc. The number of the first question-and-answer statements can be at least one, usually multiple, to provide a large amount of data support for subsequent model training.
[0107] In some optional embodiments, at least one first question-answering sentence input by the user may be obtained through the user input interface 280 of the display device 200 for use in subsequent model training. Alternatively, the communication device of the server may obtain at least one first question-answering sentence transmitted by the communication device 220 of the display device 200.
[0108] Optionally, the first question and answer statements may be input in parallel, and each first question and answer statement may be processed separately later. Alternatively, the first question and answer statements may be input in series, and after one first question and answer statement is executed, the next first question and answer statement may be obtained and processed.
[0109] S520: For each first question and answer sentence, call the intention decomposition model to perform cascade decomposition on the first question and answer sentence until there is no composite intention sentence in the decomposition result.
[0110] The intent disassembly model can be generated based on a traditional machine learning model, a deep learning model or a large language model, and no limitation is imposed on the specific network structure of the intent disassembly model. In some optional embodiments, the intent disassembly model can directly reuse an existing generative model.
[0111] Among them, the cascade disassembly process may include disassembling a first question and answer sentence to obtain at least one sentence, and continuing to disassemble the obtained compound intent sentence until there are no more compound intent sentences, terminating the disassembly process, and finally obtaining at least one disassembly result, as well as the disassembly relationship between different intent sentences in the disassembly result.
[0112] Exemplarily, the intent statement can be divided into a single intent statement and a compound intent statement according to whether it can be processed for question and answer based on a single application tool. Among them, the single intent statement can be processed for question and answer based on a single application tool, that is, the intent statement can be independently executed by a single application tool; the compound intent statement cannot be processed for question and answer based on a single application tool, that is, it can be processed for question and answer based on at least two single application tools, and accordingly, the intent statement is prohibited from being independently executed by a single application tool.
[0113] It is worth noting that a single intent statement can be set or adjusted by a technician according to actual needs, and this embodiment does not impose any restrictions on this. It is understandable that in order to avoid the uncertainty of the execution process of a single intent statement, multiple single intent statements with relatively fixed execution processes can also be integrated into a single intent statement without performing more fine-grained statement decomposition.
[0114] For example, the search statement corresponding to the search task can be split into a network search statement for obtaining relevant web page links, which can be further split into a single intent statement for searching each web page link for relevant content. However, since this search process is relatively fixed, the network search statement can be directly processed as a single intent statement.
[0115] In order to ensure the smooth execution of the disassembly process and improve the disassembly efficiency, in some embodiments, each statement obtained in the disassembly process includes at least one single intention statement.
[0116] In some embodiments, when a single intent statement is decomposed, an application tool matching the single intent statement is run in a display device to perform question and answer processing on the corresponding single intent statement.
[0117] Exemplarily, when a single intent statement is deconstructed, a first tool call request of a corresponding application tool is generated for the single intent statement. The first tool call request includes an application tool to be called, which is used to instruct to call a corresponding application tool in a display device to process the corresponding single intent statement. There is at least one first tool call request, and the order in which different first tool call requests are generated corresponds to the order in which the corresponding single intent statement is deconstructed.
[0118] In some optional embodiments, for each disassembly process, a task thread can be created, and the corresponding first question and answer statement or the disassembled compound intent statement can be split under the task thread to obtain a disassembly result; when there is only one intent statement in the disassembly result, continue to execute the next disassembly process based on the task thread; when there are at least two intent statements in the disassembly result, execute the disassembly process of one of the compound intent statements based on the task thread, or execute the application tool calling process of one of the single intent statements; and, create a new task thread for each of the remaining intent statements, and perform intent disassembly or application tool calling process on the corresponding intent statement based on each task thread.
[0119] It is worth noting that at least two single intent statements can be in a serial execution relationship; accordingly, the application tools corresponding to each single intent statement can be serially called according to the disassembly order of the single intent statements. Alternatively, at least two single intent statements can be in a parallel execution relationship; accordingly, different threads or devices can be used to call the application tools corresponding to each single intent statement in parallel, thereby improving the efficiency of question and answer processing.
[0120] Exemplarily, the execution relationship between different single intent statements can be determined based on the result dependency between the corresponding intent statements.
[0121] Optionally, the display device 200 may be controlled to display the processing result of each question and answer statement on the display.
[0122] S530, generating a derived intent statement according to the decomposition result corresponding to each first question and answer statement and the decomposition relationship between different intent statements in the corresponding decomposition result.
[0123] The decomposition relationship can be understood as which intention statement a certain intention statement is decomposed from, or which intention statements are included in the decomposition result of decomposing the intention statement. The derived intention statement is a compound intention statement that allows independent execution.
[0124] Exemplarily, the disassembly relationship between different intention statements in the disassembly result can be sorted out to determine the execution order between different intention statements; obtain the intention statement group whose repetition frequency is greater than the preset frequency threshold, and generate the derived intention statement based on each intention statement in the intention statement group. Among them, the intention statement group includes at least two intention statements executed sequentially, and at least two intention statement groups are obtained by disassembling the same intention statement. Among them, the preset frequency threshold can be set by the technician according to needs or experience, or determined through a large number of experiments.
[0125] In some optional embodiments, the processing logic of each intention statement in the intention statement group can be combined according to the question and answer execution order of each intention statement in the intention statement group to generate a derived intention statement.
[0126] S540: Update the decomposition results of each first question and answer statement according to the derived intention statement.
[0127] Exemplarily, a derived intent statement may be directly used to replace the intent statement used to generate the derived intent statement in the decomposition result of the first question and answer statement.
[0128] S550: Use the updated disassembly results as training samples to train the intention disassembly model.
[0129] In some embodiments, the intent statements in the updated disassembly results can be combined into training samples in a preset format according to the disassembly order; based on the training samples, the intent disassembly model is trained to adjust the network parameters in the intent disassembly model to gradually improve the sentence disassembly capability and question-answering capability of the intent disassembly model.
[0130] It is worth noting that the last training process can be a process of secondary training of a pre-trained intent disassembly model. Accordingly, the first training process of the intent disassembly model can be implemented in the following manner: obtain a sample question and answer sentence; and artificially split the sample question and answer sentence into at least one intent sentence, and generate a first training sample based on a preset format for each intent sentence, the processing order of the intent sentence, and the question and answer result of the intent sentence; based on the first training sample, train the pre-constructed large language model to obtain an intent disassembly model that gradually has sentence disassembly capabilities and question and answer capabilities. Furthermore, each sentence in the aforementioned updated disassembly result is combined into a second training sample in a preset format according to the sentence disassembly order; based on the second training sample, the aforementioned trained intent disassembly model is trained for the second time to adjust the network parameters in the intent disassembly model to gradually improve the sentence disassembly capabilities and question and answer capabilities of the intent disassembly model.
[0131] In some of the above-mentioned embodiments, the first question and answer statement is cascadedly decomposed by introducing an intent decomposition model, so that there are no more compound intent statements in the decomposition results that cannot be processed for question and answer based on a single application tool. This realizes the decomposition of the complex first question and answer statement into simple intent statements, ensures the atomicity of the question and answer processing process, avoids only referring to the corresponding semantics of the first question and answer statement itself, and ignores the task characteristics of the first question and answer statement, thereby improving the accuracy of the question and answer execution results. In addition, based on the disassembly results corresponding to each first question and answer sentence and the disassembly relationship between different intent sentences in the corresponding disassembly results, derived intent sentences are generated, thereby realizing the merging of some sentences. Subsequently, based on the generated derived intent sentences, the disassembly results of the first question and answer sentences are updated, and the updated disassembly results are used as training samples to train the intent disassembly model, so that the trained intent disassembly model can not only disassemble a single intent sentence for question and answer processing based on a single application tool, but also has the ability to disassemble derived intent sentences, avoiding the model from mechanically disassembling a single intent sentence, in order to affect the correlation between single intent sentences, thereby affecting the model's question and answer efficiency and the accuracy of the question and answer results, and further improving the accuracy of the disassembly results and the question and answer efficiency of the intent disassembly model.
[0132] Based on the technical solutions of the above embodiments, some optional embodiments are also provided. In these optional embodiments, the step of generating the derived intention statement in S530 is refined.
[0133] See also Fig. 6A The steps for generating the derived intent statement shown include:
[0134] S610, for each first question and answer sentence, construct an intent decomposition tree corresponding to the first question and answer sentence according to the decomposition result corresponding to the first question and answer sentence and the decomposition relationship between different intent sentences in the corresponding decomposition result.
[0135] Exemplarily, the root node in the intention decomposition tree corresponds to the first question and answer sentence; the child nodes in the intention decomposition tree correspond to the intention sentences in the decomposition result; the intention sentences corresponding to the child nodes in the intention decomposition tree are the sentence decomposition results of the intention sentences corresponding to the corresponding parent nodes;
[0136] S620, select a target subtree according to the frequency of occurrence of each subtree in the intention decomposition tree corresponding to at least one first question and answer statement.
[0137] For example, a subtree with a frequency greater than a preset frequency threshold can be selected from the intention disassembly tree corresponding to at least one first question-and-answer statement as the target subtree, so as to effectively identify the same intention statement group obtained in the disassembly process. The preset frequency threshold can be set or adjusted by the technician according to needs or experience, or repeatedly determined through a large number of experiments, and is not limited here.
[0138] S630, taking the intention statement corresponding to the root node in the target subtree as the derived intention statement.
[0139] Exemplarily, the corresponding question and answer processing logics may be combined in the order of execution of the intent statements corresponding to the target subtree to generate derived intent statements.
[0140] It is worth noting that since the derived intent statements are obtained by integrating the intent statements in the same intent statement group with a higher frequency of occurrence, these intent statement groups can be regarded as a single intent statement that can be executed independently, without the need to further disassemble, thus avoiding the mechanization of the disassembly process, improving the disassembly flexibility and efficiency, and thus helping to improve the efficiency of subsequent question and answer processing. At the same time, the above disassembly process not only ensures the atomicity of the execution of the intent statement, but also takes into account the correlation between different intent statements, and also reduces the uncertainty in the question and answer processing process, thereby helping to improve the accuracy of the question and answer processing results.
[0141] See also Figure 6B The schematic diagram of the intent decomposition tree is shown in the figure, where 1-4 are all leaf nodes, corresponding to a single intent statement. Since the subtree structure corresponding to 1-3-2 is repeated many times, the subtree corresponding to 1-3-2 can be combined into a new derived intent statement 5, which can be used as a new single intent statement in the cascade decomposition process and executed independently. Accordingly, the original intent decomposition tree is changed to Figure 6C Intended disassembly tree shown.
[0142] In the above optional embodiment, by introducing the intention decomposition tree to assist in the execution of the derived intention statement, the determination efficiency of the target subtree and the accuracy of the determination result are guaranteed, thereby improving the generation efficiency and the accuracy of the derived intention statement.
[0143] The above content has detailedly explained the training process of the intent disassembly model. The following will explain in detail the use process of the intent disassembly model based on the technical solutions of the above embodiments.
[0144] See also Figure 7 The model usage method shown is applied to the controller 110 of the display device 200, or other devices connected to the display device 200 in communication, such as a server, including:
[0145] S710, obtaining a second question and answer statement; the second question and answer statement is used to instruct question and answer processing to be performed based on at least two single application tools.
[0146] The second question and answer statement can be understood as a question and answer statement in the use stage of the intention disassembly model. The second question and answer statement can be a voice question and answer statement or a text question and answer statement, etc. This embodiment does not impose any limitation on the specific presentation form of the second question and answer statement.
[0147] In some optional embodiments, the second question and answer sentence input by the user may be obtained through the user input interface 280 in the display device 200. Alternatively, the second question and answer sentence transmitted by the communication device 220 of the display device 200 may be obtained by the communication device of the server.
[0148] S720, calling the intent decomposition model to perform cascade decomposition on the second question-answering task until there are no more compound intent sentences in the decomposition result.
[0149] Among them, the intention disassembly model is trained based on the model training method provided in the aforementioned embodiments.
[0150] Among them, the cascade disassembly process may include disassembling the second question and answer sentence to obtain at least one intent statement, and continuing to disassemble the obtained compound intent statement until there are no more compound intent statements, terminating the disassembly process, and finally obtaining at least one disassembly result, as well as the disassembly relationship between different intent statements in the disassembly result.
[0151] Exemplarily, the intent statement can be divided into a single intent statement and a compound intent statement according to whether it can be processed for question and answer based on a single application tool. Among them, the single intent statement can be processed for question and answer based on a single application tool, that is, the intent statement can be independently executed by a single application tool; the compound intent statement cannot be processed for question and answer based on a single application tool, that is, it can be processed for question and answer based on at least two single application tools, and accordingly, the intent statement is prohibited from being independently executed by a single application tool.
[0152] It is worth noting that a single intent statement can be set or adjusted by a technician according to actual needs, and this embodiment does not impose any restrictions on this. It is understandable that in order to avoid the uncertainty of the execution process of a single intent statement, multiple single intent statements with relatively fixed execution processes can also be preset into a single intent statement without performing more fine-grained statement decomposition.
[0153] In order to ensure the smooth execution of the disassembly process and improve the disassembly efficiency, in some embodiments, the intent statement obtained in each disassembly process includes at least one single intent statement.
[0154] It is worth noting that the intention disassembly model can be stored in the memory 190 of the display device 200, and the controller 110 in the display device 200 processes the second question and answer statement based on the locally stored intention disassembly model.
[0155] Of course, in order to reduce the computing power requirements and hardware costs of the display device 200, the intention disassembly model can usually be stored in the server, and when necessary, the controller 110 of the display device 200 can remotely call the intention disassembly model in the server, and feed the call result back to the display device 200, which will be obtained by the communication device 220 in the display device 200.
[0156] In some optional embodiments, for each disassembly process, a task thread can be created, and the corresponding second question and answer statement or the disassembled compound intent statement can be split under the task thread to obtain at least one disassembly result; when there is only one intent statement in the disassembly result, continue to execute the next disassembly process based on the task thread; when there are at least two intent statements in the disassembly result, execute the disassembly process of one of the compound intent statements based on the task thread, or execute the application tool calling process of one of the single intent statements; and, create a new task thread for each of the remaining intent statements, and perform intent disassembly or application tool calling process on the corresponding intent statement based on each task thread.
[0157] It is worth noting that at least two single intent statements can be in a serial execution relationship; accordingly, the application tools corresponding to each single intent statement can be serially called according to the disassembly order of the single intent statements. Alternatively, at least two single intent statements can be in a parallel execution relationship; accordingly, different threads or devices can be used to call the application tools corresponding to each single intent statement in parallel, thereby improving the efficiency of question and answer processing.
[0158] Exemplarily, the execution relationship between different single intent statements can be determined based on the result dependency between the corresponding intent statements.
[0159] Optionally, the display device 200 may be controlled to display the processing result of each intention statement on the display.
[0160] S730, when a single intent statement is disassembled, an application tool matching the single intent statement in the display device is run to perform question and answer processing on the corresponding single intent statement.
[0161] Exemplarily, when a single intent statement is decomposed, a second tool call request of a corresponding application tool is generated for the single intent statement. The second tool call request includes an application tool to be called, which is used to instruct to call a corresponding application tool in a display device to process the corresponding single intent statement. There is at least one second tool call request, and the order in which different second tool call requests are generated corresponds to the order in which the corresponding single intent statement is decomposed.
[0162] In some optional embodiments, when there is a failed question and answer processing result, the intention disassembly model is called to re-disassemble the second question and answer statement in cascade; wherein the re-disassembly process of the second question and answer statement is at least partially different from the previous disassembly process of the second question and answer statement.
[0163] It is understandable that in order to avoid repeated and invalid calls to the intent disassembly model, when the execution result indicates that the corresponding question and answer statement has failed to execute, the intent disassembly model is re-called to re-process the second question and answer task. A different disassembly method is usually used to avoid the situation where the intent statement is disassembled exactly the same as the historical disassembly process of the second question and answer task.
[0164] In other optional embodiments, the second question and answer statement can also be semantically parsed to obtain intent analysis data of the second question and answer statement; for the intent statement in the decomposition result of the second question and answer task, the intent analysis data is used as a system prompt word to perform question and answer processing on the corresponding intent statement.
[0165] It is worth noting that since the intent analysis data can be used as system prompts during the processing of the disassembled intent statements, it is ensured that the generated content of the intent statement does not deviate from the task theme of the second question and answer statement, further improving the accuracy of the processing results of the logically complex second question and answer statement.
[0166] In some other optional embodiments, when a derived intent statement is decomposed, at least two target application tools corresponding to the derived intent statement are obtained; the target application tools are used to perform question and answer processing on a single intent statement in the decomposition result of the derived intent statement; in accordance with a target order, each target application tool in the display device is run to perform question and answer processing on the derived intent statement; the target order is opposite to the generation order of each single intent statement in the decomposition result of the derived intent statement.
[0167] Among them, the derived intent statement is generated based on at least two single intent statements, and the derived intent statement can also be decomposed into the aforementioned at least two single intent statements; accordingly, each single intent statement corresponds to an application tool for question and answer processing. In order to distinguish the application tools here from other application tools, the application tool of the single intent statement corresponding to the derived intent statement is called the target application tool.
[0168] Correspondingly, the generation order of each single intent statement decomposed from the derived intent statement is obtained; the order opposite to the generation order is used as the target order; and each target application tool in the display device 200 is run according to the target order, thereby realizing question and answer processing of the derived intent statement.
[0169] It is understandable that since the derived intent statements are obtained by integrating the intent statements in the same intent statement group with a higher frequency of occurrence, these intent statement groups can be regarded as a single intent statement that can be executed independently, and there is no need to further disassemble them. Therefore, after disassembling the derived intent statements, the target application tools in the target application tool group corresponding to the derived intent statements can be directly run, taking into account the correlation between different intent statements and reducing the uncertainty in the question and answer processing process, thereby helping to improve the accuracy of the question and answer processing results.
[0170] In some of the above-mentioned embodiments, by introducing a trained intent disassembly model, the second question and answer statement is cascaded and disassembled until there is no longer a complex intent statement in the disassembly result, and when a single intent statement that can be processed for question and answer based on a single application tool is disassembled, an application tool matching the single intent statement in the display device is run to perform question and answer processing on the corresponding single intent statement, thereby disassembling the complex second question and answer statement into a single intent statement that allows independent execution to perform atomic calls to the application tool, thereby improving the accuracy of question and answer processing.
[0171] Based on the technical solutions of the above embodiments, some optional embodiments are also provided. In the optional embodiments, the question-answering process is described in detail from the perspective of the interaction between the display device 200 and the server. It is worth noting that in the question-answering process, some steps can be migrated from the display device 200 to the server, or from the server to the display device 200, as long as the computing power in the display device 200 allows.
[0172] See also Fig. 8A The question-answering processing method shown is applied to the training phase of the intent decomposition model, including:
[0173] S811, the display device obtains at least one first question and answer sentence;
[0174] S812, the display device calls the intent decomposition model trained in the server for each first question and answer sentence, processes the first question and answer sentence, and obtains a first question and answer reply and a first tool call request;
[0175] S813, when processing the first question and answer statement, the server cascades and decomposes the first question and answer statement until there are no more compound intent statements in the decomposition result, and when a single intent statement is decomposed, generates a first tool call request corresponding to the single intent statement; at least one intent statement obtained in each decomposition process includes at least one single intent statement;
[0176] S814, the display device runs a corresponding application tool in the display device in response to the first tool calling request to process the corresponding single intent statement;
[0177] S815, the display device displays the first question and answer statement and the first question and answer reply;
[0178] S816, the server constructs an intent decomposition tree corresponding to the first question and answer sentence according to at least one intent sentence obtained by cascading decomposition of the first question and answer sentence and the decomposition relationship between different intent sentences;
[0179] S817, the server determines a target subtree according to the occurrence frequency of each subtree in the intention decomposition tree corresponding to the at least one first question and answer statement;
[0180] S818, the server generates a derived intent statement according to the intent statement corresponding to the target subtree;
[0181] S819, the server updates the decomposition results of each first question and answer statement according to the derived intent statement, and uses the updated decomposition results as training samples to train the intent decomposition model.
[0182] It should be noted that, in some embodiments, there is no limitation on the execution order of S814 to S815 and S816 to S819, and the two can be executed sequentially, crosswise, or in parallel.
[0183] See also Figure 8B The question-answering processing method shown is applied to the use phase of the intent decomposition model, including:
[0184] S821, the display device obtains a second question and answer statement;
[0185] S822, the display device calls the intent decomposition model trained in the server to process the second question and answer statement to obtain a second question and answer reply and a second tool calling request;
[0186] S823, when processing the second question and answer statement, the server cascades and decomposes the second question and answer statement until no compound intent statement exists;
[0187] S824, when a single intent statement is decomposed, sending a second tool call request corresponding to the single intent statement to the display device;
[0188] Exemplarily, at least one intention statement obtained from each decomposition process includes at least one single intention statement.
[0189] S825, the display device runs a corresponding application tool in the display device in response to the second tool calling request to process the corresponding single intent statement;
[0190] S826, when the derived intent statement is disassembled, sending a third tool call request corresponding to the derived intent statement to the display device;
[0191] S827, the display device runs at least two application tools in the display device in response to the third tool calling request to process the corresponding derived intent statement;
[0192] S828: The display device displays the second question-and-answer task and the second question-and-answer reply.
[0193] It should be noted that, in some embodiments, only the execution order of S824 to S835 and S826 to S827 is exemplified, and the two can be executed in sequence according to the generation order.
[0194] In some optional embodiments, the intent decomposition model can be implemented based on the Transformer decoder. Figure 8C The processing diagram of the intent disassembly model shown in the figure shows that when the intent disassembly model is called to process the question and answer sentence, the question and answer description text corresponding to the question and answer sentence can be embedded and encoded through the embedding layer to generate an embedding vector under the preset role architecture; the training sample data is subjected to task cascade disassembly and subtask association processing through the Transformer decoder to obtain the question and answer reply. Among them, the preset role architecture can adopt the SUA (System-User-Assistant) architecture of "system role-user role-assistant role", or the UATRA (User-Analyze-Tools-Return-Assistant) architecture of "user role-analysis role-tool role-feedback role-assistant role". Accordingly, the embedding vector can include tokens of different categories (the basic processing unit of text data in deep learning), can include ordinary tokens in traditional deep learning, and can also include role tokens corresponding to the aforementioned different roles.
[0195] For example, the Analyze role is used to analyze user intentions and determine whether to use tools, which is equivalent to the model thinking process; the Tools role is the interface for calling application tools, and the content is generated by the model; the Return role is the tool return interface, and the content is generated by external tools. It can be understood that, see Fig.8D From the role distribution interaction diagram under the UATRA architecture shown in the figure, it can be seen that the UATRA architecture expands the interaction between "AI-human" to the interaction between "tools-AI-human" based on the traditional SUA architecture. It is worth noting that the content of the Analyze role, Tools role and Assistant role under the UATRA architecture is generated by the model, and the content of the User and Tools roles is input from the outside.
[0196] To implement the above architecture, the model vocabulary under the SUA architecture can be adjusted, and the three newly added roles can be inserted into the original vocabulary as special tokens. During model inference, text generation can be performed together with ordinary text.
[0197] It is worth noting that since the model under the traditional SUA architecture only considers the conversation data of human-AI interaction, but lacks the corresponding "tools-return-analyze" type conversation data between AI and tools, this type of data needs to be supplemented during the training process of the intent decomposition model.
[0198] Taking the question-and-answer sentence "What time is it now?" as an example, the content distribution corresponding to different roles is as follows:
[0199] <|User|>What time is it now?
[0200] <|Analyze|> If the user wants to check the current time, but the assistant does not know the current time, the clock tool can be used.
[0201] <|Tools|>{"tools":"clock.get",params:{}} / / Call the clock.get tool, params:{} is the calling parameter;
[0202] <|Return|>Current time: 2024-9-20 14:25:16
[0203] <|Analyze|> From the result returned by the clock tool, we can see that it is now September 20, 2024, 2:25:16 pm.
[0204] <|Assistant|>: It's 2:25 PM. How can I help you?
[0205] The above briefly describes the structure of the intent decomposition model and the construction process of the corresponding training samples. The following will describe the cascade decomposition process in detail with specific examples.
[0206] See also Fig. 8E The schematic diagram of the cascade disassembly process is shown. Among them, Task 1 corresponds to a question-and-answer statement (corresponding to a compound intent statement for question-and-answer processing based on at least two application tools). Accordingly, the above compound intent statement is split through the Analyze role to gradually obtain independent task single intent statements, that is, the smallest executable unit that can be executed independently. Among them, the smallest executable unit refers to a task that can be completed using one tool. For non-minimum executable units, the model needs to assign the task to itself for secondary splitting to achieve autonomous task dispatching.
[0207] Specifically, Task 1 can be split into three subtasks, including Task 1.1, Task 1.2, and Task 1.3. Among them, Task 1.1 is an independent task that allows independent execution (corresponding to a single intent statement), and the corresponding tool 1 (that is, the application tool) can be directly called to perform task processing (that is, question and answer processing) of the independent task. Among them, Task 1.2 and Task 1.3 are both composite tasks (corresponding to composite intent statements) that are prohibited from independent execution. Therefore, Task 1.2 and Task 1.3 need to be further disassembled separately. Among them, task 1.2 can be decomposed into independent task 1.2.1, which can be processed by directly calling tool 2; task 1.3 can be decomposed into composite task 1.3.1 and independent task 1.3.2; independent task 1.3.2 can be processed by calling tool 3; composite task 1.3.1 needs to continue to be decomposed to obtain independent tasks 1.3.1.1 and 1.3.1.2; independent task 1.3.1.1 can be processed by calling tool 4; independent task 1.3.1.2 can be processed by calling tool 5.
[0208] It is worth noting that the calling processes of the above-mentioned tasks 1.1, 1.2.1, 1.3.2, 1.3.1.1 and 1.3.1.2 have no data dependency relationship with each other, so they can be implemented in parallel to improve the computing efficiency. However, the processing of task 1.2 depends on the processing result of task 1.2.1, the processing of task 1.3.1 depends on the processing results of tasks 1.3.1.1 and 1.3.1.2, and the processing of task 1.3 depends on the processing results of tasks 1.3.1 and 1.3.2, so they can only be implemented in series to ensure the smooth operation of question-answering processing.
[0209] For example, in the compound task "<|User|>Remind me to watch the Olympic Shooting Finals", through analysis, we can know that its intention is "<|Analyze|>The user wants to set an alarm reminder to watch the Olympic Shooting Finals, and I need to query the specific time of this game". Accordingly, an independent task is disassembled, and the time of the Olympic Shooting Finals can be queried by directly calling the web search tool "<|Tools|>{"tools":"web.search","query":"Time of the Olympic Shooting Finals"}". After calling the corresponding tool, the following return result is obtained "<|Return|>xxxx Viewing Guide-xxx.com: xxx Shooting xxx Finals will be held at 3 pm on August 2 xxx". Analyze the intent again and break down an independent task, whose intent is as follows "<|Analyze|>According to the search results, the alarm needs to be set at 3 pm on August 2." Accordingly, you can set the alarm by calling the alarm setting tool "<|tools|>{"tools":"alarm","params":{"time":"2024-08-02 15:00:00","name":"Watch the Olympic Shooting Finals"}}". Furthermore, after the execution is successful, feedback is given "<|Return|>Success". Analyze the intent again and learn that "<|Analyze|>According to the returned results, the alarm has been set successfully". Correspondingly, the assistant role feedbacks the following content "<|Assistant|>The finals will be held at 3 pm on August 2, and an alarm reminder has been set for you". Finally, the composite task is executed.
[0210] It is understandable that since the calling of the alarm setting tool needs to rely on the search results of the web search tool, the processing of the two can only be executed serially and cannot be implemented in parallel.
[0211] It is worth noting that the Analyze role after the Return role also undertakes the role of result analysis. When the tool call is unreasonable or the return does not meet expectations, the Analyze role will not be able to draw the correct conclusion and will re-decompose the intent and call the tool, thereby ensuring the rationality of the results of each intent decomposition and effectively guaranteeing the execution of complex tasks.
[0212] Based on the same inventive concept, in some embodiments, a model training device for implementing the model training method involved above is also provided. The implementation solution provided by the device to solve the problem is similar to the implementation solution recorded in the above method, so the specific limitations in one or more model training device embodiments provided below can refer to the limitations on the model training method above, and will not be repeated here.
[0213] In some embodiments, Fig. 9 As shown, a model training device is provided, including: a first acquisition module 910, a first calling module 920, a first generation module 930, a first update module 940 and a model training module 950. Among them:
[0214] A first acquisition module 910 is used to acquire at least one first question-and-answer sentence; the first question-and-answer sentence is used to instruct question-and-answer processing based on at least two single application tools;
[0215] The first calling module 920 is used to call the intention disassembly model for each first question and answer sentence, so as to cascade disassemble the first question and answer sentence until there is no compound intention sentence in the disassembly result; the compound intention sentence cannot be processed for question and answer based on a single application tool;
[0216] A first generating module 930, configured to generate a derived intent statement according to the decomposition result corresponding to each first question-answer statement and the decomposition relationship between different intent statements in the corresponding decomposition result;
[0217] A first updating module 940 is used to update the decomposition results of each first question and answer statement according to the derived intention statement;
[0218] The model training module 950 is used to train the intention disassembly model using the updated disassembly results as training samples.
[0219] In some embodiments, the first generation module 930 includes: a construction unit, which is used to construct an intention decomposition tree corresponding to the first question and answer statement for each first question and answer statement according to the decomposition result corresponding to the first question and answer statement and the decomposition relationship between different intention statements in the corresponding decomposition result; the root node in the intention decomposition tree corresponds to the first question and answer statement; the child nodes in the intention decomposition tree correspond to the intention statements in the decomposition result; the intention statements corresponding to the child nodes in the intention decomposition tree are the statement decomposition results of the intention statements corresponding to the corresponding parent nodes; a selection unit, which is used to select a target subtree according to the frequency of occurrence of each subtree in the intention decomposition tree corresponding to at least one first question and answer statement; a generation unit, which is used to use the intention statement corresponding to the root node in the target subtree as a derived intention statement.
[0220] In some embodiments, the selection unit is specifically used to select a subtree whose occurrence frequency is greater than a preset frequency threshold as a target subtree from the intention decomposition tree corresponding to at least one first question and answer statement.
[0221] Each module in the above-mentioned model training device can be implemented in whole or in part by software, hardware and a combination thereof. Each of the above-mentioned modules can be embedded in or independent of the processor in the computer device in the form of hardware, or can be stored in the memory in the computer device in the form of software, so that the processor can call and execute the operations corresponding to each of the above modules.
[0222] Based on the same inventive concept, in some embodiments, a model using device for implementing the model using method involved above is also provided. The implementation scheme for solving the problem provided by the device is similar to the implementation scheme recorded in the above method, so the specific limitations in one or more embodiments of the model using device provided below can refer to the limitations of the model using method above, and will not be repeated here.
[0223] In other embodiments, Fig.10 As shown, a model using device is provided, including: a second acquisition module 1010, a second calling module 1020 and a first operation module 1030. Among them:
[0224] The second acquisition module 1010 is used to acquire a second question and answer statement; the second question and answer statement is used to instruct to perform question and answer processing based on at least two single application tools;
[0225] The second calling module 1020 is used to call the intention disassembly model to perform cascade disassembly on the second question-answering task until there is no composite intention sentence in the disassembly result; the intention disassembly model is obtained by training based on the device of claim 8;
[0226] The first running module 1030 is used to run the application tool matching the single intent statement in the display device when a single intent statement is disassembled, so as to perform question and answer processing on the corresponding single intent statement; the single intent statement can be processed with questions and answers based on a single application tool.
[0227] In some embodiments, the second calling module 1020 is also used to call the intention disassembly model and re-cascade disassembly the second question and answer statement when there is a failed question and answer processing result; wherein the re-disassembly process of the second question and answer statement is at least partially different from the previous disassembly process of the second question and answer statement.
[0228] In some embodiments, the device also includes: a parsing module, which is used to perform semantic parsing on the second question and answer statement to obtain intent analysis data of the second question and answer statement; a processing module, which is used to use the intent analysis data as a system prompt word for the intent statement in the disassembly result of the second question and answer task, and perform question and answer processing on the corresponding intent statement.
[0229] In some embodiments, the device also includes: a third acquisition module, which is used to acquire at least two target application tools corresponding to the derived intent statement when the derived intent statement is disassembled; the target application tool is used to perform question and answer processing on the single intent statement in the disassembly result of the derived intent statement; a second running module, which is used to run each target application tool in the display device in a target order to perform question and answer processing on the derived intent statement; the target order is opposite to the generation order of each single intent statement in the disassembly result of the derived intent statement.
[0230] Each module in the above model using device can be implemented in whole or in part by software, hardware and their combination. Each module can be embedded in or independent of the processor in the computer device in the form of hardware, or can be stored in the memory of the computer device in the form of software, so that the processor can call and execute the operations corresponding to each module above.
[0231] In an exemplary embodiment, a computer device is provided. The computer device may be a terminal, and its internal structure diagram may be as shown in FIG. Fig.11 As shown. The computer device includes a processor, a memory, an input / output interface, a communication interface, a display unit and an input device. Among them, the processor, the memory and the input / output interface are connected through a system bus, and the communication interface, the display unit and the input device are connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The input / output interface of the computer device is used to exchange information between the processor and an external device. The communication interface of the computer device is used to communicate with an external terminal in a wired or wireless manner, and the wireless manner can be implemented through WIFI, a mobile cellular network, near field communication (Near Field Communication, NFC) or other technologies. When the computer program is executed by the processor, a model training and / or use method is implemented. The display unit of the computer device is used to form a visually visible picture, which can be a display screen, a projection device or a virtual reality imaging device. The display screen can be a liquid crystal display screen or an electronic ink display screen, and the input device of the computer device can be a touch layer covering the display screen, or a button, trackball or touchpad set on the computer device shell, or an external keyboard, touchpad or mouse.
[0232] Those skilled in the art will understand that Fig.11The structure shown in the figure is only a block diagram of a part of the structure related to the present embodiment, and does not constitute a limitation on the computer device to which the present embodiment is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.
[0233] In an alternative embodiment, Fig.11 The computer device shown may be the aforementioned display device or server.
[0234] In an exemplary embodiment, a computer device is provided, which may be, for example, a display device or a server, including a memory and a processor, wherein a computer program is stored in the memory, and when the processor executes the computer program, the model training and / or use methods provided in each optional embodiment are implemented.
[0235] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the model training and / or use methods provided in each optional embodiment are implemented.
[0236] In one embodiment, a computer program product is provided, including a computer program, which, when executed by a processor, implements the model training and / or use methods provided in various optional embodiments.
[0237] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this embodiment are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data need to comply with relevant regulations.
[0238] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiments can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to the memory, database or other medium used in the embodiments provided in this embodiment can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. As an illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM). The database involved in each embodiment provided in this embodiment may include at least one of a relational database and a non-relational database. Non-relational databases may include distributed databases based on blockchains, etc., but are not limited to this. The processor involved in each embodiment provided in this embodiment may be a general-purpose processor, a central processing unit, a graphics processor, a digital signal processor, a programmable logic device, a data processing logic device based on quantum computing, an artificial intelligence (AI) processor, etc., but are not limited to this.
[0239] The technical features of the above embodiments may be arbitrarily combined. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of the description of this embodiment.
[0240] The above embodiments only express several implementation methods of this embodiment, and the descriptions are relatively specific and detailed, but they cannot be understood as limiting the scope of the patent of this embodiment. It should be pointed out that for ordinary technicians in this field, several modifications and improvements can be made without departing from the concept of this embodiment, which all belong to the protection scope of this embodiment. Therefore, the protection scope of this embodiment shall be based on the attached claims.
Claims
1. A model training method, characterized in that: include: Acquire at least one first question-and-answer statement; the first question-and-answer statement is used to instruct question-and-answer processing to be performed based on at least two single application tools; For each first question and answer statement, call the intention disassembly model to cascade disassemble the first question and answer statement until there is no composite intention statement in the disassembly result; the composite intention statement cannot be processed for question and answer based on a single application tool; Generate a derived intent statement according to the decomposition result corresponding to each first question-answer statement and the decomposition relationship between different intent statements in the corresponding decomposition result; According to the derived intention statements, updating the decomposition results of each of the first question and answer statements; The intention disassembly model is trained using the updated disassembly results as training samples.
2. The method according to claim 1, characterized in that: The generating of the derived intention statement according to the decomposition result corresponding to each first question-and-answer statement and the decomposition relationship between different statements in the corresponding decomposition result includes: For each first question and answer statement, according to the disassembly result corresponding to the first question and answer statement and the disassembly relationship between different intention statements in the corresponding disassembly result, construct an intention disassembly tree corresponding to the first question and answer statement; the root node in the intention disassembly tree corresponds to the first question and answer statement; the child nodes in the intention disassembly tree correspond to the intention statements in the disassembly result; the intention statements corresponding to the child nodes in the intention disassembly tree are the sentence disassembly results of the intention statements corresponding to the corresponding parent nodes; Selecting a target subtree according to the occurrence frequency of each subtree in the intention decomposition tree corresponding to the at least one first question-and-answer sentence; The intention statement corresponding to the root node in the target subtree is used as the derived intention statement.
3. The method according to claim 2, characterized in that The selecting a target subtree according to the occurrence frequency of each subtree in the intention decomposition tree corresponding to the at least one first question-and-answer sentence includes: From the intention decomposition tree corresponding to the at least one first question and answer statement, a subtree whose occurrence frequency is greater than a preset frequency threshold is selected as the target subtree.
4. A method for using a model, characterized in that: include: Acquire a second question-and-answer statement; the second question-and-answer statement is used to instruct to perform question-and-answer processing based on at least two single application tools; Calling the intention disassembly model to perform cascade disassembly on the second question-answering task until there are no more compound intention sentences in the disassembly result; the intention disassembly model is trained based on the method described in any one of claims 1 to 3; When a single intent statement is disassembled, an application tool matching the single intent statement in the display device is run to perform question-answering processing on the corresponding single intent statement; The single intent statement enables question-answering processing based on a single application tool.
5. The method according to claim 4, characterized in that The method further comprises: In the case where there is a question-and-answer processing result of execution failure, calling the intention disassembly model to re-perform cascade disassembly on the second question-and-answer statement; Among them, the re-decomposition process of the second question and answer statement is at least partially different from the previous decomposition process of the second question and answer statement.
6. The method according to claim 4, characterized in that The method further comprises: Performing semantic analysis on the second question and answer sentence to obtain intention analysis data of the second question and answer sentence; For the intent sentences in the decomposition results of the second question-and-answer task, the intent analysis data is used as a system prompt word to perform question-and-answer processing on the corresponding intent sentences.
7. The method according to any one of claims 4 to 6, characterized in that: The method further comprises: When the derived intent statement is deconstructed, at least two target application tools corresponding to the derived intent statement are obtained; the target application tools are used to perform question-answering processing on a single intent statement in the deconstructed result of the derived intent statement; In accordance with a target order, each of the target application tools in the display device is run to perform question-and-answer processing on the derived intent statement; the target order is opposite to the generation order of each single intent statement in the decomposition result of the derived intent statement.
8. A model training device, characterized in that: include: A first acquisition module is used to acquire at least one first question and answer sentence; the first question and answer sentence is used to instruct question and answer processing based on at least two single application tools; A first calling module is used to call the intention disassembly model for each first question and answer sentence, so as to cascade disassemble the first question and answer sentence until there is no composite intention sentence in the disassembly result; the composite intention sentence cannot be processed for question and answer based on a single application tool; A first generation module is used to generate a derived intent statement according to the disassembly result corresponding to each first question and answer statement and the disassembly relationship between different intent statements in the corresponding disassembly result; A first updating module, configured to update the decomposition results of each of the first question-and-answer statements according to the derived intention statements; The model training module is used to train the intention disassembly model using the updated disassembly results as training samples.
9. A model using device, characterized in that: include: A second acquisition module is used to acquire a second question and answer statement; the second question and answer statement is used to instruct to perform question and answer processing based on at least two single application tools; A second calling module is used to call the intention disassembly model to perform cascade disassembly on the second question-answering task until there are no more compound intention sentences in the disassembly result; the intention disassembly model is trained based on the device described in claim 8; A first running module is used to run an application tool matching the single intent statement in a display device when a single intent statement is disassembled, so as to perform question-answering processing on the corresponding single intent statement; The single intent statement enables question-answering processing based on a single application tool.
10. A server, characterized in that: include: A communication device configured to be communicatively connected with the display device; and at least one processor connected to the communication device and configured to execute the steps of the method according to any one of claims 1 to 7.