Natural Language Generation Method, Apparatus, and Storage Medium
By encoding and decoding the system action text, combining slot judgment and neural network, the problem of semantic information loss in task-based dialogue systems is solved, and the accuracy and stability of natural language generation is improved.
Patent Information
- Application Number
- CN202110213834.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-02-25
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2041-02-25
AI Technical Summary
The existing natural language generation technology has the problem of semantic information loss in task-based dialogue systems, especially keyword loss, which leads to inaccurate and instability of the generated natural language.
By encoding the system action text, generating an encoded vector, and determining whether the decoded text meets the decoded end condition during the decoding process, ensuring that the decoded text contains all preset slots or slot values in the system action text template, using a neural network to implement encoding and decoding, combining phrase vectors and attention information to guide the decoding process.
It improves the accuracy and stability of natural language generation, ensures that key information is fully expressed during the decoding process, avoids the loss of semantic information, and the generated natural language is more compact and diversified.
Smart Images

Figure CN114970555B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and in particular, to a natural language generation method, apparatus, and storage medium. Background Art
[0002] As a business assistant in the vertical field, the task-based dialogue system can efficiently process cumbersome and repetitive high-frequency questions and answers, and then complete the user's target tasks, such as booking tickets, querying the weather, etc.
[0003] Natural language generation (NLG) is an important part of the task-based dialogue system. Natural language generation aims to convert the semantic expressions (meaning representations, MR) in the machine representation system into natural language for people to understand. However, existing natural language generation technologies usually have the problem of semantic information loss (such as keyword loss). Summary of the Invention
[0004] In view of this, a natural language generation method, apparatus, and storage medium are proposed.
[0005] In a first aspect, an embodiment of this application provides a natural language generation method applied to a dialogue system. The method includes: encoding a system action text to obtain an encoded vector, where the system action text is generated according to user query information and a preset system action text template, the system action text template includes at least one preset slot, and the system action text includes at least one slot value; decoding the encoded vector, and determining whether the decoded text meets a decoding end condition during the decoding process, where the decoded text is the text obtained by decoding the encoded vector, and the decoding end condition is that the decoded text includes all the preset slots in the system action text template, or the decoded text includes all the slot values in the system action text; and when the decoded text meets the decoding end condition, determining the decoded text as the natural language text for replying to the user query information.
[0006] The natural language generation method according to an embodiment of the present application is applied to a dialogue system. It can encode a system action text generated according to user query information and a preset system action text template to obtain an encoded vector, and decode the encoded vector. During the decoding process, it is determined whether the decoded text meets the decoding end condition. The decoding end condition is that the decoded text includes all preset slots in the system action text template or the decoded text includes all the slot values in the system action text. When the decoded text meets the decoding end condition, the decoded text is determined as the natural language text for replying to the user query information. Thus, when the dialogue system generates natural language, it can guide the decoding process of the encoded vector according to the constraint information (such as the included slot values) implicit in the system action text or according to the constraint information (such as the included slots) implicit in the system action text template, so that the key information in the system action text is fully expressed during the decoding process, avoiding the loss of semantic information (such as the loss of keywords), and further improving the accuracy and stability of natural language generation.
[0007] According to the first aspect, in the first possible implementation manner of the natural language generation method, the encoding of the system action text to obtain an encoded vector specifically includes: encoding each word in the system action text according to a preset dictionary to obtain a word vector; determining a target phrase from the system action text according to a preset phrase library, and encoding the target phrase to obtain a phrase vector, where the phrase library includes multiple phrases, and the target phrase is a phrase that is simultaneously included in the phrase library and the system action text; fusing the word vector and the phrase vector to obtain the encoded vector.
[0008] In this embodiment, by adding the phrase vector to the encoded vector, the encoded vector includes vectors of multiple granularities (word granularity, phrase granularity), so as to improve the accuracy of the encoded vector in expressing the key semantic information in the system action text, and further improve the accuracy and stability of the decoded natural language text.
[0009] According to the first possible implementation manner of the first aspect, in the second possible implementation manner of the natural language generation method, the encoding of the target phrase to obtain a phrase vector specifically includes: averaging or weighted averaging the word vectors corresponding to the words in the target phrase to obtain a phrase vector.
[0010] In this embodiment, by averaging or weighted averaging the word vectors corresponding to the words in the target phrase to obtain a phrase vector, not only is the calculation convenient, but also the encoder parameters are not increased, so as to improve the processing efficiency.
[0011] According to the first possible implementation manner or the second possible implementation manner of the first aspect, in the third possible implementation manner of the natural language generation method, the step of fusing the word vector and the phrase vector to obtain the encoding vector specifically includes: concatenating the word vector and the phrase vector to obtain the encoding vector.
[0012] In this embodiment, by means of concatenation, the word vector and the phrase vector are fused to obtain the encoding vector, so that the encoding vector includes vectors of multiple granularities (word granularity, phrase granularity), thereby improving the accuracy of the encoding vector in expressing the key semantic information in the system action text.
[0013] According to the first aspect or any one of the first to third possible implementation manners of the first aspect, in the fourth possible implementation manner of the natural language generation method, the system action text template includes at least one short text template, each short text template includes at least one slot, the system action text includes at least one short text, and each short text includes at least one slot value. The step of determining whether the decoded text meets the decoding end condition during the decoding process specifically includes: determining whether the decoded text meets a first condition during the decoding process, where the first condition is that the decoded text includes at least one slot in the system action text template, or the decoded text includes at least one slot value in the system action text; when the decoded text meets the first condition, determining the target short text template to which the slot belongs, or determining the target short text to which the slot value belongs; determining whether the decoded text meets a second condition, where the second condition is that the decoded text includes all slots in the target short text template, or the decoded text includes all slot values in the system action text; when the decoded text meets the second condition, determining that the decoded text does not meet the decoding end condition.
[0014] In this embodiment, by judging whether the decoded text includes all slots in the target short text module, or by judging whether the decoded text includes all slot values in the target short text, it is possible to prevent slot information misalignment, making the expression of the natural language text generated for replying to the user's query information more compact.
[0015] According to the fourth possible implementation manner of the first aspect, in the fifth possible implementation manner of the natural language generation method, the step of determining whether the decoded text meets the decoding end condition during the decoding process further includes: when the decoded text meets the second condition, determining whether the decoded text meets the third condition, where the third condition is that the decoded text includes all the preset slots in the system action text template, or the decoded text includes all the slot values in the system action text; when the decoded text does not meet the third condition, determining that the decoded text does not meet the decoding end condition.
[0016] In this embodiment, when the decoded text includes all the slots in the target short text module or the decoded text includes all the slot values in the target short text, by determining whether the decoded text includes all the slots in the system action text template or by determining whether the decoded text includes all the slot values in the system action text, all the short texts in the system action text can be involved in the decoding, so that the expression accuracy of the key semantic information in the system action text in the generated natural language text can be improved, and further the accuracy and stability of natural language generation can be improved.
[0017] According to the fifth possible implementation manner of the first aspect, in the sixth possible implementation manner of the natural language generation method, the step of determining whether the decoded text meets the decoding end condition during the decoding process further includes: when the decoded text meets the third condition, determining that the decoded text meets the decoding end condition.
[0018] In this embodiment, the decoded text that meets the third condition also meets the first condition and the second condition. Determining that the decoded text that meets the third condition meets the decoding end condition can not only prevent slot information misalignment during the decoding process, making the expression of the natural language text generated for replying to the user's query information more compact, but also enable all the short texts in the system action text to be involved in the decoding, so that the expression accuracy of the key semantic information in the system action text in the generated natural language text can be improved, and further the accuracy and stability of natural language generation can be improved.
[0019] According to the first aspect or any one of the first to sixth possible implementation manners of the first aspect, in the seventh possible implementation manner of the natural language generation method, the method further includes: determining a system action according to user query information, where the system action is an action executed when the dialogue system replies to the user query information, and the system action includes an action category and at least one slot-slot value pair; determining a system action text according to the system action and a system action text template, where the system action text template corresponds to the system action, and the slots in the system action text template are the same as the slots in the system action.
[0020] In this embodiment, by determining the system action according to the user query information and determining the system action text according to the system action and the system action text module, the processing efficiency when generating the system action text can be improved.
[0021] According to the seventh possible implementation manner of the first aspect, in the eighth possible implementation manner of the natural language generation method, the step of determining the system action text according to the system action and the system action text template specifically includes: filling the slot values in the system action into the corresponding slots of the system action text template to obtain the system action text.
[0022] In this embodiment, by filling the slot values in the system action into the corresponding slots of the system action text template to obtain the system action text, it is simple, convenient and not easy to make mistakes. It can not only improve the processing efficiency, but also improve the accuracy of the system action text.
[0023] According to the first aspect, in the ninth possible implementation manner of the natural language generation method, the decoding of the encoded vector includes: determining attention information according to the encoded vector; decoding the encoded vector according to the attention information.
[0024] In this embodiment, by obtaining the attention information and combining the attention information with the decoding process, key semantic information and common phrases (i.e., common text segments) can be more fully utilized in the decoding process, thereby improving the accuracy of natural language generation.
[0025] According to the first aspect or any one of the first to ninth possible implementation manners of the first aspect, in the tenth possible implementation manner of the natural language generation method, the encoding is completed by an encoder, the decoding is completed by a decoder, and the encoder and the decoder are neural networks.
[0026] In this embodiment, encoding is performed by an encoder implemented by a neural network and decoding is performed by a decoder implemented by a neural network, which can make full use of the powerful computing and processing capabilities of the neural network, thereby improving the processing efficiency during natural language generation.
[0027] According to the tenth possible implementation manner of the first aspect, in the eleventh possible implementation manner of the natural language generation method, the encoder is a bidirectional long short-term memory network, and the decoder is a long short-term memory network.
[0028] In this embodiment, the encoder is a bidirectional long short-term memory network and the decoder is a long short-term memory network. Both the encoder and the decoder use a sequence-based model, which can improve the diversity of the generated natural language text.
[0029] In a second aspect, an embodiment of the present application provides a natural language generation device, including a processor and a memory for storing processor-executable instructions. Wherein, the processor is configured to implement the natural language generation method according to one or several of the above first aspect or various possible implementation manners of the first aspect when executing the instructions.
[0030] The natural language generation device according to the embodiment of the present application can encode the system action text generated according to the user query information and the preset system action text template to obtain an encoded vector, and decode the encoded vector. During the decoding process, it is judged whether the decoded text meets the decoding end condition. The decoding end condition is that the decoded text includes all the preset slots in the system action text template or the decoded text includes all the slot values in the system action text; when the decoded text meets the decoding end condition, it is determined that the decoded text is the natural language text for replying to the user query information, so that when the dialogue system generates natural language, it can guide the decoding process of the encoded vector according to the constraint information (such as the included slot values) implicit in the system action text or according to the constraint information (such as the included slots) implicit in the system action text template, so that the key information in the system action text is fully expressed during the decoding process, avoiding loss of semantic information (such as loss of keywords), and further improving the accuracy and stability of natural language generation.
[0031] In a third aspect, an embodiment of the present application provides a non-volatile computer-readable storage medium, on which computer program instructions are stored. When the computer program instructions are executed by a processor, the natural language generation method according to one or several of the above first aspect or various possible implementation manners of the first aspect is implemented.
[0032] Embodiments of the present application can encode a system action text generated according to user query information and a preset system action text template to obtain an encoded vector, and decode the encoded vector. During the decoding process, it is determined whether the decoded text meets the decoding end condition, where the decoding end condition is that the decoded text includes all the preset slots in the system action text template or the decoded text includes all the slot values in the system action text; when the decoded text meets the decoding end condition, it is determined that the decoded text is a natural language text for replying to the user query information, so that when the dialogue system generates natural language, it can guide the decoding process of the encoded vector according to the constraint information (such as the included slot values) implicit in the system action text or according to the constraint information (such as the included slots) implicit in the system action text template, enabling the key information in the system action text to be fully expressed during the decoding process, avoiding loss of semantic information (such as keyword loss), and thus improving the accuracy and stability of natural language generation.
[0033] In a fourth aspect, embodiments of the present application provide a computer program product, including computer-readable code or a non-volatile computer-readable storage medium carrying the computer-readable code. When the computer-readable code runs on an electronic device, the processor in the electronic device executes the natural language generation method according to any one or several of the above first aspect or various possible implementations of the first aspect.
[0034] Embodiments of the present application can encode a system action text generated according to user query information and a preset system action text template to obtain an encoded vector, and decode the encoded vector. During the decoding process, it is determined whether the decoded text meets the decoding end condition, where the decoding end condition is that the decoded text includes all the preset slots in the system action text template or the decoded text includes all the slot values in the system action text; when the decoded text meets the decoding end condition, it is determined that the decoded text is a natural language text for replying to the user query information, so that when the dialogue system generates natural language, it can guide the decoding process of the encoded vector according to the constraint information (such as the included slot values) implicit in the system action text or according to the constraint information (such as the included slots) implicit in the system action text template, enabling the key information in the system action text to be fully expressed during the decoding process, avoiding loss of semantic information (such as keyword loss), and thus improving the accuracy and stability of natural language generation.
[0035] These and other aspects of the present application will become more readily apparent in the following description of the (multiple) embodiments. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] The drawings included in and forming a part of the specification illustrate exemplary embodiments, features, and aspects of the present application together with the specification, and are used to explain the principles of the present application.
[0037] Figure 1 Exemplarily shown is a schematic diagram of a dialogue system provided by an embodiment of the present application.
[0038] Figure 2A Shown is a schematic structural diagram of an electronic device according to an embodiment of the present application.
[0039] Figure 2B Shown is a schematic software structure diagram of an electronic device according to an embodiment of the present application.
[0040] Figure 3 Shown is a schematic diagram of natural language generation based on a generative model according to an embodiment of the present application.
[0041] Figure 4 Shown is a flowchart of a natural language generation method according to an embodiment of the present application.
[0042] Figure 5 Shown is a schematic diagram of an encoding vector in a natural language generation method according to an embodiment of the present application.
[0043] Figure 6 Shown is a schematic diagram of the processing process of a natural language generation method according to an embodiment of the present application.
[0044] Figure 7 Shown is a flowchart of a natural language generation method according to an embodiment of the present application. Detailed implementation manners
[0045] Various exemplary embodiments, features, and aspects of the present application will be described in detail below with reference to the drawings. The same reference numerals in the drawings denote elements having the same or similar functions. Although various aspects of the embodiments are shown in the drawings, the drawings are not necessarily drawn to scale unless otherwise specified.
[0046] The special term "exemplary" here means "serving as an example, embodiment, or illustrative". Any embodiment described as "exemplary" here need not be construed as superior to or better than other embodiments.
[0047] In addition, in order to better illustrate the present application, numerous specific details are given in the following detailed implementation manners. Those skilled in the art should understand that the present application can be implemented without some specific details. In some instances, methods, means, elements, and circuits well-known to those skilled in the art are not described in detail so as to highlight the gist of the present application.
[0048] As a software system capable of conversing with users through natural language, a dialogue system usually needs to recognize and understand the natural language input by users in the forms of text, voice, etc., to determine the user's intention. Then, according to the user's intention, it determines the system actions to be taken, and converts the system actions into natural language to reply to the user or perform corresponding operations.
[0049] Figure 1 Exemplarily shows a schematic diagram of a dialogue system provided by an embodiment of the present application. As Figure 1 shown, user 10 inputs query information in voice mode. The automatic speech recognition (ASR) module 21 in the dialogue system 20 recognizes the query information in voice form input by the user, obtains the query information in text form, and sends it to the natural language understanding (NLU) module 22. Optionally, the query information input by the user can also be in text form, and the dialogue system 20 can directly send the query information in text form to the natural language understanding module 22 for processing. Exemplarily, the above query information can be statements in the form of, for example, "What's the weather like in New York today?", "What movies are showing recently?", etc.
[0050] The natural language understanding module 22 can parse the query information in text form, obtain the intention category to which the user's query belongs and the slot values included in the query information, and send the intention category to which the user's query belongs and the slot values included in the query information to the dialogue state tracking (DST) module 23. Exemplarily, if the query information is "What's the weather like in New York today?", then after calculation and processing, the natural language understanding module 22 can obtain that the intention category is "query weather", and the slots can include "time: today", "city: New York", etc.
[0051] The dialogue state tracking module 23 can calculate the distribution probability of the slot values according to the user's historical query information and the current query information, that is, infer the user's current dialogue state and dialogue goal according to the historical dialogue; then send the dialogue information such as the slot information, dialogue state, and dialogue goal related to the current query to the policy learning (PL) module 24.
[0052] The policy learning module 24 can process the above dialogue information, obtain the system actions to be generated by the dialogue system 20, and send the system actions to the natural language generation module 25, where the system actions include slot information.
[0053] The natural language generation module 25 can decode the above system actions, convert the system actions into corresponding natural languages to reply to the user's query, for example, display the natural language on the display screen and / or play the natural language using a speaker to reply to the user's query; optionally, the system actions can also be sent to the corresponding modules of the operating system to perform corresponding operations, such as calling a contact, opening an application, turning on the camera, playing music, searching using a search engine, etc. Exemplarily, the natural language generation module 25 can generate a natural language reply "It is cloudy in New York today. The highest temperature is 23 degrees Celsius. The lowest temperature is 10 degrees Celsius."
[0054] It should be clear that the dialogue system provided by the embodiments of the present application may include more or fewer modules than Figure 1 the shown dialogue system, and the connection between each module may also be different from Figure 1 the shown connection method. Each module can be implemented in software and / or hardware. The calculation and processing processes of each module can be completed on the cloud side (server side) or on the end side (user device side). It should be clear that any dialogue system that can generate a natural language reply based on the query information is not beyond the scope covered by the present application.
[0055] The embodiments of the present application provide a natural language generation method. Using "template + slot" to generate a natural language reply can make the natural language reply of the dialogue system reasonable. Exemplarily, for the intention category of "querying weather", the natural language generation module 25 can use the preset template "${location} ${time} weather ${weather}. The highest temperature is ${highest temperature} $ degrees Celsius. The lowest temperature is ${lowest temperature} $ degrees Celsius.", and generate a natural language reply by filling in the slots. This template contains the template and a total of five slots: location, time, weather, highest temperature, and lowest temperature. The embodiments of the present application can use "${...}" to schematically represent the slots. The natural language generation module 25 can obtain the "slot - slot value" pair information corresponding to the above slots (i.e., "location: New York", "time: today", "weather: cloudy", "highest temperature: 28", "lowest temperature: 9"), and thus fill the slot values corresponding to the above slots into the corresponding slots in the template to generate a natural language reply "It is cloudy in New York today. The highest temperature is 28 degrees Celsius. The lowest temperature is 9 degrees Celsius.". The above method of using "template + slot" to generate a natural language reply can make the reply presented to the user always have fluency and naturalness on the one hand, but on the other hand, the reply diversity based on the template is relatively low.
[0056] Embodiments of the present application provide a natural language generation method. By using a sequence-to-sequence (seq2seq) model to generate natural language, the natural language responses of the dialogue system can be made diverse. The seq2seq model can be a deep learning model implemented based on a recurrent neural network (RNN), a long short-term memory (LSTM) network, etc. Optionally, the seq2seq model can also be combined with an attention mechanism, enabling the model to focus on the information that the user pays attention to in the user query, making the generated natural language response closer to the response expected by the user. The natural language responses generated by this method are, on the one hand, relatively accurate and usually have good diversity. However, on the other hand, the controllability of the generation process (or decoding process) is relatively low, sometimes resulting in phenomena such as repetition, meaningless words, or loss of keywords in the generated natural language, that is, the stability of natural language generation is relatively low.
[0057] In view of this, embodiments of the present application provide a natural language generation method. By restricting the decoding process through a decoding end condition, semantic information loss (such as keyword loss) can be avoided, and the stability of natural language generation in the dialogue system can be improved. Optionally, the natural language generation method provided by the embodiments of the present application can add a phrase vector to the encoded vector and restrict the decoding process through a decoding end condition to improve the accuracy and stability of natural language generation. That is to say, the natural language generation method of the embodiments of the present application can be a natural language generation method based on restricted decoding.
[0058] The natural language generation method of the embodiments of the present application can be applied to a dialogue system. The dialogue system can output natural language to reply to the query information input by the user. The dialogue system can include a task-based dialogue system, a chit-chat dialogue system, a retrieval-based dialogue system, etc. It should be understood that the specific type and application scenario of the dialogue system are not limited in the present application.
[0059] In a possible implementation, the dialogue system can be applied to an electronic device on the terminal side or to a server. Among them, the electronic device can be touch-screen, non-touch-screen, or without a screen. The touch-screen electronic device can be controlled by clicking, swiping, etc. on the display screen with a finger, a stylus, etc. The non-touch-screen electronic device can be connected to input devices such as a mouse, a keyboard, a touch panel, etc. and be controlled through the input devices. The electronic device without a screen can be, for example, a smart speaker without a screen.
[0060] Figure 2A FIG. 13 shows a schematic structural diagram of an electronic device 100 according to an embodiment of the present application.
[0061] The electronic device 100 may include at least one of a mobile phone, a foldable electronic device, a tablet computer, a desktop computer, a laptop computer, a handheld computer, a notebook computer, an ultra-mobile personal computer (UMPC), a netbook, a cellular phone, a personal digital assistant (PDA), an augmented reality (AR) device, a virtual reality (VR) device, an artificial intelligence (AI) device, a wearable device, a vehicle-mounted device, a smart home device, or a smart city device. Embodiments of the present application do not impose any special restrictions on the specific type of the electronic device 100.
[0062] As Figure 2A shown, the electronic device 100 may include a processor 110, an external memory interface 120, an internal memory 121, a universal serial bus (USB) connector 130, a charging management module 140, a power management module 141, a battery 142, an antenna 1, an antenna 2, a mobile communication module 150, a wireless communication module 160, an audio module 170, a speaker 170A, a receiver 170B, a microphone 170C, a headphone jack 170D, a sensor module 180, a button 190, a motor 191, an indicator 192, a camera 193, a display screen 194, and a subscriber identification module (SIM) card interface 195, etc. The sensor module 180 may include a pressure sensor 180A, a gyroscope sensor 180B, a barometric pressure sensor 180C, a magnetic sensor 180D, an acceleration sensor 180E, a distance sensor 180F, a proximity light sensor 180G, a fingerprint sensor 180H, a temperature sensor 180J, a touch sensor 180K, an ambient light sensor 180L, a bone conduction sensor 180M, etc.
[0063] It can be understood that the structure illustrated in the embodiments of the present application does not constitute a specific limitation on the electronic device 100. In other embodiments of the present application, the electronic device 100 may include more or fewer components than those shown in the figures, or combine certain components, or split certain components, or have different component arrangements. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.
[0064] The processor 110 may include one or more processing units. For example, the processor 110 may include an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU), etc. Among them, different processing units may be independent devices or integrated in one or more processors.
[0065] The processor may generate operation control signals according to the instruction operation code and timing signals to complete the control of fetching and executing instructions.
[0066] In the embodiments of the present application, the processor 110 may be used to execute computer instructions for implementing the functions of each module in the Figure 1 shown dialogue system 20. Optionally, when the dialogue system 20 is implemented by a neural network model (such as a seq2seq model), the processor 110 may include an NPU, and the NPU may be used to perform data operations in the above neural network model to generate a natural language text for replying to user query information.
[0067] A memory may also be provided in the processor 110 for storing instructions and data. In some embodiments, the memory in the processor 110 may be a cache memory. This memory may store instructions or data that have been used by the processor 110 or are used frequently. If the processor 110 needs to use this instruction or data, it can be directly called from this memory. This avoids repeated accesses, reduces the waiting time of the processor 110, and thus improves the efficiency of the system.
[0068] In some embodiments, the processor 110 may include one or more interfaces. The interfaces may include an inter-integrated circuit (I2C) interface, an inter-integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general-purpose input / output (GPIO) interface, a subscriber identity module (SIM) interface, and / or a universal serial bus (USB) interface, etc. The processor 110 may be connected to modules such as a touch sensor, an audio module, a wireless communication module, a display, a camera, etc. through at least one of the above interfaces.
[0069] It can be understood that the interface connection relationships among the modules illustrated in the embodiments of the present application are only illustrative descriptions and do not constitute a structural limitation on the electronic device 100. In some other embodiments of the present application, the electronic device 100 may also adopt different interface connection methods in the above embodiments, or a combination of multiple interface connection methods.
[0070] The wireless communication function of the electronic device 100 may be implemented through antenna 1, antenna 2, a mobile communication module 150, a wireless communication module 160, a modulation and demodulation processor, a baseband processor, etc.
[0071] Antenna 1 and antenna 2 are used to transmit and receive electromagnetic wave signals. Each antenna in the electronic device 100 can be used to cover a single or multiple communication frequency bands. Different antennas can also be multiplexed to improve the utilization rate of the antennas. For example: Antenna 1 can be multiplexed as a diversity antenna for a wireless local area network. In some other embodiments, the antenna can be used in combination with a tuning switch.
[0072] The mobile communication module 150 may provide solutions for wireless communications such as 2G / 3G / 4G / 5G applied to the electronic device 100. The mobile communication module 150 may include at least one filter, switch, power amplifier, low noise amplifier (LNA), etc. The mobile communication module 150 may receive electromagnetic waves through the antenna 1, filter and amplify the received electromagnetic waves, and then transmit them to the modulation and demodulation processor for demodulation. The mobile communication module 150 may also amplify the signal modulated by the modulation and demodulation processor and convert it into electromagnetic waves through the antenna 1 for radiation. In some embodiments, at least some functional modules of the mobile communication module 150 may be provided in the processor 110. In some embodiments, at least some functional modules of the mobile communication module 150 and at least some modules of the processor 110 may be provided in the same device.
[0073] The modulation and demodulation processor may include a modulator and a demodulator. Among them, the modulator is used to modulate the low-frequency baseband signal to be transmitted into a medium-high frequency signal. The demodulator is used to demodulate the received electromagnetic wave signal into a low-frequency baseband signal. Subsequently, the demodulator transmits the demodulated low-frequency baseband signal to the baseband processor for processing. After being processed by the baseband processor, the low-frequency baseband signal is transmitted to the application processor. The application processor outputs a sound signal through an audio device (not limited to the speaker 170A, receiver 170B, etc.), or displays an image or video through the display screen 194. In some embodiments, the modulation and demodulation processor may be an independent device. In other embodiments, the modulation and demodulation processor may be independent of the processor 110 and be provided in the same device as the mobile communication module 150 or other functional modules.
[0074] The wireless communication module 160 may provide wireless communication solutions applied to the electronic device 100, including wireless local area networks (WLANs) (such as wireless fidelity (Wi-Fi) networks), Bluetooth (BT), Bluetooth low energy (BLE), ultra-wide band (UWB), global navigation satellite system (GNSS), frequency modulation (FM), near field communication (NFC), infrared technology (IR), etc. The wireless communication module 160 may be one or more devices integrating at least one communication processing module. The wireless communication module 160 receives electromagnetic waves via the antenna 2, performs frequency modulation and filtering processing on the electromagnetic wave signals, and sends the processed signals to the processor 110. The wireless communication module 160 may also receive the signals to be sent from the processor 110, perform frequency modulation and amplification on them, and convert them into electromagnetic waves through the antenna 2 for radiation.
[0075] In some embodiments, antenna 1 of electronic device 100 is coupled to mobile communication module 150, and antenna 2 is coupled to wireless communication module 160, such that electronic device 100 can communicate with a network and other electronic devices through wireless communication technologies. The wireless communication technologies may include Global System for Mobile Communications (GSM), General Packet Radio Service (GPRS), Code Division Multiple Access (CDMA), Wideband Code Division Multiple Access (WCDMA), Time-Division Code Division Multiple Access (TD-SCDMA), Long Term Evolution (LTE), BT, GNSS, WLAN, NFC, FM, and / or IR technology, etc. The GNSS may include Global Positioning System (GPS), Global Navigation Satellite System (GLONASS), BeiDou Navigation Satellite System (BDS), Quasi-Zenith Satellite System (QZSS), and / or Satellite Based Augmentation Systems (SBAS).
[0076] In the embodiments of the present application, dialogue system 20 can communicate with a network and other electronic devices through the above wireless communication technologies. Exemplarily, when the slot value in the system action of dialogue system 20 needs to be obtained from the network, dialogue system 20 can communicate with the network through mobile communication module 150 and antenna 1 or wireless communication module 160 and antenna 2 to obtain the slot value. Exemplarily, when the user uses a wireless headset (such as a Bluetooth headset), electronic device 100 can be connected to the wireless headset through wireless communication module 160 and antenna 2, so that the user can listen to the reply information of dialogue system 20 (i.e., the natural language text generated by dialogue system 20 for replying to the user's query information) through the wireless headset. Figure 1
[0077] The electronic device 100 can implement the display function through the GPU, the display screen 194, the application processor, etc. The GPU is a microprocessor for image processing, connected to the display screen 194 and the application processor. The GPU is used to execute mathematical and geometric calculations for graphics rendering. The processor 110 may include one or more GPUs, which execute program instructions to generate or change the display information.
[0078] The display screen 194 is used to display images, videos, etc. The display screen 194 includes a display panel. The display panel can adopt a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active matrix organic light-emitting diode (AMOLED), a flexible light-emitting diode (FLED), a Miniled, a MicroLed, a Micro-oLed, a quantum dot light-emitting diode (QLED), etc. In some embodiments, the electronic device 100 may include one or more display screens 194.
[0079] In the embodiments of the present application, the display screen 194 can be used to display the reply information of the dialogue system 20 (i.e., the natural language text generated by the dialogue system 20 for replying to the user's query information) for the user to view.
[0080] The electronic device 100 can implement the audio function through the audio module 170, the speaker 170A, the receiver 170B, the microphone 170C, the headphone jack 170D, and the application processor, etc. Such as music playback, recording, etc.
[0081] The audio module 170 is used to convert the digital audio information into an analog audio signal for output, and is also used to convert the analog audio input into a digital audio signal. The audio module 170 can also be used to encode and decode the audio signal. In some embodiments, the audio module 170 can be set in the processor 110, or some functional modules of the audio module 170 can be set in the processor 110.
[0082] The speaker 170A, also called the "loudspeaker", is used to convert the audio electrical signal into a sound signal. The electronic device 100 can listen to music through the speaker 170A or output the audio signal of a hands-free call. In the embodiments of the present application, the speaker 170A can be used to play the reply information of the dialogue system 20 (i.e., the natural language text generated by the dialogue system 20 for replying to the user's query information) for the user to listen to.
[0083] The receiver 170B, also known as the "earpiece", is used to convert audio electrical signals into sound signals. When the electronic device 100 answers a call or a voice message, the user can listen to the voice by bringing the receiver 170B close to the ear. In the embodiments of the present application, the receiver 170B can be used to play the reply message of the dialogue system 20 (i.e., the natural language text generated by the dialogue system 20 for replying to the user's query message) for the user to listen to.
[0084] The microphone 170C, also known as the "microphone" or "transmitter", is used to convert sound signals into electrical signals. When making a call or sending a voice message, the user can speak by bringing the mouth close to the microphone 170C to input the sound signal into the microphone 170C. The electronic device 100 can be provided with at least one microphone 170C. In some other embodiments, the electronic device 100 can be provided with two microphones 170C, which can not only collect sound signals but also implement a noise reduction function. In some other embodiments, the electronic device 100 can also be provided with three, four or more microphones 170C to collect sound signals, reduce noise, identify the sound source, and implement functions such as directional recording. In the embodiments of the present application, the microphone 170C can be used to record the user's voice to obtain the user's query message, and then the user's query message can be input Figure 1 into the ASR module 21 in the dialogue system 20 shown for speech recognition.
[0085] The headphone jack 170D is used to connect a wired headphone. The headphone jack 170D can be a USB interface 130, or a 3.5mm open mobile terminal platform (OMTP) standard interface, or a cellular telecommunications industry association of the USA (CTIA) standard interface. In the embodiments of the present application, when the user uses a wired headphone, the electronic device can be connected to the wired headphone through the headphone jack 170D so that the user can listen to the reply message of the dialogue system 20 (i.e., the natural language text generated by the dialogue system 20 for replying to the user's query message) through the wired headphone.
[0086] The software system of the electronic device 100 can adopt a layered architecture, an event-driven architecture, a microkernel architecture, a microservices architecture, or a cloud architecture. In the embodiments of the present application, taking the Android system with a layered architecture as an example, the software structure of the electronic device 100 is exemplarily described.
[0087] Figure 2B Shows a schematic diagram of the software structure of the electronic device 100 according to an embodiment of the present application.
[0088] The layered architecture divides software into several layers, each layer having a clear role and division of labor. The layers communicate with each other through software interfaces. In some embodiments, the Android system is divided into five layers, from top to bottom: the application layer, the application framework layer, the Android runtime (ART) and native C / C++ libraries, the Hardware Abstract Layer (HAL), and the kernel layer.
[0089] The application layer may include a series of application packages.
[0090] As Figure 2B shown, the application packages may include applications such as the camera, gallery, calendar, call, map, navigation, WLAN, Bluetooth, music, video, short message, etc. In the embodiments of the present application, the dialogue system may be an independent application in an electronic device, such as a voice assistant in a smart phone, a voice assistant in a smart speaker, etc., and the dialogue system may also be a subsystem in a certain application, such as a voice assistant in a navigation application.
[0091] The application framework layer provides application programming interfaces (APIs) and programming frameworks for the applications in the application layer. The application framework layer includes some predefined functions.
[0092] As Figure 2B shown, the application framework layer may include a window manager, a content provider, a view system, a resource manager, a notification manager, an activity manager, an input manager, etc.
[0093] Among them, the content provider can be used to store and obtain data, and make this data accessible to applications. The data may include video, image, audio, dialed and received calls, browsing history and bookmarks, phone book, etc. In the embodiments of the present application, the content provider can be used to store and obtain data related to the dialogue system such as the user's historical query information and current query information.
[0094] The view system includes visual controls, such as controls for displaying text, controls for displaying pictures, etc. The view system can be used to construct the display interface of an application. The display interface can be composed of one or more views. For example, a display interface including a short message notification icon may include a view for displaying text and a view for displaying pictures. In the embodiments of the present application, the view system can be used to construct the display interface of the dialogue system.
[0095] The notification manager enables an application to display notification information in the status bar. It can be used to convey informative messages, which can automatically disappear after a short stay without user interaction. For example, the notification manager is used to inform that a download is complete, a message reminder, etc. The notification manager can also be a notification that appears in the system top status bar in the form of a chart or scroll bar text, such as a notification of a background-running application, or a notification that appears on the screen in the form of a dialogue window. For example, it can prompt text information in the status bar, emit a prompt tone, vibrate the electronic device, blink the indicator light, etc. In the embodiment of this application, when receiving a reply message from the dialogue system, the notification manager can notify the user by displaying text in the status bar, emitting a prompt tone, vibrating the electronic device, blinking the indicator light, etc., so that the user can view it in time.
[0096] The input manager can provide an Input Manager Service (IMS). The IMS can be used to manage the input of the system, such as touch screen input, key input, sensor input, etc. The IMS retrieves events from the input device node and distributes the events to the appropriate window through interaction with the WMS. In the embodiment of this application, the IMS provided by the input manager can be used to manage the input of the dialogue system, that is, the user's query information.
[0097] The Android Runtime includes a core library and the Android Runtime. The Android Runtime is responsible for converting the source code into machine code. The Android Runtime mainly includes the Ahead-of-Time (AOT) compilation technology and the Just-in-Time (JIT) compilation technology.
[0098] The core library is mainly used to provide the functions of the basic Java class library, such as libraries for basic data structures, mathematics, IO, tools, databases, networks, etc. The core library provides APIs for users to develop Android applications.
[0099] The native C / C++ libraries can include multiple functional modules. For example: surface manager, Media Framework, libc, OpenGL ES, SQLite, Webkit, etc.
[0100] The Hardware Abstraction Layer runs in the user space, encapsulates the kernel layer drivers, and provides call interfaces to the upper layer.
[0101] The kernel layer is the layer between hardware and software. The kernel layer at least includes a display driver, a camera driver, an audio driver, and a sensor driver.
[0102] Next, in combination with the dialogue system, the working processes of the software and hardware of the electronic device 100 will be exemplarily described.
[0103] When a user inputs voice query information through the microphone 170C on the display interface of the dialogue system, the corresponding hardware interrupt is sent to the kernel layer, and the kernel layer sends the user query information to the audio module of the hardware abstraction layer; the audio module converts the analog audio input of the user query information into a digital audio signal; the application framework layer obtains the data audio signal from the hardware abstraction layer and stores it in the input manager; the dialogue system at the program layer calls the interface of the application framework layer to obtain the user query information and generates a natural language text for replying to the user query information; the dialogue system sends the generated natural language text to the audio module of the hardware abstraction layer through the interface of the application framework layer; the audio module converts the digital audio signal into an analog audio signal and starts the audio driver by calling the kernel layer to play the reply information of the dialogue system through the speaker 170A.
[0104] In a possible implementation manner, the dialogue system can also be a terminal-server architecture. The terminal (including the above-mentioned electronic device) and the server jointly complete Figure 1 the functions of each module of the dialogue system shown. For example, the ASR module 21 of the dialogue system 20 can be deployed on the terminal side, and the NLU module 22, DST module 23, PL module 24, and NLG module 25 of the dialogue system 20 can be deployed on the server side. When the terminal receives the user query information in the form of voice (such as weather query, hotel reservation, ticket query and reservation, etc.), the ASR module 21 can identify the user query information in the form of voice to obtain the user query information in the form of text and send the user query information in the form of text to the server. After receiving the user query information sent by the terminal, the server can perform parsing, calculation and other processing on the user query information through the NLU module 22, DST module 23 and PL module 24 to determine the system action that the dialogue system 20 will generate, and according to the system action and the preset system action text template, determine the system action text through the NLG module 25, encode and decode the system action text, generate a natural language text corresponding to the system action text, and send the generated natural language text to the terminal so that the terminal can display or perform voice broadcast of the natural language text as the reply information of the dialogue system 20.
[0105] It should be understood that the above only exemplarily describes the deployment of the dialogue system on the terminal side and the server side. Those skilled in the art can set the specific deployment of the dialogue system on the terminal side and the server side according to the actual situation, and this application does not limit this.
[0106] In one possible implementation, the natural language generation module in the dialogue system can be implemented by a generation model, that is, the natural language generation method of the embodiment of the present application can be applied to the generation model. The generation model can be used to convert the system action text into natural language, and its processing process may include an encoding process for analyzing the input text and a decoding process for generating the output text. The generation model can be an encoder-decoder (Enc-Dec) model, a seq2seq model, or other natural language generation models.
[0107] Figure 3 FIG. 1 is a schematic diagram showing natural language generation based on a generation model according to an embodiment of the present application. Figure 3 As shown, the query information 210 input by the user is: "Is it a good restaurant? Are they expensive?" After receiving the query information 210 input by the user, the dialogue system 220 first parses the query information to determine the system action 221 that needs to be responded to: "OFFER_INTENT(ReserveRestaurant), INFORM(price_range="expensive"), INFORM(rating="4")"; then, according to the system action 221, a short text 222 corresponding to the system action is determined: "Would you like to make a reservation at a restaurant? It is an expensive restaurant. Its rating is 4."; then, the short text 222 corresponding to the system action is input into the generation model 223 for natural language generation, and a natural language 224 corresponding to the system action 222 is obtained: "It is a fancy restaurant with a 4-star rating. Would you like me to reserve a table is there? (This is a 4-star high-end restaurant. Do you want to reserve a table?)".
[0108] It can be seen that the natural language 224 output by the generation model 223 is closer to the natural language in human conversation, more colloquial, and more semantically coherent than the short text 222 input into the generation model 223.
[0109] Among them, the generation model 223 can be a seq2seq model. The seq2seq model can be, for example, a text-to-text transfer transformer (T5) model, etc.
[0110] It should be understood that those skilled in the art can set the generation model according to the actual situation, and this application does not limit it.
[0111] Figure 4 The flowchart of the natural language generation method according to an embodiment of the present application is shown. As Figure 4 shown, the natural language generation method may include the following steps:
[0112] Step S301: Encode the system action text to obtain an encoded vector, where the system action text is generated according to the user query information and a preset system action text template, the system action text template includes at least one preset slot, and the system action text includes at least one slot value.
[0113] In practical applications, a dialogue system (such as a task-oriented dialogue system) usually includes a limited number of system actions (dialogue action, DA) in a vertical domain (i.e., a vertical field) to respond to user input, that is, the system action is an action performed by the dialogue system when replying to user query information. Each system action usually includes an action category and at least one "slot-slot value" pair. Among them, the action category of the system action may include INFORM, REQUEST, ACK, etc. This application does not limit the action category of the system action.
[0114] If there are a limited number of system actions in the dialogue system, a corresponding system action text template can be set for each system action in advance, that is, the system action text template corresponds to the system action. The slots in the system action text template are also the same as the slots in the system action. Table 1 below shows an example of system actions and related texts.
[0115] Table 1 Example of System Actions and Related Texts
[0116]
[0117] In Table 1, in the form of INFORM(slot=value), INFORM indicates the system action, slot indicates the slot, and value indicates the slot value. For example, in INFORM(maxTemp=28), INFORM indicates that the system action is "notify", maxTemp indicates the slot, and "28" indicates the value of the slot maxTemp. INFORM(maxTemp=28) can be used to indicate that the dialog system needs to notify the user that the maximum temperature is 28 degrees.
[0118] As shown in Table 1, the system action (dialogue action) includes an action category and at least one slot-slot value pair. The system action can be determined based on the user query information. For example, the user's input information is voice, and the ASR module can identify the query information corresponding to the user's voice. If the query information is "How is the weather in New York today?", the natural language understanding module can identify that the intent category of the query is "query weather", and obtain the slot value "city: NewYork (city: New York), origin: today (time: today)" contained in the query information. According to the intent category to which the query belongs and the slot value contained in the query information, the dialogue state tracking module and the strategy learning module determine the system action: INFORM (city = New York, origin = Today, weatherId = Slightly Cloudy), INFORM (maxTemp = 28), INFORM (minTemp = 9).
[0119] Among them, the slot values of the weather type (weatherId), the maximum temperature (maxTemp), the minimum temperature (minTemp) and the like can be queried by calling an interface, etc. Optionally, if the user query information is "How is the weather today?", the interface can be called to obtain the geographic location information of the user's device, so as to know which location the user wants to query for today's weather.
[0120] The system action text template (dialogue action text template) includes at least one slot, and the slot type and / or position in the template can be preset; the system action text template can also include a preset text template. For example, in Table 1, "The weather in${city}is${weatherId}${origin}." contains three slots: city (city), weather type (weatherId) and time (origin), and also contains the text template "The weather in...is...".
[0121] The system action text (dialogue action text) can be determined according to the system action and the system action text template. Corresponding to at least one slot in the system action text template, the system action text includes at least one slot value. The system action text can also include a preset text template. For example, the system action text can be constructed by filling the slot values in the system action into the corresponding slots in the system action text template. The system action text shown in Table 1 can be obtained by filling the slot values "New York", "slightly cloudy", "today", "28" and "9" into the slots ${city}, ${weatherId}, ${origin}, ${maxTemp} and ${minTemp} in the system action text template respectively.
[0122] The natural language reference text (reference) is the expected output, which can be used to evaluate the natural language text generated by the natural language generation method of the embodiment of the present application. Specifically, the natural language reference text can be determined by manual definition, based on corpus statistics, etc.; compared with the system action text, the natural language reference text is usually more natural, fluent and coherent, and easier for humans to understand; based on the natural language reference text, the accuracy of the generated natural language text can be evaluated.
[0123] Optionally, the accuracy of the generated natural language text can be evaluated by a bilingual evaluation under study (BLEU) value. Specifically, the natural language text generated by the dialogue system can be compared with a preset natural language reference text, the similarity between the natural language text generated by the dialogue system and the natural language reference text can be calculated, and the BLEU value can be determined based on the similarity. The higher the similarity, the larger the BLEU value, and it can be considered that the accuracy of the natural language text generated by the dialogue system is higher.
[0124] In step S301, the system action text can be encoded to obtain an encoded vector. Optionally, the system action text can be encoded by an encoder to obtain an encoded vector. Among them, the encoder can be, for example, a neural network such as a bi-directional recurrent neural network (BRNN), a long short-term memory network (LSTM), or a bi-directional long short-term memory network (BLSTM). The encoded vector can include a word vector (enc_word_emb) and / or a phrase vector (enc_phrase_emb). The embodiments of the present application do not limit the specific types of the encoder and the encoded vector.
[0125] When encoding the system action text, each word in the system action text can be encoded according to a preset dictionary to obtain a word vector. For example, when encoding the system action text through a bi-directional long short-term memory network BLSTM (including a forward long short-term memory sub-network and a backward long short-term memory sub-network), each word in the system action text can be input into the forward long short-term memory sub-network and the backward long short-term memory sub-network respectively for processing according to the preset dictionary, to obtain the forward hidden vector and the backward hidden vector of each word in the system action text, and then the forward hidden vector and the backward hidden vector of each word are concatenated to obtain the word vector of each word.
[0126] The target phrase can be determined from the system action text according to a preset phrase library, and the target phrase can be encoded to obtain a phrase vector. Among them, the phrase library can include multiple phrases. The phrase library can be preset according to the system action text. Optionally, common phrase expressions (at least two consecutive words or characters) in the system action text can be pre-added to the phrase library, such as the highest temperature, the lowest temperature. Optionally, common phrase expressions in the slot values can also be added to the phrase library, such as New York, slightly cloudy, to expand the coverage of the phrase library and improve the diversity of the phrases in the phrase library. The content and length of the phrases in the phrase library can be adaptively determined according to the content in the system action text and / or the slot values. The present application does not limit the content and length of the phrases.
[0127] The target phrase is a phrase that is included in both the phrase library and the system action text. When determining the target phrase, the target phrase can be determined from the system action text by matching the system action text with a preset phrase library, and the target phrase is encoded to obtain a phrase vector. For example, if the system action text includes phrase A and phrase B, and the phrase library includes phrase B, phrase C, and phrase D, then by matching the system action text with the phrase library, the target phrase is determined to be phrase B.
[0128] When encoding the target phrase, for any target phrase, the phrase vector of the target phrase can be obtained by vector fusion according to the word vectors of the respective words in the target phrase. Optionally, the word vectors corresponding to the respective words in the target phrase can be averaged or weighted averaged to obtain the phrase vector, that is, the phrase vector of the target phrase can be the average vector or weighted average vector of the word vectors of the respective words constituting the target phrase. For example, the phrase vector of the target phrase "the highest temperature" can be the average vector of the word vectors of the three words "the", "highest", and "temperature" (i.e., the sum of the three word vectors divided by three). Another example, the phrase vector of the target phrase "slightly cloudy" can be the weighted average vector of the word vectors of the two words "slightly" and "cloudy". By determining the phrase vector in this way, not only is the calculation convenient, but also the encoder parameters are not increased, thereby improving the processing efficiency.
[0129] The word vectors and phrase vectors determined in the above manner can be fused by means such as addition, dot multiplication, splicing, etc. to obtain an encoded vector. Optionally, the word vector and the phrase vector can be spliced to obtain the encoded vector. For example, the phrase vector can be spliced after the word vector, before the word vector, or interspersed between the word vectors. The embodiments of the present application do not make any limitations in this regard. Figure 5 The schematic diagram of the encoded vector in the natural language generation method according to an embodiment of the present application is shown. As Figure 5 shown, the encoded vector of the system action text "The weather in New York is slightly cloudy today. The highest temperature is 28 degrees. The lowest temperature is 9 degrees." includes the word vector 41 and the phrase vector 42. Among them, the word vector 41 includes the word vectors of the respective words in the system action text, and the phrase vector 42 includes the phrase vectors of 4 target phrases "New York", "slightly cloudy", "the highest temperature", and "the lowest temperature". The phrase vector 42 is spliced after the word vector 41.
[0130] By adding the phrase vector to the encoding vector, the encoding vector includes vectors of multiple granularities (word granularity, phrase granularity), so as to improve the accuracy of the encoding vector in expressing the key semantic information in the system action text, and further improve the accuracy and stability of the decoded natural language text.
[0131] Step S302: Decode the encoding vector, and during the decoding process, determine whether the decoded text meets the decoding end condition, where the decoded text is the text obtained by decoding the encoding vector, and the decoding end condition is: the decoded text includes all the preset slots in the system action text template, or the decoded text includes all the slot values in the system action text.
[0132] In a possible implementation manner, the encoding vector can be decoded by a decoder, where the decoder can be, for example, a neural network such as a bi-directional recurrent neural network (BRNN), a long short-term memory network (LSTM), or a bidirectional long short-term memory network (BLSTM).
[0133] In a possible implementation manner, when decoding the encoding vector, attention information can be determined according to the encoding vector (including word vectors and / or phrase vectors) by means of attention distribution, probability calculation, etc., and then, in combination with the attention information, the encoding vector can be decoded by a decoder (such as the long short-term memory network LSTM). By obtaining the attention information and combining the attention information with the decoding process, the key semantic information and common phrases (i.e., common text segments) can be more fully utilized during the decoding process, thereby improving the accuracy of natural language generation.
[0134] In a possible implementation manner, the decoding end condition may include that the decoded text has all the slot values in the system action text. Optionally, a slot value list can be established according to all the slot values in the system action text to facilitate the judgment of the decoding end condition during the decoding process. When determining whether the decoded text meets the decoding end condition during the decoding process, it can be determined whether the decoded text has all the slot values in the system action text by means of comparison, search, etc. according to the slot value list corresponding to the system action text.
[0135] For example, in the system action text “The weather in New York is slightly cloudytoday.The highest temperature is 28degrees.The lowest temperature is9degrees.”, there are a total of 5 slot values including “New York”, “slightly cloudy”, “today”, “28” and “9”. Then, a slot value list can be established based on these 5 slot values. During the decoding process, for the decoded text, it can be determined whether the decoded text contains the 5 slot values of “New York”, “slightly cloudy”, “today”, “28” and “9” by means of comparison, searching, etc. according to this slot value list. In the case where a certain one or some of these 5 slot values are missing in the decoded text, for example, when the decoded text only contains 3 slot values of “New York”, “slightly cloudy” and “9” and lacks the 2 slot values of “today” and “28”, it can be considered that the decoded text does not meet the decoding end condition. Therefore, decoding needs to continue until the decoded text contains the 5 slot values of “New York”, “slightly cloudy”, “today”, “28” and “9”. In the case where the decoded text contains the 5 slot values of “New York”, “slightly cloudy”, “today”, “28” and “9”, it can be considered that the decoded text meets the decoding end condition, and thus the decoding can be ended.
[0136] In this way, taking that the decoded text includes all the slot values in the system action text as the decoding end condition, it is possible to determine whether the decoded text meets the decoding end condition during the decoding process by judging the slot values included in the decoded text, which is simple and effective and convenient for understanding and implementation.
[0137] In a possible implementation, the decoding end condition may include that the decoded text has all the slots in the system action text template. For example, for the system action text module corresponding to the system action text “The weather in New York is slightly cloudytoday.The highest temperature is 28degrees.The lowest temperature is9degrees.” which is “The weather in${city}is${weatherId}${origin}.The highest temperature is${maxTemp}degrees.The lowest temperatureis${minTemp}degrees.”, it includes a total of 5 slots, namely “${city}”, “${weatherId}”, “${origin}”, “${minTemp}” and “${maxTemp}”. Optionally, a slot list may be established according to all the slots in the system action text template to facilitate the judgment of the decoding end condition during the decoding process.
[0138] In the case where there are multiple slot values for the slots in the system action text template, for multiple system action texts generated according to the same system action text template, although the slot values in each system action text may be different, each system action text corresponds to the same system action text template. Therefore, multiple system action texts generated according to the same system action text template can share a slot list.
[0139] For example, according to the same system action text template “The weather in${city}is${weatherId}${origin}.The highest temperature is${maxTemp}degrees.The lowest temperatureis${minTemp}degrees.”, the following two system action texts are generated:
[0140] “The weather in New York is slightly cloudy today.The highesttemperature is28degrees.The lowest temperature is 9degrees.”
[0141] "The weather in Beijing is sunny today. The highest temperature is 11 degrees. The lowest temperature is 1 degree."
[0142] Although the slot values in the above two system action texts are different, the above two system action texts both correspond to a system action text template. Therefore, the above two system action texts can share the same slot list.
[0143] When determining whether the decoded text meets the decoding end condition during the decoding process, based on the slot list corresponding to the system action text template and the preset correspondence between slots and slot values, through methods such as comparison and search, it can be determined whether the decoded text includes all the slots in the system action text template. Among them, the correspondence between slots and slot values can be set in advance. For example, when presetting the system action text template, the correspondence between slots and slot values can be preset.
[0144] For example, during the process of decoding the encoding vector of the system action text "The weather in New York is slightly cloudy today. The highest temperature is 28 degrees. The lowest temperature is 9 degrees.", there are a total of 5 slots in the system action text template corresponding to this system action text, namely "${city}", "${weatherId}", "${origin}", "${minTemp}", and "${maxTemp}". A slot list can be established based on these 5 slots.
[0145] During the decoding process, based on this slot list and the preset correspondence between slots and slot values, through methods such as search and comparison, it can be determined whether the decoded text includes these 5 slots. If the above 5 slots are not included in the decoded text, for example, if the decoded text only includes 3 slots, namely "${city}", "${origin}", and "${minTemp}", and lacks the 2 slot values of "${weatherId}" and "${maxTemp}", it can be considered that the decoded text does not meet the decoding end condition. Therefore, decoding needs to continue until the decoded text includes the above 5 slots; if the decoded text includes the above 5 slots, it can be considered that the decoded text meets the decoding end condition, and thus decoding can be ended.
[0146] By using all the slots in the system action text template included in the decoded text as the decoding end condition, when determining whether the decoded text meets the decoding end condition, multiple system action texts generated according to the same system action text template can share a slot list, thereby reducing the number of lists and improving the processing efficiency.
[0147] Step S303: When the encoded text meets the decoding end condition, determine the decoded text as the natural language text for replying to the user's query information.
[0148] When the decoded text meets the decoding end condition, decoding can be ended or terminated, the decoded text is determined as the natural language text for replying to the user's query information, and the natural language text is output through text display and / or voice broadcast, etc., for the user to view and / or listen to.
[0149] The embodiments of the present application can encode the system action text generated according to the user's query information and the preset system action text template to obtain an encoded vector, and decode the encoded vector. During the decoding process, it is judged whether the decoded text meets the decoding end condition. The decoding end condition is that the decoded text includes all the preset slots in the system action text template or the decoded text includes all the slot values in the system action text; when the decoded text meets the decoding end condition, the decoded text is determined as the natural language text for replying to the user's query information, so that when the dialogue system generates natural language, it can guide the decoding process of the encoded vector according to the constraint information (such as the included slot values) implicit in the system action text or according to the constraint information (such as the included slots) implicit in the system action text template, so that the key information in the system action text is fully expressed during the decoding process, avoiding the loss of semantic information (such as keyword loss), and further improving the accuracy and stability of natural language generation.
[0150] Figure 6 A schematic diagram showing the processing process of the natural language generation method according to an embodiment of the present application. As Figure 6As shown, the system action text generated according to the user query information "How is the weather in New York today?" and the system action text template is: "The weather in New York is slightly cloudy today. The highest temperature is 28 degrees. The lowest temperature is 9 degrees." According to the preset dictionary, each word in the system action text can be encoded by the encoder 51 to obtain word vectors, and according to the preset phrase library, target phrases (including New York, slightly cloudy, the highest temperature, etc.) can be determined from the system action text, and each target phrase can be encoded by the encoder 52 to obtain phrase vectors. Then, the word vectors and the phrase vectors can be concatenated to obtain an encoded vector. Among them, the encoder 51 and the encoder 52 can be bidirectional long short-term memory networks. The encoder 51 and the encoder 52 can be the same network or two networks, and this application does not make any restrictions on this.
[0151] After obtaining the encoded vector, the decoder 54 can decode the encoded vector. Among them, the decoder 54 can be a unidirectional long short-term memory network. During the decoding process, the attention information can be determined through the attention module 53 according to the encoded vector and the word vector of the current word in the decoded text, and the encoded vector can be decoded in combination with the attention information. Optionally, during the decoding process, the decoding process can also be controlled through the constrained condition in the constraint module 55, that is, it is judged whether the decoded text meets the constrained condition. The constrained condition can include a decoding end condition, and the decoding end condition is: the decoded text includes all the preset slots in the system action text template, or the decoded text includes all the slot values in the system action text.
[0152] Optionally, the constrained condition can also include a coverage mechanism. The coverage mechanism can dynamically adjust the attention information of each word vector and phrase vector at the current moment according to the attention information and the decoded text at the previous moment during the decoding process, so as to reduce the weight of the words that have been generated in the attention information (that is, the words included in the decoded text), thereby reducing the repeated segments in the generated natural language text. It should be understood that the constrained condition can also include other conditions, and this application does not make any restrictions on this.
[0153] When the decoded text meets the restrictive conditions, the decoded text can be determined as the output text 56, serving as the response to the user's query "How is the weather in New York today?".
[0154] It should be understood that the encoder and decoder in this embodiment are only examples, and the encoder and decoder can also be other models or other neural networks. For example, adding phrase vectors and restrictive conditions to the T5 model to obtain an improved T5 model; or neural networks such as fully connected networks, RNNs, BRNNs, LSTMs, BLSTMs, etc. The present application does not limit the specific types of the encoder and decoder.
[0155] The above description only uses the system action text as an English text as an example. The natural language generation method of the embodiments of the present application can also be applicable to other languages, such as Spanish, French, German, Chinese, etc. The present application does not limit this.
[0156] In a possible implementation manner, the system action text template may include at least one short text template, and each short text template may include at least one slot. Correspondingly, the system action text may include at least one short text, and each short text may include at least one slot value.
[0157] For example, the system action text template "The weather in ${city} is ${weatherId} ${origin}. The highest temperature is ${maxTemp} degrees. The lowest temperature is ${minTemp} degrees." includes 3 short text templates, namely: "The weather in ${city} is ${weatherId} ${origin}.", "The highest temperature is ${maxTemp} degrees.", and "The lowest temperature is ${minTemp} degrees.". Among them, the short text template "The weather in ${city} is ${weatherId} ${origin}." includes 3 slots: "${city}", "${weatherId}", and "${origin}"; the short text template "The highest temperature is ${maxTemp} degrees." includes 1 slot: "${maxTemp}"; the short text template "The lowest temperature is ${minTemp} degrees." includes 1 slot: "${minTemp}".
[0158] Accordingly, the system action text "The weather in New York is slightly cloudy today. The highest temperature is 28 degrees. The lowest temperature is 9 degrees." generated according to the system actions "INFORM(city = New York, origin = Today, weatherId = Slightly Cloudy), INFORM(maxTemp = 28), INFORM(minTemp = 9)" and the above system action text template includes 3 short texts, namely: "The weather in New York is slightly cloudy today.", "The highest temperature is 28 degrees.", and "The lowest temperature is 9 degrees.". Among them, the short text "The weather in New York is slightly cloudy today." includes 3 slot values: "New York", "slightly cloudy", and "today"; the short text "The highest temperature is 28 degrees." includes 1 slot value: "28"; the short text "The lowest temperature is 9 degrees." includes 1 short text: "9".
[0159] Figure 7 A flowchart showing a natural language generation method according to an embodiment of the present application. As Figure 7 shown, the natural language generation method of this embodiment includes step S301, step S3021, step S3022, step S3023, step S3024, step S3025, and step S303. Among them, steps S3021 to SS3025 are Figure 4 a possible more refined implementation manner of step S302 in the shown embodiment. By hierarchically judging multiple times whether the decoded text meets the corresponding conditions in steps S3021 to S3025, it is not only possible to prevent slot information misalignment during the decoding process, making the expression of the natural language text generated for replying to the user's query information more compact, but also enabling all short texts in the system action text to participate in the decoding, thereby improving the accuracy of the expression of the key semantic information in the system action text in the generated natural language text, and further improving the accuracy and stability of natural language generation.
[0160] Step S301: Encode the system action text to obtain an encoded vector. Among them, the system action text is generated according to the user query information and a preset system action text template, the system action text template includes at least one preset slot, and the system action text includes at least one slot value. Optionally, Figure 7 Step S301 in the illustrated embodiment is similar to Figure 4 Step S301 in the illustrated embodiment, and no repetitive description will be made here.
[0161] Step S3021: Decode the encoded vector.
[0162] Step S3022: During the decoding process, determine whether the decoded text meets the first condition. Among them, the first condition is that the decoded text includes at least one slot in the system action text template, or the decoded text includes at least one slot value in the system action text.
[0163] In a possible implementation manner, when the first condition is that the decoded text includes at least one slot in the system action text template, when determining whether the decoded text meets the first condition during the decoding process, a slot list can be established according to the system action text template, and according to the slot list and the preset correspondence between slots and slot values, it can be determined whether the decoded text includes at least one slot in the system action text template by means of searching, comparison, etc.
[0164] For example, the system action text template “The weather in${city}is${weatherId}${origin}.The highest temperature is${maxTemp}degrees.The lowest temperature is${minTemp}degrees.” includes 5 slots: “${city}”, “${weatherId}”, “${origin}”, “${minTemp}” and “${maxTemp}”. A slot list can be established according to these 5 slots, and according to the slot list and the preset correspondence between slots and slot values, it can be determined whether the decoded text includes at least one of the 5 slots “${city}”, “${weatherId}”, “${origin}”, “${minTemp}” and “${maxTemp}”. When the decoded text includes at least one of these 5 slots, for example, when the decoded text includes 2 slots “${city}” and “${weatherId}”, it can be considered that the decoded text meets the first condition.
[0165] In a possible implementation, when the first condition is that the decoded text includes at least one slot value in the system action text, when determining whether the decoded text meets the first condition during the decoding process, a slot value list can be established based on the system action text, and based on this slot value list, it can be determined whether the decoded text includes at least one slot value in the system action text by means of searching, comparison, etc.
[0166] For example, in the system action text “The weather in New York is slightly cloudytoday.The highest temperature is 28degrees.The lowest temperature is9degrees.”, there are a total of 5 slot values including “New York”, “slightly cloudy”, “today”, “28” and “9”. Then, a slot value list can be established based on these 5 slot values. During the decoding process, based on this slot value list, it can be determined whether the decoded text includes at least one of the 5 slot values “New York”, “slightly cloudy”, “today”, “28” and “9” by means of searching, comparison, etc. When the decoded text includes at least one of these 5 slot values, for example, when the decoded text includes 2 slot values “New York” and “slightly cloudy”, it can be considered that the decoded text meets the first condition.
[0167] If the decoded text does not meet the first condition, return to execute step S3021 and continue to decode the encoded vector;
[0168] If the decoded text meets the first condition, execute step S3023.
[0169] Step S3023, determine the target short text template to which the slot belongs, or determine the target short text to which the slot value belongs.
[0170] When the first condition is that the decoded text includes at least one slot in the system action text template, when the decoded text meets the first condition, the target short text template to which the slot included in the decoded text belongs can be determined from at least one short text template included in the system action text template.
[0171] For example, the decoded text includes one slot "${city}". From the three short text templates included in the system action text template, namely, "The weather in ${city} is ${weatherId} ${origin}.", "The highest temperature is ${maxTemp} degrees.", and "The lowest temperature is ${minTemp} degrees.", it can be determined that the target short text template to which this one slot "${city}" belongs is: "The weather in ${city} is ${weatherId} ${origin}.".
[0172] In the case where the decoded text includes multiple slots, the target short text template to which each slot belongs can be determined separately. The specific method is similar to the above and will not be described repetitively here.
[0173] In the case where the first condition is that the decoded text includes at least one slot value in the system action text, when the decoded text meets the first condition, the target short text to which the slot value included in the decoded text belongs can be determined from at least one short text included in the system action text.
[0174] For example, the decoded text includes one slot value "New York". From the three short texts included in the system action text, namely, "The weather in New York is slightly cloudy today.", "The highest temperature is 28 degrees.", and "The lowest temperature is 9 degrees.", it can be determined that the target short text to which this one slot value "New York" belongs is: "The weather in New York is slightly cloudy today".
[0175] In the case where the decoded text includes multiple slot values, the target short text to which each slot value belongs can be determined separately. The specific method is similar to the above and will not be described repetitively here.
[0176] Step S3024, determine whether the decoded text meets the second condition, where the second condition is that the decoded text includes all the slots in the target short text template, or the decoded text includes all the slot values in the target short text.
[0177] In a possible implementation, when the second condition is that all slots in the target short text template are included in the decoded text, when determining whether the decoded text meets the second condition, all slots included in the target short text template and all slots included in the decoded text can be determined respectively, and then by means of slot searching, comparison, etc., it is determined whether all slots in the target short text template are included in the decoded text.
[0178] For example, all slots in the target short text template "The weather in ${city} is ${weatherId} ${origin}." are the three slots "${city}", "${weatherId}" and "${origin}"; there is only one slot "${city}" included in the decoded text; through slot comparison, it can be determined that the slots "${weatherId}" and "${origin}" in the target short text template are not included in the decoded text, and it can be considered that the decoded text does not meet the second condition.
[0179] For another example, there are two target short text templates. All slots in the first target short text template "The weather in ${city} is ${weatherId} ${origin}." are the three slots "${city}", "${weatherId}" and "${origin}", and all slots in the second target short text template "The lowest temperature is ${minTemp} degrees." are the one slot "${minTemp}". Then all slots in the target segment text template are the four slots "${city}", "${weatherId}", "${origin}" and "${minTemp}"; the decoded text includes the four slots "${city}", "${weatherId}", "${origin}" and "${minTemp}"; through slot comparison, it can be determined that all slots in the target short text template are included in the decoded text, and it can be considered that the decoded text meets the second condition.
[0180] In a possible implementation, when the second condition is that all slot values in the target short text are included in the decoded text, when determining whether the decoded text meets the second condition, all slot values included in the target short text and all slot values included in the decoded text can be determined respectively, and then by means of slot value searching, comparison, etc., it is determined whether all slot values in the target short text are included in the decoded text.
[0181] For example, all slot values in the target short text "The weather in New York is slightly cloudy today." are the three slot values of "New York", "slightly cloudy", and "today"; the decoded text includes the one slot value of "New York"; by comparing the slot values, it can be determined that the decoded text does not include the slot values of "slightly cloudy" and "today" in the target short text, and it can be considered that the decoded text does not meet the second condition.
[0182] For another example, there are two target short texts. All slot values in the first target short text "The weather in New York is slightly cloudy today." are the three slot values of "New York", "slightly cloudy", and "today", and all slots in the second target short text "The highest temperature is 28 degrees." are the one slot value of "28". Then all slots in the target segment text are the four slot values of "New York", "slightly cloudy", "today", and "28"; the decoded text includes the four slot values of "New York", "slightly cloudy", "today", and "28"; by comparing the slot values, it can be determined that the decoded text includes all slot values in the target short text, and it can be considered that the decoded text meets the second condition.
[0183] It should be understood that the above only gives an exemplary description of judging whether the decoded text meets the second condition. Those skilled in the art can also make judgments in other ways, and this application does not limit this.
[0184] If the decoded text does not meet the second condition, then return to execute step S3021 to continue decoding the encoded vector;
[0185] If the decoded text meets the second condition, then execute step S3025.
[0186] Step S3025, judge whether the decoded text meets the third condition, where the third condition is: the decoded text includes all the preset slots in the system action text template, or the decoded text includes all the slot values in the system action text.
[0187] Optionally, Figure 7 The specific method for judging whether the decoded text meets the third condition in step S3025 in the illustrated embodiment is the same as Figure 4The specific method for determining whether the decoded text meets the decoding end condition in step S302 of the illustrated embodiment is similar, and no repetitive description will be given here.
[0188] In a possible implementation manner, when the decoded text meets the second condition, before executing step S3025, it is also possible to determine whether the decoded text meets the sentence pause symbol generation condition according to a preset sentence pause symbol (i.e., a full stop). If the decoded text meets the sentence pause symbol generation condition, after generating the sentence pause symbol, step S3025 is executed; if the decoded text does not meet the sentence pause symbol generation condition, no sentence pause symbol is generated, and step S3025 is directly executed.
[0189] If the decoded text does not meet the third condition, return to execute step S3021 to continue decoding the encoded vector;
[0190] If the decoded text meets the third condition, execute step S303.
[0191] Step 303, determining that the decoded text is a natural language text for replying to the user's query information.
[0192] Optionally, Figure 7 Step S303 in the illustrated embodiment is similar to Figure 4 Step S303 in the illustrated embodiment, and no repetitive description will be given here.
[0193] The embodiments of the present application can encode the system action text generated according to the user's query information and the preset system action text template to obtain an encoded vector, and decode the encoded vector. During the decoding process, it is sequentially determined whether the decoded text meets the first condition, the second condition, and the third condition. When the decoded text meets all the above conditions, it is determined that the decoded text meets the decoding end condition, and it is determined that the decoded text is a natural language text for replying to the user's query information. Therefore, in the decoding process of the embodiments of the present application, by judging whether the decoded text includes all the slots in the target short text module, or by judging whether the decoded text includes all the slot values in the target short text, it is possible to prevent slot information misalignment during the decoding process, making the expression of the generated natural language text for replying to the user's query information more compact; by judging whether the decoded text includes all the slots in the system action text template, or by judging whether the decoded text includes all the slot values in the system action text, it is possible to make all the short texts in the system action text participate in the decoding, thereby improving the accuracy of the expression of the key semantic information in the system action text in the generated natural language text, and further improving the accuracy and stability of natural language generation.
[0194] The natural language generation method according to the embodiments of the present application can adaptively generate a phrase vector based on the system action text (generated according to the user query information and the preset system action text template). It can not only enable the attention module to better focus on the key semantic information in the encoded vector without adding extra learning parameters, improving the efficiency of the attention module, but also obtain an encoded vector including vectors with multiple granularities. During the process of decoding the encoded vector, the decoding process of the encoded vector can be guided according to the constraint information (such as the included slots) implicit in the system action text, so that the key information in the system action text can be fully expressed during the decoding process, thereby solving problems such as possible semantic information loss and inaccurate expression during the natural language generation process, and improving the accuracy and stability of natural language generation.
[0195] The natural language generation method according to the embodiments of the present application makes the expression of key semantic information more accurate during the natural language generation process by adding a phrase vector to the encoded vector. Thus, it can not only improve the diversity and accuracy of the generated natural language text, but also reduce the dependence of the encoder-decoder model on the training data to a certain extent, improve the stability of the encoder-decoder model, and enable the encoder-decoder model to have a certain degree of generalization ability for new input data, thereby better improving the effect of natural language generation in a dialogue system (such as a task-based dialogue system).
[0196] The natural language generation method according to the embodiments of the present application is applicable not only to various types of dialogue systems, but also to scenarios that require natural language generation in other fields, such as abstract generation, comment generation, machine translation, etc. The present application does not limit this.
[0197] In a possible implementation manner, the natural language text generated according to the embodiments of the present application can also be evaluated by indicators such as the accuracy of the slots in the generated natural language text. Among them, the accuracy of the slots in the generated natural language text can be that the natural language text includes all the slots in the input text; the accuracy of the slots in the generated natural language text can also be that the slots in the same short text in the input text are also in the same sentence in the generated natural language text. Those skilled in the art can set the evaluation indicators of the generated natural language text according to the actual situation, and the present application does not limit this.
[0198] An embodiment of the present application provides a natural language generation device, including: a processor and a memory for storing instructions executable by the processor; wherein, the processor is configured to implement the above method when executing the instructions.
[0199] Embodiments of the present application provide a non-volatile computer-readable storage medium, on which computer program instructions are stored, and when the computer program instructions are executed by a processor, the above-mentioned method is implemented.
[0200] Embodiments of the present application provide a computer program product, including computer-readable code, or a non-volatile computer-readable storage medium carrying computer-readable code. When the computer-readable code runs in a processor of an electronic device, the processor in the electronic device executes the above-mentioned method.
[0201] A computer-readable storage medium may be a tangible device that can hold and store instructions used by an instruction execution device. A computer-readable storage medium may be, for example, but not limited to, an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the above. More specific examples (non-exhaustive list) of computer-readable storage media include: portable computer disks, hard disks, random access memory (RAM), read only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), portable compact disc read-only memory (CD-ROM), digital versatile disc (DVD), memory stick, floppy disk, mechanically encoded device, such as punched cards or raised structures in grooves storing instructions thereon, and any suitable combination of the above.
[0202] The computer-readable program instructions or code described herein can be downloaded from a computer-readable storage medium to various computing / processing devices, or downloaded to an external computer or external storage device through a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network may include copper transmission cables, fiber optic transmissions, wireless transmissions, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions for storage in the computer-readable storage media in each computing / processing device.
[0203] The computer program instructions for performing the operations of this application may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-related instructions, microcode, firmware instructions, state setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages such as Smalltalk, C++, etc., and conventional procedural programming languages such as the "C" language or similar programming languages. The computer-readable program instructions may be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., by using an Internet service provider to connect through the Internet). In some embodiments, by using the state information of the computer-readable program instructions to customize an electronic circuit, such as a programmable logic circuit, a field-programmable gate array (FPGA), or a programmable logic array (PLA), the electronic circuit can execute the computer-readable program instructions to implement various aspects of this application.
[0204] Aspects of the present application are described herein with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present application. It should be understood that each block of the flowcharts and / or block diagrams, and the combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer-readable program instructions.
[0205] These computer-readable program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, thereby producing a machine such that when these instructions are executed by the processor of the computer or other programmable data processing device, a device is produced that implements the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium, and these instructions cause the computer, programmable data processing device, and / or other devices to work in a specific manner. Thus, the computer-readable medium storing the instructions includes a manufactured article that includes instructions for implementing various aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.
[0206] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device, causing a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process such that the instructions executed on the computer, other programmable data processing apparatus, or other device implement the functions / acts specified in one or more boxes of the flowchart and / or block diagram.
[0207] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of apparatus, systems, methods, and computer program products according to various embodiments of the present application. In this regard, each box in the flowchart or block diagram may represent a module, a segment of a program, or a portion of an instruction, which contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions noted in the boxes may occur in a different order than noted in the figures. For example, two consecutive boxes may in fact be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending on the functionality involved.
[0208] It should also be noted that each box of the block diagrams and / or flowcharts, and combinations of boxes in the block diagrams and / or flowcharts, can be implemented by hardware (such as circuits or an ASIC (Application Specific Integrated Circuit)) that performs the corresponding functions or acts, or can be implemented by a combination of hardware and software, such as firmware.
[0209] Although the present invention has been described in conjunction with the various embodiments, it will be understood by those skilled in the art that various changes in the disclosed embodiments can be understood and effected while practicing the claimed invention, by referring to the figures, the disclosure, and the appended claims. In the claims, the word "comprising" does not exclude other elements or steps, and the singular "a" or "an" does not exclude a plurality. A single processor or other unit may implement several functions recited in the claims. Certain measures are recited in mutually different dependent claims, but this does not indicate that these measures cannot be combined to advantage.
[0210] The embodiments of the present application have been described above. The above description is exemplary and not exhaustive, and is also not limited to the disclosed embodiments. Many modifications and variations are obvious to those of ordinary skill in the art without departing from the scope and spirit of the described embodiments. The choice of terms used herein is intended to best explain the principles of the embodiments, the practical application, or the improvement of the technology in the market, or to enable other ordinary skill in the art to understand the embodiments disclosed herein.
Claims
1. A natural language generation method, applied to a dialogue system, characterized in that The method includes: Encoding the system action text to obtain an encoded vector; wherein, the system action text is generated according to the user query information and a preset system action text template, the system action text template includes at least one preset slot, and the system action text includes at least one slot value; Decoding the encoded vector, and during the decoding process, judging whether the decoded text meets the decoding end condition according to the slot list corresponding to the system action text template or the slot value list corresponding to the system action text; wherein, the slot list is established according to all the preset slots in the system action text template, and multiple system action texts generated from the same system action text template share the same slot list, the slot value list is established according to all the slot values in the system action text, the decoded text is the text obtained by decoding the encoded vector, and the decoding end condition is: the decoded text includes all the preset slots in the system action text template, or the decoded text includes all the slot values in the system action text; When the decoded text meets the decoding end condition, determining the decoded text as the natural language text for replying to the user query information; The encoding the system action text to obtain an encoded vector specifically includes: Encoding each word in the system action text according to a preset dictionary to obtain word vectors; Determining a target phrase from the system action text according to a preset phrase library, and encoding the target phrase to obtain a phrase vector; wherein, the phrase library includes multiple phrases, and the target phrase is a phrase that is included in both the phrase library and the system action text; Fusing the word vectors and the phrase vector to obtain the encoded vector.
2. The method according to claim 1, characterized in that The encoding the target phrase to obtain a phrase vector specifically includes: Averaging or weighted-averaging the word vectors corresponding to each word in the target phrase to obtain a phrase vector.
3. The method according to claim 1, wherein The fusing the word vectors and the phrase vector to obtain the encoded vector specifically includes: Concatenating the word vectors and the phrase vector to obtain the encoded vector.
4. The method according to claim 1, wherein The system action text template includes at least one short text template, each short text template includes at least one slot, the system action text includes at least one short text, and each short text includes at least one slot value, The judging whether the decoded text meets the decoding end condition during the decoding process specifically includes: Judging whether the decoded text meets the first condition during the decoding process, wherein the first condition is: the decoded text includes at least one slot in the system action text template, or the decoded text includes at least one slot value in the system action text; When the decoded text meets the first condition, determining the target short text template to which the slot belongs, or determining the target short text to which the slot value belongs; Determine whether the decoded text satisfies a second condition, where the second condition is that the decoded text includes all slots in the target short text template, or the decoded text includes all slot values in the system action text; When the decoded text satisfies the second condition, determine that the decoded text does not satisfy the decoding end condition.
5. The method according to claim 4, characterized in that, The determining whether the decoded text satisfies the decoding end condition during the decoding process further includes: When the decoded text satisfies the second condition, determine whether the decoded text satisfies a third condition, where the third condition is that the decoded text includes all preset slots in the system action text template, or the decoded text includes all the slot values in the system action text; When the decoded text does not satisfy the third condition, determine that the decoded text does not satisfy the decoding end condition.
6. The method according to claim 5, characterized in that, The determining whether the decoded text satisfies the decoding end condition during the decoding process further includes: When the decoded text satisfies the third condition, determine that the decoded text satisfies the decoding end condition.
7. The method according to claim 1, characterized in that, The method further includes: Determine a system action according to the user query information, where the system action is an action executed when the dialogue system replies to the user query information, and the system action includes an action category and at least one slot-slot value pair; Determine a system action text according to the system action and the system action text template, where the system action text template corresponds to the system action, and the slots in the system action text template are the same as the slots in the system action.
8. The method according to claim 7, characterized in that The determining the system action text according to the system action and the system action text template specifically includes: Fill the slot values in the system action into the corresponding slots of the system action text template to obtain the system action text.
9. The method according to claim 1, characterized in that The decoding of the encoded vector includes: Determine attention information according to the encoded vector; Decode the encoded vector according to the attention information.
10. The method according to any one of claims 1-9, characterized in that, The encoding is completed by an encoder, the decoding is completed by a decoder, and the encoder and the decoder are neural networks.
11. The method according to claim 10, wherein The encoder is a bidirectional long short-term memory network, and the decoder is a long short-term memory network.
12. A natural language generation device, characterized in that, Includes: A processor; A memory for storing instructions executable by the processor; Wherein, when the processor is configured to execute the instructions, the method described in any one of claims 1-11 is implemented.
13. A non-volatile computer-readable storage medium having computer program instructions stored thereon, characterized in that, When the computer program instructions are executed by the processor, the method described in any one of claims 1-11 is implemented.
Citation Information
Patent Citations
Natural language generation method, device and equipment and readable storage medium
CN109815486A
Voice interaction method and system, terminal and computer readable storage medium
CN111368538A