Conversational processing method and apparatus
Patent Information
- Application Number
- CN202110219068.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-02-26
- Publication Date
- 2026-09-29
- Estimated Expiration
- 2041-02-26
AI Technical Summary
通过对输入语句和输入图像进行共同编码,得到融合有输入语句的序列信息以及输入图像所表达的情感信息的上下文变量,接着,对得到的上下文变量进行解码处理,得到用于响应输入语句、且与情感信息适配的回复语句,如此,实现了基于输入图像识别察觉用户的情绪,以根据用户的情绪生成应景的回复语句,提高了回复语句的准确度,提升了用户体验。
Smart Images

Figure CN113704419B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and in particular to a dialogue processing method, apparatus, electronic device, and computer-readable storage medium. Background Technology
[0002] Artificial intelligence (AI) is the theory, methods, technology, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results. In other words, AI is a comprehensive technology within computer science that attempts to understand the essence of intelligence and produce a new kind of intelligent machine that can react in a way similar to human intelligence. AI studies the design principles and implementation methods of various intelligent machines, enabling them to possess the functions of perception, reasoning, and decision-making.
[0003] Human-computer dialogue systems are an important branch of artificial intelligence. Their main goal is to enable machines to understand and use natural language, thus allowing them to interact with users like "humans." Human-computer dialogue systems have a wide range of applications, such as human-computer dialogue interfaces for various robots, intelligent customer service systems, and personal assistants.
[0004] However, the human-computer dialogue systems provided by related technologies typically rely on corpora and templates to interpret user input and then select appropriate responses. For example, simply searching for relevant responses based on the literal meaning of user input lacks human emotion, leading to poor communication, low accuracy, and a subpar user experience. Summary of the Invention
[0005] This application provides a dialogue processing method, apparatus, electronic device, and computer-readable storage medium, which can generate appropriate response statements based on the user's emotions, thereby improving the accuracy of the response statements.
[0006] The technical solution of this application embodiment is implemented as follows: This application provides a dialogue processing method, including: Get the input statements and input images from the dialogue; The input statement and the input image are encoded to obtain context variables, which incorporate the sequence information of the input statement and the emotional information expressed by the input image. The context variables are decoded to obtain a response statement that is used in response to the input statement and is adapted to the emotional information.
[0007] This application provides a dialogue processing apparatus, including: The acquisition module is used to acquire input statements and input images from the dialogue; An encoding module is used to encode the input statement and the input image to obtain context variables, wherein the context variables are fused with the sequence information of the input statement and the emotional information expressed by the input image; The decoding module is used to decode the context variables to obtain a response statement that is used in response to the input statement and is adapted to the emotional information.
[0008] In the above scheme, the acquisition module is further used to acquire the sequence information of the input statement; the device further includes an image representation module for extracting corresponding image features from the input image; the device further includes an object representation module for extracting object features of objects included in the input image; the encoding module is further used to fuse the sequence information of the input statement, the image features of the input image, and the object features of objects in the input image to obtain multimodal features; and to call an encoder to encode the multimodal features to obtain context variables that fuse the sequence information of the input statement and the emotional information expressed by the input image.
[0009] In the above scheme, the acquisition module is further configured to perform word segmentation on the input statement, concatenate the word embedding vectors corresponding to each word obtained from the word segmentation, and determine the concatenation result as the sequence information of the input statement; the image representation module is further configured to scale the input image to a fixed size to obtain a standard image; perform convolution processing on the standard image, perform pooling processing on the obtained convolution features to obtain pooling features; perform fully connected processing on the pooling features to obtain the image features corresponding to the input image; the object representation module in the image is configured to perform object detection processing on the input image to obtain the bounding box of each object in the input image; and extract the object features of the corresponding object from the bounding box of each object.
[0010] In the above scheme, the context variable also incorporates the fluency of the input statement; the decoding module is further used to decode the context variable to obtain a response statement that is used to respond to the input statement and is adapted to the emotional information and the fluency.
[0011] In the above scheme, the decoding module is further configured to decode the context variable to obtain the response word corresponding to the context variable and the selection probability of the response word; and select at least one response word to form a response statement that is used to respond to the input statement and is adapted to the emotional information based on the selection probability of the response word.
[0012] In the above scheme, the encoding and decoding processes are implemented through a dialogue model; the acquisition module is further used to acquire training images and training corpora, the training corpora including sample input statements and target response statements corresponding to the sample input statements; the encoding and decoding modules are further used to jointly encode and decode the training images and sample input statements through the dialogue model to obtain a predicted response statement for responding to the sample input statement; the device further includes a text sentiment classification module, used to perform sentiment prediction processing on the predicted response statement through a sentiment classification model to obtain a sentiment score representing positive sentiment; the device further includes an update module, used to substitute the sentiment score and the loss function of the dialogue model itself into a first objective function to determine the parameters of the dialogue model when the first objective function reaches its minimum value, and update the dialogue model based on the parameters; wherein, the first objective function is used to perform a weighted summation of the sentiment score and the loss function of the dialogue model itself.
[0013] In the above scheme, the encoding and decoding processes are implemented through a dialogue model. The acquisition module is further used to acquire training images and training corpora, the training corpora including sample input statements and target response statements corresponding to the sample input statements. The encoding and decoding modules are further used to jointly encode and decode the training images and sample input statements through the dialogue model to obtain a predicted response statement for responding to the sample input statement. The text sentiment classification module is further used to perform sentiment prediction processing on the predicted response statement through a sentiment classification model to obtain a sentiment score representing positive sentiment. The device also includes a language fluency module, used to perform fluency prediction processing on the predicted response statement through a language fluency model to obtain a fluency score representing the fluency of the statement. The update module is further used to substitute the sentiment score, the fluency score, and the loss function of the dialogue model itself into a second objective function to determine the parameters of the dialogue model when the second objective function reaches its minimum value, and update the dialogue model based on the parameters. The second objective function is used to perform a weighted summation of the sentiment score, the fluency score, and the loss function of the dialogue model itself.
[0014] In the above scheme, the language fluency module is further used to train the language fluency model in the following ways: obtaining normally expressed sentences as positive examples for training the language fluency model; obtaining abnormally expressed sentences as negative examples for training the language fluency model; constructing a training set based on the positive examples and the negative examples; and training the language fluency model based on the training set.
[0015] In the above scheme, the language fluency module is further configured to perform at least one of the following processes on the normally expressed statement: delete some words in the normally expressed statement and take the statement obtained after deletion as the abnormal statement; replace some words in the normally expressed statement with irrelevant words and take the statement obtained after replacement as the abnormal statement.
[0016] This application provides an electronic device, including: Memory, used to store executable instructions; The processor, when executing executable instructions stored in the memory, implements the dialogue processing method provided in the embodiments of this application.
[0017] This application provides a computer-readable storage medium storing executable instructions, which, when executed by a processor, implement the dialogue processing method provided in this application.
[0018] The embodiments of this application have the following beneficial effects: By co-encoding the input statement and the input image, a context variable is obtained that integrates the sequence information of the input statement and the emotional information expressed by the input image. Then, the obtained context variable is decoded to obtain a response statement that is appropriate to the input statement and matches the emotional information. In this way, the user's emotions can be detected based on the input image, and appropriate response statements can be generated according to the user's emotions, which improves the accuracy of the response statements and enhances the user experience. Attached Figure Description
[0019] Figure 1 This is a schematic diagram of the architecture of the dialogue processing system 100 provided in an embodiment of this application; Figure 2 This is a schematic diagram of the structure of the server 200 provided in the embodiments of this application; Figure 3 This is a flowchart illustrating the dialogue processing method provided in an embodiment of this application; Figure 4 This is a flowchart illustrating the dialogue processing method provided in an embodiment of this application; Figure 5 This is a schematic diagram illustrating the principle of Fast R-CNN provided in the embodiments of this application; Figure 6 This is a schematic diagram of the encoder provided in the embodiments of this application; Figure 7 This is a schematic diagram of the decoder provided in the embodiments of this application; Figure 8 This is a schematic diagram illustrating the principle of CNN provided in the embodiments of this application; Figure 9 This is a schematic diagram illustrating the principle of obtaining object features from an input image, as provided in an embodiment of this application. Figure 10 This is a schematic diagram illustrating the basic dialogue model provided by related technologies; Figure 11 This is a schematic diagram illustrating the principle of the improved dialogue model provided in the embodiments of this application. Detailed Implementation
[0020] To make the objectives, technical solutions, and advantages of this application clearer, the application will be further described in detail below with reference to the accompanying drawings. The described embodiments should not be regarded as limitations on this application. All other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0021] In the following description, references are made to “some embodiments,” which describe a subset of all possible embodiments. However, it is understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.
[0022] In the following description, the terms "first" and "second" are used merely to distinguish similar objects and do not represent a specific ordering of objects. It is understood that "first" and "second" may be interchanged in a specific order or sequence where permitted, so that the embodiments of this application described herein can be implemented in an order other than that illustrated or described herein.
[0023] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing embodiments of this application only and is not intended to limit this application.
[0024] Before providing a further detailed description of the embodiments of this application, the nouns and terms involved in the embodiments of this application will be explained, and the nouns and terms involved in the embodiments of this application shall be interpreted as follows.
[0025] 1) Human-Machine Conversation System: This refers to a system capable of human-computer dialogue, including task-oriented and non-task-oriented dialogue systems (also known as chatbots). Taking non-task-oriented dialogue systems as an example, they can engage in casual conversation with users, providing reasonable responses and entertainment functions. Typically, they focus on open-ended topics when conversing with users.
[0026] 2) Encoder: Its function is to transform a variable-length input sequence into a fixed-length vector, which encodes the sequence information of the input sequence. A common encoder is a recurrent neural network (RNN).
[0027] 3) Decoder: Its function is to map the fixed-length vector generated by the encoder to a variable-length output sequence. Common decoders are also recurrent neural networks, such as Long Short-Term Memory (LSTM) neural networks.
[0028] In related technologies, human-computer dialogue is usually implemented by using corpora and templates to determine the user's input information (such as user-inputted sentences or images), and then selecting the appropriate response.
[0029] In human communication, emotion is a crucial factor. Humans adjust their communication strategies based on the other person's current mood to achieve smooth communication. However, human-computer dialogue systems often only search for relevant responses based on the literal meaning of the user's input or extract non-emotional information from images. They fail to select appropriate responses based on the user's emotions, resulting in poor communication, low accuracy, and a subpar user experience.
[0030] To address the aforementioned technical problems, embodiments of this application provide a dialogue processing method, apparatus, electronic device, and computer-readable storage medium, which can generate appropriate response statements based on the user's emotions, thereby improving the accuracy of the response statements.
[0031] The following describes exemplary applications of the electronic devices provided in the embodiments of this application. These electronic devices can be implemented as various types of user terminals such as laptops, tablets, desktop computers, set-top boxes, and mobile devices (e.g., mobile phones, portable music players, personal digital assistants, dedicated messaging devices, portable gaming devices), or as servers, or implemented through a collaborative approach between servers and user terminals. The following will describe exemplary applications when the electronic device is implemented as a server.
[0032] See Figure 1 , Figure 1 This is a schematic diagram of the architecture of the dialogue processing system 100 provided in this application embodiment, designed to support the generation of appropriate response statements based on the user's emotions, thereby improving the accuracy of the response statements. For example... Figure 1 As shown, the dialogue processing system 100 includes a server 200, a network 300, and a terminal device 400, which will be described below.
[0033] Server 200 runs a trained dialogue model to encode the input statements and images received by terminal device 400, obtaining context variables. These context variables incorporate the sequence information of the input statements and the emotional information expressed by the input images. Next, server 200 calls the trained dialogue model to decode the context variables, obtaining a response statement that matches the emotional information and is used to respond to user input. Finally, server 200 returns the response statement output by the dialogue model to terminal device 400, enabling terminal device 400 to call the chat page provided by client 410 for presentation.
[0034] Network 300 is used as a medium for communication between server 200 and terminal device 400, and can be a wide area network or a local area network, or a combination of both.
[0035] Terminal device 400 runs a client 410, which can be a chat application, such as a chat assistant or intelligent customer service. Terminal device 400 uses client 410 to obtain images and statements entered by the user on the chat page, and sends the obtained input images and statements to server 200 via network 300. Terminal device 400 is also used to invoke the chat page of client 410 for presentation after receiving a reply statement from server 200.
[0036] It should be noted that the dialogue processing method provided in this application embodiment can be implemented independently by the server, independently by the terminal device, or jointly by the server and the terminal device. For example, the server trains a dialogue model and sends the trained dialogue model to the terminal device, so that the terminal device can directly conduct human-computer dialogue based on the received dialogue model.
[0037] The following describes an exemplary application of the dialogue processing method provided in the embodiments of this application when the electronic device is a terminal device.
[0038] In some embodiments, with Figure 1 Taking terminal device 400 as an example, a client 410 runs on terminal device 400. The client 410 can be implemented as a functional module integrated into the operating system, or as an independent application (such as a chat assistant, intelligent customer service, etc.), or as an application programming interface (API) integrated into an application for other applications to call. Taking the implementation of client 410 as a chat assistant as an example, in this embodiment, the chat assistant implements the dialogue processing method in an offline manner (i.e., without relying on server 200) in terminal device 400. That is, the process of encoding and decoding input statements and input images and outputting reply statements is all completed independently by the chat assistant in terminal device 400.
[0039] For example, when the chat assistant receives an image and a sentence input by the user on the chat page, it calls the trained dialogue model stored in the terminal device 400 to encode the received input image and input sentence, obtaining context variables that integrate the sequence information of the input sentence and the emotional information expressed by the input image. Then, the chat assistant continues to call the dialogue model to decode the context variables, obtaining a reply sentence that is used to respond to the input sentence and is adapted to the emotional information, and displays the reply sentence output by the dialogue model on the chat page.
[0040] In some embodiments, server 200 may be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. Terminal device 400 may be a smartphone, tablet, laptop, desktop computer, smart speaker, smartwatch, etc., but is not limited to these. Terminal device 400 and server 200 can be directly or indirectly connected via wired or wireless communication, which is not limited in this embodiment.
[0041] The following is about Figure 1 The structure of server 200 in the document will be explained. See [link / reference]. Figure 2 , Figure 2 This is a schematic diagram of the structure of the server 200 provided in the embodiments of this application. Figure 2 The server 200 shown includes at least one processor 210, memory 240, and at least one network interface 220. The various components of server 200 are coupled together via a bus system 230. It is understood that the bus system 230 is used to implement communication between these components. In addition to a data bus, the bus system 230 also includes a power bus, a control bus, and a status signal bus. However, for clarity, ... Figure 2 The general labeled all buses as Bus System 230.
[0042] Processor 210 can be an integrated circuit chip with signal processing capabilities, such as a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among them, the general-purpose processor can be a microprocessor or any conventional processor, etc.
[0043] The memory 240 may be removable, non-removable, or a combination thereof. Exemplary hardware devices include solid-state storage, hard disk drives, optical disk drives, etc. The memory 240 may optionally include one or more storage devices physically located away from the processor 210.
[0044] The memory 240 may include volatile memory or non-volatile memory, or both. The non-volatile memory may be read-only memory (ROM), and the volatile memory may be random access memory (RAM). The memory 240 described in this application embodiment is intended to include any suitable type of memory.
[0045] In some embodiments, memory 240 is capable of storing data to support various operations, examples of which include programs, modules, and data structures or subsets or supersets thereof, as illustrated below.
[0046] Operating system 241 includes system programs for handling various basic system services and performing hardware-related tasks, such as the framework layer, core library layer, driver layer, etc., for implementing various basic business functions and handling hardware-based tasks; The network communication module 242 is used to reach other computing devices via one or more (wired or wireless) network interfaces 220, such as Bluetooth, WiFi, and Universal Serial Bus (USB). In some embodiments, the apparatus provided in this application can be implemented in software. Figure 2 A dialogue processing device 243 stored in memory 240 is shown. This device can be software in the form of programs and plugins, including the following software modules: an acquisition module 2431, an encoding module 2432, a decoding module 2433, an image representation module 2434, an object representation module 2435, a text sentiment classification module 2436, an update module 2437, and a speech fluency module 2438. These modules are logically connected and can therefore be arbitrarily combined or further separated according to the functions they implement. It should be noted that... Figure 2 For ease of explanation, all the above modules are shown at once, but this should not be interpreted as excluding the implementation of the dialogue processing device 243 which may only include the acquisition module 2431, the encoding module 2432, and the decoding module 2433. The functions of each module will be described below.
[0047] The following examples illustrate different software implementations of the dialogue processing device 243.
[0048] Example 1: The dialogue processing device can be an application or module running on a terminal device. This application provides software modules designed using programming languages such as C / C++ and Java, which can be embedded into various terminal apps (such as game applications) based on systems such as Android or iOS (with executable instructions stored in the terminal's storage medium and executed by the terminal's processor). This allows the terminal to directly use its own computing resources to complete related machine model training, application, and other tasks, and to periodically or irregularly transmit the model training and application results to a remote server through various network communication methods, or to save them locally on the mobile device.
[0049] Example 2: The dialogue processing device can be a server application and platform. This application provides a dedicated software module in application software or large software systems designed using programming languages such as C / C++ and Java. It runs on the server side (stored in the server-side storage medium as executable instructions and executed by the server-side processor). It integrates at least one of various raw data, intermediate data at various levels, and final results received from other devices with some existing data or results on the server to train the model, and uses the trained model to identify transactions. Then, it outputs the model or transaction identification results to other applications or modules in real time or non-real time, or it can write them to a server-side database or file for storage.
[0050] This application embodiment can also provide a UI design platform for individuals, groups, or enterprises, which is formed by mounting a customized, easy-to-interact web interface or other user interfaces on a distributed, parallel computing platform composed of multiple servers. Users can upload existing data packets to this platform in batches to obtain various calculation results, or they can transmit real-time data streams to this platform to calculate and refresh results at all levels in real time.
[0051] Example 3: The dialogue processing device can be a server-side application programming interface (API) and plugins. The embodiments of this application can provide server-side implementations of model training functions, APIs for generating abnormal transaction identification based on models, software development kits (SDKs) or plugins, which can be called by other server-side application developers and embedded into various applications.
[0052] Example 4: The dialogue processing device can be a terminal device client API and plugin. This application embodiment may also provide APIs, SDKs, or plugins for implementing model training functions on terminal devices, generating abnormal transaction identification based on machine learning or deep learning models, for other terminal application developers to call and embed into various applications.
[0053] Example 5: Dialogue processing devices can be cloud-based open services. This application embodiment can provide cloud services for UI interface design for AI-based abnormal transaction processing. This application embodiment can also provide application packages (APK, Android Application Package), software development kits (SDK, Software Development Kit), and plugins for UI interface design cloud services, packaged and encapsulated into cloud services that can be used by people inside and outside the enterprise, or display various results in an appropriate form on various terminal display devices for use by individuals, groups, or enterprises.
[0054] In other embodiments, the dialogue processing apparatus provided in this application can be implemented in hardware. As an example, the dialogue processing apparatus provided in this application can be a processor in the form of a hardware decoding processor, which is programmed to execute the dialogue processing method provided in this application. For example, the processor in the form of a hardware decoding processor can be one or more application-specific integrated circuits (ASICs), DSPs, programmable logic devices (PLDs), complex programmable logic devices (CPLDs), field-programmable gate arrays (FPGAs), or other electronic components.
[0055] The dialogue processing method provided in the embodiments of this application will be described in detail below with reference to the accompanying drawings. See also Figure 3 , Figure 3 This is a flowchart illustrating the dialogue processing method provided in the embodiments of this application. In some embodiments, the dialogue processing method provided in the embodiments of this application can be implemented by the server or the terminal device alone, or it can be implemented collaboratively by the server and the terminal device. In the following steps, the dialogue processing method provided in the embodiments of this application is illustrated using the example of the server implementing it alone.
[0056] In step S101, the input statements and input images in the dialogue are obtained.
[0057] In some embodiments, a trained dialogue model runs on the server to obtain input statements and input images from the dialogue. The input images can be images uploaded to the server by the user via a terminal device (e.g., images pre-stored on the terminal device or images captured in real-time by the user through the terminal device's built-in camera), images pre-stored on the server (e.g., images stored on the server that are adapted to existing dialogue scenarios), or images from an existing image set. This application embodiment does not specifically limit the type of input image.
[0058] For example, taking a standalone application as the client-side implementation, a chat application, such as a chat assistant, runs on the user's terminal device. When the user clicks to enter the chat page provided by the chat assistant, they can enter an opening statement (e.g., "How's the weather today?"). Simultaneously, the user can also enter images (e.g., a selfie taken in real-time by the terminal device's built-in camera, or a target image selected from a pre-stored image set (e.g., a photo album) on the terminal device). After receiving the statements and images entered by the user on the chat page, the chat assistant uploads the input statements and images to the server, allowing the server to invoke a trained dialogue model to perform subsequent processing on the received input statements and images.
[0059] In step S102, the input statement and the input image are encoded to obtain context variables, wherein the context variables are fused with the sequence information of the input statement and the emotional information expressed by the input image.
[0060] In some embodiments, Figure 3 The illustrated step S102 can be achieved through Figure 4 Steps S1021 to S1024 shown are implemented, and will be combined with Figure 4 The steps shown are explained.
[0061] In step S1021, the sequence information of the input statement is obtained.
[0062] In some embodiments, the server may obtain the sequence information of the input statement by performing word segmentation on the input statement, concatenating the word embedding vectors corresponding to each word obtained from the word segmentation, and determining the concatenation result as the sequence information of the input statement.
[0063] For example, to analyze input statements using a dialogue model, the words in the corresponding text need to be converted into vectors, i.e., used as numerical input to the dialogue model. Word embedding is a method for converting words in text into numerical vectors. The word embedding process involves embedding a high-dimensional space containing the number of words into a much lower-dimensional continuous vector space. Each word or phrase is mapped to a vector in the real number field, and the result of word embedding is a word vector. Common word embedding methods include one-hot encoding, distributed representation, skip-gram model, and continuous bag of words (CBOW).
[0064] One-hot encoding is the most basic vector representation method, which represents words in the text using word-size vectors, where only the term corresponding to the word is 1 and all other terms are 0. Distributed representation aims to find a transformation function to convert each word into its associated vector; that is, distributed representation is to convert words into vectors, where the similarity between vectors is related to the semantic similarity between words. The basic idea of skip-word model is to predict the window function of the order of use of each central function and correct the vector of the central function based on the prediction results. The basic idea of continuous bag-of-words model is to predict the vector of the central function by using the vector of the window function of the order of use of each function.
[0065] For example, for the input statement "How's the weather today?", the server first performs word segmentation to obtain the corresponding word sequence: "today", "weather", and "how". Then, it uses word embedding to obtain the feature vector corresponding to each word in the word sequence, that is, it obtains the feature vectors corresponding to "today", "weather", and "how". Subsequently, the server concatenates the feature vectors corresponding to each word and uses the concatenated vector as the sequence information of the input statement "How's the weather today?".
[0066] It should be noted that the embodiments of this application do not specifically limit the method of obtaining the sequence information of the input statement (that is, converting the input statement into the corresponding vector sequence), and any of the word embedding methods mentioned above can be used to achieve this.
[0067] In step S1022, the corresponding image features are extracted from the input image.
[0068] In some embodiments, the server may extract corresponding image features from the input image by: scaling the input image to a fixed size to obtain a standard image; performing convolution processing on the standard image, performing pooling processing on the obtained convolution features to obtain pooled features; and performing fully connected processing on the pooled features to obtain the image features corresponding to the input image.
[0069] For example, given an input image, the server can extract corresponding image features (i.e., a vectorized representation of the input image) from the input image using a trained Convolutional Neural Network (CNN). The basic structure of a CNN includes convolutional layers, pooling layers, and fully connected layers. Convolutional layers are primarily used for feature extraction; the input of each neuron is connected to the local receptive field of the previous layer, and features from that local field are extracted. Pooling layers are mainly used to compress the convolutional features input to the convolutional layers. This reduces the size of the convolutional features, simplifying the network's computational complexity, and also performs feature compression to extract the main features from the convolutional features, which are then used as pooling features. There are generally two types of pooling operations: average pooling (Avg Pooling) and max pooling (Max Pooling). Fully connected layers are mainly used for feature mapping. Each computational layer of the network consists of multiple feature maps, each of which is a plane where all neurons have equal weights. Users can customize the structure of a CNN (such as the number of layers, the connectivity between layers, etc.), then determine the parameters of each layer through training, and then use the trained CNN to extract features from the input image. The extracted features can be represented in vector space to obtain the image feature vector, thereby mapping the input image into a low-dimensional vector space.
[0070] For example, given an input image A, the server can first use a spatial transformation matrix to scale the received input image A to a fixed size, obtaining a standard image (assuming the standard image size is 32). 32 3, of which 32 32 is the width Height, 3 is the depth (i.e., R, G, B) of the standard image; then, the server inputs the standard image into the trained CNN, so that the CNN's convolutional layers can extract features from the standard image. The convolutional layer can be a 5x5... 5 A 3x3 receptive field (filter) is generated, where the depth of the receptive field must be the same as the depth of the standard image. A 2x8 filter can be obtained by convolving the receptive field with the standard image. 28 A feature map of size 1 is then input into a pooling layer after convolution to compress the input feature map. For example, max pooling can be used to compress the input feature map to size 13. The size is 13; finally, the compressed feature map is input into the fully connected layer for feature mapping, thereby mapping the input image A into a low-dimensional vector space, that is, obtaining the image features corresponding to the input image A.
[0071] In step S1023, object features of objects included in the input image are extracted from the input image.
[0072] In some embodiments, the server may extract object features of objects included in the input image from the input image by performing object detection processing on the input image to obtain the bounding box of each object in the input image; and extracting the object features of the corresponding object from the bounding box of each object.
[0073] For example, the server can first use a Regions with CNN features (R-CNN) to perform object detection processing on the input image, obtaining the bounding box of each object in the input image. The process of R-CNN for object detection on the input image is as follows: First, the model input is an image (e.g., the input image). Then, a predetermined number (e.g., 200) of regions to be detected are extracted from the image. Features are extracted one by one (i.e., serially) from the predetermined number of regions to be detected by the convolutional neural network. The extracted features are then classified by a Support Vector Machine (SVM) to determine the category of the object, and the size of the target bounding box is adjusted by bounding box regression. After obtaining the bounding box of each object in the input image, for each bounding box, the corresponding object features (i.e., the vectorized representation of the object) can be extracted from the bounding box of each object in a manner similar to step S1022.
[0074] For example, the server can also use Fast R-CNN to perform object detection on the input image. See also Figure 5 , Figure 5 This is a schematic diagram illustrating the principle of Fast R-CNN provided in the embodiments of this application. Figure 5As shown, firstly, a predetermined number (e.g., 100) of regions to be detected are determined on the image (e.g., the input image). Then, a convolutional neural network is used to extract features from these predetermined number of regions. Next, a Region of Interest Pooling Layer (ROI Pooling Layer) is used to extract features corresponding to each ROI from the full image features. Finally, a Fully Connected Layer (FCLayer) is used for classification and bounding box correction. After obtaining the bounding box of each object in the input image, for each bounding box, the object features of the corresponding object can be extracted from the bounding box in a manner similar to step S1022. This embodiment will not be elaborated further here.
[0075] It should be noted that steps S1021 to S1023 can be executed synchronously or sequentially. For example, step S1021 can be executed first, then step S1022, and finally step S1023. Alternatively, step S1022 can be executed first, then step S1023, and finally step S1021. This application embodiment does not specifically limit this.
[0076] In step S1024, the sequence information of the input statement, the image features of the input image, and the object features of the objects in the input image are fused to obtain multimodal features; the encoder is called to encode the multimodal features to obtain context variables that fuse the sequence information of the input statement and the emotional information expressed by the input image.
[0077] In some embodiments, after obtaining the sequence information of the input statement, the image features of the input image, and the object features of the objects in the input image, the server performs a fusion process (also known as a concatenation process) on the sequence information of the input statement, the image features of the input image, and the object features of the objects in the input image to obtain multimodal features. For example, the image features of the input image and the object features of the objects in the input image can be directly concatenated after the sequence information of the input statement. Then, an encoder (e.g., RNN) is called to encode the fused multimodal features to obtain context variables that fuse the sequence information of the input statement and the emotional information expressed by the input image.
[0078] For example, taking the input statement "How's the weather today?" and the input image A as an example, the server first maps the text in the input statement into word vectors through word embedding processing, thereby obtaining the sequence information of the input statement. Let's assume the sequence information of the input statement is {a1, a2, a3}, where "a1" is the word vector corresponding to "today", "a2" is the word vector corresponding to "weather", and "a3" is the word vector corresponding to "how's". Next, the server obtains the image features of image A through CNN, let's assume it's vector B. Subsequently, the server obtains the object features of objects in image A through Fast R-CNN, let's assume it's {b1, b2}, where "b1" is the feature vector corresponding to object c in image A, and "b2" is the feature vector corresponding to object d in image A. After obtaining the sequence information of the input statement, the image features of image A, and the object features of objects in image A, the server concatenates these features to obtain the multimodal features {a1, a2, a3, B, b1, b2}.
[0079] See Figure 6 , Figure 6 This is a schematic diagram of the encoder provided in the embodiments of this application, as shown below. Figure 6 As shown, after obtaining multimodal features through fusion processing, the server can input the multimodal features into an encoder (e.g., an RNN) for encoding processing to transform the input multimodal features into a fixed-length vector c (i.e., a context variable that integrates the sequence information of the input statement and the emotional information expressed by image A). There are several ways to obtain vector c. For example, the last hidden state of the encoder can be directly assigned to vector c, i.e., c = h6; a transformation can be performed on the last hidden state to obtain vector c, i.e., c = q(h6); or a transformation can be performed on all hidden states to obtain vector c, i.e., c = q(h1, h2, h3, h4, h5, h6), where "h0" is the initial hidden state, "h1" is the hidden state obtained by encoding vector "a1", "h2" is the hidden state obtained by encoding vector "a2", and so on, with "h6" being the hidden state obtained by encoding vector "b2".
[0080] Alternatively, a bidirectional recurrent neural network can be used to construct the encoder. In this case, the hidden state of the encoder at each time step depends on both the subsequences before and after that time step (including the input at the current time step), and encodes information about the entire sequence. Then, by concatenating the hidden states of each time step, a context variable is obtained that integrates the sequence information of the input statement and the emotional information expressed by the input image.
[0081] In step S103, the context variables are decoded to obtain a response statement that is used to respond to the input statement and is adapted to the emotional information.
[0082] In some embodiments, the server can decode the context variables to obtain a response statement that is appropriate to the input statement and matches the sentiment information by: decoding the context variables to obtain the response words corresponding to the context variables and the probability of the response words being selected; and selecting at least one response word to form a response statement that is appropriate to the input statement and matches the sentiment information based on the probability of the response words being selected.
[0083] For example, after the server calls the encoder to encode the concatenated multimodal features to obtain context variables that integrate the sequence information of the input statement and the emotional information expressed by the input image, it calls the decoder (e.g., another RNN) to decode the context variables to obtain the response words corresponding to the context variables and the selection probability of each response word. Then, based on the selection probability of each response word, at least one response word is selected to form a response statement that is used to respond to the input statement and is adapted to the emotional information expressed by the input image. The RNN can be a Long Short-Term Memory (LSTM) model or a Gated Recurrent Unit (GRU) Recurrent Neural Network.
[0084] For example, see Figure 7 , Figure 7 This is a schematic diagram of the decoder provided in the embodiments of this application, as shown below. Figure 7 As shown, after the server calls the encoder to encode the multimodal features and obtains a fixed-length vector c (i.e., a context variable that integrates the sequence information of the input statement and the emotional information expressed by the input image), it inputs vector c into the decoder to obtain the output sequence. The output of the previous time step is used as the input of the current time step, and vector c only participates in the calculation as the initial state; subsequent calculations are independent of vector c. Of course, vector c can also participate in the calculations at all time steps of the sequence; that is, the output of the previous time step is still used as the input of the current time step, but vector c participates in the calculations at all time steps.
[0085] In other embodiments, the context variable may also incorporate the fluency of the input statement. The server can then decode the context variable as described above to obtain a response statement that is appropriate for the input statement and matches the emotional information by decoding the context variable.
[0086] In some embodiments, the above encoding and decoding processes can be implemented through a dialogue model. Before the server performs encoding and decoding processes on the input statement and input image through the dialogue model, it can also perform the following operations: acquire training images and training corpora, wherein the training corpora include sample input statements and target response statements corresponding to the sample input statements; perform encoding and decoding processes on the training images and sample input statements together through the dialogue model to obtain a predicted response statement for responding to the sample input statement; perform sentiment prediction processing on the predicted response statement through a sentiment classification model to obtain a sentiment score representing positive sentiment; substitute the sentiment score and the dialogue model's own loss function (e.g., a loss function obtained based on the difference between the predicted response statement and the target response statement output by the dialogue model) into a first objective function to determine the parameters of the dialogue model when the first objective function reaches its minimum value, and update the dialogue model based on these parameters; wherein the first objective function is used to perform a weighted summation of the sentiment score and the dialogue model's own loss function.
[0087] For example, the server can train the dialogue model as follows: After the dialogue model outputs a predicted response to the sample input statement, the server calls a sentiment classification model (e.g., a CNN-based text sentiment classification model; sentiment classification is essentially a text classification task, so a classic CNN architecture for text classification can be used) to perform sentiment prediction on the predicted response statement, obtaining a sentiment score representing positive sentiment (the sentiment score can range from 0 to 1, where positive sentiment is 1 and negative sentiment is 0). Then, the server substitutes the sentiment score and the dialogue model's own loss function into the first objective function to determine the parameters of the dialogue model when the first objective function reaches its minimum value, and updates the dialogue model based on the obtained parameters. The dialogue model's own loss function can use the error between the predicted response statement and the target response statement as a difference factor. The type of loss function can include Mean Squared Error (MSE), Hinge Loss Function (HLF), and Cross Entropy, etc. For example, taking the square loss function as an example, when the number of samples is n, the loss function can be expressed as:
[0088] in, Y This indicates the target response statement. This represents the error between the target response and the predicted response, with as the difference factor. The entire formula represents the sum of squared errors, so the first objective function can be expressed as:
[0089] in, Indicates emotional rating. This represents the loss function of the dialogue model itself. Represents the loss function L The corresponding weight value, Indicates emotional score The corresponding weight values, and the ultimate goal of training is to minimize the first objective function. The function value.
[0090] In other embodiments, the above encoding and decoding processes can be implemented through a dialogue model. Before the server encodes and decodes the input statement and input image using the dialogue model, it can also perform the following operations: acquire training images and training corpora, wherein the training corpora include sample input statements and corresponding target response statements; encode and decode the training images and sample input statements together using the dialogue model to obtain a predicted response statement for responding to the sample input statement; perform sentiment prediction processing on the predicted response statement using a sentiment classification model to obtain a sentiment score representing positive sentiment; perform fluency prediction processing on the predicted response statement using a language fluency model to obtain a fluency score representing the fluency of the statement; substitute the sentiment score, fluency score, and the dialogue model's own loss function into a second objective function to determine the parameters of the dialogue model when the second objective function reaches its minimum value, and update the dialogue model based on these parameters; wherein the second objective function is used to perform a weighted summation of the sentiment score, fluency score, and the dialogue model's own loss function.
[0091] For example, when training the dialogue model, the server can also incorporate fluency information. That is, when the dialogue model outputs a predicted response to the input statement, in addition to calling the sentiment classification model to perform sentiment prediction on the predicted response and obtain a sentiment score representing positive sentiment, it also calls the trained language fluency model to perform fluency prediction on the predicted response and obtain a fluency score representing the fluency of the statement. Then, the server simultaneously substitutes the sentiment score, fluency score, and the dialogue model's own loss function into the second objective function to determine the parameters of the dialogue model when the second objective function reaches its minimum value, and updates the dialogue model based on the obtained parameters. In this case, the second objective function can be expressed as:
[0092] in, Indicates emotional rating. S 2 indicates a smoothness score. This represents the loss function of the dialogue model itself. 1 indicates an emotional score.S The weight value corresponding to 1 2 indicates a smoothness score. S The weight value corresponding to 2, 1 represents the loss function. L The corresponding weight values, and the ultimate goal of training is to minimize the second objective function. The function value.
[0093] In some embodiments, before the server performs fluency prediction processing on the predicted response statement using the language fluency model, it may also perform the following operations: obtain statements with normal expression as positive examples for training the language fluency model; obtain statements with abnormal expression as negative examples for training the language fluency model; construct a training set based on the positive and negative examples, and train the language fluency model based on the constructed training set.
[0094] In other embodiments, following the above embodiments, the server can obtain statements with abnormal expressions in the following ways: delete some words from statements with normal expressions and use the statements obtained after deletion as statements with abnormal expressions; or, replace some words from statements with normal expressions with irrelevant words and use the statements obtained after replacement as statements with abnormal expressions.
[0095] For example, the server can train a language fluency model as follows: First, obtain normally expressed sentences and use them as positive examples for training the language fluency model (these sentences are fluent and should have high scores). Then, add "noise" to certain words in the normally expressed sentences to make them fluent, and use these as negative examples for training the language fluency model (these sentences are not fluent and should have low scores). Adding "noise" to words in normally expressed sentences can be done by randomly masking some words or randomly replacing some words with irrelevant words. Next, a training sample set is constructed based on the positive and negative examples, and the language fluency model is trained using this training sample set to obtain the trained language fluency model. In this way, the trained language fluency model can be used to predict the fluency of predicted responses and obtain the fluency score for the predicted response.
[0096] The dialogue processing method provided in this application integrates the emotional information expressed by the input image into the encoded context variables, enabling the recognition and perception of the user's emotions based on the input image. This allows for the decoding of response statements that match the user's emotions (and since the dialogue model is trained based on positive emotion scoring, the response statements output by the dialogue model are positive in emotion, thus encouraging the user, making the user happy, and improving the user experience). Simultaneously, the integration of the fluency of the input statement into the context variables also makes the decoded response statements more fluent, further improving the accuracy of the response statements.
[0097] The following will describe an exemplary application of the embodiments of this application in a real-world application scenario.
[0098] In related technologies, human-computer dialogue is typically implemented using corpora and templates to determine the user's input information (such as user-inputted sentences or images), and then selecting an appropriate response. In other words, the solutions provided by these technologies do not consider the user's emotions; they merely analyze the user's input text data or extract non-emotional information from user-inputted images to generate the corresponding response.
[0099] However, since emotion is a crucial factor in human communication, humans adjust their dialogue strategies based on the other person's current mood to achieve smooth communication. But current human-computer dialogue systems typically only search for relevant responses based on the literal meaning of the user's input or extract non-emotional information from images, failing to select appropriate responses based on the user's emotions. This results in disjointed communication, low accuracy, and a poor user experience.
[0100] In view of this, embodiments of this application provide a dialogue processing method that can recognize and perceive the user's emotions through images input by the user (such as the user's selfie), and generate appropriate response statements based on the user's emotions. In this way, by engaging in dialogue with the user, the user is encouraged, made happy, and thus generates positive emotions.
[0101] The dialogue processing method provided in this application can be applied to multimodal (e.g., including both image and text modalities) dialogue scenarios, such as intelligent customer service and chat assistants. A user speaks a sentence to the intelligent customer service (e.g., inputting text on the chat page or via voice), accompanied by an image (e.g., the user's facial expression). The intelligent customer service can recognize the user's emotions based on the image input and return a response that matches the user's emotions.
[0102] The dialogue processing method provided in the embodiments of this application will be described in detail below.
[0103] The dialogue processing method provided in this application can be run by a server or a terminal device. Taking a terminal device as an example, the terminal device runs the following five functional modules: an image representation module, an object representation module in an image, a text sentiment classification module, a language fluency model, and a basic dialogue model. Each of the above functional modules will be described in detail below.
[0104] I. Image Representation Module The image representation module is mainly used to represent an image as a vector, which represents the information of the image. For example, a trained convolutional neural network (CNN) can be used to process images and extract image features. The extracted features can be represented in vector space to obtain the image feature vector, thereby mapping the image into a low-dimensional vector space.
[0105] For example, see Figure 8 , Figure 8 This is a schematic diagram of the principle of CNN provided in the embodiments of this application, such as... Figure 8 As shown, the image passes through convolutional layers, pooling layers, and fully connected layers in sequence to obtain a vector representation of the image. This representation will be input into the subsequent dialogue model (e.g., a dialogue model built based on the transformer model).
[0106] II. Object Representation Module in Images The object representation module in an image is mainly used to identify multiple objects contained in an image and represent each identified object as a corresponding vector, where each vector represents information about the corresponding object. For example... Figure 9As shown, Faster R-CNN can be used first to perform object detection on the user-input image, obtaining bounding boxes for each object in the input image. Then, for each bounding box, a trained CNN is used to map the object within the bounding box into a corresponding vector representation. These vectors are also input into the subsequent transformer model. The process of using Faster R-CNN to perform object detection on the user-input image is as follows: First, shared convolutional layers are used to extract features from the entire image. Then, the obtained feature map is fed into a Region Proposal Network (RPN). The RPN generates detection boxes (specifying the location of the ROI) and performs the first correction on the bounding boxes of the ROIs. After that, the Fast R-CNN architecture is used. The ROI Pooling Layer selects the features corresponding to each ROI on the feature map based on the output of the RPN and sets the dimension to a fixed value. Finally, fully connected layers are used to classify the bounding boxes and perform a second correction on the bounding boxes.
[0107] III. Text Sentiment Classification Module The function of the text sentiment classification module is to take a sentence as input and output its sentiment score. The sentiment score ranges from 0 to 1, with 1 representing positive sentiment and 0 representing negative sentiment. Text sentiment classification is essentially a text classification task; therefore, classic CNN architectures used for text classification can be used, such as CNN-based text sentiment classification models. After the dialogue model outputs a response, the text sentiment classification module scores the sentiment of the response, and this score is used to train the dialogue model.
[0108] IV. Language Fluency Module The function of the language fluency module is to input a sentence and give a fluency score for that sentence.
[0109] For example, a language fluency model can be trained as follows: Normal sentences are used as positive examples (these sentences are fluent and should have high scores); "noise" is added to some words in normal sentences to make them fluent, and these fluent sentences are used as negative examples (these sentences are not fluent and should have low scores). Adding "noise" to normal sentences can be done by randomly masking some words (Mask) or randomly replacing words in normal sentences with irrelevant words (Replace).
[0110] The language fluency module can be pre-trained before the entire dialogue model begins training. After pre-training, the language fluency module will not be used further; it will only serve as a tool to provide "language fluency" scores throughout the dialogue model.
[0111] V. Basic Dialogue Model For example, see Figure 10 , Figure 10 This is a schematic diagram illustrating the basic dialogue model provided by related technologies. For example... Figure 10 As shown, the basic dialogue model provided by the related technology is a dialogue model based on plain text information. For example, after a user inputs the text "The weather is really nice today!", the basic dialogue model generates the response "Yes! Let's go play together." The structure of the basic dialogue model can be selected according to different types as needed. For example, it can be a sequence-to-sequence (seq2seq) model (suitable for cases with a small amount of data, such as less than 100,000 data points) or a transformer model (suitable for cases with a large amount of data, such as more than 100,000 data points).
[0112] In other words, a basic dialogue model can take a sentence as input and generate a corresponding response (i.e., another sentence). The specific process for generating the response is as follows: An input sentence is called, which encodes it to generate a corresponding vector. Then, a decoder is called to decode it, generating another sentence. Here, the input sentence is the sentence entered by the user, and the generated sentence is the response given by the basic dialogue model.
[0113] The following is based on Figure 10 Based on the basic dialogue model shown, this paper explains how the four modules provided in the embodiments of this application are integrated into the basic dialogue model.
[0114] For example, see Figure 11 , Figure 11 This is a schematic diagram illustrating the principle of the improved dialogue model provided in the embodiments of this application, such as... Figure 11 As shown, the image representation module and the object representation module process the user-input image. The image representation module generates an image-level vector for the entire image; the object representation module generates an object-level vector for each object in the image. For an image containing multiple objects, the object representation module generates a series of object vectors. These vectors (including the image vector generated by the image representation module and the series of object vectors generated by the object representation module) are concatenated with the input text and fed into the basic dialogue model.
[0115] After the basic dialogue model outputs a response, this embodiment of the application can further utilize reinforcement learning to provide two reward functions for the generated response, and use the evaluation results obtained from these functions for feedback learning. Specifically, the text sentiment classification module scores the positiveness of the sentiment expressed in the response output by the basic dialogue model; this score serves as the evaluation for the text sentiment classification module, and a higher score is desirable. Similarly, the language fluency module scores the fluency of the response output by the basic dialogue model; this score also serves as the evaluation for the language fluency module, and a higher score is desirable. Reinforcement learning guides the training of the basic dialogue model based on these two scores, enabling the model to extract positive emotional information from image information and reflect it in the output response.
[0116] Furthermore, the basic dialogue model's own loss function can be used for training (the optimization objective is maximum likelihood estimation, which means making the generated response as similar as possible to the "standard response" in the dataset). Therefore, the final training objective function can be expressed as: The final objective function = sentiment score obtained from the text sentiment classification module + fluency score obtained from the language fluency module + loss function (maximum likelihood estimation) of the basic dialogue model itself.
[0117] The dialogue processing method provided in this application can improve the performance of multimodal dialogue systems, making the emotional tone of the responses generated by the dialogue system more positive, making users feel more comfortable, and thus improving the user experience.
[0118] The following continues to describe the exemplary structure of the dialogue processing device 243 provided in the embodiments of this application as a software module. In some embodiments, such as Figure 2 As shown, the software modules stored in the dialogue processing device 243 of the memory 240 may include: an acquisition module 2431, an encoding module 2432, a decoding module 2433, an image representation module 2434, an object representation module 2435, a text sentiment classification module 2436, an update module 2437, and a language fluency module 2438.
[0119] The acquisition module 2431 is used to acquire the input statement and input image in the dialogue; the encoding module 2432 is used to encode the input statement and input image to obtain context variables, which are fused with the sequence information of the input statement and the emotional information expressed by the input image; the decoding module 2433 is used to decode the context variables to obtain a response statement that is used to respond to the input statement and is adapted to the emotional information.
[0120] In some embodiments, the acquisition module 2431 is further configured to acquire sequence information of the input statement; the dialogue processing device 243 further includes an image representation module 2434, configured to extract corresponding image features from the input image; the dialogue processing device 243 further includes an object representation module 2435, configured to extract object features of objects included in the input image from the input image; the encoding module 2432 is further configured to fuse the sequence information of the input statement, the image features of the input image, and the object features of objects in the input image to obtain multimodal features; and call the encoder to encode the multimodal features to obtain context variables that fuse the sequence information of the input statement and the emotional information expressed by the input image.
[0121] In some embodiments, the acquisition module 2431 is further configured to perform word segmentation on the input statement, concatenate the word embedding vectors corresponding to each word obtained from the word segmentation, and determine the concatenation result as the sequence information of the input statement; the image representation module 2434 is further configured to scale the input image to a fixed size to obtain a standard image; perform convolution processing on the standard image, perform pooling processing on the obtained convolution features to obtain pooling features; perform fully connected processing on the pooling features to obtain the image features corresponding to the input image; the object representation module 2435 is configured to perform object detection processing on the input image to obtain the bounding box of each object in the input image; and extract the object features of the corresponding object from the bounding box of each object.
[0122] In some embodiments, the context variable also incorporates the fluency of the input statement; the decoding module 2433 is further configured to decode the context variable to obtain a response statement that is used to respond to the input statement and is adapted to the emotional information and fluency.
[0123] In some embodiments, the decoding module 2433 is further configured to decode the context variable to obtain the response word corresponding to the context variable and the probability of the response word being selected; and to select at least one response word to form a response statement that is used to respond to the input statement and is adapted to the emotional information based on the probability of the response word being selected.
[0124] In some embodiments, encoding and decoding are implemented through a dialogue model; the acquisition module 2431 is further configured to acquire training images and training corpora, the training corpora including sample input statements and target response statements corresponding to the sample input statements; the encoding module 2432 and the decoding module 2433 are further configured to jointly encode and decode the training images and sample input statements through the dialogue model to obtain predicted response statements for responding to sample input statements; the dialogue processing device 243 further includes a text sentiment classification module 2436, configured to perform sentiment prediction processing on the predicted response statements through a sentiment classification model to obtain a sentiment score representing positive sentiment; the dialogue processing device 243 further includes an update module 2437, configured to substitute the sentiment score and the loss function of the dialogue model itself into a first objective function to determine the parameters of the dialogue model when the first objective function reaches its minimum value, and update the dialogue model based on the parameters; wherein, the first objective function is used to perform weighted summation processing on the sentiment score and the loss function of the dialogue model itself.
[0125] In some embodiments, encoding and decoding are implemented through a dialogue model; the acquisition module 2431 is further configured to acquire training images and training corpora, the training corpora including sample input statements and target response statements corresponding to the sample input statements; the encoding module 2432 and the decoding module 2433 are further configured to jointly encode and decode the training images and sample input statements through the dialogue model to obtain a predicted response statement for responding to the sample input statement; the text sentiment classification module 2436 is further configured to perform sentiment prediction processing on the predicted response statement through the sentiment classification model to obtain a sentiment score representing positive sentiment; the dialogue processing device 243 further includes a language fluency module 2438, configured to perform fluency prediction processing on the predicted response statement through the language fluency model to obtain a fluency score representing the fluency of the statement; the update module 2437 is further configured to substitute the sentiment score, fluency score, and the dialogue model's own loss function into the second objective function to determine the parameters of the dialogue model when the second objective function reaches its minimum value, and update the dialogue model based on the parameters; wherein, the second objective function is used to perform weighted summation processing on the sentiment score, fluency score, and the dialogue model's own loss function.
[0126] In some embodiments, the language fluency module 2438 is further configured to train the language fluency model by: obtaining normally expressed sentences as positive examples for training the language fluency model; obtaining abnormally expressed sentences as negative examples for training the language fluency model; constructing a training set based on the positive and negative examples; and training the language fluency model based on the training set.
[0127] In some embodiments, the language fluency module 2438 is further configured to perform at least one of the following processes on a normally expressed statement: deleting some words from the normally expressed statement and treating the statement obtained after deletion as a statement with abnormal expression; replacing some words from the normally expressed statement with irrelevant words and treating the statement obtained after replacement as a statement with abnormal expression.
[0128] It should be noted that the description of the device embodiments in this application is similar to the description of the method embodiments described above, and has similar beneficial effects as the method embodiments; therefore, it will not be repeated. For any technical details not covered in the dialogue processing device provided in the embodiments of this application, please refer to... Figure 3-5 The meaning is understood in accordance with the description of any of the accompanying drawings.
[0129] This application provides a computer program product or computer program that includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the dialogue processing method described above in this application.
[0130] This application provides a computer-readable storage medium storing executable instructions. When these executable instructions are executed by a processor, they cause the processor to perform the method provided in this application, for example... Figure 3 ,or Figure 4 The dialogue processing method is shown.
[0131] In some embodiments, the computer-readable storage medium may be a memory such as FRAM, ROM, PROM, EPROM, EEPROM, flash memory, magnetic surface memory, optical disk, or CD-ROM; or it may be a variety of devices including one or any combination of the above-mentioned memories.
[0132] In some embodiments, executable instructions may take the form of a program, software, software module, script, or code, written in any form of programming language (including compiled or interpreted languages, or declarative or procedural languages), and may be deployed in any form, including as a standalone program or as a module, component, subroutine, or other unit suitable for use in a computing environment.
[0133] As an example, executable instructions may, but do not necessarily, correspond to files in a file system. They may be stored as part of a file that holds other programs or data, for example, in one or more scripts in a Hyper Text Markup Language (HTML) document, in a single file dedicated to the program in question, or in multiple collaborating files (e.g., a file that stores one or more modules, subroutines, or code sections).
[0134] As an example, executable instructions can be deployed to execute on a single computing device, or on multiple computing devices located in one location, or on multiple computing devices distributed across multiple locations and interconnected via a communication network.
[0135] In summary, this application embodiment obtains context variables by co-encoding the input statement and the input image, which integrate the sequence information of the input statement and the emotional information expressed by the input image. Then, the context variables are decoded to obtain a response statement that is used to respond to the input statement and is adapted to the emotional information. In this way, the user's emotions can be detected based on the input image, and appropriate response statements can be generated according to the user's emotions, thereby improving the accuracy of the response statements and enhancing the user experience.
[0136] The above description is merely an embodiment of this application and is not intended to limit the scope of protection of this application. Any modifications, equivalent substitutions, and improvements made within the spirit and scope of this application are included within the scope of protection of this application.
Claims
1. A dialogue processing method, characterized in that, The method includes: Get the input statements and input images from the dialogue; Obtain the sequence information of the input statement; Extract corresponding image features from the input image, and extract object features of the objects included in the input image from the input image; The sequence information of the input statement, the image features of the input image, and the object features of the objects in the input image are fused together to obtain multimodal features; Train the dialogue model using the following methods: Acquire training images and training corpus, wherein the training corpus includes sample input statements and target response statements corresponding to the sample input statements; The dialogue model encodes and decodes the training image and the sample input statement together to obtain a predicted response statement for responding to the sample input statement. The predicted response statement is processed by a sentiment classification model to obtain a sentiment score representing positive sentiment. The predicted response statement is processed by a language fluency model to obtain a fluency score that represents the fluency of the statement. The emotion score, the fluency score, and the loss function of the dialogue model itself are substituted into the second objective function to determine the parameters of the dialogue model when the second objective function reaches its minimum value, and the dialogue model is updated based on the parameters; wherein, the second objective function is used to perform a weighted summation of the emotion score, the fluency score, and the loss function of the dialogue model itself; Before the dialogue model performs joint encoding and decoding processing on the training images and the sample input statements, the language fluency model is trained in the following manner: Obtain sentences that are expressed correctly as positive examples for training the language fluency model; Extract sentences with abnormal expressions and use them as negative examples to train the language fluency model; A training set is constructed based on the positive and negative examples, and the language fluency model is trained based on the training set. The statement for obtaining the expression anomaly includes: Perform at least one of the following processes on the statements that are expressed correctly: Delete some words from the normally expressed statement, and take the statement after deletion as the abnormal statement; Replace some words in the normally expressed statement with irrelevant words, and take the statement after the replacement process as the abnormal statement. The encoder is invoked to encode the multimodal features to obtain context variables that integrate the sequence information of the input statement, the fluency of the input statement, and the emotional information expressed by the input image. The context variables are decoded to obtain the corresponding response words and the selection probability of each response word. Based on the selection probability of each response word, a response statement is obtained that is used to respond to the input statement and is adapted to the emotional information and the fluency. The encoding and decoding processes are implemented through the dialogue model.
2. The method according to claim 1, characterized in that, The step of obtaining the sequence information of the input statement includes: The input statement is segmented into words, and the word embedding vectors corresponding to each word obtained from the segmentation are concatenated. The concatenation result is then used as the sequence information of the input statement. Extracting corresponding image features from the input image includes: The input image is scaled to a fixed size to obtain a standard image; The standard image is subjected to convolution processing, and the resulting convolution features are subjected to pooling processing to obtain pooled features; The pooled features are processed by a fully connected layer to obtain the image features corresponding to the input image. Extracting object features of objects included in the input image from the input image includes: The input image is subjected to object detection processing to obtain the bounding box of each object in the input image; Extract the object features of the corresponding object from the bounding box of each object.
3. A dialogue processing device, characterized in that, The device includes: The acquisition module is used to acquire input statements and input images from the dialogue; The acquisition module is further configured to acquire the sequence information of the input statement; The image representation module is used to extract corresponding image features from the input image; An object representation module in an image is used to extract object features of objects included in the input image from the input image; The encoding module is used to fuse the sequence information of the input statement, the image features of the input image, and the object features of the objects in the input image to obtain multimodal features; The device trains the dialogue model in the following manner: The acquisition module is further configured to acquire training images and training corpus, wherein the training corpus includes sample input statements and target response statements corresponding to the sample input statements; The encoding module and the decoding module are used to encode and decode the training image and the sample input statement together through the dialogue model to obtain a predicted response statement for responding to the sample input statement; The text sentiment classification module is used to perform sentiment prediction processing on the predicted response statement through a sentiment classification model to obtain a sentiment score representing positive sentiment. The language fluency module is used to perform fluency prediction processing on the predicted response statement through the language fluency model to obtain a fluency score that represents the fluency of the statement. An update module is used to substitute the sentiment score, the fluency score, and the loss function of the dialogue model itself into a second objective function to determine the parameters of the dialogue model when the second objective function reaches its minimum value, and to update the dialogue model based on the parameters; wherein, the second objective function is used to perform a weighted summation of the sentiment score, the fluency score, and the loss function of the dialogue model itself; The language fluency module is also used to train the language fluency model in the following ways: Obtain sentences that are expressed correctly as positive examples for training the language fluency model; Extract sentences with abnormal expressions and use them as negative examples to train the language fluency model; A training set is constructed based on the positive and negative examples, and the language fluency model is trained based on the training set. The language fluency module is also used for: Perform at least one of the following processes on the statements that are expressed correctly: Delete some words from the normally expressed statement, and take the statement after deletion as the abnormal statement; Replace some words in the normally expressed statement with irrelevant words, and take the statement after the replacement process as the abnormal statement. The encoding module is also used to call the encoder to encode the multimodal features to obtain context variables that integrate the sequence information of the input statement, the fluency of the input statement, and the emotional information expressed by the input image. The decoding module is further configured to decode the context variable to obtain the response words corresponding to the context variable and the selection probability of each response word, and based on the selection probability of each response word, obtain a response statement that is used to respond to the input statement and is adapted to the emotional information and the fluency. The encoding and decoding processes are implemented through the dialogue model.
Citation Information
Patent Citations
Intelligent dialogue method and system based on emotion and semantics
CN106683672A
Automatic question answering method and system
CN108345692A
Conversation generation method and device, video comment method and device, equipment and storage medium
CN111625660A
Visual dialogue generation system based on semantic alignment
CN111967272A