Text Increment Method, Apparatus and Terminal Device
By extracting the feature matrix of the text to be incremented and determining its theme, and combining with the variational autoencoder to generate incremental text, the problem of excessive randomness of the traditional text generation model is solved, and the relevance and quality of text generation is improved.
Patent Information
- Application Number
- CN202010019294.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-01-08
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2040-01-08
AI Technical Summary
The text generated by the traditional text generation model is too random, resulting in insufficient correlation between incremental text and text to be incremented.
By extracting the feature matrix of the text to be incremented, determining its text theme, and inputting the feature matrix to the Variational Autoencoder (VAE) corresponding to the topic, generating incremental text.
The correlation between incremental text and text to be incremented is improved, the randomness of text generation is reduced, and the quality of text generation is significantly improved.
Smart Images

Figure CN111241815B_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the technical field of natural language processing, and particularly relates to a text increment method, apparatus, terminal device, and computer-readable storage medium. Background Art
[0002] Currently, in many artificial intelligence fields such as question-and-answer systems and machine translation, there is a need to generate other text data based on original text data. For example, in a human-machine question-and-answer system, when a user asks a robot, the robot's answer needs to be relevant to the user's question. That is to say, it is required that the answer text data generated by the robot is associated with the text data asked by the user.
[0003] However, the challenge faced by traditional text generation models is that the generated text has too much randomness. Therefore, there is an urgent need to provide a new text increment solution. Summary of the Invention
[0004] Embodiments of this application provide a text increment method, apparatus, terminal device, and computer-readable storage medium, providing a new text increment solution and improving the relevance between the increment text and the text to be incremented.
[0005] In a first aspect, embodiments of this application provide a text increment method, including:
[0006] Obtain the text to be incremented;
[0007] Extract features from the text to be incremented to obtain a feature matrix corresponding to the text to be incremented;
[0008] Determine the text theme of the text to be incremented;
[0009] Input the feature matrix into a variational autoencoder corresponding to the text theme to obtain an increment text of the text to be incremented.
[0010] In a second aspect, embodiments of this application provide a text increment apparatus, including:
[0011] An obtaining module, configured to obtain the text to be incremented;
[0012] An extraction module, configured to extract features from the text to be incremented to obtain a feature matrix corresponding to the text to be incremented;
[0013] A determination module, configured to determine the text theme of the text to be incremented;
[0014] An increment module, configured to input the feature matrix into a variational autoencoder corresponding to the text theme to obtain an increment text of the text to be incremented.
[0015] In a third aspect, an embodiment of the present application provides a terminal device, including: a memory, a processor, and a computer program stored in the memory and executable on the processor, where when the processor executes the computer program, the text increment method described in the first aspect is implemented.
[0016] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium storing a computer program, where when the computer program is executed by a processor, the text increment method described in the first aspect is implemented.
[0017] In a fifth aspect, an embodiment of the present application provides a computer program product, which when running on a terminal device causes the terminal device to execute the text increment method described in the first aspect.
[0018] In the embodiment of the present application, by first extracting the feature matrix of the text to be incremented to determine the text theme of the text to be incremented, and then combining with the VAE corresponding to the text theme to generate the incremented text. On the one hand, the VAE corresponding to the text theme is used to generate the incremented text, and a different VAE is set for different themes; on the other hand, since the distribution calculated by the VAE depends on the input variables, all samplings of this distribution will generate outputs similar or related to the input, and it can itself help achieve determinism when generating text. Therefore, through the dual effects of these two aspects, the complete randomness when generating text is avoided, the relevance between the incremented text and the text to be incremented is improved, and thus the quality of text generation can be greatly improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] To more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0020] Figure 1 It is a schematic structural diagram of a mobile phone to which the text increment method provided in an embodiment of the present application is applicable;
[0021] Figure 2 It is a schematic flowchart of the text increment method provided in an embodiment of the present application;
[0022] Figure 3 It is a schematic flowchart of step 202 in the text increment method provided in an embodiment of the present application;
[0023] Figure 4 It is a schematic structural diagram of the VAE in the text increment method provided in an embodiment of the present application;
[0024] Figure 5 It is a schematic structural diagram of a text increment device provided by an embodiment of the present application;
[0025] Figure 6 It is a schematic structural diagram of a terminal device to which the text increment method provided by an embodiment of the present application is applicable. Detailed implementation manners
[0026] In the following description, specific details such as specific system structures and technologies are presented for the purpose of illustration rather than limitation, so as to thoroughly understand the embodiments of the present application.
[0027] In order to enable those skilled in the art to better understand the solutions of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, for those of ordinary skill in the art, all other embodiments obtained without creative efforts shall fall within the scope of protection of the present application. It should be noted that, without conflict, the embodiments in the present application and the features in the embodiments may be combined with each other.
[0028] In the following description, specific details such as specific system structures and technologies are presented for the purpose of illustration rather than limitation, so as to thoroughly understand the embodiments of the present application. However, those skilled in the art should understand that the present application can also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid unnecessary details from interfering with the description of the present application.
[0029] It should be understood that when used in the specification of the present application and the appended claims, the term "comprising" indicates the presence of the described features, wholes, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or their combinations.
[0030] It should also be understood that the term " / and" as used in the specification of the present application and the appended claims refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.
[0031] As used in the specification of this application and the appended claims, the term "if" may be construed, depending on the context, as "when", "once", "in response to determining", or "in response to detecting". Similarly, the phrase "if determined" or "if [the described condition or event] is detected" may be construed, depending on the context, to mean "once determined", "in response to determining", "once [the described condition or event] is detected", or "in response to detecting [the described condition or event]".
[0032] In addition, in the description of the specification of this application and the appended claims, the terms "first", "second", "third", etc. are used only for distinguishing descriptions and should not be construed as indicating or implying relative importance.
[0033] The reference to "one embodiment" or "some embodiments" or the like described in the specification of this application means that a specific feature, structure, or characteristic described in connection with that embodiment is included in one or more embodiments of this application. Thus, the statements "in one embodiment", "in some embodiments", "in other some embodiments", "in still other embodiments", etc. that appear in different places in this specification do not necessarily all refer to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized in other ways. The terms "comprising", "including", "having", and their variants all mean "including but not limited to", unless otherwise specifically emphasized in other ways.
[0034] The text increment method provided by the embodiments of this application can be applied to terminal devices such as mobile phones, tablet computers, wearable devices, in-vehicle devices, augmented reality (AR) / virtual reality (VR) devices, laptop computers, ultra-mobile personal computers (UMPCs), netbooks, personal digital assistants (PDAs), or servers. The embodiments of this application do not impose any restrictions on the specific types of terminal devices. Among them, the server includes, but is not limited to, independent servers, cloud servers, distributed servers, and server clusters, etc.
[0035] For example, the terminal device may be a station (STAION, ST) in a WLAN, a cellular phone, a cordless phone, a Session Initiation Protocol (SIP) phone, a Wireless Local Loop (WLL) station, a PDA, a handheld device with wireless communication capabilities, a computing device, or other processing devices connected to a wireless modem, a vehicle-mounted device, a vehicle networking terminal, a computer, a laptop computer, a handheld communication device, a handheld computing device, a satellite wireless device, a wireless modem card, a television set top box (STB), a customer premise equipment (CPE), and / or other devices for communicating on a wireless system, as well as next-generation communication systems, such as a mobile terminal in a 5G network or a mobile terminal in a future evolved Public Land Mobile Network (PLMN) network, etc.
[0036] By way of example and not limitation, when the terminal device is a wearable device, the wearable device may also be a general term for devices that apply wearable technology to the intelligent design of daily wear and develop wearable devices, such as glasses, gloves, watches, clothing, and shoes, etc. A wearable device is a portable device that is either directly worn on the body or integrated into the user's clothes or accessories. A wearable device is not just a hardware device, but also realizes powerful functions through software support, data interaction, and cloud interaction. Broadly speaking, wearable intelligent devices include those with complete functions and large sizes that can realize complete or partial functions without relying on a smart phone, such as smart watches or smart glasses, etc., and those that only focus on a certain type of application function and need to cooperate with other devices such as smart phones, such as various smart bracelets and smart jewelry for physical sign monitoring.
[0037] Taking the terminal device as a mobile phone as an example. Figure 1 Shown is a block diagram of a part of the structure of the mobile phone provided by an embodiment of the present application. Refer to Figure 1 , the mobile phone includes: a Radio Frequency (RF) circuit 110, a memory 120, an input unit 130, a display unit 140, a sensor 150, an audio circuit 160, a wireless fidelity (WiFi) module 170, a processor 180, and a power supply 190, etc. Those skilled in the art can understand that Figure 1 the structure of the mobile phone shown in
[0038] is not a limitation on the mobile phone, and it may include more or fewer components than shown in the figure, or combine some components, or have different component arrangements. Figure 1Specifically introduce each component of the mobile phone:
[0039] The RF circuit 110 can be used for receiving and transmitting signals during information reception or call processes. Specifically, after receiving the downlink information from the base station, it is given to the processor 180 for processing; in addition, the designed uplink data is sent to the base station. Generally, the RF circuit includes but is not limited to antennas, at least one amplifier, a transceiver, a coupler, a low noise amplifier (LNA), a duplexer, etc. In addition, the RF circuit 110 can also communicate with the network and other devices through wireless communication. The above wireless communication can use any communication standard or protocol, including but not limited to the Global System of Mobile communication (GSM), General Packet Radio Service (GPRS), Code Division Multiple Access (CDMA), Wideband Code Division Multiple Access (WCDMA), Long Term Evolution (LTE), email, and Short Messaging Service (SMS), etc.
[0040] The memory 120 can be used to store software programs and modules. The processor 180 executes various functional applications and data processing of the mobile phone by running the software programs and modules stored in the memory 120. The memory 120 mainly includes a program storage area and a data storage area. Among them, the program storage area can store the operating system, application programs required for at least one function (such as the sound playback function, image playback function, etc.), a boot loader, etc.; the data storage area can store data created according to the use of the mobile phone (such as audio data, phone book, etc.). In addition, the memory 120 can include high-speed random access memory and can also include non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other non-volatile solid-state storage devices. It can be understood that in the embodiments of the present application, a program for text increment is stored in the memory 120.
[0041] The input unit 130 can be used to receive input numerical or character information and generate key signal inputs related to the user settings and function controls of the mobile phone 100. Specifically, the input unit 130 can include a touch panel 131 and other input devices 132. The touch panel 131, also known as a touch screen, can collect touch operations of the user on or near it (such as operations of the user using any suitable object or accessory like a finger, a stylus, etc. on or near the touch panel 131), and drive corresponding connection devices according to a pre-set program. Optionally, the touch panel 131 can include two parts: a touch detection device and a touch controller. Among them, the touch detection device detects the touch position of the user and detects the signal brought by the touch operation, and transmits the signal to the touch controller; the touch controller receives the touch information from the touch detection device, converts it into contact coordinates, and then sends it to the processor 180, and can receive and execute commands sent by the processor 180. In addition, various types such as resistive, capacitive, infrared, and surface acoustic wave can be used to implement the touch panel 131. In addition to the touch panel 131, the input unit 130 can also include other input devices 132. Specifically, the other input devices 132 can include, but are not limited to, one or more of a physical keyboard, function keys (such as volume control keys, switch keys, etc.), a trackball, a mouse, a joystick, etc.
[0042] The display unit 140 can be used to display information input by the user or information provided to the user and various menus of the mobile phone. The display unit 140 can include a display panel 141. Optionally, the display panel 141 can be configured in the form of a liquid crystal display (LCD), an organic light-emitting diode (OLED), etc. Further, the touch panel 131 can cover the display panel 141. When the touch panel 131 detects a touch operation on or near it, it is transmitted to the processor 180 to determine the type of touch event. Subsequently, the processor 180 provides corresponding visual output on the display panel 141 according to the type of touch event. Although in Figure 1 the touch panel 131 and the display panel 141 are implemented as two independent components to realize the input and input functions of the mobile phone, in some embodiments, the touch panel 131 and the display panel 141 can be integrated to realize the input and output functions of the mobile phone.
[0043] The mobile phone 100 may also include at least one sensor 150, such as a light sensor, a motion sensor, and other sensors. Specifically, the light sensor may include an ambient light sensor and a proximity sensor. Among them, the ambient light sensor can adjust the brightness of the display panel 141 according to the brightness of the ambient light, and the proximity sensor can turn off the display panel 141 and / or the backlight when the mobile phone is moved to the ear. As a kind of motion sensor, the accelerometer sensor can detect the magnitude of acceleration in various directions (generally three axes), and can detect the magnitude and direction of gravity when stationary, and can be used in applications for identifying the posture of the mobile phone (such as horizontal and vertical screen switching, related games, magnetometer posture calibration), vibration recognition related functions (such as pedometer, tapping), etc.; as for other sensors such as gyroscopes, barometers, hygrometers, thermometers, and infrared sensors that the mobile phone can also be configured with, they will not be elaborated here.
[0044] The audio circuit 160, the speaker 161, and the microphone 162 can provide an audio interface between the user and the mobile phone. The audio circuit 160 can transmit the electrical signal converted from the received audio data to the speaker 161, and the speaker 161 converts it into a sound signal for output; on the other hand, the microphone 162 converts the collected sound signal into an electrical signal, which is received by the audio circuit 160 and then converted into audio data. After the audio data is output to the processor 180 for processing, it is sent through the RF circuit 110 to, for example, another mobile phone, or the audio data is output to the memory 120 for further processing.
[0045] WiFi belongs to short - range wireless transmission technology. The mobile phone can help users send and receive emails, browse the web, and access streaming media through the WiFi module 170. It provides users with wireless broadband Internet access. Although Figure 1 the WiFi module 170 is shown, it can be understood that it does not belong to an essential component of the mobile phone 100 and can be omitted completely within the scope of not changing the essence of the invention according to needs.
[0046] The processor 180 is the control center of the mobile phone, connecting various parts of the entire mobile phone through various interfaces and circuits. By running or executing software programs and / or modules stored in the memory 120, and by calling the data stored in the memory 120, it executes various functions of the mobile phone and processes data, thereby monitoring the mobile phone as a whole. Optionally, the processor 180 may include one or more processing units; preferably, the processor 180 may integrate an application processor and a modem processor. Among them, the application processor mainly processes the operating system, user interface, application programs, etc., and the modem processor mainly processes wireless communications. It can be understood that the above-mentioned modem processor may not be integrated into the processor 180. It can be understood that in the embodiments of the present application, a program for text increment is stored in the memory 120, and the processor 180 can be used to call and execute the program for text increment stored in the memory 120 to implement the text increment method of the embodiments of the present application.
[0047] The mobile phone 100 further includes a power supply 190 (such as a battery) for supplying power to each component. Preferably, the power supply can be logically connected to the processor 180 through a power management system, so as to implement functions such as management of charging, discharging, and power consumption management through the power management system.
[0048] Although not shown, the mobile phone 100 may further include a camera. Optionally, the position of the camera on the mobile phone 100 can be front-facing, rear-facing, or built-in (which can extend out of the body when in use). The embodiments of the present application do not make any limitations in this regard.
[0049] Optionally, the mobile phone 100 may include a single camera, a dual camera, or a triple camera, etc. The embodiments of the present application do not make any limitations in this regard. The camera includes but is not limited to a wide-angle camera, a telephoto camera, or a depth camera, etc.
[0050] For example, the mobile phone 100 may include a triple camera, among which, one is the main camera, one is the wide-angle camera, and one is the telephoto camera.
[0051] Optionally, when the mobile phone 100 includes multiple cameras, these multiple cameras may all be front-facing, or all be rear-facing, or all be built-in, or at least part of them be front-facing, or at least part of them be rear-facing, or at least part of them be built-in, etc. The embodiments of the present application do not make any limitations in this regard.
[0052] In addition, although not shown, the mobile phone 100 may further include a Bluetooth module, etc., which will not be elaborated here.
[0053] Figure 2The figure shows an implementation flowchart of a text increment method provided by an embodiment of the present application. The text increment method is applied to a terminal device. By way of example and not limitation, the text increment method may be applied to the mobile phone 100 having the above hardware structure. The text increment method includes steps S201 to S204, and the specific implementation principles of each step are as follows.
[0054] S201, obtain the text to be incremented.
[0055] In an embodiment of the present application, the text to be incremented is the object for text increment, such as sentence text, etc.
[0056] The text to be incremented may be text instantaneously input by a user through an input unit of the terminal device; it may also be voice data instantaneously collected by the user through an audio collection unit of the terminal device; it may also be a picture including text instantaneously captured by the user through a camera of the terminal device; it may also be a picture including text instantaneously scanned by the user through a scanning device of the terminal device; it may also be text originally stored in the terminal device; or even text obtained by the terminal device from other terminal devices through a wired or wireless network, etc.
[0057] It should be noted that for a picture including text, the text in the picture needs to be extracted as the text to be incremented by enabling the picture recognition function of the terminal device; for voice data, the text in the voice data needs to be recognized as the text to be incremented by starting the audio-to-text function of the terminal device.
[0058] In a non-limiting usage scenario of the present application, when a user collects a piece of voice data input by the user through the audio collection unit of the terminal device, enables the audio-to-text function, and obtains the text input by the user, if the user wants to perform text increment at this time, the user can enable the text increment function of the terminal device by clicking a specific physical button or virtual button of the terminal device. In this mode, the terminal device will automatically process the text input by the user according to the processes of steps S202 to S204 to obtain the incremented text. It should be noted here that the order of the user inputting text and clicking the button can be interchanged, that is, the button can also be clicked first, then the text input by the user is obtained, and finally the text input by the user is automatically processed according to the processes of steps S202 to S204.
[0059] In another non-limiting usage scenario of the present application, when a user wants to increment the text already stored in the terminal device, the text increment function of the terminal device can be enabled by clicking a specific physical button or virtual button, and the text to be incremented is selected. Then, the terminal device will automatically process the selected text to be incremented according to the process from step S202 to step S204 to obtain the incremented text. It should be noted here that the order of clicking the button and selecting the text can be interchanged, that is, the text can also be selected first, and then the text increment function of the terminal device is enabled.
[0060] S202. Extract features from the text to be incremented to obtain a feature matrix corresponding to the text to be incremented.
[0061] Step S202 is a step of extracting features from the text to be incremented, obtaining a feature matrix corresponding to the text to be incremented, and realizing representing the text by a low-dimensional matrix.
[0062] In some embodiments of the present application, the text to be incremented can be extracted by a word vector model to obtain a feature matrix corresponding to the text to be incremented. That is to say, the text to be incremented is converted into a feature matrix by the word vector model.
[0063] The word vector model includes, but is not limited to, models such as word2vec (word to vector), ELMo, and BERT (Bidirectional Encoder Representation from Transformers). In the embodiments of the present application, through step S202, using the word vector model, the text that abstractly exists in the real world is converted into a vector or matrix that can perform mathematical formula operations. The data is processed into data that can be processed by a machine, enabling the present application to be implemented.
[0064] It should be noted that before using the word vector model, the training of the word vector model needs to be completed to pre-train and generate word vectors. In addition, during the training process of the word vector model, in order to obtain more accurate feature extraction results, punctuation marks in the text to be incremented can be retained, and the complete text to be incremented is extracted.
[0065] Exemplarily, during the training process of the BERT model, in order to perform unsupervised pre-training on a large dataset, 15% of the words in the training sentences used for training are randomly selected as the words to be masked during the training process. Such a masking design is to enable the BERT model to fill in the masked positions and achieve unsupervised training.
[0066] As a non-limiting example of the present application, step S202 includes: extracting features of the text to be incremented through a preset BERT model to obtain a feature matrix corresponding to the text to be incremented.
[0067] Exemplarily, the text to be incremented is converted into a feature matrix of N×768 dimensions through a preset BERT model. The preset BERT model includes 24 encoding layers, that is, the BERT Large model is adopted, and the number of Transformer blocks in this model is 24; wherein, the text to be incremented includes N characters, and N is a positive integer.
[0068] Each row in the feature matrix corresponds to a foreign character (or a Chinese character) of the text to be incremented. Obviously, the feature matrix can reflect the semantic features of the text to be incremented.
[0069] For example, the text to be incremented is "I love my motherland!", and the number of Chinese characters including punctuation marks is 7, and the number of Chinese characters excluding punctuation marks is 6.
[0070] Another exemplarily, a BERT model including 12 encoding layers is adopted, that is, the BERT Base model, and the number of Transformer blocks in this model is 12.
[0071] As another non-limiting example of the present application, as Figure 3 shown, step S202 includes steps S2021 to S2023.
[0072] S2021, obtain the keywords of the text to be incremented.
[0073] Among them, for the text to be incremented, word segmentation and part-of-speech tagging are first performed, then stop words are removed according to a preset stop word dictionary, and according to the part-of-speech of the words after word segmentation, non-feature words such as prepositions, locative words, and modal particles are removed to obtain a keyword set of the text to be incremented.
[0074] By obtaining the keywords of the text to be incremented in step S2021, some noise data is filtered, while ensuring the result accuracy, the data volume is also appropriately reduced, the processing efficiency is improved, the system resource occupation is reduced, and the computing power cost is reduced.
[0075] S2022, obtain the feature vector corresponding to each keyword.
[0076] Among them, the terminal device pre-stores the corresponding relationship between keywords and feature vectors, and by looking up the corresponding relationship between keywords and feature vectors, the feature vector corresponding to each keyword is obtained.
[0077] It should be noted that before step S2022, a correspondence between keywords and feature vectors is established in advance. The method for establishing the correspondence is as follows:
[0078] First, use web crawler technology to crawl corpora from various channels and organize them into a document collection.
[0079] Then, use an open-source word segmentation tool to perform word segmentation and part-of-speech tagging on each document. Then, remove stop words according to a preset stop word dictionary, and remove non-feature words such as prepositions, locative words, and modal particles according to the part-of-speech of the words after word segmentation, to obtain a keyword set.
[0080] Finally, use the open-source word vector training tool Word2Vec to train the above keyword set to obtain feature vectors corresponding to different keywords, and store the correspondence between keywords and feature vectors in a word vector database. Exemplarily, each feature vector has the same dimension, and N-dimensional (N is a positive integer) word vectors are used, and the values of each word vector are all between 0 and 1, or between -1 and 1.
[0081] The correspondence between keywords and feature vectors is established through the above method. By looking up the correspondence, the feature vector corresponding to a keyword can be obtained, so as to convert each keyword into a feature vector.
[0082] S2023, combine the feature vectors corresponding to all the keywords to generate a feature matrix.
[0083] Among them, combining the feature vectors corresponding to all keywords means concatenating the feature vectors of all keywords to generate a feature matrix.
[0084] For example, when the feature vector is 1×N dimension and the preset quantity is M (M is a positive integer), the feature matrix obtained by combining M 1×N-dimensional feature vectors can be M×N dimension, or 1×(M + N) dimension.
[0085] In some embodiments of the present application, a deep learning network model is used to extract features from the text to be incremented, and a feature matrix corresponding to the text to be incremented is obtained.
[0086] The deep learning network model is used to extract features of the text to be incremented. When the text to be incremented is input into the deep learning network model, the deep learning network model outputs a feature matrix corresponding to the text to be incremented. The deep learning network model can be a deep learning network model based on machine learning technology, including but not limited to a deep convolutional neural network model, and a deep residual convolutional neural network model (Res Net), etc. Among them, the deep convolutional neural network model includes but not limited to AlexNet, VGG-Net, and DenseNet, etc.
[0087] It can be understood that before using a deep learning network model, it is necessary to complete the training of the deep learning network model. During the process of training the deep learning network model, the loss function adopted can be one of the 0-1 loss function, absolute value loss function, logarithmic loss function, exponential loss function, and hinge loss function, or a combination of at least two of them.
[0088] It should be noted that the process of training the model, including the process of training the deep learning network model and the word vector model, can be implemented on the terminal device or on other terminal devices that are communicatively connected to the terminal device. After the terminal device stores the trained model, or other terminal devices push the trained model to the terminal device, the feature extraction of the obtained text to be incremented can be realized on the terminal device. It should be noted that the text to be incremented obtained by the terminal device during the text increment process can also be used to increase the data in the sample database for training the model, perform further optimization of the model on the terminal device or other terminal devices, and then the terminal device or other terminal devices store the further optimized model in the terminal device to replace the previous model. By optimizing the model in this way, the data breadth of the model is increased, thereby improving the applicable scope of the solution of the present application.
[0089] S203. Determine the text theme of the text to be incremented.
[0090] In step 203, the text theme of the text to be incremented is determined, so as to realize the increment of the feature matrix through the variational autoencoder (VAE) corresponding to the text theme in the subsequent step S204.
[0091] In some embodiments of the present application, a document topic generation model (Latent Dirichlet Allocation, LDA) is used to identify the text theme of the text to be incremented. LDA is an unsupervised machine learning technique that can be used to identify latent topic information in a large-scale document collection or corpus.
[0092] It can be understood that this is only an exemplary illustration here and cannot be construed as a specific limitation of the present application. All methods that can realize the determination of the text theme of the text to be incremented are applicable to the present application.
[0093] It should be noted that although there is a sequence in the description between step S202 and step S203, and there is also a difference in the labels, neither the sequence in the description nor the difference in the labels represents a specific limitation on the chronological relationship of the steps. In the embodiments of the present application, step S202 can be executed before step S203, can also be executed after step S203, or can be executed simultaneously with step S203. The present application does not specifically limit the chronological relationship between step S202 and S203.
[0094] S204, input the feature matrix into the variational autoencoder (VAE) corresponding to the text theme to obtain the incremental text of the text to be incremented.
[0095] In the embodiments of the present application, the terminal device pre-stores multiple VAEs, and each VAE corresponds to a text theme. After determining the text theme of the text to be incremented in step 203, the VAE corresponding to the text theme of the text to be incremented is determined from the pre-stored multiple VAEs, so as to perform increment on the text to be incremented based on the determined VAE.
[0096] Input the feature matrix into the VAE corresponding to the text theme to obtain the incremental text of the text to be incremented. Performing text increment based on the VAE corresponding to the text theme of the text to be incremented greatly improves the relevance between the incremental text and the text to be incremented, and significantly improves the accuracy of text generation.
[0097] As Figure 4 shown, the VAE consists of two parts, including an encoder and a decoder. The encoder of the VAE does not directly output the encoding, but assumes that all encodings conform to a normal distribution. The mean and variance calculation module of the encoder calculates the mean and variance of the normal distribution. Based on the mean and variance, a normal distribution can be determined, and a sampled encoding is obtained by sampling from the determined normal distribution. Then, this sampled encoding is input into the generator of the decoder to generate incremental text data. That is to say, in the embodiments of the present application, it can be considered that each text to be incremented corresponds to an encoding in the normal distribution. First, estimate this normal distribution through the existing training data, and then only need to sample from the normal distribution to obtain a new encoding to generate incremental text data.
[0098] As a non - restrictive example of this application, a relatively simple Recurrent Neural Network (RNN) is used as the encoder and decoder. The encoder receives a feature matrix as input and outputs the variance and mean. The decoder determines a normal distribution based on the variance and mean, and samples within the normal distribution to obtain a sampled encoding. The sampled encoding vector obtained by sampling from the normal distribution is input at each time step of the RNN of the decoder. In this way, the output at each time step, after passing through a fully - connected layer and a softmax function, generates the probability of each word appearing at this position. We select the word with the highest probability as the word appearing at this time step. It should be noted that if the actual length of the generated text does not have that many time steps, the part exceeding the length will generate a preset character to represent padding.
[0099] Exemplarily, for the N×768 - dimensional feature matrix output by the BERT Large model mentioned above, after the encoder receives this feature matrix, it returns a vector with a dimension of 1×256; this vector is then respectively connected to two fully - connected layers, and the two fully - connected layers respectively output two vectors with a size of 1×256, and these two outputs are the mean and variance. Based on the mean and variance, a normal distribution is determined, and sampling is performed within the normal distribution to obtain a sampled encoding. The sampled encoding is added with a variance and then input into the decoder, and the decoder generates incremental text character by character. It should be noted that the embodiments of this application can generate incremental text because the variance is added to the sampled encoding, so completely identical incremental text will not be generated.
[0100] In the above example, on the one hand, the high - dimensional text vector output by the BERT model contains extremely rich information, which is very suitable for the encoder of VAE to process it into semantic encoding. On the other hand, the distribution calculated by VAE depends on the input variables, and all samplings of this distribution will generate outputs similar or related to the input. It can itself help to achieve determinism when generating text. Therefore, through the dual role of the combination of the BERT model and VAE, the randomness when generating text is avoided, and the relevance between the incremental text and the text to be incremented is greatly improved, thus significantly enhancing the quality of text generation.
[0101] It can be understood that before using VAE for text increment, the training of VAE needs to be completed.
[0102] In a non - restrictive example of this application, after obtaining a large - scale corpus for training the model, the corpus in the corpus is first classified by text theme, and then for the corpus of each category, a VAE is trained respectively, so as to obtain multiple VAEs corresponding to different text themes.
[0103] In another non-limiting example of the present application, after obtaining a large-scale corpus for training a model, first train a basic VAE based on the corpus in the corpus; then, on the basis of text topic classification of the corpus in the corpus, for the corpus of each category, retrain based on the basic VAE to obtain a VAE, thereby obtaining multiple VAEs corresponding to different text topics.
[0104] It can be understood that in the above two non-limiting examples, in order to improve the accuracy of the VAE text increment result, for each text topic, there is a corresponding large-scale corpus in the corpus.
[0105] In the embodiment of the present application, first extract the feature matrix of the text to be incremented, determine the text topic of the text to be incremented, and then combine the VAE corresponding to the text topic to generate the incremented text. On the one hand, use the VAE corresponding to the text topic to generate the incremented text, and set a different VAE for different topics; on the other hand, since the distribution calculated by the VAE depends on the input variables, all samplings of this distribution will generate outputs similar or related to the input, and it can itself help to achieve determinacy when generating text. Therefore, through the dual effects of these two aspects, the complete randomness when generating text is avoided, the relevance between the incremented text and the text to be incremented is improved, and thus the quality of text generation can be greatly improved.
[0106] It should be understood that the magnitudes of the sequence numbers of the steps in the above embodiments do not mean the order of execution. The order of execution of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application.
[0107] Corresponding to the text increment method described in the above embodiments, Figure 5 The structural block diagram of the text increment device provided by the embodiment of the present application is shown. For the sake of convenience of description, only the parts related to the embodiment of the present application are shown.
[0108] Referring to Figure 5 , the device includes:
[0109] An acquisition module 51, configured to acquire the text to be incremented;
[0110] An extraction module 52, configured to perform feature extraction on the text to be incremented to obtain a feature matrix corresponding to the text to be incremented;
[0111] A determination module 53, configured to determine the text topic of the text to be incremented;
[0112] An increment module 54, configured to input the feature matrix into a variational autoencoder corresponding to the text topic to obtain an incremented text of the text to be incremented.
[0113] Among them, the extraction module 52 is specifically configured to:
[0114] Convert the text to be incremented into a feature matrix through a preset word vector model.
[0115] Among them, the extraction module 52 is specifically configured to:
[0116] Extract features from the text to be incremented through a preset BERT model to obtain a feature matrix corresponding to the text to be incremented.
[0117] Among them, the extraction module 52 is specifically configured to:
[0118] Convert the text to be incremented into an N×768-dimensional feature matrix through a preset BERT model, and the preset BERT model includes 24 encoding layers; where the text to be incremented includes N characters, and N is a positive integer.
[0119] Among them, the extraction module 52 is specifically configured to:
[0120] Obtain the keywords of the text to be incremented;
[0121] Obtain the feature vectors corresponding to each of the keywords;
[0122] Combine the feature vectors corresponding to all the keywords to generate a feature matrix.
[0123] Among them, the increment module 54 is specifically configured to:
[0124] Input the feature matrix into the encoder of the variational autoencoder corresponding to the text theme to obtain the mean and variance of the feature matrix;
[0125] Determine a normal distribution according to the mean and the variance, and sample from the normal distribution to obtain a sampled code;
[0126] Input the sampled code into the decoder of the variational autoencoder to generate an increment text of the text to be incremented.
[0127] It should be noted that for the information interaction, execution process, etc. between the above modules / units, since they are based on the same concept as the method embodiment of the present application, their specific functions and the technical effects brought about can be specifically referred to in the method embodiment part, and will not be elaborated here.
[0128] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the above-mentioned division of each functional unit and module is used as an example. In actual applications, the above-mentioned functions can be allocated to different functional units and modules according to needs, that is, the internal structure of the device is divided into different functional units or modules to complete all or part of the functions described above. Each functional unit and module in the embodiments can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of a software functional unit. In addition, the specific names of each functional unit and module are only for the convenience of mutual distinction and do not limit the protection scope of this application. The specific working processes of the units and modules in the above-mentioned system can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated here.
[0129] Figure 6 FIG. is a schematic structural diagram of a terminal device provided by an embodiment of the present application. As Figure 6 shown, the terminal device 6 of this embodiment includes: at least one processor 60 ( Figure 6 only one processor is shown in the figure), a memory 61, and a computer program 62 stored in the memory 61 and executable on the at least one processor 60. When the processor 60 executes the computer program 62, the steps in the foregoing method embodiments are implemented.
[0130] An embodiment of the present application further provides a computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, the steps in the foregoing method embodiments can be implemented.
[0131] An embodiment of the present application provides a computer program product, and when the computer program product runs on a mobile terminal, the mobile terminal can implement the steps in the foregoing method embodiments when executed.
[0132] When the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, to implement all or part of the processes in the above-described embodiment methods of this application, a computer program can be used to instruct relevant hardware to complete. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above-described method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file, or some intermediate form, etc. The computer-readable medium can at least include: any entity or device that can carry the computer program code to the photographing device / terminal device, recording medium, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, and software distribution medium. For example, a USB flash drive, a mobile hard disk, a magnetic disk, or an optical disc, etc. In some jurisdictions, according to legislation and patent practice, the computer-readable medium cannot be an electrical carrier signal and a telecommunication signal.
[0133] In the above embodiments, the descriptions of the respective embodiments have their own emphases. For the parts not detailed or recorded in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0134] Those of ordinary skill in the art can realize that the units and algorithm steps of the examples described in conjunction with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods for each specific application to implement the described functions, but such implementation should not be considered to exceed the scope of this application.
[0135] In the embodiments provided in this application, it should be understood that the disclosed device / network device and method can be implemented in other ways. For example, the device / network device embodiments described above are merely illustrative. For example, the division of the modules or units is only a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of the devices or units can be in electrical, mechanical, or other forms.
[0136] The unit described as a separation component may or may not be physically separated. The component shown as a unit may or may not be a physical unit, that is, it may be located in one place or distributed over multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0137] The above-described embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the various embodiments of the present application, and should all be included in the protection scope of the present application.
Claims
1. A text increment method for a human-machine question-answering system, characterized in that, Including: Obtain the text to be incremented; The text to be incremented is the text asked by the user; Extract features from the text to be incremented to obtain the feature matrix corresponding to the text to be incremented; Determine the text theme of the text to be incremented; Input the feature matrix into the variational autoencoder corresponding to the text theme to obtain the incremented text of the text to be incremented; wherein, the incremented text is generated as an answer text according to the text to be incremented; Before inputting the feature matrix into the variational autoencoder corresponding to the text theme, the method further includes: Obtain the corpus for training the variational autoencoder; Train the basic variational autoencoder based on the corpus in the corpus; Perform text theme classification on the corpus in the corpus to obtain corpora of multiple categories; For each category of corpus, retrain based on the basic variational autoencoder to obtain the variational autoencoder corresponding to the text theme of each category of corpus.
2. The text increment method according to claim 1, characterized in that, Extracting features from the text to be incremented to obtain the feature matrix corresponding to the text to be incremented includes: Convert the text to be incremented into a feature matrix through a preset word vector model.
3. The text increment method according to claim 2, characterized in that, The converting the text to be incremented into a feature matrix through a preset word vector model includes: Extract features from the text to be incremented through a preset BERT model to obtain the feature matrix corresponding to the text to be incremented.
4. The text increment method according to claim 3, characterized in that, The extracting features from the text to be incremented through a preset BERT model to obtain the feature matrix corresponding to the text to be incremented includes: Convert the text to be incremented into an N×768-dimensional feature matrix through a preset BERT model, and the preset BERT model includes 24 encoding layers; wherein, the text to be incremented includes N characters, and N is a positive integer.
5. The text increment method according to claim 2, characterized in that, Converting the text to be incremented into a feature matrix through a preset word vector model includes: Obtain the keywords of the text to be incremented; Obtain the feature vectors corresponding to each keyword; Combine the feature vectors corresponding to all the keywords to generate a feature matrix.
6. The text increment method according to claim 1, characterized in that, The inputting the feature matrix into the variational autoencoder corresponding to the text theme to obtain the incremented text of the text to be incremented includes: Input the feature matrix into the encoder of the variational autoencoder corresponding to the text theme to obtain the mean and variance of the feature matrix; Determine a normal distribution according to the mean and the variance, and sample from the normal distribution to obtain a sampled code; Input the sampled code into the decoder of the variational autoencoder to generate the incremented text of the text to be incremented.
7. A text increment device for a human-machine question-answering system, characterized in that, Including: An acquisition module for obtaining the text to be incremented; The text to be incremented is the text asked by the user; An extraction module for extracting features from the text to be incremented to obtain the feature matrix corresponding to the text to be incremented; A determination module for determining the text theme of the text to be incremented; An increment module for inputting the feature matrix into the variational autoencoder corresponding to the text theme to obtain the incremented text of the text to be incremented; wherein, the incremented text is generated as an answer text according to the text to be incremented; The increment module is further used for: Obtain a corpus for training a variational autoencoder; Train a basic variational autoencoder based on the corpus in the corpus; Perform text topic classification on the corpus in the corpus to obtain corpora of multiple categories; For each category of corpus, retrain based on the basic variational autoencoder to obtain a variational autoencoder corresponding to the text topic of each category of corpus.
8. The text increment device according to claim 7, characterized in that, The extraction module is specifically configured to: Extract features from the text to be incremented through a preset BERT model to obtain a feature matrix corresponding to the text to be incremented.
9. A terminal device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, the text increment method described in any one of claims 1 to 6 is implemented.
10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, the text increment method described in any one of claims 1 to 6 is implemented.
Citation Information
Patent Citations
A title generation method based on a variational neural network topic model
CN108984524A