Text generation method and device, equipment and storage medium
By using global encoder and local encoder in the text generation model to extract semantic features and fusion, the problem of logical incoherence or duplication of content when generating long text is solved, efficient text generation is achieved and deployment cost is reduced.
Patent Information
- Application Number
- CN202510017001.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-12-31
- Filing Date
- 2025-01-06
- Publication Date
- 2025-05-06
AI Technical Summary
Existing text generation models may have logical incoherence or duplication of content when generating long text, and require additional algorithms for post-processing, increasing deployment costs and processing time.
By obtaining the source text and sorting word segmentation, converting it into word embedding representation based on the word embedding matrix, semantic features are extracted using a global encoder and a local encoder, and feature fusion and filtering are performed through the global gated unit to obtain a context semantic vector, and finally the target text is generated by the decoder.
This method does not require post-processing algorithms to perform operations such as deduplication, reduces cost and deployment time, and can generate text with higher consistency and semantic accuracy, solving the problems of logical incoherence or duplication of content.
Smart Images

Figure CN119940310A_ABST
Abstract
Description
Technical Field
[0001] The present application belongs to the technical field of text generation models, and specifically relates to a text generation method, device, equipment and storage medium. Background Art
[0002] The text generation task is an important research field in natural language processing (NLP). Its goal is to generate text that conforms to semantic logic and grammatical rules based on input content. With the rise of deep learning, neural network-based methods have become the mainstream technical route in the field of text generation.
[0003] Existing text generation models may have logical incoherence or repeated content when generating long texts. The main reason is that the decoder has limited ability to remember previous content during generation, which makes it difficult to maintain a high level of consistency and semantic accuracy in long text tasks. In addition, the generated text often requires additional algorithms for post-processing such as deduplication and grammar correction, which increases deployment costs and processing time. Summary of the invention
[0004] The purpose of this application is to provide a text generation method, device, equipment and storage medium to solve the technical problem that the text generation model in the prior art may have logical incoherence or content duplication when generating long texts, and requires additional algorithms for post-processing, which increases deployment costs and processing time.
[0005] In order to achieve the above-mentioned purpose, the first aspect of the present application provides a text generation method, comprising:
[0006] Get the source text, perform word segmentation and sorting, and obtain a text sequence;
[0007] Based on the word embedding matrix, convert the text sequence into a word embedding representation;
[0008] Extracting semantic features from the word embedding representation based on a global encoder including n encoding layers to obtain n output vectors, and fusing the output vectors of the last several encoding layers based on an attention mechanism to obtain first semantic encoding information;
[0009] Extracting semantic features from the word embedding representation based on a local encoder to obtain second semantic encoding information;
[0010] Based on the global gating unit, feature fusion and filtering are performed on the first semantic encoding information and the second semantic encoding information to obtain a context semantic vector;
[0011] The context semantic vector is decoded based on a decoder to generate a target text.
[0012] In one or more implementations, the step of converting the text sequence into a word embedding representation based on a word embedding matrix is specifically as follows:
[0013] E=W E X
[0014] In the formula, word embedding represents E = (e 1 ,e 2 …,e n ), text sequence Word Embedding Matrix voc is the vocabulary size and emb is the dimension of the word embedding layer.
[0015] In one or more embodiments, the step of extracting semantic features from the word embedding representation based on a global encoder including n encoding layers to obtain n output vectors, and fusing the output vectors of the last several encoding layers based on an attention mechanism includes:
[0016] Each encoding layer linearly transforms the output vector of the previous encoding layer based on the multi-head self-attention mechanism to obtain the query vector, key vector and value vector, calculates the attention vector based on the query vector and key vector, and obtains the output vector based on the attention vector and value vector;
[0017] Based on the attention mechanism, different weights are assigned to the output vectors of the last several encoding layers;
[0018] The output vectors of the last several coding layers are fused based on the weights to obtain the first semantic coding information.
[0019] In one or more embodiments, the local encoder includes a Bi-LSTM, a local convolution extractor, and a gating unit;
[0020] The step of extracting local semantic features from the word embedding representation based on a local encoder to obtain second semantic encoding information comprises:
[0021] Extracting global semantic information from the word embedding representation based on Bi-LSTM to obtain time series information;
[0022] Based on the local convolution extractor, convolution operations are performed on the word embedding representation using convolution kernels of different sizes to learn n-gram features of different sizes to obtain multiple convolution outputs;
[0023] The multiple convolution outputs are concatenated, and then vector optimization is performed through residual connection to obtain a feature extraction vector;
[0024] The time series information and the feature extraction vector are fused based on the gating unit to obtain second semantic coding information.
[0025] In one or more embodiments, the step of concatenating the multiple convolution outputs and then performing vector optimization through residual connection to obtain the feature extraction vector is specifically as follows:
[0026] H T =σ(W T T+b T )+H
[0027] In the formula, σ is the sigmoid function, H is the hidden layer state, and W T and b T is a learnable weight matrix, T is a vector obtained by concatenating multiple convolution outputs;
[0028] The step of fusing the time series information and the feature extraction vector based on the gating unit to obtain the second semantic coding information is specifically as follows:
[0029] R=(W T H T +b T )⊙(W B B+b B )
[0030] Where W T , b T , W B and b B is the learnable weight matrix, H T is the feature extraction vector, and B is the time series information.
[0031] In one or more embodiments, the step of performing feature fusion and filtering on the first semantic encoding information and the second semantic encoding information based on the global gating unit to obtain a context semantic vector includes:
[0032] splicing the first semantic coding information and the second semantic coding information to obtain spliced coding information;
[0033] Performing weight calculation and nonlinear calculation on the spliced coding information through a global gating unit to obtain a tube selection probability;
[0034] The first semantic coding information and the second semantic coding information are weighted based on the control selection probability, and the first semantic coding information and the second semantic coding information are fused based on the weights assigned to obtain a context semantic vector.
[0035] In one or more embodiments, the step of performing weight calculation and nonlinear calculation on the spliced coding information through the global gating unit to obtain the control selection probability is specifically as follows:
[0036] g t=σ(W g H')
[0037] In the formula, g t is the probability of tube selection, W g is a learnable weight matrix, σ is a sigmoid function, H' is the concatenated coding information, that is, H'=concat(G,R), G is the first semantic coding information, and R is the second semantic coding information;
[0038] The calculation method of the context semantic vector is as follows:
[0039] O t =(1-g t )G+g t R
[0040] In the formula, Q t is the context semantic vector.
[0041] In order to achieve the above-mentioned purpose, the second aspect of the present application provides a text generation device, comprising:
[0042] The word segmentation module is used to obtain the source text, perform word segmentation and sorting, and obtain a text sequence;
[0043] An embedding conversion module, used for converting the text sequence into a word embedding representation based on a word embedding matrix;
[0044] A first feature extraction module is used to extract semantic features from the word embedding representation based on a global encoder including n encoding layers to obtain n output vectors, and to fuse the output vectors of the last several encoding layers based on an attention mechanism to obtain first semantic encoding information;
[0045] A second feature extraction module is used to extract semantic features from the word embedding representation based on a local encoder to obtain second semantic encoding information;
[0046] A global gating module, used for performing feature fusion and filtering on the first semantic encoding information and the second semantic encoding information based on a global gating unit to obtain a context semantic vector;
[0047] A decoding module is used to decode the context semantic vector based on a decoder to generate a target text.
[0048] In order to achieve the above-mentioned object, the third aspect of the present application provides an electronic device, including:
[0049] at least one processor; and
[0050] A memory storing instructions, which, when executed by the at least one processor, enables the at least one processor to execute the text generation method as described in any of the above embodiments.
[0051] In order to achieve the above-mentioned purpose, the fourth aspect of the present application provides a machine-readable storage medium, which stores executable instructions, and when the instructions are executed, the machine executes the text generation method as described in any of the above-mentioned embodiments.
[0052] Different from the prior art, the beneficial effects of this application are:
[0053] In the text generation method of the present application, the context semantic vector is obtained by fusing and filtering the semantic features extracted by the global encoder and the local encoder, and redundant features are removed. Therefore, there is no need for a post-processing algorithm to perform deduplication operations, which reduces costs on the one hand and deployment time on the other.
[0054] In the text generation method of the present application, the global encoder retains more comprehensive semantic features by fusing the output vectors of the last several encoding layers based on the attention mechanism;
[0055] In the text generation method of the present application, the local encoder performs convolution operations with different convolution kernel sizes on the word embedding representation through the local convolution extractor, and concatenates and filters the different convolution outputs and optimizes the vectors, so as to learn the n-gram information with different byte fragment sizes and obtain more key information;
[0056] In the text generation method of the present application, the local encoder fuses the time series information B extracted by the Bi-LSTM with the feature extraction vector extracted by the local convolution extractor through a gating unit, thereby obtaining richer key information features.
[0057] The contextual semantic vector in the text generation method of the present application can retain richer semantic features and key features, which helps the decoder to generate text with higher consistency and semantic accuracy, solves the problems of logical incoherence or content duplication in existing models, and maintains a high level in long text tasks. BRIEF DESCRIPTION OF THE DRAWINGS
[0058] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0059] Figure 1 It is a flowchart of an implementation method of the text generation method of the present application;
[0060] Figure 2 yes Figure 1 A schematic diagram of a flow chart of an implementation method corresponding to S400;
[0061] Figure 3 yes Figure 1 A schematic flow chart of an implementation method corresponding to S500;
[0062] Figure 4 It is a structural schematic diagram of an implementation method of the text generation device of the present application;
[0063] Figure 5 It is a schematic structural diagram of an embodiment of the electronic device of the present application. DETAILED DESCRIPTION
[0064] In order to enable those skilled in the art to better understand the technical solutions in the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in the field without creative work should fall within the scope of protection of the present application.
[0065] Currently, the commonly used text generation technologies mainly include: generation methods based on statistical language models, generation methods based on recurrent neural networks (RNNs), and generation methods based on Transformer models.
[0066] Among them, the generation method based on statistical language model needs to calculate the joint probability of word sequences to generate text, relies on a large amount of artificial feature engineering, has limited generation capabilities and is difficult to process long texts.
[0067] RNN and its improved versions (such as LSTM and GRU) realize text generation through time series modeling and can process contextual information. However, due to its sequential limitations, RNN is prone to gradient disappearance and information loss problems when processing long texts.
[0068] The Transformer model processes input sequences in parallel through the self-attention mechanism, overcoming the shortcomings of RNN and becoming the basis of the current mainstream text generation technology. In particular, the emergence of large-scale pre-trained language models (such as GPT, BERT, T5, etc.) has made significant progress in the fluency and semantic consistency of text generation tasks.
[0069] However, when generating long texts, the existing Transformer model may have logical incoherence or content duplication, and it is difficult to maintain a high level of consistency and semantic accuracy in long text tasks. In addition, in order to ensure the accuracy and simplicity of the generated results, a post-processing algorithm is required, which leads to higher costs and more deployment time.
[0070] In order to solve the above problems, the applicant has developed a new text generation method, which can perform content filtering in the generation algorithm without the need for a post-processing algorithm, and through the optimization of the encoder, it is possible to extract more comprehensive semantic information, which helps to improve the consistency and semantic accuracy of the generated text.
[0071] Specifically, see Figure 1 , Figure 1 It is a flow chart of an implementation method of the text generation method of the present application.
[0072] like Figure 1 As shown, the method includes:
[0073] S100: Obtain source text, perform word segmentation and sorting, and obtain a text sequence.
[0074] First, the user can input the source text from which the text summary needs to be extracted, and by segmenting and sorting the source text, a text sequence sorted by position can be obtained.
[0075] For example, the text sequence Among them, voc is the vocabulary size, and n represents the number of words in the source text, that is, the length of the text.
[0076] The source text is the text of the information to be extracted, for example, it can be but not limited to product description text, introduction text of objects or scenes, news text or articles, etc.
[0077] S200, based on the word embedding matrix, convert the text sequence into a word embedding representation.
[0078] Furthermore, the text sequence can be converted into a vector form acceptable to the model through the word embedding matrix, namely the word embedding representation.
[0079] Exemplarily, the conversion formula may be as follows:
[0080] E=W E X
[0081] Among them, word embedding represents E = (e 1 ,e 2 …,e n ), text sequence Word Embedding Matrix voc is the vocabulary size and emb is the dimension of the word embedding layer.
[0082] In the above formula, the word embedding matrix is a learnable matrix. In practical applications, the word embedding matrix can be trained based on application scenarios in specific fields. For example, large-scale field-specific data sets can be used for pre-training, which can significantly enhance the model's adaptability to scenarios such as product descriptions, spoken copy, and intelligent customer service.
[0083] S300, extracting semantic features from word embedding representation based on a global encoder including n coding layers to obtain n output vectors, and fusing the output vectors of the last several coding layers based on an attention mechanism to obtain first semantic coding information.
[0084] After obtaining the word embedding representation, the word embedding representation can be input into the global encoder for semantic feature extraction.
[0085] Among them, the global encoder includes multiple layers of encoding layers, each of which can include two sub-layers: a multi-head self-attention mechanism and a position feedforward network. Each encoding layer can extract features from the output vector of the previous layer based on the self-attention mechanism and output a vector.
[0086] Therefore, a global encoder based on n encoding layers can obtain n output vectors.
[0087] In order to obtain more comprehensive semantic information, in this embodiment, the output vectors of the last several coding layers can be adaptively weighted and fused based on the attention mechanism, so as to obtain more comprehensive first semantic coding information.
[0088] As described in detail below, each encoding layer can linearly transform the output vector of the previous encoding layer based on the multi-head self-attention mechanism to obtain the query vector, key vector and value vector.
[0089] Specifically, the vector can be mapped to three spaces through three matrices to obtain the word e i The query vector q i , key vector k i , value vector v i , as follows:
[0090] q i =e i W Q
[0091] k i =e i W K
[0092] v i =e i W V
[0093] In the above formula: is a trainable matrix.
[0094] After obtaining the three vectors, you can query the vector q i Access key vector k i , get the attention weight a of a word vector for the rest of the word vectors ij , normalize it, and the calculation formula is:
[0095]
[0096] In the above formula: t is the dimension of the Transformer output vector; exp is the exponential function with the natural constant e as the base.
[0097] Use the attention weight a ij With the value vector v i Multiply them together to get the output vector z i for:
[0098]
[0099] In one embodiment, when the global encoder includes 12 encoding layers, the output vector of each layer can be expressed as:
[0100] Z k =(z 1 ,z 2 ,…,z n ), where k∈[1,12].
[0101] In order to further obtain comprehensive semantic information, the output vectors of the last several layers are adaptively weighted and fused. In one embodiment, when the global encoder includes 12 coding layers, the output vectors of the last 4 layers can be adaptively weighted, and different weights can be assigned to each layer. The formula can be as follows:
[0102]
[0103] In the formula, ω j That is, the weight of the jth layer, j is selected from 9 to 12, and exp is an exponential function with the natural constant e as the base.
[0104] After obtaining the weights of each of the final four layers, vector fusion can be performed based on the weights to obtain the first semantic coding information, namely:
[0105]
[0106] It can be understood that in other implementations, other numbers of output vectors of coding layers may be selected for fusion, and the above formula may be adjusted based on actual needs to achieve the effect of this implementation.
[0107] S400: Extract semantic features of the word embedding representation based on the local encoder to obtain second semantic encoding information.
[0108] In order to extract more comprehensive semantic features, this embodiment not only inputs the word embedding representation into the global encoder, but also inputs the word embedding representation into the local encoder in parallel for semantic feature extraction, thereby obtaining more key information, which helps to enrich the semantic information of the generated text.
[0109] Specifically, in this embodiment, the local encoder may include a Bi-LSTM, a local convolution extractor and a gating unit. The Bi-LSTM can obtain global semantic information, the local convolution extractor can extract information of n-grams with a byte segment size of n, and the gating unit will fuse the information extracted by the two, so as to extract the internal relationship of the language and capture richer semantic information.
[0110] Specifically, see Figure 2 , Figure 2 yes Figure 1 A flow chart of an implementation method corresponding to S400 in FIG.
[0111] like Figure 2 As shown, the method for extracting semantic features by a local encoder may include:
[0112] S401. Extract global semantic information from word embedding representation based on Bi-LSTM to obtain time series information.
[0113] Bi-LSTM consists of a forward and backward network, and its output is B = {h 1 ,h 2 ,…,h n}. The output at each moment contains hidden states in two directions The calculation formula is:
[0114] B=Bi-LSTM(E).
[0115] Where B is the time series information.
[0116] The Bi-LSTM model is a feature extraction model well known in the art, and its specific structure and algorithm flow will not be described here.
[0117] S402, based on the local convolution extractor, convolution operations are performed on the word embedding representation using convolution kernels of different sizes to learn n-gram features of different sizes and obtain multiple convolution outputs.
[0118] The local convolution extractor performs convolution through convolution modules with different receptive fields, which can learn n-gram features of different byte fragment sizes, that is, n-gram information, to obtain more key information.
[0119] In one embodiment, three convolution kernels with sizes of 1, 3, and 5 can be used to perform convolution operations. Combined with the input hidden layer state H, three convolution operations with different granularity sizes can be used to obtain the corresponding output T k=1 、T k=3 、T k=5 , where k is the convolution kernel size.
[0120] S403, concatenating multiple convolution outputs, and then performing vector optimization through residual connection to obtain a feature extraction vector.
[0121] In order to integrate the different granularity features of word-level information, the above three outputs are concatenated, that is: T = concat(T k=1 ,T k=3 ,T k=5 ).
[0122] In order to avoid the gradient vanishing problem caused by deep convolutional neural networks, this module adds a residual connection to optimize the output vector. The final joint output - feature extraction vector H T for:
[0123] H T =σ(W T T+b T )+H
[0124] In the formula, σ is the sigmoid function, H is the hidden layer state, and W T and b T is a learnable weight matrix, and T is a vector obtained by concatenating multiple convolution outputs.
[0125] S404: Based on the gating unit, the time series information and the feature extraction vector are integrated to obtain second semantic coding information.
[0126] This embodiment combines the gating unit to extract the H T It is fused with the time series information B extracted by Bi-LSTM to obtain richer key information features, namely the second semantic coding information.
[0127] Specifically, the calculation formula of the second semantic coding information is as follows:
[0128] R=(W T H T +b T)⊙(W B B+b B )
[0129] Where W T 、b T , W B and b B is the learnable weight matrix, H T is the feature extraction vector, and B is the time series information.
[0130] By continuously adjusting and optimizing W during the training process T 、b T , W B and b B The value of can improve the performance of the model and obtain the optimal second semantic encoding information.
[0131] S500 , performing feature fusion and filtering on the first semantic coding information and the second semantic coding information based on a global gating unit to obtain a context semantic vector.
[0132] In the text generation method of this embodiment, word embedding representations are input into the global encoder and the local encoder in parallel. By combining the feature extraction vectors of the global encoder and the local encoder, more global semantics can be captured.
[0133] However, the combination process usually includes too many repeated words, and only a small number of words can be called key information of the source text, and these key information are the elements that really need to be extracted. Therefore, in this implementation, the feature vectors are fused and filtered through a global gating unit to avoid redundant and repeated words in the final generated text.
[0134] Specifically, see Figure 3 , Figure 3 yes Figure 1 A flow chart of an implementation method corresponding to S500 in FIG.
[0135] like Figure 3 As shown, the method for feature fusion and filtering of the first semantic coding information and the second semantic coding information includes:
[0136] S501: Concatenate the first semantic coding information and the second semantic coding information to obtain concatenated coding information.
[0137] In order to improve the global semantic information expression of the input sequence, first, the semantic coding information extracted by the two encoders can be concatenated, which can be specifically expressed as follows: H'=concat(G,R), where G is the first semantic coding information and R is the second semantic coding information.
[0138] S502 , performing weight operation and nonlinear operation on the spliced coding information through a global gating unit to obtain a tube selection probability.
[0139] Inputting the splicing coding information into the global gating unit can generate the tube selection probability g t
[0140] Specifically, the tube selection probability g t The calculation formula can be as follows:
[0141] g t =σ(W g H')
[0142] In the formula, g t is the probability of tube selection, W g is the learnable weight matrix, σ is the sigmoid function, and H' is the concatenated coding information.
[0143] S503 , weights are assigned to the first semantic coding information and the second semantic coding information based on the control selection probability, and the first semantic coding information and the second semantic coding information are fused based on the assigned weights to obtain a context semantic vector.
[0144] The control selection probability generated by the global gating unit represents the weight distribution of the semantic coding information of the two encoders. Therefore, the two semantic coding information can be fused based on the control selection probability, and the redundant information can be filtered to obtain the context semantic vector.
[0145] Specifically, the calculation method of the context semantic vector is as follows:
[0146] O t =(1-g t )G+g t R
[0147] In the formula, Q t is the context semantic vector.
[0148] S600: Decode the context semantic vector based on the decoder to generate a target text.
[0149] After obtaining the context semantic vector in the above steps, the context semantic vector can be decoded based on the decoder. The decoder generates coherent and semantically consistent output content by modeling the current context and learning the semantic information.
[0150] The decoder in this embodiment can be any decoder commonly used in the art, for example, it can include an attention mechanism and an LSTM decoder, and it can also include a pointer generator network commonly used in the art to avoid the problem of unregistered words. Both can achieve the effect of this embodiment and will not be repeated here.
[0151] Based on the methods of the above-mentioned implementation modes, the context semantic vector is obtained by fusing and filtering the semantic features extracted by the global encoder and the local encoder, and redundant features are removed. Therefore, there is no need for post-processing algorithms to perform deduplication operations, which reduces costs on the one hand and deployment time on the other.
[0152] In addition, the global encoder retains more comprehensive semantic features by fusing the output vectors of the last several encoding layers based on the attention mechanism;
[0153] The local encoder performs convolution operations with different convolution kernel sizes on the word embedding representation through the local convolution extractor, and concatenates and filters the different convolution outputs and optimizes the vectors. It can learn the n-gram information of different byte fragment sizes and obtain more key information.
[0154] In addition, the local encoder fuses the time series information B extracted by Bi-LSTM with the feature extraction vector extracted by the local convolution extractor through the gating unit, which can obtain richer key information features.
[0155] Therefore, the contextual semantic vector can retain richer semantic features and key features, help the decoder generate text with higher consistency and semantic accuracy, solve the problems of logical incoherence or content repetition in existing models, and maintain a high level in long text tasks.
[0156] This application also provides a text generation device, see Figure 4 , Figure 4 It is a structural diagram of an implementation method of the text generation device of the present application.
[0157] like Figure 4 As shown, the device includes a word segmentation module 21, an embedding conversion module 22, a first feature extraction module 23, a second feature extraction module 24, a global gating module 25 and a decoding module 26.
[0158] The word segmentation module 21 is used to obtain the source text, and perform word segmentation and sorting to obtain a text sequence;
[0159] The embedding conversion module 22 is used to convert the text sequence into a word embedding representation based on the word embedding matrix;
[0160] The first feature extraction module 23 is used to extract semantic features from the word embedding representation based on a global encoder including n encoding layers to obtain n output vectors, and fuse the output vectors of the last several encoding layers based on an attention mechanism to obtain first semantic encoding information;
[0161] The second feature extraction module 24 is used to extract semantic features from the word embedding representation based on the local encoder to obtain second semantic encoding information;
[0162] The global gating module 25 is used to perform feature fusion and filtering on the first semantic encoding information and the second semantic encoding information based on the global gating unit to obtain a context semantic vector;
[0163] The decoding module 26 is used to decode the context semantic vector based on the decoder to generate the target text.
[0164] As above Figures 1 to 3 , a method for generating text according to an embodiment of this specification is described. The details mentioned in the above description of the method embodiment are also applicable to the text generating device of the embodiment of this specification. The above text generating device can be implemented by hardware, or by software, or by a combination of hardware and software.
[0165] This application also provides an electronic device, see Figure 5 , Figure 5 Schematic diagram of the structure of an electronic device of the present application. Figure 5 As shown, the electronic device 30 may include at least one processor 31, a memory 32 (e.g., a non-volatile memory), a memory 33, and a communication interface 34, and the at least one processor 31, the memory 32, the memory 33, and the communication interface 34 are connected together via a bus 35. At least one processor 31 executes at least one computer-readable instruction stored or encoded in the memory 32.
[0166] It should be understood that the computer executable instructions stored in the memory 32, when executed, cause at least one processor 31 to perform the above combined operations in various embodiments of the present specification. Figure 1-Figure 4 Describes the various operations and functions.
[0167] In the embodiments of the present specification, the electronic device 30 may include, but is not limited to, personal computers, server computers, workstations, desktop computers, laptop computers, notebook computers, mobile electronic devices, smart phones, tablet computers, cellular phones, personal digital assistants (PDAs), handheld devices, messaging devices, wearable electronic devices, consumer electronic devices, and the like.
[0168] According to one embodiment, a program product such as a machine-readable medium is provided. The machine-readable medium may have instructions (i.e., the above-mentioned elements implemented in the form of software), which, when executed by a machine, causes the machine to perform the above-mentioned combination of various embodiments of this specification. Figure 1-Figure 5 Specifically, a system or device equipped with a readable storage medium may be provided, on which a software program code implementing the functions of any of the above-mentioned embodiments is stored, and a computer or processor of the system or device reads and executes the instructions stored in the readable storage medium.
[0169] In this case, the program code itself read from the machine-readable medium can realize the function of any one of the above embodiments, and thus the machine-readable code and the machine-readable storage medium storing the machine-readable code constitute part of this specification.
[0170] Examples of readable storage media include floppy disks, hard disks, magneto-optical disks, optical disks (such as CD-ROM, CD-R, CD-RW, DVD-ROM, DVD-RAM, DVD-RW, DVD-RW), magnetic tapes, non-volatile memory cards, and ROMs. Alternatively, the program code may be downloaded from a server computer or a cloud via a communication network.
[0171] Those skilled in the art should understand that the various embodiments disclosed above can be modified and altered in various ways without departing from the essence of the invention. Therefore, the protection scope of this specification should be defined by the appended claims.
[0172] It should be noted that not all steps and units in the above-mentioned processes and system structure diagrams are necessary, and some steps or units can be ignored according to actual needs. The execution order of each step is not fixed and can be determined as needed. The device structure described in the above-mentioned embodiments can be a physical structure or a logical structure, that is, some units may be implemented by the same physical client, or some units may be implemented by multiple physical clients, or some components in multiple independent devices may be implemented together.
[0173] In the above embodiments, the hardware unit or module can be realized by mechanical or electrical means. For example, a hardware unit, module or processor can include permanent dedicated circuit or logic (such as special processor, FPGA or ASIC) to complete the corresponding operation. The hardware unit or processor can also include programmable logic or circuit (such as general-purpose processor or other programmable processor), which can be temporarily set by software to complete the corresponding operation. Specific implementation (mechanical method or dedicated permanent circuit or temporary circuit) can be determined based on cost and time consideration.
[0174] The specific embodiments described above in conjunction with the accompanying drawings describe exemplary embodiments, but do not represent all embodiments that can be implemented or fall within the scope of protection of the claims. The term "exemplary" used throughout this specification means "used as an example, instance or illustration" and does not mean "preferred" or "having advantages" over other embodiments. For the purpose of providing an understanding of the described technology, the specific embodiments include specific details. However, these technologies can be implemented without these specific details. In some instances, in order to avoid making the concepts of the described embodiments difficult to understand, well-known structures and devices are shown in block diagram form.
[0175] The above description of the present disclosure is provided to enable any person of ordinary skill in the art to implement or use the present disclosure. Various modifications to the present disclosure will be apparent to those of ordinary skill in the art, and the general principles corresponding to the present disclosure may be applied to other variations without departing from the scope of protection of the present disclosure. Therefore, the present disclosure is not limited to the examples and designs described herein, but is consistent with the widest range of principles and novel features disclosed herein.
Claims
1. A text generation method, characterized in that: include: Get the source text, perform word segmentation and sorting, and obtain a text sequence; Based on the word embedding matrix, convert the text sequence into a word embedding representation; Extracting semantic features from the word embedding representation based on a global encoder including n encoding layers to obtain n output vectors, and fusing the output vectors of the last several encoding layers based on an attention mechanism to obtain first semantic encoding information; Extracting semantic features from the word embedding representation based on a local encoder to obtain second semantic encoding information; Based on the global gating unit, feature fusion and filtering are performed on the first semantic encoding information and the second semantic encoding information to obtain a context semantic vector; The context semantic vector is decoded based on a decoder to generate a target text.
2. The text generation method according to claim 1, characterized in that: The step of converting the text sequence into a word embedding representation based on the word embedding matrix is specifically as follows: E=W E X In the formula, word embedding represents E = ( e1,e2…,e n) , text sequence Word Embedding Matrix voc is the vocabulary size and emb is the dimension of the word embedding layer.
3. The text generation method according to claim 1, characterized in that: The steps of extracting semantic features from the word embedding representation based on a global encoder including n encoding layers to obtain n output vectors, and fusing the output vectors of the last several encoding layers based on an attention mechanism include: Each encoding layer linearly transforms the output vector of the previous encoding layer based on the multi-head self-attention mechanism to obtain the query vector, key vector and value vector, calculates the attention vector based on the query vector and key vector, and obtains the output vector based on the attention vector and value vector; Based on the attention mechanism, different weights are assigned to the output vectors of the last several encoding layers; The output vectors of the last several coding layers are fused based on the weights to obtain the first semantic coding information.
4. The text generation method according to claim 1, characterized in that: The local encoder includes a Bi-LSTM, a local convolution extractor and a gating unit; The step of extracting local semantic features from the word embedding representation based on a local encoder to obtain second semantic encoding information comprises: Extracting global semantic information from the word embedding representation based on Bi-LSTM to obtain time series information; Based on the local convolution extractor, convolution operations are performed on the word embedding representation using convolution kernels of different sizes to learn n-gram features of different sizes to obtain multiple convolution outputs; The multiple convolution outputs are concatenated, and then vector optimization is performed through residual connection to obtain a feature extraction vector; The time series information and the feature extraction vector are fused based on the gating unit to obtain second semantic coding information.
5. The text generation method according to claim 4, characterized in that: The step of concatenating the multiple convolution outputs and then performing vector optimization through residual connection to obtain the feature extraction vector is specifically as follows: H T =σ(W T T+b T )+H In the formula, σ is the sigmoid function, H is the hidden layer state, and W T and b T is a learnable weight matrix, T is a vector obtained by concatenating multiple convolution outputs; The step of fusing the time series information and the feature extraction vector based on the gating unit to obtain the second semantic coding information is specifically as follows: R=(W T H T +b T )⊙(W B B+b B ) Where W T , b T , W B and b B is the learnable weight matrix, H T is the feature extraction vector, and B is the time series information.
6. The text generation method according to claim 1, characterized in that: The step of performing feature fusion and filtering on the first semantic encoding information and the second semantic encoding information based on the global gating unit to obtain a context semantic vector comprises: splicing the first semantic coding information and the second semantic coding information to obtain spliced coding information; Performing weight calculation and nonlinear calculation on the spliced coding information through a global gating unit to obtain a tube selection probability; The first semantic coding information and the second semantic coding information are weighted based on the control selection probability, and the first semantic coding information and the second semantic coding information are fused based on the weights assigned to obtain a context semantic vector.
7. The text generation method according to claim 6, characterized in that: The step of performing weight calculation and nonlinear calculation on the spliced coding information through the global gating unit to obtain the control selection probability is specifically as follows: g t =s ( W g H In the formula, g t is the probability of tube selection, W g is a learnable weight matrix, σ is a sigmoid function, H' is the concatenated coding information, that is, H'=concat(G,R), G is the first semantic coding information, and R is the second semantic coding information; The calculation method of the context semantic vector is as follows: O t = ( 1-g t ) G+g t R In the formula, Q t is the context semantic vector.
8. A text generation device, characterized in that: include: The word segmentation module is used to obtain the source text, perform word segmentation and sorting, and obtain a text sequence; An embedding conversion module, used for converting the text sequence into a word embedding representation based on a word embedding matrix; A first feature extraction module is used to extract semantic features from the word embedding representation based on a global encoder including n encoding layers to obtain n output vectors, and to fuse the output vectors of the last several encoding layers based on an attention mechanism to obtain first semantic encoding information; A second feature extraction module is used to extract semantic features from the word embedding representation based on a local encoder to obtain second semantic encoding information; A global gating module, used for performing feature fusion and filtering on the first semantic encoding information and the second semantic encoding information based on a global gating unit to obtain a context semantic vector; A decoding module is used to decode the context semantic vector based on a decoder to generate a target text.
9. An electronic device, comprising: at least one processor; as well as A memory storing instructions, which, when executed by the at least one processor, enables the at least one processor to execute the text generation method according to any one of claims 1 to 8.
10. A machine-readable storage medium storing executable instructions, wherein when the instructions are executed, the machine executes the text generation method according to any one of claims 1 to 8.
Citation Information
Cited By
Path planning method and system based on user text travel demand analysis
CN120782090A
Text recognition method and device, server and medium
CN120877298A
Text recognition method and device, server and medium
CN120877298B