Generative network allowing for rapid training and low-computing-power prediction, and prediction method

By optimizing the structure and calculation methods of the generative network, the problems of long training time and low prediction efficiency of the generative network are solved, and fast training and low computing power prediction are achieved, which are suitable for natural language processing and other fields.

WO2025138383A1PCT designated stage expired Publication Date: 2025-07-03ROCK AI
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/074192
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-28
Filing Date
2024-01-26
Publication Date
2025-07-03

AI Technical Summary

Technical Problem

The long training time of generative networks and low prediction efficiency limit their wide application in practical applications.

Method used

The combination of sequence conversion layer, symbolic word vector encoding layer, position word vector encoding layer, first linear conversion layer, mixed word vector encoding layer and mixed word vector processing layer is used to optimize the training process of the generative network through matrix multiplication and star multiplication calculation to generate predicted word vectors.

Benefits of technology

It improves the training speed and prediction efficiency of generative networks, reduces computing and storage overhead, improves the generalization ability of large models, and can quickly predict natural language processing tasks on the CPU.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024074192_03072025_PF_FP_ABST
    Figure CN2024074192_03072025_PF_FP_ABST
Patent Text Reader

Abstract

The present invention provides a generative network allowing for rapid training and low-computing-power prediction, comprising: a sequence conversion layer, used for encoding a natural language into an input sequence; a token word vector encoding layer, used for encoding the input sequence into a first token word vector; a location word vector encoding layer, used for encoding the input sequence into a location word vector; a first linear conversion layer, used for performing linear conversion on the first token word vector to generate a second token word vector; a word vector mixing and encoding layer, used for mixing the second token word vector and the location word vector and performing encoding into a mixed word vector; a mixed word vector processing layer, used for dividing the mixed word vector into three portions, respectively performing linear conversion, cumulative summation and normalization processing, and then performing association calculation to generate a prediction word vector; and an output layer, used for performing forward calculation on the prediction word vector to generate an output sequence. The present invention can improve the training speed and efficiency of a generative task, and improves the generalization capability of a large model.
Need to check novelty before this filing date? Find Prior Art

Description

A generative network and prediction method with fast training and low computing power for prediction Technical Field The present invention relates to the technical field of generative networks, and in particular to a generative network and a prediction method with fast training and low computing power for prediction. Background Art Generative networks have achieved remarkable success in many applications, but the long training cycle and computing resource requirements limit their wide application in practical applications. At the same time, the prediction of current large models (models after the training of generative networks) still requires computing power support, and the requirements for the training resources of large models also limit the scope of the actual application of the models. Therefore, it is necessary to propose a generative network with fast training to reduce the training time, improve the performance of large models, and prediction efficiency. Summary of the Invention The present invention provides a generative network and a prediction method with fast training and low computing power for prediction to solve the technical problems of long training time of generative networks, low performance of large models, and low prediction efficiency in the prior art. One aspect of the present invention lies in providing a generative network with fast training and low computing power for prediction, and the generative network includes: A sequence conversion layer for encoding natural language into an input sequence; A morpheme word vector encoding layer for encoding the natural language words corresponding to the input sequence into first morpheme word vectors; A position word vector encoding layer for encoding the natural language words corresponding to the input sequence into position word vectors; A first linear conversion layer for linearly converting the first morpheme word vectors to generate second morpheme word vectors; A mixed word vector encoding layer for mixing and encoding the second morpheme word vectors and the position word vectors into mixed word vectors; A mixed word vector processing layer for dividing the mixed word vectors into three paths, respectively performing linear conversion, cumulative summation, and normalization processing, and then performing correlation calculation to generate prediction word vectors; An output layer for performing forward calculation on the prediction word vectors to generate an output sequence. In a preferred embodiment, the first linear conversion layer linearly converts the first morpheme word vectors by matrix multiplication calculation to generate second morpheme word vectors. In a preferred embodiment, the mixed word vector encoding layer mixes and encodes the second morpheme word vectors and the position word vectors into mixed word vectors by star multiplication calculation. In a preferred embodiment, the hybrid word vector processing layer includes: a second linear transformation layer, a third linear transformation layer, and a fourth linear transformation layer, a first cumulative summation layer, a second cumulative summation layer, and a third cumulative summation layer, as well as a first normalization layer, a second normalization layer, and a third normalization layer; The second linear transformation layer is configured to perform a linear transformation on the hybrid word vector to generate a first linearly transformed word vector; The third linear transformation layer is configured to perform a linear transformation on the hybrid word vector to generate a second linearly transformed word vector; The fourth linear transformation layer is configured to perform a linear transformation on the hybrid word vector to generate a third linearly transformed word vector; The first cumulative summation layer is configured to perform cumulative summation on the first linearly transformed word vector to generate a first cumulative summation word vector; The second cumulative summation layer is configured to perform cumulative summation on the second linearly transformed word vector to generate a second cumulative summation word vector; The third cumulative summation layer is configured to perform cumulative summation on the third linearly transformed word vector to generate a third cumulative summation word vector; The first normalization layer is configured to perform a normalization process on the first cumulative summation word vector to generate a first normalized word vector; The second normalization layer is configured to perform a normalization process on the second cumulative summation word vector to generate a second normalized word vector; The third normalization layer is configured to perform a normalization process on the third cumulative summation word vector to generate a third normalized word vector. In a preferred embodiment, the hybrid word vector processing layer further includes: a correlation calculation layer; The correlation calculation layer is configured to perform a correlation calculation on the first normalized word vector, the second normalized word vector, the third normalized word vector, as well as the first morpheme word vector and the second morpheme word vector to generate a predicted word vector. Another aspect of the present invention lies in providing a prediction method for a generative network capable of fast training and low-computing power prediction. The prediction method uses a generative network capable of fast training and low-computing power prediction provided by the present invention to predict natural language, and includes the following method steps: S1. Select a segment of natural language and input it into the sequence conversion layer of the generative network to encode the natural language into an input sequence; S2. Use the input sequence as the input and the output sequence as the output to train the generative network; S3. Select a natural language word to be predicted and input it into the sequence conversion layer of the generative network to encode the natural language word to be predicted into an input sequence; S4. Encode the natural language words corresponding to the input sequence into a first ideogram word vector, and at the same time, encode the natural language words corresponding to the input sequence into a position word vector; S5. Perform a linear transformation on the first ideogram word vector to generate a second ideogram word vector; S6. Mix and encode the second ideogram word vector and the position word vector into a mixed word vector; S7. Divide the mixed word vector into three paths, perform linear transformation, cumulative summation, and normalization processing respectively, and then perform an association calculation to generate a predicted word vector; S8. Perform a forward calculation on the predicted word vector to generate an output sequence, and predict the next word of the natural language word. In a preferred embodiment, in step S5, the first ideogram word vector is linearly transformed by matrix multiplication to generate a second ideogram word vector. In a preferred embodiment, in step S6, the second ideogram word vector and the position word vector are mixed and encoded into a mixed word vector by star multiplication calculation. In a preferred embodiment, in step S7, the mixed word vector is linearly transformed to generate a first linearly transformed word vector, a second linearly transformed word vector, and a third linearly transformed word vector respectively; Perform cumulative summation on the first linearly transformed word vector to generate a first cumulative summation word vector; perform cumulative summation on the second linearly transformed word vector to generate a second cumulative summation word vector; perform cumulative summation on the third linearly transformed word vector to generate a third cumulative summation word vector; Perform normalization processing on the first cumulative summation word vector to generate a first normalized word vector; perform normalization processing on the second cumulative summation word vector to generate a second normalized word vector; perform normalization processing on the third cumulative summation word vector to generate a third normalized word vector. In a preferred embodiment, an association calculation is performed on the first normalized word vector, the second normalized word vector, the third normalized word vector, and the first ideogram word vector and the second ideogram word vector to generate a predicted word vector. Compared with the prior art, the present invention has the following beneficial effects: A generative network and prediction method capable of fast training and low-computing power prediction provided by the present invention, by combining deep learning and optimization algorithms, enables the generative network to significantly improve the convergence speed and performance of the large model during training, and can directly and quickly predict on the CPU without quantization, pruning, etc., and can be widely used in fields such as natural language processing. A generative network and prediction method capable of rapid training and low-computing-power prediction provided by the present invention can improve the training speed and efficiency of generative tasks, improve the generalization ability of large models, reduce the machine hallucination of the model, and reduce the computing and storage overhead of the network structure. The entire computing process can cache the intermediate output results and quickly calculate the next character through the cache when predicting the next character. BRIEF DESCRIPTION OF THE DRAWINGS In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for the description of the specific embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts. FIG. 1 is a structural block diagram of a generative network capable of rapid training and low-computing-power prediction according to the present invention. structural block diagram. FIG. 2 is a schematic diagram of natural language word prediction by a generative network capable of rapid training and low-computing-power prediction according to the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS In order to make the above and other features and advantages of the present invention clearer, the present invention will be further described below with reference to the drawings. It should be understood that the specific embodiments given herein are for the purpose of explaining to those skilled in the art and are merely exemplary, not restrictive. In the description of the present invention, it should be understood that the terms "center", "longitudinal", "lateral", "length", "width", "thickness", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", "clockwise", "counterclockwise", "axial", "radial", "circumferential", etc. indicate the orientation or positional relationship based on the orientation or positional relationship shown in the drawings, and are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and thus cannot be understood as a limitation of the present invention. In addition, the terms "first" and "second" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly indicating the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one of such features. In the description of the present invention, "a plurality" means at least two, such as two, three, etc., unless otherwise specifically defined. In combination with Figures 1 and 2, according to an embodiment of the present invention, a generative network 100 that can be quickly trained and predicted with low computing power is provided, including: an input layer, a sequence conversion layer 101, a shape word vector encoding layer 102, a position word vector encoding layer 103, a first linear conversion layer 104, a mixed word vector encoding layer 105, a mixed word vector processing layer and an output layer. The input layer is used to input natural language. The sequence conversion layer 101 is used to convert natural language Encoded as input sequence (xs sequence). For example, a natural language [["Today", "Weather", "Great"]] is input into the generative network 100 through the input layer. The sequence conversion layer 101 encodes the natural language [["Today", "Weather", "Great"]] into an input sequence [[36,78,69]], that is, the xs sequence is [[36,78,69]]. The shape-to-word vector encoding layer 102 is used to encode the natural language words corresponding to the input sequence into a first shape-to-word vector. For example, the input sequence [[36,78,69]] (xs sequence [[36,78,69]])

[0036] ] corresponds to the natural language word “today”, the natural language word “weather” corresponding to the input sequence [

[0078] ] in the input sequence [[36,78,69]] (xs sequence [[36,78,69]]), and the natural language word “great” corresponding to the input sequence [

[0069] ] in the input sequence [[36,78,69]] (xs sequence [[36,78,69]]). The token word vector encoding layer 102 encodes the natural language word "today" corresponding to the input sequence [

[0036] ] in the input sequence [[36,78,69]] (xs sequence [[36,78,69]]) into the first token word vector (token_embs). Similarly, the shape-token word vector encoding layer 102 encodes the natural language word "weather" corresponding to the input sequence [

[0078] ] in the input sequence [[36,78,69]] (xs sequence [[36,78,69]]) into the first shape-token word vector (token_embs). The token word vector encoding layer 102 encodes the natural language word "really good" corresponding to the input sequence [

[0069] ] in the input sequence [[36,78,69]] (xs sequence [[36,78,69]]) into the first token word vector (token_embs). The position word vector encoding layer 103 is used to encode the natural language words corresponding to the input sequence into position word vectors. For example, the positional word vector encoding layer 103 encodes the natural language word "today" corresponding to the input sequence [

[0036] ] in the input sequence [[36, 78, 69]] (xs sequence [[36, 78, 69]]) as positional word vector (position_embs). Similarly, the positional word vector encoding layer 103 encodes the natural language word "weather" corresponding to the input sequence [

[0078] ] in the input sequence [[36, 78, 69]] (xs sequence [[36, 78, 69]]) as the positional word vector (position_embs). The positional word vector encoding layer 103 encodes the natural language word "so nice" corresponding to the input sequence [

[0069] ] in the input sequence [[36, 78, 69]] (xs sequence [[36, 78, 69]]) as the positional word vector (position_embs). The first linear transformation layer 104 is used to perform a linear transformation on the first morpheme word vector to generate a second morpheme word vector. Furthermore, the first linear transformation layer 104 performs a linear transformation on the first morpheme word vector by means of matrix multiplication to generate a second morpheme word vector. For example, a linear transformation is performed on a morpheme word vector through the linear calculation function j(x) to generate a second morpheme word vector. For example, the first linear transformation layer 104 performs a linear transformation on the first morpheme word vector (token_embs) of the natural language word "today" corresponding to the input sequence [

[0036] ] to generate a second morpheme word vector (j_token_embs). Similarly, the first linear transformation layer 104 encodes the natural language word "weather" corresponding to the input sequence [

[0078] ] as the first morpheme word vector (token_embs) and performs a linear transformation to generate a second morpheme word vector (j_token_embs). The first linear transformation layer 104 encodes the natural language word "so nice" corresponding to the input sequence [

[0069] ] as the first morpheme word vector (token_embs) and performs a linear transformation to generate a second morpheme word vector (j_token_embs). The hybrid word vector encoding layer 105 is used to hybridly encode the second morpheme word vector (j_token_embs) and the positional word vector (position_embs) into a hybrid word vector (embs). Furthermore, the hybrid word vector encoding layer 105, the second morpheme word vector (j_token_embs) and position word vectors (position_embs) are mixed and encoded into mixed word vectors (embs) through the method of element-wise multiplication. In some embodiments, the first linear transformation layer 104 may not be provided, and the first morpheme word vectors encoded by the morpheme word vector encoding layer 102 and the position word vectors encoded by the position word vector encoding layer 103 may be directly mixed and encoded into mixed word vectors. The mixed word vector processing layer is used to divide the mixed word vectors embs into three paths, perform linear transformation, cumulative summation, and normalization processing respectively, and then perform correlation calculation to generate predicted word vectors. Further, as shown in FIG. 1, the mixed word vector processing layer includes: a second linear transformation layer 106, a third linear transformation layer 107, and a fourth linear transformation layer 108, a first cumulative summation layer 109, a second cumulative summation layer 110, and a third cumulative summation layer 111, and a first normalization layer 112, a second normalization layer 113, and a third normalization layer 114. The second linear transformation layer 106 is used to perform linear transformation on the mixed word vectors (embs) to generate first linearly transformed word vectors. The third linear transformation layer 107 is used to perform linear transformation on the mixed word vectors (embs) to generate second linearly transformed word vectors. The fourth linear transformation layer 108 is used to perform linear transformation on the mixed word vectors (embs) to generate third linearly transformed word vectors. In one embodiment, the second linear transformation layer 106 performs linear transformation on the mixed word vectors (embs) using the linear transformation function q(x) to generate first linearly transformed word vectors q(embs). The third linear transformation layer 107 performs linear transformation on the mixed word vectors (embs) using the linear transformation function k(x) to generate second linearly transformed word vectors k(embs). The fourth linear transformation layer 108 performs linear transformation on the mixed word vectors (embs) using the linear transformation function v(x) to generate third linearly transformed word vectors v(embs). The linear transformation functions of the second linear transformation layer 106, the third linear transformation layer 107, and the fourth linear transformation layer 108 of the present invention are not specifically limited, and those skilled in the art can select a suitable matrix multiplication function according to actual needs. The first cumulative summation layer 109 is used to perform cumulative summation (cumsum) on the first linearly transformed word vectors q(embs) to generate first cumulative summation word vectors (q_embs); The second cumulative summation layer 110 is used to perform cumulative summation (cumsum) on the second linearly transformed word vectors k(embs) to generate second cumulative summation word vectors (k_embs); The third cumulative summation layer 111 is used to perform cumulative summation (cumsum) on the third linearly transformed word vector v(embs) to generate a third cumulative summation word vector (v_embs). The cumulative summation functions of the first cumulative summation layer 109, the second cumulative summation layer 110, and the third cumulative summation layer 111 of the present invention are not specifically limited. Those skilled in the art can select a suitable cumulative summation function to perform cumulative summation calculation on the vector according to actual needs. The first normalization layer 112 is used to perform normalization processing (layer_norm) on the first cumulative summation word vector (q_embs) to generate a first normalized word vector a; The second normalization layer 113 is used to perform normalization processing (layer_norm) on the second cumulative summation word vector (k_embs) to generate a second normalized word vector b; The third normalization layer 114 is used to perform normalization processing (layer_norm) on the third cumulative summation word vector (v_embs) to generate a third normalized word vector c. According to an embodiment of the present invention, the hybrid word vector processing layer further includes: a correlation calculation layer 115. The correlation calculation layer 115 is used to perform correlation calculation on the first normalized word vector a, the second normalized word vector b, the third normalized word vector c, and the first token word vector (token_embs) and the second token word vector (j_token_embs) to generate a predicted word vector d. The correlation function adopted by the correlation calculation layer 115 of the present invention is not specifically limited. Those skilled in the art can select a suitable correlation function according to actual needs. In this embodiment, the predicted word vector d is calculated by the following exemplary correlation function: d = (a + a * b - b * token_embs - c * j_token_embs). The output layer is used to perform forward calculation on the predicted word vector d to generate an output sequence (ys sequence). By performing forward calculation on the predicted word vector d through the output layer, the generated output sequence (ys sequence) is used to predict the next word. As shown in Figure 2, the natural language word [[“today”]] is encoded as an input sequence [

[0036] ] (xs sequence [

[0036] ]) by the sequence conversion layer 101 and input into the generative network 100. The predicted word vector d is forward calculated in the output layer to output an output sequence [

[0078] ] (ys sequence [

[0078] ]), thereby predicting the next word [[“weather”]] of the natural language word “today”. For another example, natural language words [[“today”, “weather”]] are encoded by the sequence conversion layer 101 into an input sequence [[36, 78]] (xs sequence [[36, 78]]), which is input into the generative network 100. At the output layer, the word vector d is predicted and forward calculation is performed to output an output sequence [

[0069] ] (ys sequence [

[0069] ]), so as to predict the next word [[“so nice”]] of the natural language words [[“today”, “weather”]]. For another example, natural language words [[“today”, “weather”, “so nice”]] are encoded by the sequence conversion layer 101 into an input sequence [[36, 78, 69]] (xs sequence [[36, 78, 69]]), which is input into the generative network 100. At the output layer, the word vector d is predicted and forward calculation is performed to output an output sequence <eos>(ys sequence <eos>) so as to predict the next word of the natural language words [[ "today", "weather", "nice" ]] <eos>, wherein, <eos>is the End Of Sequence. According to an embodiment of the present invention, there is provided a prediction method for a generative network that can be quickly trained and has low computing power for prediction. Using a generative network that can be quickly trained and has low computing power provided by the present invention to perform prediction on natural language, the method includes the following steps: Step S1: Select a piece of natural language and input it into the sequence conversion layer 101 of the generative network 100 to encode the natural language into an input sequence. For example, select a piece of natural language [["Today", "weather", "is nice"]], input it into the sequence conversion layer 101 of the generative network 100, and encode the natural language into an input sequence [[36, 78, 69]]. Step S2: Use the input sequence as the input and the output sequence as the output to train the generative network 100. For example, use the input sequence [

[0036] ] as the input and the output sequence [

[0078] ] as the output; use the input sequence [[36, 78]] as the input and the output sequence [

[0069] ] as the output; use the input sequence [[36, 78, 69]] as the input and the output sequence <eos>As output, the generative network 100 is trained. In one embodiment, the cross-entropy loss function is used as the loss function during the training process. Step S3: Select the natural language word to be predicted, input it into the sequence conversion layer 101 of the generative network 100, and encode the natural language word to be predicted into an input sequence. For example, select the natural language word ["today"], input it into the sequence conversion layer 101 of the generative network 100, and encode the natural language word to be predicted into the input sequence [

[0036] ]. Step S4: Encode the natural language word ["today"] corresponding to the input sequence [

[0036] ] into the first token embedding (token_embs), and at the same time encode the natural language word ["today"] corresponding to the input sequence [

[0036] ] into the position embedding (position_embs). Step S5: Perform a linear transformation on the first token embedding (token_embs) to generate a second token embedding (j_token_embs). Furthermore, perform a linear transformation on the first token embedding (token_embs) by means of matrix multiplication to generate a second token embedding (j_token_embs). Step S6: Mix and encode the second token embedding (j_token_embs) and the position embedding (position_embs) into a mixed embedding (embs). Furthermore, the second token embedding (j_token_embs) and the position embedding (position_embs) are mixed and encoded into a mixed embedding (embs) by means of star multiplication. Step S7: Divide the mixed embedding (embs) into three paths, perform linear transformation, cumulative summation, and normalization respectively, and then perform correlation calculation to generate a predicted embedding. Specifically, perform a linear transformation on the mixed embedding (embs) to generate a first linearly transformed embedding q(embs), a second linearly transformed embedding k(embs), and a third linearly transformed embedding v(embs) respectively. Perform cumulative summation on the first linearly transformed embedding q(embs) to generate a first cumulative summation embedding (q_embs); perform cumulative summation on the second linearly transformed embedding k(embs) to generate a second cumulative summation embedding (k_embs); perform cumulative summation on the third linearly transformed embedding v(embs) to generate a third cumulative summation embedding (v_embs). Normalize the first cumulative sum word vector (q_embs) to generate the first normalized word vector a; normalize the second cumulative sum word vector (k_embs) to generate the second normalized word vector b; normalize the third cumulative sum word vector (v_embs) to generate the third normalized word vector c. Perform an association calculation on the first normalized word vector a, the second normalized word vector b, the third normalized word vector c, as well as the first token vector (token_embs) and the second token vector (j_token_embs) to generate the predicted word vector d. Step S8: Perform a forward calculation on the predicted word vector d to generate the output sequence (ys sequence) to predict the next word of the natural language word. As shown in Figure 2, the natural language word [[“today”]] is encoded by the sequence conversion layer 101 The input sequence [

[0036] ] (xs sequence [

[0036] ]) is input into the generative network 100, and the predicted word vector d is calculated forward at the output layer to output the output sequence [

[0078] ] (ys sequence [

[0078] ]), thereby predicting the next word [[“weather”]] of the natural language word "today". Although the embodiments of the present invention have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present invention.< / eos> < / eos> < / eos> < / eos> < / eos>

Claims

1. A generative network capable of rapid training and low-computing-power prediction, characterized in that, The generative network includes: A sequence conversion layer for encoding natural language into an input sequence; A grapheme word vector encoding layer for encoding the natural language words corresponding to the input sequence into first grapheme word vectors; A positional word vector encoding layer for encoding the natural language words corresponding to the input sequence into positional word vectors; A first linear conversion layer for linearly converting the first grapheme word vectors to generate second grapheme word vectors; A mixed word vector encoding layer for mixing and encoding the second grapheme word vectors and the positional word vectors into mixed word vectors; A mixed word vector processing layer for dividing the mixed word vectors into three paths, performing linear conversion, cumulative summation, and normalization processing respectively, and then performing correlation calculation to generate predicted word vectors; An output layer for performing forward calculation on the predicted word vectors to generate an output sequence.

2. The generative network according to claim 1, wherein The first linear conversion layer linearly converts the first grapheme word vectors by means of matrix multiplication calculation to generate second grapheme word vectors.

3. The generative network according to claim 1, characterized in that, The mixed word vector encoding layer mixes and encodes the second grapheme word vectors and the positional word vectors into mixed word vectors by means of star multiplication calculation.

4. The generative network according to claim 1, wherein The mixed word vector processing layer includes: a second linear conversion layer, a third linear conversion layer, and a fourth linear conversion layer, a first cumulative summation layer, a second cumulative summation layer, and a third cumulative summation layer, and a first normalization layer, a second normalization layer, and a third normalization layer; The second linear conversion layer is used for linearly converting the mixed word vectors to generate first linearly converted word vectors; The third linear conversion layer is used for linearly converting the mixed word vectors to generate second linearly converted word vectors; The fourth linear conversion layer is used for linearly converting the mixed word vectors to generate third linearly converted word vectors; A first cumulative summation layer for cumulatively summing the first linearly converted word vectors to generate first cumulatively summed word vectors; A second cumulative summation layer for cumulatively summing the second linearly converted word vectors to generate second cumulatively summed word vectors; A third cumulative summation layer for cumulatively summing the third linearly converted word vectors to generate third cumulatively summed word vectors; The first normalization layer is used for normalizing the first cumulatively summed word vectors to generate first normalized word vectors; The second normalization layer is used for normalizing the second cumulatively summed word vectors to generate second normalized word vectors; The third normalization layer is used for normalizing the third cumulatively summed word vectors to generate third normalized word vectors.

5. The generative network according to claim 4, wherein The mixed word vector processing layer further includes: a correlation calculation layer; The correlation calculation layer is used for performing correlation calculation on the first normalized word vectors, the second normalized word vectors, the third normalized word vectors, and the first grapheme word vectors and the second grapheme word vectors to generate predicted word vectors.

6. A prediction method for a generative network that can be quickly trained and has low computing power for prediction, characterized in that, The prediction method uses the generative network according to any one of claims 1 to 5 to predict natural language, and includes the following method steps: S1. Select a segment of natural language and input it into the sequence conversion layer of the generative network to encode the natural language into an input sequence; S2. Use the input sequence as the input and the output sequence as the output to train the generative network; S3. Select the natural language word to be predicted, input it into the sequence conversion layer of the generative network, and encode the natural language word to be predicted into an input sequence; S4. Encode the natural language word corresponding to the input sequence into a first morpheme word vector, and at the same time encode the natural language word corresponding to the input sequence into a position word vector; S5. Perform a linear transformation on the first morpheme word vector to generate a second morpheme word vector; S6. Mix and encode the second morpheme word vector and the position word vector into a mixed word vector; S7. Divide the mixed word vector into three paths, perform linear transformation, cumulative summation, and normalization respectively, and then perform an association calculation to generate a predicted word vector; S8. Perform a forward calculation on the predicted word vector to generate an output sequence and predict the next word of the natural language word.

7. The prediction method according to claim 6, characterized in that In step S5, the first morpheme word vector is linearly transformed by matrix multiplication to generate a second morpheme word vector.

8. The prediction method according to claim 6, characterized in that In step S6, the second morpheme word vector and the position word vector are mixed and encoded into a mixed word vector by star multiplication.

9. The prediction method according to claim 6, characterized in that In step S7, perform a linear transformation on the mixed word vector to generate a first linearly transformed word vector, a second linearly transformed word vector, and a third linearly transformed word vector respectively; Perform a cumulative summation on the first linearly transformed word vector to generate a first cumulative summation word vector; Perform a cumulative summation on the second linearly transformed word vector to generate a second cumulative summation word vector; Perform a cumulative summation on the third linearly transformed word vector to generate a third cumulative summation word vector; Perform a normalization process on the first cumulative summation word vector to generate a first normalized word vector; Perform a normalization process on the second cumulative summation word vector to generate a second normalized word vector; Perform a normalization process on the third cumulative summation word vector to generate a third normalized word vector.

10. The prediction method according to claim 9, wherein Perform an association calculation on the first normalized word vector, the second normalized word vector, the third normalized word vector, the first morpheme word vector, and the second morpheme word vector to generate a predicted word vector.

Citation Information

Patent Citations

  • Machine writing system based on natural language processing technology

    CN110287478A

  • Text sequence generation method and system

    CN111046178A

  • Super-long sequence processing method based on improved Transform model

    CN115510812A

  • Information processing method and system based on big data

    CN117236323A

  • Language processing method and apparatus

    US20170372696A1