Operation control method of generative pre-training GPT model and electronic device
By generating at least two probability vectors to predict the initial characters of the GPT model and as output when the conditions are met, the problem of low output efficiency of the GPT model is solved, and the effect of reducing time cost and improving output efficiency is achieved.
Patent Information
- Application Number
- CN202311762770.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-19
- Publication Date
- 2025-06-20
AI Technical Summary
When reasoning, the GPT model requires one character to reason, which leads to high time cost for output characters and low output efficiency.
By generating at least two probability vectors, at least two initial characters with an associated relationship are predicted, and when these initial characters satisfy the output conditions, these initial characters are used as output characters of the GPT model.
By reducing the number of runs of the GPT model, the time cost is reduced and the output efficiency of the GPT model is improved.
Smart Images

Figure CN120181071A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of machine learning technology, and particularly relates to a method for controlling the operation of a generative pre-trained GPT model and an electronic device. Background Art
[0002] With the development of technology, the Generative Pre-trained Transformer (GPT) model series has become a paradigm model in the field of artificial intelligence related to natural language. It is an abbreviation of (Generative Pre-trained Transformer transformation model). It is based on the Transformer architecture and uses a large amount of text data for training to achieve the understanding and generation of natural language. The "unidirectionality" of GPT comes from the fact that in the decoder of the Transformer, masked self-attention is used to block the characters behind the current token, preventing "seeing" what the next character is when predicting the next character, and predicting the next character based on the history before the current character.
[0003] However, when the GPT model is inferring, it needs to infer one character at a time. For each character, the GPT model needs to run completely once, resulting in a high time cost for outputting characters and a low output efficiency of the GPT model.
[0004] Therefore, there is an urgent need for a technical solution that can improve the output efficiency of the GPT model. Summary of the Invention
[0005] In view of this, this application provides a method for controlling the operation of a generative pre-trained GPT model and an electronic device to solve the technical defect that the GPT model has a low efficiency in outputting characters in the prior art, as follows:
[0006] A method for controlling the operation of a generative pre-trained GPT model, the method comprising:
[0007] Obtaining an input text by using the GPT model;
[0008] Generating at least two probability vectors according to the input text, the at least two probability vectors corresponding to a first vocabulary, the first vocabulary including a marker character and a plurality of first candidate characters, the marker character corresponding to a second vocabulary, the second vocabulary including a plurality of second candidate characters, and the statistical usage frequency of the first candidate characters being greater than the statistical usage frequency of the second candidate characters;
[0009] Based on the at least two probability vectors, at least two initial characters are obtained, each of the initial characters corresponding to one of the probability vectors, each of the initial characters being from the first vocabulary, and there is an association relationship between adjacent initial characters among the at least two initial characters;
[0010] When the at least two initial characters meet the output conditions, the at least two initial characters are determined as the output characters of the GPT model.
[0011] Preferably, for the above method, obtaining at least two initial characters according to the at least two probability vectors includes:
[0012] Based on the at least two probability vectors, first label data is obtained, the first label data including a plurality of candidate characters, the candidate characters being from the first vocabulary, and the candidate characters corresponding to first probability values;
[0013] Based on the first probability values, at least two initial characters having an association relationship are obtained among the candidate characters.
[0014] Preferably, for the above method, the at least two initial characters meeting the output conditions include:
[0015] All of the at least two initial characters are characters among the plurality of first candidate characters.
[0016] Preferably, for the above method, the at least two initial characters meeting the output conditions include:
[0017] There are candidate characters associated with the at least two initial characters among the plurality of first candidate characters.
[0018] Preferably, for the above method, when the at least two initial characters do not meet the output conditions, the method further includes:
[0019] If the first character among the at least two initial characters is the identification character, a target character is obtained among the plurality of second candidate characters, and the target character is determined as the output character of the GPT model;
[0020] If the last character among the at least two initial characters is the identification character, the characters among the at least two initial characters other than the last character are determined as the output characters of the GPT model.
[0021] Preferably, for the above method, obtaining a target character among the plurality of second candidate characters includes:
[0022] According to the configuration parameters, multiple first characters are filtered out from the second vocabulary according to the probability values corresponding to each of the second candidate characters in the second vocabulary, and the number of the first characters matches the configuration parameters;
[0023] Among the multiple first characters, a target character is obtained.
[0024] In the above method, preferably, the GPT model at least includes a target layer, and the target layer is used to obtain at least two initial characters according to the at least two probability vectors; when the at least two initial characters meet the output conditions, the at least two initial characters are determined as the output characters of the GPT model;
[0025] Wherein, the method further includes:
[0026] Using a first input sample and a first output sample to optimize the target layer;
[0027] Wherein, the first input sample includes at least two probability vector samples, and the first output sample includes at least two character samples.
[0028] In the above method, preferably, the method further includes:
[0029] Using a second input sample and a second output sample to optimize the GPT model;
[0030] Wherein, the second input sample includes a text sample, and the second output sample includes at least two character samples.
[0031] A computer-readable storage medium, the computer-readable storage medium includes a stored program, wherein the program executes the method described in any one of the above when running.
[0032] An electronic device includes a memory and a processor, a computer program is stored in the memory, and the processor is configured to execute the method described in any one of the above through the computer program.
[0033] As can be seen from the above technical solutions, in a method for controlling the operation of a generative pre-trained GPT model and an electronic device disclosed in the present application, the GPT model is used to obtain an input text, and then at least two probability vectors are generated according to the input text to predict at least two initial characters having an association relationship. When these initial characters meet the output conditions, these initial characters are used as the output characters of the GPT model. It can be seen that in the present application, by outputting at least two probability vectors in the GPT model, at least two characters can be output in one run, thereby reducing the time cost by reducing the number of runs of the GPT model, and thus improving the output efficiency of the GPT model. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] The drawings herein are incorporated into and form a part of this specification, showing embodiments consistent with the present application, and are used together with the specification to explain the principles of the present application.
[0035] To more clearly illustrate the technical solutions in the embodiments of the present application or in the prior art, the following briefly introduces the drawings required for use in the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0036] Figure 1a Schematic diagram of the hardware environment constituted by the terminal device and the server applicable to the present application;
[0037] Figure 1b Flowchart of a method for controlling the operation of a generative pre-trained GPT model provided in Embodiment 1 of the present application;
[0038] Figure 2 Example diagram of generating characters by the GPT model;
[0039] Figure 3 Schematic diagram of the GPT model in the present application outputting two characters through the CRF layer;
[0040] Figure 4 Another flowchart of a method for controlling the operation of a generative pre-trained GPT model provided in Embodiment 1 of the present application;
[0041] Figure 5 Schematic diagram of the structure of a device for controlling the operation of a generative pre-trained GPT model provided in Embodiment 2 of the present application;
[0042] Figure 6 and Figure 7 Another schematic diagram of the structure of a device for controlling the operation of a generative pre-trained GPT model provided in Embodiment 2 of the present application;
[0043] Figure 8 Schematic diagram of the structure of an electronic device provided in Embodiment 3 of the present application;
[0044] Figure 9 Schematic diagram of the GPT model in the present application generating characters;
[0045] Figure 10 Schematic diagram of the GPT model in the present application generating two characters based on a probability vector through the CRF layer;
[0046] Figure 11 Schematic diagram of the operation process of the GPT model in the prior art generating characters;
[0047] Figure 12 This is a schematic diagram of the operation process for generating characters by the GPT model in this application. Specific implementation manners
[0048] In order to enable those skilled in the art of this technology to better understand the solutions of this application, the technical solutions in the embodiments of this application will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all the embodiments. Based on the embodiments in this application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of this application.
[0049] It should be noted that the terms "first", "second", etc. in the specification, claims and the above-mentioned drawings of this application are used to distinguish similar objects, and do not necessarily need to be used to describe a specific order or sequence. It should be understood that such used data can be interchanged under appropriate circumstances so that the embodiments of this application described here can be implemented in an order other than those illustrated or described here. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device including a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0050] This application proposes an operation control method and an electronic device for a generative pre-trained GPT model, which can be applied to whole-house intelligent digital control application scenarios such as Smart Home, smart home, smart home appliance ecosystem, and Intelligence House ecosystem. Optionally, in this embodiment, the above-mentioned operation control method for the generative pre-trained GPT model can be applied to a hardware environment composed of a terminal device and a server, such as Figure 1a inside. The server is connected to the terminal device through a network and can be used to provide services such as text prediction for the terminal or the client installed on the terminal. Users can log in to the client to use the text prediction service on the server. In addition, a database is set up on the server or independently of the server to provide data storage services for the server, and cloud computing and / or edge computing services can be configured on the server or independently of the server to provide data operation services for the server.
[0051] The above network may include, but is not limited to, at least one of the following: a wired network, a wireless network. The above wired network may include, but is not limited to, at least one of the following: a wide area network, a metropolitan area network, a local area network. The above wireless network may include, but is not limited to, at least one of the following: WIFI (Wireless Fidelity), Bluetooth. The terminal device is not limited to a PC, a mobile phone, a tablet computer, etc.
[0052] Reference Figure 1b As shown in the figure, it is a flowchart of the implementation of a method for controlling the operation of a generative pre-training GPT model provided in Embodiment 1 of the present application. This method can be applied to the GPT model and can specifically run in the Conditional Random Fields (CRF) layer in the GPT model. The GPT model runs in an electronic device such as a computer or a server. The GPT model further includes at least an input layer input, a processing layer, and an output layer finally output. The processing layer may at least include a decoding layer block and a hidden layer hiddenstate, as Figure 2 shown in the figure. Among them, the input layer obtains the input character w, and the decoding layer (including N decoder blocks) decodes the input character w(1-n). N is a positive integer greater than or equal to 1, and n is the number of input characters w, and n is a positive integer greater than or equal to 2. The hidden layer outputs a probability vector h(1-n) for the decoded data, and the output layer outputs a predicted character according to the probability vector. Based on this, the GPT model sequentially executes the corresponding inference process from the input layer to the output layer based on the previous character "jie" to obtain the next character "shao". The technical solution in this embodiment is mainly used to improve the output efficiency of the GPT model.
[0053] Specifically, the method in this embodiment may include the following steps:
[0054] Step 101: Use the GPT model to obtain the input text.
[0055] Specifically, in this embodiment, the input text received by the input layer in the GPT model is obtained.
[0056] Step 102: Generate at least two probability vectors according to the input text.
[0057] Specifically, in this embodiment, at least two probability vectors output by the processing layer for the input text of the input layer can be obtained.
[0058] Among them, at least two probability vectors in this embodiment correspond to a first vocabulary. The first vocabulary contains identification characters and multiple first candidate characters. The identification characters correspond to a second vocabulary, and the second vocabulary contains multiple second candidate characters. The statistical usage frequency of the first candidate characters is greater than that of the second candidate characters.
[0059] Specifically, in this embodiment, the first probability vector and the second probability vector are taken as examples for illustration. For example, the first probability vector and the second probability vector correspond to a first vocabulary. The first vocabulary contains an identification character and multiple first candidate characters. The identification character corresponds to a second vocabulary. The first probability vector contains probability values respectively corresponding to each first candidate character and the identification character. The second probability vector contains probability values respectively corresponding to each first candidate character and the identification character. The probability values in the first probability vector correspond to the first character that the GPT model is about to output. The probability values in the second probability vector correspond to the second character that the GPT model is about to output. The first character refers to the next character generated by the GPT model for the input text of the input layer. The second character refers to the next character that the GPT model is about to output after the first character. In addition, the processing layer also outputs a third probability vector. The third probability vector corresponds to the second vocabulary. The second vocabulary contains multiple second candidate characters. The third probability vector contains probability values respectively corresponding to each second candidate character.
[0060] It should be noted that the statistical usage frequency of the first candidate characters is greater than that of the second candidate characters. For example, the first vocabulary is a common vocabulary, and the second vocabulary is a rare vocabulary.
[0061] Specifically, in this embodiment, characters involved in the current running scenario of the GPT model can be obtained in large quantities, and then counted according to the usage frequency in historical usage, and a part of the characters with higher statistical usage frequency and the identification characters representing the second vocabulary are selected and filled into the first vocabulary, and a part of the characters with lower statistical usage frequency are selected and filled into the second vocabulary. Or, in this embodiment, the first vocabulary and the second vocabulary can be pre-configured.
[0062] Step 103: Obtain at least two initial characters according to at least two probability vectors.
[0063] Among them, each initial character corresponds to a probability vector, and each initial character is from the first vocabulary. That is to say, each initial character is obtained according to the first candidate character in the first vocabulary or the identification character representing the second vocabulary. Moreover, there is an association relationship between adjacent two of at least two initial characters. For example, taking two initial characters as an example, both of the two initial characters are from the first candidate characters; or, the first of the two initial characters is from the first candidate character and the second character is an identification character; or, the first of the two initial characters is an identification character and the second character is from the first candidate character.
[0064] Specifically, the CRF layer can learn the bigram model of characters. Since there is a strong association between phrases in a specific field, in this embodiment, the CRF layer is used to make there be an association relationship between adjacent two initial characters, and this association relationship satisfies the rationality of combined words in the corresponding field.
[0065] Step 104: Determine whether at least two initial characters meet the output condition. If at least two initial characters meet the output condition, execute Step 105.
[0066] Specifically, the situation where the at least two initial characters meet the output condition can be: the at least two initial characters are from the first candidate characters, that is, there is no identification character among the at least two initial characters. The identification character is a character representing the second vocabulary. For example, the identification character "OTHERS" represents the second vocabulary. Based on this, when there is no identification character among the at least two initial characters, execute Step 105.
[0067] Step 105: Determine at least two initial characters as the output characters of the GPT model.
[0068] Specifically, taking two initial characters as an example, take the first of the two initial characters as the first character output by the output layer in the GPT model, and take the second of the two initial characters as the second character output by the output layer in the GPT model. Thus, for this second character, the GPT model does not need to execute an inference process from the input layer to the output layer once. Then, the GPT model executes an inference process from the input layer to the output layer again according to the second character to obtain the next character.
[0069] For example, as Figure 3 shown, take "artificial" as the first to second output characters of the GPT model after ":", and the GPT model does not need to execute an inference process from the input layer to the output layer again for "person", and an output character "worker" can be obtained. As Figure 3In the process shown by the seventh dotted line in the figure, the GPT model actually does not perform the reasoning process, that is, the GPT model needs to perform the reasoning process from w7 to h7 and to the output layer. Furthermore, the GPT model performs the next reasoning process from the input layer to the output layer based on the output characters such as "工" to obtain the next character. As a result, the GPT model saves the time of executing one reasoning process.
[0070] It can be seen from the above technical solution that in a method for controlling the operation of a GPT model provided in Example 1 of the present application, a GPT model is used to obtain input text, and then at least two probability vectors are generated based on the input text to predict at least two initial characters with an associated relationship. When these initial characters meet the output conditions, these initial characters are used as the output characters of the GPT model. It can be seen that in the present application, at least two characters can be output at one time by outputting at least two probability vectors in the GPT model, thereby reducing the time cost by reducing the number of operations of the GPT model, thereby improving the output efficiency of the GPT model.
[0071] In one implementation, when obtaining at least two initial characters in step 103, it can be implemented in the following manner:
[0072] First, first label data is obtained based on at least two probability vectors. The first label data may include multiple characters to be selected, which are derived from the first vocabulary, that is, the characters to be selected are derived from the first candidate characters and the identification characters, and the characters to be selected have corresponding first probability values. The first probability value is the probability value that the character to be selected may be the character to be output by the GPT model.
[0073] For example, the label output by the CRF layer contains common domain words and a unified symbol for uncommon words "OTHERS". Common domain words come from the constructed common vocabulary C in the domain, that is, the first vocabulary. The size of C can be 2000. Based on this, the label contains each word in C (that is, the character to be selected) and its corresponding first probability value.
[0074] Afterwards, at least two initial characters having an associated relationship are obtained from the characters to be selected according to the first probability value.
[0075] At this time, the at least two initial characters may include an identification character, that is, only one initial character is obtained from the first candidate characters; or, the at least two initial characters may not include an identification character, that is, both initial characters are obtained from the first candidate characters.
[0076] In one implementation, the at least two initial characters satisfy an output condition, including:
[0077] The at least two initial characters are all characters among the multiple first candidate characters, that is, the initial characters all originate from the first candidate characters and do not include identification characters.
[0078] In another implementation, the at least two initial characters satisfy the output condition, including:
[0079] The at least two initial characters are associated with candidate characters among the multiple first candidate characters, that is, the initial characters are associated with the first candidate characters.
[0080] Based on this, when the at least two initial characters satisfy the output condition, step 104 is executed, that is: the two initial characters are determined as the output characters of the output layer in the GPT model, and the output characters are used for the GPT model to output the next character.
[0081] Further, when the at least two initial characters do not satisfy the output condition, such as when these initial characters contain identification characters, the method in this embodiment can also include the following processing, such as Figure 4 as shown in
[0082] Step 106: Determine whether the identification character in the at least two initial characters is the first character or the second character; if the first character in the at least two initial characters is the identification character, execute step 107, if the last character in the at least two initial characters is the identification character, execute step 108;
[0083] Step 107: Obtain a target character among the multiple second candidate characters included in the second vocabulary, and use the target character as the output character of the output layer in the GPT model, and ignore other characters. At this time, only one character, that is, the target character, is output by the output layer.
[0084] Step 108: Determine the other characters in the at least two initial characters except the last character as the output characters of the output layer in the GPT model.
[0085] At this time, if there are only two initial characters, then only the first character among these initial characters is output by the output layer, and other characters are ignored.
[0086] Among them, if the initial characters contain identification characters, then the probability that the candidate characters in the first vocabulary are the output characters of the GPT model is lower than that of the candidate characters in the second vocabulary, that is, the probability that the candidate characters in the second vocabulary are the output characters of the GPT model is higher than that of the candidate characters in the first vocabulary. At this time, the output characters of the output layer are determined according to the position of the identification character in the initial characters.
[0087] Specifically, if the first character in the initial character is an identification character, a candidate character can be obtained from the second vocabulary as the target character, and then this target character is used as the output character of the output layer. If the last character in the initial character is an identification character, the identification character can be removed, and only the other characters in the initial character except the last character are used as the output character of the output layer.
[0088] In one implementation, in this embodiment, according to the configuration parameters and the probability values corresponding to each second candidate character in the second vocabulary, multiple first characters can be screened out from the second vocabulary, and then a target character is obtained from these multiple first characters. For example, a character is randomly selected from these multiple first characters as the target character to increase the diversity of the text output by the GPT model, or the character with the largest probability value is selected from these multiple first characters as the target character.
[0089] It should be noted that the number of first characters matches the configuration parameters. That is to say, in this embodiment, according to the screening quantity indicated in the configuration parameters, the corresponding number of second candidate characters whose probability values meet the screening conditions are screened out from the multiple second candidate characters in the second vocabulary, that is, the first characters. The probability value meeting the screening conditions can be understood as: the probability values are sorted from large to small. That is to say, in this embodiment, after sorting the second candidate characters in the order of probability values from large to small, the corresponding number of second candidate characters indicated by the configuration parameters are screened out, that is, the first characters.
[0090] For example, according to the screening quantity k indicated in the configuration parameters, the characters ranked in the top k in the order of probability values from large to small are screened out from the rare vocabulary as the first characters, and then a character is randomly selected or the character with the largest probability value is selected from these k first characters as the target character. k is a positive integer greater than 1.
[0091] Furthermore, in this embodiment, the configuration parameters can be adjusted according to requirements. Specifically, in this embodiment, a modification input operation for the configuration parameters can be obtained first, and then, according to the modification parameters in the modification input operation, the parameter values of the configuration parameters are adjusted.
[0092] For example, in this embodiment, a configuration modification interface can be provided for the user. In the configuration modification interface, a modification input box corresponding to the configuration parameter is output, and the user can enter the required parameter value in the modification input box. Based on this, on the device where the configuration modification interface is located, a modification input operation can be generated based on the parameter value input by the user. In this embodiment, after receiving the modification input operation of the user for the configuration parameter, the modification parameter in the modification input operation, such as the parameter value input by the user, can be parsed. Then, according to the parsed parameter value input by the user, the parameter value of the configuration parameter is adjusted, that is, the screening quantity indicated in the configuration parameter is modified to the parameter value input by the user.
[0093] In one implementation, the GPT model includes at least a target layer such as the CRF layer. The target layer is used to obtain at least two initial characters according to at least two probability vectors; when the at least two initial characters meet the output condition, the at least two initial characters are determined as the output characters of the output layer in the GPT model. Based on this, in this embodiment, only the target layer, that is, the CRF layer, can be optimized using the first input sample and the first output sample.
[0094] Among them, the first input sample includes at least two probability vector samples, and the first output sample includes at least two character samples. Specifically, the first input sample is a probability vector sample and is three probability vectors. For example, the first input sample contains a first vector sample, a second vector sample, and a third vector sample. The first vector sample and the second vector sample correspond to the first vocabulary. The first vector sample contains sample probability values corresponding to the identification character and each first candidate character respectively. The second vector sample contains sample probability values corresponding to the identification character and each first candidate character respectively. The third vector sample corresponds to the second vocabulary, and the third vector sample contains probability values corresponding to each second candidate character respectively. And the first output sample is a character sample and is at least two. These character samples may contain an identification character, and the identification character is a character representing the second vocabulary; or these character samples may not contain an identification character.
[0095] Thus, by calculating the loss value of the characters output by the CRF layer and the first output sample according to the corresponding loss function, and then adjusting the parameters in the CRF layer according to the loss value, the separate training of the CRF layer can be realized. Based on this, after optimizing the CRF layer using the first input sample and the first output sample, the CRF layer can output two characters for the input three probability vectors, and one of the characters may be an identification character.
[0096] In another implementation, in this embodiment, the GPT model can also be globally optimized using the second input sample and the second output sample. The process of global optimization includes optimizing the target layer, i.e., the CRF layer. The second input sample includes text samples, and the second output sample includes at least two character samples.
[0097] Thus, by calculating the loss value between the characters output by the GPT model and the second output sample according to the corresponding loss function, and then optimizing and adjusting the parameters in the GPT model according to the loss value, the global optimization of the GPT model is achieved. Based on this, after optimizing the GPT model using the second input sample and the second output sample, the CRF layer in the GPT model can output two characters for the input three probability vectors, and one of the characters may be an identification character.
[0098] Reference Figure 5 , is a schematic structural diagram of a running control device for a generative pre-trained GPT model provided in the second embodiment of the present application. This device can be applied to the GPT model and can specifically run in the CRF layer of the GPT model. The GPT model runs in an electronic device such as a computer or a server, and the GPT model is as Figure 2 shown. The technical solution in this embodiment is mainly used to improve the output efficiency of the GPT model.
[0099] Specifically, the device in this embodiment may include the following units:
[0100] The text acquisition unit 501 is used to obtain the input text using the GPT model.
[0101] The vector acquisition unit 502 is used to generate at least two probability vectors according to the input text. The at least two probability vectors correspond to a first vocabulary. The first vocabulary contains an identification character and a plurality of first candidate characters. The identification character corresponds to a second vocabulary, and the second vocabulary contains a plurality of second candidate characters. The statistical usage frequency of the first candidate characters is greater than the statistical usage frequency of the second candidate characters;
[0102] The character acquisition unit 503 is used to obtain at least two initial characters according to the at least two probability vectors. Each initial character corresponds to one of the probability vectors. Each initial character is from the first vocabulary, and there is an association relationship between adjacent initial characters among the at least two initial characters;
[0103] The character determination unit 504 is used to determine the at least two initial characters as the output characters of the GPT model when the at least two initial characters meet the output conditions.
[0104] As can be seen from the above technical solution, in the operation control device of a GPT model provided in the second embodiment of the present application, an input text is obtained by using the GPT model, and then at least two probability vectors are generated according to the input text to predict at least two initial characters with an associated relationship. When these initial characters meet the output conditions, these initial characters are used as the output characters of the GPT model. It can be seen that in the present application, by outputting at least two probability vectors in the GPT model, at least two characters can be output in one run, thereby reducing the time cost by reducing the number of runs of the GPT model, and thus improving the output efficiency of the GPT model.
[0105] In one implementation manner, the character obtaining unit 503 is specifically configured to: obtain first label data according to the at least two probability vectors, where the first label data includes a plurality of candidate characters, the candidate characters are from the first vocabulary, and the candidate characters correspond to first probability values; obtain at least two initial characters with an associated relationship among the candidate characters according to the first probability values.
[0106] Among them, the at least two initial characters meeting the output conditions include: the at least two initial characters are all characters in the plurality of first candidate characters. Or, the at least two initial characters meeting the output conditions include: the at least two initial characters are associated with candidate characters among the plurality of first candidate characters.
[0107] In one implementation manner, the character determining unit 504 is further configured to: when the at least two initial characters do not meet the output conditions, if the first character among the at least two initial characters is the identification character, obtain a target character among the plurality of second candidate characters, and determine the target character as the output character of the GPT model; if the last character among the at least two initial characters is the identification character, determine the characters other than the last character among the at least two initial characters as the output characters of the GPT model.
[0108] In one implementation manner, when the character determining unit 504 obtains a target character among the plurality of second candidate characters, it is specifically configured to: according to the configuration parameter, screen out a plurality of first characters in the second vocabulary according to the probability values corresponding to each of the second candidate characters in the second vocabulary, where the number of the first characters matches the configuration parameter; obtain a target character among the plurality of first characters.
[0109] In one implementation manner, the present embodiment may further include the following units, as Figure 6 shown in
[0110] A parameter adjustment unit 505 is configured to: obtain a modification input operation for the configuration parameter; and adjust the parameter value of the configuration parameter according to the modification parameter in the modification input operation.
[0111] In one implementation, the GPT model at least includes a target layer, and the target layer is configured to obtain at least two initial characters according to the at least two probability vectors; when the at least two initial characters meet the output condition, determine the at least two initial characters as the output characters of the GPT model; this embodiment may further include at least one of the following units, such as Figure 7 as shown in
[0112] A first training unit 506 is configured to perform an optimization process on the target layer by using a first input sample and a first output sample; wherein, the first input sample includes at least two probability vector samples, and the first output sample includes at least two character samples.
[0113] A second training unit 507 is configured to perform an optimization process on the GPT model by using a second input sample and a second output sample; wherein, the second input sample includes a text sample, and the second output sample includes at least two character samples.
[0114] It should be noted that the specific implementation of each unit in this embodiment may refer to the corresponding content in the foregoing, and will not be elaborated herein.
[0115] In addition, an embodiment of the present application further claims protection for a computer-readable storage medium, and the computer-readable storage medium includes a stored program, wherein when the program runs, it executes the operation control method of the generative pre-training GPT model as described in any of the foregoing embodiments.
[0116] Refer to Figure 8 , which is a schematic structural diagram of an electronic device provided in Embodiment 3 of the present application. The electronic device may include the following structure:
[0117] A memory 801 is configured to store a computer program and data generated by the running of the computer program;
[0118] The processor 802 is configured to implement, via a computer program: obtaining an input text by using the GPT model; generating at least two probability vectors according to the input text, the at least two probability vectors corresponding to a first vocabulary, the first vocabulary including identification characters and a plurality of first candidate characters, the identification characters corresponding to a second vocabulary, the second vocabulary including a plurality of second candidate characters, and the statistical usage frequency of the first candidate characters being greater than that of the second candidate characters; obtaining at least two initial characters according to the at least two probability vectors, each initial character corresponding to one of the probability vectors, each initial character being from the first vocabulary, and there being an association relationship between adjacent initial characters among the at least two initial characters; and determining the at least two initial characters as the output characters of the GPT model when the at least two initial characters meet the output conditions.
[0119] As can be seen from the above technical solution, in the electronic device provided in the third embodiment of the present application, an input text is obtained by using the GPT model, and then at least two probability vectors are generated according to the input text to predict at least two initial characters having an association relationship. When these initial characters meet the output conditions, these initial characters are used as the output characters of the GPT model. It can be seen that in the present application, at least two characters can be output in one run by outputting at least two probability vectors in the GPT model, thereby reducing the time cost by reducing the number of runs of the GPT model, and thus improving the output efficiency of the GPT model.
[0120] Taking the scenario of generating a description text of artificial intelligence through the GPT model as an example, the following is an example of the specific application of the present application:
[0121] In view of the current situation that the GPT model needs to infer one character at a time during inference, considering the high time cost of single-step generation, the present application proposes a two-step generation algorithm based on CRF, and it is speculated that the time cost and computing resource consumption can be greatly reduced.
[0122] Reference Figure 9 As shown, it is the algorithm schematic diagram of the GPT model in the present application. The Emission score is the input layer of the CRF, specifically as follows:
[0123] In the GPT model, for the iteration process at h5, 3 vectors will be generated at h5, namely the vocabulary probability vector of the next word next, the vocabulary probability vector of the word after the next word after next, and the vocabulary probability vector of rare words OTHERS. Among them, the vocabulary probability vector of next and the vocabulary probability vector of after next are used as the input of the CRF layer, and then the CRF layer will output 2 characters simultaneously as the additional items for the next iteration input;
[0124] If the CRF layer outputs <others>, then the top-k (the top k most probable word samplings) of the OTHERS vector is adopted as the output. If OTHERS is generated at the next position of the CRF layer, the label output by the CRF layer at the position after next is not adopted, and directly enters the next iteration, that is, single-step generation, and double-step generation will not occur, avoiding the problem of low accuracy of the CRF layer in processing OTHERS words.
[0125] Among them, the label output by the CRF layer is the unified symbol of common domain words and rare words <others>, a common vocabulary C is constructed here. The size of C can be set to about 2000, and the label is C+ <others>, the quantity is the length of C + 1, <others>It is a unified symbol for other non-common words.
[0126] It should be noted that the CRF layer can learn the bi-gram model of words. Due to the strong correlation of phrases in a specific domain, the CRF layer is used to ensure the rationality of combined words in a small domain.
[0127] In addition, in this application, by configuring the nbest parameter of CRF decoding, that is, the sequence path with the top nbest probabilities, the sampling effect of topp-topk can be achieved to ensure the diversity of the generated text.
[0128] Specifically, in the pre-training model stage of this application, the training of the CRF layer is increased. The loss functions of the label of the next word and the label of the word after the next word will be accumulated. The increase in the number of parameters can be ignored, and the increase in the training cost can basically also be ignored.
[0129] Such as Figure 10 As shown, it is the schematic diagram of generating two characters Y1 and Y2 based on two probability vectors X1 and X2 in the CRF layer:
[0130] First of all, the parameters of the CRF layer can be trained and learned separately, or jointly trained with GPT on domain data to increase the training flexibility.
[0131] Secondly, the number of output labels of the CRF layer is related to the size C of the common word list in the domain and can be configured separately for different domains. C should not be too large, otherwise the accuracy of the CRF layer will be affected.
[0132] It can be seen that if the technical solution of this application is not adopted, the execution process of the GPT model when outputting characters is as Figure 11 shown. For the GPT model to output one character, it needs to execute a processing process once; if the technical solution of this application is adopted, the GPT model can execute a processing process to output two related characters, as Figure 12 shown. Thus, compared with the GPT model running 5 times to output "Introduce AI: Artificial Intelligence", the template-based GPT only runs 3 times of GPT, greatly reducing the computational resources and time consumption.
[0133] It should be noted that the computational consumption and time cost of the CRF layer are basically negligible compared to the GPT model.
[0134] In summary, in this application, by adding a CRF layer to the GPT model and utilizing the effectiveness of the next and after next probability vectors and the accuracy of the CRF, a double-step fast generation algorithm of GPT within the domain is realized, which can multiply the GPT generation efficiency and reduce the consumption of model inference computing resources.
[0135] The above are only the preferred embodiments of the present application. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present application, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present application.< / others> < / others> < / others> < / others>
Claims
1. A method for controlling the operation of a generative pre-trained GPT model, characterized in that, The method includes: Obtaining an input text by using the GPT model; Generating at least two probability vectors according to the input text, the at least two probability vectors corresponding to a first vocabulary, the first vocabulary including identification characters and a plurality of first candidate characters, the identification characters corresponding to a second vocabulary, the second vocabulary including a plurality of second candidate characters, and the statistical usage frequency of the first candidate characters being greater than that of the second candidate characters; Obtaining at least two initial characters according to the at least two probability vectors, each of the initial characters corresponding to one of the probability vectors, each of the initial characters being from the first vocabulary, and there being an association relationship between adjacent initial characters among the at least two initial characters; When the at least two initial characters meet the output condition, determining the at least two initial characters as the output characters of the GPT model.
2. The method according to claim 1, characterized in that, Obtaining at least two initial characters according to the at least two probability vectors includes: Obtaining first label data according to the at least two probability vectors, the first label data including a plurality of candidate characters, the candidate characters being from the first vocabulary, and the candidate characters corresponding to first probability values; Obtaining at least two initial characters with an association relationship among the candidate characters according to the first probability values.
3. The method according to claim 2, characterized in that, The at least two initial characters meeting the output condition includes: The at least two initial characters are all characters among the plurality of first candidate characters.
4. The method according to claim 2, characterized in that, The at least two initial characters meeting the output condition includes: There are candidate characters associated among the at least two initial characters in the plurality of first candidate characters.
5. The method according to claim 2, characterized in that, When the at least two initial characters do not meet the output condition, the method further includes: If the first character among the at least two initial characters is the identification character, obtaining a target character among the plurality of second candidate characters and determining the target character as the output character of the GPT model; If the last character among the at least two initial characters is the identification character, determining the other characters except the last character among the at least two initial characters as the output characters of the GPT model.
6. The method according to claim 5, characterized in that, Obtaining a target character among the plurality of second candidate characters includes: According to configuration parameters, screening out a plurality of first characters in the second vocabulary according to the probability values corresponding to each of the second candidate characters in the second vocabulary, the number of the first characters matching the configuration parameters; Obtaining a target character among the plurality of first characters.
7. The method according to claim 1 or 2, characterized in that, The GPT model at least includes a target layer for obtaining at least two initial characters according to the at least two probability vectors; When the at least two initial characters meet the output condition, determining the at least two initial characters as the output characters of the GPT model; Wherein, the method further includes: Using a first input sample and a first output sample to perform an optimization process on the target layer; Wherein, the first input sample includes at least two probability vector samples, and the first output sample includes at least two character samples.
8. The method according to claim 1 or 2, characterized in that, The method further includes: using a second input sample and a second output sample to perform an optimization process on the GPT model; wherein, the second input sample includes text samples, and the second output sample includes at least two character samples.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a stored program, wherein the program, when running, executes the method according to any one of claims 1 to 8.
10. An electronic device, comprising a memory and a processor, characterized in that, A computer program is stored in the memory, and the processor is configured to execute the method according to any one of claims 1 to 8 through the computer program.
Citation Information
Cited By
Space-time backtracking personal simulation image generation method and system based on real person reality
CN120852602A