Operation control method of generative pre-training GPT model and electronic device

By finding matching target keywords in the keyword set to obtain the output characters of the GPT model, the problem of low output efficiency of the GPT model is solved, and the effect of reducing the number of runs and improving the output efficiency is achieved.

CN120181074APending Publication Date: 2025-06-20QINGDAO HAIER TECH +2
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202311762771.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-19
Publication Date
2025-06-20

AI Technical Summary

Technical Problem

When reasoning, the GPT model requires one character to reason, which leads to high time cost for output characters and low output efficiency.

Method used

By finding the target keywords that match the first character output by the GPT model in the preset keyword set, the second character in the target keyword is obtained as the output character of the GPT model, thereby reducing the number of runs of the GPT model.

Benefits of technology

The output characters can be obtained without running the GPT model again, saving time in executing the inference process and improving the output efficiency of the GPT model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120181074A_ABST
    Figure CN120181074A_ABST
Patent Text Reader

Abstract

The invention discloses an operation control method for a generative pre-training GPT model and an electronic device, and relates to the technical field of smart home / smart home, and the method comprises the steps: obtaining a first character outputted by a GPT model; in a preset keyword set, searching for a target keyword matched with the first character; the keyword set comprises a plurality of keywords to be selected, and the keywords to be selected comprise a plurality of characters to be selected; obtaining at least one second character at least according to a character to be selected in the target keyword; and determining the at least one second character as an output character of the GPT model, so that the GPT model outputs a next character at least according to the second character.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of machine learning, and in particular, to a method for controlling the operation of a generative pre-trained GPT model and an electronic device. Background Art

[0002] With the development of technology, the generative pre-trained GPT (Generative Pre-trained Transformer) model series has become a paradigm model in the field of artificial intelligence related to natural language. It is an abbreviation of (Generative Pre-trained Transformer transformation model). It is based on the Transformer architecture and uses a large amount of text data for training to achieve the understanding and generation of natural language. The "unidirectionality" of GPT comes from the fact that in the decoder of the Transformer, it uses masked self-attention to block the characters behind the current character when renting a house, preventing it from "seeing" what the next character is when predicting the next character, and predicting the next character based on the history before the current character.

[0003] However, when the GPT model is reasoning, it needs to reason character by character. For each character, the GPT model needs to run completely once, resulting in a high time cost for outputting characters and a low output efficiency of the GPT model.

[0004] Therefore, there is an urgent need for a technical solution that can improve the output efficiency of the GPT model. Summary of the Invention

[0005] In view of this, this application provides a method and device for controlling the operation of a generative pre-trained GPT model to solve the technical defect that the output efficiency of the GPT model in the prior art is low, as follows:

[0006] A method for controlling the operation of a generative pre-trained GPT model, the method comprising:

[0007] Obtain a first character output by the GPT model;

[0008] In a preset keyword set, search for a target keyword that matches the first character; the keyword set contains multiple candidate keywords, and the candidate keywords contain multiple candidate characters;

[0009] Obtain at least one second character at least according to the candidate characters in the target keyword;

[0010] Determine the at least one second character as the output character of the GPT model, so that the GPT model outputs the next character at least according to the second character.

[0011] Preferably, for the above method, the candidate keywords in the keyword set are related to at least one of the following:

[0012] The running scenario of the GPT model, which is determined based on the input text of the GPT model;

[0013] The output parameters of the GPT model, where the output parameters at least characterize the text output format of the GPT model.

[0014] Preferably, for the above method, the keyword set is obtained by the following method:

[0015] Obtain the text output format of the GPT model;

[0016] Extract the format description words from the text output format;

[0017] Add the format description words as candidate keywords to the keyword set.

[0018] Preferably, for the above method, the keyword set is obtained by the following method:

[0019] Obtain the input text of the GPT model;

[0020] Parse the input text to obtain the running scenario of the GPT model;

[0021] According to the running scenario, obtain multiple candidate keywords, where the meanings of the candidate keywords are related to the running scenario;

[0022] Add the candidate keywords to the keyword set.

[0023] Preferably, for the above method, the keyword set is obtained by the following method:

[0024] Obtain the keyword configuration information for the GPT model;

[0025] Extract the keywords from the keyword configuration information to obtain multiple candidate keywords;

[0026] Add the candidate keywords to the keyword set.

[0027] Preferably, for the above method, the target keyword matching the first character includes:

[0028] The first character in the target keyword is the same as the first character.

[0029] Preferably, for the above method, obtaining at least one second character according to at least the candidate character in the target keyword includes:

[0030] Extract the other candidate characters except the first character in the target keyword, and the other candidate characters are the second characters.

[0031] In the above method, preferably, when no target keyword matching the first character is found in the keyword set, the method further includes:

[0032] Determine the first character as the output character of the GPT model, so that the GPT model outputs the next character at least based on the first character.

[0033] An operating control device for a generative pre-trained GPT model, the device includes:

[0034] A character acquisition unit for acquiring a first character output by the GPT model;

[0035] A keyword search unit for searching for a target keyword matching the first character in a preset keyword set; the keyword set includes multiple candidate keywords, and the candidate keywords include multiple candidate characters;

[0036] A character extraction unit for obtaining at least one second character at least based on the candidate characters in the target keyword;

[0037] A character determination unit for determining the at least one second character as the output character of the GPT model, so that the GPT model outputs the next character at least based on the second character.

[0038] A computer-readable storage medium, the computer-readable storage medium includes a stored program, wherein the program executes the method described in any one of the above when running.

[0039] An electronic device includes a memory and a processor, a computer program is stored in the memory, and the processor is configured to execute the method described in any one of the above through the computer program.

[0040] As can be seen from the above technical solution, in a method for controlling the operation of a generative pre-trained GPT model and an electronic device disclosed in the present application, after obtaining the first character output by the GPT model and before the GPT model performs the next operation, first search for a target keyword that matches the first character in a preset keyword set, and then use the second character in the target keyword as the output character of the GPT model. In this way, the output character can be obtained without the GPT model running again. After that, the GPT model can use the second character as the output character of the previous time to perform the inference and prediction of the next character. It can be seen that in the present application, the output character of the GPT model is predicted by configuring the keyword set, without having to let the GPT model run one more time, thereby reducing the time cost by reducing the number of times the GPT model runs, and thus improving the output efficiency of the GPT model. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] The accompanying drawings herein are incorporated into and constitute a part of this specification, showing embodiments consistent with the present application and, together with the specification, are used to explain the principles of the present application.

[0042] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the accompanying drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can be obtained based on these drawings.

[0043] Figure 1a It is a schematic diagram of the hardware environment composed of the terminal device and the server applicable to the present application;

[0044] Figure 1b It is a flowchart of a method for controlling the operation of a generative pre-trained GPT model provided in Embodiment 1 of the present application;

[0045] Figure 2 It is an example diagram for the GPT model to generate characters;

[0046] Figure 3 and Figure 4 They are respectively schematic diagrams of the generative pre-trained GPT model in the present application for determining the output character according to the keyword set;

[0047] Figure 5 It is another flowchart of a method for controlling the operation of a generative pre-trained GPT model provided in Embodiment 1 of the present application;

[0048] Figure 6 It is a schematic structural diagram of a device for controlling the operation of a generative pre-trained GPT model provided in Embodiment 2 of the present application;

[0049] Figure 7A schematic structural diagram of an electronic device provided in Embodiment 3 of the present application;

[0050] Figure 8 A schematic diagram of the GPT model in the present application generating characters;

[0051] Figure 9 A schematic diagram of the operation process of the GPT model in the prior art generating demo text;

[0052] Figure 10 A schematic diagram of the operation process of the GPT model in the present application generating text. Detailed implementation manners

[0053] In order to enable those skilled in the art to better understand the solution of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.

[0054] It should be noted that the terms "first", "second", etc. in the specification, claims and above-mentioned drawings of the present application are used to distinguish similar objects, and do not have to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present application described here can be implemented in an order different from those illustrated or described here. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device including a series of steps or units does not have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0055] The present application proposes an operation control method and an electronic device for a generative pre-trained GPT model, which can be applicable to whole-house intelligent digital control application scenarios such as Smart Home, smart home, smart home appliance ecosystem, and Intelligence House ecosystem. Optionally, in this embodiment, the above operation control method for the generative pre-trained GPT model can be applied to a hardware environment composed of a terminal device and a server, such as Figure 1aAs shown. The server is connected to the terminal device through a network and can be used to provide services such as text prediction for the terminal or the client installed on the terminal. Users can log in to the client to use the text prediction service on the server. In addition, a database is set up on the server or independently of the server to provide data storage services for the server, and cloud computing and / or edge computing services can be configured on the server or independently of the server to provide data operation services for the server.

[0056] The above network can include, but is not limited to, at least one of the following: wired network, wireless network. The above wired network can include, but is not limited to, at least one of the following: wide area network, metropolitan area network, local area network. The above wireless network can include, but is not limited to, at least one of the following: WIFI (Wireless Fidelity), Bluetooth. The terminal device is not limited to a PC, mobile phone, tablet computer, etc.

[0057] Refer to Figure 1b As shown, it is a flowchart of the implementation of a method for controlling the operation of a generative pre-trained GPT model provided in Embodiment 1 of the present application. This method can be applied to an electronic device capable of running the GPT model, such as a computer or a server, etc. The technical solution in this embodiment is mainly used to improve the output efficiency of the GPT model.

[0058] Specifically, the method in this embodiment can include the following steps:

[0059] Step 101: Obtain the first character output by the GPT model.

[0060] Among them, the first character is the character obtained by the GPT model performing a corresponding inference process at least based on the previous character output by the GPT model.

[0061] For example, as Figure 2 shown, the GPT model has the following structure: input layer input, decoder layer decoder block, hidden layer hidden state, output layer finally output, etc. Among them, the input layer obtains the input character w, and the decoder layer (including N decoder blocks) decodes the input character w(1-n), where N is a positive integer greater than or equal to 1, n is the number of input characters w, and n is a positive integer greater than or equal to 2. The hidden layer outputs a probability vector h(1-n) for the decoded data, and the output layer outputs a predicted character according to the probability vector. Based on this, the GPT model performs a corresponding inference process in sequence from the input layer to the output layer based on the previous character "fen" to obtain the output character, that is, the first character "xi".

[0062] Step 102: In the preset keyword set, search for the target keyword that matches the first character; the keyword set contains multiple candidate keywords.

[0063] Among them, the candidate keywords contain multiple candidate characters. For example, the keyword set at least contains candidate keywords: "start", "end", "theme", and "theme support point". Taking the candidate keyword "theme" as an example, it contains two candidate characters "main" and "topic".

[0064] Specifically, in this embodiment, according to the first character, the first character is compared with each candidate character included in each candidate keyword in the keyword set, and then the candidate keyword whose comparison result indicates a match with the first character is used as the target keyword.

[0065] In one implementation, the target keyword matching the first character may include:

[0066] The first character in the target keyword is consistent with the first character.

[0067] For example, taking the first character as "main" as an example, "main" is compared with the first character of each candidate keyword in the keyword set, and thus the candidate keyword "theme" whose first character is "main" is found and determined as the target keyword.

[0068] It should be noted that there may be multiple candidate keywords that match the first character. At this time, one keyword can be selected as the target keyword from these multiple candidate keywords that match the first character according to multiple historical output characters of the GPT model, and the semantics of the target keyword match the semantics of the multiple historical output characters of the GPT model.

[0069] For example, taking the first character as "main" as an example, "main" is compared with the first character of each candidate keyword in the keyword set, and thus the candidate keywords "theme" and "theme support point" whose first character is "main" are found. Then, combined with the character "Analyze the theme ***;" that the GPT model has already output, it can be inferred that "theme support point" is the keyword that matches the first character and also matches the historical output characters of the GPT model in semantics, and it is determined as the target keyword.

[0070] Step 103: Obtain at least one second character based on at least the candidate characters in the target keyword.

[0071] In one implementation, in step 103, all candidate characters in the target keyword can be determined as the second characters;

[0072] In another implementation, in step 103, other candidate characters in the target keyword that are different from the first character can be determined as the second character.

[0073] Specifically, when the first character of the target keyword is the same as the first character, in step 103, other candidate characters in the target keyword except the first character can be extracted, and these extracted other candidate characters are the second characters.

[0074] For example, taking the first character as "main" as an example, the target keyword "theme" is found in the keyword set. Based on this, "theme" in the "topic" is used as the second character.

[0075] Another example, taking the first character as "main" as an example, the target keyword "theme support point" is found in the keyword set. Based on this, at least one of "topic", "support", "hold", and "point" in "theme support point" is used as the second character.

[0076] Step 104: Determine at least one second character as the output character of the GPT model, so that the GPT model outputs the next character at least based on the second character.

[0077] For example, as Figure 3 shown, the second character "topic" is used as the first to fourth output characters of the GPT model after "main", and there is no need for the GPT model to perform an inference process from the input layer to the output layer for "main" again to obtain an output character. As Figure 3 shown in the fourth dotted line process, actually the GPT model does not perform the inference process. Further, the GPT model then performs the next inference process from the input layer to the output layer based on output characters such as "topic" to obtain the next character. Thus, the GPT model saves the time of performing an inference process.

[0078] Another example, as Figure 4 shown, the second characters "topic", "support", "hold", and "point" are used as the first to fourth output characters of the GPT model after "main", and there is no need for the GPT model to perform multiple inference processes from the input layer to the output layer for "main" again to obtain multiple output characters. Further, the GPT model then performs the next inference process from the input layer to the output layer based on output characters such as "point" to obtain the next character. Thus, the GPT model saves the time of performing four inference processes.

[0079] As can be seen from the above solution, in the operation control method of the generative pre-trained GPT model provided in the first embodiment of the present application, after obtaining the first character output by the GPT model, before the GPT model performs the next operation, first search for the target keyword that matches the first character in the keyword set, and then use the second character in the target keyword as the output character of the GPT model. In this way, the output character can be obtained without the GPT model running again. After that, the GPT model can use the second character as the output character of the previous time to perform the inference and prediction of the next character. It can be seen that in this embodiment, the output character of the GPT model is predicted by configuring the keyword set, without having to let the GPT model run one more time, thereby reducing the time cost by reducing the number of runs of the GPT model, and thus improving the output efficiency of the GPT model.

[0080] In one implementation, the candidate keywords in the keyword set are related to at least one of the following:

[0081] The running scenario of the GPT model, which is determined based on the input text of the GPT model;

[0082] The output parameters of the GPT model, and the output parameters at least characterize the text output format of the GPT model.

[0083] Specifically, in one implementation, the keyword set is obtained in the following manner in this embodiment:

[0084] First, obtain the text output format of the GPT model, and the format description words of the output text are configured in the text output format. For example, the text output format is: "Start: ***; End: ***" and "Theme: ***; Theme support points: ***".

[0085] After that, extract the format description words in the text output format, such as "Start", "End", "Theme" and "Theme support points".

[0086] Finally, add these format description words as candidate keywords to the keyword set.

[0087] It should be noted that after the text output format of the GPT model is adjusted, the candidate keywords in the corresponding keyword set of the GPT model are modified accordingly, so that the GPT model can improve the output accuracy while improving the output efficiency.

[0088] In another implementation, the keyword set can also be obtained in the following manner in this embodiment:

[0089] First, obtain the input text of the GPT model. For example, the user evaluation text that needs to be processed by the GPT model is "The installation master is really too hard alone. Good review."

[0090] Secondly, parse the input text to obtain the running scenario of the GPT model. Specifically, in this embodiment, a machine learning model can be pre-trained to process the text and obtain corresponding running scenarios, such as a theme analysis scenario, a shopping consultation scenario, etc. Taking the input text as a user evaluation text as an example, the running scenario of the GPT model is identified as a "theme analysis scenario".

[0091] After that, according to the running scenario, obtain multiple candidate keywords, and the meanings of the candidate keywords are related to the running scenario. For example, taking the "theme analysis scenario" as an example, keywords such as "start", "end", "theme", and "theme support point" that are semantically related to the theme analysis scenario are determined as candidate keywords.

[0092] Finally, add these candidate keywords to the keyword set.

[0093] In another implementation manner, in this embodiment, the keyword set can also be obtained in the following way:

[0094] First, obtain the keyword configuration information for the GPT model. These keyword configuration information can be obtained by the user through the configuration interface provided by the electronic device for configuration operations. For example, the user sequentially sets keyword configuration information such as "start", "end", "theme", and "theme support point" on the configuration interface.

[0095] After that, extract the keywords in the keyword configuration information to obtain multiple candidate keywords. For example, extract keywords such as "start", "end", "theme", and "theme support point" in the keyword configuration information configured by the user on the configuration interface as candidate keywords.

[0096] Finally, add these candidate keywords to the keyword set.

[0097] In one implementation manner, in step 102, if no target keyword matching the first character is found in the keyword set, then the method in this embodiment may further include the following steps, as Figure 5 shown in

[0098] Step 105: Determine the first character as the output character of the GPT model, so that the GPT model outputs the next character at least based on the first character.

[0099] That is to say, in this embodiment, after obtaining an output character of the GPT model, first search in the keyword set to find a target keyword that matches this output character. If there is a match, then the other characters in the target keyword except this output character can be directly used as the output character of the GPT model. At this time, the GPT model does not need to execute the inference process. If no keyword that matches this output character is found in the keyword set, then the GPT model still needs to execute the inference process based on this output character to infer the next character. It can be seen that in this embodiment, the keyword set can reduce the number of runs of the GPT model to improve the output efficiency. However, if no matching target keyword is found in the keyword set, the GPT model can also infer the next character through running, thereby ensuring the reliability of the GPT model.

[0100] Reference Figure 6 , which is a schematic structural diagram of a running control device for a generative pre-trained GPT model provided in the second embodiment of the present application. This device can be configured in an electronic device capable of running the GPT model, such as a computer or a server, etc. The technical solution in this embodiment is mainly used to improve the output efficiency of the GPT model.

[0101] Specifically, the device in this embodiment may include the following units:

[0102] A character acquisition unit 601, configured to acquire a first character output by the GPT model;

[0103] A keyword search unit 602, configured to search in a preset keyword set for a target keyword that matches the first character; the keyword set includes multiple candidate keywords, and the candidate keywords include multiple candidate characters;

[0104] A character extraction unit 603, configured to obtain at least one second character at least according to the candidate characters in the target keyword;

[0105] A character determination unit 604, configured to determine the at least one second character as the output character of the GPT model, so that the GPT model outputs the next character at least according to the second character.

[0106] As can be seen from the above solution, in a running control device for a generative pre-trained GPT model provided in the second embodiment of the present application, after obtaining the first character output by the GPT model and before the GPT model runs next time, first search for a target keyword matching the first character in the keyword set, and then use the second character in the target keyword as the output character of the GPT model. In this way, the output character can be obtained without the GPT model running again. After that, the GPT model can use the second character as the output character of the previous time to perform inference and prediction on the next character. It can be seen that in this embodiment, the output character of the GPT model is predicted by configuring the keyword set, without having to let the GPT model run one more time, thereby reducing the time cost by reducing the number of runs of the GPT model, and thus improving the output efficiency of the GPT model.

[0107] In one implementation, the candidate keywords in the keyword set are related to at least one of the following:

[0108] The running scenario of the GPT model, where the running scenario is determined based on the input text of the GPT model;

[0109] The output parameters of the GPT model, where the output parameters at least characterize the text output format of the GPT model.

[0110] In one implementation, the keyword set is obtained in the following way:

[0111] Obtain the text output format of the GPT model; extract the format description words in the text output format; add the format description words as candidate keywords to the keyword set.

[0112] In one implementation, the keyword set is obtained in the following way:

[0113] Obtain the input text of the GPT model; parse the input text to obtain the running scenario of the GPT model; according to the running scenario, obtain multiple candidate keywords, where the meanings of the candidate keywords are related to the running scenario; add the candidate keywords to the keyword set.

[0114] In one implementation, the keyword set is obtained in the following way:

[0115] Obtain the keyword configuration information for the GPT model; extract the keywords in the keyword configuration information to obtain multiple candidate keywords; add the candidate keywords to the keyword set.

[0116] In one implementation, the target keyword matching the first character includes: the first character in the target keyword is consistent with the first character.

[0117] In one implementation, the character extraction unit 603 is specifically configured to: extract other candidate characters except the first character from the target keyword, and the other candidate characters are the second characters.

[0118] In one implementation, the character determination unit 604 is further configured to: when the keyword search unit 602 fails to find a target keyword matching the first character in the keyword set, determine the first character as the output character of the GPT model, so that the GPT model outputs the next character at least based on the first character.

[0119] It should be noted that the specific implementation of each unit in this embodiment can refer to the corresponding content in the foregoing text, and will not be elaborated here.

[0120] In addition, the embodiments of the present application also claim to protect a computer-readable storage medium, where the computer-readable storage medium includes a stored program, and when the program runs, it executes the running control method of the generative pre-trained GPT model as described in any of the foregoing embodiments.

[0121] Reference Figure 7 , which is a schematic structural diagram of an electronic device provided in Embodiment 3 of the present application. The electronic device may include the following structure:

[0122] A memory 701 for storing a computer program and data generated by the running of the computer program;

[0123] A processor 702 is configured to implement, through a computer program: obtain a first character output by the GPT model; search for a target keyword matching the first character in a keyword set; the keyword set includes multiple candidate keywords, and the candidate keywords include multiple candidate characters; obtain at least one second character at least based on the candidate characters in the target keyword; determine the second character as the output character of the GPT model, so that the GPT model outputs the next character at least based on the second character.

[0124] As can be seen from the above solution, in an electronic device provided in the third embodiment of the present application, after obtaining the first character output by the GPT model and before the GPT model performs the next run, first search for a target keyword that matches the first character in the keyword set, and then use the second character in the target keyword as the output character of the GPT model. In this way, the output character can be obtained without the GPT model running again. After that, the GPT model can use the second character as the output character of the previous time to perform the inference and prediction of the next character. It can be seen that in this embodiment, the output character of the GPT model is predicted by configuring the keyword set, without having to let the GPT model run one more time, thereby reducing the time cost by reducing the number of runs of the GPT model, and thus improving the output efficiency of the GPT model.

[0125] Taking the scenario of theme analysis based on user evaluation text as an example, in order to improve the output efficiency of the GPT model, the present application proposes, for scenarios with a fixed output format that require formatted output data: combining the output template words to reduce or avoid the GPT model from outputting known format content, and efficiently reducing the time cost of GPT speculation and generation. The specific solution is as follows:

[0126] As Figure 8 shown in, the GPT model is an intelligent assistant for theme parsing, key point extraction, and sentiment discrimination of user comment content. In each iteration of the GPT model, after generating a character in the output logit layer and the origin output layer between the hidden layer and the output layer, add a process of rule checker trigger judgment in the output layer. If the rule judgment is TRUE, perform template filling to directly obtain the subsequent characters without the GPT executing the processing process.

[0127] Among them, the user evaluation text is as follows:

[0128] "The installation master is really too hard alone. Good review. The washing machine is grand and good-looking. Although it has not been installed yet. Wrong, the washing machine has not been installed yet. It was the delivery master who delivered it alone. Although there is an elevator, it is really hard. Give a thumbs up! __ It has been installed in time, the service is timely, and everything is fine __ The quality is good, the price is low, and the cost performance is high. The clothes are washed very clean"

[0129] The example of the text format required for the GPT model to output is as follows:

[0130] {

[0131] "Start: Installation master; End: Good review": "Theme: Installation service -> Positive; Theme support point: None",

[0132] "Start: Washing machine; End: Grand and good-looking": "Theme: Appearance -> Positive; Theme support point: Grand and good-looking",

[0133] "Start: Although still; End: Give a like": "Subject: Delivery problem -> Positive; Subject support point: None",

[0134] "Start: In time; End: Everything is fine": "Subject: Installation service -> Positive; Subject support point: Timely installation; Service is timely",

[0135] "Start: Good quality; End: Very clean": "Subject: Product quality -> Positive; Subject support point: Clothes are washed very cleanly",

[0136] "Start: Low price; End: High cost performance": "Subject: Price fluctuation -> Positive; Subject support point: High cost performance; Low price"

[0137] }

[0138] According to the output example, words such as "Start:", "End:", "Subject:", "Subject support point:", etc. can be used as supplementary content for the rule engine, that is, the keyword set in the previous text. There is a lot of repetitive content. It can be seen that by applying this application, the efficiency improvement effect will be very obvious.

[0139] Among them, the native GPT will generate "main" at w n+1 and then iteratively generate "topic", "support", "point", ":". This application will generate the 6 words "Subject support point:" at w n+1 , saving the computational resource cost and time cost of the native GPT model for generating "topic", "support", "point", ":", which reflects the advantage of this application.

[0140] The construction time cost of the rule engine can be ignored, and the content of the rule engine can be adaptively modified according to different output formats, which reflects the flexibility and expandability of this method.

[0141] Moreover, the pre-trained model architecture in this application remains unchanged, only the processing flow of the GPT model in the inference stage is changed, and the time amount and the number of parameters during training will not be changed.

[0142] It can be seen that if the technical solution of this application is not adopted, the execution process of the GPT model when outputting characters is as shown in Figure 9 . For the GPT model to output one character, it needs to execute a processing flow once; if the technical solution of this application is adopted, after the GPT model outputs the previous character, it first judges according to the rule engine trigger. When the corresponding keyword is matched in the rule engine, that is, when it is TRUE, the characters in the keyword can be used as the output characters, and there is no need for the GPT model to execute the processing flow again. When the corresponding keyword is not matched in the rule engine, that is, when it is FALSE, the GPT model needs to execute the processing flow to infer the next character, as shown in Figure 10As shown in []. Thus, compared with the need to run 7 times to generate demo text for the GPT model, the GPT model based on the rule engine in the technical solution of this application only needs to run 3 times, greatly reducing the computing resources and time consumption.

[0143] In summary, the technical solution of this application can efficiently improve the GPT generation efficiency and reduce the consumption of model inference computing resources.

[0144] The above are only the preferred embodiments of this application. It should be noted that for those of ordinary skill in the art, without departing from the principle of this application, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of this application.

Claims

1. A method for controlling the operation of a generative pre-trained GPT model, characterized in that, The method includes: Obtaining a first character output by the GPT model; Searching in a preset keyword set for a target keyword that matches the first character; the keyword set contains multiple candidate keywords, and the candidate keywords contain multiple candidate characters; Obtaining at least one second character based at least on the candidate characters in the target keyword; Determining the at least one second character as the output character of the GPT model, so that the GPT model outputs the next character based at least on the second character.

2. The method according to claim 1, characterized in that, The candidate keywords in the keyword set are related to at least one of the following: The running scenario of the GPT model, which is determined based on the input text of the GPT model; The output parameters of the GPT model, and the output parameters at least characterize the text output format of the GPT model.

3. The method according to claim 1 or 2, characterized in that, The keyword set is obtained by the following method: Obtaining the text output format of the GPT model; Extracting format description words from the text output format; Adding the format description words as candidate keywords to the keyword set.

4. The method according to claim 1 or 2, characterized in that, The keyword set is obtained by the following method: Obtaining the input text of the GPT model; Parsing the input text to obtain the running scenario of the GPT model; Obtaining multiple candidate keywords according to the running scenario, and the meanings of the candidate keywords are related to the running scenario; Adding the candidate keywords to the keyword set.

5. The method according to claim 1 or 2, characterized in that, The keyword set is obtained by the following method: Obtaining keyword configuration information for the GPT model; Extracting keywords from the keyword configuration information to obtain multiple candidate keywords; Adding the candidate keywords to the keyword set.

6. The method according to claim 1 or 2, characterized in that, The target keyword matching the first character includes: The first character in the target keyword is consistent with the first character.

7. The method according to claim 6, characterized in that, The obtaining at least one second character based at least on the candidate characters in the target keyword includes: Extracting the other candidate characters in the target keyword except the first character, and the other candidate characters are the second characters.

8. The method according to claim 1 or 2, characterized in that, In the case where no target keyword matching the first character is found in the keyword set, the method further includes: Determining the first character as the output character of the GPT model, so that the GPT model outputs the next character based at least on the first character.

9. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a stored program, wherein the program, when running, executes the method according to any one of claims 1 to 8.

10. An electronic device, comprising a memory and a processor, characterized in that, A computer program is stored in the memory, and the processor is configured to execute the method according to any one of claims 1 to 8 through the computer program.

Citation Information

Cited By

  • Space-time backtracking personal simulation image generation method and system based on real person reality

    CN120852602A