Method, device and storage medium for reducing a set of prompt words

By optimizing the initial prompt word set through iterative processing and scenario constraints, the problem of low prompt word quality in the field of information security is solved, the quality of synthetic data is improved, and the workload of operations personnel is reduced.

CN119721013BActive Publication Date: 2025-11-18CHINA TELECOM CORP LTD TECHNOLOGY INNOVATION CENTER +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411542679.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-10-31
Publication Date
2025-11-18
Estimated Expiration
2044-10-31

AI Technical Summary

Technical Problem

The quality of prompt words used in the current information security field is not high, resulting in poor quality of synthetic data after processing by large language models, and operators need to spend time and effort processing logs and alarms.

Method used

The initial set of prompt words is processed iteratively, and the synthesized data is scored using a preset algorithm. Combined with scenario constraints, the set of prompt words is simplified to ensure that it matches the application scenario.

Benefits of technology

The quality of the prompt words has been improved to better suit the current application scenarios, thus enhancing the quality of the synthesized data and reducing the workload of operations personnel.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119721013B_ABST
    Figure CN119721013B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of data processing, and discloses a method and device for simplifying a prompt word set and a storage medium. The method is applied to a server and comprises the following steps: inputting a pre-set initial prompt word set into a large language model for learning; in each iteration process, determining whether the number of initial input data changes according to a first score value and / or a second score value of the current round; each iteration process is as follows: adopting a pre-set first algorithm to score each first synthetic data to obtain the first score value; and / or, adopting a pre-set second algorithm to score each second synthetic data to obtain the second score value, wherein the second synthetic data is obtained by learning each initial input data included in the initial prompt word set by using the large language model after adding a scene constraint condition; the above method improves the quality of the prompt word, makes the prompt word more suitable for the current application scene, and further improves the quality of the generated synthetic data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of data processing technology, and provides a method, apparatus and storage medium for simplifying a set of prompt words. Background Technology

[0002] Large language models, with their powerful language understanding and knowledge emergence capabilities, have begun to empower many fields, and the information security field has also started to study the benefits they can bring. Taking network operations as an example, the current mainstream approach is to filter out a large amount of log information through various rule engines and threat intelligence databases. For those log alerts that cannot be filtered out, alert tickets are generated through aggregation and correlation rules, and then the operations personnel make judgments, which is time-consuming and labor-intensive.

[0003] Furthermore, the quality of the prompt words used in the field of information security is not high. For example, the domain to which the data in the prompt words belongs is poorly related to the security domain, the data itself does not conform to grammatical rules, and the quality is poor, etc., which leads to the poor quality of the synthetic data after the above prompt words are processed by the large language model. Summary of the Invention

[0004] This application provides a method, apparatus, and storage medium for streamlining a set of prompt words, in order to improve the quality of each initial input data in the prompt words and make the prompt words more suitable for the current application scenario.

[0005] The specific technical solution provided in this application is as follows:

[0006] In a first aspect, embodiments of this application provide a method for simplifying a set of prompt words, applied to a server, the method comprising:

[0007] The pre-set initial prompt word set is input into the large language model for learning, and the multiple initial input data included in the initial prompt word set are processed in an iterative manner;

[0008] During each iteration, the number of initial input data points is determined based on the first and / or second scores of the current round. The iteration process stops when the number of input data points no longer changes. The iteration process for each round is as follows:

[0009] A preset first algorithm is used to score each piece of the first synthesized data to obtain a first score value. The first synthesized data is obtained by the large language model after learning from each initial input data included in the initial prompt word set; and / or

[0010] The second algorithm is used to score each second synthetic data to obtain a second score value. The second synthetic data is obtained by learning each initial input data included in the initial prompt word set using a large language model with added scene constraints.

[0011] Optionally, the first synthesized data is determined in the following manner:

[0012] A pre-set set of initial prompt words is input into a large language model for learning, resulting in multiple first output data. The initial prompt word set includes one initial input data that corresponds to at least one first output data.

[0013] The first synthetic data is determined based on the initial input data and at least one corresponding first output data.

[0014] Optionally, the second synthetic data is determined in the following manner:

[0015] Add scene constraints to the large language model, where the scene constraints are different for different rounds of iteration;

[0016] Each initial input data included in the initial prompt word set is input into the large language model after adding scene constraints for learning, and multiple second output data are obtained. Among them, one initial input data included in the initial prompt word set corresponds to at least one second output data.

[0017] The second synthetic data is determined based on the initial input data and at least one corresponding second output data.

[0018] Optionally, the first algorithm includes a first specialization sub-algorithm, a first syntax sub-algorithm, and a first anomaly detection sub-algorithm. A preset first algorithm is used to score each of the first synthetic data points to obtain a first score value, including:

[0019] Each piece of the first synthetic data is scored using a pre-defined first professionalism sub-algorithm to obtain a professionalism score;

[0020] Each first synthetic data is scored using a pre-defined first grammar sub-algorithm to obtain a grammar score;

[0021] The first anomaly detection sub-algorithm is used to score each of the first synthetic data to obtain anomaly detection scores;

[0022] The professionalism score, grammar score, and anomaly detection score are added together to obtain the first score.

[0023] Optionally, a preset second algorithm is used to score each of the second synthetic data to obtain a second score value, including:

[0024] The quality of each second synthetic data is scored using a pre-set second algorithm to obtain a second score value.

[0025] Optionally, before scoring each of the second synthetic data using a preset second algorithm, the method further includes:

[0026] Determine the relevant values ​​of each initial input data and scene constraints included in the initial prompt word set;

[0027] Initial input data with a relevance value less than a preset relevance threshold will be removed from the initial prompt word set.

[0028] Optionally, determine whether the number of initial input data has changed based on the first score and / or second score of this round, including:

[0029] Compare the first score of this round with the preset first standard score of this round;

[0030] If the first score is less than the first standard score, the initial input data that is below the first threshold will be removed from the initial prompt word set;

[0031] Compare the second score of this round with the preset second standard score of this round;

[0032] If the second score is less than the second standard score, the initial input data that is below the second threshold will be removed from the initial prompt word set;

[0033] If the first score is not less than the first standard score and / or the second score is not less than the second standard score, then the number of initial input data in the initial prompt word set corresponding to this round will no longer change.

[0034] Secondly, embodiments of this application also provide an apparatus for simplifying a set of prompt words, comprising:

[0035] The input unit is used to input a pre-set set of initial prompt words into the large language model for learning, and processes multiple initial input data included in the initial prompt word set in an iterative manner;

[0036] An iterative unit is used in each iteration to determine whether the number of initial input data has changed based on the first score and / or the second score of the current iteration, and stops the iteration process when the number no longer changes; the iteration process in each round is as follows:

[0037] The first processing unit is configured to score each piece of first synthetic data using a preset first algorithm to obtain a first score value, wherein the first synthetic data is obtained by the large language model after learning from each initial input data included in the initial prompt word set; and / or

[0038] The second processing unit is used to score each of the second synthetic data using a preset second algorithm to obtain a second score value. The second synthetic data is obtained by learning each of the initial input data included in the initial prompt word set using a large language model with added scene constraints.

[0039] Optionally, the first synthesized data is determined in the following manner:

[0040] A pre-set set of initial prompt words is input into a large language model for learning, resulting in multiple first output data. The initial prompt word set includes one initial input data that corresponds to at least one first output data.

[0041] The first synthetic data is determined based on the initial input data and at least one corresponding first output data.

[0042] Optionally, the second synthetic data is determined in the following manner:

[0043] Add scene constraints to the large language model, where the scene constraints are different for different rounds of iteration;

[0044] Each initial input data included in the initial prompt word set is input into the large language model after adding scene constraints for learning, and multiple second output data are obtained. Among them, one initial input data included in the initial prompt word set corresponds to at least one second output data.

[0045] The second synthetic data is determined based on the initial input data and at least one corresponding second output data.

[0046] Optionally, the first algorithm includes a first specialization sub-algorithm, a first syntax sub-algorithm, and a first anomaly detection sub-algorithm. A preset first algorithm is used to score each of the first synthetic data to obtain a first score value. The first processing unit is used for:

[0047] Each piece of the first synthetic data is scored using a pre-defined first professionalism sub-algorithm to obtain a professionalism score;

[0048] Each first synthetic data is scored using a pre-defined first grammar sub-algorithm to obtain a grammar score;

[0049] The first anomaly detection sub-algorithm is used to score each of the first synthetic data to obtain anomaly detection scores;

[0050] The professionalism score, grammar score, and anomaly detection score are added together to obtain the first score.

[0051] Optionally, a preset second algorithm is used to score each of the second synthetic data to obtain a second score value. The second processing unit is used for:

[0052] The quality of each second synthetic data is scored using a pre-set second algorithm to obtain a second score value.

[0053] Optionally, before scoring each of the second synthetic data using a preset second algorithm, the method further includes:

[0054] Determine the relevant values ​​of each initial input data and scene constraints included in the initial prompt word set;

[0055] Initial input data with a relevance value less than a preset relevance threshold will be removed from the initial prompt word set.

[0056] Optionally, the number of initial input data points is determined based on the first and / or second score values ​​of this round. The iterative unit is used to:

[0057] Compare the first score of this round with the preset first standard score of this round;

[0058] If the first score is less than the first standard score, the initial input data that is below the first threshold will be removed from the initial prompt word set;

[0059] Compare the second score of this round with the preset second standard score of this round;

[0060] If the second score is less than the second standard score, the initial input data that is below the second threshold will be removed from the initial prompt word set;

[0061] If the first score is not less than the first standard score and / or the second score is not less than the second standard score, then the number of initial input data in the initial prompt word set corresponding to this round will no longer change.

[0062] Thirdly, a server includes:

[0063] Memory, used to store executable instructions;

[0064] A processor for reading and executing executable instructions stored in memory to implement the method as described in any of the first aspects.

[0065] Fourthly, a computer-readable storage medium, when instructions in the storage medium are executed by a processor, enables the processor to perform the method described in any of the first aspects above.

[0066] The beneficial effects of this application are as follows:

[0067] In summary, this application provides a method, apparatus, and storage medium for simplifying a prompt word set. The method is applied to a server and includes: inputting a pre-set initial prompt word set into a large language model for learning; processing multiple initial input data included in the initial prompt word set iteratively; determining whether the number of initial input data has changed based on the first score and / or second score of each iteration; and stopping the iteration process when the number no longer changes. Each iteration process is as follows: scoring each first synthesized data using a pre-set first algorithm to obtain a first score value, wherein the first synthesized data is a large language model. The model learns from each initial input data in the initial prompt word set; and / or, a preset second algorithm is used to score each second synthesized data to obtain a second score value. The second synthesized data is obtained by learning from each initial input data in the initial prompt word set using a large language model with added scene constraints. The above-mentioned methods of simplifying each initial input data using the first score value and / or the second score value, and learning from each initial input data using a large language model with added scene constraints, effectively improve the quality of the prompt words, making the prompt words more in line with the current application scenario, thereby improving the quality of the generated synthesized data.

[0068] Other features and advantages of this application will be set forth in the description which follows, and will be apparent in part from the description, or may be learned by practicing the application. The objectives and other advantages of this application may be realized and obtained by means of the structures particularly pointed out in the written description, claims, and drawings. Attached Figure Description

[0069] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings:

[0070] Figure 1 This is a schematic diagram of the system architecture for simplifying prompt words in an embodiment of this application;

[0071] Figure 2 This is a flowchart illustrating the process of simplifying prompts in an embodiment of this application;

[0072] Figure 3 This is a schematic diagram of a process for determining the first synthetic data in an embodiment of this application;

[0073] Figure 4 This is a schematic diagram of a process for determining the second synthetic data in an embodiment of this application;

[0074] Figure 5This is a schematic diagram of a process for obtaining a first score value using a first algorithm in an embodiment of this application;

[0075] Figure 6 This is a flowchart illustrating a method for simplifying the set of prompt words in an embodiment of this application.

[0076] Figure 7 This is a schematic diagram of the logical architecture of a device for simplifying prompts according to an embodiment of this application;

[0077] Figure 8 This is a schematic diagram of the physical architecture of a server according to an embodiment of this application. Detailed Implementation

[0078] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of this application will be clearly and completely described below with reference to the accompanying drawings of the embodiments of this application. Obviously, the described embodiments are only some embodiments of the technical solutions of this application, and not all embodiments. Based on the embodiments recorded in this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the technical solutions of this application.

[0079] The terms "first," "second," etc., used in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that embodiments of the invention described herein can be implemented in sequences other than those illustrated or described herein.

[0080] The following is an explanation of the technical terms used in this application.

[0081] Large Language Models (LLMs) are artificial intelligence models designed to understand and generate human language. They are trained on massive amounts of text data and can perform a wide range of tasks, including text summarization, translation, sentiment analysis, and more. LLMs are characterized by their massive scale, containing billions of parameters that help them learn complex patterns in language data.

[0082] Synthetic data is a type of non-human-created data that mimics real-world data. It is generated through computational algorithms and simulations based on generative artificial intelligence techniques. Synthetic prompts possess the same mathematical properties as the actual data they are based on, but do not contain the same information. In the fields of machine learning and artificial intelligence, synthetic data can provide training material for models, helping them learn, understand, and predict.

[0083] The preferred embodiments of this application will now be described in detail with reference to the accompanying drawings.

[0084] See Figure 1 As shown in the embodiments of this application, the system includes at least one server. Figure 1 In this process, the initial prompt word set 1, initial prompt word set 2, and initial prompt word set n can be input into the server for simplification, thereby obtaining the simplified initial prompt word set 1, initial prompt word set 2, and initial prompt word set n. The aforementioned initial prompt word sets usually include multiple prompt texts input into the large language model, but the correlation between the prompt texts and the large language model is uncertain. After simplifying the aforementioned initial prompt word sets, the final prompt texts that are more suitable for the large language model are obtained.

[0085] In this embodiment of the application, the implementation of a method for simplifying a set of prompt words is mainly executed on the server side. The following describes the process of simplifying an initial set of prompt words in detail.

[0086] See Figure 2 As shown in the embodiments of this application, the specific process of a simplified prompt word set is as follows:

[0087] Step 201: Input the pre-set initial prompt word set into the large language model for learning, and process the multiple initial input data included in the initial prompt word set in an iterative manner.

[0088] First, it should be noted that the initial prompt word set includes multiple initial input data, which are usually pre-set manually. However, the relevance of each initial input data to the application scenario in which the large language model is used is uncertain. Therefore, in this embodiment, the large language model is used to simplify the initial input data included in the initial prompt word set, thereby resulting in higher quality synthetic data obtained after processing the simplified initial input data through the large language model.

[0089] During implementation, an iterative approach is used to simplify the multiple initial input data included in the initial prompt word set.

[0090] Step 202: During each iteration, determine whether the number of initial input data has changed based on the first score and / or the second score of this round, and stop the iteration process when the number no longer changes; wherein, the iteration process of each round is as follows:

[0091] Considering the uncertainty of the number of initial input data in the initial prompt word set, and the uncertainty of the relevance of each initial input data to the application scenario used by the large language model, in each iteration, after calculating the first score and / or the second score, it is determined whether to reduce the number of initial input data in the initial prompt word set based on either the first score, the second score, or both. This means determining whether the number of initial input data in the initial prompt word set has changed. The iteration process stops when the number of initial input data no longer changes, resulting in a simplified initial prompt word set. The execution steps in each iteration are detailed below through steps 203 and 204.

[0092] Step 203: The first synthetic data is scored using a pre-defined first algorithm to obtain a first score value. The first synthetic data is obtained by the large language model after learning from the initial input data included in the initial prompt word set; and / or

[0093] First, it should be noted that the reference Figure 3 As shown, the first synthetic data is determined in the following manner:

[0094] Step 101: Input the pre-set initial prompt word set into the large language model for learning to obtain multiple first output data, wherein the initial prompt word set includes one initial input data corresponding to at least one first output data.

[0095] During implementation, the initial input data, including the pre-set set of initial prompt words, are fed into a large language model for learning. For each initial input data, the large language model can obtain at least one first output data. Thus, for each initial input data included in the initial prompt word set, multiple first output data can be obtained, and each initial input data corresponds to at least one first output data.

[0096] Step 102: Determine the first synthesized data based on the initial input data and at least one corresponding first output data.

[0097] After obtaining multiple first output data, each initial input data is combined with at least one corresponding first output data to determine a first composite data. It should be noted that when an initial input data and multiple first output data constitute a composite data, a set method can be used to combine an initial input data and multiple first output data.

[0098] It should be noted that, for reference Figure 4 As shown, the second synthetic data is determined in the following manner:

[0099] Step 101': Add scene constraints to the large language model, where the scene constraints are different for different rounds of iteration.

[0100] To achieve targeted optimization and rewriting of the initial prompt word set, different scenario constraints can be added to the large language model in different iterations, making each initial input data more aligned with the preset usage scenario. To achieve gradual optimization of the initial prompt word set, relatively broad scenario constraints can be added to the large language model initially, with stricter constraints gradually added in subsequent iterations.

[0101] Step 102': Input each initial input data included in the initial prompt word set into the large language model after adding scene constraints for learning, and obtain multiple second output data, wherein one initial input data included in the initial prompt word set corresponds to at least one second output data.

[0102] After adding scene constraints to the large language model, the initial input data from the initial prompt word set are fed into the large language model with added scene constraints for learning. For each initial input data, the large language model can obtain at least one second output data. In this way, multiple second output data can be obtained for each initial input data from the initial prompt word set, and each initial input data corresponds to at least one second output data.

[0103] It should be further explained that the initial prompt word set input into the large language model after adding scene constraints can be a pre-set initial prompt word set, or an initial prompt word set after simplifying each initial input data with a first score value, or an initial prompt word set after simplifying each initial input data with a second score value, or an initial prompt word set after simplifying each initial input data with the first score value and the second score value for N rounds of iteration. That is, in one embodiment, the second synthetic data can be obtained directly from the pre-set initial prompt word set, and each initial input data can be simplified according to the second score value; in another embodiment, the first synthetic data is first obtained from the pre-set initial prompt word set, and each initial input data is initially simplified according to the first score value obtained by scoring the first synthetic data. Based on this, the initially simplified initial prompt word set is input into the large language model after adding scene constraints; in a third embodiment, the initial prompt word set after simplifying each initial input data with the second score value for at least one round, or the initial prompt word set after simplifying each initial input data with the first score value and the second score value for N rounds of iteration, is then input into the large language model after adding scene constraints.

[0104] Step 103': Determine the second synthetic data based on the initial input data and at least one corresponding second output data.

[0105] After obtaining multiple second output data, each initial input data is combined with at least one corresponding second output data to determine a second composite data. It should be noted that when an initial input data and multiple second output data constitute a composite data, a set method can be used to combine an initial input data and multiple second output data.

[0106] During implementation, after obtaining multiple first synthetic data sets, a first algorithm is used to score each first synthetic data set. This first algorithm includes a first professionalism sub-algorithm, a first syntax sub-algorithm, and a first anomaly detection sub-algorithm. The preset first algorithm is used to score each first synthetic data set to obtain a first score value. (See reference...) Figure 5 As shown, it includes:

[0107] Step 2031: Use the preset first professionalism sub-algorithm to score each of the first synthetic data to obtain a professionalism score.

[0108] It should be noted that the aforementioned first-level expertise sub-algorithm can be pre-set and integrated into the large language model. During implementation, the first-level expertise sub-algorithm can be implemented by scoring the historical first-synthetic data based on prior professional knowledge. Implementation methods include, but are not limited to, scoring the historical first-synthetic data based on prior professional knowledge to obtain different levels of scores, and combining the weights corresponding to each level to determine the first-level expertise sub-algorithm. For example, when scoring the historical first-synthetic data based on prior professional knowledge to obtain three levels of scores, and setting the weights of the scores from highest to lowest to 100%, 50%, and 0 respectively, the total score corresponding to the first-level expertise sub-algorithm can be expressed by the formula: Where N1 represents the number of initial input data corresponding to the high score level, N2 represents the number of initial input data corresponding to the medium score level, and N3 represents the number of initial input data corresponding to the low score level.

[0109] During implementation, after determining the first professionalism sub-algorithm, the first professionalism sub-algorithm is used to score each first synthetic data, and then the scored data are weighted or weighted averaged to obtain the professionalism score.

[0110] Step 2032: Use the preset first grammar sub-algorithm to score each first synthetic data to obtain a grammar score.

[0111] It should be noted that the aforementioned first grammar sub-algorithm can be pre-set and integrated into the large language model. During implementation, it can be achieved by scoring the historical first synthetic data based on grammatical knowledge. The aforementioned grammar can be the rules set in various fields or application scenarios. For example, when the large language model is applied to the field of information security, the aforementioned grammar is the regulation related to the field of information security, and the first grammar sub-algorithm is constructed based on this regulation.

[0112] After determining the first sub-algorithm, the first sub-algorithm is used to score each first synthetic data, and then the scored data are weighted or weighted averaged to obtain the grammar score.

[0113] Step 2033: Use the preset first anomaly detection sub-algorithm to score each of the first synthetic data to obtain anomaly detection scores.

[0114] It should be noted that the aforementioned first anomaly detection sub-algorithm can also be pre-set and integrated into the large language model. During implementation, it is inevitable that some anomalies will occur in the first synthesized data. Common anomalies include: long-tail generation, empty replies, abnormal characters, and whether instructions are followed. In the industry, the first synthesized data with anomalies is usually called anomaly samples, and the first synthesized data without anomalies is called normal samples. Thus, the larger the ratio of normal samples to the total number of samples, the fewer the anomaly samples in this set of first synthesized data, and the lower the corresponding anomaly detection score.

[0115] After determining the first anomaly detection sub-algorithm, the first anomaly detection sub-algorithm is used to score each of the first synthetic data, and then the scored data are weighted or weighted averaged to obtain the anomaly detection score.

[0116] Step 2034: Add the professionalism score, grammar score, and anomaly detection score together to obtain the first score.

[0117] To comprehensively evaluate the quality of the first synthesized data, the professionalism score, grammar score, and anomaly detection score obtained above are summed during the implementation process, and the total score is the first score value. It should be noted that the average of the total scores can also be calculated to obtain the first score value. Each of these first score values ​​corresponds one-to-one with a first synthesized data point, and multiple first score values ​​can be obtained by scoring each first synthesized data point separately. Furthermore, multiple first score values ​​can be obtained in a single iteration.

[0118] The following section continues with the process of determining the second score. First, it should be noted that before scoring each piece of the second composite data using the preset second algorithm, the following steps are also included:

[0119] (1) Determine the relevant values ​​of each initial input data and scene constraints included in the initial prompt word set.

[0120] Considering that scenario constraints are used to limit the correlation between the input and output of a large language model, if the initial input data has a poor correlation with the scenario constraints, for example, when the scenario constraints are related to the field of information security, but the initial input data is vegetable planting, the correlation between the field of information security and vegetable planting is poor. Even if the initial input data is input into the large language model, it will not be possible to obtain results related to information security.

[0121] Based on this, during the implementation process, after adding scene constraints to the large language model, the correlation values ​​between each initial input data included in the initial prompt word set and the scene constraints are calculated separately.

[0122] (2) Remove initial input data with a relevance value less than the preset relevance threshold from the initial prompt word set.

[0123] After obtaining multiple correlation values, each correlation value is compared with a preset correlation threshold, which represents the minimum correlation value with the added scene constraints. When the comparison finds that the correlation value is less than the preset correlation threshold, it indicates that the initial input data corresponding to the correlation value has a poor correlation with the scene constraints. The initial input data corresponding to that correlation value is then deleted from the initial prompt word set, thereby simplifying the initial prompt word set.

[0124] Step 204: Use the preset second algorithm to score each of the second synthetic data to obtain the second score value. The second synthetic data is obtained by learning each of the initial input data included in the initial prompt word set using a large language model with added scene constraints.

[0125] During implementation, after inputting the initial input data from the initial prompt word set into a large language model with added scene constraints for learning, multiple second synthesized data are obtained. A pre-defined second algorithm is used to score each second synthesized data, yielding a second score value, including:

[0126] The quality of each second synthetic data is scored using a pre-set second algorithm to obtain a second score value.

[0127] It should be noted that the aforementioned second algorithm can be pre-configured and integrated into the large language model. During implementation, the second algorithm is used to score the quality of each second synthesized data. The quality constraints of the second synthesized data can be specifically set according to the scenario constraints. For example, the second algorithm may include scoring the quality of the second synthesized data based on grammar, scenario relevance, etc., so that a corresponding second score value can be obtained for each second synthesized data, i.e., multiple second score values ​​can be obtained.

[0128] The determination of whether the number of initial input data has changed based on the first score and / or second score of this round includes:

[0129] First, it should be noted that in this embodiment, each iteration process can be divided into three cases: First, only the first synthesized data is scored to obtain a first score value; second, only the second synthesized data is scored to obtain a second score value; third, after scoring the first synthesized data to obtain a first score value, the second synthesized data is then scored to obtain a second score value. Correspondingly, during the iteration process, the number of initial input data can be determined by referring only to the first score value; the number of initial input data can be determined by referring only to the second score value; or the number of initial input data can be determined by referring to both the first and second score values ​​simultaneously.

[0130] Preferably, during implementation, both the first score and the second score are referenced to determine whether the number of initial input data has changed. (See [reference]) Figure 6 As shown:

[0131] Step 2041: Compare the first score of this round with the preset first standard score of this round.

[0132] During implementation, after scoring the first synthesized data and obtaining a first score, this first score can be used to determine whether the number of initial input data has changed. Specifically, the first score can be compared with a preset first standard score. It should be noted that the aforementioned first standard score is usually the lowest score that the initial input data must meet the requirements, based on pre-set criteria such as professionalism score, grammar score, and anomaly detection score. That is, only initial input data with a score higher than or equal to the first standard score can be retained in the initial prompt word set; initial input data with a score lower than the first standard score must be deleted from the initial prompt word set.

[0133] Step 2042: If the first score is less than the first standard score, then the initial input data that is below the first threshold is deleted from the initial prompt word set.

[0134] During implementation, if the first score of this round is less than the preset first standard score of this round, it means that the initial input data corresponding to the first score does not meet the requirements. In this case, the initial input data will be deleted from the initial prompt word set, thereby simplifying the initial input data in the initial prompt word set.

[0135] Step 2043: Compare the second score of this round with the preset second standard score of this round.

[0136] During implementation, after scoring the second synthesized data and obtaining a second score value in this iteration, the second score value can be used to determine whether the number of initial input data has changed. Specifically, the second score value can be compared with a preset second standard score value. It should be noted that the aforementioned second standard score value is usually the lowest score value set after the quality scoring of the second synthesized data, which is the minimum score value that the initial input data must meet the requirements. That is, only initial input data with a score higher than or equal to the second standard score value can be retained in the initial prompt word set, and initial input data with a score lower than the second standard score value must be deleted from the initial prompt word set.

[0137] Step 2044: If the second score is less than the second standard score, then the initial input data that is below the second threshold will be removed from the initial prompt word set.

[0138] During implementation, if the second score value of this round is less than the preset second standard score value of this round, it means that the initial input data corresponding to the second score value does not meet the requirements. In this case, the initial input data will be deleted from the initial prompt word set, so that the initial input data in the initial prompt word set is more relevant to the added scene constraints.

[0139] Step 2045: If the first score is not less than the first standard score and / or the second score is not less than the second standard score, then determine that the number of initial input data in the initial prompt word set corresponding to this round will no longer change.

[0140] In one embodiment, if it is determined after comparison that the first score is not less than the first standard score, then the initial input data corresponding to the first score does not need to be deleted from the initial prompt word set, that is, it is determined that the number of initial input data in the initial prompt word set corresponding to this round will no longer change.

[0141] In another embodiment, if it is determined after comparison that the second score is not less than the second standard score, then the initial input data corresponding to the second score does not need to be deleted from the initial prompt word set, that is, it is determined that the number of initial input data in the initial prompt word set corresponding to this round will no longer change.

[0142] In the third embodiment, if the comparison determines that the first score is not less than the first standard score, then the initial input data corresponding to the first score does not need to be deleted from the initial prompt word set. Based on this, if the comparison determines that the second score is not less than the second standard score, then the initial input data corresponding to the second score does not need to be deleted from the initial prompt word set. That is, it is determined that the number of initial input data in the initial prompt word set corresponding to this round will no longer change.

[0143] Based on the same inventive concept, see [reference] Figure 7 As shown in the embodiment of this application, an apparatus for simplifying a set of prompt words is provided, comprising:

[0144] The input unit 701 is used to input a pre-set set of initial prompt words into the large language model for learning, and to process the multiple initial input data included in the initial prompt word set in an iterative manner.

[0145] Iteration unit 702 is used to determine whether the number of initial input data has changed based on the first score value and / or the second score value of the current iteration during each iteration, and to stop the iteration process when the number no longer changes; wherein, the iteration process of each round is as follows:

[0146] The first processing unit 703 is configured to score each piece of first synthetic data using a preset first algorithm to obtain a first score value, wherein the first synthetic data is obtained by the large language model after learning from each initial input data included in the initial prompt word set; and / or

[0147] The second processing unit 704 is used to score each of the second synthetic data using a preset second algorithm to obtain a second score value. The second synthetic data is obtained by learning each of the initial input data included in the initial prompt word set using a large language model with added scene constraints.

[0148] Optionally, the first synthesized data is determined in the following manner:

[0149] A pre-set set of initial prompt words is input into a large language model for learning, resulting in multiple first output data. The initial prompt word set includes one initial input data that corresponds to at least one first output data.

[0150] The first synthetic data is determined based on the initial input data and at least one corresponding first output data.

[0151] Optionally, the second synthetic data is determined in the following manner:

[0152] Add scene constraints to the large language model, where the scene constraints are different for different rounds of iteration;

[0153] Each initial input data included in the initial prompt word set is input into the large language model after adding scene constraints for learning, and multiple second output data are obtained. Among them, one initial input data included in the initial prompt word set corresponds to at least one second output data.

[0154] The second synthetic data is determined based on the initial input data and at least one corresponding second output data.

[0155] Optionally, the first algorithm includes a first specialization sub-algorithm, a first syntax sub-algorithm, and a first anomaly detection sub-algorithm. A preset first algorithm is used to score each of the first synthetic data to obtain a first score value. The first processing unit 703 is used for:

[0156] Each piece of the first synthetic data is scored using a pre-defined first professionalism sub-algorithm to obtain a professionalism score;

[0157] Each first synthetic data is scored using a pre-defined first grammar sub-algorithm to obtain a grammar score;

[0158] The first anomaly detection sub-algorithm is used to score each of the first synthetic data to obtain anomaly detection scores;

[0159] The professionalism score, grammar score, and anomaly detection score are added together to obtain the first score.

[0160] Optionally, a preset second algorithm is used to score each of the second synthetic data to obtain a second score value, and the second processing unit 704 is used for:

[0161] The quality of each second synthetic data is scored using a pre-set second algorithm to obtain a second score value.

[0162] Optionally, before scoring each of the second synthetic data using a preset second algorithm, the method further includes:

[0163] Determine the relevant values ​​of each initial input data and scene constraints included in the initial prompt word set;

[0164] Initial input data with a relevance value less than a preset relevance threshold will be removed from the initial prompt word set.

[0165] Optionally, the iteration unit 702 determines whether the number of initial input data has changed based on the first score and / or the second score of this round.

[0166] Compare the first score of this round with the preset first standard score of this round;

[0167] If the first score is less than the first standard score, the initial input data that is below the first threshold will be removed from the initial prompt word set;

[0168] Compare the second score of this round with the preset second standard score of this round;

[0169] If the second score is less than the second standard score, the initial input data that is below the second threshold will be removed from the initial prompt word set;

[0170] If the first score is not less than the first standard score and / or the second score is not less than the second standard score, then the number of initial input data in the initial prompt word set corresponding to this round will no longer change.

[0171] Based on the same inventive concept, see [reference] Figure 8 As shown, this application embodiment provides a server, including: a memory 801 for storing executable instructions; and a processor 802 for reading and executing the executable instructions stored in the memory, and executing any of the methods described in the first aspect above.

[0172] Based on the same inventive concept, embodiments of this application provide a computer-readable storage medium that, when the instructions in the storage medium are executed by a processor, enables the processor to perform the method described in any of the first aspects above.

[0173] In summary, this application provides a method, apparatus, and storage medium for simplifying a prompt word set. The method is applied to a server and includes: inputting a pre-set initial prompt word set into a large language model for learning; processing multiple initial input data included in the initial prompt word set iteratively; determining whether the number of initial input data has changed based on the first score and / or second score of each iteration; and stopping the iteration process when the number no longer changes. Each iteration process is as follows: scoring each first synthesized data using a pre-set first algorithm to obtain a first score value, wherein the first synthesized data is a large language model. The model learns from each initial input data in the initial prompt word set; and / or, a preset second algorithm is used to score each second synthesized data to obtain a second score value. The second synthesized data is obtained by learning from each initial input data in the initial prompt word set using a large language model with added scene constraints. The above-mentioned methods of simplifying each initial input data using the first score value and / or the second score value, and learning from each initial input data using a large language model with added scene constraints, effectively improve the quality of the prompt words, making the prompt words more in line with the current application scenario, thereby improving the quality of the generated synthesized data.

[0174] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program product systems. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product system implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0175] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program product systems according to this application. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0176] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0177] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0178] Obviously, those skilled in the art can make various modifications and variations to this application without departing from the spirit and scope of this application. Therefore, if such modifications and variations fall within the scope of the claims of this application and their equivalents, this application also intends to include such modifications and variations.

Claims

1. A method for simplifying a set of prompt words, characterized in that, Applied to a server, the method includes: A pre-set set of initial prompt words is input into a large language model for learning, and the multiple initial input data included in the set of initial prompt words are processed iteratively. During each iteration, the number of initial input data points is determined based on the first score and / or second score of the current round, and the iteration process stops when the number no longer changes. The iteration process for each round is as follows: The first algorithm includes a first professionalism sub-algorithm, a first grammar sub-algorithm, and a first anomaly detection sub-algorithm. The first professionalism sub-algorithm is used to score each first synthesized data point to obtain a professionalism score. The first grammar sub-algorithm is used to score each first synthesized data point to obtain a grammar score. The first anomaly detection sub-algorithm is used to score each first synthesized data point to obtain an anomaly detection score. The professionalism score, the grammar score, and the anomaly detection score are added together to obtain a first score value. The first synthesized data is obtained by the large language model after learning from each of the initial input data points included in the initial prompt word set; and / or The quality of each second synthetic data is scored using a preset second algorithm to obtain a second score value. The second synthetic data is obtained by learning each of the initial input data included in the initial prompt word set using a large language model with added scene constraints.

2. The method as described in claim 1, characterized in that, The first synthesized data is determined in the following manner: A pre-set set of initial prompt words is input into a large language model for learning, resulting in multiple first output data. The initial prompt word set includes one of the initial input data that corresponds to at least one of the first output data. The first synthesized data is determined based on the initial input data and at least one corresponding first output data.

3. The method as described in claim 1, characterized in that, The second synthetic data was determined in the following manner: Add scene constraints to the large language model, wherein the scene constraints are different for different rounds of iteration; Each of the initial input data included in the initial prompt word set is input into the large language model after adding the scene constraints for learning, to obtain multiple second output data, wherein one of the initial input data included in the initial prompt word set corresponds to at least one second output data; The second synthesized data is determined based on the initial input data and at least one corresponding second output data.

4. The method as described in claim 1, characterized in that, Before scoring each of the second synthetic data using a preset second algorithm, the method further includes: Determine the correlation values ​​between each of the initial input data included in the initial prompt word set and the scene constraints; The initial input data whose relevance value is less than a preset relevance threshold will be deleted from the initial prompt word set.

5. The method according to any one of claims 1 to 4, characterized in that, The step of determining whether the number of initial input data has changed based on the first score and / or the second score of this round includes: Compare the first score value of this round with the preset first standard score of this round; If the first score is less than the first standard score, then the initial input data that is below the first threshold will be deleted from the initial prompt word set; Compare the second score value of this round with the preset second standard score of this round; If the second score is less than the second standard score, then the initial input data that is below the second threshold will be deleted from the initial prompt word set; If the first score is not less than the first standard score and / or the second score is not less than the second standard score, then it is determined that the number of initial input data in the initial prompt word set corresponding to this round will no longer change.

6. A device for condensing a set of prompt words, characterized in that, include: The input unit is used to input a pre-set set of initial prompt words into the large language model for learning, and to process the multiple initial input data included in the set of initial prompt words in an iterative manner; An iterative unit is used to determine whether the number of initial input data has changed based on the first score value and / or the second score value of the current iteration during each iteration, and to stop the iteration process when the number no longer changes; wherein, the iteration process of each round is as follows: A first processing unit is configured to: use a first algorithm including a first professionalism sub-algorithm, a first grammar sub-algorithm, and a first anomaly detection sub-algorithm; use a preset first professionalism sub-algorithm to score each first synthesized data to obtain a professionalism score; use a preset first grammar sub-algorithm to score each first synthesized data to obtain a grammar score; use a preset first anomaly detection sub-algorithm to score each first synthesized data to obtain an anomaly detection score; and add the professionalism score, the grammar score, and the anomaly detection score to obtain a first score value. The first synthesized data is obtained by the large language model after learning from each of the initial input data included in the initial prompt word set; and / or The second processing unit is used to score the quality of each second synthetic data using a preset second algorithm to obtain a second score value. The second synthetic data is obtained by learning each of the initial input data included in the initial prompt word set using a large language model with added scene constraints.

7. A server, characterized in that, include: Memory, used to store executable instructions; A processor for reading and executing executable instructions stored in the memory to implement the method as described in any one of claims 1 to 5.

8. A computer-readable storage medium, characterized in that, When the instructions in the storage medium are executed by the processor, the processor is able to perform the method as described in any one of claims 1 to 5.

Citation Information

Patent Citations

  • Generation method and device of thinking chain expansion data based on large language model

    CN117217202A

  • Method and device for recommending articles based on large language model and search engine

    CN117391824A