Code generation task oriented DocString prompt compression method and device
By preprocessing DocString and dynamically adjusting the compression rate, the problem of low compression efficiency and degraded generation quality of DocString is solved, and efficient code generation tasks are achieved.
Patent Information
- Application Number
- CN202510533682.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-25
- Publication Date
- 2025-07-25
AI Technical Summary
In the code generation task, the compression efficiency of DocString is limited, resulting in high computing costs and degradation of the quality of the generated code. The existing methods are difficult to adapt to different scenarios and lack flexibility.
By obtaining the initial DocString, removing the target symbols and stop words, calculating the token importance score based on the principles of information theory, building a search space and determining the compression prompt through the constraint model, dynamically adjusting the compression rate to maintain the quality of the generated code.
Significantly reduce the length of input prompts, improve model efficiency, maintain the quality of generated code, adapt to different code generation scenarios, and reduce calculation costs.
Smart Images

Figure CN120371316A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of this specification relate to the field of computer technology, and particularly to a DocString prompt compression method for code generation tasks. Background Art
[0002] Large language models (LLMs) have become key tools in software development due to their powerful language understanding and generation capabilities. Especially in the field of code generation, they can convert function / method signatures and DocStrings into executable code. However, their widespread deployment brings challenges in improving inference efficiency and reducing computational resource requirements, mainly reflected in the computational overhead of model inference and the cost of calling third-party LLMs' APIs.
[0003] As a string literal in source code, DocString is used to document specific functions or features, which is crucial for enhancing code readability, maintainability, and usability. It is usually located at the beginning of a function or method and guides the model to generate compliant code by describing the code function, parameters, return values, and exceptions in code generation. However, in the code generation applications of LLMs, DocString as a prompt has quality differences. Some DocStrings are verbose and contain a large amount of redundant information, which not only increases the computational cost of model inference but also may hinder the model from accurately understanding user requirements. For example, OpenAI's GPT-4o processes approximately 200 billion tokens per day, and Baidu's ERNIE model processes more than 1.5 billion calls per day. Under large-scale processing, a 25% to 40% compression of user input can significantly reduce the overall computational cost and energy consumption. In addition, shorter prompts can also reduce the financial cost of calling third-party LLMs' APIs because input tokens are usually 1 to 5 times cheaper than output tokens, with differences among different LLM providers.
[0004] DocString compression can complement other efficiency improvement techniques (such as efficient decoding strategies and optimized programming language syntax), and can also provide a theoretical basis for more complex retrieval-augmented generation (RAG)-based code generation tasks. However, applying existing NLP prompt compression techniques (such as Selective_Context and LLMLingua) to DocString faces challenges: one is the limited compression efficiency, and a compression rate exceeding 10% will lead to a significant decline in the quality of generated code because these methods fail to extract code-related semantic information; the other is the lack of flexibility, as the compression rate needs to be set manually, making it difficult to adapt to different code generation scenarios and difficult to balance the compression rate and model efficiency.
[0005] How to solve the above technical problems has become the subject of this application. Summary of the Invention
[0006] In view of this, the embodiments of this specification provide a DocString hint compression method for code generation tasks. One or more embodiments of this specification are also related to a DocString hint compression device for code generation tasks, a computing device, a computer-readable storage medium, and a computer program, so as to solve the technical defects existing in the prior art.
[0007] According to the first aspect of the embodiments of this specification, a DocString hint compression method for code generation tasks is provided, including: Obtain the initial DocString, remove the target symbols in the initial DocString, and perform semantic similarity checking on the stop words to determine the target DocString; Convert the target DocString into tokens, calculate the importance score of the tokens for the conditional entropy of the sequence based on the information theory principle, and sort the tokens based on the importance score to determine the sorting result; Select a preset number of tokens as candidate compression objects based on the sorting result, and construct a search space based on the continuous sequence of the candidate compression objects; Determine the compression hint based on the search space and the constraint model.
[0008] In a possible implementation, removing the target symbols in the initial DocString and performing semantic similarity checking on the stop words to determine the target DocString includes: Remove the line breaks and tab characters in the initial DocString; Identify the stop words in the initial DocString, and calculate the semantic similarity after removing the stop words using cosine similarity; In the case where the similarity exceeds the similarity threshold, remove the stop words to determine the target DocString.
[0009] In a possible implementation, converting the target DocString into tokens includes: Use a data compression algorithm to decompose the DocString into sub-word level tokens.
[0010] In a possible implementation, calculating the importance score of the tokens for the conditional entropy of the sequence based on the information theory principle and sorting the tokens based on the importance score to determine the sorting result includes: Calculate the negative logarithm probability of the tokens based on the information theory principle; Determine the importance score based on the negative logarithm probability; Sort the tokens in ascending order based on the importance score to determine the sorting result.
[0011] In a possible implementation, a preset number of tokens are selected as candidate compression objects based on the sorting result, and a search space is constructed based on the consecutive sequence of the candidate compression objects, including: Select the first preset number of tokens in the sorting result as candidate compression objects; Generate N-gram sequences based on the candidate compression objects; Merge based on the candidate compression objects and the N-gram sequences to determine the search space.
[0012] In a possible implementation, determining a compression hint based on the search space and a constraint model includes: Based on the search space, use a comparison model to calculate the similarity between the output distributions of the original and compressed DocStrings; Determine a compression hint when the similarity of the output distributions is lower than a distribution threshold.
[0013] In a possible implementation, it further includes: Verify the compression efficiency and generation quality through effectiveness evaluation, efficiency evaluation, and generalization evaluation.
[0014] According to the second aspect of the embodiments of this specification, a DocString hint compression device for code generation tasks is provided, including: A preprocessing module configured to obtain an initial DocString, remove target symbols from the initial DocString, and perform a semantic similarity check on stop words to determine a target DocString; A token importance sorting module configured to convert the target DocString into tokens, calculate the importance scores of the tokens for the conditional entropy of the sequence based on information theory principles, and sort the tokens based on the importance scores to determine a sorting result; A search space construction module configured to select a preset number of tokens as candidate compression objects based on the sorting result and construct a search space based on the consecutive sequence of the candidate compression objects; A constraint definition module configured to determine a compression hint based on the search space and a constraint model.
[0015] According to the third aspect of the embodiments of this specification, a computing device is provided, including: A memory and a processor; The memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions, and when the computer-executable instructions are executed by the processor, the steps of the above-mentioned DocString hint compression method for code generation tasks are implemented.
[0016] According to a fourth aspect of the embodiments of the present specification, a computer-readable storage medium is provided, which stores computer-executable instructions that, when executed by a processor, implement the steps of the above-mentioned DocString hint compression method for code generation tasks.
[0017] According to a fifth aspect of the embodiments of the present specification, a computer program is provided, wherein when the computer program is executed on a computer, the computer is made to execute the steps of the above-mentioned DocString hint compression method for code generation tasks.
[0018] The embodiments of the present specification provide a DocString hint compression method and device for code generation tasks. The DocString hint compression method for code generation tasks includes: obtaining an initial DocString, removing target symbols from the initial DocString, and performing semantic similarity checks on stop words to determine a target DocString; converting the target DocString into tokens, calculating the importance scores of the tokens for the conditional entropy of the token pair sequence based on information theory principles, and sorting the tokens based on the importance scores to determine a sorting result; selecting a preset number of the tokens as candidate compression objects based on the sorting result, and constructing a search space based on the continuous sequence of the candidate compression objects; and determining a compressed hint based on the search space and a constraint model. Compressing the DocString in code generation tasks can dynamically adjust the compression rate, significantly reduce the length of the input hint while maintaining the quality of the generated code, and improve the efficiency of the model. Description of the Drawings
[0019] Figure 1 is a flowchart of a DocString hint compression method for code generation tasks provided by an embodiment of the present specification; Figure 2 is a schematic diagram of a DocString hint compression method for code generation tasks provided by an embodiment of the present specification; Figure 3 is a schematic structural diagram of a DocString hint compression device for code generation tasks provided by an embodiment of the present specification; Figure 4 is a block diagram of the structure of a computing device provided by an embodiment of the present specification. Detailed Embodiments
[0020] In the following description, numerous specific details are set forth in order to provide a thorough understanding of this specification. However, this specification can be implemented in many other ways different from those described herein, and those skilled in the art can make similar extensions without departing from the spirit of this specification. Therefore, this specification is not limited by the specific implementations disclosed below.
[0021] The terms used in one or more embodiments of this specification are for the purpose of describing specific embodiments only and are not intended to limit one or more embodiments of this specification. The singular forms "a" and "the" used in one or more embodiments of this specification and the appended claims are also intended to include the plural forms unless the context clearly dictates otherwise. It should also be understood that the term "and / or" used in one or more embodiments of this specification refers to and encompasses any and all possible combinations of one or more of the associated listed items.
[0022] It should be understood that although the terms first, second, etc. may be used in one or more embodiments of this specification to describe various information, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from each other. For example, without departing from the scope of one or more embodiments of this specification, the first may also be referred to as the second, and similarly, the second may also be referred to as the first. Depending on the context, the word "if" as used herein may be interpreted as "when" or "while" or "in response to determining".
[0023] In this specification, a DocString hint compression method for code generation tasks is provided. This specification also relates to a DocString hint compression device for code generation tasks, a computing device, and a computer-readable storage medium, which will be described in detail one by one in the following embodiments.
[0024] See Figure 1 , Figure 1 which shows a flowchart of a DocString hint compression method for code generation tasks provided according to an embodiment of this specification, specifically including the following steps.
[0025] Step 101: Obtain an initial DocString, remove the target symbols from the initial DocString, and perform a semantic similarity check on stop words to determine the target DocString; In a possible implementation, remove the target symbols from the initial DocString and perform a semantic similarity check on the stop words to determine the target DocString, including: removing line breaks and tab characters from the initial DocString; identifying the stop words in the initial DocString and calculating the semantic similarity after removing the stop words using cosine similarity; and if the similarity exceeds the similarity threshold, remove the stop words to determine the target DocString.
[0026] In practical applications, refer to Figure 2 , first perform preprocessing, remove line breaks and tab characters from the DocString, and perform a semantic similarity check on the stop words (such as "the", "a", etc.) to ensure that the semantics of the compressed DocString are highly consistent with the original version. The specific steps are as follows: Remove line breaks and tab characters from the DocString to simplify the text structure; Identify and mark the stop words (such as "the", "a", etc.), calculate the semantic similarity after removing these stop words using cosine similarity, and if the similarity exceeds the threshold (such as 0.999), remove the stop word.
[0027] Step 102: Convert the target DocString into tokens, calculate the importance score of the tokens for the conditional entropy of the sequence based on the information theory principle, and sort the tokens based on the importance score to determine the sorting result; In a possible implementation, convert the target DocString into tokens, including: using a data compression algorithm to decompose the DocString into sub-word level tokens.
[0028] In practical applications, use BPE (Byte Pair Encoding) to decompose the DocString into sub-word level tokens to adapt to the input format of modern large language models LLMs.
[0029] In a possible implementation, calculate the importance score of the tokens for the conditional entropy of the sequence based on the information theory principle, and sort the tokens based on the importance score to determine the sorting result, including: calculating the negative log probability of the tokens based on the information theory principle; determining the importance score based on the negative log probability; sorting the tokens in ascending order based on the importance score to determine the sorting result.
[0030] In practical applications, calculate the contribution of each token to the conditional entropy of the sequence based on the information theory principle, quantify its importance, and sort the tokens according to the importance. The specific steps are as follows: Using the information theory principle, calculate the contribution of each token to the conditional entropy of the sequence, which can be done by calculating each token wi Formally, the negative log probability is expressed as , which reflects the probability that a token w will i The amount of information carried. The higher the value, the more information the token carries, and therefore the more important it is to the overall meaning of the sequence; According to the calculated importance scores, all tokens are sorted in ascending order, with the lower importance tokens at the front.
[0031] Step 103: selecting a preset number of tokens as candidate compression objects based on the sorting result, and constructing a search space based on a continuous sequence of the candidate compression objects; In one possible implementation, a preset number of tokens are selected as candidate compression objects based on the sorting results, and a search space is constructed based on a continuous sequence of the candidate compression objects, including: selecting the first preset number of tokens in the sorting results as candidate compression objects; generating an N-gram sequence based on the candidate compression objects; and determining the search space based on merging the candidate compression objects and the N-gram sequence.
[0032] In practical applications, the N tokens with the lowest importance are selected as candidate compression objects, and the continuous sequence (N-gram) of these tokens is considered to construct a complete search space, which includes the following steps: Select the N tokens with the lowest importance from the sorted token list as candidates for compression; For each candidate token, generate its possible N-gram sequence (such as 1-gram, 2-gram, 3-gram, etc.) to explore the compression effect of consecutive tokens. Formally, for each 1≤k≤N, consider k-gram, that is , where each G k represents the set of all possible consecutive k-grams among the top N tokens with the lowest importance; All candidate tokens and their N-gram sequences are merged to form a complete search space for subsequent compression operations.
[0033] Step 104: Determine compression hints based on the search space and the constraint model.
[0034] In a possible implementation, determining a compression hint based on a search space and a constraint model includes: comparing the similarity of output distributions of original and compressed DocStrings based on the search space through a comparison model; and determining a compression hint when the similarity of the output distributions is lower than a distribution threshold.
[0035] In practical applications, by constraining the similarity of the model output distribution, the compressed DocString is ensured to be consistent with the original DocString when generating code, and the specific steps are as follows: By comparing the similarity of the model output distributions for the original and compressed DocStrings (such as cosine similarity), the compressed DocString is ensured to be consistent with the original DocString when generating code; Set a similarity threshold (such as 0.999). Only when the similarity between the compressed DocString and the original DocString is higher than this threshold is the compression considered effective.
[0036] In a possible implementation, it also includes: verifying the compression efficiency and generation quality through effectiveness evaluation, efficiency evaluation, and generalization evaluation.
[0037] In practical applications, through effectiveness evaluation (such as Pass@1 metric), efficiency evaluation (such as GPU memory usage, FLOPs analysis, and inference time metric), and generalization evaluation (such as performance on different programming languages and datasets), the advantages of the method (ShortenDoc) of this application in terms of compression efficiency and generation quality are verified, and the specific steps are as follows: Measure the accuracy and reliability of the compressed DocString in the code generation task through the Pass@1 metric to ensure that the code generated by the model can pass all test cases; To comprehensively evaluate the efficiency of ShortenDoc, evaluate it from three aspects: GPU memory usage, FLOPs analysis, and inference time metric respectively; Test the performance of ShortenDoc on multiple programming languages and different datasets to verify its applicability and cross - language ability in different scenarios.
[0038] Compared with the prior art, the beneficial effects of this solution are as follows: A DocString prompt compression method for code generation tasks proposed in this application first preprocesses the DocString, removes formatting characters such as line breaks and tab characters therein, and removes some high-frequency but semantically less contributing stop words based on semantic similarity analysis; then calculates the importance score of each token in the DocString, and sorts them according to the contribution of the token to the conditional entropy of the sequence; secondly, constructs a search space based on the importance score, adopts a token importance sorting strategy to select the token sequences that may be compressed, and considers the compression effect of consecutive token sequences; finally, dynamically adjusts the compression rate during the compression process, and ensures that the compressed DocString is highly consistent with the original DocString semantically through semantic similarity constraints, so as to achieve efficient DocString compression while maintaining the quality of the generated code.
[0039] Furthermore, through experiments, as shown in Table 1, Table 1 is a performance comparison table of the method of this embodiment and other baseline methods. Among them, DaraSet is the data set, Method is the method, and Ratio is the ratio. The DocString prompt compression method ShortenDoc for code generation tasks proposed in this embodiment is superior to other methods on all data sets and models, especially at higher compression rates. For example, on the HumanEval data set, ShortenDoc achieved the highest Pass@1 score on all models. Even when the DocString length was reduced by 30%, it could effectively maintain or even improve the quality of code generation. This shows that ShortenDoc can significantly compress the input prompt while maintaining or even improving the model's Pass@1 score.
[0040] Table 1
[0041] In addition, in some cases, ShortenDoc even outperformed the uncompressed baseline. For example, on the HumanEval data set, the Pass@1 scores of DeepSeekCoder-6.7B, CodeQwen, and CodeGeeX4 were even higher after being compressed by ShortenDoc. This surprising result indicates that ShortenDoc not only retains the key information but also may eliminate the redundant or interfering content in the DocString, thus enabling the model to focus more on the relevant information during the code generation process.
[0042] On other datasets such as CodeHarmony, MBPP, Subtle, Creative, and BigCodeBench, ShortenDoc also consistently matched or outperformed competing methods, especially in aggressive compression scenarios. The general advantages of ShortenDoc across different datasets and models highlight its versatility and robustness. These results suggest that ShortenDoc effectively balances compression and information retention, making it a powerful tool for optimizing prompt inputs in code generation tasks.
[0043] Furthermore, ShortenDoc maintains high Pass@1 scores across models of different sizes and architectures, indicating its broad applicability. It can effectively assist both open-source and closed-source models, including state-of-the-art systems such as GPT-4o. This broad applicability enhances the practical value of ShortenDoc in real-world code generation applications.
[0044] In another embodiment, the improvement of the method of the present application in terms of efficiency is further verified. Specifically, it includes the following: (1) GPU memory usage: GPU memory consumption is a key resource limitation when deploying LLMs for code generation tasks. To evaluate the memory efficiency of the present application, the GPU memory usage when using compressed and uncompressed DocStrings was compared; (2) FLOPs analysis: FLOPs (floating-point operations per second) refers to the total number of floating-point operations required during model execution and is an indicator of model complexity. FLOPs is closely related to the input length and model structure. The goal of DocString compression is to reduce computational costs and improve model efficiency. Therefore, the FLOPs metric was used to evaluate the computational load of different models before and after compression, involving multiple datasets; (3) Inference time: The end-to-end inference time is a key metric in practical applications and directly affects the user experience and system throughput. To evaluate the actual benefits of the compression method, comprehensive timing measurements were conducted on an NVIDIA RTX 3090 GPU. All models were configured to use torch.bfloat16 precision, and the input and output lengths were both set to 512 tokens.
[0045] Table 2
[0046] Experiments have shown that, as shown in Table 2 (Table 2 shows the comparison of floating-point operation amounts (FLOPs) for different datasets and models), the method of the present application is significantly more efficient than the baseline model in multiple aspects (the FLOPs required for ShortenDoc calculation are shown in parentheses).
[0047] GPU Memory Usage: This application uses CodeGPT-py-adapted as the base model for compression, but its memory footprint is negligible compared to the deployment of LLMs. Specifically, CodeGPT-py-adapted has only 124 million parameters and requires only 0.23 GB of GPU memory, an overhead that is almost negligible in the face of the large amounts of memory required by modern LLMs.
[0048] FLOPs Analysis: The results in Table 3 show that the computational load of the model generally decreases after compression. For example, on the HumanEval dataset, the FLOPs decreased from 0.34 to 0.28, a reduction of approximately 17.6%; on the CodeHarmony dataset, the FLOPs decreased from 0.15 to 0.12, a reduction of approximately 20%. This reduction in computational load due to compression means that the model can process data faster or process more data in the same amount of time under the same hardware conditions.
[0049] Inference Time: The results in Table 3 indicate that the inference times for different model sizes vary significantly. The DeepSeekCoder model with 1.3 billion parameters takes approximately 9.21 seconds to generate each time, while the large model with approximately 7 billion parameters takes approximately 17.77 seconds. In contrast, when using the GPT-4 API, the generation time is typically between 5 and 10 seconds, although this may fluctuate depending on server load. The time overhead introduced by the ShortenDoc compression process is relatively small, averaging approximately 2 seconds per DocString. This additional preprocessing time is a worthwhile investment. In practical applications, the entire processing pipeline (compression plus inference) remains very efficient.
[0050] Overall, the method of this application is significantly more efficient than the baseline model in multiple aspects. This further demonstrates the improvements of the method of this application in terms of storage requirements, GPU usage, computational efficiency, and environmental impact.
[0051] Furthermore, in another embodiment, the performance of ShortenDoc is tested on multiple programming languages and different datasets to verify its applicability and cross-language ability in different scenarios. Specifically, it includes the following: (1) Given that the datasets selected in Example 1 are all in the Python language, in order to explore the generalization ability of ShortenDoc, Example 3 selects four other programming languages in the HumanEval-X dataset, namely C++, Go, Java, and JavaScript, for experiments; (2)The main reason for choosing HumanEval-X is that most of the publicly available code generation datasets in non-Python languages are currently variants of HumanEval, which has become the de facto benchmark in this field. Although it is recognized that using variants of a single dataset may introduce potential biases, HumanEval-X is one of the few comprehensive and validated multilingual code generation benchmarks currently available in the research community.
[0052] (3)Since these four datasets share HumanEval with DocString, Example 3 directly migrates the compressed DocString obtained for HumanEval in Example 1 to these datasets; Table 3
[0053] Through experiments, the results in Table 3 show that Table 3 presents the generalization ability of ShortenDoc of this solution in different programming languages. In all four programming languages, compared with the baseline methods (such as Random, SelectiveContext, and LLMLingua2), ShortenDoc consistently maintains superior performance (for example, in the C++ dataset, the average Pass@1 score of ShortenDoc is 51.71%, while those of the other methods are 46.10%, 33.66%, and 42.93% respectively). Despite the differences in structure and syntax among these languages, ShortenDoc is still able to maintain its ability to retain key information.
[0054] It is worth noting that in the C++ and Java datasets, ShortenDoc significantly outperforms other methods and is closer to the uncompressed baseline, indicating that its evaluation of token importance and compression strategy when dealing with these languages do not tightly depend on a specific programming language. In the JavaScript dataset (where the use of DocString is usually more diverse), ShortenDoc still achieves considerable benefits, further demonstrating its versatility. The Go dataset also shows a similar trend, and ShortenDoc performs well in adapting to the more concise and function-oriented syntax of Go.
[0055] However, it is also observed that when directly migrating the compressed DocString obtained in RQ1 (Python) to other languages (such as C++, Go, Java, and JavaScript), there is a slight decrease in the Pass@1 score. Although ShortenDoc outperforms the baseline methods in all cases, its performance does not fully reach the results achieved by the uncompressed DocString in these new languages.
[0056] This performance degradation can be attributed to the fact that the base language model CodeGPT used in the method is specifically trained and optimized for Python. The model's understanding of the syntax, structure, and semantics of Python code affects the way DocStrings are compressed. When these compressed DocStrings are applied to other programming languages, certain language-specific nuances may not be effectively preserved, resulting in a slight decrease in code generation performance.
[0057] Despite this limitation, ShortenDoc still demonstrates strong generalization ability and can retain the critical information required for code generation tasks in multiple programming languages. This indicates that although the best results may be obtained when used in conjunction with language-specific models, ShortenDoc still performs well even when used across languages with little adjustment. This shows that this application has strong generalization ability and can retain key information in multiple programming languages.
[0058] The embodiments of this specification provide a DocString prompt compression method and apparatus for code generation tasks. The DocString prompt compression method for code generation tasks includes: obtaining an initial DocString, removing target symbols from the initial DocString, and performing a semantic similarity check on stop words to determine a target DocString; converting the target DocString into tokens, calculating the importance scores of the tokens for the conditional entropy of the token pair sequence based on information theory principles, and sorting the tokens based on the importance scores to determine a sorting result; selecting a preset number of the tokens as candidate compression objects based on the sorting result, and constructing a search space based on the continuous sequences of the candidate compression objects; and determining a compression prompt based on the search space and a constraint model. Compressing the DocString in code generation tasks can dynamically adjust the compression rate, significantly reduce the length of the input prompt while maintaining the quality of the generated code, and improve the efficiency of the model.
[0059] Corresponding to the above method embodiments, this specification also provides embodiments of a DocString prompt compression apparatus for code generation tasks. Figure 3 Fig. shows a schematic structural diagram of a DocString prompt compression apparatus for code generation tasks provided by an embodiment of this specification. As Figure 3 shown, the apparatus includes: A preprocessing module 301, configured to obtain an initial DocString, remove target symbols from the initial DocString, and perform a semantic similarity check on stop words to determine a target DocString; The token importance ranking module 302 is configured to convert the target DocString into tokens, calculate the importance scores of the tokens for the conditional entropy of the sequence based on the information theory principle, and rank the tokens based on the importance scores to determine the ranking result; The search space construction module 303 is configured to select a preset number of tokens as candidate compression objects based on the ranking result and construct a search space based on the consecutive sequences of the candidate compression objects; The constraint definition module 304 is configured to determine the compression hint based on the search space and the constraint model.
[0060] In a possible implementation, removing the target symbols from the initial DocString and performing a semantic similarity check on the stop words to determine the target DocString includes: Removing the line breaks and tab characters from the initial DocString; Identifying the stop words in the initial DocString and calculating the semantic similarity after removing the stop words using cosine similarity; In the case where the similarity exceeds the similarity threshold, removing the stop words and determining the target DocString.
[0061] In a possible implementation, converting the target DocString into tokens includes: Using a data compression algorithm to decompose the DocString into sub-word level tokens.
[0062] In a possible implementation, calculating the importance scores of the tokens for the conditional entropy of the sequence based on the information theory principle and ranking the tokens based on the importance scores to determine the ranking result includes: Calculating the negative log probability of the tokens based on the information theory principle; Determining the importance scores based on the negative log probability; Performing an ascending order ranking on the tokens based on the importance scores to determine the ranking result.
[0063] In a possible implementation, selecting a preset number of tokens as candidate compression objects based on the ranking result and constructing a search space based on the consecutive sequences of the candidate compression objects includes: Selecting the first preset number of tokens in the ranking result as candidate compression objects; Generating N-gram sequences based on the candidate compression objects; Merging the candidate compression objects and the N-gram sequences to determine the search space.
[0064] In a possible implementation, determining the compression hint based on the search space and the constraint model includes: Based on the search space, comparing the output distributions of the original and compressed DocStrings by a comparison model for similarity; When the similarity of the output distributions is lower than the distribution threshold, determining a compression hint.
[0065] In a possible implementation, it further includes: Verifying the compression efficiency and generation quality through effectiveness evaluation, efficiency evaluation, and generalization evaluation.
[0066] The embodiments of this specification provide a DocString hint compression method and apparatus for code generation tasks. The DocString hint compression apparatus for code generation tasks includes: obtaining an initial DocString, removing target symbols from the initial DocString, and performing semantic similarity checking on stop words to determine a target DocString; converting the target DocString into tokens, calculating the importance scores of the tokens for the conditional entropy of the token pair sequence based on information theory principles, and sorting the tokens based on the importance scores to determine a sorting result; selecting a preset number of the tokens as candidate compression objects based on the sorting result, and constructing a search space based on the continuous sequences of the candidate compression objects; determining a compression hint based on the search space and a constraint model. Compressing the DocString in code generation tasks can dynamically adjust the compression rate, significantly reducing the length of the input hint while maintaining the quality of the generated code and improving the efficiency of the model.
[0067] The above is a schematic solution of a DocString hint compression apparatus for code generation tasks in this embodiment. It should be noted that the technical solution of the DocString hint compression apparatus for code generation tasks and the technical solution of the above-mentioned DocString hint compression method for code generation tasks belong to the same concept. For the details not described in detail in the technical solution of the DocString hint compression apparatus for code generation tasks, reference can be made to the description of the technical solution of the above-mentioned DocString hint compression method for code generation tasks.
[0068] Figure 4 FIG. shows a structural block diagram of a computing device 400 according to an embodiment of this specification. The components of the computing device 400 include but are not limited to a memory 410 and a processor 420. The processor 420 is connected to the memory 410 through a bus 430, and a database 450 is used to store data.
[0069] The computing device 400 also includes an access device 440, which enables the computing device 400 to communicate via one or more networks 460. Examples of such networks include the Public Switched Telephone Network (PSTN), Local Area Network (LAN), Wide Area Network (WAN), Personal Area Network (PAN), or a combination of communication networks such as the Internet. The access device 440 may include one or more of any type of wired or wireless network interfaces (e.g., a network interface card (NIC)), such as an IEEE 802.11 Wireless Local Area Network (WLAN) wireless interface, a Worldwide Interoperability for Microwave Access (Wi-MAX) interface, an Ethernet interface, a Universal Serial Bus (USB) interface, a cellular network interface, a Bluetooth interface, a Near Field Communication (NFC) interface.
[0070] In one embodiment of the present specification, the above components of the computing device 400, as well as Figure 4 other components not shown, may also be connected to each other, for example, via a bus. It should be understood that Figure 4 the block diagram of the computing device shown is merely for illustrative purposes and is not a limitation on the scope of the present specification. Those skilled in the art may add or replace other components as needed.
[0071] The computing device 400 can be any type of stationary or mobile computing device, including a mobile computer or mobile computing device (e.g., a tablet computer, a personal digital assistant, a laptop computer, a notebook computer, a netbook, etc.), a mobile phone (e.g., a smartphone), a wearable computing device (e.g., a smartwatch, smart glasses, etc.), or other types of mobile devices, or a stationary computing device such as a desktop computer or a Personal Computer (PC). The computing device 400 can also be a mobile or stationary server.
[0072] Among them, the processor 420 is used to execute the following computer-executable instructions. When the computer-executable instructions are executed by the processor, the steps of the above-mentioned DocString hint compression method for code generation tasks are implemented. The above is a schematic solution of a computing device according to this embodiment. It should be noted that the technical solution of this computing device and the technical solution of the above-mentioned DocString hint compression method for code generation tasks belong to the same concept. For the detailed content not described in the technical solution of the computing device, reference can be made to the description of the technical solution of the above-mentioned DocString hint compression method for code generation tasks.
[0073] An embodiment of this specification also provides a computer-readable storage medium, which stores computer-executable instructions. When the computer-executable instructions are executed by the processor, the steps of the above-mentioned DocString hint compression method for code generation tasks are implemented.
[0074] The above is a schematic solution of a computer-readable storage medium according to this embodiment. It should be noted that the technical solution of this storage medium and the technical solution of the above-mentioned DocString hint compression method for code generation tasks belong to the same concept. For the detailed content not described in the technical solution of the storage medium, reference can be made to the description of the technical solution of the above-mentioned DocString hint compression method for code generation tasks.
[0075] An embodiment of this specification also provides a computer program. When the computer program is executed on a computer, the computer is made to execute the steps of the above-mentioned DocString hint compression method for code generation tasks.
[0076] The above is a schematic solution of a computer program according to this embodiment. It should be noted that the technical solution of this computer program and the technical solution of the above-mentioned DocString hint compression method for code generation tasks belong to the same concept. For the detailed content not described in the technical solution of the computer program, reference can be made to the description of the technical solution of the above-mentioned DocString hint compression method for code generation tasks.
[0077] The above describes specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be executed in a different order than in the embodiments and still achieve the desired result. Additionally, the processes depicted in the figures do not necessarily require the particular order or sequential order shown to achieve the desired result. In certain implementations, multitasking and parallel processing are also possible or may be advantageous.
[0078] The computer instructions include computer program code, which may be in the form of source code, object code, executable files or some intermediate forms, etc. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, removable hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, and software distribution medium, etc. It should be noted that the content included in the computer-readable medium can be appropriately increased or decreased according to the requirements of legislation and patent practice within the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, the computer-readable medium does not include electrical carrier signals and telecommunication signals.
[0079] It should be noted that for the foregoing method embodiments, for the sake of simplicity of description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the embodiments of this specification are not limited by the described action sequence, because according to the embodiments of this specification, some steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to the embodiments of this specification.
[0080] In the above embodiments, the descriptions of the various embodiments have their own focuses. For the parts not detailed in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0081] The preferred embodiments of this specification disclosed above are only used to help explain this specification. The optional embodiments do not elaborate on all the details, nor do they limit the invention to the specific embodiments described. Obviously, many modifications and changes can be made according to the content of the embodiments of this specification. This specification selects and specifically describes these embodiments to better explain the principles and practical applications of the embodiments of this specification, so that those skilled in the art can well understand and utilize this specification. This specification is only limited by the claims and their full scope and equivalents.
Claims
1. A DocString hint compression method for code generation tasks, characterized in that, Including: Obtain an initial DocString, remove target symbols from the initial DocString, and perform a semantic similarity check on stop words to determine a target DocString; Convert the target DocString into tokens, calculate the importance score of the tokens for the conditional entropy of the sequence based on the information theory principle, and sort the tokens based on the importance score to determine a sorting result; Select a preset number of the tokens as candidate compression objects based on the sorting result, and construct a search space based on the continuous sequences of the candidate compression objects; Determine a compression hint based on the search space and a constraint model.
2. The method according to claim 1, wherein The removing target symbols from the initial DocString and performing a semantic similarity check on stop words to determine a target DocString includes: Remove line breaks and tab characters from the initial DocString; Identify stop words in the initial DocString, and calculate the semantic similarity after removing the stop words using cosine similarity; If the similarity exceeds a similarity threshold, then remove the stop words to determine a target DocString.
3. The method according to claim 1, wherein The converting the target DocString into tokens includes: Use a data compression algorithm to decompose the DocString into sub-word level tokens.
4. The method according to claim 1, wherein The calculating the importance score of the tokens for the conditional entropy of the sequence based on the information theory principle and sorting the tokens based on the importance score to determine a sorting result includes: Calculate the negative log probability of the tokens based on the information theory principle; Determine the importance score based on the negative log probability; Sort the tokens in ascending order based on the importance score to determine a sorting result.
5. The method according to claim 4, wherein The selecting a preset number of the tokens as candidate compression objects based on the sorting result and constructing a search space based on the continuous sequences of the candidate compression objects includes: Select the tokens in the front of the sorting result with the preset number as candidate compression objects; Generate N-gram sequences based on the candidate compression objects; Merge based on the candidate compression objects and the N-gram sequences to determine a search space.
6. The method according to claim 1, wherein The determining a compression hint based on the search space and a constraint model includes: Based on the search space, compare the similarity of the output distributions of the original and compressed DocStrings through a comparison model; If the similarity of the output distributions is lower than a distribution threshold, determine the compression hint.
7. The method according to claim 1, characterized in that, Also including: Verify the compression efficiency and generation quality through effectiveness evaluation, efficiency evaluation, and generalization evaluation.
8. A DocString prompt compression device for code generation tasks, characterized in that Including: A preprocessing module configured to obtain an initial DocString, remove target symbols from the initial DocString, and perform a semantic similarity check on stop words to determine a target DocString; A token importance ranking module, configured to convert the target DocString into tokens, calculate the importance scores of the tokens for the conditional entropy of the sequence based on the information theory principle, and rank the tokens based on the importance scores to determine a ranking result; A search space construction module, configured to select a preset number of the tokens as candidate compression objects based on the ranking result, and construct a search space based on the continuous sequences of the candidate compression objects; A constraint definition module, configured to determine compression hints based on the search space and a constraint model.
9. A computing device, characterized in that, Comprising: A memory and a processor; The memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions. When the computer-executable instructions are executed by the processor, the steps of the DocString hint compression method for code generation tasks described in any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium storing computer-executable instructions, which when executed by a processor implement the steps of the DocString hint compression method for code generation tasks described in any one of claims 1 to 7.
Citation Information
Patent Citations
Log analysis method and device based on prompt optimization, equipment and storage medium
CN118427294A
Speculation decoding method and device for large language model and medium
CN119150848A
Multi-modal program inference
US20230176829A1