Cue word compression method and device, electronic equipment, storage medium and program product

By constructing the maximum spanning tree and semantic segmentation, combining the importance score filtering and original order combination, the problem of semantic information loss in the existing prompt word compression method is solved, efficient prompt word compression is achieved, and important semantic information is maintained.

CN120031123APending Publication Date: 2025-05-23INST OF AUTOMATION CHINESE ACAD OF SCI
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411859790.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-17
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

The existing prompt word compression method assumes that there is strong independence between tag tokens, resulting in the loss of important semantic information of compressed prompt words.

Method used

By constructing the maximum spanning tree based on the self-attention between tokens in the query document, a community detection algorithm is applied to semantic segmentation of the maximum spanning tree to obtain multiple semantic units. Then, multiple semantic units are filtered according to the importance score of each semantic unit, and finally the tokens in the filtered semantic unit are combined in the original order to obtain the compressed prompt words.

Benefits of technology

Ensure that the compressed prompt words can maintain important semantic information, solve the problem of semantic information loss caused by the independence assumption, and improve the compression efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120031123A_ABST
    Figure CN120031123A_ABST
Patent Text Reader

Abstract

The invention provides a cue word compression method and device, electronic equipment, a storage medium and a program product, and relates to the technical field of artificial intelligence, the method comprises the following steps: constructing a maximum spanning tree based on self-attention between tokens in a query document; the maximum spanning tree is the spanning tree with the maximum self-attention sum; performing semantic segmentation on the maximum spanning tree based on a community detection algorithm to obtain a plurality of semantic units; each semantic unit comprises a plurality of tokens; filtering the plurality of semantic units according to the importance score of each semantic unit; the importance score of each semantic unit is the average value of the importance scores of all tokens in each semantic unit; and combining the tokens in the filtered semantic units according to an original sequence to obtain a compressed cue word. According to the invention, the compressed cue word can maintain important semantic information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a prompt word compression method, device, electronic equipment, storage medium and program product. Background Art

[0002] A prompt word is a text provided to a large language model to guide the large language model to better complete query tasks. In some query tasks, in order to reduce query costs and improve query efficiency, the prompt word can be compressed before use.

[0003] Existing prompt word compression methods usually assume that there is strong independence between tokens, that is, the importance of each token is evaluated independently. However, the independence assumption does not hold true in actual operations, and this method may cause the compressed prompt word to lose important semantic information. Summary of the invention

[0004] The present invention provides a prompt word compression method, device, electronic device, storage medium and program product, which are used to solve the defect of the compressed prompt word losing important semantic information in the prior art, and realize that the compressed prompt word can retain important semantic information.

[0005] The present invention provides a prompt word compression method, comprising: Based on the self-attention between the marked tokens in the query document, a maximum spanning tree is constructed; the maximum spanning tree is a spanning tree with the largest sum of self-attention; Performing semantic segmentation on the maximum spanning tree based on a community detection algorithm to obtain a plurality of semantic units; each of the semantic units includes a plurality of tokens; Filtering the multiple semantic units according to the importance score of each semantic unit; the importance score of each semantic unit is the average of the importance scores of all tokens in each semantic unit; The tokens in the filtered semantic units are combined in the original order to obtain compressed prompt words.

[0006] In some embodiments, before filtering the plurality of semantic units according to the importance score of each semantic unit, the method further includes: Determine the attention heads associated with contextual information in large language models; Determine the maximum cross attention corresponding to each token in the query document according to the cross attention corresponding to each token in the query document under different attention heads; The maximum cross attention corresponding to each token in the query document is used as the importance score of each token in the query document.

[0007] In some embodiments, the method further comprises: Filtering the tokens in the query document according to the importance score of each token in the query document; The filtered tokens in the query document are merged to obtain compressed prompt words.

[0008] In some embodiments, constructing a maximum spanning tree based on self-attention between tokens in the query document includes: Construct an undirected graph by taking each token in the query document as a vertex and the self-attention between the tokens in the query document as the edge weight between the vertices; The undirected graph is optimized with the maximum sum of self-attention as the optimization goal to obtain the constructed maximum spanning tree.

[0009] In some embodiments, before constructing the maximum spanning tree based on the self-attention between the marked tokens in the query document, the method further includes: Inputting the query document and the query question into a causal language model to obtain causal attention; Based on the causal attention and the length of the query document, the self-attention between tokens in the query document is obtained.

[0010] In some embodiments, the method further comprises: Based on the causal attention, the length of the query document and the length of the prompt word, a cross attention between the tokens in the query question and the tokens in the query document is obtained.

[0011] The present invention also provides a prompt word compression device, comprising: A construction module is used to construct a maximum spanning tree based on the self-attention between the marked tokens in the query document; the maximum spanning tree is a spanning tree with the largest sum of self-attention; A segmentation module, used for performing semantic segmentation on the maximum spanning tree based on a community detection algorithm to obtain a plurality of semantic units; each of the semantic units includes a plurality of tokens; A first filtering module, configured to filter the plurality of semantic units according to an importance score of each semantic unit; the importance score of each semantic unit is an average of importance scores of all tokens in each semantic unit; The combination module is used to combine the tokens in the filtered semantic units according to the original order to obtain compressed prompt words.

[0012] The present invention also provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, any one of the prompt word compression methods described above is implemented.

[0013] The present invention also provides a non-transitory computer-readable storage medium on which a computer program is stored. When the computer program is executed by a processor, the prompt word compression method described in any one of the above is implemented.

[0014] The present invention also provides a computer program product, comprising a computer program, wherein when the computer program is executed by a processor, the computer program implements any of the above-mentioned prompt word compression methods.

[0015] The prompt word compression method, device, electronic device, storage medium and program product provided by the present invention construct a maximum spanning tree based on the self-attention between tokens in the query document, and apply a community detection algorithm to perform semantic segmentation on the maximum spanning tree to achieve grouping of tokens into semantic units. Such a clustering method can ensure that tokens within the same semantic unit have strong semantic dependencies, while the semantic dependencies between tokens of different unit semantics are weak, thus solving the problem of independence assumption; then, multiple semantic units are filtered according to the importance scores of the semantic units to reduce irrelevant content and improve compression efficiency; finally, the tokens in the filtered semantic units are combined in the original order to obtain compressed prompt words, so that the compressed prompt words can maintain important semantic information. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0017] Figure 1 It is a flow chart of the prompt word compression method provided by the present invention; Figure 2 It is a structural schematic diagram of the prompt word compression device provided by the present invention; Figure 3 It is a structural schematic diagram of the electronic device provided by the present invention. DETAILED DESCRIPTION

[0018] In order to make the purpose, technical solution and advantages of the present invention clearer, the technical solution of the present invention will be clearly and completely described below in conjunction with the drawings of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0019] Cue word compression faces the challenge of independence assumption. To further address this challenge, a semantic unit identification algorithm is proposed. This method aims to ensure that there is a strong semantic dependency between tokens within the same semantic unit, while the semantic dependency between tokens in different semantic units is weak. Therefore, each semantic unit can be independently evaluated whether to keep or delete.

[0020] The method of the present invention is based on the assumption that the attention distribution within the language model essentially encodes the division mechanism of semantic units. Specifically, the self-attention between tokens captures their conditional dependencies, where tokens with strong mutual attention are semantically related and should be classified as the same semantic unit. In contrast, tokens with weaker attention have lower semantic correlation and should be assigned to different semantic units. Therefore, the task can be transformed into finding an optimal token grouping that maximizes the attention value within the group and minimizes the attention value between groups.

[0021] Suppose the set of all tokens in the query document is , Indicates the first Tokens, Indicates the number of tokens in the query document. express semantic unit, express The kth semantic unit among the semantic units, each ,and .

[0022] Based on the above assumptions, finding an optimal token grouping that maximizes the attention value within the group and minimizes the attention value between groups can be formulated as: in, In the formula, express semantic unit, i.e. Token groups, Indicates The total attention value corresponding to the semantic unit, that is, the total attention value within the group, Indicates The semantic unit and The total attention value corresponding to the semantic units, that is, the total attention value between groups, represents the weighted hyperparameter, Indicates tokentp, Indicates token tq, Represents the self-attention between token tp and token tq.

[0023] However, this optimization problem is a highly complex combinatorial optimization problem, and direct computation is infeasible considering the time constraints of the compression algorithm.

[0024] Figure 1 FIG. 1 is a flow chart of the prompt word compression method provided by the present invention. Figure 1 As shown, the present invention provides a prompt word compression method, comprising the following steps: Step 110, construct a maximum spanning tree based on the self-attention between the marked tokens in the query document; the maximum spanning tree is the spanning tree with the largest sum of self-attention.

[0025] Specifically, after determining the self-attention between tokens in the query document, a spanning tree with the largest sum of self-attention is constructed based on the self-attention between tokens in the query document, that is, a maximum spanning tree is constructed. Using the maximum spanning tree, the simple sequential connection of tokens in the query document is converted into a tree topology structure, highlighting the semantic connection between tokens.

[0026] Step 120, semantic segmentation is performed on the maximum spanning tree based on a community detection algorithm to obtain multiple semantic units; each semantic unit includes multiple tokens.

[0027] Specifically, a community detection algorithm (such as the Louvain algorithm) is used to perform semantic segmentation on the maximum spanning tree, that is, semantic dependencies are identified on the tokens in the maximum spanning tree, tokens with strong semantic dependencies are divided into the same semantic unit, and tokens with weak semantic dependencies are divided into different semantic units, thereby obtaining multiple semantic units, each of which includes multiple tokens.

[0028] It can be formulated as: In the formula, express semantic unit, Represents a maximum spanning tree.

[0029] Step 130 , filtering multiple semantic units according to the importance score of each semantic unit; the importance score of each semantic unit is the average of the importance scores of all tokens in each semantic unit.

[0030] Specifically, after obtaining multiple semantic units, the average of the importance scores of all tokens in each semantic unit is taken as the importance score of each semantic unit.

[0031] It can be formulated as: In the formula, Indicates The importance score of a semantic unit, Indicates The number of tokens in a semantic unit, Indicates The first semantic unit Tokens, Indicates The first semantic unit The importance score of a token.

[0032] After obtaining the importance score of each semantic unit, the obtained multiple semantic units are filtered according to the importance score of each semantic unit, and some semantic units with lower importance scores are filtered out.

[0033] For example, a percentage filtering method is used to sort multiple semantic units in descending order of their importance scores, and the semantic units ranked at the bottom z% are filtered out. The specific value of z can be determined according to the target compression constraint.

[0034] Step 140, combining the tokens in the filtered semantic units according to the original order to obtain compressed prompt words.

[0035] Specifically, since the tokens in the query document are grouped into semantic units, the original order of the tokens in the query document is destroyed. Therefore, the tokens in the filtered semantic units cannot be simply merged directly. Instead, the tokens in the filtered semantic units need to be combined according to the original order in the query document to obtain the compressed prompt words.

[0036] The prompt word compression method provided by the embodiment of the present invention constructs a maximum spanning tree based on the self-attention between tokens in the query document, and applies a community detection algorithm to perform semantic segmentation on the maximum spanning tree, so as to group tokens into semantic units. Such a clustering method can ensure that tokens within the same semantic unit have strong semantic dependencies, while the semantic dependencies between tokens of different unit semantics are weak, thus solving the problem of independence assumption. Then, multiple semantic units are filtered according to the importance scores of the semantic units, so as to reduce irrelevant content and improve compression efficiency. Finally, the tokens in the filtered semantic units are combined in the original order to obtain compressed prompt words, so that the compressed prompt words can maintain important semantic information.

[0037] In some embodiments, before filtering the multiple semantic units according to the importance score of each semantic unit, the method further includes: Determine the attention heads in the large language model that are relevant to the query document; According to the cross attention corresponding to each token in the query document under different attention heads, the maximum cross attention corresponding to each token in the query document is determined; The maximum cross attention corresponding to each token in the query document is used as the importance score of each token in the query document.

[0038] Specifically, recent research has found that certain attention heads (also called retrieval heads) in large language models (such as LLMs) become active when processing context. This finding not only reveals the underlying mechanism of large language models utilizing context, but also provides a new way to enhance the hint word compression method by decomposing the contextual ability of large language models. Based on this, it is proposed to use cross attention as a compression metric to decide whether to keep or discard tokens in the query document.

[0039] Different attention heads show unique attention distribution in the large language model, and each distribution reflects different information utilization patterns. Apply the maximum strategy to all attention heads, that is, first determine the attention head related to the query document in the large language model; then determine the cross attention corresponding to each token in the query document under different attention heads, that is, under one attention head, each token in the query document corresponds to a cross attention, and from the multiple cross attentions corresponding to each token in the query document, determine the maximum cross attention corresponding to each token in the query document; finally, use the maximum cross attention corresponding to each token in the query document as the importance score of each token in the query document.

[0040] It can be formulated as: In the formula, Indicates that the query document The importance score of the query document, H represents the set of attention heads related to the query document, Indicates that under the attention head h, the query document Corresponding cross attention.

[0041] The prompt word compression method provided by the embodiment of the present invention evaluates the importance of each token by querying the maximum cross-attention corresponding to each token in the document. This method can more effectively retain the information integration pattern of the attention head, thereby retaining important semantic information during the compression process.

[0042] In some embodiments, the prompt word compression method provided by the present invention further includes: Filter the tokens in the query document according to the importance score of each token in the query document; Merge the filtered tokens in the query document to get the compressed prompt words.

[0043] Specifically, after obtaining the importance score of each token in the query document, the importance score of the token can be directly used to filter the tokens in the query document, and then the filtered tokens in the query document are merged to obtain the compressed prompt words.

[0044] For example, a percentage filtering method is used to adaptively select the most informative content. Specifically, the importance scores of the tokens in the query document are sorted and the tokens ranked in the bottom b% are filtered out. The value of b is determined by the target compression constraint.

[0045] The prompt word compression method provided by the embodiment of the present invention filters the tokens in the query document according to the importance scores of the tokens, obtains the compressed prompt words according to the filtered tokens, provides a concise prompt word compression method, and further improves the prompt word compression efficiency.

[0046] In some embodiments, constructing a maximum spanning tree based on self-attention between tokens in a query document includes: Take each token in the query document as a vertex, and the self-attention between tokens in the query document as the edge weight between vertices to construct an undirected graph; Taking the maximum sum of self-attention as the optimization goal, the undirected graph is optimized to obtain the constructed maximum spanning tree.

[0047] Specifically, an undirected graph is constructed with each token in the query document as a vertex and the self-attention between tokens in the query document as the edge weight between vertices.

[0048] An undirected graph can be represented as: In the formula, represents an undirected graph, represents a vertex, Represents the edge weights between vertices.

[0049] To simplify the problem, the direction of the edge is ignored and the token to token The self-attention of is equivalent to the edge weight between them. It can be formulated as: In the formula, Represents a vertex With Vertex The edge weight between Represents a vertex With Vertex The edge weight between Indicates that from token to token The self-attention between .

[0050] In a connected graph, a spanning tree is a subgraph that contains all vertices and does not form any cycles. In order to construct the maximum spanning tree, the undirected graph is optimized with the maximum sum of self-attention as the optimization goal, and the final optimized undirected graph is used as the constructed maximum spanning tree.

[0051] For example, taking the maximum sum of self-attention as the optimization goal, input the undirected graph Model, we can get the maximum spanning tree. It can be formulated as: In the formula, represents the maximum spanning tree, Represents an undirected graph.

[0052] The prompt word compression method provided by the embodiment of the present invention constructs an undirected graph by taking each token in the query document as a vertex and the self-attention between the tokens in the query document as the edge weight between the vertices; then the undirected graph is optimized with the maximum sum of self-attention as the optimization goal to obtain a constructed maximum spanning tree, which further highlights the semantic connection between the tokens and facilitates the subsequent semantic segmentation of the maximum spanning tree.

[0053] In some embodiments, before constructing the maximum spanning tree based on the self-attention between the marked tokens in the query document, the method further includes: Input the query document and query question into the causal language model to obtain causal attention; Based on the causal attention and the length of the query document, the self-attention between tokens in the query document is obtained.

[0054] Specifically, in long context scenarios, the prompt word usually consists of two parts: the query document and the query question. It can be formulated as: In the formula, Indicates prompt words, Indicates the query question. Represents a query document.

[0055] The key information in the context is usually related to the query, and a small language model can effectively capture this relationship. Since a causal language model is used, the query question is appended to the end of the query document and input into the causal language model to obtain causal attention. It can be formulated as: In the formula, represents causal attention, Represents a query document. Indicates a query question.

[0056] Based on causal attention and the length of the query document, the self-attention between tokens in the query document is obtained. The specific formula is as follows: In the formula, represents the self-attention between tokens in the query document, Indicates the first token in the query document, and c indicates the length of the query document.

[0057] In some embodiments, the prompt word compression method provided by the present invention further includes: Based on causal attention, the length of the query document and the length of the prompt word, we obtain the cross attention between the tokens in the query question and the tokens in the query document.

[0058] Specifically, In the formula, represents the cross attention between the tokens in the query question and the tokens in the query document, represents causal attention, 1 represents the first token in the query document, c represents the length of the query document, and n represents the length of the prompt word.

[0059] The cross attention between the tokens in the query question and the tokens in the query document, that is, the cross attention corresponding to the tokens in the query document.

[0060] The prompt word compression method provided by the embodiment of the present invention extracts self-attention and cross-attention, which is conducive to the subsequent use of self-attention to identify semantic units and the use of cross-attention to derive compression metrics.

[0061] The prompt word compression device provided by the present invention is described below. The prompt word compression device described below and the prompt word compression method described above can be referred to each other.

[0062] Figure 2 Schematic diagram of the structure of the prompt word compression device provided by the present invention. Figure 2 As shown, the present invention provides a prompt word compression device, comprising: A construction module 210 is used to construct a maximum spanning tree based on the self-attention between the marked tokens in the query document; the maximum spanning tree is a spanning tree with the largest sum of self-attention; A segmentation module 220 is used to perform semantic segmentation on the maximum spanning tree based on a community detection algorithm to obtain a plurality of semantic units; each of the semantic units includes a plurality of tokens; A first filtering module 230, configured to filter the plurality of semantic units according to an importance score of each semantic unit; the importance score of each semantic unit is an average of importance scores of all tokens in each semantic unit; The combining module 240 is used to combine the tokens in the filtered semantic units according to the original order to obtain compressed prompt words.

[0063] In some embodiments, the apparatus further comprises: A first determination module is used to determine the attention heads related to the context information in the large language model; A second determination module is used to determine the maximum cross attention corresponding to each token in the query document according to the cross attention corresponding to each token in the query document under different attention heads; The first acquisition module is used to use the maximum cross attention corresponding to each token in the query document as the importance score of each token in the query document.

[0064] In some embodiments, the apparatus further comprises: A second filtering module, configured to filter the tokens in the query document according to the importance score of each token in the query document; The merging module is used to merge the filtered tokens in the query document to obtain compressed prompt words.

[0065] In some embodiments, the construction module 210 is specifically used to: construct an undirected graph with each token in the query document as a vertex and the self-attention between the tokens in the query document as the edge weight between the vertices; The undirected graph is optimized with the maximum sum of self-attention as the optimization goal to obtain the constructed maximum spanning tree.

[0066] In some embodiments, the apparatus further comprises: A second acquisition module is used to input the query document and the query question into a causal language model to obtain causal attention; The third acquisition module is used to obtain the self-attention between tokens in the query document based on the causal attention and the length of the query document.

[0067] In some embodiments, the apparatus further comprises: The fourth acquisition module is used to obtain the cross attention between the token in the query question and the token in the query document based on the causal attention, the length of the query document and the length of the prompt word.

[0068] It should be noted here that the above-mentioned prompt word compression device provided by the present invention can implement all the method steps implemented by the above-mentioned method embodiment, and can achieve the same technical effect. The parts and beneficial effects that are the same as the method embodiment in this embodiment will not be described in detail here.

[0069] Figure 3 is a schematic diagram of the structure of the electronic device provided by the present invention, such as Figure 3As shown, the electronic device may include: a processor 310, a communications interface 320, a memory 330, and a communication bus 340. Among them, the processor 310, the communications interface 320, and the memory 330 complete communication with each other through the communication bus 340. The processor 310 may call the logical instructions in the memory 330 to execute the prompt word compression method, which includes: constructing a maximum spanning tree based on the self-attention between tokens marked in the query document; the maximum spanning tree is the spanning tree with the largest sum of self-attention; performing semantic segmentation on the maximum spanning tree based on a community detection algorithm to obtain multiple semantic units; each of the semantic units includes multiple tokens; filtering the multiple semantic units according to the importance score of each semantic unit; the importance score of each semantic unit is the average of the importance scores of all tokens in each semantic unit; combining the tokens in the filtered semantic units in the original order to obtain the compressed prompt word.

[0070] In addition, when the logical instructions in the above-mentioned memory 330 are implemented in the form of software functional units and sold or used as an independent product, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.

[0071] On the other hand, the present invention also provides a computer program product. The computer program product includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the prompt word compression method provided by the above-mentioned various methods. The method includes: Based on the self-attention between the marked tokens in the query document, a maximum spanning tree is constructed; the maximum spanning tree is a spanning tree with the largest sum of self-attention; semantic segmentation is performed on the maximum spanning tree based on a community detection algorithm to obtain multiple semantic units; each of the semantic units includes multiple tokens; the multiple semantic units are filtered according to the importance score of each of the semantic units; the importance score of each of the semantic units is the average of the importance scores of all tokens in each of the semantic units; the tokens in the filtered semantic units are combined in the original order to obtain compressed prompt words.

[0072] In another aspect, the present invention further provides a non-transitory computer-readable storage medium having a computer program stored thereon, and when the computer program is executed by a processor, the prompt word compression method provided by the above methods is implemented, and the method includes: Based on the self-attention between the marked tokens in the query document, a maximum spanning tree is constructed; the maximum spanning tree is a spanning tree with the largest sum of self-attention; semantic segmentation is performed on the maximum spanning tree based on a community detection algorithm to obtain multiple semantic units; each of the semantic units includes multiple tokens; the multiple semantic units are filtered according to the importance score of each of the semantic units; the importance score of each of the semantic units is the average of the importance scores of all tokens in each of the semantic units; the tokens in the filtered semantic units are combined in the original order to obtain compressed prompt words.

[0073] The device embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. Ordinary technicians in this field can understand and implement it without paying creative labor.

[0074] Through the description of the above implementation methods, those skilled in the art can clearly understand that each implementation method can be implemented by means of software plus a necessary general hardware platform, and of course, can also be implemented by hardware. Based on this understanding, the above technical solution is essentially or the part that contributes to the prior art can be embodied in the form of a software product, and the computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a disk, an optical disk, etc., including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0075] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A prompt word compression method, characterized in that: include: Construct a maximum spanning tree based on the self-attention between tokens in the query document; The maximum spanning tree is a spanning tree with the largest sum of self-attention; Performing semantic segmentation on the maximum spanning tree based on a community detection algorithm to obtain a plurality of semantic units; each of the semantic units includes a plurality of tokens; Filtering the multiple semantic units according to the importance score of each semantic unit; the importance score of each semantic unit is the average of the importance scores of all tokens in each semantic unit; The tokens in the filtered semantic units are combined in the original order to obtain compressed prompt words.

2. The prompt word compression method according to claim 1, characterized in that: Before filtering the plurality of semantic units according to the importance score of each semantic unit, the method further includes: Determine the attention heads associated with contextual information in large language models; Determine the maximum cross attention corresponding to each token in the query document according to the cross attention corresponding to each token in the query document under different attention heads; The maximum cross attention corresponding to each token in the query document is used as the importance score of each token in the query document.

3. The prompt word compression method according to claim 2, characterized in that: The method further comprises: Filtering the tokens in the query document according to the importance score of each token in the query document; The filtered tokens in the query document are merged to obtain compressed prompt words.

4. The prompt word compression method according to claim 1, characterized in that: The method of constructing a maximum spanning tree based on self-attention between tokens in the query document includes: Construct an undirected graph by taking each token in the query document as a vertex and the self-attention between the tokens in the query document as the edge weight between the vertices; The undirected graph is optimized with the maximum sum of self-attention as the optimization goal to obtain the constructed maximum spanning tree.

5. The prompt word compression method according to claim 1, characterized in that: Before constructing the maximum spanning tree based on the self-attention between the marked tokens in the query document, the method further includes: Inputting the query document and the query question into a causal language model to obtain causal attention; Based on the causal attention and the length of the query document, the self-attention between tokens in the query document is obtained.

6. The prompt word compression method according to claim 5, characterized in that: The method further comprises: Based on the causal attention, the length of the query document and the length of the prompt word, a cross attention between the tokens in the query question and the tokens in the query document is obtained.

7. A prompt word compression device, characterized in that: include: A construction module for constructing a maximum spanning tree based on the self-attention between tokens in the query document; The maximum spanning tree is a spanning tree with the largest sum of self-attention; A segmentation module, used for performing semantic segmentation on the maximum spanning tree based on a community detection algorithm to obtain a plurality of semantic units; each of the semantic units includes a plurality of tokens; A first filtering module, configured to filter the plurality of semantic units according to an importance score of each semantic unit; the importance score of each semantic unit is an average of importance scores of all tokens in each semantic unit; The combination module is used to combine the tokens in the filtered semantic units according to the original order to obtain compressed prompt words.

8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that: When the processor executes the computer program, the prompt word compression method according to any one of claims 1 to 6 is implemented.

9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the prompt word compression method as claimed in any one of claims 1 to 6 is implemented.

10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the prompt word compression method as claimed in any one of claims 1 to 6 is implemented.