Context compression method and device, electronic equipment, storage medium and program product

By acquiring and utilizing the difference information of context information to optimize the compression model, the problem of poor accuracy of context compression in existing technologies is solved, and more efficient information compression and a wider range of application scenarios are achieved.

CN121997992APending Publication Date: 2026-05-08BEIJING ZITIAO NETWORK TECH CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
BEIJING ZITIAO NETWORK TECH CO LTD
Filing Date
2026-01-04
Publication Date
2026-05-08

AI Technical Summary

Technical Problem

Existing context compression methods suffer from poor accuracy in multimodal scenarios, mainly due to significant loss of detail caused by specific training methods.

Method used

By acquiring the difference between the first and second context information, the compression model is optimized using the difference information, enabling it to learn the lost details. This allows it to fit the lost details in subsequent compression processes, ensuring accurate compression of the context information.

Benefits of technology

It improves the accuracy and completeness of contextual information compression, reduces information loss, and enhances the training speed and generalization ability of the compression model in various application scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121997992A_ABST
    Figure CN121997992A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of computers, and discloses a context compression method and device, electronic equipment, a storage medium and a program product. The method comprises the following steps: acquiring first context information and second context information corresponding to the first context information; respectively compressing the first context information and the second context information by using a compression model to obtain a first compression feature corresponding to the first context information and a second compression feature corresponding to the second context information; determining difference information between the first context information and the second context information based on the first compression feature and the second compression feature; and optimizing the compression model based on the difference information to enable the compression model to perform context compression. According to the technical scheme, it is guaranteed that the compression model can learn lost detail information, and accurate compression of context information is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of computer technology, and specifically to context compression methods, apparatus, electronic devices, storage media, and program products. Background Technology

[0002] In multimodal scenarios, compressing long contexts into a few tokens can reduce the sequence length in the question-answering (QA) process, saving on model encoding and inference overhead. However, current compression schemes employ specific training methods, making only simple modifications to task instruction diversity and loss functions, resulting in significant loss of detail and affecting the accuracy of context information compression. Summary of the Invention

[0003] This disclosure provides a context compression method, apparatus, electronic device, storage medium, and program product to address the problem of poor accuracy in context compression.

[0004] In a first aspect, this disclosure provides a context compression method, comprising: acquiring first context information and second context information corresponding to the first context information; compressing the first context information and the second context information respectively using a compression model to obtain a first compression feature corresponding to the first context information and a second compression feature corresponding to the second context information; determining the difference information between the first context information and the second context information based on the first compression feature and the second compression feature; and optimizing the compression model based on the difference information so that the compression model performs context compression.

[0005] Secondly, this disclosure provides a context compression apparatus, comprising: an acquisition module for acquiring first context information and second context information corresponding to the first context information; a first compression module for compressing the first context information and the second context information respectively using a compression model to obtain a first compression feature corresponding to the first context information and a second compression feature corresponding to the second context information; a difference information determination module for determining difference information between the first context information and the second context information based on the first compression feature and the second compression feature; and a second compression module for optimizing the compression model based on the difference information so that the compression model performs context compression.

[0006] Thirdly, this disclosure provides an electronic device, including: a memory and a processor, which are communicatively connected to each other. The memory stores computer instructions, and the processor executes the computer instructions to perform the context compression method of the first aspect or any corresponding embodiment described above.

[0007] Fourthly, this disclosure provides a computer-readable storage medium storing computer instructions for causing a computer to execute the context compression method of the first aspect or any corresponding embodiment described above.

[0008] Fifthly, this disclosure provides a computer program product, including computer instructions for causing a computer to execute the context compression method of the first aspect or any corresponding embodiment thereof.

[0009] The context compression method, apparatus, electronic device, storage medium, and program product provided in this disclosure obtain first context information and second context information that differs from the first context information. By combining first compression features corresponding to the first context information and second compression features corresponding to the second context information, the difference information between the first and second context information is determined. This difference information characterizes the detailed information lost by the second context information compared to the first context information. The difference information is then used to optimize the compression model, enabling the model to learn the lost detailed information and fit it during subsequent compression processes, thus ensuring accurate compression of the context information. Attached Figure Description

[0010] To more clearly illustrate the technical solutions in the specific embodiments of this disclosure or the prior art, the drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this disclosure. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0011] Figure 1 This is a schematic diagram illustrating an application scenario according to an embodiment of this disclosure; Figure 2 This is a schematic flowchart of a first type of context compression method according to an embodiment of the present disclosure; Figure 3 This is a schematic diagram of context compression for text difference detection according to an embodiment of the present disclosure; Figure 4 This is a schematic diagram of a second flow of the context compression method according to an embodiment of the present disclosure; Figure 5 This is a schematic diagram of a third process of the context compression method according to an embodiment of the present disclosure; Figure 6 This is a schematic diagram of context compression for feature representation difference detection according to an embodiment of the present disclosure; Figure 7 This is a structural block diagram of a context compression apparatus according to an embodiment of the present disclosure; Figure 8This is a schematic diagram of the hardware structure of an electronic device according to an embodiment of this disclosure. Detailed Implementation

[0012] To make the objectives, technical solutions, and advantages of the embodiments of this disclosure clearer, the technical solutions of the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this disclosure, and not all embodiments. Based on the embodiments of this disclosure, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this disclosure.

[0013] It is understood that before using the technical solutions disclosed in the various embodiments of this disclosure, users should be informed of the types, scope of use, and usage scenarios of the personal information involved in this disclosure in an appropriate manner in accordance with relevant laws and regulations, and user authorization should be obtained.

[0014] For example, upon receiving a user's active request, a prompt message is sent to the user to explicitly inform them that the requested operation will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the software or hardware, such as the electronic device, application, server, or storage medium performing the operations of this disclosed technical solution, based on the prompt message.

[0015] As an optional but non-limiting implementation, in response to a user's active request, sending a prompt message to the user can be done via a pop-up window, where the prompt message can be presented in text format. Furthermore, the pop-up window can also include a selection control allowing the user to choose "agree" or "disagree" to provide personal information to the electronic device.

[0016] It is understood that the above notification and user authorization process are merely illustrative and do not constitute a limitation on the implementation of this disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of this disclosure.

[0017] It is understood that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) shall comply with the requirements of relevant laws, regulations and related provisions.

[0018] The terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature. In the description of this disclosure, "a plurality of" means two or more, unless otherwise expressly specified.

[0019] In multimodal scenarios, compressing long contexts into a few tokens can reduce the sequence length during question-and-answer (QA) processes. Currently, the main compression schemes used include: (1) In-Context Autoencoder (ICAE): Compresses long context into “memory slots”, adds about 1% of parameters. Based on Llama-7b, 128 memory slots can recover 512-length context, which improves the long text processing capability of large language models and reduces inference costs.

[0020] (2) 500x Compression Framework: Compresses a hint of about 500 tokens to a minimum of 1 token, with a compression ratio ranging from 6x to 480x, while retaining 62.26% to 72.89% of the original large language model, thus verifying the high compressibility of natural language hints.

[0021] (3) Retrieval Enhanced Generative Compression (xRAG): Based on modal fusion, it replaces the original document text with only one token. It achieves an average improvement of over 10% in X knowledge-intensive tasks, reduces the total number of floating point operations (FLOPs) by 3.53 times, reduces the effective inference time on the GPU, and improves inference time efficiency by 1.64 times.

[0022] (4) Efficient compression and distillation framework for long context LLM (LongLLMLingua): Compresses context in discrete space through problem-aware coarse-fine granular compression, document reordering, dynamic compression ratio and subsequence recovery.

[0023] However, the above methods all use specific training methods and only make some minor modifications to the task instruction diversity and loss function. They still have a large loss of detail, and the research on these methods is relatively limited, often only conducted on specific tasks.

[0024] Based on this, this disclosure uses difference information to characterize the details lost by the second context information compared to the first context information, and uses the difference information to optimize the compression model so that the compression model can learn the lost details information. This eliminates the need for the compression model to always retain the ability to follow general instructions, and the training speed is faster. As a result, the lost details information can be fitted in the subsequent compression process, ensuring accurate compression of the context information.

[0025] As one optional application scenario of this disclosure embodiment, such as Figure 1 As shown, application 101 is installed in electronic device 110, and user 130 can interact with application 101 through electronic device 110 and / or access device of electronic device 110.

[0026] For example, application 101 can be arbitrary, and it includes a compression model and a decompression model to provide question-and-answer related services. For instance, application 101 could be a question-and-answer interactive application, etc. Figure 1 In the application scenario shown, if application 101 is active, electronic device 110 can display the interface 102 of application 101. Interface 102 may include various pages that application 101 can provide, such as interactive pages, settings pages, query pages, etc.

[0027] In some embodiments, electronic device 110 is communicatively connected to server 120 to provide services to application 101. Electronic device 110 may be a mobile terminal, fixed terminal, or portable terminal, etc., including but not limited to mobile phones, desktop computers, laptop computers, multimedia tablets, e-book devices, gaming devices, or any combination thereof, including accessories and peripherals of these devices, or any combination thereof. In some embodiments, electronic device 110 may also support any type of interface, and server 120 may be various types of computing systems or servers capable of providing computing power, including but not limited to mainframes, edge computing nodes, computing devices in cloud environments, etc.

[0028] It should be noted that, Figure 1 This is merely an example of an application scenario and does not limit the scope of protection of this disclosure.

[0029] The embodiments of this disclosure will now be described with reference to the accompanying drawings. It should be understood that the pages shown in the drawings are merely examples, and various page designs are possible in practice. The various graphic elements on the page may have different arrangements and different visual representations, one or more elements may be omitted or replaced, and one or more other elements may also be present; no limitations are imposed on the embodiments described in this disclosure. Furthermore, the embodiments are primarily described below with reference to electronic device 110. It should be understood that the actions described with respect to electronic device 110 can be performed by application 101 on electronic device 110, or can be performed by application 101 in conjunction with its server (e.g., server 120).

[0030] According to an embodiment of this disclosure, a context compression method embodiment is provided. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Furthermore, although a logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than that shown here.

[0031] This embodiment provides a context compression method that can be used in the aforementioned electronic devices, such as computers and tablets. Figure 2This is a flowchart of a context compression method according to an embodiment of the present disclosure, such as... Figure 2 As shown, the process includes the following steps: Step S201: Obtain first context information and second context information corresponding to the first context information.

[0032] The first context information is the original context information input to the compression model, such as text information, visual images, and multimodal information containing text and images. The second context information is information that differs from the first context information. Specifically, the second context information can be obtained by recovering the compressed first context information using the decompression model, or it can be obtained by modifying the first context information. No specific limitation is made here.

[0033] In a specific example, a compression command and corresponding first context information are input through an interactive page provided by an application in an electronic device. After receiving the compression command and the first context information, the application inputs the first context to the compression model according to the compression command. Accordingly, the compression model can obtain the first context information.

[0034] Based on the first context information obtained by the compression model, the first context information is processed accordingly to obtain second context information that differs from the first context information.

[0035] Step S202: Compress the first context information and the second context information using a compression model to obtain the first compressed feature corresponding to the first context information and the second compressed feature corresponding to the second context information.

[0036] like Figure 3 As shown, after receiving the first context information, the compression model performs vectorized compression processing on the internal structure of the first context information according to the compression instruction, thereby obtaining the first compressed feature of the first context information at the feature representation level.

[0037] Similarly, by using the same compression model to vectorize and compress the content of the second context information, we can obtain the second compressed feature of the second context information at the feature representation level.

[0038] It should be noted that for the compression of multimodal context information, compression can be completed simply by replacing the second context information (i.e., the visual image in the multimodal scene) with the image caption text.

[0039] Step S203: Based on the first compression feature and the second compression feature, determine the difference information between the first context information and the second context information.

[0040] The difference information represents the content difference between the first and second contextual information, that is, the information lost by the second contextual information compared to the first contextual information. Based on the first and second compressed features, a decompression model is used to detect the feature differences between the first and second compressed features, and the difference information is predicted based on the feature differences, thereby obtaining the corresponding difference information.

[0041] Step S204: Optimize the compression model based on the difference information so that the compression model can perform context compression.

[0042] Since the difference information can characterize the information lost by the second context information compared to the first context information, it can be used to optimize the compression model. This allows the compression model to continuously learn, reducing information loss during compression and resulting in an optimized compression model. The optimized compression model is then used to compress the received new context information.

[0043] Furthermore, by utilizing the downstream decompression model to perform decompression tasks associated with context information, the associated content matching the context information can be obtained. For example, if the context information expresses the generation process of a certain video, and the decompression instruction corresponding to the decompression model is "what needs to be paid attention to during the video generation process," then the decompression model can decompress the context information compressed by the compression model and summarize the response content matching the decompression instruction.

[0044] The context compression method provided in this embodiment obtains first context information and second context information that differs from the first context information. It then combines a first compression feature corresponding to the first context information with a second compression feature corresponding to the second context information to determine the difference information between the first and second context information. This difference information characterizes the detailed information lost by the second context information compared to the first context information. The compression model is then optimized using this difference information so that it can learn the lost detailed information and fit it during subsequent compression processes, ensuring accurate compression of the context information.

[0045] This embodiment provides a context compression method that can be used in the aforementioned electronic devices, such as computers and tablets. Figure 4 This is a flowchart of a context compression method according to an embodiment of the present disclosure, such as... Figure 4 As shown, the process includes the following steps: Step S301: Obtain first context information and second context information corresponding to the first context information.

[0046] Specifically, step S301 includes: Step S3011: Obtain first context information. For details, please refer to the relevant descriptions of the corresponding steps in the embodiments shown above, which will not be repeated here.

[0047] Step S3012: Compress the first context information using a compression model to obtain the text features corresponding to the first context.

[0048] like Figure 3 As shown, the first context information in text form is input into the compression model, and the compression model itself is used for compression processing to obtain the text features of the first context information at the feature representation level inside the model.

[0049] Step S3013: Using the decompression model corresponding to the compression model, information is recovered from the text features to obtain the second context information.

[0050] The decompression model corresponds to the compression model. The text features output by the compression model are input into the decompression model, and the decompression parameters are used to recover the text features, resulting in the second context information recovered by the decompression model, such as... Figure 3 As shown.

[0051] Step S302: Compress the first context information and the second context information using a compression model to obtain the first compressed feature corresponding to the first context information and the second compressed feature corresponding to the second context information.

[0052] The first compression feature here is the text feature mentioned above. After compressing the first context information using the compression model to obtain the corresponding first compressed feature, the decompression model uses this first compressed feature to recover the corresponding second context information. Then, the second context information is input into the compression model again for compression processing to obtain the corresponding second compressed feature, such as... Figure 3 As shown.

[0053] It should be noted that if the first compression feature is the same as the second compression feature, it means that the first context information and the second context information are completely consistent; if the first compression feature is different from the second compression feature, it means that the first context information and the second context information are different.

[0054] Step S303: Based on the first compression feature and the second compression feature, determine the difference information between the first context information and the second context information.

[0055] Specifically, the difference information includes content difference information and feature difference information. Content difference information refers to the difference between the first context information and the second context information at the level of direct comparison of text content; feature difference information refers to the difference between the first context information and the second context information at the level of feature representation within the model. Accordingly, step S303 includes: Step S3031: Based on the text content represented by the first context information and the second context information, obtain the content difference information between the first context information and the second context information.

[0056] The first context information and the second context information are input into the decompression model so that the decompression model can compare the differences in text content between the first context information and the second context information and output the corresponding content difference information.

[0057] Step S3032: Based on the text features represented by the first compression feature and the second compression feature, obtain the feature difference information between the first compression feature and the second compression feature.

[0058] While comparing the differences in content, the first compressed feature and the second compressed feature are input into the decompression model so that the decompression model can compare the feature differences of the first context information and the second context information at the feature representation level and output the corresponding feature difference information.

[0059] Step S304: Optimize the compression model based on the difference information so that the compression model can perform context compression.

[0060] The content difference information is used as the training label for the optimized compression model, and the feature difference information is used as the model prediction value for the optimized compression model. The second context information recovered by the decompression model is used as the hard negative example sample (i.e., a negative sample that is highly similar to the first context information but has key differences). A corresponding language modeling task is constructed. The context representation ability of the compression model is optimized through the language modeling task to obtain the optimized compression model, and the optimized compression model is used for context compression processing.

[0061] It should be noted that this compression model can perform self-learning iterations during the execution of context compression tasks to improve its context representation capabilities.

[0062] The context compression method provided in this embodiment compresses first context information to obtain corresponding text features, then restores information from these text features to generate corresponding second context information. This ensures that the second context information is similar to the first context information, facilitating subsequent difference detection. By acquiring content and feature difference information between the first and second context information, the compression model is optimized to retain more and more details during compression, reducing the loss of detail in the text-based context information during compression.

[0063] This embodiment provides a context compression method that can be used in the aforementioned electronic devices, such as computers and tablets. Figure 5 This is a flowchart of a context compression method according to an embodiment of the present disclosure, such as... Figure 5 As shown, the process includes the following steps: Step S401: Obtain first context information and second context information corresponding to the first context information.

[0064] Specifically, step S401 includes: Step S4011: Obtain first context information. For details, please refer to the relevant descriptions of the corresponding steps in the embodiments shown above, which will not be repeated here.

[0065] Step S4012: Compress the first context information using a compression model to obtain the spatial feature token set corresponding to the first context.

[0066] The spatial feature token set is a spatial feature representation of the first context information. The spatial feature token set includes multiple feature tokens (tokens) obtained by encoding the first context information.

[0067] like Figure 6 As shown, the first context information is input into the compression model, and the compression model's own compression parameters are used to compress the first context information to obtain a spatial feature token set composed of multiple feature tokens. For example, this spatial feature token set can be a token sequence containing the first context information.

[0068] Step S4013: Obtain the token to be modified from the spatial feature token set.

[0069] The token to be modified is one that violates the model's own generation probability distribution when the compression model performs a compression task. When reconstructing information using the decompression model, it is also easier to lose the details represented by the token to be modified if there is insufficient supplementary information.

[0070] In some optional implementations, step S4013 above includes: Step a1: Based on the prediction strategy corresponding to the compression model, select the first tokens with the lowest prediction generation probability from the spatial feature token set.

[0071] Step a2: The first token is identified as the token to be modified.

[0072] The prediction strategy is the method by which the compression model predicts the generated tokens during the compression task. This prediction strategy is used to select the token with the lowest prediction probability. k One token is used as the token to be modified.

[0073] Specifically, assuming the compression model generates the first i The predicted probability of each token is expressed as follows: , in, S represents the compression model; S represents the token sequence containing context information, i.e., the spatial feature token set.

[0074] Based on the calculated prediction probabilities, select the token with the lowest prediction probability from the spatial feature token set. k One token is used as the token to be modified, as follows:

[0075]

[0076] in, r The scaling factor is between 0 and 1. This indicates rounding up to the nearest integer.

[0077] By selecting the token to be modified from the spatial feature tokens corresponding to the first context information, the easily lost details information of the compressed model can be obtained, which facilitates subsequent optimization of the compressed model.

[0078] Step S4014: Based on the position of the token to be modified in the spatial feature token set, adjust the token to be modified into the target token.

[0079] The target token is a pseudo-token located at the position of the token to be modified. Specifically, the tokens in the spatial feature token set are arranged in a corresponding order, and each token has a corresponding token identifier. Based on the position of the token to be modified in the spatial feature token set, the target token corresponding to the position of the token to be modified is predicted, given other tokens.

[0080] In some optional implementations, step S4014 above includes: Step b1: Based on the positional relationship between the token to be modified and the unmodified token, token sampling is performed on the position of the token to be modified to obtain the target token.

[0081] Step b2: Replace the token to be modified with the target token.

[0082] The target token is sampled given the other unmodified tokens, and the specific determination of the target token can be represented as follows:

[0083] in, Indicates the first i The target token for each location.

[0084] After obtaining the target token corresponding to each token to be modified, replace each token to be modified with the corresponding target token.

[0085] Given the remaining non-modified tokens, sample the target token corresponding to the position of the token to be modified to generate pseudo-context information that matches the first context information, which facilitates accurate optimization of the compression model.

[0086] Step S4015: Generate second context information based on the fusion result of the target token and the unmodified token.

[0087] The target token and the unmodified token are sorted in the corresponding order to form the modified pseudo-context information, which is the second context information.

[0088] Step S402: The first context information and the second context information are compressed using a compression model to obtain a first compressed feature corresponding to the first context information and a second compressed feature corresponding to the second context information. For details, please refer to the relevant descriptions of the steps in the embodiments shown above; they will not be repeated here.

[0089] Step S403: Based on the first compression feature and the second compression feature, determine the difference information between the first context information and the second context information.

[0090] The topic drift problem exists in the long context generation process, indicating that although the compressed features of the compressed model already contain all the necessary information, the recovered context may still contain additional / lost information due to autoregression. Therefore, difference detection can be performed in the representation space here.

[0091] Specifically, the difference information includes a first feature difference and a second feature difference. Accordingly, step S403 above includes: Step S4031: Perform differential processing on the first compressed feature and the second compressed feature to obtain the first feature difference.

[0092] The first compressed feature and the second compressed feature are input into the difference model for difference processing to capture the difference between the first compressed feature and the second compressed feature, and are marked in the first context information.

[0093] Specifically, the first compression feature can be expressed as: The second compression feature can be expressed as The difference between the first and second compressed features is calculated using a difference module and then expressed in the representation space. The first compressed feature can be represented in the representation space as:

[0094]

[0095] The second compression feature can be represented in the representation space as:

[0096]

[0097] in, and Represents the projection operator. ctx represents the representation corresponding to the contextual difference information.

[0098] The difference between the second compressed feature corresponding to the second context information and itself is detected on the first compressed feature corresponding to the first context information, that is:

[0099]

[0100] in, This represents the temperature coefficient.

[0101] Step S4032: Perform differential processing on the second compressed feature and the first compressed feature to obtain the second feature difference.

[0102] The difference between the first compressed feature corresponding to the first context information and itself is detected on the second compressed feature corresponding to the second context information, that is:

[0103]

[0104] in, This represents the temperature coefficient.

[0105] Step S404: Optimize the compression model based on the difference information so that the compression model can perform context compression.

[0106] Specifically, step S404 includes: Step S4041: Obtain the modeling loss information corresponding to the compressed model.

[0107] The modeling loss information is determined based on the probability distribution of the natural language being modeled, i.e., the probability of predicting the next word token given the context. This modeling loss information is used to quantify the difference between the predicted probabilities of the compression model and the true labels, driving the optimization of the compression model.

[0108] In a specific example, modeling loss information can be based on maximizing the log-likelihood, such as the negative log-likelihood.

[0109] Step S4042: Obtain first difference loss information based on the first feature difference information.

[0110] Specifically, based on the first feature difference information, the corresponding first difference loss information is determined as follows: .

[0111] Step S4043: Obtain second difference loss information based on the second feature difference information.

[0112] Specifically, based on the second feature difference information, the corresponding second difference loss information is determined as follows: .

[0113] Step S4044: Obtain target loss information based on modeling loss information, first difference loss information, and second difference loss information.

[0114] The modeling loss information, the first difference loss information, and the second difference loss information are fused to narrow the feature distance of semantically similar texts / images and widen the feature distance of semantically similar texts / images, thereby obtaining the corresponding target loss information.

[0115] In some optional implementations, step S4044 above includes: Step c1: Determine the target difference loss information based on the fusion result of the first difference loss information and the second difference loss information.

[0116] Step c2: Overlay the target difference loss information with the modeling loss information to obtain the target loss information.

[0117] The first and second difference loss information are fused to obtain the corresponding difference loss information: .

[0118] The difference loss information is superimposed with the modeling loss information to obtain the target loss information: .

[0119] in, Model loss information.

[0120] By fusing the first and second difference loss information, the corresponding difference loss information is obtained. Then, the modeling loss information is superimposed on the difference loss information to obtain the complete target loss information, so that the target loss information can accurately represent the difference features before and after compression.

[0121] Step S4045: Optimize the compression model using the target loss information.

[0122] Under the constraint of target loss information, the parameters of the compression model are optimized to minimize the target loss information, ensuring that the compression model can compress the first context information without losing information.

[0123] The context compression method provided in this embodiment selects the token to be modified from the spatial feature token set corresponding to the first context information. Based on the position of the token to be modified in the spatial feature token set, it modifies it into a target token different from the original content, ensuring the similarity and difference between the second context information and the first context information, which helps to improve the training speed of the compression model. By detecting the differences between the first and second compressed features and themselves, the method ensures that both the first and second compressed features can be recognized, thus improving the accuracy of feature recognition.

[0124] By fusing modeling loss information, first difference loss information corresponding to the first feature difference information, and second difference loss information corresponding to the second feature difference information, target loss information is constructed. Optimizing the compressed model according to the target loss information significantly improves its ability to represent contextual information, making its contextual information representation capabilities not limited to specific tasks and broadening its application scenarios.

[0125] As a specific application embodiment of this disclosure, compared with the use of open-source methods, experiments were conducted on multiple publicly available academic reading comprehension datasets using the technical solution of this disclosure. The results show that the context compression method adopted in this disclosure achieves better results, with an average F1 score increase of 10.66 percentage points and an exact match (EM) increase of 3.23 percentage points. This indicates that the compression model has optimized both accuracy and completeness when performing context information compression tasks. Meanwhile, ablation experiments also demonstrate that the discriminant detection task brings positive benefits to compression.

[0126] This embodiment also provides a context compression device for implementing the above embodiments and preferred embodiments; details already described will not be repeated. As used below, the term "module" can refer to a combination of software and / or hardware that performs a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, hardware implementations, or combinations of software and hardware, are also possible and contemplated.

[0127] This embodiment provides a context compression device, such as... Figure 7 As shown, it includes: The acquisition module 501 is used to acquire first context information and second context information corresponding to the first context information.

[0128] The first compression module 502 is used to compress the first context information and the second context information respectively using a compression model to obtain the first compressed feature corresponding to the first context information and the second compressed feature corresponding to the second context information.

[0129] The difference information determination module 503 is used to determine the difference information between the first context information and the second context information based on the first compression feature and the second compression feature.

[0130] The second compression module 504 is used to optimize the compression model based on the difference information so that the compression model can perform context compression.

[0131] In some optional implementations, the acquisition module 501 includes: The first compression unit is used to compress the first context information using a compression model to obtain the text features corresponding to the first context.

[0132] The recovery unit is used to recover information from text features using the decompression model corresponding to the compression model, and obtain the second context information.

[0133] In some optional implementations, the difference information includes content difference information and feature difference information. Accordingly, the difference information determination module 503 includes: The content difference determination unit is used to obtain content difference information between the first context information and the second context information based on the text content represented by the first context information and the second context information.

[0134] The feature difference determination unit is used to obtain feature difference information between the first compressed feature and the second compressed feature based on the text features represented by the first compressed feature and the second compressed feature.

[0135] In some optional implementations, the acquisition module 501 includes: The feature token acquisition unit is used to compress the first context information using a compression model to obtain the spatial feature token set corresponding to the first context.

[0136] The modified token acquisition unit is used to acquire the token to be modified from the spatial feature token set.

[0137] The adjustment unit is used to adjust the token to be modified into the target token based on the position of the token to be modified in the spatial feature token set.

[0138] The fusion unit is used to generate second context information based on the fusion result of the target token and the unmodified token.

[0139] In some optional implementations, the modified token acquisition unit includes: The first token determination subunit is used to select multiple first tokens with the lowest prediction generation probability from the spatial feature token set based on the prediction strategy corresponding to the compression model.

[0140] The Modify Token Determination Subunit is used to determine the first token as the token to be modified.

[0141] In some optional implementations, the adjustment unit includes: The sampling subunit is used to sample the position of the token to be modified based on the positional relationship between the token to be modified and the unmodified token, so as to obtain the target token.

[0142] The replacement sub-unit is used to replace the token to be modified with the target token.

[0143] In some optional implementations, the difference information includes a first feature difference and a second feature difference. Accordingly, the difference information determination module 503 includes: The first difference unit is used to perform difference processing on the first compressed feature and the second compressed feature to obtain the first feature difference.

[0144] The second difference unit is used to perform difference processing on the second compressed feature and the first compressed feature to obtain the second feature difference.

[0145] In some alternative implementations, the second compression module 504 includes: The modeling loss acquisition unit is used to acquire the modeling loss information corresponding to the compressed model.

[0146] The first difference loss acquisition unit is used to acquire first difference loss information based on the first feature difference information.

[0147] The second difference loss acquisition unit is used to acquire second difference loss information based on the second feature difference information.

[0148] The target loss determination unit is used to obtain target loss information based on modeling loss information, first difference loss information, and second difference loss information.

[0149] The optimization unit is used to optimize the compression model using the target loss information.

[0150] In some optional implementations, the target loss determination unit includes: The difference loss fusion subunit is used to determine the target difference loss information based on the fusion result of the first difference loss information and the second difference loss information.

[0151] The loss superposition subunit is used to superimpose the target difference loss information with the modeling loss information to obtain the target loss information.

[0152] The context compression apparatus provided in this disclosure can execute the context compression method provided in any embodiment of this disclosure, and has the corresponding functional modules and beneficial effects of the method. By acquiring first context information and second context information that differs from the first context information, and combining the first compression feature corresponding to the first context information and the second compression feature corresponding to the second context information, the difference information between the first and second context information is determined. This difference information can characterize the detailed information lost by the second context information compared to the first context information. The compression model is optimized using the difference information so that the compression model can learn the lost detailed information and fit the lost detailed information in subsequent compression processes, ensuring accurate compression of the context information.

[0153] Further functional descriptions of the above modules and units are the same as those in the corresponding embodiments described above, and will not be repeated here.

[0154] Figure 8 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this disclosure.

[0155] The following is a detailed reference. Figure 8 The diagram illustrates a structural schematic suitable for implementing an electronic device according to embodiments of the present disclosure. The electronic device may include a processor (e.g., a central processing unit, graphics processor, etc.) 801, which can perform various appropriate actions and processes according to a program stored in read-only memory (ROM) 802 or a program loaded from memory 808 into random access memory (RAM) 803. The RAM 803 also stores various programs and data required for the operation of the electronic device. The processor 801, ROM 802, and RAM 803 are interconnected via a bus 804. An input / output (I / O) interface 805 is also connected to the bus 804.

[0156] Typically, the following devices can be connected to I / O interface 805: input devices 806 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 807 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; memory devices 808 including, for example, magnetic tapes, hard disks, etc.; and communication devices 809. Communication device 809 allows electronic devices to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 8 Electronic devices with various devices are shown, but it should be understood that it is not required to implement or have all of the devices shown, and more or fewer devices may be implemented or have instead.

[0157] In particular, according to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device 809, or installed from a memory 808, or installed from a ROM 802. When the computer program is executed by the processor 801, it performs the functions defined in the context compression method of embodiments of this disclosure.

[0158] Figure 8 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments disclosed herein.

[0159] This disclosure also provides a computer-readable storage medium in which the methods described in this disclosure can be implemented in hardware or firmware, or implemented as recordable on a storage medium, or implemented as computer code originally stored on a remote storage medium or a non-transitory machine-readable storage medium and subsequently stored on a local storage medium after being downloaded over a network. Thus, the methods described herein can be processed by software stored on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. The storage medium can be a magnetic disk, optical disk, read-only memory, random access memory, flash memory, hard disk, or solid-state drive, etc.; further, the storage medium may also include combinations of the above types of memory. It is understood that computers, processors, microprocessor controllers, or programmable hardware include storage components capable of storing or receiving software or computer code, which, when accessed and executed by the computer, processor, or hardware, implements the context compression method shown in the above embodiments.

[0160] A portion of this disclosure can be applied to computer program products, such as computer program instructions, which, when executed by a computer, can invoke or provide methods and / or technical solutions according to this disclosure through the operation of the computer. Those skilled in the art will understand that the forms in which computer program instructions exist in a computer-readable medium include, but are not limited to, source files, executable files, and installation package files. Accordingly, the ways in which computer program instructions are executed by a computer include, but are not limited to: the computer directly executing the instructions; the computer compiling the instructions and then executing the corresponding compiled program; the computer reading and executing the instructions; or the computer reading and installing the instructions and then executing the corresponding installed program. Here, the computer-readable medium can be any available computer-readable storage medium or communication medium accessible to a computer.

[0161] Although embodiments of the present disclosure have been described in conjunction with the accompanying drawings, those skilled in the art can make various modifications and variations without departing from the spirit and scope of the present disclosure, and such modifications and variations all fall within the scope defined by the appended claims.

Claims

1. A context compression method, characterized in that, The method includes: Obtain first context information and second context information corresponding to the first context information; The first context information and the second context information are compressed using a compression model to obtain a first compressed feature corresponding to the first context information and a second compressed feature corresponding to the second context information. Based on the first compression feature and the second compression feature, determine the difference information between the first context information and the second context information; The compression model is optimized based on the difference information to enable contextual compression.

2. The method according to claim 1, characterized in that, Obtaining the second context information corresponding to the first context information includes: The compression model is used to compress the first context information to obtain the text features corresponding to the first context information; Using the decompression model corresponding to the compression model, information is recovered from the text features to obtain the second context information.

3. The method according to claim 1 or 2, characterized in that, The step of determining the difference information between the first context information and the second context information based on the first compression feature and the second compression feature includes: Based on the text content represented by the first context information and the second context information, content difference information between the first context information and the second context information is obtained; Based on the text features represented by the first compression feature and the second compression feature, feature difference information between the first compression feature and the second compression feature is obtained; The difference information includes the content difference information and the feature difference information.

4. The method according to claim 1, characterized in that, The step of obtaining the second context information corresponding to the first context information includes: The compression model is used to compress the first context information to obtain the spatial feature token set corresponding to the first context; Obtain the token to be modified from the spatial feature token set; Based on the position of the token to be modified in the spatial feature token set, the token to be modified is adjusted to the target token; The second context information is generated based on the fusion result of the target token and the unmodified token.

5. The method according to claim 4, characterized in that, The step of obtaining the token to be modified from the spatial feature token set includes: Based on the prediction strategy corresponding to the compression model, select multiple first tokens with the lowest prediction generation probability from the spatial feature token set; The first token is identified as the token to be modified.

6. The method according to claim 4, characterized in that, The step of adjusting the token to be modified into the target token based on the position of the token to be modified in the spatial feature token set includes: Based on the positional relationship between the token to be modified and the unmodified token, token sampling is performed on the position of the token to be modified to obtain the target token; Replace the token to be modified with the target token.

7. The method according to any one of claims 4-6, characterized in that, The step of determining the difference information between the first context information and the second context information based on the first compression feature and the second compression feature includes: The first compressed feature and the second compressed feature are differentially processed to obtain the first feature difference; The second compressed feature is compared with the first compressed feature to obtain the second feature difference; The difference information includes the first feature difference and the second feature difference.

8. The method according to claim 7, characterized in that, The optimization of the compression model based on the difference information includes: Obtain the modeling loss information corresponding to the compressed model; Based on the first feature difference information, obtain the first difference loss information; Based on the second feature difference information, obtain the second difference loss information; Based on the modeling loss information, the first difference loss information, and the second difference loss information, the target loss information is obtained; The compression model is optimized using the target loss information.

9. The method according to claim 8, characterized in that, The step of obtaining target loss information based on the modeling loss information, the first difference loss information, and the second difference loss information includes: Based on the fusion result of the first difference loss information and the second difference loss information, the target difference loss information is determined; The target difference loss information is superimposed with the modeling loss information to obtain the target loss information.

10. A context compression device, characterized in that, The device includes: The acquisition module is used to acquire first context information and second context information corresponding to the first context information; The first compression module is used to compress the first context information and the second context information respectively using a compression model to obtain a first compression feature corresponding to the first context information and a second compression feature corresponding to the second context information. The difference information determination module is used to determine the difference information between the first context information and the second context information based on the first compression feature and the second compression feature; The second compression module is used to optimize the compression model based on the difference information, so that the compression model can perform context compression.

11. An electronic device, characterized in that, include: A memory and a processor are communicatively connected, the memory storing computer instructions, and the processor executing the computer instructions to perform the context compression method of any one of claims 1 to 9.

12. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions for causing a computer to perform the context compression method according to any one of claims 1 to 9.

13. A computer program product, characterized in that, Includes computer instructions for causing a computer to perform the context compression method according to any one of claims 1 to 9.