A text compression method, a text decompression method, a model training method, an apparatus, and a device

Through deep learning technology and q-former quantization processing module, text features are extracted and token sequences are generated, which solves the problems of low text compression rate and poor decompression effect in the existing technology, and achieves efficient text compression and high-quality restoration effects.

CN119250020BActive Publication Date: 2025-05-30SHANGHAI SOULGATE TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411794658.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-09
Publication Date
2025-05-30
Estimated Expiration
2044-12-09

AI Technical Summary

Technical Problem

When existing text compression technology deals with long text and complex language structures, the compression rate is low and the decompression effect is poor, making it difficult to effectively process text content with rich context and high language dependencies.

Method used

Deep learning technology is used combined with q-former quantization processing module to extract text features through feature extraction modules (such as GPT or XLNet), and the embedded vector sequence is quantized using q-former module to generate token sequences to achieve text compression.

Benefits of technology

It significantly improves the efficiency and compression rate of text compression, while ensuring high-quality restoration capabilities of text information, especially in scenarios where long text and complex language structures are handled.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119250020B_ABST
    Figure CN119250020B_ABST
Patent Text Reader

Abstract

This application relates to the technical field of text processing, and particularly to a text compression method, a text decompression method, a model training method, a device, and a device. The text compression method includes: obtaining a target text to be compressed, inputting the target text to be compressed into the feature extraction module of a text compression model, using deep learning technology to capture the semantic and syntactic features of the text, and representing the extracted key information in the form of an embedded vector sequence; using a q-former quantization processing module to perform quantization processing on the embedded vector sequence, converting the continuous embedded vector sequence into a discrete token sequence, thereby achieving effective compression of text information. By adopting the form of combining the feature extraction module with the q-former structure, it is possible to significantly reduce the redundancy of text data while maintaining the ability to restore the original text with high quality.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of text processing, and in particular, to a text compression method, a text decompression method, a model training method, a device, and a device. Background Art

[0002] With the rapid development of information technology, the global data volume has shown an explosive growth. Against this background, text data, as an important form of information transmission and knowledge storage, its processing and compression technologies have become a research hotspot.

[0003] Conventional text compression technologies, such as Huffman coding and Lempel-Ziv-Welch (LZW) coding, achieve compression by analyzing the occurrence frequencies of characters in the text to construct an optimal prefix code or through string matching. For example, Huffman coding constructs a binary tree based on character frequencies and assigns a unique variable-length code to each character to achieve compression. The LZW algorithm, on the other hand, constructs a string dictionary and replaces repeatedly occurring strings with indices in the dictionary to achieve compression.

[0004] However, with the explosion of text data volume and the complexity of language structures, the compression rate of conventional text compression technologies is low and the decompression effect is poor. Summary of the Invention

[0005] The purpose of this application is to provide a text compression method, a text decompression method, a model training method, a device, and a device, which can improve the compression rate and the decompression effect.

[0006] In the first aspect, a text compression method is provided, including:

[0007] Obtain a target text to be compressed;

[0008] Obtain a text compression model, where the text compression model includes: a feature extraction module, a q-former quantization processing module;

[0009] Input the target text to be compressed into the text compression model, use the feature extraction module to perform feature extraction on the target text to be compressed to obtain an embedded vector sequence; use the q-former quantization processing module to perform quantization processing on the embedded vector sequence to obtain a token sequence, so as to achieve text compression.

[0010] In a preferred example of this application, it can be further configured that: the text compression model further includes: a preprocessing module;

[0011] Before using the feature extraction module to perform feature extraction on the target text to be compressed to obtain an embedded vector sequence, it further includes:

[0012] Use the preprocessing module to preprocess the target text to be compressed to obtain the processed text;

[0013] Correspondingly, using the feature extraction module to perform feature extraction on the target text to be compressed to obtain an embedded vector sequence, including:

[0014] Use the feature extraction module to perform feature extraction on the processed text to obtain an embedded vector sequence.

[0015] In a preferred example of the present application, it can be further configured that: the feature extraction module includes: GPT or XLNet.

[0016] In a preferred example of the present application, it can be further configured that: using the q-former quantization processing module to perform quantization processing on the embedded vector sequence to obtain a token sequence to achieve text compression, including:

[0017] Obtain the quantization granularity; based on the quantization granularity, use the q-former quantization processing module to perform quantization processing on the embedded vector sequence to obtain a token sequence to achieve text compression.

[0018] In a second aspect, a text decompression method is provided, including:

[0019] Obtain a token sequence and a request language, where the token sequence is obtained by the text compression method described in any item of the first aspect;

[0020] Obtain a large language model for decompression;

[0021] Input the token sequence and the request language into the large language model to use the large language model to restore the token sequence according to the request language to obtain an output text.

[0022] In a third aspect, a model training method is provided, including:

[0023] Obtain a plurality of training original texts, a request language, and token sequences respectively corresponding to the plurality of training original texts;

[0024] Obtain a training model, where the training model includes a compression training model and a large language training model, and the compression training model includes: a feature extraction training module and a q-former quantization processing training module;

[0025] Input the original training text into the compression training model, and use the feature extraction training module to extract features from the original training text to obtain a training embedded vector sequence; use the q-former quantization processing training module to perform quantization processing on the training embedded vector sequence to obtain a training token sequence;

[0026] Input the training token sequence and the request language into the large language training model, and use the large language training model to restore the training token sequence according to the request language to obtain a training output text;

[0027] Iteratively train the training model according to the training output text and the original training text to obtain a text processing model, where the text processing model includes: a text compression model and a large language model, the text compression model is used to compress text, and the large language model is used to decompress the token sequence obtained from the compressed text.

[0028] In a fourth aspect, a text compression device is provided, including:

[0029] A first acquisition module, configured to acquire a target text to be compressed; acquire a text compression model, where the text compression model includes: a feature extraction module and a q-former quantization processing module;

[0030] A compression module, configured to input the target text to be compressed into the text compression model, and use the feature extraction module to extract features from the target text to be compressed to obtain an embedded vector sequence; use the q-former quantization processing module to perform quantization processing on the embedded vector sequence to obtain a token sequence, so as to implement text compression.

[0031] In a fifth aspect, a text decompression device is provided, including:

[0032] A second acquisition module, configured to acquire a token sequence and a request language, where the token sequence is obtained by the text compression method according to any item in the first aspect; acquire a large language model for decompression;

[0033] A decompression module, configured to input the token sequence and the request language into the large language model, and use the large language model to restore the token sequence according to the request language to obtain an output text.

[0034] In a sixth aspect, a model training device is provided, including:

[0035] A third acquisition module, configured to acquire a plurality of training original texts, a request language, and token sequences respectively corresponding to the plurality of training original texts; acquire a training model, where the training model includes a compression training model and a large language training model, and the compression training model includes: a feature extraction training module and a q-former quantization processing training module;

[0036] A training module, configured to input the training original texts into the compression training model, so as to perform feature extraction on the training original texts by using the feature extraction training module to obtain a training embedded vector sequence; perform quantization processing on the training embedded vector sequence by using the q-former quantization processing training module to obtain a training token sequence; input the training token sequence and the request language into the large language training model, so as to perform text restoration on the training token sequence by using the large language training model according to the request language to obtain a training output text; perform iterative training on the training model according to the training output text and the training original texts to obtain a text processing model, where the text processing model includes: a text compression model and a large language model, and the text compression model is configured to compress text, and the large language model is configured to decompress the token sequence obtained from the compressed text.

[0037] In a seventh aspect, there is provided an electronic device, including:

[0038] One or more processors;

[0039] A memory;

[0040] One or more applications, where one or more applications are stored in the memory and are configured to be executed by one or more processors, and the one or more programs are configured to: execute the operations corresponding to the method shown in any possible implementation manner of the first aspect, or execute the operations corresponding to the method shown in the implementation manner of the second aspect, or execute the operations corresponding to the method shown in the implementation manner of the third aspect.

[0041] In an eighth aspect, there is provided a computer-readable storage medium, where the storage medium stores at least one instruction, at least one segment of program, a code set or an instruction set, and the at least one instruction, at least one segment of program, the code set or the instruction set is loaded and executed by a processor to execute the operations corresponding to the method shown in any possible implementation manner of the first aspect, or execute the operations corresponding to the method shown in the implementation manner of the second aspect, or execute the operations corresponding to the method shown in the implementation manner of the third aspect.

[0042] In a ninth aspect, a computer program product is provided, including a computer program which, when executed by a processor, implements the operations corresponding to the methods shown in any possible implementation manner of the first aspect, or, the operations corresponding to the methods shown in the implementation manner of the second aspect, or, the operations corresponding to the methods shown in the implementation manner of the third aspect.

[0043] In summary, the text compression method provided in this application has the following beneficial technical effects:

[0044] Obtain the target text to be compressed, input the target text to be compressed into the feature extraction module of the text compression model, capture the semantic and syntactic features of the text using deep learning technology, and represent the extracted key information in the form of an embedded vector sequence; use the q-former quantization processing module to perform quantization processing on the embedded vector sequence, convert the continuous embedded vector sequence into a discrete token sequence, and achieve effective compression of text information. By adopting the form of combining the feature extraction module with the q-former structure, it can significantly reduce the redundancy of text data while maintaining the ability to restore the original text with high quality.

[0045] In addition, this application also provides a text decompression method, a model training method, a device, and a device, all of which have the above beneficial technical effects. BRIEF DESCRIPTION OF THE DRAWINGS

[0046] In order to more clearly illustrate the technical solutions of the embodiments of this application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of this application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0047] Figure 1 is a schematic diagram of an application scenario of a text compression method provided by an embodiment of this application;

[0048] Figure 2 is a schematic flow chart of a text compression method provided by an embodiment of this application;

[0049] Figure 3 is a schematic flow chart of a text decompression method provided by an embodiment of this application;

[0050] Figure 4 is a schematic flow chart of a text compression + decompression method provided by an embodiment of this application;

[0051] Figure 5 is a schematic flow chart of a model training method provided by an embodiment of this application;

[0052] Figure 6 It is a schematic structural diagram of a text compression device provided by an embodiment of the present application;

[0053] Figure 7 It is a schematic structural diagram of a text decompression device provided by an embodiment of the present application;

[0054] Figure 8 It is a schematic structural diagram of a model training device provided by an embodiment of the present application;

[0055] Figure 9 It is a schematic structural diagram of an electronic device provided by an embodiment of the present application. Detailed implementation manners

[0056] This specific embodiment is only an interpretation of the present application, and it does not limit the present application. After reading this specification, those skilled in the art can make modifications to this embodiment without creative contributions as needed, but as long as it is within the scope of the present application, it is protected by the patent law.

[0057] It should be noted that in the optional embodiments of the present application, for relevant data such as object information, when the embodiments in the present application are applied to specific products or technologies, object permission or consent is required, and the collection, use, and processing of relevant data need to comply with relevant laws, regulations, and standards of relevant countries and regions. That is to say, if the embodiments in the present application involve data related to objects, they need to be obtained under the authorization and consent of the objects, the authorization and consent of relevant departments, and compliance with relevant laws, regulations, and standards of relevant countries and regions. If personal information is involved in the embodiments, the acquisition of all personal information requires the consent of the individual. If sensitive information is involved, the separate consent of the information subject needs to be obtained, and the embodiments also need to be implemented under the authorization and consent of the objects.

[0058] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are some, but not all, of the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application without creative efforts shall fall within the protection scope of the present application.

[0059] In addition, the term "and / or" in this article is only a description of the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " in this article generally represents an "or" relationship between the associated objects before and after, unless otherwise specified.

[0060] Currently, conventional text compression algorithms, such as Huffman coding and LZW coding, achieve compression by analyzing the frequency of character occurrences in the text to construct an optimal prefix code or through string matching. For example, Huffman coding constructs a binary tree based on character frequencies and assigns a unique variable-length code to each character to achieve compression. The LZW algorithm achieves compression by constructing a string dictionary and replacing repeated strings with indices in the dictionary.

[0061] These algorithms work well for short texts and simple language structures, but when dealing with long texts and complex language structures, their compression ratios and decompression effects are often unsatisfactory. In addition, these algorithms are difficult to handle text content with rich context and high language dependence, resulting in information loss or compression distortion.

[0062] Therefore, the inventors have found the following deficiencies in the conventional techniques:

[0063] Limited compression ratio: It is difficult to further improve the compression ratio when dealing with long texts and complex language structures.

[0064] Poor decompression effect: During the compression process, conventional algorithms may lose some key context information, resulting in a decline in the quality of the decompressed text.

[0065] Difficulty in handling complex language structures: Conventional algorithms mainly rely on character frequencies and local pattern matching and are difficult to effectively handle text content with rich context and high language dependence.

[0066] In recent years, the development of artificial intelligence technology, especially large language models (LLMs), has provided a new perspective for text processing. By deeply learning a large amount of text data, LLMs can understand complex language structures and context information and demonstrate excellent capabilities in fields such as text generation, translation, and summarization. However, despite the significant progress made by LLMs in text generation and other aspects, their application in the field of text compression is still limited. In particular, there is still a lack of effective solutions for achieving high compression ratios and high-quality restoration.

[0067] Based on this, aiming to solve the problems of low compression ratio and poor decompression effect of existing text compression technologies when facing long texts and complex language structures, embodiments of this application propose a new text compression technology by leveraging the natural language processing capabilities of large language models (LLMs), achieving a text compression technology with high compression ratio and high-quality restoration. Specifically, by extracting text features through a deep learning encoder, performing quantization processing in combination with the q-former module, and prediction generation using an autoregressive language model (LLM), the present invention can effectively compress text data while maintaining a high restoration quality, especially in scenarios of processing long texts and complex language structures.

[0068] Embodiments of this application not only improve the efficiency of text compression but also ensure the quality of the text after decompression. This has important practical application value for application scenarios that need to process a large amount of text data, such as cloud computing, big data analysis, mobile communication, and digital storage. Through the technical solution of embodiments of this application, while maintaining the integrity of text information, the resources required for data storage and transmission can be significantly reduced, providing a more efficient and economical text processing solution for users.

[0069] Please refer to Figure 1 , Figure 1 FIG. is a schematic diagram of an application scenario of a text compression method provided by an embodiment of this application. This text compression method can be applied to a text compression system. In some embodiments, the text compression system includes an electronic device and a terminal device. The terminal device can send the target text to be compressed. Of course, it can also be other servers that send the target file to the electronic device or read the target file of the electronic device itself. Embodiments of this application do not limit this further. Among them, the terminal device includes but is not limited to: mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Tablet Computers), PMPs (Portable Multimedia Players), in-vehicle terminals (such as in-vehicle navigation terminals), etc., and fixed terminals such as digital TVs and desktop computers. The electronic device includes but is not limited to a server or a terminal device. The server can be an independent physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server providing cloud computing services. The electronic device is used to implement the text compression method. The terminal device and the electronic device can be directly or indirectly connected through wired or wireless communication methods. Embodiments of this application do not limit this here. It can be understood that the above is only an example, and this embodiment does not limit it here.

[0070] Embodiments of this application provide a text compression method, as Figure 2 shown, this method includes:

[0071] S101. Obtain the target text to be compressed;

[0072] In the embodiments of the present application, the target text to be compressed is the target text uploaded by the terminal device or the target text in the electronic device.

[0073] Exemplarily, in the field of cloud computing, the target file can be a large number of log files generated by cloud computing platform servers, applications, etc. The log files include but are not limited to running status, error information, and user access records; the target file can also be user data stored in cloud storage services, such as documents, pictures, and videos. In the field of big data analysis, the target file can be data from various sources, such as data from social media, Internet of Things devices, and enterprise databases. When the data is transmitted or stored, the large amount of data needs to be compressed to save bandwidth and storage space. In the field of mobile communication, the target file can be SMS, MMS content, or voice call records, etc., which may need to be compressed before transmission to save transmission bandwidth and costs; the target file can also be data transmitted by mobile users when accessing web pages.

[0074] It should be noted that the embodiments of the present application can effectively process texts in different languages and formats, and do not limit the language and format of the target text.

[0075] S102. Obtain a text compression model, where the text compression model includes: a feature extraction module and a q-former quantization processing module;

[0076] S103. Input the target text to be compressed into the text compression model, and use the feature extraction module to extract features from the target text to be compressed to obtain an embedded vector sequence; use the q-former quantization processing module to perform quantization processing on the embedded vector sequence to obtain a token sequence, so as to achieve text compression.

[0077] Among them, the feature extraction module is used to extract the features of the target file. In some embodiments, the feature extraction module can use a deep learning encoder to process the text; the encoder can be a model based on the Transformer architecture, such as BERT. Through the feature extraction module, the deep semantic and syntactic features of the text can be captured; the text is converted into an embedded vector sequence in a high-dimensional space. It can be understood that the embedded vector sequence contains the key information of the text and can provide a basis for subsequent compression.

[0078] The q-former quantization processing module refers to a neural network module specifically designed for quantizing embedded vector sequences. It can convert continuous vector values into discrete token values, thereby achieving text compression. The q-former quantization processing module receives the embedded vector sequence output by the encoder and performs quantization processing on it. Among them, by learning the distribution of the embedded vector sequence, the q-former quantization processing module can use the embedded vector sequence output by the feature extraction module as the input data of this module, discretize the continuous values, and map them into a finite token set to obtain a token sequence. The quantized token sequence has a high compression rate while retaining the key information of the original text. The token sequence refers to the result after quantization processing. It is a series of discrete symbols or tokens used to represent the approximate representation of the original text data under quantization, thereby achieving compressed storage of the text. In the embodiments of this application, the size of the obtained token sequence is much smaller than the target text. However, the embodiments of this application adopt an innovative feature extraction module, namely the encoder-q-former structure, which can significantly reduce the redundancy of text data while maintaining the ability to restore the original text with high quality. The encoder uses deep learning techniques to capture the semantic and syntactic features of the text, while the q-former module efficiently compresses these features, so that even though the token sequence is small, it is still sufficient to support high-quality text restoration.

[0079] It can be seen that in the embodiments of this application, the target text to be compressed is obtained, and the target text to be compressed is input into the feature extraction module of the text compression model. Deep learning techniques are used to capture the semantic and syntactic features of the text, and the extracted key information is represented in the form of an embedded vector sequence; the q-former quantization processing module is used to perform quantization processing on the embedded vector sequence, converting the continuous embedded vector sequence into a discrete token sequence, achieving effective compression of text information. Adopting the form of combining the feature extraction module and the q-former structure can significantly reduce the redundancy of text data while maintaining the ability to restore the original text with high quality.

[0080] A possible implementation manner of the embodiments of this application is that the text compression model further includes: a preprocessing module;

[0081] Before using the feature extraction module to perform feature extraction on the target text to be compressed to obtain an embedded vector sequence, it further includes: using the preprocessing module to perform preprocessing on the target text to be compressed to obtain a processed text;

[0082] Correspondingly, using the feature extraction module to perform feature extraction on the target text to be compressed to obtain an embedded vector sequence includes: using the feature extraction module to perform feature extraction on the processed text to obtain an embedded vector sequence.

[0083] Among them, the preprocessing module is responsible for preliminarily processing the input text data (target text) to improve the efficiency and accuracy of subsequent processing. The preliminary processing includes at least one of the following:

[0084] Word segmentation processing: Decompose the continuous text string of the target text into meaningful lexical units.

[0085] Stop word removal processing: Delete common and insignificant words in the target text, such as "de" (of), "shi" (is), etc.

[0086] Punctuation processing: Identify and retain the punctuation marks in the target text because they have an important impact on the text structure.

[0087] It can be seen that in the embodiments of the present application, by preprocessing the target text, cleaning and processing the target text, redundant information can be removed, the data quality can be improved, and thus the compression quality can be improved.

[0088] A possible implementation manner of the embodiments of the present application is that the feature extraction module includes: GPT (Generative Pre-trained Transformer) or XLNet.

[0089] The extraction module based on the Transformer architecture directly captures the dependencies in the text through the self-attention mechanism, greatly improving the efficiency of the model in processing text data. Among them, GPT is a language model based on the Transformer architecture. XLNet is another language model based on the Transformer architecture, which adopts the method of the permutation language model.

[0090] It can be seen that in the embodiments of the present application, in the feature extraction stage, different pre-trained models, such as GPT or XLNet, can be used to adapt to different types of text data.

[0091] A possible implementation manner of the embodiments of the present application is to use the q-former quantization processing module to perform quantization processing on the embedded vector sequence to obtain a token sequence to achieve text compression, including:

[0092] Obtain the quantization granularity; Based on the quantization granularity, use the q-former quantization processing module to perform quantization processing on the embedded vector sequence to obtain a token sequence to achieve text compression.

[0093] The quantization granularity represents the fineness or resolution adopted when quantizing embedded vectors during text compression, and determines the detail level of the information that the quantized tokens can represent. In the embodiments of the present application, when a large amount of text data needs to be efficiently stored or transmitted, the text first needs to be embedded to convert it into a sequence of embedded vectors; according to the preset quantization granularity, the q-former quantization processing module is used to perform quantization processing on these vectors, dividing the continuous vector space into a series of discrete intervals, and mapping each vector to the token corresponding to the interval it belongs to, thus realizing the compression of the text. In the quantization processing stage, the number of compressed tokens can be changed by adjusting the output dimension of the q-former, so as to adjust the quantization granularity and achieve different compression ratios and restoration qualities. The quantization granularity can be adjusted according to actual needs. For example, the embedded vectors can be quantized into a set of 256, 512, or 1024 tokens.

[0094] Specifically, the quantization granularity can be determined according to the characteristics of the text, the compression requirements, or the default granularity. For example, for texts that need to highly preserve semantic information (such as legal documents, medical reports, etc.), a finer quantization granularity is set; while for texts with low requirements for semantic information (such as log data, social media content, etc.), a coarser quantization granularity can be set in exchange for a higher compression ratio.

[0095] Therefore, in some embodiments, the step of obtaining the quantization granularity can be implemented in multiple ways:

[0096] The first way is to determine the quantization granularity through experience or expert knowledge, and a suitable quantization granularity can be manually set. The second way is to dynamically determine the quantization granularity through an automated algorithm. For example, analyze features such as the vocabulary distribution and sentence length in the text, and automatically adjust the quantization granularity according to the corresponding relationship between the features, the preset features, and the granularity to optimize the compression effect, so as to more flexibly adapt to various compression requirements. It can be understood that other ways can also be used to implement the step of obtaining the quantization granularity, such as an adjustment mechanism based on user feedback, etc. It is not limited here, and the specific implementation methods will vary according to different application scenarios and requirements.

[0097] Furthermore, the embodiments of the present application provide a text decompression method, which includes: obtaining a token sequence and a request language, where the token sequence is obtained through the above-mentioned text compression method; obtaining a decompression model; inputting the token sequence into the decompression model to use the decompression model to restore the text of the token sequence to obtain an output text.

[0098] Further, the decompression model is a large language model for decompression; see Figure 3 , Figure 3A text decompression method provided by an embodiment of the present application includes:

[0099] S201. Obtain a token sequence and a request language, where the token sequence is obtained through a text compression method;

[0100] Among them, the request language is used to instruct a large language model to perform a decoding operation. Exemplarily, "Please perform text decoding according to the provided token sequence".

[0101] S202. Obtain a large language model for decompression;

[0102] S203. Input the token sequence and the request language into the large language model to restore the token sequence into text according to the request language by using the large language model, and obtain an output text.

[0103] In the decoding stage, the large language model LLM uses its powerful language understanding ability to gradually predict and restore the original text from the compressed token sequence and the request language to obtain the output text. The structure of the large language model is not limited in the embodiment of the present application, and the user can set it according to actual needs. Exemplarily, the large language model is the QWEN2 structure.

[0104] The embodiment of the present application proposes a text compression technology based on a large language model (LLM), aiming to solve the challenges faced by existing text compression algorithms when processing long and complex language structure texts. See Figure 4 , this technical solution extracts the features of the input text through a deep learning encoder, converts them into an embedded vector sequence, and then uses the q-former module to perform quantization processing on these vectors to generate a token sequence with a high compression rate. These tokens can retain the key information of the original text and are decoded through an autoregressive language model (LLM) to predict and restore the original text content.

[0105] It can be seen that the decompression method provided by the embodiment of the present application gradually expands the compressed token sequence through the LLM to restore the original long text, which not only improves the compression rate but also ensures the integrity and accuracy of the text information.

[0106] Furthermore, an embodiment of the present application provides a model training method, including: obtaining a plurality of training original texts and the token sequences corresponding to each of the plurality of training original texts; obtaining a training model, where the training model includes a compression training model and a decompression training model, and the compression training model includes: a feature extraction training module, a q-former quantization processing training module; inputting the training original texts into the compression training model to perform feature extraction on the training original texts using the feature extraction training module to obtain a training embedded vector sequence; performing quantization processing on the training embedded vector sequence using the q-former quantization processing training module to obtain a training token sequence; inputting the training token sequence into the decompression training model to perform text restoration on the training token sequence using the decompression training model to obtain a training output text; and iteratively training the training model based on the training output text and the training original texts to obtain a text processing model, where the text processing model includes: a text compression model and a decompression model, and the text compression model is used to compress text, and the decompression model is used to decompress the token sequence obtained from the compressed text.

[0107] Further, the decompression training model is a large language training model; see Figure 5 , Figure 5 A model training method provided by an embodiment of the present application includes:

[0108] S301. Obtain a plurality of training original texts, the request language, and the token sequences corresponding to each of the plurality of training original texts;

[0109] Based on multiple in-site and off-site data from multiple fields, construct a training format: request language (such as please restore the original text) + compressed token sequence + original text.

[0110] S302. Obtain a training model, where the training model includes a compression training model and a large language training model, and the compression training model includes: a feature extraction training module, a q-former quantization processing training module;

[0111] S303. Input the training original texts into the compression training model to perform feature extraction on the training original texts using the feature extraction training module to obtain a training embedded vector sequence; perform quantization processing on the training embedded vector sequence using the q-former quantization processing training module to obtain a training token sequence;

[0112] S304. Input the training token sequence and the request language into the large language training model to perform text restoration on the training token sequence using the large language training model according to the request language to obtain a training output text;

[0113] S305. Iteratively train the training model based on the training output text and the training original text to obtain a text processing model, which includes: a text compression model and a large language model. The text compression model is used to compress the text, and the large language model is used to decompress the token sequence obtained from the compressed text.

[0114] During the training phase, the large language training model learns how to predict and generate the original text based on the compressed token sequence and the requested language. For the model parameters: the hidden layer size of the LLM can be set between 128 and 512 to balance model complexity and performance. A large amount of text data is used to adjust the model parameters; learn the generation rules of the text in order to accurately recover the structure and semantic information of the original text from the compressed token sequence.

[0115] In an implementable way, during training, only calculate the loss for the original text part to achieve iterative training. Therefore, calculate the difference between the training output text and the expected output, and adjust the model parameters according to the difference until the difference meets the requirements and the training stops.

[0116] The training process includes a compression process and a decompression process, where:

[0117] Compression process:

[0118] Step a101: Input the preprocessed text.

[0119] Step a102: Extract features through the encoder.

[0120] Step a103: Perform quantization processing through the q-former module.

[0121] Step a104: Generate a compressed token sequence.

[0122] Decompression process: The decompression process is the reverse process of the compression process, including the following steps:

[0123] Step a201: Input the compressed token sequence.

[0124] Step a202: Use the autoregressive language model LLM to predict the text based on the token sequence and context information.

[0125] Step a203: Gradually expand the token sequence to restore the original text.

[0126] Furthermore, during the autoregressive language model training phase, different optimization algorithms, such as Adam or RMSprop, can be adopted to improve the training efficiency.

[0127] In summary, the present embodiment provides an efficient and flexible text compression technology that can significantly reduce the resources required for data storage and transmission while maintaining high restoration quality.

[0128] Specifically, 1. High compression ratio: High compression ratio is achieved through quantization processing and the prediction ability of the LLM. 2. High restoration quality: The q-former for quantization processing has a high compression ratio while retaining the key information of the original text, and the natural language processing ability of the LLM ensures high-quality restoration of the text. 3. Wide applicability: It is applicable to texts in different languages and formats and has broad application prospects.

[0129] It should be noted that the text compression method, text decompression method, and model training method provided in the embodiments of the present application can be referred to each other, and will not be elaborated in this embodiment.

[0130] Next, a text compression device provided in the embodiments of the present application will be introduced. The text compression device described below can be correspondingly referred to the text compression method described above. The text compression device of this embodiment is set in an electronic device. Refer to Figure 6 , Figure 6 which is the structural block diagram of the text compression device of one embodiment of the present application, including:

[0131] The first acquisition module 510 is configured to acquire the target text to be compressed; acquire a text compression model, where the text compression model includes: a feature extraction module and a q-former quantization processing module;

[0132] The compression module 520 is configured to input the target text to be compressed into the text compression model, so as to perform feature extraction on the target text to be compressed by using the feature extraction module to obtain an embedded vector sequence; perform quantization processing on the embedded vector sequence by using the q-former quantization processing module to obtain a token sequence, so as to achieve text compression.

[0133] In an implementable manner, the text compression model further includes: a preprocessing module;

[0134] It further includes:

[0135] The preprocessing module is configured to preprocess the target text to be compressed by using the preprocessing module to obtain a processed text;

[0136] Correspondingly, the compression module 520 is configured to:

[0137] Perform feature extraction on the processed text by using the feature extraction module to obtain an embedded vector sequence.

[0138] In an implementable manner, the feature extraction module includes: GPT or XLNet.

[0139] In an implementable manner, the compression module 520 is configured to:

[0140] Obtain a quantization granularity; based on the quantization granularity, use the q-former quantization processing module to perform quantization processing on the embedded vector sequence to obtain a token sequence, so as to implement text compression.

[0141] Next, a text decompression device provided in an embodiment of the present application will be introduced. The text decompression device described below can be correspondingly referred to the text decompression method described above. The text decompression device of this embodiment is set in an electronic device. Refer to Figure 7 , Figure 7 is a structural block diagram of a text decompression device according to an embodiment of the present application, including:

[0142] The second acquisition module 610 is configured to acquire a token sequence and a requested language, where the token sequence is obtained by a text compression method; acquire a large language model for decompression;

[0143] The decompression module 620 is configured to input the token sequence and the requested language into the large language model, so as to use the large language model to restore the token sequence according to the requested language to obtain an output text.

[0144] Next, a model training device provided in an embodiment of the present application will be introduced. The model training device described below can be correspondingly referred to the model training method described above. The model training device of this embodiment is set in an electronic device. Refer to Figure 8 , Figure 8 is a structural block diagram of a model training device according to an embodiment of the present application, including:

[0145] The third acquisition module 710 is configured to acquire a plurality of training original texts, a requested language, and token sequences respectively corresponding to the plurality of training original texts; acquire a training model, where the training model includes a compression training model and a large language training model, and the compression training model includes: a feature extraction training module, a q-former quantization processing training module;

[0146] A training module 720, configured to input the original training text into a compression training model, perform feature extraction on the original training text using a feature extraction training module to obtain a training embedded vector sequence; perform quantization processing on the training embedded vector sequence using a q-former quantization processing training module to obtain a training token sequence; input the training token sequence and the requested language into a large language training model, perform text restoration on the training token sequence using the large language training model according to the requested language to obtain a training output text; perform iterative training on the training model according to the training output text and the original training text to obtain a text processing model, where the text processing model includes: a text compression model and a large language model, the text compression model is used to compress text, and the large language model is used to decompress the token sequence obtained from the compressed text.

[0147] In an embodiment of the present application, an electronic device is provided, such as Figure 9 shown. Figure 9 The electronic device 300 shown in the figure includes: a processor 301 and a memory 303. Among them, the processor 301 and the memory 303 are connected, such as connected through a bus 302. Optionally, the electronic device 300 may further include a transceiver 304. It should be noted that in practical applications, the transceiver 304 is not limited to one, and the structure of the electronic device 300 does not constitute a limitation on the embodiments of the present application.

[0148] The processor 301 may be a CPU (Central Processing Unit, central processor), a general-purpose processor, a DSP (Digital Signal Processor, data signal processor), an ASIC (Application Specific Integrated Circuit, application-specific integrated circuit), an FPGA (Field Programmable Gate Array, field programmable gate array) or other programmable logic devices, transistor logic devices, hardware components or any combination thereof. It can implement or execute various exemplary logic blocks, modules and circuits described in conjunction with the disclosure of the present application. The processor 301 may also be a combination of computing functions, such as a combination of one or more microprocessors, a combination of a DSP and a microprocessor, etc.

[0149] The bus 302 may include a path for transmitting information between the above components. The bus 302 can be a PCI (Peripheral Component Interconnect) bus or an EISA (Extended Industry Standard Architecture) bus, etc. The bus 302 can be divided into an address bus, a data bus, a control bus, etc. For the sake of convenience of representation, Figure 9 it is only represented by a thick line in Figure 9 , but it does not mean that there is only one bus or one type of bus.

[0150] The memory 303 can be a ROM (Read Only Memory) or other types of static storage devices that can store static information and instructions, a RAM (Random Access Memory) or other types of dynamic storage devices that can store information and instructions, or it can also be an EEPROM (Electrically Erasable Programmable Read Only Memory), a CD-ROM (Compact Disc Read Only Memory) or other optical disc storage, optical disc storage (including compact discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), magnetic disk storage media or other magnetic storage devices, or any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto.

[0151] The memory 303 is used to store the application program code for implementing the solution of this application, and is controlled and executed by the processor 301. The processor 301 is used to execute the application program code stored in the memory 303 to implement the content shown in the foregoing method embodiments.

[0152] Figure 9 The illustrated electronic device is only an example and should not impose any limitations on the functions and usage scope of the embodiments of this application.

[0153] The embodiments of this application provide a computer-readable storage medium on which a computer program is stored. When the computer program runs on a computer, it enables the computer to execute the corresponding content in the foregoing method embodiments.

[0154] The embodiments of this application provide a computer program product, including a computer program that implements the corresponding content in the foregoing method embodiments when executed by a processor.

[0155] It should be understood that although the steps in the flowchart of the accompanying drawings are shown sequentially according to the indication of the arrows, these steps are not necessarily executed sequentially in the order indicated by the arrows. Unless there is a clear indication in this document, there is no strict order restriction for the execution of these steps, and they can be executed in other orders. Moreover, at least a part of the steps in the flowchart of the accompanying drawings may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily executed at the same time, but can be executed at different times, and their execution order is not necessarily sequential, but can be executed alternately or in turn with at least a part of other steps or sub-steps or stages of other steps.

[0156] The above are only partial embodiments of the present application. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present application, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present application.

Claims

1. A text compression method, characterized in that: include: Get the target text to be compressed; Acquire a text compression model, wherein the text compression model includes: a feature extraction module based on a Transformer architecture and a q-former quantization processing module; The target text to be compressed is input into the text compression model, and the feature extraction module is used to extract features according to the target text to be compressed, so as to capture the deep semantic and syntactic features of the text and obtain an embedded vector sequence; the features of vocabulary distribution and sentence length in the target text are analyzed, and the quantization granularity is determined according to the correspondence between the features and the preset features and the granularity, and the quantization granularity can be dynamically adjusted according to user feedback; based on the quantization granularity, the embedded vector sequence is quantized by the q-former quantization processing module to obtain a token sequence to achieve text compression, and the token sequence has a high compression rate while retaining the key information of the target text; The q-former quantization processing module can convert continuous vector values ​​into discrete token values, thereby achieving text compression; The token sequence and the request language are input into a large language model, so as to perform text restoration on the token sequence using the large language model according to the request language to obtain an output text, wherein the hidden layer size of the large language model is set between 128 and 512.

2. The text compression method according to claim 1, characterized in that: The text compression model also includes: a preprocessing module; Before extracting features using the feature extraction module according to the target text to be compressed to obtain the embedded vector sequence, the method further includes: Using the preprocessing module to preprocess the target text to be compressed to obtain a processed text; Accordingly, the feature extraction module is used to extract features according to the target text to be compressed to obtain an embedded vector sequence, including: The feature extraction module is used to extract features from the processed text to obtain an embedded vector sequence.

3. The text compression method according to claim 1, characterized in that: The feature extraction module includes: GPT or XLNet.

4. A model training method, characterized in that: include: Get multiple training original texts, request language, and token sequences corresponding to each of the multiple training original texts; Obtaining a training model, wherein the training model includes a compression training model and a large language training model, wherein the compression training model includes: a feature extraction training module based on a Transformer architecture, and a q-former quantization processing training module, wherein the q-former quantization processing training module can convert continuous vector values ​​into discrete token values, thereby achieving text compression; wherein the hidden layer size of the large language model is set between 128 and 512; The original training text is input into the compression training model, so as to perform feature extraction using the feature extraction training module according to the original training text, capture the deep semantic and syntactic features of the text, and obtain a training embedded vector sequence; the training embedded vector sequence is quantized using the q-former quantization processing training module to obtain a training token sequence, wherein the training token sequence has a high compression rate and retains key information of the original training text; Inputting the training token sequence and the request language into the large language training model, so as to perform text restoration on the training token sequence using the large language training model according to the request language to obtain a training output text; The training model is iteratively trained according to the training output text and the training original text to obtain a text processing model, wherein the text processing model includes: a text compression model and a large language model, so as to implement the method described in any one of claims 1 to 3.

5. A text compression device, characterized in that: include: A first acquisition module, used for acquiring a target text to be compressed; Acquire a text compression model, wherein the text compression model includes: a feature extraction module based on a Transformer architecture and a q-former quantization processing module, wherein the q-former quantization processing module can convert continuous vector values ​​into discrete token values, thereby achieving text compression; A compression module is used to input the target text to be compressed into the text compression model, and use the feature extraction module to perform feature extraction according to the target text to be compressed, capture the deep semantic and syntactic features of the text, and obtain an embedded vector sequence; analyze the features of vocabulary distribution and sentence length in the target text, and determine the quantization granularity according to the correspondence between the features and the preset features and granularity, and the quantization granularity can be dynamically adjusted according to user feedback; based on the quantization granularity, use the q-former quantization processing module to quantize the embedded vector sequence to obtain a token sequence to achieve text compression, and the token sequence has a high compression rate while retaining the key information of the target text; A decompression module is used to input the token sequence and the request language into a large language model, so as to perform text restoration on the token sequence using the large language model according to the request language to obtain an output text, wherein the hidden layer size of the large language model is set between 128 and 512.

6. A model training device, characterized in that: include: The third acquisition module is used to obtain multiple training original texts, request languages, and token sequences corresponding to each of the multiple training original texts; Obtaining a training model, wherein the training model includes a compression training model and a large language training model, wherein the compression training model includes: a feature extraction training module based on a Transformer architecture, and a q-former quantization processing training module, wherein the q-former quantization processing training module can convert continuous vector values ​​into discrete token values, thereby achieving text compression, wherein the hidden layer size of the large language model is set between 128 and 512; A training module is used to input the original training text into the compression training model, so as to use the feature extraction training module to perform feature extraction according to the original training text, capture the deep semantic and syntactic features of the text, and obtain a training embedded vector sequence; use the q-former quantization processing training module to quantize the training embedded vector sequence to obtain a training token sequence, which has a high compression rate and retains the key information of the original training text; input the training token sequence and the request language into the large language training model, so as to use the large language training model to perform text restoration on the training token sequence according to the request language to obtain a training output text; iteratively train the training model according to the training output text and the original training text to obtain a text processing model, wherein the text processing model includes: a text compression model and a large language model, so as to implement the method described in any one of claims 1 to 3.

7. An electronic device, characterized in that: include: one or more processors; Memory; One or more applications, wherein the one or more applications are stored in the memory and configured to be executed by the one or more processors, and the one or more applications are configured to: perform the steps of the method according to any one of claims 1 to 4.