One-way data transmission method based on Tokenizer
Through the Tokenizer model, multimodal data is converted into Token format and compressed, which solves the shortcomings of the existing one-way data transmission system in terms of flexibility, efficiency, security and multimodal support, and realizes efficient and secure one-way transmission of data in a cross-net environment.
Patent Information
- Application Number
- CN202411958754.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-30
- Publication Date
- 2025-05-27
AI Technical Summary
The existing one-way data transmission systems have significant shortcomings in flexibility, efficiency, security, multimodal support and cost, especially in cross-network environments that may cause information leakage risks.
The Tokenizer-based one-way data transmission method is adopted, and the multimodal data is converted into a unified Token format through the Tokenizer model, and compressed and decompressed to achieve safe and efficient one-way transmission of data in a cross-net environment.
It significantly improves the efficiency and security of data transmission, enhances the applicability of the system and multimodal data support capabilities, reduces transmission time and improves bandwidth utilization.
Smart Images

Figure CN120050337A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data processing, and particularly to a one-way data transmission method based on Tokenizer. Background Art
[0002] Tokenizer is a tool or algorithm in natural language processing (NLP) used to split input data (usually text) into smaller units called tokens. These tokens can be words, sub-words, characters, or even higher-level semantic units and are one of the fundamental components of modern NLP models. The role of Tokenizer is to segment and perform vocabulary mapping on a piece of text data to obtain the token sequence of this text.
[0003] Tokenizer is the core component of the one-way data transmission system in this example, responsible for splitting and encoding multi-modal data (text, images, audio, etc.) into a unified token sequence to achieve a standardized and structured representation of the data. Through efficient segmentation and compression mechanisms, Tokenizer not only reduces the transmission burden but also enhances data security and context retention capabilities. Its dynamic adjustment and modality unification functions enable the system to adapt to complex data requirements and diverse application scenarios, while providing an efficient interface for decoding and restoration, which is the key to achieving efficient, secure, and intelligent one-way data transmission.
[0004] Although existing one-way data transmission systems have advantages in terms of security and data isolation, they have significant deficiencies in terms of flexibility, efficiency, security, multi-modal support, and cost. Especially when data is transmitted across networks, that is, between the internal network and the external network, it may pose a risk of information leakage. Specifically, the data transmission protocol and format are fixed, making it difficult to adapt to complex requirements; the transmission efficiency is low and the real-time performance is poor; relying on dedicated hardware results in high deployment and maintenance costs; there is a lack of protection against covert channels and malicious data; there is insufficient support for multi-modal data processing and context information retention, and at the same time, there is a lack of intelligent optimization capabilities, making it difficult to meet the needs of modern networks and multi-modal applications. In addition, traditional character encoding methods are difficult to achieve a good balance between the compression rate and the recovery rate. These drawbacks make traditional one-way systems ineffective in complex scenarios of high efficiency, security, and intelligence. Summary of the Invention
[0005] In view of the above problems, a one-way data transmission method based on Tokenizer is provided. Through a unified process of Token format conversion, compression, decompression, and restoration, the present invention realizes secure and efficient one-way transmission of data in a cross-network environment; by analyzing the structure and content of data through the Tokenizer model, high-frequency data is extracted as Token data, and redundant information is further compressed; through the Tokenizer model, multimodal data such as text, audio, and images are uniformly encoded into common Token data, significantly enhancing the applicability of the system; by converting and compressing the original data through the Tokenizer model, the specific content of the data can be effectively hidden, significantly enhancing the security of data transmission and the privacy protection of data.
[0006] To achieve the above object, the technical solution adopted by the present invention is as follows.
[0007] A Token compression method for data includes the following steps: Step S1: The original data is converted into a Token sequence through the Tokenizer module and output to the Token compression module; Step S2: The Token compression module slices the Token sequence into fixed-size N pieces and distributes them to different threads with corresponding quantities; Step S3: Parallelly count the frequencies of each Token data in each thread N and output the counted frequency data to the central node; Step S4: The central node performs global grouping and frequency merging according to the Token data, and assigns indexes to the high-frequency Token data to construct global dictionary data; Step S5: Through the global dictionary data, index replacement is performed on the high-frequency Token data in each thread, and the low-frequency Token data in each thread is retained, and finally the compressed Token data is obtained.
[0008] Preferably, the fixed size N > 4096.
[0009] Preferably, the fixed size N is sliced according to the number of rows of the Token sequence.
[0010] Preferably, in step S5, the high-frequency Token data refers to the frequency of the Token data > 3 times.
[0011] A one-way data transmission method based on Tokenizer, adopting the Token compression method according to any one of claims 1-4, includes the following steps: Step 1: The original data in the external network is converted into a uniformly formatted Token sequence by the Tokenizer module and output to the Token compression module; Step 2: The Token compression module converts the Token sequence into global dictionary data and compressed Token data. The global dictionary data and the compressed Token data are sequentially output to the Token decompression module after passing through the sending module and the receiving module; Step 3: The Token decompression module traverses the compressed Token data one by one according to the indexes in the global dictionary data; Step 4: If the Token decompression module determines that the compressed Token data is an index in the global dictionary data, it replaces the index with the corresponding Token data; otherwise, it retains the non-index data to obtain the decompressed Token sequence; Step 5: The decompressed Token sequence is translated into the original data by the Tokenizer module and then output to the internal network.
[0012] Preferably, a unidirectional channel is used to connect the sending module and the receiving module.
[0013] Preferably, the training steps of the Tokenizer module are as follows: Step S1: Collect multimodal data and preprocess the multimodal data to obtain a corpus; Step S2: The corpus is split into test data and model training data in a ratio of 99:1; Step S3: Each character, each sampling point, or each batch in the model training data is regarded as a Token data, and special markers are added at the corresponding positions to obtain an initial vocabulary; Step S4: The initial vocabulary is expanded, compressed, and iterated through a subword algorithm to obtain a Tokenizer model; Step S5: Input the test data into the Tokenizer model. If the processing speed < 1000 token data / second, or the coverage rate of the expansion of the initial vocabulary for the corpus ≤ 99%, clean the Token data in the corpus and expand the corpus; otherwise, it indicates that the compression conversion and translation effects of the Tokenizer model are good. At this time, output the final Tokenizer model.
[0014] Preferably, the multimodal data includes text files, audio files, and video files.
[0015] Preferably, the special markers are to insert special markers respectively between different functional classifications and different types of Token data <cls> 、 <pad>。
[0016] Preferably, the special mark is to replace the Token data of unknown category with <unk>。
[0017] Preferably, the preprocessing is to convert the text file into a txt format after unifying the case, removing duplicates, and splitting sentences.
[0018] Preferably, the preprocessing is to convert the audio file into a wav format after denoising and normalizing.
[0019] Preferably, the preprocessing is to convert the video file into an mp4 format.
[0020] Due to the adoption of the above technical solutions, the present invention has the following beneficial effects.
[0021] (1) Through the unified Token format conversion, compression, decompression, and restoration processes, the present invention realizes the safe and efficient one-way transmission of data in a cross-network environment; by analyzing the structure and content of the data through the Tokenizer model, high-frequency data is extracted as Tokens, and redundant information is further compressed. Therefore, under the same bandwidth conditions, the present invention can significantly reduce the transmission time, improve the bandwidth utilization rate, has a high data compression efficiency, and optimizes the transmission performance.
[0022] (2) By using the Tokenizer model to uniformly encode multimodal data such as text, audio, images, etc. into a common Token representation, the present invention significantly enhances the applicability of the system, can be flexibly deployed in different application scenarios, thereby realizing multimodal data support and efficient cross-modal data transmission, and has strong applicability.
[0023] (3) By converting and compressing the original data through the Tokenizer model, the present invention can effectively hide the specific content of the data. Even if the transmitted data is intercepted, the un-decoded Tokens cannot directly restore the original information, which can significantly enhance the security of data transmission and the privacy protection of data, and provide guarantee for the secure transmission of sensitive information.
[0024] (4) The Token decompression link of the present invention adopts a predefined global dictionary data and a segmented verification mechanism for traversing the compressed Token sequence one by one. Even if some data is lost during the transmission process, the data can be restored through redundant information, which is more reliable than traditional character-by-character transmission systems, has high decoding reliability, and accurate data recovery. Description of the Drawings
[0025] The following provides a detailed discussion on the fabrication and application of the preferred embodiments of the present invention. It should be understood, however, that the present invention provides numerous applicable inventive concepts, which can be embodied in various specific environments. The specific embodiments discussed are merely for illustrating the specific ways of manufacturing and using the present invention and do not limit the scope of the present invention. For those of ordinary skill in the art, other drawings can also be obtained based on these drawings without creative efforts.
[0026] Figure 1 This is the flowchart of the present invention. Detailed implementation manners
[0027] The following provides a detailed discussion on the fabrication and application of the preferred embodiments of the present invention. It should be understood, however, that the present invention provides numerous applicable inventive concepts, which can be embodied in various specific environments. The specific embodiments discussed are merely for illustrating the specific ways of manufacturing and using the present invention and do not limit the scope of the present invention.
[0028] The present invention realizes the secure and efficient one-way transmission of data in a cross-network environment through a unified Token format conversion, compression, decompression, and restoration process; analyzes the structure and content of data through the Tokenizer model, extracts high-frequency data as Tokens, and further compresses redundant information. Therefore, under the same bandwidth conditions, the present invention can significantly reduce the transmission time, improve the bandwidth utilization rate, has a high data compression efficiency, and optimizes the transmission performance.
[0029] The following will be further elaborated in conjunction with the attached Figure 1 As shown in Figure 1 A one-way data transmission method based on Tokenizer is shown.
[0030] A one-way data transmission method based on Tokenizer includes the following steps: Step 1: The original data in the external network is converted into a Token sequence with a unified format by the Tokenizer module and output to the Token compression module.
[0031] The training steps of the Tokenizer module are as follows: Step S1: Collect multi-modal data and preprocess the multi-modal data to obtain a corpus; the multi-modal data includes text files, audio files, and video files; the preprocessing is to convert the text files into txt format after unifying the case, removing duplicates, and sentence segmentation, convert the audio files into wav format after denoising and normalization, and convert the video files into mp4 format.
[0032] Step S2: Split the corpus into test data and model training data at a ratio of 99:1.
[0033] Step S3: Treat each character, each sampling point, or each batch in the model training data as a Token data, and insert special tokens between different functional classifications and different types of Token data respectively <cls> 、 <pad>; or replace the Token data of unknown categories with <unk>, an initial vocabulary is obtained.
[0034] Step S4: Expand, compress, and iterate the initial vocabulary through a subword algorithm to obtain a Tokenizer model.
[0035] The core idea of the subword algorithm is to gradually combine common characters or subword pairs in the text into larger subword units through frequency-driven merging operations until the set vocabulary size is met. The subword algorithm includes the following steps: First, decompose the initial vocabulary into individual characters and add delimiters, such as spaces, between each character.
[0036] In each round of merging, the frequencies of all subword pairs (bigrams) in the corpus need to be counted. Let C represent the corpus, P ( a , b ) represent the frequency of the subword pair ( a , b ): where w represents a word sequence in the corpus. Count w ( a , b ) represents w the number of occurrences of the subword pair ( a , b ) in
[0037] Then find the subword pair with the highest frequency, merge it into a new subword, continuously merge the highest-frequency subword pairs, and update the corpus until the target vocabulary size is reached or there are no more subpairs to merge.
[0038] Step S5: Input the test data into the Tokenizer model. If the processing speed < 1000 token data / second, or the coverage rate of the expansion of the initial vocabulary in the corpus ≤ 99%, clean the Token data in the corpus and expand the corpus; otherwise, it means that the compression transformation and translation effect of the Tokenizer model is good, enabling it to efficiently support text encoding and processing in subsequent tasks. At this time, output the final Tokenizer model.
[0039] Step 2: The Token compression module converts the Token sequence into global dictionary data and compressed Token data, and the global dictionary data and the compressed Token data are sequentially output to the Token decompression module after passing through the sending module and the receiving module; a unidirectional channel is used to connect the sending module and the receiving module.
[0040] The Token compression method adopted by the Token compression module includes the following steps: Step S1: The original data is converted into a Token sequence by the Tokenizer module and output to the Token compression module.
[0041] Step S2: The Token compression module slices the Token sequence into pieces according to a fixed size N The fixed size N > 4096 or slices according to the number of rows of the Token sequence, and distributes them to different threads with corresponding quantities.
[0042] Step S3: Parallelly count the frequencies of each Token data in each thread N and output the counted frequency data to the central node.
[0043] Step S4: The central node performs grouping and frequency merging of Token data within the global scope according to Token data, and assigns indexes to high-frequency Token data to construct global dictionary data.
[0044] Step S5: Index replacement is performed on high-frequency Token data in each thread, that is, Token data with a frequency of occurrence > 3 times, through the global dictionary data, and low-frequency Token data in each thread is retained, and finally the compressed Token data is obtained.
[0045] In the present invention, the Tokenizer model uniformly encodes multimodal data such as text, audio, images, etc. into a general Token representation, significantly enhancing the applicability of the system, enabling flexible deployment in different application scenarios, thereby realizing multimodal data support and efficient transmission of cross-modal data, with strong applicability.
[0046] Step 3: The Token decompression module traverses the compressed Token data one by one according to the indexes in the global dictionary data.
[0047] Step 4: If the Token decompression module determines that the compressed Token data is an index in the global dictionary data, it replaces the index with the corresponding Token data; otherwise, it retains the non-index data to obtain the decompressed Token sequence.
[0048] The Token decompression link of the present invention adopts a predefined global dictionary data and a segmented verification mechanism for traversing the compressed Token sequence one by one. Even if some data is lost during the transmission process, data recovery can be achieved through redundant information, which is more reliable than traditional systems based on character-by-character transmission, with high decoding reliability and accurate data recovery.
[0049] Step 5: The decompressed Token sequence is translated into the original data by the Tokenizer module and then output to the internal network.
[0050] In the present invention, the Tokenizer model is used to convert and compress the original data, which can effectively hide the specific content of the data. Even if the transmitted data is intercepted, the undecoded Tokens cannot directly restore the original information, which can significantly enhance the security of data transmission and the privacy protection of data, providing guarantee for the secure transmission of sensitive information.
[0051] Although the description has been made in detail, it should be understood that various changes, substitutions and alterations can be made without departing from the spirit and scope of the invention defined by the appended claims. In addition, the specific embodiments described do not limit the scope of the present invention. It is easy for those of ordinary skill in the art to understand based on the present invention that the processes, machines, manufactures, compositions of matter, means, methods, or steps that currently exist or will be developed in the future can perform functions substantially the same as those of the embodiments of the present invention or achieve results substantially the same as those of the embodiments of the present invention. Therefore, the appended claims are intended to include such processes, machines, manufactures, compositions of matter, means, methods or steps within their scope.< / unk> < / pad> < / cls> < / unk> < / pad> < / cls>
Claims
1. A data Token compression method, characterized in that: The following steps are involved: Step S1: The original data is converted into a Token sequence through the Tokenizer module and output to the Token compression module; Step S2: The Token compression module converts the Token sequence into a fixed size N Cut the slices and distribute them to the corresponding number of different threads; Step S3: Parallel statistics of each thread N The frequency of each token data appearing, and the statistical frequency data is output to the central node; Step S4: The central node performs global grouping and frequency merging according to the Token data, assigns indexes to high-frequency Token data, and constructs global dictionary data; Step S5: The high-frequency Token data in each thread is indexed and replaced by the global dictionary data, and the low-frequency Token data in each thread is retained, so as to finally obtain the compressed Token data.
2. A data Token compression method as claimed in claim 1, characterized in that: The fixed size N >4096.
3. A data Token compression method as claimed in claim 2, characterized in that: The fixed size N The token sequence is split into pieces according to the number of rows.
4. A data Token compression method as claimed in claim 1, characterized in that: In step S5, high-frequency Token data refers to Token data that appears more than 3 times.
5. A Tokenizer-based one-way data transmission method, using the Token compression method according to any one of claims 1 to 4, characterized in that: The following steps are involved: Step 1: The original data in the external network is converted into a Token sequence with a unified format through the Tokenizer module and output to the Token compression module; Step 2: The Token compression module converts the Token sequence into global dictionary data and compressed Token data, and the global dictionary data and compressed Token data are output to the Token decompression module after passing through the sending module and the receiving module in sequence; Step 3: The Token decompression module traverses the compressed Token data one by one according to the index in the global dictionary data; Step 4: If the Token decompression module determines that the compressed Token data is an index in the global dictionary data, the index is replaced with the corresponding Token data; otherwise, the non-index data is retained to obtain the decompressed Token sequence; Step 5: The decompressed Token sequence is translated into raw data by the Tokenizer module and then output to the internal network.
6. A Tokenizer-based one-way data transmission method as claimed in claim 5, characterized in that: The sending module and the receiving module are connected by a unidirectional channel.
7. A Tokenizer-based one-way data transmission method as claimed in claim 5, characterized in that: The training steps of the Tokenizer module are as follows: Step S1: Collect multimodal data and preprocess the multimodal data to obtain a corpus; Step S2: Split the corpus into test data and model training data at a ratio of 99:1; Step S3: Treat each character or each sampling point or each batch in the model training data as a token data, and add a special tag at the corresponding position to obtain the initial vocabulary; Step S4: expanding, compressing and iterating the initial vocabulary by a subword algorithm to obtain a Tokenizer model; Step S5: Input the test data into the Tokenizer model. If the processing speed is <1000 token data / second, or the coverage of the corpus by the expansion of the initial vocabulary is ≤99%, the Token data in the corpus is cleaned and the corpus is expanded. Otherwise, it means that the compression, conversion and translation effects of the Tokenizer model are good. At this time, the final Tokenizer model is output.
8. A Tokenizer-based one-way data transmission method as claimed in claim 7, characterized in that: The multimodal data includes text files, audio files and video files.
9. A Tokenizer-based one-way data transmission method as claimed in claim 7, characterized in that: The special mark is to insert special marks between different functional categories and different types of Token data. <cls> 、 <pad>, or the special marker is to replace the unknown category of Token data with <unk> 。< / unk> < / pad> < / cls> 10. A Tokenizer-based one-way data transmission method as claimed in claim 7, characterized in that: The preprocessing is to convert the text file into txt format after unifying the upper and lower case, removing duplicates, and segmenting sentences; convert the audio file into wav format after denoising and normalization; and convert the video file into mp4 format.