Data processing method and device, electronic equipment and storage medium

By constructing a data feature element library and using data feature elements to characterize the distribution of characters in the data, and splitting data blocks for matching processing, the problem of low data compression efficiency in existing technologies is solved, and more efficient data compression and encryption are achieved.

CN116662278BActive Publication Date: 2026-05-29BEIJING YOUZHUJU NETWORK TECH CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
BEIJING YOUZHUJU NETWORK TECH CO LTD
Filing Date
2023-06-02
Publication Date
2026-05-29

AI Technical Summary

Technical Problem

Existing data compression algorithms suffer from low efficiency in data center storage, especially due to the limitations of a single compression method, resulting in poor compression ratios.

Method used

By constructing a data feature element library, the distribution of characters in the data is characterized by data feature elements, and the data to be processed is matched for compression or encryption. This includes feature elements for repetitive descriptions, regular descriptions, and constant data information, and the data blocks are split for matching processing.

Benefits of technology

It improves data processing efficiency, avoids the inefficiency caused by a single compression method, and achieves higher compression ratio and encryption effect.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116662278B_ABST
    Figure CN116662278B_ABST
Patent Text Reader

Abstract

The present disclosure relates to a data processing method and device, electronic equipment and storage medium. The method comprises: obtaining to-be-processed data; matching the to-be-processed data with feature elements in a data feature element library; wherein the data feature element library comprises a plurality of different data feature elements, and the data feature elements represent the distribution of characters in the data; and in the case that there is a target data feature element in the data feature element library that matches the to-be-processed data, processing the to-be-processed data based on the target data feature element to obtain processed data. Since the data feature elements represent the distribution of characters in the data, the to-be-processed data can be compressed or encrypted by the target data feature element matched during processing, which can avoid the phenomenon of low processing efficiency caused by a single compression method, thereby improving the processing efficiency of the data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of computer technology, and in particular to data processing methods, apparatus, electronic devices and storage media. Background Technology

[0002] With the booming development of cloud computing and big data industries, the larger the data center, the more data it stores. Larger data volumes and longer storage cycles mean higher storage costs. Therefore, it is necessary to compress the data stored in data centers to reduce storage requirements.

[0003] In the process of data compression, it is usually necessary to pre-set the required compression algorithm (such as LZ77 or BZIP2) to compress the data before storage, thereby reducing the amount of data to be stored. Since various compression algorithms have their own advantages and disadvantages, data compressed using a pre-set compression algorithm may result in a low compression ratio. Summary of the Invention

[0004] This disclosure provides a data processing method, apparatus, electronic device, and storage medium.

[0005] According to one aspect of this disclosure, a data processing method is provided, the method comprising:

[0006] Obtain the data to be processed;

[0007] The data to be processed is matched with feature elements in the data feature element library; wherein, the data feature element library includes multiple different data feature elements, and the data feature elements represent the distribution of characters in the data;

[0008] If a target data feature element that matches the data to be processed exists in the data feature element library, the data to be processed is processed based on the target data feature element to obtain processed data; the processing method includes compression or encryption.

[0009] According to another aspect of this disclosure, a data processing apparatus is provided, the apparatus comprising:

[0010] The data acquisition module is used to acquire the data to be processed.

[0011] A matching module is used to match the data to be processed with feature elements in a data feature element library; wherein, the data feature element library includes multiple different data feature elements, and the data feature elements represent the distribution of characters in the data;

[0012] The processing module is used to process the data to be processed based on the target data feature element when there is a target data feature element in the data feature element library that matches the data to be processed, so as to obtain processed data; the processing method includes compression or encryption.

[0013] According to a third aspect of this disclosure, an electronic device is provided. The electronic device includes a memory and a processor, wherein the memory stores a computer program, and the processor executes the program to implement the method described above.

[0014] According to a fourth aspect of this disclosure, a computer-readable storage medium is provided having a computer program stored thereon that, when executed by a processor, implements the methods described above.

[0015] The data processing method, apparatus, electronic device, and storage medium provided in this disclosure match the data to be processed with feature elements in a data feature element library. If a corresponding target feature data element is matched, the data to be processed can be processed based on that target feature data element to achieve compression or encryption of the data. Since data feature elements represent the distribution of characters in the data, compression or encryption can be performed based on the matched target data feature elements during the processing of the data. This avoids the low processing efficiency caused by a single compression or encryption method, thereby improving data processing efficiency. Attached Figure Description

[0016] Further details, features, and advantages of this disclosure are disclosed in the following description of exemplary embodiments in conjunction with the accompanying drawings, in which:

[0017] Figure 1 A flowchart illustrating a data processing method provided in an exemplary embodiment of this disclosure;

[0018] Figure 2 A schematic block diagram of the functional modules of a data processing apparatus provided in an exemplary embodiment of this disclosure;

[0019] Figure 3 A structural block diagram of an electronic device provided as an exemplary embodiment of this disclosure;

[0020] Figure 4 A block diagram of a computer system provided for an exemplary embodiment of this disclosure. Detailed Implementation

[0021] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.

[0022] It should be understood that the steps described in the method embodiments of this disclosure may be performed in different orders and / or in parallel. Furthermore, the method embodiments may include additional steps and / or omit the steps shown. The scope of this disclosure is not limited in this respect.

[0023] The term "comprising" and its variations as used herein are open-ended, meaning "including but not limited to". The term "based on" means "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". Definitions of other terms will be given in the description below. It should be noted that the concepts of "first", "second", etc., used in this disclosure are only used to distinguish different devices, modules, or units, and are not intended to limit the order of functions performed by these devices, modules, or units or their interdependencies.

[0024] It should be noted that the terms "a" and "a plurality of" used in this disclosure are illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".

[0025] The names of messages or information exchanged between multiple devices in the embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of such messages or information.

[0026] In data compression, related technologies typically require pre-setting the desired compression algorithm to compress the data to be stored, and then storing the compressed data to reduce its size. However, since various existing compression algorithms have their own advantages and disadvantages and limitations, if all data to be compressed uses a single specific algorithm, the limitations of that algorithm will result in low data compression efficiency.

[0027] Therefore, when obtaining the data to be compressed in this embodiment of the present disclosure, if the data to be compressed is large, for example, if the size of the data exceeds a certain value, it can be split into multiple small data blocks. For example, it can be split into data blocks of 512B, 1K, 4K or 4M as needed. The size of each data block can be the same or different.

[0028] This disclosure allows for the pre-construction of a data feature element library, which may include multiple different data feature elements used to characterize the distribution of characters in the data. For example, the data feature elements in this library may include repetitive description information, regular description information, or constant data information.

[0029] The repetition description indicates the presence of duplicate data. This duplicate data can be characterized by a high frequency of repeated characters, such as characters having a repetition rate exceeding a certain threshold (e.g., over 20%). The duplicate data can also be one or more characters appearing repeatedly in certain positions within the data. Furthermore, the duplicate data can exhibit one or more of these patterns, such as the repeated occurrences of the numbers "168" or the repeated appearance of the character "ye".

[0030] This pattern description indicates that the data exhibits a certain regularity, such as increasing or decreasing patterns. For example, the data might contain the characters: increasing + other + decreasing. Increasing and decreasing patterns in the data can be considered as patterns.

[0031] The constant data can be known data, such as the value corresponding to π, the value corresponding to the speed of light, or the data corresponding to a musical score, etc.

[0032] In this embodiment, the data feature element library can be continuously improved during its creation process. For example, more known data can be added to the library, such as commonly used known data, like the dates corresponding to important historical events. In this embodiment, the corresponding data feature elements in the library can be abstracted into corresponding feature values. These feature values ​​can be specific numbers, characters, etc. Thus, when the data to be compressed is matched with data feature elements in the library, at least one of the following—repeating characters, pattern descriptions, and constant data—can be replaced with its corresponding feature value, thereby significantly improving the data compression rate.

[0033] In the embodiments provided in this disclosure, after the data feature element library is created, it can be split into multiple data blocks as needed, for example, when the data to be processed is large, and each of these multiple data blocks can be matched with the data feature elements in the data feature element library. Otherwise, it is not necessary to split it; if it is not split, the data to be processed can be directly matched with the data feature elements in the data feature element library.

[0034] In this embodiment, if it is necessary to split the data to be processed, the data can be identified and split according to its distribution. For example, highly repetitive data, data with strong regularity, or constant data can be kept in the same data block. For instance, the data to be processed, "123123123123891654321", can be split into "123123123", "123891", and "654321". This avoids destroying the repetitive, regular, or constant data in the data to be compressed, thus facilitating preparation for subsequent compression.

[0035] In this embodiment, the example of data to be processed being split into multiple data blocks is used for illustration. A data block is matched against feature elements in a data feature element library. If a matching target data feature element is found, the data block can be compressed based on that target feature element. For example, if the data block is "18888888888", and the target data feature element is repetitive descriptive information, then when compressing the data block, the corresponding repetitive string can be used as a repetitive feature value. This repetitive feature value is multiplied by the number of repetitions to obtain the data description of the data block. The data block can be described as "1 followed by 10 eights". Similarly, a phone number like "12828282828" can be described as "1 followed by 5 28s". By processing the data block or data to be compressed using repetitive feature values, the position of repetitive characters in the data, and non-repetitive characters as its data description, compression of the data block or data to be processed is achieved. In this embodiment, the repetitive feature value can also be replaced with other data, achieving data encryption simultaneously with data compression.

[0036] When matching a data block with feature elements in a data feature element library, if the matched target data feature element is a regular descriptive information, the data block can be compressed based on that target data feature element. For example, if a data string exhibits obvious increasing or decreasing characteristics, it can also be identified as an increasing or decreasing feature descriptor element.

[0037] For example, if the data block corresponds to a string of numbers: 1234567891011, it can be described as the first number being 1, and the following 10 numbers being incremented by 1; if the data block corresponds to a string of numbers: 1110987654321, it can be described as the first number being 11, and the following 10 numbers being decremented by 1.

[0038] For example, when processing data blocks, if the data string is relatively stable overall, an element extraction method using difference feature description can be used. For instance:

[0039] The sequence 137713771377137713771378137813781378137813761376137613761379137913791379 can be described as follows: the first group of numbers is 1377, which is kept for 7 groups; then +1 is added and kept for 6 groups; then -2 is added and kept for 5 groups; then +3 is added and kept for 6 groups.

[0040] This allows the data block to be abstracted as initial data plus a regular characteristic value, which represents the data distribution pattern within the data. Of course, in this embodiment, the regular characteristic value can also be represented by other numerical values ​​to achieve an encryption effect.

[0041] During the matching process of a data block with feature elements in the data feature element library, if the matched target data feature element is a constant description, the data block can be compressed based on this target data feature element. For example, if a data string exhibits a pattern conforming to a specific constant over a certain continuous length, it can be directly described using the constant's feature element pattern and length. For instance, a string of numbers 3141592653299792458 can be described as the first 10 digits of π plus the first 9 digits of the speed of light. Another example: a string of numbers 33455432112332211556654433221 can be described as the 15 notes of the sheet music for "Ode to Joy" plus the 14 notes of the sheet music for "Twinkle Twinkle Little Star." In this way, data containing constant features can be described using data containing constant feature values. For example, π, "Ode to Joy," and the sheet music for "Twinkle Twinkle Little Star" can all be used as constant feature values. Of course, in this embodiment, constant feature values ​​can also be represented by other numerical values ​​to achieve an encryption effect.

[0042] In the embodiment, during the process of matching a data block with feature elements in the data feature element library, if no corresponding target data feature element is found, for example: 2357111317192329313741434753596167717379838997, it can be identified through data training as an element described by 19 prime numbers starting from 2 and added to the data element feature library as a data feature element.

[0043] Correspondingly, during the decompression or decryption of the data processed above, when data needs to be read, the decryption module or decompression module can be started to call the matching feature elements in the data feature element library to restore the data to the original data format and convert it into data that can be directly recognized for the upper-layer application that needs the data.

[0044] Based on the above embodiments, in another embodiment provided in this disclosure, a data processing method is provided, such as... Figure 1 As shown, the method may include the following steps:

[0045] In step S110, the data to be processed is obtained.

[0046] In this embodiment, the data to be processed can be data that needs to be compressed or data that needs to be encrypted. The data to be processed can specifically be audio or video data, etc. For clarity, this embodiment uses a string as an example, but it is not limited to this.

[0047] In step S120, the data to be processed is matched with feature elements in the data feature element library. The data feature element library includes multiple different data feature elements, which characterize the distribution of characters in the data.

[0048] In this embodiment, a data feature element library can be pre-built, and data feature elements can be added to the library in advance. These data feature elements include repetitive description information, pattern description information, or constant data information, etc. Data feature elements can be further subdivided; for example, pattern description information can be divided into increasing, decreasing, or conforming to a certain function distribution, etc. In this embodiment, the priority of data feature elements can be set, with more common or high-matching-rate data feature elements having a higher priority, in order to improve the matching success rate.

[0049] In step S130, if a target data feature element matching the data to be processed exists in the data feature element library, the data to be processed is processed based on the target data feature element to obtain processed data. This processing may include compression or encryption.

[0050] In this embodiment, the data to be processed can be compressed or encrypted based on target data feature elements. For example, during the matching process between the data to be processed and feature elements in the data feature element library, if the matched target data feature element is a regular descriptive information, the data to be processed can be compressed based on that target data feature element. If the matched target data feature element is a constant descriptive information, the data block can be compressed based on that target data feature element. For details, please refer to the description of the above embodiments.

[0051] In the embodiments provided in this disclosure, when no target data feature element matches the data to be processed in the data feature element library, character distribution information can be extracted from the data to be processed, and the character distribution information can be stored as a feature element in the data feature element library. In these embodiments, the data feature elements in the data feature element library can also be continuously enriched through methods such as data training.

[0052] The data processing method provided in this disclosure matches the data to be processed with feature elements in a data feature element library. If a corresponding target feature data element is matched, the data to be processed can be processed based on that target feature data element to achieve compression or encryption of the data. Since data feature elements represent the distribution of characters in the data, compression or encryption can be performed based on the matched target data feature elements during the processing of the data. This avoids the low processing efficiency caused by a single compression or encryption method, thereby improving the data processing efficiency.

[0053] Based on the above embodiments, in another embodiment provided in this disclosure, when the target data feature element is repetitive descriptive information, the above step S130 may further include the following steps:

[0054] S131, obtain the duplicate and non-duplicate characters in the data to be processed.

[0055] S132, determine the repeating feature value corresponding to the repeating character and the position information of the repeating character in the data to be processed.

[0056] S133, process the data to be processed into data containing repeating feature values, position information and non-repeating characters.

[0057] In this embodiment, when the target data feature element is repetitive descriptive information, it indicates that there is some repetitive data or repetitive parts in the data to be processed. In this case, it is necessary to identify the repetitive characters in the data to be processed, leaving only the non-repetitive characters. By obtaining the position of the repetitive character in the data to be processed and the corresponding repetitive feature value, and combining this with the non-repetitive characters in the data to be processed, the data to be processed can be transformed into containing the repetitive feature value, the position of the repetitive feature value, and the non-repetitive characters. This allows for the processing of the data to be processed, thereby improving data compression efficiency.

[0058] For example, in the above embodiment, the red phone number is 12828282828, which can be processed as: 1 followed by 5 28s. Here, "1" represents a non-repeating character in the data to be processed, "5 28s" represents the position of the repeating character in the data to be processed, i.e., the position of the repeating feature value, and "28" represents the repeating feature value. See the description of the above embodiment for further details.

[0059] Based on the above embodiments, in another embodiment provided in this disclosure, when the target data feature element is regular descriptive information, the above step S130 may further include the following steps:

[0060] S134, Obtain the regularity feature value in the data to be processed. This regularity feature value characterizes the data distribution pattern of the data to be processed.

[0061] S135, process the data to be processed into data containing regular feature values.

[0062] In this embodiment, when the target data feature element is a regular descriptive information, it indicates that at least a portion of the data to be processed contains regular data. Therefore, the regular data in the data to be processed can be obtained through the target data feature element, and then the regular feature value of the regular data can be determined. In this way, by describing the data with regularity in the data to be processed, the data compression effect can be well achieved. For details, please refer to the description in the above embodiment, which will not be repeated here.

[0063] Based on the above embodiments, in another embodiment provided in this disclosure, when the target data feature element is constant data information, the above step S130 may further include the following steps:

[0064] S136, Obtain the constant feature value in the data to be processed. This constant feature value represents the known data in the data to be processed.

[0065] S137, process the data to be processed into data containing constant feature values.

[0066] In this embodiment, when the target data feature elements are regular descriptive information, it indicates that there are at least some constant data in the data to be processed. Therefore, the constant data in the data to be processed can be obtained through the target data feature elements, and then the constant feature value of the constant data can be determined. In this way, by describing the data containing the constant part in the data to be processed, the data compression effect can be well achieved. For details, please refer to the description in the above embodiment, which will not be repeated here.

[0067] Based on the above embodiments, in another embodiment provided in this disclosure, step S120 may further include the following steps:

[0068] S121, the data to be processed is split into multiple data blocks according to a preset method.

[0069] S122, Match multiple data blocks with feature elements in the data feature element library respectively to obtain matching results corresponding to multiple data blocks respectively.

[0070] In this embodiment, if the data to be compressed is large, for example, if the data size exceeds a certain value, it can be split into multiple smaller data blocks. For example, it can be split into data blocks of sizes such as 512B, 1K, 4K, or 4M, as needed. The size of each data block can be the same or different. By splitting the data to be processed into multiple data blocks, and by matching each of these split data blocks with feature elements in the data feature element library, the matching results corresponding to each of these data blocks can be obtained.

[0071] Therefore, in this embodiment, the preset method can be to split the data based on the size of the data blocks, for example, splitting them into multiple data blocks of the same or different sizes; it can also be split according to the type or distribution pattern of the data blocks, for example, splitting data of the same type in the data to be processed into one data block, and splitting data of different types in the data to be processed into different data blocks, etc. In this way, when processing the data to be processed, the multiple split data blocks can be compressed or encrypted separately, thereby improving the processing efficiency of the data to be processed. For example, in this embodiment, each data block can be compressed separately, and the compressed data blocks can be merged to obtain compressed data.

[0072] By dividing each functional module according to its corresponding function, this disclosure provides a data processing device, which can be a server or a chip applied to a server. Figure 2 This is a schematic block diagram of the functional modules of a data processing apparatus provided for an exemplary embodiment of this disclosure. Figure 2 As shown, the data processing device includes:

[0073] Data acquisition module 10 is used to acquire data to be processed;

[0074] The matching module 20 is used to match the data to be processed with feature elements in the data feature element library; wherein, the data feature element library includes multiple different data feature elements, and the data feature elements represent the distribution of characters in the data;

[0075] Processing module 30 is used to process the data to be processed based on the target data feature element when there is a target data feature element in the data feature element library that matches the data to be processed, so as to obtain processed data; the processing method includes compression or encryption.

[0076] In another embodiment provided in this disclosure, the target data feature elements include repetitive description information, regular description information, or constant data information.

[0077] In another embodiment provided in this disclosure, when the target data feature element is repetitive descriptive information, the processing module is specifically used for:

[0078] Obtain duplicate and non-duplicate characters from the data to be processed;

[0079] Determine the repetition feature value corresponding to the repetition character and the position information of the repetition character in the data to be processed;

[0080] The data to be processed is processed into data containing the repeating feature value, the location information, and the non-repeating character.

[0081] In another embodiment provided in this disclosure, when the target data feature elements are regular descriptive information, the processing module is further configured to:

[0082] Obtain regularity feature values ​​from the data to be processed, wherein the regularity feature values ​​characterize the data distribution pattern of the data to be processed;

[0083] The data to be processed is then processed into data containing the regular feature values.

[0084] In another embodiment provided in this disclosure, where the target data feature element is constant data information, the processing module is further configured to:

[0085] Obtain constant feature values ​​from the data to be processed, wherein the constant feature values ​​characterize known data in the data to be processed;

[0086] The data to be processed is then processed into data containing the constant feature values.

[0087] In yet another embodiment provided in this disclosure, the apparatus further includes:

[0088] The information extraction module is used to extract character distribution information from the data to be processed when there is no target data feature element in the data feature element library that matches the data to be processed.

[0089] The character distribution information is stored as a feature element in the data feature element library.

[0090] In another embodiment provided in this disclosure, the matching module is specifically used for:

[0091] The data to be processed is divided into multiple data blocks according to a preset method;

[0092] The multiple data blocks are matched with feature elements in the data feature element library to obtain matching results for each of the multiple data blocks.

[0093] The device part corresponds to the above embodiment, and the specific details are described in the above method embodiment, which will not be repeated here.

[0094] The data processing apparatus provided in this disclosure matches the data to be processed with feature elements in a data feature element library. If a corresponding target feature data element is matched, the data to be processed can be processed based on that target feature data element to achieve compression or encryption of the data. Since data feature elements represent the distribution of characters in the data, compression or encryption can be performed based on the matched target data feature elements during the processing of the data. This avoids the low processing efficiency caused by a single compression or encryption method, thereby improving the data processing efficiency.

[0095] This disclosure also provides an electronic device, including: at least one processor; a memory for storing processor-executable instructions; wherein the at least one processor is configured to execute the instructions to implement the methods disclosed in this disclosure.

[0096] Figure 3 This is a schematic diagram of the structure of an electronic device provided as an exemplary embodiment of this disclosure. For example... Figure 3 As shown, the electronic device 1800 includes at least one processor 1801 and a memory 1802 coupled to the processor 1801. The processor 1801 can perform the corresponding steps in the methods disclosed in the embodiments of this disclosure.

[0097] The processor 1801 described above can also be called a central processing unit (CPU), which can be an integrated circuit chip with signal processing capabilities. Each step in the method disclosed in this embodiment can be implemented by the integrated logic circuitry in the processor 1801 or by software instructions. The processor 1801 can be a general-purpose processor, a digital signal processor (DSP), an ASIC, a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the method disclosed in this embodiment can be directly implemented by a hardware decoding processor, or by a combination of hardware and software modules in the decoding processor. The software modules can be located in the memory 1802, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, or other mature storage media in the art. The processor 1801 reads information from the memory 1802 and, in conjunction with its hardware, completes the steps of the method described above.

[0098] Furthermore, various operations / processes according to this disclosure, implemented via software and / or firmware, can be transmitted from a storage medium or network to a computer system with a dedicated hardware architecture, such as... Figure 4 The computer system 1900 shown is equipped with the programs that constitute the software. When various programs are installed, the computer system is able to perform various functions, including those described above. Figure 4 A block diagram of a computer system provided for an exemplary embodiment of this disclosure.

[0099] Computer System 1900 is intended to represent various forms of digital electronic computer devices, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. Electronic devices can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0100] like Figure 4As shown, the computer system 1900 includes a computing unit 1901, which can perform various appropriate actions and processes based on a computer program stored in a read-only memory (ROM) 1902 or a computer program loaded from a storage unit 1908 into a random access memory (RAM) 1903. The RAM 1903 may also store various programs and data required for the operation of the computer system 1900. The computing unit 1901, ROM 1902, and RAM 1903 are interconnected via a bus 1904. An input / output (I / O) interface 1905 is also connected to the bus 1904.

[0101] Multiple components in computer system 1900 are connected to I / O interface 1905, including: input unit 1906, output unit 1907, storage unit 1908, and communication unit 1909. Input unit 1906 can be any type of device capable of inputting information into computer system 1900. Input unit 1906 can receive input digital or character information and generate key signal inputs related to user settings and / or function control of the electronic device. Output unit 1907 can be any type of device capable of presenting information and may include, but is not limited to, a monitor, speaker, video / audio output terminal, vibrator, and / or printer. Storage unit 1908 may include, but is not limited to, hard disks and optical disks. Communication unit 1909 allows computer system 1900 to exchange information / data with other devices via a network such as the Internet, and may include, but is not limited to, modems, network cards, infrared communication devices, wireless communication transceivers, and / or chipsets, such as Bluetooth™ devices, WiFi devices, WiMax devices, cellular communication devices, and / or the like.

[0102] The computing unit 1901 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 1901 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 1901 performs the various methods and processes described above. For example, in some embodiments, the methods disclosed in this disclosure can be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 1908. In some embodiments, part or all of the computer program can be loaded and / or installed on the electronic device 1900 via ROM 1902 and / or communication unit 1909. In some embodiments, the computing unit 1901 can be configured to perform the methods disclosed in this disclosure by any other suitable means (e.g., by means of firmware).

[0103] This disclosure also provides a computer-readable storage medium, wherein when the instructions in the computer-readable storage medium are executed by a processor of an electronic device, the electronic device is able to perform the methods disclosed in this disclosure.

[0104] The computer-readable storage medium in this disclosure can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. The aforementioned computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specifically, the aforementioned computer-readable storage medium may include electrical connections based on one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0105] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device.

[0106] This disclosure also provides a computer program product, including a computer program, wherein when the computer program is executed by a processor, it implements the methods disclosed in the embodiments of this disclosure.

[0107] In embodiments of this disclosure, computer program code for performing the operations of this disclosure can be written in one or more programming languages ​​or a combination thereof. These programming languages ​​include, but are not limited to, object-oriented programming languages ​​such as Java, Smalltalk, and C++, as well as conventional procedural programming languages ​​such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network (including a local area network (LAN) or a wide area network (WAN)), or it can be connected to an external computer.

[0108] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0109] The modules, components, or units described in the embodiments of this disclosure can be implemented in software or hardware. The names of the modules, components, or units do not necessarily constitute a limitation on the module, component, or unit itself.

[0110] The functions described above in this document can be performed at least in part by one or more hardware logic components. For example, without limitation, exemplary hardware logic components that can be used include: field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), system-on-a-chip (SoCs), complex programmable logic devices (CPLDs), and so on.

[0111] The above description is merely an embodiment of this disclosure and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of this disclosure is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features disclosed in this disclosure that have similar functions.

[0112] While specific embodiments of this disclosure have been described in detail by way of example, those skilled in the art should understand that the examples are for illustrative purposes only and not intended to limit the scope of this disclosure. Those skilled in the art should understand that modifications can be made to the above embodiments without departing from the scope and spirit of this disclosure. The scope of this disclosure is defined by the appended claims.

Claims

1. A data processing method, characterized in that, The method includes: Obtain the data to be processed; The data to be processed is matched with data feature elements in the data feature element library; wherein, the data feature element library includes multiple different data feature elements, and the data feature elements represent the distribution of characters in the data; If a target data feature element that matches the data to be processed exists in the data feature element library, the data to be processed is processed based on the target data feature element to obtain processed data; the processing method includes compression or encryption, and the target data feature element includes repetitive description information, regular description information, or constant data information; Wherein, when the target data feature element is repeated descriptive information, the step of processing the data to be processed based on the target data feature element includes: obtaining repeated characters and non-repeating characters in the data to be processed; determining the repeated feature value corresponding to the repeated character and the position information of the repeated character in the data to be processed; and processing the data to be processed into data containing the repeated feature value, the position information and the non-repeating character.

2. The method according to claim 1, characterized in that, When the target data feature elements are regular descriptive information, the processing of the data to be processed based on the target data feature elements includes: Obtain regularity feature values ​​from the data to be processed, wherein the regularity feature values ​​characterize the data distribution pattern of the data to be processed; The data to be processed is then processed into data containing the regular feature values.

3. The method according to claim 1, characterized in that, When the target data feature elements are constant data information, the processing of the data to be processed based on the target data feature elements includes: Obtain constant feature values ​​from the data to be processed, wherein the constant feature values ​​characterize known data in the data to be processed; The data to be processed is then processed into data containing the constant feature values.

4. The method according to claim 1, characterized in that, The method further includes: If no target data feature element matching the data to be processed exists in the data feature element library, extract the character distribution information from the data to be processed. The character distribution information is stored as a data feature element in the data feature element library.

5. The method according to any one of claims 1 to 4, characterized in that, The step of matching the data to be processed with data feature elements in the data feature element library includes: The data to be processed is divided into multiple data blocks according to a preset method; The multiple data blocks are matched with data feature elements in the data feature element library to obtain matching results for each of the multiple data blocks.

6. A data processing apparatus, characterized in that, The device includes: The data acquisition module is used to acquire the data to be processed. A matching module is used to match the data to be processed with data feature elements in a data feature element library; wherein, the data feature element library includes multiple different data feature elements, and the data feature elements represent the distribution of characters in the data; The processing module is used to process the data to be processed based on the target data feature element when there is a target data feature element in the data feature element library that matches the data to be processed, so as to obtain processed data; the processing method includes compression or encryption, and the target data feature element includes repetitive description information, regular description information or constant data information; Specifically, when the target data feature element is repeated descriptive information, the processing module is used to obtain repeated characters and non-repeating characters in the data to be processed; determine the repeated feature value corresponding to the repeated character and the position information of the repeated character in the data to be processed; and process the data to be processed into data containing the repeated feature value, the position information and the non-repeating character.

7. An electronic device, characterized in that, include: At least one processor; Memory for storing the at least one processor-executable instruction; The at least one processor is configured to execute the instructions to implement the method as described in any one of claims 1-5.

8. A computer-readable storage medium, characterized in that, When the instructions in the computer-readable storage medium are executed by the processor of the electronic device, the electronic device is enabled to perform the method as described in any one of claims 1-5.