Dictionary-based data compression method and device, electronic equipment and medium
By determining the compression dictionary based on the service request and response data associated with the device set of client devices, and encoding the service request and response data, the problem of low compression rate in multiple application service scenarios in the existing technology is solved, and efficient data compression of multiple application services is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-27
- Publication Date
- 2026-03-27
AI Technical Summary
Existing compression dictionaries cannot achieve a high overall compression rate across multiple application service scenarios; they can only achieve a high compression rate for a single application service.
By obtaining the service request code corresponding to the target application, and using a compression dictionary determined based on the service request and service response data associated with the device set to which the client device belongs, the service request code is decompressed, and the service response data is compressed to generate encoded data.
It achieves a high compression rate in multiple application service scenarios with centralized devices, thus improving the overall compression efficiency of data transmission.
Smart Images

Figure CN121749989A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to computer technology, and more particularly to a dictionary-based data compression method, apparatus, electronic device, and medium. Background Technology
[0002] A compression dictionary is used to map characters in the data to be transmitted to encoded sequences, thereby reducing the amount of data that needs to be transmitted by compressing the data. The higher the compression ratio, the less data needs to be transmitted.
[0003] In dictionary-based data compression, the compression dictionary is searched to determine the encoding sequence corresponding to the character sequence in the data to be transmitted, and the character sequence in the data to be transmitted is replaced with the encoding sequence to achieve data compression. In dictionary-based data decompression, the corresponding character in the compression dictionary is looked up based on the encoding sequence and replaced to achieve data decompression. However, the compression dictionaries in related technologies can achieve a high compression rate for the interactive data of a single application service. For interactive data from multiple online application services, a high compression rate cannot be achieved uniformly. Summary of the Invention
[0004] This disclosure provides a dictionary-based data compression method, apparatus, electronic device, and medium, which can improve the data compression rate when client devices access application services.
[0005] In a first aspect, embodiments of this disclosure provide a dictionary-based data compression method, including:
[0006] Obtain the service request code corresponding to the target application, wherein the service request code represents the compression result obtained by compressing the service request of the target application using a compression dictionary, the compression dictionary is determined based on the service request and service response data associated with the device set to which the client device belongs, and the device set is determined based on the service access attributes corresponding to the client device and the target application;
[0007] Based on the device identifier of the client device corresponding to the service request code, obtain the compression dictionary corresponding to the service request code, and use the compression dictionary to decompress the service request code to obtain the service request.
[0008] Obtain the service response data corresponding to the service request, and compress the service response data using the compression dictionary to obtain the encoded data of the service response data.
[0009] Secondly, embodiments of this disclosure provide a dictionary-based data compression apparatus, comprising:
[0010] The request acquisition module is used to acquire the service request code corresponding to the target application. The service request code represents the compression result obtained by compressing the service request of the target application using a compression dictionary. The compression dictionary is determined based on the service request and service response data associated with the device set to which the client device belongs. The device set is determined based on the service access attributes of the client device and the target application.
[0011] The decompression module is used to obtain the compression dictionary corresponding to the service request code based on the device identifier of the client device corresponding to the service request code, and to decompress the service request code using the compression dictionary to obtain the service request.
[0012] The compression module is used to obtain the service response data corresponding to the service request, and to compress the service response data using the compression dictionary to obtain the encoded data of the service response data.
[0013] Thirdly, embodiments of this disclosure also provide an electronic device, including:
[0014] One or more processors;
[0015] Storage device for storing one or more programs.
[0016] When the one or more programs are executed by the one or more processors, the one or more processors implement the dictionary-based data compression method as described in any embodiment of this disclosure.
[0017] Fourthly, embodiments of this disclosure also provide a storage medium containing computer-executable instructions, characterized in that the computer-executable instructions, when executed by a computer processor, are used to perform the dictionary-based data compression method as described in any embodiment of this disclosure.
[0018] This disclosure provides a dictionary-based data compression method. It obtains the service request code corresponding to the target application, which is the compression result obtained by compressing the service request using a compression dictionary. Then, it decompresses the service request code using a compression dictionary corresponding to the device identifier of the client device to obtain the service request. The service response data corresponding to the service request is obtained and compressed using a compression dictionary to obtain the encoded data of the service response data. Since the compression dictionary is determined based on the service request and service response data associated with the device set to which the client device belongs, it can achieve a high compression rate for the interaction data of multiple application services of the target application associated with the device set. This disclosure solves the problem that existing dictionaries can only achieve a high compression rate for a single application service and cannot achieve a high overall compression rate in scenarios involving multiple application services. Attached Figure Description
[0019] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and the originals and elements are not necessarily drawn to scale.
[0020] Figure 1 A flowchart illustrating a dictionary-based data compression method provided in an embodiment of this disclosure;
[0021] Figure 2 This is a data interaction architecture diagram of a target application provided in an embodiment of the present disclosure;
[0022] Figure 3 A schematic diagram of the dictionary training process for a dictionary-based data compression method provided in this embodiment of the present disclosure;
[0023] Figure 4 This is a schematic diagram of the dictionary distribution process in a dictionary-based data compression method provided in an embodiment of this disclosure;
[0024] Figure 5 This is a schematic diagram of the dictionary distribution process in another dictionary-based data compression method provided in this embodiment of the present disclosure;
[0025] Figure 6 This is a schematic diagram of the structure of a dictionary-based data compression device provided in an embodiment of the present disclosure;
[0026] Figure 7 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this disclosure. Detailed Implementation
[0027] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.
[0028] It should be understood that the steps described in the method embodiments of this disclosure may be performed in different orders and / or in parallel. Furthermore, the method embodiments may include additional steps and / or omit the steps shown. The scope of this disclosure is not limited in this respect.
[0029] The term "comprising" and its variations as used herein are open-ended inclusions, meaning "including but not limited to". The term "based on" means "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". Definitions of other terms will be given in the description below.
[0030] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are used only to distinguish different devices, modules or units, and are not used to limit the order of functions performed by these devices, modules or units or their interdependencies.
[0031] It should be noted that the terms "a" and "a plurality of" used in this disclosure are illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".
[0032] The names of messages or information exchanged between multiple devices in the embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of such messages or information.
[0033] It is understood that before using the technical solutions disclosed in the various embodiments of this disclosure, users should be informed of the types, scope of use, and usage scenarios of the personal information involved in this disclosure in an appropriate manner in accordance with relevant laws and regulations, and user authorization should be obtained.
[0034] For example, upon receiving a user's active request, a prompt message is sent to the user to explicitly inform them that the requested operation will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the software or hardware, such as the electronic device, application, server, or storage medium performing the operations of this disclosed technical solution, based on the prompt message.
[0035] As an optional but non-limiting implementation, in response to a user's active request, sending a prompt message to the user can be done via a pop-up window, where the prompt message can be presented in text format. Furthermore, the pop-up window can also include a selection control allowing the user to choose "agree" or "disagree" to provide personal information to the electronic device.
[0036] It is understood that the above notification and user authorization process are merely illustrative and do not constitute a limitation on the implementation of this disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of this disclosure.
[0037] It is understood that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) shall comply with the requirements of relevant laws, regulations and related provisions.
[0038] Figure 1 This is a flowchart illustrating a dictionary-based data compression method provided in an embodiment of this disclosure. This embodiment is applicable to situations requiring optimization of the overall compression rate of network transmission data, such as improving the compression rate of interactive data between multiple application services accessed by a device set in a microservice architecture. The method can be executed by a dictionary-based data compression device, which can be implemented in software and / or hardware, optionally through an electronic device such as a PC, server, or server cluster.
[0039] like Figure 1 As shown, the method includes:
[0040] S110. Obtain the service request code corresponding to the target application.
[0041] The service request encoding characterization is obtained by compressing the service requests of the target application using a compression dictionary. The compression dictionary is determined based on the service request and service response data associated with the device set to which the client device belongs. The device set is determined based on the service access attributes corresponding to the client device and the target application.
[0042] The target application is the application that provides application services. Application services refer to services within an application built on a microservices architecture. Microservices are a software development architecture where an application consists of small, independent services that communicate through well-defined APIs.
[0043] A compression dictionary can be a data structure that stores keys and compressed characters in key-value pairs. The string in the service request is matched against the keys in the compression dictionary to obtain at least one compressed character. This at least one compressed character is then used to replace the corresponding string in the service request, resulting in the service request code. At this point, the service request code represents the compressed character sequence composed of the compressed characters.
[0044] Service access attributes include service access frequency and / or service access duration, etc. This disclosure does not limit the specific meaning of service access attributes. For example, devices accessing the target application can be categorized based on the frequency of client devices accessing the target application's services, resulting in a device set. Since different client devices access application services differently, the frequency of client devices accessing the target application's application services can be determined based on the client devices' historical access information. For example, if client device A accesses application service a 10 times and application service b 30 times in one day, then client device A's access frequency for that day is 40 times. If client device B accesses application service a 8 times, application service c 20 times, and application service d 50 times in the same day, then client device B's access frequency for that day is 78 times. Client devices can be divided according to the access frequency of application services, resulting in a device set. For example, client devices accessing the target application's application services 10-100 times in one day can be grouped together; then client device A and client device B belong to the same device set. Since the frequency of client device access to application services changes in real time, if the service access frequency of a client device increases to the point that it exceeds the upper limit of the service access frequency of its device set, the client device can be reassigned to a device set at the next higher level. For example, if the service access frequency of client device B becomes 120 times, client device B can be moved to a device set with a service access frequency between 101 and 500. If the service access frequency of a client device decreases to the point that it falls below the lower limit of the service access frequency of its device set, the client device can be reassigned to a device set at the next lower level. For example, D1, D2, ..., Dn represent n device sets. Assume D1 represents the device set with an application service access frequency between 1 and 10, and D2 represents the device set with an application service access frequency between 10 and 100. If the service access frequency of a client device corresponding to D1 increases to 50 times due to user activity, the client device can be moved to D2.
[0045] Optionally, devices accessing the target application can be categorized based on the duration of their access to the target application's services, thus obtaining a device set. Alternatively, devices accessing the target application can be categorized based on the frequency and duration of their access to the target application's services, thus obtaining a device set.
[0046] Based on the service request and response data of the target applications corresponding to the client devices in the device set, determine the service request and response data associated with the device set. Update the dictionary configuration parameters according to the dictionary update cycle to obtain the updated compressed dictionary. Verify whether the global compression ratio of the updated compressed dictionary is higher than the global compression ratio of the compressed dictionary in the previous dictionary update cycle based on the service request and response data associated with the device set. If so, save the updated compressed dictionary to the dictionary database. Distribute the updated compressed dictionary to the client devices in the device set and record the association between the client device's device identifier and the dictionary identifier.
[0047] The dictionary database includes compressed dictionaries and configuration information corresponding to different device sets. The configuration information includes application service levels, domain information, dictionary update cycles, and training target levels. For example, application service levels can include the priority of application services. Domain information includes application services corresponding to fixed domains and application services corresponding to dynamic domains. Service request and response data for application services within a fixed domain can be grouped into an application service set, corresponding to a compressed dictionary. This compressed dictionary does not need to be updated after training. Application services corresponding to fixed domains are typically important but infrequently accessed application services, such as APIs (Application Programming Interfaces) that must be accessed during client startup. Application services corresponding to dynamic domains typically represent APIs related to specific usage scenarios. To reduce the amount of data updated and the computational cost of training, inherent key-value pairs are determined based on application service data that all client devices need to access. These inherent key-value pairs remain unchanged during dictionary updates and do not require further updates or training. Furthermore, the compressed dictionary corresponding to the fixed domain is used as the fixed dictionary. The fixed dictionary remains unchanged during dictionary updates and does not require further updates or training.
[0048] The server distributes compressed dictionaries to client devices in different device sets. The compressed dictionaries for different device sets may be the same or different. The dictionary update cycles for different device sets can also be the same or different. For example, a uniform dictionary update cycle can be set for the compressed dictionaries of all device sets. Alternatively, different dictionary update cycles can be set for different device sets based on the numerical range of application service access frequency or performance attributes corresponding to the device set. The dictionary update cycle can be different time periods such as month, week, day, or hour. If a device set meets the dictionary update cycle, the dictionary configuration parameters of the compressed dictionary corresponding to that device set are adjusted to perform dictionary training.
[0049] To reduce the amount of data used for dictionary training, and thus reduce the computational consumption of services, service request and response data associated with the device set can be sampled to obtain a training dataset. Optionally, the sampling rate is determined based on the proportion of the service request and response data volume of an application service in the total amount of service request and response data associated with the device set. For ease of understanding, assume that client devices in the device set access three application services of the target application. The access frequency of each application service can be multiplied by its data volume and then summed to obtain the data volume corresponding to each application service. The data volumes corresponding to each application service are then summed to obtain the total data volume for the device set. Based on the proportion of the data volume corresponding to each application service in the total data volume, the sampling rate corresponding to each application service is determined. Service request and response data corresponding to that application service in the device set are sampled according to the sampling rate corresponding to each application service, so that application services with high data volumes also have relatively high data volumes in the training dataset.
[0050] Optionally, if the application service has a higher priority, the sampling rate of the application service will be higher, which can ensure that more service request and service response data of important application services are sampled.
[0051] Optionally, the sampling rate can be set according to actual needs, and this disclosure does not impose specific limitations.
[0052] For example, the service request and service response data of the application services associated with the device set are sampled according to a set sampling rate to obtain the training dataset P1, P2, ..., Pi, ..., Pn. Pi represents the service request and service response data of the i-th application service in the application services associated with the device set, i∈[1,n].
[0053] For example, the server associated with the target application obtains the service request code corresponding to the target application sent by the client device.
[0054] S120. Based on the device identifier of the client device corresponding to the service request code, obtain the compression dictionary corresponding to the service request code, and use the compression dictionary to decompress the service request code to obtain the service request.
[0055] Decompression refers to restoring the encoded service request to its original string form.
[0056] Since the compression dictionary required for decompressing the service request code needs to be the same as the compression dictionary used by the client device, the version of the client device's compression dictionary needs to be determined after obtaining the service request code. Because the server records the association between the client device's device identifier and the dictionary identifier, the dictionary identifier can be retrieved based on the client identifier. Therefore, the compression dictionary corresponding to the service request code can be obtained based on the dictionary identifier. The compression dictionary used during compression can be retrieved based on the compressed character sequence in the service request code to obtain the corresponding string. Arranging these strings sequentially constitutes the service request.
[0057] S130. Obtain the service response data corresponding to the service request, and compress the service response data using the compression dictionary to obtain the encoded data of the service response data.
[0058] Service response data refers to the execution result data of the application service in fulfilling the service request. The server needs to send the service response data corresponding to the service request to the client. Encoded data refers to the compressed result obtained by compressing the service response data using a compression dictionary. Since the data needs to be compressed before being sent, a compression dictionary can be used to compress the service response data to obtain the encoded service response data.
[0059] For example, the strings in the service response data are matched against the keys in the compression dictionary to obtain at least one target compressed character. The corresponding strings in the service response data are then replaced with the at least one target compressed character to obtain the encoded service response data. The encoded service response data corresponding to the service request of the target application is then sent to the client device.
[0060] In some embodiments, the server may store service request and response data for the application service in a database. The stored service request and response data are then used for dictionary training. Figure 2 This is a data interaction architecture diagram for a target application provided in an embodiment of this disclosure. For example... Figure 2As shown, the data interaction process between client device 210 and server 220 includes: compressing the service request based on the compression dictionary 230 built into client device 210 to obtain a service request code. Then, the service request code is sent to server 220 through network library 280. Server 220 aggregates the service request codes from client device 210 through gateway 240. Gateway 240 decompresses the service request codes based on the compression dictionary corresponding to client device 210 to obtain the service request. Then, the service request is sent to the corresponding application service 250. Application service 250 responds to the service request from client device 210, returning service response data. Server 220 copies the service response data generated by each application service 250 and the service request sent by client device 210, and stores the service request and service response data in database 260. The service request and service response data in database 260 are used for dictionary training 270.
[0061] This embodiment of the disclosure obtains the service request code corresponding to the target application. The service request code is the result of compressing the service request using a compression dictionary. Then, the service request code is decompressed using a compression dictionary corresponding to the device identifier of the client device to obtain the service request. Service response data corresponding to the service request is obtained, and the service response data is compressed using a compression dictionary to obtain the encoded data of the service response data. Since the compression dictionary is determined based on the service request and service response data associated with the device set to which the client device belongs, a high compression rate can be achieved for the interaction data of multiple application services of the target application associated with the device set. This embodiment of the disclosure solves the problem that existing dictionaries can only achieve a high compression rate for a single application service and cannot achieve a high overall compression rate in multiple application service scenarios.
[0062] Figure 3 This is a schematic diagram of the dictionary training process of a dictionary-based data compression method provided in this embodiment of the present disclosure. Based on the above embodiments, this embodiment of the present disclosure specifically defines the training method of the compressed dictionary.
[0063] like Figure 3 As shown, the method includes:
[0064] S310. Based on the service access attributes corresponding to the client devices and the target application, the client devices associated with the target application are grouped to obtain at least one set of devices.
[0065] The client devices associated with the target application can represent the client devices that have accessed the target application. The device set to which a client device belongs can be determined based on the service access attributes corresponding to the client device and the target application.
[0066] Pre-define the numerical range of service access frequency for each device set. For example, set device sets with service access frequencies of 5-10, 10-100, 100-1000, 1000-10000, and more than 10000. For newly installed clients, the client device has zero service access frequency for the target application and does not belong to any device set. However, as users use the client, the service access frequency for the target application increases. If the current service access frequency exceeds 5 times, the client device is assigned to the device set with a service access frequency of 5-10. Then, as the client is used, if the current service access frequency exceeds 10 times, the client device is moved to the device set with a service access frequency of 10-100.
[0067] The server determines the device set to which a client device belongs based on the frequency of service access at set intervals. For example, the server counts the frequency of service accesses between client devices and the target application, and groups the client devices according to the service access frequency to obtain at least one device set.
[0068] S320. Based on the service request and service response data of the application service associated with any device set, generate a service data subset of the application service, and construct the application service dataset based on the service data subset.
[0069] Here, a service data subset represents a collection of service request and service response data for the same application service. Since client devices in a device set may access different application services, there are multiple service data subsets associated with the device set. These multiple service data subsets constitute the application service dataset.
[0070] If Pi represents the service request and service response data of the i-th application service in the application services associated with the device set, where i∈[1,n], then Pi is a subset of the service data. The application service dataset can be represented as P1, P2, ..., Pi, ..., Pn.
[0071] For example, the server stores the service request and service response data corresponding to the application service in a database, and determines the device set to which the client device belongs based on the service access frequency of the application service at set intervals. The server then categorizes the service request and service response data associated with the device set according to the application service, obtaining the service request and service response data for each application service, thus generating a service data subset for each application service. This service data subset constitutes the application service dataset.
[0072] S330. For any set of devices, train the compressed dictionary based on the application service dataset and dictionary configuration parameters associated with the set of devices.
[0073] The dictionary configuration parameters include the length of the index values and the length of the data search window. Optionally, dictionary training objectives can be pre-set, such as compression ratio priority, speed priority, or a balance between compression ratio and speed. Different training objectives can correspond to different ranges of dictionary configuration parameters. Optionally, more granular objectives can be defined based on the dictionary training objectives to achieve fine-tuning. For example, under the compression ratio priority objective, multiple granular objectives at different levels can be set, with different granular objectives corresponding to different ranges of dictionary configuration parameters. Different dictionary training objectives can also be set for the compressed dictionaries corresponding to different device sets based on the service access attributes of different device sets.
[0074] If the compressed dictionary corresponding to at least one device set meets the dictionary update cycle, the dictionary training target for the corresponding client cluster is obtained, and the corresponding dictionary configuration parameter range is determined based on the dictionary training target. The length of the index values and the length of the data search window in the compressed dictionary are adjusted based on the dictionary configuration parameter range, and the compressed dictionary is updated based on different combinations of index value lengths and data search window lengths. Then, the updated compressed dictionary is used to compress the application service dataset associated with the device set. The compression rate of the application service dataset associated with the device set is determined based on the compression rate of each service data subset in the application service dataset. If it meets the preset training termination condition, training is terminated; otherwise, the dictionary configuration parameters are updated, and the process returns to update the compressed dictionary according to the updated dictionary configuration parameters.
[0075] For example, for any set of devices, training the compressed dictionary based on the application service dataset associated with the set of devices and dictionary configuration parameters includes:
[0076] For any device set, if the dictionary configuration parameters are updated, the compressed dictionary is updated according to the updated dictionary configuration parameters. The updated compressed dictionary is used to compress the service data subset associated with the device set to obtain a service data subset encoding. An individual compression ratio is determined based on the service data subset encoding and the service data subset itself. A global compression ratio is determined based on the weight of the service data subset and the individual compression ratio, wherein the individual compression ratio represents the compression ratio associated with the service data subset, and the global compression ratio represents the compression ratio associated with the application service dataset. If the global compression ratio does not meet a preset training termination condition, the dictionary configuration parameters are updated, wherein the preset training termination condition includes the global compression ratio being higher than the global compression ratio associated with the previous dictionary update cycle.
[0077] Among them, the service data subset encoding represents the compressed data of the service request and service response data corresponding to the service data subset.
[0078] For example, updating the dictionary configuration parameters includes updating the key value length and / or search window length according to a set step size.
[0079] In this embodiment of the disclosure, updating the compressed dictionary according to the updated dictionary configuration parameters includes: searching a subset of service data according to the updated dictionary configuration parameters to obtain candidate key values and their occurrence frequencies. For example, if the dictionary configuration parameters include an index value length of 3 characters and a data search window length of 100 characters, then for the service request and service response data corresponding to the service data subset, a search is performed within the search range of 100 characters, with 3 characters as the search object, to obtain the occurrence frequency of the search object. The target key value is determined from the candidate key values based on the occurrence frequency, and the target compressed character is determined based on the target key value. The compressed dictionary is then updated based on the target key value and the target compressed character.
[0080] It should be noted that the dictionary training method employs unsupervised reinforcement learning and heuristic search techniques. In the initial stage, a baseline compressed dictionary is generated based on historical interaction data of each application service within the target application. The server distributes this baseline compressed dictionary to newly installed client devices of the target application. As users utilize the target application, the frequency of accesses accumulates. When the server detects that a client device's access frequency to the application service falls within the service access frequency range of a certain device set, it adds the client device to that device set and distributes the corresponding compressed dictionary to that device. When the dictionary update cycle for the device set is met, the dictionary configuration parameters are adjusted, and dictionary training is performed based on the application service dataset associated with that device set and the dictionary configuration parameters to obtain a new compressed dictionary. The new compressed dictionary is used to compress the service request and service response data of each service data subset in the application service dataset, resulting in service data subset encodings. The individual compression rate is determined based on the data volume of the service data subset encodings and the data volume of the service data subsets themselves.
[0081] Optionally, determining the global compression ratio based on the weights of the service data subsets and the individual compression ratios includes: determining the weight of the service data subset based on its proportion of the data volume in the application service dataset; and performing a weighted fusion process based on the weights of the service data subsets and the individual compression ratios to obtain the global compression ratio.
[0082] Understandably, the greater the usage of an application service, the higher the proportion of data corresponding to the service data subset within the total data volume of the application service dataset, and thus the greater the weight of the service data subset. The overall compression rate is obtained by weighted summation of the individual compression rates based on the weights of the service data subsets.
[0083] Optionally, after determining the weight of the service data subset, the method further includes updating the weight according to the priority of the application service corresponding to the service data subset. For example, for application services in the startup phase, the access frequency is relatively low, resulting in a relatively low weight for the service data subset. However, application services in the startup phase have a high priority, and correspondingly, the weight of the service data subset of application services in the startup phase can be increased.
[0084] Since the preset training termination condition includes the global compression ratio of the current dictionary update cycle being higher than the global compression ratio associated with the previous dictionary update cycle, indicating that the global compression ratio of the compressed dictionary obtained through current training is improved compared to the previous dictionary update cycle, dictionary training can be terminated. If the global compression ratio of the compressed dictionary obtained through current training is not higher than the global compression ratio of the previous dictionary update cycle, then the dictionary configuration parameters need to be updated, and dictionary training should continue.
[0085] It should be noted that there may be cases where the global compression rate is lower than the compression rate of a particular individual dictionary. However, as long as the global compression rate of the compressed dictionary obtained by the current training is improved compared to the previous dictionary update cycle, the training result is considered to meet the preset training termination condition.
[0086] It should be noted that if a new compressed dictionary trained based on all dictionary configuration parameters within the dictionary configuration parameter range fails to improve the global compression rate, then the historical dictionary before the new dictionary is issued is considered optimal for the current scenario, and there is no need to update the dictionary.
[0087] This disclosure embodiment groups client devices associated with the target application based on the service access attributes corresponding to the client devices and the target application, obtaining at least one device set. Then, a service data subset is generated based on the service request and service response data of the application services associated with the device set. An application service dataset is constructed based on the service data subset, and a compressed dictionary corresponding to the device set is trained according to the application service dataset and dictionary configuration parameters. This ensures that the global compression rate of the trained compressed dictionary associated with the application service dataset is higher than the global compression rate associated with the previous dictionary update cycle, achieving a higher compression rate for the client devices in the device set. This disclosure embodiment solves the problem that existing dictionaries can only achieve a high compression rate in a single application service scenario, and cannot achieve a high compression rate in multiple application service scenarios.
[0088] Figure 4This is a schematic diagram of the dictionary distribution process in a dictionary-based data compression method provided by an embodiment of the present disclosure. Based on the above embodiments, this embodiment of the present disclosure specifically defines the step of distributing the compressed dictionary to the client device.
[0089] like Figure 4 As shown, the method includes:
[0090] S410. Based on the service access attributes corresponding to the client devices and the target application, the client devices associated with the target application are grouped to obtain at least one set of devices.
[0091] S420. Based on the service request and service response data of the application service associated with any device set, generate a service data subset of the application service, and construct the application service dataset based on the service data subset.
[0092] S430. For any device set, if the dictionary configuration parameters are updated, the compressed dictionary is updated according to the updated dictionary configuration parameters, and the service data subset associated with the device set is compressed using the updated compressed dictionary to obtain the service data subset encoding.
[0093] S440. Determine the individual compression ratio based on the service data subset encoding and the service data subset, and determine the global compression ratio based on the weight of the service data subset and the individual compression ratio.
[0094] Wherein, the individual compression rate represents the compression rate associated with the service data subset, and the global compression rate represents the compression rate associated with the application service dataset.
[0095] S450: Determine whether the global compression rate meets the preset training termination condition. If not, proceed to S460; otherwise, proceed to S470.
[0096] If the global compression rate associated with the compressed dictionary obtained from the current training is higher than the global compression rate associated with the previous dictionary update cycle, then the global compression rate is determined to meet the preset training termination condition; otherwise, the global compression rate is determined not to meet the preset training termination condition.
[0097] S460. Update the dictionary configuration parameters.
[0098] For example, the key value length is updated according to a set step size. Optionally, the search window length is updated according to a set step size. Optionally, both the key value length and the search window length are updated according to a set step size.
[0099] S470. Sort the service data subsets in descending order according to their weights, and determine the target service data subset based on the ranking.
[0100] Since the method for determining the weights of service data subsets has been described in the above embodiments, it will not be repeated here. Based on the weights of the service data subsets associated with the device set, the service data subsets are sorted in descending order to obtain a ranking. The top n service data subsets are determined as the target service data subsets. That is, the n service data subsets with higher weights are determined as the target service data subsets.
[0101] S480. Determine whether the individual compression rate of the target service data subset is higher than the individual compression rate corresponding to the previous dictionary update cycle. If yes, execute S490; otherwise, execute S460.
[0102] S490. Send the updated compressed dictionary to the client devices in the device set.
[0103] For example, if the individual compression rate of at least one subset of target service data is higher than the individual compression rate corresponding to the previous dictionary update cycle, then the updated compressed dictionary is sent to the client devices in the device set.
[0104] Optionally, if the individual compression ratio of all target service data subsets is higher than the individual compression ratio corresponding to the previous dictionary update cycle, then the updated compressed dictionary is sent to the client devices in the device set.
[0105] This embodiment of the disclosure determines whether to send an updated compression dictionary to the client devices in the device set by judging whether the individual compression rate of the service data subset with higher weight has improved compared to the previous dictionary update cycle. This can ensure that the service request and service response data of the application service with higher weight can achieve a higher compression rate. Since the access volume of the application service with higher weight is usually large, the bandwidth consumption of the encoded data corresponding to the application service with higher weight can be reduced by the higher compression rate.
[0106] Figure 5 This is a schematic diagram of the dictionary distribution process in another dictionary-based data compression method provided in this disclosure embodiment. Based on the above embodiments, this disclosure embodiment defines the dictionary rollback method.
[0107] S501. Based on the service access attributes corresponding to the client devices and the target application, the client devices associated with the target application are grouped to obtain at least one set of devices.
[0108] S502. Based on the service request and service response data of the application service associated with any device set, generate a service data subset of the application service, and construct the application service dataset based on the service data subset.
[0109] S503. For any device set, if the dictionary configuration parameters are updated, the compressed dictionary is updated according to the updated dictionary configuration parameters, and the service data subset associated with the device set is compressed using the updated compressed dictionary to obtain the service data subset encoding.
[0110] S504. Determine the individual compression ratio based on the service data subset encoding and the service data subset, and determine the global compression ratio based on the weight of the service data subset and the individual compression ratio.
[0111] S505. Determine whether the global compression rate meets the preset training termination condition. If not, execute S506; otherwise, execute S507.
[0112] S506. Update the dictionary configuration parameters.
[0113] S507. Sort the service data subsets in descending order according to their weights, and determine the target service data subset based on the ranking.
[0114] S508. Determine whether the individual compression rate of the target service data subset is higher than the individual compression rate corresponding to the previous dictionary update cycle. If yes, execute S509; otherwise, execute S506.
[0115] S509. Send the updated compressed dictionary to the client devices in the device set.
[0116] S510. Based on the business and performance indicators of the client device before and after the compression dictionary is distributed, determine the indicator deviation value associated with the updated compression dictionary.
[0117] Because higher compression ratios consume more resources on the client device—for example, requiring more resources for data decompression—a higher compression ratio is not always better. It's necessary to consider both the client device's performance metrics and business metrics. Business metrics include video playback duration and startup time. Startup time represents the time it takes to open the first frame of the video. Performance metrics include CPU utilization, memory usage, stuttering, crashes, and memory overflows.
[0118] S511. Determine whether the deviation value of the indicator is less than the preset indicator change threshold. If yes, execute S512; otherwise, execute S506.
[0119] For example, after the compression dictionary is distributed, the performance metrics generated during the data compression and decompression processes performed by the client device using the compression dictionary are obtained. Additionally, the business metrics of the client device after using the compression dictionary are obtained. The obtained performance and business metrics are compared with the performance and business metrics associated with the corresponding compression dictionary from the previous dictionary update cycle stored in the database. For any one of the performance and business metrics, the metric deviation value is obtained by subtracting the metric value before the compression dictionary distribution from the metric value after the compression dictionary distribution.
[0120] S512. Obtain the compressed dictionary corresponding to the previous dictionary update cycle of the current dictionary update cycle, and send the compressed dictionary corresponding to the previous dictionary update cycle to the client device.
[0121] For example, if the deviation between the performance metric and a specific business metric is less than a preset metric change threshold, then when the dictionary update cycle is met, the server sends the historical compressed dictionary from before the current dictionary update to the client device, so that the client device uses the compressed dictionary from before the new dictionary update. That is, the compressed dictionary associated with the previous dictionary update cycle is sent directly, without the need for dictionary training.
[0122] Alternatively, if the number of indicators with deviation values less than a preset indicator change threshold exceeds a set threshold, then when the dictionary update cycle is met, the server will send the historical compressed dictionary before the current update to the client device.
[0123] Optionally, if the indicator deviation value is greater than or equal to the preset indicator change threshold, then return to execute S506 to update the dictionary configuration parameters. The updated compressed dictionary and configuration parameters generate a new compressed dictionary.
[0124] In one specific embodiment, the server establishes background statistics to count the frequency of service accesses to the application service from different client devices. Client devices are then divided into different device sets based on their service access frequency. Service request and response data for the application service are sampled from these different device sets to obtain an application service dataset. The application service dataset includes at least one subset of service data from an application service.
[0125] A DCBR (Dict Compression Boost Rate) can be set. Assume the individual compression rates of P1, ..., Pn are CRp1~n, and the global compression rate is ΣRi*CRpi. Here, each Pi corresponds to the service request and response data of the i-th application service in the application service associated with the device set, and Ri is the proportion (weight) of Pi's data volume in the application service dataset. Ri can be adjusted according to requirements, such as based on the request volume of different application services; a higher request volume results in a higher proportion and weight for Pi. Alternatively, the weight can be adjusted according to the priority of the application service; important application services, such as critical data acquisition services during the startup phase, have higher priority and greater weight.
[0126] DCBR(P1, ..., Pn) = ΣRi*CRpi / CRpi, where i = 1, 2, ..., n, and CRpi represents the individual compression ratio of Pi. The global compression ratio is obtained by merging the individual compression ratios of each service data subset based on weights. The dictionary training objective can be set to a global compression ratio higher than the global compression ratio before the new dictionary was distributed. If the compressed dictionary obtained through dictionary training results in a global compression ratio higher than the global compression ratio of the compressed dictionary determined in the previous dictionary update cycle, then the global compression ratio satisfies the training objective, and training can end. It is also possible that the global compression ratio is smaller than the compression ratio of a single individual component, i.e., DCBR(P1, ..., Pn) < 1, but as long as the global compression ratio is higher than the global compression ratio of the compressed dictionary determined in the previous dictionary update cycle, training can end.
[0127] It should be noted that, in addition to using the global compression rate as the training objective, training can also be combined with real-time online data feedback, using the improvement of business indicators of the centralized client devices as the training objective. That is, even if the global compression rate does not increase, but the decompression cost of the client devices is saved, and the local performance of the client devices is improved, thus achieving benefits, then the global compression rate at this time can be considered to meet the training objective, and training can be terminated.
[0128] The training process includes: In the initial stage, a baseline compressed dictionary is obtained based on API service request data, and this is provided to newly installed devices of the target application. As new devices use the target application, the frequency of their access to different APIs within the application gradually changes. When the access frequency falls within a certain range corresponding to a specific device set according to a clustering standard, the compressed dictionary corresponding to that device set is sent to the new device. Since the data generated and received by the application service is constantly changing, different device sets need to periodically train and update the dictionary. The amount of data for the corresponding API's Pi will vary depending on the characteristics of the device set (based on access frequency, API importance, or a combination of both). Furthermore, the training result must satisfy the requirement that the CRpi of Pi with high Ri (e.g., the top 10 Ri, the range of which can be fine-tuned) is higher than before training. During training, adjustable parameters include, but are not limited to, adjusting the key length in the dictionary (e.g., from 1 byte to 100 bytes) and the search window length for the training data. The current training round can be stopped when the global compression rate associated with the new compressed dictionary obtained through multiple training iterations is better than the global compression rate before the new dictionary was distributed. Alternatively, if no better result is obtained, the current dictionary is considered optimal for the current scenario. Specifically, data reported from client devices can be statistically analyzed to determine whether the deviation of metrics before and after applying the new compressed dictionary on different client devices exceeds a preset metric change threshold. If so, when the next round of dictionary updates does not improve the corresponding performance and business, the system will revert to the more advantageous dictionary (the old dictionary) and distribute it to the client devices.
[0129] This embodiment of the disclosure compares the business metrics and performance metrics before and after the compressed dictionary is distributed to determine the metric deviation value associated with the updated compressed dictionary. If the metric deviation value is less than a preset metric change threshold, the compressed dictionary corresponding to the previous dictionary update cycle of the current dictionary update cycle is distributed to the client device to reduce the impact of the use of the compressed dictionary on the business and performance of the client device.
[0130] Figure 6 This is a schematic diagram of a dictionary-based data compression device provided in an embodiment of the present disclosure. The device can be implemented in software and / or hardware, and optionally, it can be implemented in an electronic device, such as a PC, a server, or a server cluster.
[0131] like Figure 6 As shown, the device includes: a request acquisition module 610, a decompression module 620, and a compression module 630.
[0132] The request acquisition module 610 is used to acquire the service request code corresponding to the target application. The service request code represents the compression result obtained by compressing the service request of the target application using a compression dictionary. The compression dictionary is determined based on the service request and service response data associated with the device set to which the client device belongs. The device set is determined based on the service access attributes of the client device and the target application.
[0133] The decompression module 620 is used to obtain the compression dictionary corresponding to the service request code based on the device identifier of the client device corresponding to the service request code, and use the compression dictionary to decompress the service request code to obtain the service request.
[0134] Compression module 630 is used to obtain service response data corresponding to the service request, and to compress the service response data using the compression dictionary to obtain the encoded data of the service response data.
[0135] Optionally, the training method for the compressed dictionary includes:
[0136] Based on the service access attributes corresponding to the client devices and the target application, the client devices associated with the target application are grouped to obtain at least one set of devices.
[0137] Based on the service request and service response data of the application service associated with any device set, a service data subset of the application service is generated, and the application service dataset is constructed based on the service data subset.
[0138] For any set of devices, the compressed dictionary is trained based on the application service dataset and dictionary configuration parameters associated with the set of devices.
[0139] Further, for any set of devices, training the compressed dictionary based on the application service dataset associated with the set of devices and dictionary configuration parameters includes:
[0140] For any set of devices, if the dictionary configuration parameters are updated, the compressed dictionary is updated according to the updated dictionary configuration parameters, and the service data subset associated with the set of devices is compressed using the updated compressed dictionary to obtain the service data subset encoding.
[0141] The individual compression ratio is determined based on the service data subset encoding and the service data subset, and the global compression ratio is determined based on the weight of the service data subset and the individual compression ratio, wherein the individual compression ratio represents the compression ratio associated with the service data subset, and the global compression ratio represents the compression ratio associated with the application service dataset;
[0142] If the global compression ratio does not meet the preset training termination condition, the dictionary configuration parameters are updated, wherein the preset training termination condition includes the global compression ratio being higher than the global compression ratio associated with the previous dictionary update cycle.
[0143] Further, determining the global compression rate based on the weights of the service data subset and the individual compression rates includes:
[0144] The weight of the service data subset is determined based on the proportion of the data volume corresponding to the service data subset in the data volume corresponding to the application service dataset.
[0145] The global compression rate is obtained by weighted fusion processing based on the weight of the service data subset and the individual compression rate.
[0146] Optionally, after determining the weights of the service data subset, the method further includes:
[0147] The weights are updated based on the priority of the application services corresponding to the subset of service data.
[0148] Optionally, the dictionary configuration parameters include key-value length and search window length;
[0149] The updating of the dictionary configuration parameters includes:
[0150] Update the key value length and / or search window length according to the set step size.
[0151] Optionally, it also includes:
[0152] If the global compression rate meets the preset training termination condition, the service data subset is sorted in descending order according to its weight, and the target service data subset is determined based on the ranking.
[0153] If the individual compression rate of the target service data subset is higher than the individual compression rate corresponding to the previous dictionary update cycle, then the updated compressed dictionary is sent to the client devices in the device set.
[0154] Optionally, after sending the updated compressed dictionary to the client devices in the device set, the method further includes:
[0155] Based on the business and performance indicators of the client device before and after the compression dictionary is distributed, determine the indicator deviation value associated with the updated compression dictionary;
[0156] If the deviation value of the indicator is less than the preset indicator change threshold, then the compressed dictionary corresponding to the previous dictionary update cycle of the current dictionary update cycle is obtained, and the compressed dictionary corresponding to the previous dictionary update cycle is sent to the client device.
[0157] The dictionary-based data compression apparatus provided in this disclosure can execute the dictionary-based data compression method provided in any embodiment of this disclosure, and has the corresponding functional modules and beneficial effects of executing the method.
[0158] It is worth noting that the various units and modules included in the above-mentioned device are only divided according to functional logic, but are not limited to the above division, as long as the corresponding functions can be realized; in addition, the specific names of each functional unit are only for easy differentiation and are not used to limit the protection scope of the embodiments of this disclosure.
[0159] Figure 7 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this disclosure. Reference is made below. Figure 7 It illustrates an electronic device suitable for implementing embodiments of the present disclosure (e.g., Figure 7 The diagram below shows the structure of the terminal device or server 700. The terminal device in this embodiment may include, but is not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), and vehicle terminals (e.g., vehicle navigation terminals), as well as fixed terminals such as digital TVs and desktop computers. Figure 7 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments disclosed herein.
[0160] like Figure 7 As shown, the electronic device 700 may include a processing unit (e.g., a central processing unit, a graphics processing unit, etc.) 701, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 702 or a program loaded from a storage device 708 into a random access memory (RAM) 703. The RAM 703 also stores various programs and data required for the operation of the electronic device 700. The processing unit 701, ROM 702, and RAM 703 are interconnected via a bus 704. An edit / output (I / O) interface 705 is also connected to the bus 704.
[0161] Typically, the following devices can be connected to I / O interface 705: input devices 706 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 707 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 708 including, for example, magnetic tapes, hard disks, etc.; and communication devices 709. Communication device 709 allows electronic device 700 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 7An electronic device 700 with various devices is shown; however, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively.
[0162] In particular, according to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via communication device 709, or installed from storage device 708, or installed from ROM 702. When the computer program is executed by processing device 701, it performs the functions defined in the methods of embodiments of this disclosure.
[0163] The names of messages or information exchanged between multiple devices in the embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of such messages or information.
[0164] The electronic device provided in this embodiment belongs to the same inventive concept as the method provided in the above embodiments. Technical details not described in detail in this embodiment can be found in the above embodiments, and this embodiment has the same beneficial effects as the above embodiments.
[0165] This disclosure provides a computer storage medium storing a computer program that, when executed by a processor, implements the dictionary-based data compression method provided in the above embodiments.
[0166] It should be noted that the computer-readable medium described in this disclosure can be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this disclosure, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in connection with an instruction execution system, apparatus, or device. In this disclosure, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.
[0167] In some implementations, clients and servers can communicate using any currently known or future-developed network protocol such as HTTP (Hypertext Transfer Protocol) and can interconnect with digital data communication (e.g., communication networks) of any form or medium. Examples of communication networks include local area networks (“LANs”), wide area networks (“WANs”), the Internet (e.g., the Internet of Things), and peer-to-peer networks (e.g., ad hoc peer-to-peer networks), as well as any currently known or future-developed networks.
[0168] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device.
[0169] The aforementioned computer-readable medium carries one or more programs that, when executed by the electronic device, cause the electronic device to:
[0170] Obtain the service request code corresponding to the target application, wherein the service request code represents the compression result obtained by compressing the service request of the target application using a compression dictionary, the compression dictionary is determined based on the service request and service response data associated with the device set to which the client device belongs, and the device set is determined based on the service access attributes corresponding to the client device and the target application;
[0171] Based on the device identifier of the client device corresponding to the service request code, obtain the compression dictionary corresponding to the service request code, and use the compression dictionary to decompress the service request code to obtain the service request.
[0172] Obtain the service response data corresponding to the service request, and compress the service response data using the compression dictionary to obtain the encoded data of the service response data.
[0173] Computer program code for performing the operations of this disclosure can be written in one or more programming languages or a combination thereof, including but not limited to object-oriented programming languages such as Java, Smalltalk, and C++, as well as conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0174] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0175] The units described in the embodiments of this disclosure can be implemented in software or hardware. The names of the units are not, in some cases, intended to limit the specific unit.
[0176] The functions described above in this document can be performed, at least in part, by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: Field Programmable Gate Arrays (FPGAs), Application-Specific Integrated Circuits (ASICs), Application Standard Products (ASSPs), System-on-Chip (SoCs), Complex Programmable Logic Devices (CPLDs), and so on.
[0177] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0178] The above description is merely a preferred embodiment of this disclosure and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of this disclosure is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features disclosed in this disclosure that have similar functions.
[0179] Furthermore, while the operations are described in a specific order, this should not be construed as requiring these operations to be performed in the specific order shown or in a sequential order. In certain environments, multitasking and parallel processing may be advantageous. Similarly, while several specific implementation details are included in the above discussion, these should not be construed as limiting the scope of this disclosure. Certain features described in the context of individual embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented individually or in any suitable sub-combination in multiple embodiments.
[0180] Although the subject matter has been described using language specific to structural features and / or methodological logic, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or actions described above. Rather, the specific features and actions described above are merely illustrative examples of implementing the claims.
Claims
1. A dictionary-based data compression method, characterized in that, include: Obtain the service request code corresponding to the target application, wherein the service request code represents the compression result obtained by compressing the service request of the target application using a compression dictionary, the compression dictionary is determined based on the service request and service response data associated with the device set to which the client device belongs, and the device set is determined based on the service access attributes corresponding to the client device and the target application; Based on the device identifier of the client device corresponding to the service request code, obtain the compression dictionary corresponding to the service request code, and use the compression dictionary to decompress the service request code to obtain the service request. Obtain the service response data corresponding to the service request, and compress the service response data using the compression dictionary to obtain the encoded data of the service response data.
2. The method according to claim 1, characterized in that, The training methods for the compressed dictionary include: Based on the service access attributes corresponding to the client devices and the target application, the client devices associated with the target application are grouped to obtain at least one set of devices. Based on the service request and service response data of the application service associated with any device set, a service data subset of the application service is generated, and the application service dataset is constructed based on the service data subset. For any set of devices, the compressed dictionary is trained based on the application service dataset and dictionary configuration parameters associated with the set of devices.
3. The method according to claim 2, characterized in that, For any set of devices, training the compressed dictionary based on the application service dataset and dictionary configuration parameters associated with the set of devices includes: For any set of devices, if the dictionary configuration parameters are updated, the compressed dictionary is updated according to the updated dictionary configuration parameters, and the service data subset associated with the set of devices is compressed using the updated compressed dictionary to obtain the service data subset encoding. The individual compression ratio is determined based on the service data subset encoding and the service data subset, and the global compression ratio is determined based on the weight of the service data subset and the individual compression ratio, wherein the individual compression ratio represents the compression ratio associated with the service data subset, and the global compression ratio represents the compression ratio associated with the application service dataset; If the global compression ratio does not meet the preset training termination condition, the dictionary configuration parameters are updated, wherein the preset training termination condition includes the global compression ratio being higher than the global compression ratio associated with the previous dictionary update cycle.
4. The method according to claim 3, characterized in that, The step of determining the global compression ratio based on the weights of the service data subsets and the individual compression ratios includes: The weight of the service data subset is determined based on the proportion of the data volume corresponding to the service data subset in the data volume corresponding to the application service dataset. The global compression rate is obtained by weighted fusion processing based on the weight of the service data subset and the individual compression rate.
5. The method according to claim 4, characterized in that, After determining the weights of the service data subset, the process further includes: The weights are updated based on the priority of the application services corresponding to the subset of service data.
6. The method according to claim 3, characterized in that, The dictionary configuration parameters include key length and search window length; The updating of the dictionary configuration parameters includes: Update the key value length and / or search window length according to the set step size.
7. The method according to claim 3, characterized in that, Also includes: If the global compression rate meets the preset training termination condition, the service data subset is sorted in descending order according to its weight, and the target service data subset is determined based on the ranking. If the individual compression rate of the target service data subset is higher than the individual compression rate corresponding to the previous dictionary update cycle, then the updated compressed dictionary is sent to the client devices in the device set.
8. The method according to claim 7, characterized in that, After sending the updated compressed dictionary to the client devices in the device set, the process also includes: Based on the business and performance indicators of the client device before and after the compression dictionary is distributed, determine the indicator deviation value associated with the updated compression dictionary; If the deviation value of the indicator is less than the preset indicator change threshold, then the compressed dictionary corresponding to the previous dictionary update cycle of the current dictionary update cycle is obtained, and the compressed dictionary corresponding to the previous dictionary update cycle is sent to the client device.
9. A dictionary-based data compression device, characterized in that, include: The request acquisition module is used to acquire the service request code corresponding to the target application. The service request code represents the compression result obtained by compressing the service request of the target application using a compression dictionary. The compression dictionary is determined based on the service request and service response data associated with the device set to which the client device belongs. The device set is determined based on the service access attributes of the client device and the target application. The decompression module is used to obtain the compression dictionary corresponding to the service request code based on the device identifier of the client device corresponding to the service request code, and to decompress the service request code using the compression dictionary to obtain the service request. The compression module is used to obtain the service response data corresponding to the service request, and to compress the service response data using the compression dictionary to obtain the encoded data of the service response data.
10. An electronic device, characterized in that, include: One or more processors; Storage device for storing one or more programs. When the one or more programs are executed by the one or more processors, the one or more processors implement the dictionary-based data compression method as described in any one of claims 1-8.
11. A storage medium containing computer-executable instructions, characterized in that, The computer-executable instructions, when executed by a computer processor, are used to perform the dictionary-based data compression method as described in any one of claims 1-8.