Data processing method and device, electronic equipment and storage medium
Patent Information
- Application Number
- CN202510350584.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-21
- Publication Date
- 2026-09-22
AI Technical Summary
[0003]进行边云数据传输或者在设备端内存储这些数据,若不考虑数据本身的数据特点进行压缩等数据处理,则很难有较好的压缩效果
[0050]在上述方案中,考虑到字段的数据在多条数据流中如果变化程度较低,则说明很多条数据流中该字段的数据都相同,若变化程度较高,则说明很多条数据流中该字段的数据基本都不相同,通过对各字段的数据在多条数据流中的变化程度构建包含变化程度不同的至少两个原始数据组,然后对至少一组原始数据组进行压缩处理,并根据压缩处理后的压缩数据组得到目标批次数据,相对于直接传输多条数据流,本方案考虑了待传输的数据流自身的特点进行分组压缩,能够提高压缩效果,进而提高后续数据传输的效率。
Smart Images

Figure CN122801958A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the energy field, and in particular to a data processing method, apparatus, electronic device, and storage medium. Background Technology
[0002] In some fields, a lot of duplicate data may be sent during data transmission. For example, energy data generally includes telemetry and teleindication, which occupy a large amount of data space. The frequency of data collection and reporting by equipment in the energy field is basically 1 second, 5 seconds, or 10 seconds. The higher the frequency, the greater the probability of data duplication. For example, the data such as SOC, SOH, voltage, current, and charging and discharging power reported by the BMS of energy storage batteries are all the same in a very short time.
[0003] When transmitting data between the edge and cloud or storing this data on the device, it is difficult to achieve good compression results without considering the characteristics of the data itself and performing data processing such as compression. Summary of the Invention
[0004] This application provides at least one data processing method, apparatus, electronic device, and storage medium.
[0005] This application provides a data processing method, including acquiring multiple data streams to be transmitted, wherein each data stream includes data corresponding to multiple fields; constructing at least two sets of original data groups based on the degree of change of each field's data in each data stream, wherein each set of original data groups includes data corresponding to at least one field in the multiple data streams, the fields contained in different sets of original data groups are different, and the degree of change of the fields contained in each set of original data groups in the multiple data streams is different; compressing at least one set of original data groups to obtain corresponding target compressed data groups; when all original data groups have been compressed, combining the target compressed data groups to obtain target batch data corresponding to the multiple data streams; when some original data groups have not been compressed, combining the target compressed data groups with other uncompressed original data groups to obtain target batch data corresponding to the multiple data streams.
[0006] In the above scheme, considering that if the data of a field changes little across multiple data streams, it means that the data of that field is the same across many data streams; if the data changes much, it means that the data of that field is basically different across many data streams, the scheme constructs at least two original data groups with different degrees of change based on the degree of change of each field's data across multiple data streams. Then, at least one original data group is compressed, and the target batch data is obtained based on the compressed data group. Compared to directly transmitting multiple data streams, this scheme considers the characteristics of the data streams to be transmitted and performs group compression, which can improve the compression effect and thus improve the efficiency of subsequent data transmission.
[0007] In some embodiments, the method further includes: for at least a portion of the original data set, determining attribute information of the original data set based on the degree of change of the data of each field in the original data set across multiple data streams; determining the correlation between the data of each field in the original data set across multiple data streams based on the attribute information of the original data set; and determining the compression method of the original data set based on the correlation between the data of each field in the original data set across multiple data streams.
[0008] In the above scheme, the data belonging to fields of different attribute information vary to different degrees in multiple data streams, and the correlation between each data in multiple data streams may also be different. In the process of determining the compression method of the obtained original data based on the correlation between each data in multiple data streams, the characteristics of each field are fully considered, thereby improving the subsequent compression effect.
[0009] In some embodiments, the types of the original data group include a first original data group and a second original data group, wherein the fields contained in the first original data group have unchanged corresponding data in multiple data streams, while the fields contained in the second original data group have changed corresponding data in multiple data streams.
[0010] In the above scheme, the data of each field in the first raw data group is the same in each data stream. By dividing the data of these fields into the same raw data group, the data can be compressed in a targeted manner, which can improve the efficiency of data compression.
[0011] In some embodiments, when it is determined that at least two sets of original data sets include the first original data set, the first original data set is compressed, including: obtaining data from any data stream corresponding to the fields contained in the first original data set in the first original data set, as a shared data set; and obtaining a target compressed data set based on the shared data set.
[0012] In the above scheme, for each field in the first original data group, the data of these fields are the same in each data stream. Therefore, the data of this field in the first original data group can be compressed into one line, so that the target compressed data group obtained by compressing the first original data group has a much larger amount of data compressed compared to the first original data group.
[0013] In some embodiments, obtaining a target compressed data group based on a shared data group includes: taking each field in the shared data group as a target field, compressing the data of the target field in the shared data group respectively to obtain the compressed data corresponding to the target field; and combining the compressed data corresponding to all target fields to obtain the target compressed data group.
[0014] In the above scheme, by compressing each field in the shared data group, the compression effect can be improved compared to compressing only some fields.
[0015] In some embodiments, the method further includes: dividing each field in the first data group template to obtain a first part and a second part corresponding to each field in the first data group template; for each field in the first data group template, determining the number of bits occupied by the second part based on the number of bits required for the compressed data of the target field to be carried by the field, and determining the number of bits occupied by the first part of the field based on the number of bits required for encoding the target number of bits; combining the compressed data corresponding to all target fields to obtain a target compressed data group, including: writing the compressed data of each target field into the corresponding second part in the first data group template, and writing the value of the number of bits occupied by each second part into each first part respectively.
[0016] In the above scheme, by dynamically planning the number of bits occupied by the compressed data of the target field, compared with using the same and longer number of bits to carry the compressed data for each field, this scheme achieves the purpose of dynamically configuring the length of each second part by encoding the value of the number of bits occupied by each second part into the first part, thereby improving the compression effect.
[0017] In some embodiments, each field in the shared data group is taken as a target field, and the data of the target field in the shared data group is compressed to obtain the compressed data corresponding to the target field. This includes: encoding the data of each target field to obtain initial data; splitting the initial data to obtain several initial data segments; adding a flag bit to each initial data segment to obtain an advanced data segment; and combining the advanced data segments to obtain the compressed data of the target field.
[0018] In the above scheme, the encoded initial data is split into multiple segments, and a flag bit is added to each segment to dynamically determine the number of bits occupied by each target field. Compared with using the same and longer number of bits to carry compressed data for each field, this scheme improves the compression effect by encoding the number of bits occupied by each second part into the first part.
[0019] In some embodiments, the method further includes: for each target field, determining a first total number of bits required for each advanced data segment of the target field, and a second total number of bits required for the data of the target field according to a preset encoding method; in response to the first total number of bits being less than the second total number of bits, performing a step of combining each advanced data segment to obtain compressed data of the target field; in response to the first total number of bits being greater than or equal to the second total number of bits, compressing the target field using a preset encoding method to obtain compressed data corresponding to the target field.
[0020] In the above scheme, by selecting the encoding method that requires fewer bits based on the number of bits needed to encode the target field data for each encoding method, the compression effect of the target field can be improved.
[0021] In some embodiments, the method further includes: in response to the fact that at least some target fields in the shared data group use different encoding methods, encoding the indication information of the encoding method used by each target field to obtain a first compression identifier; combining the compressed data of each target field with each indication information respectively, and using the combined data as the final compressed data of each target field.
[0022] In the above scheme, by encoding the indication information of the encoding method of each target field into the final compressed data of each field, the decoding efficiency can be improved in the subsequent decoding process.
[0023] In some embodiments, the method further includes: adding a second compression identifier that represents the compression method used by the shared data group to the target compressed data group to update the target compressed data group.
[0024] In the above scheme, considering that there may be multiple encoding methods for different target fields, by parsing the second compression identifier, it is possible to know whether the encoding methods of each target field are the same, which can facilitate the subsequent parsing of target batch data.
[0025] In some embodiments, when it is determined that at least two sets of original data sets include a second original data set, the second original data set is compressed, including: based on the degree of change of each field in the second original data set in multiple data streams in the second original data set, each field in the second original data set is divided into at least two sub-data sets, wherein each sub-data set includes data of at least one field in the corresponding data stream; and at least one set of sub-data sets is compressed to obtain a target compressed data set.
[0026] In the above scheme, the fields with different degrees of change in the second original data group can be further segmented to obtain at least two sub-data groups. Then, the data of the fields with different degrees of change are compressed in a targeted manner. Compared with compressing fields with different degrees of change in the same way, this scheme further considers the characteristics of data with different degrees of change for compression, which can improve the compression effect.
[0027] In some embodiments, before dividing each field of the second original data group into at least two sub-data groups based on the degree of change of the data of each field in the second original data group among multiple data streams in the second original data group, the method further includes: if it is determined that there are data streams belonging to the same event in the second original data group, compressing the data streams belonging to the same event to obtain a new second original data group.
[0028] In the above scheme, if there are data streams belonging to the same event in the second original data group, continuing to encode these data streams belonging to the same event may lead to duplicate encoding. Therefore, this scheme chooses to compress the data streams belonging to the same event and then use a new second original data group for segmentation, which can improve the subsequent compression effect.
[0029] In some embodiments, based on the degree of change of the data of each field in the second original data group in multiple data streams in the second original data group, each field in the second original data group is divided into at least two sub-data groups, including: for each data stream in the second original data group, based on the degree of change of the data of each field in the second original data group in multiple data streams in the second original data group, each field in each data stream in the second original data group is divided into at least two sub-data groups.
[0030] In the above scheme, each data stream in the second original data group is grouped, so that each data stream can be compressed in a targeted manner, thereby improving the compression effect.
[0031] In some embodiments, compressing at least one set of sub-data groups to obtain a target compressed data group includes: taking a sub-data group whose data of each field in the second original data changes to a degree that meets a preset condition in multiple data streams of the second original data group as a target sub-data group, and taking each field in the target sub-data group as a target field; compressing the data of the target field in the target sub-data group respectively to obtain compressed data corresponding to each target field; and combining the compressed data corresponding to all target fields to obtain the target compressed data group.
[0032] In the above scheme, by compressing each field in the target sub-data group, the compression effect can be improved compared to compressing only some fields.
[0033] In some embodiments, the method further includes: for each target sub-data group, dividing each field in the second data group template to obtain a first part and a second part corresponding to each field in the second data group template; for each field in the second data group template, determining the number of bits occupied by the second part based on the number of bits required for the compressed data of each target field in the target sub-data group to be carried by the field, and determining the number of bits occupied by the first part of the field based on the number of bits required for encoding the target number of bits; combining the compressed data corresponding to all target fields to obtain a target compressed data group, including: writing the compressed data of each target field into the corresponding second part in the second data group template, and writing the value of the number of bits occupied by each second part into each first part respectively.
[0034] In the above scheme, by dynamically planning the number of bits occupied by the compressed data of the target field, compared with using the same and longer number of bits to carry the compressed data for each field, this scheme achieves the purpose of dynamically configuring the length of each second part by encoding the value of the number of bits occupied by each second part into the first part, thereby improving the compression effect.
[0035] In some embodiments, the data of the target field in the target sub-data group are compressed respectively to obtain compressed data corresponding to each target field, including: for the data of each target field in the target sub-data group, the data of the target field is encoded to obtain initial data; the initial data is split to obtain several initial data segments; a flag bit is added to each initial data segment to obtain an advanced data segment, and the advanced data segments are combined to obtain compressed data of the target field.
[0036] In the above scheme, the encoded initial data is split into multiple segments, and a flag bit is added to each segment to dynamically determine the number of bits occupied by each target field. Compared with using the same and longer number of bits to carry compressed data for each field, this scheme improves the compression effect by encoding the number of bits occupied by each second part into the first part.
[0037] In some embodiments, the method further includes: for each target field in the target sub-data group, determining the third total number of bits required for each advanced data segment of the target field, and the fourth total number of bits required for the data of the target field according to a preset encoding method; in response to the third total number of bits being less than the fourth total number of bits, performing a step of combining each advanced data segment to obtain compressed data of the target field; in response to the third total number of bits being greater than or equal to the fourth total number of bits, compressing the target field using a preset encoding method to obtain compressed data corresponding to the target field.
[0038] In the above scheme, by selecting the encoding method that requires fewer bits based on the number of bits needed to encode the target field data for each encoding method, the compression effect of the target field can be improved.
[0039] In some embodiments, the method further includes: in response to the different encoding methods used by at least some target fields in the target sub-data group, encoding the indication information of the encoding method used by each target field to obtain a third compression identifier; combining the compressed data of each target field with each indication information, and using the combined data as the final compressed data of each target field.
[0040] In the above scheme, by encoding the indication information of the encoding method of each target field into the final compressed data of each field, the decoding efficiency can be improved in the subsequent decoding process.
[0041] In some embodiments, the method further includes: adding a fourth compression identifier, which characterizes the compression method used by the target sub-data group, to the target compressed data group to update the target compressed data group.
[0042] In the above scheme, considering that there may be multiple encoding methods for different target fields, by parsing the second compression identifier, it is possible to know whether the encoding methods of each target field are the same, which can facilitate the subsequent parsing of target batch data.
[0043] In some embodiments, the compression processing of other sub-data groups in the second original data group besides the target sub-data group includes: taking other sub-data groups besides the target sub-data group in each data stream of the second original data as data groups to be compressed; for each field in the data group to be compressed, in response to the data group to be compressed being a first data group to be compressed, the first data group to be compressed being the first data group to be compressed in the second original data group or the data of the field in the first data group to be compressed being different from the data of the field in the previous data group to be compressed; selecting a data as the reference data of the field, combining the reference data and the statistical value of the second data group to be compressed to obtain the compressed data of the field in the first compressed data group, wherein the second data group to be compressed is a data group to be compressed where the data of each field is the same as the reference data.
[0044] In the above scheme, by recording the mathematical statistical values of consecutive second data groups to be compressed in the first data group to be compressed, the same data in subsequent second data groups to be compressed does not need to be encoded again, which further improves the compression effect.
[0045] In some embodiments, before compressing at least one set of raw data groups to obtain a corresponding target compressed data group, the method further includes: receiving compression rules configured for multiple data streams sent by a management device, wherein the configured compression rules are determined based on the type of device that generates the multiple data streams and / or the data change characteristics of each field of the multiple data streams; compressing at least one set of raw data groups to obtain a corresponding target compressed data group includes: selecting at least one set of raw data groups for compression processing according to the configured compression rules.
[0046] In the above scheme, the management device first determines the compression rules based on the type of the device and / or the data change characteristics of each field of multiple data streams. Compared with using the same encoding method for all types of energy data, the encoding method provided by this scheme is more flexible. In addition, since the encoding rules are issued by the management device, the edge does not need to encode the compression method of this part of the energy data into the target batch data when encoding, which further improves the compression effect.
[0047] This application provides a data processing apparatus, comprising: a data acquisition module, a data group construction module, a compression module, and an assembly module; the data acquisition module is used to acquire multiple data streams to be transmitted, wherein each data stream includes data corresponding to multiple fields; the data group construction module is used to construct at least two sets of original data groups according to the degree of change of each field's data in each data stream, wherein each set of original data groups includes data corresponding to at least one field in the multiple data streams, the fields contained in different original data groups are different, and the degree of change of the fields contained in each set of original data groups in the multiple data streams is different; the compression module is used to compress at least one set of original data groups to obtain corresponding target compressed data groups; the assembly module is used to combine the target compressed data groups when all original data groups have been compressed to obtain target batch data corresponding to multiple data streams; and when some original data groups have not been compressed, to combine the target compressed data groups with other uncompressed original data groups to obtain target batch data corresponding to multiple data streams.
[0048] This application provides an electronic device, including a memory and a processor, wherein the processor is used to execute program instructions stored in the memory to implement the above-described data processing method.
[0049] This application provides a computer-readable storage medium storing program instructions thereon, which, when executed by a processor, implement any of the above-described data processing methods.
[0050] In the above scheme, considering that if the data of a field changes little across multiple data streams, it means that the data of that field is the same across many data streams; if the data changes much, it means that the data of that field is basically different across many data streams, the scheme constructs at least two original data groups with different degrees of change based on the degree of change of each field's data across multiple data streams. Then, at least one original data group is compressed, and the target batch data is obtained based on the compressed data group. Compared to directly transmitting multiple data streams, this scheme considers the characteristics of the data streams to be transmitted and performs group compression, which can improve the compression effect and thus improve the efficiency of subsequent data transmission.
[0051] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this application. Attached Figure Description
[0052] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with this application and, together with the specification, serve to explain the technical solutions of this application.
[0053] Figure 1 This is a flowchart illustrating one embodiment of the data processing method provided in some implementation examples;
[0054] Figure 2 These are schematic diagrams illustrating the number of bits required to encode the first original data group according to a fixed step size + data encoding method provided in some embodiments;
[0055] Figure 3 These are schematic diagrams illustrating the number of bits required to encode the first original data group according to a variable-length encoding method, provided in some embodiments.
[0056] Figure 4 These are schematic diagrams illustrating the number of bits required to encode the first original data group using a hybrid encoding method, provided in some embodiments.
[0057] Figure 5 These are schematic diagrams illustrating the structure of multiple data groups obtained by dividing multiple data streams, as provided in some embodiments.
[0058] Figure 6 These are schematic diagrams of the data processing apparatus provided in some embodiments;
[0059] Figure 7 These are schematic diagrams of the structure of electronic devices provided in some embodiments;
[0060] Figure 8 These are schematic diagrams of the structure of a computer-readable storage medium provided in some embodiments. Detailed Implementation
[0061] The embodiments of this application will now be described in detail with reference to the accompanying drawings.
[0062] In the following description, specific details such as particular subsystem structures, interfaces, and technologies are presented for illustrative purposes rather than for limiting purposes, in order to provide a thorough understanding of this application.
[0063] In this document, the term "and / or" is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, and B existing alone. Additionally, the character " / " generally indicates that the preceding and following related objects have an "or" relationship. Furthermore, "many" in this document means two or more. Moreover, the term "at least one" in this document means any combination of at least two of any one or more of a plurality of objects. For example, including at least one of A, B, and C can mean including any one or more elements selected from the set consisting of A, B, and C.
[0064] Considering that some data is collected at a high frequency, and the higher the frequency, the greater the probability of data duplication, such as the SOC, SOH, voltage, current, and charge / discharge power data reported by the BMS of energy storage batteries, the transmitted data content is the same within a very short period of time. If traditional binary encoding or other encoding methods are used directly, there will still be a lot of repetitive data being encoded repeatedly.
[0065] Therefore, this solution provides a data processing method that takes into account the characteristics of the data itself, refers to the degree of change of each field in multiple data streams, constructs original data groups containing different fields, and performs targeted data processing on different data information, thereby improving the compression effect of the original data stream.
[0066] Please see Figure 1The data processing method provided in this application may include the following steps S11 to S14. Step S11: Acquire multiple data streams to be transmitted. Each data stream includes data corresponding to multiple fields. Step S12: Construct at least two sets of original data groups based on the degree of change of each field's data in each data stream. Each set of original data groups includes data corresponding to at least one field in the multiple data streams. Different original data groups contain different fields, and the degree of change of the fields in each set of original data groups varies across the multiple data streams. Step S13: Compress at least one set of original data groups to obtain corresponding target compressed data groups. If all original data groups are compressed, proceed to step S14; if some original data groups are not compressed, proceed to step S15. Step S14: Combine the target compressed data groups to obtain target batch data corresponding to multiple data streams. Step S15: Combine the target compressed data groups with other uncompressed original data groups to obtain target batch data corresponding to multiple data streams.
[0067] The execution entity of the data processing method provided in this solution can be, but is not limited to, a data acquisition device, or other devices connected to the data acquisition device. For example, the execution entity can be the energy device itself, other devices connected to the energy device, or devices used to store energy data. Energy devices include, but are not limited to, energy storage stacks, energy storage clusters, photovoltaic systems, and charging piles. This embodiment takes data transmission or storage in the energy field as an example. The multiple data streams to be transmitted can include, but are not limited to, data directly collected from the energy device, or data obtained by analyzing and calculating the data collected from the energy device. Multiple fields in each data stream can be data types such as temperature, SOC, SOH, voltage, current, and charging / discharging power reported by the battery in the energy device at the same time, while the data in the fields can be specific values such as temperature, SOC, SOH, voltage, current, and charging / discharging power. For example, assuming a data stream records a "tenant ID" of "200", then "tenant ID" is a field, and "200" is the data. Step S12 can be implemented by selecting fields that have not changed or have changed infrequently across data streams to construct an original data group, and using fields that have changed significantly across data streams to construct another original data group. For example, among multiple data streams, fields that have not changed are selected as the first field, and other fields are selected as the second field. A first original data group can be constructed based on the first field, and a second original data group can be constructed based on the second field. For example, the first field may include device ID, tenant ID, model ID, model version number, and timestamp to the minute. In some application scenarios, if multiple original data streams need to upload energy data to the cloud every second or every ten seconds, the minute timestamps in several data streams are the same, and the first timestamp information can be extracted as the first original data group without repeated compression. Furthermore, the various IDs corresponding to the energy data generally do not change and can be encoded separately, eliminating the need to transmit duplicate data from each data stream. In some application scenarios, the compression method used for the original data group matches the degree of data change corresponding to the original data group. In step S13, the compression processing of different original data groups can be performed by pre-establishing compression processing methods for each original data group, and then compressing at least one set of original data groups according to the compression processing methods for each original data group. For example, the method of compressing the first field to obtain compressed data can include, but is not limited to: 1. Since the data of the first field is the same in each data stream, one can select the data of the first field from each data stream as the compressed data of the first field; 2. The compressed data can be obtained by aggregating or classifying the data of the first field in different data streams.Each data stream may further include a timestamp composed of a first timestamp and a second timestamp. The first timestamp may be a common part of the timestamps in each data stream. For example, if the timestamps down to the minute level are the same in each data stream, then the first timestamp is down to the minute level. In other embodiments, the first timestamp and each second timestamp may not be in each data stream. The first timestamp may be the time when multiple data streams are acquired, while the second timestamp may be the encoding time of one field in each data stream. Steps S14 and S15 perform assembly processing to obtain the target batch data. This can be done by directly concatenating the tail of one data group with the head of another data group, or concatenating the head of one data group with the tail of another data group, or compressing the data groups to be assembled into a compressed package to obtain the target batch data. The target batch data can represent multiple data streams to be transmitted and sent to the management device or represent multiple data streams to be transmitted and stored.
[0068] In the above scheme, considering that if the data of a field changes little across multiple data streams, it means that the data of that field is the same across many data streams; if the data changes much, it means that the data of that field is basically different across many data streams, the scheme constructs at least two original data groups with different degrees of change based on the degree of change of each field's data across multiple data streams. Then, at least one original data group is compressed, and the target batch data is obtained based on the compressed data group. Compared to directly transmitting multiple data streams, this scheme considers the characteristics of the data streams to be transmitted and performs group compression, which can improve the compression effect and thus improve the efficiency of subsequent data transmission.
[0069] In some embodiments, the method further includes: for at least a portion of the original data set, determining attribute information of the original data set based on the degree of change of the data of each field in the original data set across multiple data streams; determining the correlation between the data of each field in the original data set across multiple data streams based on the attribute information of the original data set; and determining the compression method of the original data set based on the correlation between the data of each field in the original data set across multiple data streams.
[0070] For each field, the degree of variation of its data across multiple data streams can be, but is not limited to, one or more of the following: number of variations, differences between data points, and data repetition rate. For example, the degree of variation of a field's data across multiple data streams is the number of variations. If a field's data is 'a' in the first data stream, 'b' in the second data stream to be compressed, and 'a' in the third data stream, then the field has varied twice across these three data streams. The attribute information of the original data group can be the descriptive information of the original array. For example, the attribute information of the original data group composed of fields with large data variation can be whether the data group is regular or irregular, meaning that the attribute information can characterize whether there is a pattern in the data of each field in the multiple data streams. In some application scenarios, the correspondence between different degrees of variation and several attribute information can be pre-set. In some application scenarios, the attribute information of the original data group can also be the type of device to which multiple data streams belong. For example, data collected by different devices have different degrees of variation, so the degree of variation of each field in multiple data streams can determine the type of device to which multiple data streams belong. The correlation between data in each field across multiple data streams can be defined by the patterns observed in these data streams. These patterns include, but are not limited to, periodic, linear, or other patterns. For example, for a particular field, the correlation between its data across multiple data streams could be periodic. Therefore, the compression method for this correlation could be periodic compression, or if the correlation satisfies a linear relationship, the linear relationship of the field's data across different data streams can be encoded, eliminating the need for repeated encoding of the field's data across different data streams. In other words, the attribute information of the data group can be determined based on the degree of variation of each field's data across different data streams. Then, based on the attribute information, the correlation between the data in each field of the original data group across multiple data streams can be determined, and finally, the compression method for the original data group can be determined.
[0071] In the above scheme, the data belonging to fields of different attribute information vary to different degrees in multiple data streams, and the correlation between each data in multiple data streams may also be different. In the process of determining the compression method of the obtained original data based on the correlation between each data in multiple data streams, the characteristics of each field are fully considered, thereby improving the subsequent compression effect.
[0072] In some embodiments, the types of the original data group include a first original data group and a second original data group, wherein the fields contained in the first original data group have unchanged corresponding data in multiple data streams, while the fields contained in the second original data group have changed corresponding data in multiple data streams.
[0073] In other words, data belonging to the same field in the first original data group are completely identical across multiple data streams. Data belonging to the same field in the second original data group are not completely identical across multiple data streams.
[0074] In the above scheme, the data of each field in the first raw data group is the same in each data stream. By dividing the data of these fields into the same raw data group, the data can be compressed in a targeted manner, which can improve the efficiency of data compression.
[0075] In some embodiments, when it is determined that at least two sets of original data sets include a first original data set, the method of compressing the first original data set may include the following steps: obtaining data from any data stream corresponding to the fields contained in the first original data set in the first original data set, as a shared data set; and obtaining a target compressed data set based on the shared data set.
[0076] Because the fields contained in the first original data group are identical across all corresponding data streams, these fields across multiple data streams can be compressed into a single data stream. Based on the shared data group, the target compressed data group can be obtained either by directly using the shared data group as the target compressed data group, or by compressing the shared data group first. For example, the data of one or more fields contained in the shared data group can be compressed to obtain the target compressed data group.
[0077] In the above scheme, for each field in the first original data group, the data of these fields are the same in each data stream. Therefore, the data of this field in the first original data group can be compressed into one line, so that the target compressed data group obtained by compressing the first original data group has a much larger amount of data compressed compared to the first original data group.
[0078] In some embodiments, the above method of obtaining a target compressed data group based on a shared data group may include: taking each field in the shared data group as a target field, compressing the data of the target field in the shared data group respectively to obtain the compressed data corresponding to the target field; and combining the compressed data corresponding to all target fields to obtain the target compressed data group.
[0079] The compressed data corresponding to the target field in the shared data group is compressed separately. This can be done by selecting one compression method from multiple methods. Different target fields can use different compression methods, or they can all use the same compression method. The compression method used in the compressed data corresponding to the target field can be specified by the data receiver, or it can be the encoding method for the data carried in the second part determined according to the following methods. For example, it can be the encoding method for the second part in the fixed step size + data encoding method, or Varint encoding, etc. The compressed data corresponding to all target fields is combined to obtain the target compressed data group. This can be done by directly combining the compressed data, or by further processing the compressed data before combining them.
[0080] In the above scheme, by compressing each field in the shared data group, the compression effect can be improved compared to compressing only some fields.
[0081] In some embodiments, the method further includes: dividing each field in the first data group template to obtain a first part and a second part corresponding to each field in the first data group template; for each field in the first data group template, determining the number of bits occupied by the second part based on the number of bits required for the compressed data of the target field to be carried by the field, and determining the number of bits occupied by the first part of the field based on the number of bits required for encoding the target number of bits; combining the compressed data corresponding to all target fields to obtain a target compressed data group, including: writing the compressed data of each target field into the corresponding second part in the first data group template, and writing the value of the number of bits occupied by each second part into each first part respectively.
[0082] The first data group template can be an empty template or a pre-set template with multiple fields. When the first data group template is empty or a pre-set template, the number of bits occupied by each field in the first data group template can be a preset value. The first data group template can be set by the management device and sent to the execution device of the data processing method in this scheme. Specifically, it can be determined by the management device based on the type of the execution device and / or the degree of change of each field in the historical data sent. Figure 2As shown, each field can have a first part and a second part. The first part can occupy a fixed number of bits to store the value of the target number of bits. Based on the number of bits required for the compressed data of the target field to be carried by this field, the number of bits occupied by the second part can be determined by directly configuring the number of bits occupied by the second part to be the number of bits required for the compressed data. Alternatively, in some embodiments, the number of bits occupied by the second part can be greater than the number of bits required for the compressed data. Based on the number of bits required for encoding the target number of bits, the number of bits occupied by the first part of this field can be determined by directly using the number of bits required for encoding the target number of bits as the number of bits required for the first part of this field. Alternatively, the number of bits required for the first part can be greater than the number of bits required for encoding the target number of bits. In some application scenarios, configuring the number of bits required for the first part of each field can be done by using the number of bits required to encode the value of the reference number of bits (the number of bits required for encoding the target number of bits) as the number of bits required for the first part. Alternatively, it can be added to a pre-set number of bits as the number of bits required for the first part. Please refer to [link to relevant documentation]. Figure 2 Taking tenant ID as an example, the data of tenant ID is 200. Storing the decimal data 200 as needed only requires 8 bits (8 bits can represent a maximum of 256). The number of bits required for the first part only needs to be enough to store the number of bits required for encoding "8". For example, the number of bits required for binary encoding of 8 is 3 bits, so the number of bits required for the first part can be greater than or equal to 3 bits.
[0083] For example, please refer to Figure 2 If the target field's compressed data requires 4 bits, then the value carried in the first part of the first data group template is "4". Of course, it can specifically carry the encoded value of "4", such as the binary encoded value of "4". The target bit count required for the second part carried in the first part can be used to indicate that after parsing the first part, 4 bits of data need to be read and decoded to obtain the data carried in the second part. The encoding methods for each field include, but are not limited to, binary encoding, variable-length integer encoding (Varint encoding), etc. For example, the structure of the target compressed data group corresponding to the first original data group can be referenced... Figure 2 As shown, Figure 2The target fields include tenant ID, timestamp to minute, model ID, model version number, and device ID. The tenant ID is 200, the timestamp to minute is 133024306, the model ID is 2001, the model version number is 5, and the device ID is 12003. Taking binary encoding as an example, encoding 200 requires 8 bits, so the target bit count for 200 can be 8. Encoding 133024306 requires 27 bits, so the target bit count for 133024306 is 27. The data for other fields follows the same logic and will not be elaborated here. Of course, the target bit count obtained from the number of target fields can also be the sum of the number of bits required for encoding and the preset bit count. That is, different fields can be separated by a preset bit count. For example, the target bit count for tenant ID 200 can be 9 or any number greater than 8.
[0084] In the above scheme, by dynamically planning the number of bits occupied by the compressed data of the target field, compared with using the same and longer number of bits to carry the compressed data for each field, this scheme achieves the purpose of dynamically configuring the length of each second part by encoding the value of the number of bits occupied by each second part into the first part, thereby improving the compression effect.
[0085] In some application scenarios, the method to determine the number of bits occupied by the first part of the field based on the number of bits required for encoding the target number of bits may include: encoding the maximum number of bits required for each compressed data as the number of bits required for each first part.
[0086] In other embodiments, the proportion of bits required for the first part can be determined based on the energy device. For example, the maximum number of bits required for encoding each data in the first raw data group constructed from several raw data streams that need to be encoded at different times by energy device B can be determined. For example, if the maximum number of bits is 32 bits, then encoding 32 bits requires 4 bits. That is, the number of bits required for each first part is the same, and it is determined by the maximum value among the reference bit numbers. In other application scenarios, the value of the number of bits required for encoding using fixed-length binary encoding can be used as the proportion of bits required for the first part. For example, currently, fixed 32-bit or 64-bit binary encoding can be used, so the proportion of bits required for the first part can be the number of bits required for encoding 32 or 64 bits.
[0087] For example, taking a tenant ID of 200 as an example, the first part accounts for 4 bits. The value of the 4 bits can be regarded as the step size and used to store the decimal number 8. When reading data, first read 4 bits to get the decimal number 8, and then read 8 more bits to get the decimal number 200, which is the tenant ID data. Optionally, the target number of bits can be stored using a fixed-length step size (bits).
[0088] In the above scheme, the maximum number of bits required for each compressed data is encoded as the number of bits required for each first part, so that the number of bits required for each first part is relatively fixed, which facilitates the parsing of the target batch of data.
[0089] In some embodiments, each field in the shared data group is taken as a target field, and the data of the target field in the shared data group is compressed to obtain the compressed data corresponding to the target field. This includes: encoding the data of each target field to obtain initial data; splitting the initial data to obtain several initial data segments; adding a flag bit to each initial data segment to obtain an advanced data segment; and combining the advanced data segments to obtain the compressed data of the target field.
[0090] For example, the data of each target field can be binary encoded. For instance, the initial data of the number 300 after binary encoding is 100101100. This 100101100 is split into two segments: the first segment consists of the last seven low-order bits (0101100), and the second segment consists of the remaining high-order bits (10). Adding a flag bit to each initial data segment yields the advanced data segment of the first segment (100101100) and the advanced data segment of the second segment (00000010). That is, the number of bits occupied by the advanced data segments obtained by adding flag bits to different initial data segments can be the same. If the initial data segment has fewer than seven bits, it can be padded with a flag bit of 0. In other words, when decoding first the low-order bits and then the high-order bits, the initial data segment of the low-order bits can be a non-end group, while the initial data segment containing the highest-order bits is the end group. The non-end group and the end group can use different flag bits to obtain the corresponding advanced data segments. For example, the flag bit added to the non-end group can be 1, while the flag bit added to the end group can be 0. Combining the various advanced data segments to obtain compressed data for the target field can be achieved by combining the advanced data segments in the order they were split. This encoding method, compared to fixed-length binary encoding, can be called variable-length encoding.
[0091] In the above scheme, the encoded initial data is split into multiple segments, and a flag bit is added to each segment to dynamically determine the number of bits occupied by each target field. Compared with using the same and longer number of bits to carry compressed data for each field, this scheme improves the compression effect by encoding the number of bits occupied by each second part into the first part.
[0092] In some embodiments, the method further includes: for each target field, determining a first total number of bits required for each advanced data segment of the target field, and a second total number of bits required for the data of the target field according to a preset encoding method; in response to the first total number of bits being less than the second total number of bits, performing a step of combining each advanced data segment to obtain compressed data of the target field; in response to the first total number of bits being greater than or equal to the second total number of bits, compressing the target field using a preset encoding method to obtain compressed data corresponding to the target field.
[0093] The sum of the number of bits required for each advanced data segment of each target field is used as the first total number of bits. Please refer to [link / reference]. Figure 3 In the case of tenant ID 200, the first total bit count is 16 bits, and the first total bit count for timestamp to minute is 32 bits. The preset encoding method can be based on the method specified by the data receiver, or it can be determined as described above for the compression method of the first and second parts. The second total bit count can be the sum of the bit counts of the first and second parts, for example... Figure 2 The second total bit length for tenant ID 200 is 12 bits. A method with a smaller bit length is chosen for compression of the target field. This compression method of the first part + second part can be called fixed-step + data encoding. An illustration of this fixed-step + data encoding method can be found in [reference needed]. Figure 2 Taking tenant ID 200 as an example, the number of bits required for the fixed step size + data encoding method is 4 bits + 8 bits = 12 bits. Figure 3 As shown, the variable-length encoding method requires 16 bits for a tenant ID of 200. Following this method, the number of bits required to encode other target fields in each encoding method is determined sequentially. The target encoding method used for different target fields may differ. For example, for the tenant ID, a fixed-step + data compression encoding method can be chosen, while the device ID can be encoded using a variable-length encoding method.
[0094] In the above scheme, by selecting the encoding method that requires fewer bits based on the number of bits needed to encode the target field data for each encoding method, the compression effect of the target field can be improved.
[0095] In some embodiments, the method further includes: in response to the fact that at least some target fields in the shared data group use different encoding methods, encoding the indication information of the encoding method used by each target field to obtain a first compression identifier; combining the compressed data of each target field with each indication information respectively, and using the combined data as the final compressed data of each target field.
[0096] If different fields in a shared data group use different encoding methods, failing to encode the encoding method indication for each field may lead to decoding confusion due to the inability to determine the compressed data structure of each field. The encoding method indication information for each field can be encoded to obtain a first compression identifier. For each field, the decoding priority of the first compression identifier is the highest in the final compressed data group for that field. The number of bits required for encoding the first compression identifier in each field can be determined based on the number of available encoding methods. For example, if there are two available encoding methods, it can be 1 bit, where 0 indicates fixed-step + data encoding and 1 indicates variable-length encoding. If the first compression representation of a field is 0, then the first part needs to be configured in that field to carry the target number of bits. If the field's encoding method is variable-length encoding, then the first part does not need to be configured in that field.
[0097] The structure of the final compressed data for each field in the hybrid encoded data set can be found in [reference]. Figure 4 For example, when the data group uses mixed encoding, taking the tenant ID as an example, the structure of the tenant ID field can be: first compression identifier + first part + second part. Another example is if the device ID uses variable-length encoding, then the device ID encoding structure is: first compression identifier + second part, where the second part stores the compressed data obtained by combining the various advanced data segments according to the order of their splitting. Figure 4 As shown, in hybrid encoding, when the tenant ID is 200 and encoded using the fixed-step + data compression encoding method, a total of 1 bit + 12 bits = 13 bits are required. Conversely, when the device ID is 12003 and encoded using Varint encoding, a total of 1 bit + 16 bits = 17 bits are required. In the above example, the decoding process for the tenant ID field first decodes the compressed encoding method of the tenant ID, then decodes the target number of bits for the tenant ID, then decodes the tenant ID data, then decodes the timestamp-to-minute compressed encoding method, then decodes the target number of bits for the timestamp-to-minute, then decodes the timestamp-to-minute data information, and so on.
[0098] Therefore, it is evident that hybrid encoding requires compressing the encoding method indication information for each field, resulting in an additional number of bits needed for each field to store the first compression identifier. It is understandable that if all fields in the data group are encoded using the same encoding method, the first compression identifier does not need to be added to the compressed data of each field. Therefore, in this scheme, the number of bits required for the target compressed data group obtained by using the same encoding method for all fields in the data group can be compared with the number of bits required for the target compressed data obtained by using different compression methods for each field in the data group. This comparison determines whether a unified encoding method or a hybrid encoding method should be used for encoding the fields in the data group.
[0099] In the above scheme, by encoding the indication information of the encoding method of each target field into the final compressed data of each field, the decoding efficiency can be improved in the subsequent decoding process.
[0100] In some embodiments, the method further includes: adding a second compression identifier that represents the compression method used by the shared data group to the target compressed data group to update the target compressed data group.
[0101] The second compression identifier can be added to the end of the target compressed data group that is decoded first. The number of bits required for the second compression identifier can be determined based on the number of compression encoding methods. For example, if all data is encoded using fixed-step + data, variable-length encoding, or mixed encoding, then the second compression identifier requires 2 bits. For instance, a second compression identifier of 00 indicates that the data group is encoded using the fixed-step + data compression encoding method, a second compression identifier of 01 indicates that the data group is encoded using the variable-length encoding method, or a second compression identifier of 11 indicates that the data group is encoded using mixed encoding. As mentioned above, if the data group's compression encoding method is mixed encoding, then the first compression identifier needs to be configured in each field; if the data group's compression encoding method is not mixed encoding, then the first compression identifier does not need to be configured in each field.
[0102] In the above scheme, considering that there may be multiple encoding methods for different target fields, by parsing the second compression identifier, it is possible to know whether the encoding methods of each target field are the same, which can facilitate the subsequent parsing of target batch data.
[0103] In some embodiments, when it is determined that at least two sets of original data sets include a second original data set, the second original data set is compressed, including: dividing each field in the second original data set into at least two sub-data sets based on the degree of variation of the data in each field of the second original data set across multiple data streams in the second original data set. Each sub-data set includes data for at least one field in its corresponding data stream; and compressing at least one set of sub-data sets to obtain a target compressed data set.
[0104] Understandably, the degree of variation in different fields within the second raw data may differ. Based on the degree of variation of each field's data across different data streams, the second raw data group can be further segmented, ensuring that each field in each data stream within the second raw data group is divided into at least two sub-data groups. The degree of variation of the fields in different sub-data groups varies across multiple data streams; this variation can be expressed as the number of repetitions. For example, the data in one sub-data group may have more repetitions across multiple data streams, while the data in another sub-data group may have less repetition. Dividing each field in the second raw data group into at least two sub-data groups ensures that each data stream within the second raw data group contains at least two data groups, and that sub-data groups containing fields with the same degree of variation across different data streams contain the same fields. For example, the electricity data in each data stream may be in one sub-data group, while the temperature data in each data stream may be in another sub-data group. In the compression process of at least one set of sub-data groups, compression can be performed on at least a portion of the sub-data groups of at least some data streams in the second original data group, or on all sub-data groups of the entire data stream. If all sub-data groups of the entire data stream are compressed, the compressed data groups of each sub-data group are combined to obtain the target compressed data group; or if some sub-data groups are not compressed, the compressed data corresponding to each compressed data group and the uncompressed data groups are combined to obtain the target compressed data group.
[0105] In the above scheme, the fields with different degrees of change in the second original data group can be further segmented to obtain at least two sub-data groups. Then, the data of the fields with different degrees of change are compressed in a targeted manner. Compared with compressing fields with different degrees of change in the same way, this scheme further considers the characteristics of data with different degrees of change for compression, which can improve the compression effect.
[0106] In some embodiments, before dividing each field of the second original data group into at least two sub-data groups based on the degree of change of the data of each field in the second original data group among multiple data streams in the second original data group, the method further includes: if it is determined that there are data streams belonging to the same event in the second original data group, compressing the data streams belonging to the same event to obtain a new second original data group.
[0107] If every field in two data streams has the same data value, then these two data streams are considered to belong to the same event. Multiple data streams belonging to the same event can be compressed to reduce the number of data streams belonging to the same event; for example, only one of the multiple data streams belonging to the same event can be retained. By filtering multiple data streams belonging to the same event, the number of data streams contained in the second original data group can be reduced, further improving the compression effect.
[0108] In the above scheme, if there are data streams belonging to the same event in the second original data group, continuing to encode these data streams belonging to the same event may lead to duplicate encoding. Therefore, this scheme chooses to compress the data streams belonging to the same event and then use a new second original data group for segmentation, which can improve the subsequent compression effect.
[0109] In some embodiments, based on the degree of change of the data of each field in the second original data group in multiple data streams in the second original data group, each field in the second original data group is divided into at least two sub-data groups, including: for each data stream in the second original data group, based on the degree of change of the data of each field in the second original data group in multiple data streams in the second original data group, each field in each data stream in the second original data group is divided into at least two sub-data groups.
[0110] For example, for each data stream, one data group might contain data with fields exhibiting high variability, while another data group might contain data with fields exhibiting low variability. Exemplarily, a variability threshold can be set, and the fields can be divided according to this threshold. Optionally, the variability threshold can be determined by the management device (e.g., the cloud) and sent to the execution device of this solution. In some application scenarios, the execution device can divide the sub-data groups according to compression rules sent by the management device. These compression rules are determined by the management device based on the variability of each field in multiple historical data streams sent by the execution device across different historical data streams. Considering that the variability of the same field in historical data streams is similar to that in the multiple data streams to be transmitted in this batch, the management device can directly set the compression rules, and the execution device can parse and group the multiple data streams according to these rules.
[0111] In the above scheme, each data stream in the second original data group is grouped, so that each data stream can be compressed in a targeted manner, thereby improving the compression effect.
[0112] In some embodiments, compressing at least one set of sub-data groups to obtain a target compressed data group includes: taking a sub-data group whose data of each field in the second original data changes to a degree that meets a preset condition in multiple data streams of the second original data group as a target sub-data group, and taking each field in the target sub-data group as a target field; compressing the data of the target field in the target sub-data group respectively to obtain compressed data corresponding to each target field; and combining the compressed data corresponding to all target fields to obtain the target compressed data group.
[0113] The preset condition can be that the degree of change is greater than or equal to a preset degree of change. The main purpose of selecting sub-data groups from multiple data streams within the second raw data group whose data for each field meets the preset condition as target sub-data groups is to identify sub-data groups with higher degrees of change. Taking energy data as an example, fields with higher degrees of change in the target sub-data group might include timestamps (to the second level), timestamps (to the millisecond level), event IDs, cumulative power generation, current cumulative power generation time, etc., while fields with lower degrees of change might include voltage, current, temperature, etc. Each field in the target sub-data group is then used as a target field and compressed to obtain compressed data corresponding to each target field.
[0114] In the above scheme, by compressing each field in the target sub-data group, the compression effect can be improved compared to compressing only some fields.
[0115] In some embodiments, the method further includes: for each target sub-data group, dividing each field in the second data group template to obtain a first part and a second part corresponding to each field in the second data group template. For each field in the second data group template, determining the number of bits occupied by the second part based on the number of bits required for the compressed data of each target field in the target sub-data group to be carried by the field, and determining the number of bits occupied by the first part of the field based on the number of bits required for encoding the target bit number. The step of combining the compressed data corresponding to all target fields to obtain the target compressed data group may include: writing the compressed data of each target field into the corresponding second part in the second data group template, and writing the value of the number of bits occupied by each second part into each first part respectively.
[0116] The second data group template can be an empty template or a pre-set template with multiple fields. When the second data group template is empty or a pre-set template, the number of bits occupied by each field in the second data group template can be a preset value. The second data group template can be set by the management device and sent to the execution device of the data processing method in this scheme. Specifically, it can be determined by the management device based on the type of the execution device and / or the degree of change of each field in the historical data sent. Figure 2 As shown, each field can have a first part and a second part. The first part can occupy a fixed number of bits to store the value of the target number of bits. Based on the number of bits required for the compressed data of the target field to be carried by this field, the number of bits occupied by the second part can be determined by directly configuring the number of bits occupied by the second part to be the number of bits required for the compressed data. Alternatively, in some embodiments, the number of bits occupied by the second part can be greater than the number of bits required for the compressed data. Based on the number of bits required for encoding the target number of bits, the number of bits occupied by the first part of this field can be determined by directly using the number of bits required for encoding the target number of bits as the number of bits required for the first part of this field. Alternatively, the number of bits required for the first part can be greater than the number of bits required for encoding the target number of bits. In some application scenarios, configuring the number of bits required for the first part of each field can be done by using the number of bits required to encode the value of the reference number of bits (the number of bits required for encoding the target number of bits) as the number of bits required for the first part. Alternatively, it can be added to a pre-set number of bits as the number of bits required for the first part. Please refer to [link to relevant documentation]. Figure 2 Taking a field in the target sub-data group that includes the tenant ID as an example, the tenant ID data is 200. Storing this decimal number 200 only requires 8 bits (8 bits can represent a maximum of 256). The number of bits required for the first part only needs to be enough to store the number of bits required for encoding "8". For example, the number of bits required for binary encoding of "8" is 3 bits, so the number of bits required for the first part can be greater than or equal to 3 bits. It is worth noting that the compression method of each field in the target sub-data group can be the same as the compression method of the shared data group in the first original data group. Therefore, this example only uses the tenant ID to continue the compression of each field in the target sub-data group and is not intended to represent that all fields in the shared data group and the target sub-data group include the tenant ID.
[0117] For example, please refer to Figure 2If the target field's compressed data requires 4 bits, then the value carried in the first part of the second data group template is "4". Of course, it can specifically carry the encoded value of "4", such as the binary encoded value of "4". The target bit count required for the second part carried in the first part can be used to indicate that after parsing the first part, 4 bits of data need to be read and decoded to obtain the data carried in the second part. The encoding methods for each field include, but are not limited to, binary encoding, variable-length integer encoding (Varint encoding), etc. For example, the structure of the target compressed data group corresponding to the target sub-data group can be referenced... Figure 2 As shown, Figure 2 The target fields include tenant ID, timestamp to minute, model ID, model version number, and device ID. The tenant ID is 200, the timestamp to minute is 133024306, the model ID is 2001, the model version number is 5, and the device ID is 12003. Taking binary encoding as an example, encoding 200 requires 8 bits, so the target bit count for 200 can be 8. Encoding 133024306 requires 27 bits, so the target bit count for 133024306 is 27. The data for other fields follows the same logic and will not be elaborated here. Of course, the target bit count obtained from the number of target fields can also be the sum of the number of bits required for encoding and the preset bit count. That is, different fields can be separated by a preset bit count. For example, the target bit count for tenant ID 200 can be 9 or any number greater than 8.
[0118] In the above scheme, by dynamically planning the number of bits occupied by the compressed data of the target field, compared with using the same and longer number of bits to carry the compressed data for each field, this scheme achieves the purpose of dynamically configuring the length of each second part by encoding the value of the number of bits occupied by each second part into the first part, thereby improving the compression effect.
[0119] In some application scenarios, the method to determine the number of bits occupied by the first part of the field based on the number of bits required for encoding the target number of bits may include: encoding the maximum number of bits required for each compressed data as the number of bits required for each first part.
[0120] In other embodiments, the proportion of bits required for the first part can be determined based on the energy device. For example, the maximum number of bits required for encoding each data in the first raw data group constructed from several raw data streams that need to be encoded at different times by energy device B can be determined. For example, if the maximum number of bits is 32 bits, then encoding 32 bits requires 4 bits. That is, the number of bits required for each first part is the same, and it is determined by the maximum value among the reference bit numbers. In other application scenarios, the value of the number of bits required for encoding using fixed-length binary encoding can be used as the proportion of bits required for the first part. For example, currently, fixed 32-bit or 64-bit binary encoding can be used, so the proportion of bits required for the first part can be the number of bits required for encoding 32 or 64 bits.
[0121] For example, taking a tenant ID of 200 as an example, the first part accounts for 4 bits. The value of the 4 bits can be regarded as the step size and used to store the decimal number 8. When reading data, first read 4 bits to get the decimal number 8, and then read 8 more bits to get the decimal number 200, which is the tenant ID data. Optionally, the target number of bits can be stored using a fixed-length step size (bits).
[0122] In the above scheme, the maximum number of bits required for each compressed data is encoded as the number of bits required for each first part, so that the number of bits required for each first part is relatively fixed, which facilitates the parsing of the target batch of data.
[0123] In some embodiments, the data of the target field in the target sub-data group are compressed respectively to obtain compressed data corresponding to each target field, including: for the data of each target field in the target sub-data group, the data of the target field is encoded to obtain initial data; the initial data is split to obtain several initial data segments; a flag bit is added to each initial data segment to obtain an advanced data segment, and the advanced data segments are combined to obtain compressed data of the target field.
[0124] For example, the data of each target field can be binary encoded. For instance, the initial data of the number 300 after binary encoding is 100101100. This 100101100 is split into two segments: the first segment consists of the last seven low-order bits (0101100), and the second segment consists of the remaining high-order bits (10). Adding a flag bit to each initial data segment yields the advanced data segment of the first segment (100101100) and the advanced data segment of the second segment (00000010). That is, the number of bits occupied by the advanced data segments obtained by adding flag bits to different initial data segments can be the same. If the initial data segment has fewer than seven bits, it can be padded with a flag bit of 0. In other words, when decoding first the low-order bits and then the high-order bits, the initial data segment of the low-order bits can be a non-end group, while the initial data segment containing the highest-order bits is the end group. The non-end group and the end group can use different flag bits to obtain the corresponding advanced data segments. For example, the flag bit added to the non-end group can be 1, while the flag bit added to the end group can be 0. Combining the various advanced data segments to obtain compressed data for the target field can be achieved by combining the advanced data segments in the order they were split. This encoding method, compared to fixed-length binary encoding, can be called variable-length encoding.
[0125] In the above scheme, the encoded initial data is split into multiple segments, and a flag bit is added to each segment to dynamically determine the number of bits occupied by each target field. Compared with using the same and longer number of bits to carry compressed data for each field, this scheme improves the compression effect by encoding the number of bits occupied by each second part into the first part.
[0126] In some embodiments, the method further includes: for each target field in the target sub-data group, determining the third total number of bits required for each advanced data segment of the target field, and the fourth total number of bits required for the data of the target field according to a preset encoding method; in response to the third total number of bits being less than the fourth total number of bits, performing a step of combining each advanced data segment to obtain compressed data of the target field; in response to the third total number of bits being greater than or equal to the fourth total number of bits, compressing the target field using a preset encoding method to obtain compressed data corresponding to the target field.
[0127] The sum of the number of bits required for each advanced data segment of each target field is used as the third total number of bits. Continuing with... Figure 3 For example, suppose Figure 3 This refers to the third total number of bits required for each advanced data segment of each field in the target sub-data group, as mentioned above. This is merely an example and is not intended to indicate that the target sub-data group determined from multiple data streams to be transmitted includes these fields. Please refer to [link to relevant documentation]. Figure 3In this context, assuming the target fields in the target sub-data group include tenant ID and timestamp to minute, with tenant ID being 200, the third total bit count is 16 bits, and the third total bit count for timestamp to minute is 32 bits. The preset encoding method can be based on the method specified by the data receiver, or it can be determined according to the compression method of the first and second parts as described above. The fourth total bit count can be the sum of the bit counts of the first and second parts, for example... Figure 2 The fourth bit in the tenant ID 200 is 12 bits. A method with a smaller percentage of bits is chosen for compression of the target field. The compression method of the first part + the second part can be called fixed-step + data encoding. An illustration of this fixed-step + data encoding method can be found in [reference needed]. Figure 2 Taking tenant ID 200 as an example, the number of bits required for the fixed step size + data encoding method is 4 bits + 8 bits = 12 bits. Figure 3 As shown, the variable-length encoding method requires 16 bits for a tenant ID of 200. Following this method, the number of bits required to encode other target fields in each encoding method is determined sequentially. The target encoding method used for different target fields may differ. For example, for the tenant ID, a fixed-step + data compression encoding method can be chosen, while the device ID can be encoded using a variable-length encoding method.
[0128] In the above scheme, by selecting the encoding method that requires fewer bits based on the number of bits needed to encode the target field data for each encoding method, the compression effect of the target field can be improved.
[0129] In some embodiments, the method further includes: in response to the different encoding methods used by at least some target fields in the target sub-data group, encoding the indication information of the encoding method used by each target field to obtain a third compression identifier; combining the compressed data of each target field with each indication information, and using the combined data as the final compressed data of each target field.
[0130] If different fields in the target sub-data group use different encoding methods, failing to encode the encoding method indication for each field may lead to decoding confusion due to the inability to determine the compressed data structure of each field. A third compression identifier can be encoded to represent the encoding method indication for each field. For each field, the decoding priority of the third compression identifier is the highest in the final compressed data group for that field. The number of bits required for encoding the third compression identifier in each field can be determined based on the number of available encoding methods. For example, if there are two available encoding methods, it could be 1 bit, where 0 indicates fixed-step + data encoding and 1 indicates variable-length encoding. If the first compression representation of a field is 0, then the first part needs to be configured in that field to carry the target number of bits. If the field's encoding method is variable-length encoding, then the first part does not need to be configured in that field.
[0131] In the above scheme, by encoding the indication information of the encoding method of each target field into the final compressed data of each field, the decoding efficiency can be improved in the subsequent decoding process.
[0132] In some embodiments, the method further includes: adding a fourth compression identifier, which characterizes the compression method used by the target sub-data group, to the target compressed data group to update the target compressed data group.
[0133] A fourth compression identifier can be added to the end of the target compressed data group that is decoded first. The number of bits required for the fourth compression identifier can be determined based on the number of compression encoding methods. For example, all data can use fixed-step + data, all data can use variable-length encoding, or there can be mixed encoding. Therefore, the fourth compression identifier requires 2 bits. For example, a fourth compression identifier of 00 indicates that the data group is encoded using the fixed-step + data compression encoding method, a fourth compression identifier of 01 indicates that the data group is encoded using the variable-length encoding method, or a fourth compression identifier of 11 indicates that the data group is encoded using mixed encoding. As mentioned above, if the data group is encoded using mixed encoding, a third compression identifier needs to be configured in each field; if the data group is not encoded using mixed encoding, a third compression identifier does not need to be configured in each field.
[0134] In the above scheme, considering that there may be multiple encoding methods for different target fields, by parsing the second compression identifier, it is possible to know whether the encoding methods of each target field are the same, which can facilitate the subsequent parsing of target batch data.
[0135] In some embodiments, the compression processing of other sub-data groups in the second original data group besides the target sub-data group includes: taking other sub-data groups besides the target sub-data group in each data stream of the second original data as data groups to be compressed; for each field in the data group to be compressed, in response to the data group to be compressed being a first data group to be compressed, the first data group to be compressed being the first data group to be compressed in the second original data group or the data of the field in the first data group to be compressed being different from the data of the field in the previous data group to be compressed; selecting a data as the reference data of the field, combining the reference data and the statistical value of the second data group to be compressed to obtain the compressed data of the field in the first compressed data group, wherein the second data group to be compressed is a data group to be compressed where the data of each field is the same as the reference data.
[0136] The first data group to be compressed can be the data group to be compressed from the data stream with the smallest timestamp or the smallest event identifier. Each data stream can correspond to one event identifier. In some application scenarios, the division between the first and second data groups to be compressed is based on a single field. For a specific field, if the data in the data stream for that field is different from the data in the previous data stream, then the data group to be compressed in that data stream can be considered the first data group to be compressed. For example, if the temperature in the third data stream is different from the temperature in the first data stream, but the temperature in the second data stream is the same as the temperature in the first data stream, then for the temperature field, the data groups to be compressed in the first and third data streams are both the first data groups to be compressed, while the second data stream is the second data group to be compressed from the first compressed data group in the first data stream. Furthermore, assuming that the data in the voltage field is the same in the first three data streams, then for the voltage field, the data groups to be compressed in the first data stream are the first data groups to be compressed, and the data groups to be compressed in the second and third data streams are the second data groups to be compressed. In other words, for a single field within a data set to be compressed, there can be multiple first data sets corresponding to that field. The mathematical statistical value can be the number of records in a second data set or the cumulative time between the last data stream and the first data set. Specifically, multiple fields within a data set to be compressed can be grouped together and encoded with only one mathematical statistical value. For example, if temperature, voltage, and current are grouped together, only one mathematical statistical value needs to be encoded in the compressed data set of the first data set. Alternatively, in some application scenarios, each field can be encoded with a separate mathematical statistical value. For the second data set corresponding to a field, the data for that field may not need to be encoded in that second data set.
[0137] In the above scheme, by recording the mathematical statistical values of consecutive second data groups to be compressed in the first data group to be compressed, the same data in subsequent second data groups to be compressed does not need to be encoded again, which further improves the compression effect.
[0138] In some embodiments, before compressing at least one set of raw data groups to obtain a corresponding target compressed data group, the method further includes: receiving compression rules configured for multiple data streams sent by a management device, wherein the configured compression rules are determined based on the type of device that generates the multiple data streams and / or the data change characteristics of each field of the multiple data streams; compressing at least one set of raw data groups to obtain a corresponding target compressed data group includes: selecting at least one set of raw data groups for compression processing according to the configured compression rules.
[0139] For example, the management device can set different compression rules for fields with low and high volatility. These compression rules can specify which fields need to be divided into the first raw data group, which fields need to be divided into the second raw data group, and the rules for further division of fields within the second raw data group. In some application scenarios, when building device attribute management functionality on the management device, taking energy data streams as an example, firstly, the management device must have the ability to add devices, each with a unique device ID; secondly, telemetry and teleindication attributes need to be defined for each device, with at least the English name and data type of the attribute defined. Compression rules are defined for each type of device. These rules can be implemented in the management interface of the management device, allowing grouping of all telemetry and teleindication attributes within the device, for example, grouping and tagging based on the characteristics (e.g., volatility) of the telemetry and teleindication data uploaded by the device. Some telemetry data, such as voltage, current, and temperature, show little change within 1 second, while other data, such as timestamps at the millisecond level and photovoltaic power generation, change continuously over time. This solution analyzes historical data characteristics of telemetry and tele-signaling attributes for each type of device, automatically tags and groups data, and formulates compression rules for different groups. In this scheme, compression rules are defined for each type of device, allowing for highly efficient coding schemes tailored to the characteristics of different device types. The compression rules for the managed devices are distributed to the edge devices, which are the execution devices that need to compress the data. The edge devices parse and store the compression rules locally, also storing and applying them according to device type.
[0140] In the above scheme, the management device first determines the compression rules based on the type of the device and / or the data change characteristics of each field of multiple data streams. Compared with using the same encoding method for all types of energy data, the encoding method provided by this scheme is more flexible. In addition, since the encoding rules are issued by the management device, the edge does not need to encode the compression method of this part of the energy data into the target batch data when encoding, which further improves the compression effect.
[0141] In some embodiments, the management device can be the cloud, and the execution device of this solution can be the edge. The cloud manages the compression rules for energy data. Then, the compression rules are sent to the edge, which can segment the original data stream according to the compression rules sent by the cloud. That is, the fields included in the first original data group obtained by segmentation can be defined in the compression rules sent by the cloud, and the fields included in each sub-data group contained in the second original data group can also be defined in the compression rules sent by the cloud. The first original data group can be considered as the common transmission header of all data in each second original data group. For example, the first and second original data groups obtained by dividing multiple data streams to be transmitted, and the structure of the second original data group, can be referenced... Figure 5 .
[0142] The data processing method provided in this application may include the following steps: Managing equipment categories by cloud platform. Each type of equipment has its own telemetry and teleindication data, which constitute the largest portion of the data transmission space. Due to the distinct characteristics of energy sector equipment, there is a high repetition rate in data reporting within a short period. For example, data such as SOC, SOH, voltage, current, and charge / discharge power reported by the energy storage battery BMS are identical within a very short time (e.g., 1 second). The cloud platform defines the fields to be segmented and compressed based on the reporting frequency and equipment type. The edge device obtains the compression rules from the cloud and parses them locally. The edge device performs segmented compression according to the rules. Segmentation is mainly divided into three parts (e.g., it may include the first original data group, the target sub-data group, and the data group to be compressed). Theoretically, more segments can be configured. Each segment has an independent compression scheme based on its characteristics, and multi-segment combination compression achieves the best compression effect.
[0143] For example, the first raw data group (BathHead) may include tenant ID, minute-level timestamp, model ID, site ID, and device ID. These are essential pieces of information for each energy data group from an energy device. The time in BathHead is extracted from the common timestamps of the currently reported batch of energy data groups (e.g., if 200 data entries are reported every 10 seconds, the common minute timestamps of these 200 data entries can be extracted, but it's not limited to minutes). This further improves compression capabilities. The lowest frequency of data reported by energy devices is every 10 seconds, and the highest frequency can be as low as milliseconds, making the extraction of common timestamps essential.
[0144] For the second raw data group (Bathbody), after BathHead is extracted, the remaining fields all belong to the Bathbody part (e.g., Figure 2(As shown in the diagram). For the Bathbody portion, the telemetry and telecontrol data for each type of device is currently divided into two groups based on the grouping and compression rules configured in the cloud (target sub-data group and data to be compressed group) (as shown in the diagram above; theoretically, it can be divided into multiple groups according to the rules). For cases where data changes frequently within a short period (target sub-data group), Gzip and Snappy compression methods can be used directly. In the data to be compressed group, for cases where data changes are not significant within a short period, the first data entry is kept unchanged. The last element of the first data to be compressed group can be a quantity code, which can greatly compress the data. Of course, the last element can be a quantity, a time, or a repeating rule for these repetitive data. That is, the first data to be compressed group can be the field data plus a quantity code, and subsequent data to be compressed groups do not need to encode the same data.
[0145] Then, based on the above method, determine the compression encoding parameters for the first original data group and the target sub-data group, and then encode according to the compression encoding parameters. For example, first determine whether the first original data group and the target sub-data group use a fixed step size + data (or identification information) compression encoding method, a variable length encoding method, or a hybrid encoding method. Then encode each data information (or identification information) according to the determined encoding method.
[0146] Please see Figure 6 The data processing apparatus 30 provided in this application includes a data acquisition module 31, a data group construction module 32, a compression module 33, and an assembly module 34. The data acquisition module 31 is used to acquire multiple data streams to be transmitted, wherein each data stream includes data corresponding to multiple fields. The data group construction module 32 is used to construct at least two sets of original data groups based on the degree of change of each field's data in each data stream, wherein each set of original data groups includes data corresponding to at least one field in the multiple data streams, the fields contained in different sets of original data groups are different, and the degree of change of the fields contained in each set of original data groups in the multiple data streams is different. The compression module 33 is used to compress at least one set of original data groups to obtain corresponding target compressed data groups. The assembly module 34 is used to combine the target compressed data groups when all original data groups have been compressed to obtain target batch data corresponding to multiple data streams; and to combine the target compressed data groups with other uncompressed original data groups when some original data groups have not been compressed to obtain target batch data corresponding to multiple data streams.
[0147] In the above scheme, considering that if the data of a field changes little across multiple data streams, it means that the data of that field is the same across many data streams; if the data changes much, it means that the data of that field is basically different across many data streams, the scheme constructs at least two original data groups with different degrees of change based on the degree of change of each field's data across multiple data streams. Then, at least one original data group is compressed, and the target batch data is obtained based on the compressed data group. Compared to directly transmitting multiple data streams, this scheme considers the characteristics of the data streams to be transmitted and performs group compression, which can improve the compression effect and thus improve the efficiency of subsequent data transmission.
[0148] In some embodiments, the compression module 33 is further configured to: for at least a portion of the original data group, determine the attribute information of the original data group based on the degree of change of the data of each field in the original data group in multiple data streams; determine the correlation relationship of the data of each field in the original data group in multiple data streams based on the attribute information of the original data group; and determine the compression method of the original data group based on the correlation relationship of the data of each field in the original data group in multiple data streams.
[0149] In the above scheme, the data belonging to fields of different attribute information vary to different degrees in multiple data streams, and the correlation between each data in multiple data streams may also be different. In the process of determining the compression method of the obtained original data based on the correlation between each data in multiple data streams, the characteristics of each field are fully considered, thereby improving the subsequent compression effect.
[0150] In some embodiments, the types of the original data group include a first original data group and a second original data group, wherein the fields contained in the first original data group have unchanged corresponding data in multiple data streams, while the fields contained in the second original data group have changed corresponding data in multiple data streams.
[0151] In the above scheme, the data of each field in the first raw data group is the same in each data stream. By dividing the data of these fields into the same raw data group, the data can be compressed in a targeted manner, which can improve the efficiency of data compression.
[0152] In some embodiments, when it is determined that at least two sets of original data groups include the first original data group, the compression module 33 performs compression processing on the first original data group, including: obtaining data from any data stream corresponding to the fields contained in the first original data group in the first original data group, as a shared data group; and obtaining a target compressed data group based on the shared data group.
[0153] In the above scheme, for each field in the first original data group, the data of these fields are the same in each data stream. Therefore, the data of this field in the first original data group can be compressed into one line, so that the target compressed data group obtained by compressing the first original data group has a much larger amount of data compressed compared to the first original data group.
[0154] In some embodiments, the compression module 33 obtains a target compressed data group based on a shared data group, including: taking each field in the shared data group as a target field, compressing the data of the target field in the shared data group respectively to obtain the compressed data corresponding to the target field; and combining the compressed data corresponding to all target fields to obtain the target compressed data group.
[0155] In the above scheme, by compressing each field in the shared data group, the compression effect can be improved compared to compressing only some fields.
[0156] In some embodiments, the compression module 33 is further configured to: divide each field in the first data group template to obtain a first part and a second part corresponding to each field in the first data group template; for each field in the first data group template, determine the number of bits occupied by the second part based on the number of bits required for the compressed data of the target field to be carried by the field, and determine the number of bits occupied by the first part of the field based on the number of bits required for encoding the target number of bits; combine the compressed data corresponding to all target fields to obtain a target compressed data group, including: writing the compressed data of each target field into the corresponding second part in the first data group template, and writing the value of the number of bits occupied by each second part into each first part respectively.
[0157] In the above scheme, by dynamically planning the number of bits occupied by the compressed data of the target field, compared with using the same and longer number of bits to carry the compressed data for each field, this scheme achieves the purpose of dynamically configuring the length of each second part by encoding the value of the number of bits occupied by each second part into the first part, thereby improving the compression effect.
[0158] In some embodiments, each field in the shared data group is taken as a target field, and the compression module 33 compresses the data of the target field in the shared data group to obtain the compressed data corresponding to the target field. This includes: encoding the data of each target field to obtain initial data; splitting the initial data to obtain several initial data segments; adding a flag bit to each initial data segment to obtain an advanced data segment; and combining the advanced data segments to obtain the compressed data of the target field.
[0159] In the above scheme, the encoded initial data is split into multiple segments, and a flag bit is added to each segment to dynamically determine the number of bits occupied by each target field. Compared with using the same and longer number of bits to carry compressed data for each field, this scheme improves the compression effect by encoding the number of bits occupied by each second part into the first part.
[0160] In some embodiments, the compression module 33 is further configured to: for each target field, determine the first total number of bits required for each advanced data segment of the target field, and the second total number of bits required for the data of the target field according to a preset encoding method; in response to the first total number of bits being less than the second total number of bits, perform the step of combining each advanced data segment to obtain compressed data of the target field; in response to the first total number of bits being greater than or equal to the second total number of bits, compress the target field using a preset encoding method to obtain compressed data corresponding to the target field.
[0161] In the above scheme, by selecting the encoding method that requires fewer bits based on the number of bits needed to encode the target field data for each encoding method, the compression effect of the target field can be improved.
[0162] In some embodiments, the compression module 33 is further configured to: in response to the fact that at least some target fields in the shared data group use different encoding methods, encode the indication information of the encoding method used by each target field to obtain a first compression identifier; combine the compressed data of each target field with each indication information respectively, and use the combined data as the final compressed data of each target field.
[0163] In the above scheme, by encoding the indication information of the encoding method of each target field into the final compressed data of each field, the decoding efficiency can be improved in the subsequent decoding process.
[0164] In some embodiments, the compression module 33 is further configured to: add a second compression identifier representing the compression method used by the shared data group to the target compressed data group, so as to update the target compressed data group.
[0165] In the above scheme, considering that there may be multiple encoding methods for different target fields, by parsing the second compression identifier, it is possible to know whether the encoding methods of each target field are the same, which can facilitate the subsequent parsing of target batch data.
[0166] In some embodiments, when it is determined that at least two sets of original data groups include the second original data group, the compression module 33 performs compression processing on the second original data group, including: based on the degree of change of the data of each field in the second original data group in multiple data streams in the second original data group, dividing each field in the second original data group into at least two sub-data groups, wherein each sub-data group includes the data of at least one field in the corresponding data stream; and performing compression processing on at least one set of sub-data groups to obtain the target compressed data group.
[0167] In the above scheme, the fields with different degrees of change in the second original data group can be further segmented to obtain at least two sub-data groups. Then, the data of the fields with different degrees of change are compressed in a targeted manner. Compared with compressing fields with different degrees of change in the same way, this scheme further considers the characteristics of data with different degrees of change for compression, which can improve the compression effect.
[0168] In some embodiments, before dividing each field in the second original data group into at least two sub-data groups based on the degree of change of the data in each field in the second original data group among multiple data streams in the second original data group, the compression module 33 is further configured to: compress the data streams belonging to the same event in the second original data group when it is determined that there are data streams belonging to the same event in the second original data group to obtain a new second original data group.
[0169] In the above scheme, if there are data streams belonging to the same event in the second original data group, continuing to encode these data streams belonging to the same event may lead to duplicate encoding. Therefore, this scheme chooses to compress the data streams belonging to the same event and then use a new second original data group for segmentation, which can improve the subsequent compression effect.
[0170] In some embodiments, the compression module 33 divides each field in the second original data group into at least two sub-data groups based on the degree of change of the data in each field of the second original data group in multiple data streams of the second original data group, including: for each data stream in the second original data group, dividing each field in each data stream of the second original data group into at least two sub-data groups based on the degree of change of the data in each field of the second original data group in multiple data streams of the second original data group.
[0171] In the above scheme, each data stream in the second original data group is grouped, so that each data stream can be compressed in a targeted manner, thereby improving the compression effect.
[0172] In some embodiments, the compression module 33 compresses at least one set of sub-data groups to obtain a target compressed data group, including: taking the sub-data groups whose data of each field in the second original data changes to a degree that meets a preset condition in multiple data streams in the second original data group as target sub-data groups, and taking each field in the target sub-data groups as target fields; compressing the data of the target field in the target sub-data groups respectively to obtain compressed data corresponding to each target field; and combining the compressed data corresponding to all target fields to obtain the target compressed data group.
[0173] In the above scheme, by compressing each field in the target sub-data group, the compression effect can be improved compared to compressing only some fields.
[0174] In some embodiments, the compression module 33 is further configured to: for each target sub-data group, divide each field in the second data group template to obtain a first part and a second part corresponding to each field in the second data group template; for each field in the second data group template, determine the number of bits occupied by the second part based on the number of bits required for the compressed data of each target field in the target sub-data group to be carried by the field, and determine the number of bits occupied by the first part of the field based on the number of bits required for encoding the target number of bits; combine the compressed data corresponding to all target fields to obtain a target compressed data group, including: writing the compressed data of each target field into the corresponding second part in the second data group template, and writing the value of the number of bits occupied by each second part into each first part respectively.
[0175] In the above scheme, by dynamically planning the number of bits occupied by the compressed data of the target field, compared with using the same and longer number of bits to carry the compressed data for each field, this scheme achieves the purpose of dynamically configuring the length of each second part by encoding the value of the number of bits occupied by each second part into the first part, thereby improving the compression effect.
[0176] In some embodiments, the compression module 33 compresses the data of the target field in the target sub-data group respectively to obtain compressed data corresponding to each target field, including: for the data of each target field in the target sub-data group, encoding the data of the target field to obtain initial data; splitting the initial data to obtain several initial data segments; adding a flag bit to each initial data segment to obtain an advanced data segment, and combining each advanced data segment to obtain compressed data of the target field.
[0177] In the above scheme, the encoded initial data is split into multiple segments, and a flag bit is added to each segment to dynamically determine the number of bits occupied by each target field. Compared with using the same and longer number of bits to carry compressed data for each field, this scheme improves the compression effect by encoding the number of bits occupied by each second part into the first part.
[0178] In some embodiments, the compression module 33 is further configured to: for each target field in the target sub-data group, determine the third total number of bits required for each advanced data segment of the target field, and the fourth total number of bits required for the data of the target field according to a preset encoding method; in response to the third total number of bits being less than the fourth total number of bits, perform the step of combining each advanced data segment to obtain compressed data of the target field; in response to the third total number of bits being greater than or equal to the fourth total number of bits, compress the target field using a preset encoding method to obtain compressed data corresponding to the target field.
[0179] In the above scheme, by selecting the encoding method that requires fewer bits based on the number of bits needed to encode the target field data for each encoding method, the compression effect of the target field can be improved.
[0180] In some embodiments, the compression module 33 is further configured to: in response to the different encoding methods used by at least some target fields in the target sub-data group, encode the indication information of the encoding method used by each target field to obtain a third compression identifier; combine the compressed data of each target field with each indication information respectively, and use the combined data as the final compressed data of each target field.
[0181] In the above scheme, by encoding the indication information of the encoding method of each target field into the final compressed data of each field, the decoding efficiency can be improved in the subsequent decoding process.
[0182] In some embodiments, the compression module 33 is further configured to: add a fourth compression identifier, which represents the compression method used by the target sub-data group, to the target compressed data group to update the target compressed data group.
[0183] In the above scheme, considering that there may be multiple encoding methods for different target fields, by parsing the second compression identifier, it is possible to know whether the encoding methods of each target field are the same, which can facilitate the subsequent parsing of target batch data.
[0184] In some embodiments, the compression module 33 performs compression processing on other sub-data groups in the second original data group besides the target sub-data group, including: taking other sub-data groups in each data stream of the second original data besides the target sub-data group as data groups to be compressed; for each field in the data group to be compressed, in response to the data group to be compressed being the first data group to be compressed, the first data group to be compressed being the first data group to be compressed in the second original data group or the data of the field in the first data group to be compressed being different from the data of the field in the previous data group to be compressed; selecting a data as the reference data of the field, combining the reference data and the statistical value of the second data group to be compressed to obtain the compressed data of the field in the first compressed data group, wherein the second data group to be compressed is the data group to be compressed where the data of each field is the same as the reference data.
[0185] In the above scheme, by recording the mathematical statistical values of consecutive second data groups to be compressed in the first data group to be compressed, the same data in subsequent second data groups to be compressed does not need to be encoded again, which further improves the compression effect.
[0186] In some embodiments, before compressing at least one set of raw data groups to obtain the corresponding target compressed data group, the compression module 33 is further configured to: receive compression rules configured for multiple data streams sent by the management device, wherein the configured compression rules are determined based on the type of device that generates the multiple data streams and / or the data change characteristics of each field of the multiple data streams; and compress at least one set of raw data groups to obtain the corresponding target compressed data group, including: selecting at least one set of raw data groups for compression processing according to the configured compression rules.
[0187] In the above scheme, the management device first determines the compression rules based on the type of the device and / or the data change characteristics of each field of multiple data streams. Compared with using the same encoding method for all types of energy data, the encoding method provided by this scheme is more flexible. In addition, since the encoding rules are issued by the management device, the edge does not need to encode the compression method of this part of the energy data into the target batch data when encoding, which further improves the compression effect.
[0188] Please see Figure 7 The electronic device 40 provided in this application includes a memory 41 and a processor 42. The processor 42 is used to execute program instructions stored in the memory 41 to implement the steps in any of the above data processing method embodiments. In a specific implementation scenario, the electronic device 40 may include, but is not limited to, monitoring equipment, microcomputers, and servers. In addition, the electronic device 40 may also include laptops, tablets, and other carrier devices, which are not limited here.
[0189] Specifically, processor 42 controls itself and memory 41 to implement the steps in any of the above data processing method embodiments. Processor 42 can also be referred to as a CPU (Central Processing Unit). Processor 42 may be an integrated circuit chip with signal processing capabilities. Processor 42 can also be a general-purpose processor, digital signal processor (DSP), application-specific integrated circuit (ASIC), field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. A general-purpose processor can be a microprocessor or any conventional processor. Furthermore, processor 42 can be implemented using integrated circuit chips.
[0190] Please see Figure 8 , Figure 8 This is a schematic diagram of the structure of a computer-readable storage medium provided in some embodiments. The computer-readable storage medium 50 stores program instructions 51 thereon, which, when executed by a processor, implement the steps in any of the above-described data processing method embodiments.
[0191] In the above scheme, considering that if the data of a field changes little across multiple data streams, it means that the data of that field is the same across many data streams; if the data changes much, it means that the data of that field is basically different across many data streams, the scheme constructs at least two original data groups with different degrees of change based on the degree of change of each field's data across multiple data streams. Then, at least one original data group is compressed, and the target batch data is obtained based on the compressed data group. Compared to directly transmitting multiple data streams, this scheme considers the characteristics of the data streams to be transmitted and performs group compression, which can improve the compression effect and thus improve the efficiency of subsequent data transmission.
[0192] In some embodiments, the functions or modules of the apparatus provided in this disclosure can be used to perform the methods described in the above method embodiments. The specific implementation can be referred to the description of the above method embodiments, and for the sake of brevity, it will not be repeated here.
[0193] The description of the various embodiments above tends to emphasize the differences between the various embodiments. The similarities or similarities between them can be referred to, and for the sake of brevity, they will not be repeated here.
[0194] In the several embodiments provided in this application, it should be understood that the disclosed methods and apparatus can be implemented in other ways. For example, the apparatus implementations described above are merely illustrative. For instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, units or components may be combined or integrated into another system, or some features may be ignored or not executed. In another image location, the mutual coupling or direct coupling or communication connection shown or discussed may be through some interfaces, and the indirect coupling or communication connection of devices or units may be electrical, mechanical, or other forms.
[0195] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit. If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) or processor to execute all or part of the steps of the methods in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
Claims
1. A data processing method, characterized in that, include: Acquire multiple data streams to be transmitted, where each data stream includes data corresponding to multiple fields; Based on the degree of change of the data of each field in each data stream, at least two sets of original data groups are constructed, wherein each set of original data groups includes at least one field corresponding to the data in the multiple data streams, the fields included in different sets of original data groups are different, and the degree of change of the data corresponding to the fields included in each set of original data groups is different in the multiple data streams. Compress at least one set of the original data groups to obtain the corresponding target compressed data groups; With all the original data groups compressed, the target compressed data groups are combined to obtain the target batch data corresponding to the multiple data streams. If some of the original data groups are not compressed, each of the target compressed data groups is combined with other uncompressed original data groups to obtain the target batch data corresponding to the multiple data streams.
2. The method according to claim 1, characterized in that, The method further includes: For at least a portion of the original data set, the attribute information of the original data set is determined based on the degree of change of the data of each field in the original data set in the multiple data streams; Based on the attribute information of the original data group, determine the association relationship of the data of each field in the original data group in multiple data streams; Based on the correlation between the data of each field in the original data group and multiple data streams, the compression method of the original data group is determined.
3. The method according to claim 1 or 2, characterized in that, The types of the raw data groups include a first raw data group and a second raw data group. The fields contained in the first raw data group have no change in the corresponding data in the multiple data streams, while the fields contained in the second raw data group have changed in the corresponding data in the multiple data streams.
4. The method according to claim 3, characterized in that, If it is determined that the at least two sets of raw data include the first raw data set, the first raw data set is compressed, including: Obtain the data from any data stream corresponding to the fields contained in the first original data group, and use it as a shared data group; Based on the shared data group, the target compressed data group is obtained.
5. The method according to claim 4, characterized in that, The process of obtaining the target compressed data group based on the shared data group includes: Each field in the shared data group is taken as a target field, and the data of the target field in the shared data group is compressed to obtain the compressed data corresponding to the target field. The compressed data corresponding to all the target fields are combined to obtain the target compressed data group.
6. The method according to claim 5, characterized in that, The method further includes: The fields in the first data group template are divided to obtain the first part and the second part corresponding to each field in the first data group template. For each field in the first data group template, the number of bits occupied by the second part is determined based on the number of bits required for the compressed data of the target field to be carried by the field, and the number of bits occupied by the first part of the field is determined based on the number of bits required for encoding the target number of bits. The step of combining the compressed data corresponding to all the target fields to obtain the target compressed data group includes: The compressed data of each target field is written into the corresponding second part of the first data group template, and the value of the number of bits occupied by each second part is written into each first part.
7. The method according to claim 5, characterized in that, The step of taking each field in the shared data group as a target field and compressing the data of that target field in the shared data group to obtain the compressed data corresponding to the target field includes: For each target field, the data is encoded to obtain initial data; The initial data is split into several initial data segments; A flag bit is added to each of the initial data segments to obtain an advanced data segment, and the advanced data segments are combined to obtain the compressed data of the target field.
8. The method according to claim 7, characterized in that, The method further includes: For each target field, determine the first total number of bits required for each advanced data segment of the target field, and the second total number of bits required for the data of the target field according to a preset encoding method; In response to the first total number of bits being less than the second total number of bits, the step of combining each of the advanced data segments to obtain the compressed data of the target field is performed; In response to the first total number of bits being greater than or equal to the second total number of bits, the target field is compressed using the preset encoding method to obtain compressed data corresponding to the target field.
9. The method according to claim 8, characterized in that, The method further includes: In response to the fact that at least some of the target fields in the shared data group use different encoding methods, the indication information of the encoding method used by each target field is encoded to obtain a first compression identifier; The compressed data of each target field is combined with each indication information, and the combined data is used as the final compressed data of each target field.
10. The method according to claim 9, characterized in that, The method further includes: A second compression identifier, representing the compression method used by the shared data group, is added to the target compressed data group to update the target compressed data group.
11. The method according to any one of claims 1 to 10, characterized in that, If it is determined that the at least two sets of original data groups include a second set of original data groups, the second set of original data groups is compressed, including: Based on the degree of change of the data of each field in the second original data group in multiple data streams in the second original data group, each field in the second original data group is divided into at least two sub-data groups, wherein each sub-data group includes at least one field's data in the corresponding data stream; At least one set of the sub-data groups is compressed to obtain the target compressed data group.
12. The method according to claim 11, characterized in that, Before dividing each field in the second original data group into at least two sub-data groups based on the degree of variation of the data in each field of the second original data group across multiple data streams in the second original data group, the method further includes: If it is determined that there are data streams belonging to the same event in the second original data group, the data streams belonging to the same event are compressed to obtain a new second original data group.
13. The method according to claim 11 or 12, characterized in that, The method of dividing each field in the second original data group into at least two sub-data groups based on the degree of change of the data in multiple data streams in the second original data group includes: For each data stream in the second original data group, based on the degree of change of the data of each field in the second original data group across multiple data streams in the second original data group, each field in each data stream of the second original data group is divided into at least two sub-data groups.
14. The method according to any one of claims 11 to 13, characterized in that, The step of compressing at least one group of the sub-data groups to obtain a target compressed data group includes: The sub-data group whose data of each field in the second original data meets the preset condition in the multiple data streams in the second original data group is taken as the target sub-data group, and each field in the target sub-data group is taken as the target field. The data of the target field in the target sub-data group are compressed respectively to obtain the compressed data corresponding to each target field; The compressed data corresponding to all the target fields are combined to obtain the target compressed data group.
15. The method according to claim 14, characterized in that, The method further includes: For each target sub-data group, the fields in the second data group template are divided to obtain the first part and the second part corresponding to each field in the second data group template. For each field in the second data group template, the number of bits occupied by the second part is determined based on the number of bits required for the compressed data of each target field in the target sub-data group that the field needs to carry, and the number of bits occupied by the first part of the field is determined based on the number of bits required for encoding the target number of bits. The step of combining the compressed data corresponding to all the target fields to obtain the target compressed data group includes: The compressed data of each target field is written into the corresponding second part of the second data group template, and the value of the number of bits occupied by each second part is written into each first part.
16. The method according to claim 14, characterized in that, The step of compressing the data of the target field in the target sub-data group to obtain compressed data corresponding to each target field includes: For the data of each target field in the target sub-data group, the data of the target field is encoded to obtain initial data; The initial data is split into several initial data segments; A flag bit is added to each of the initial data segments to obtain an advanced data segment, and the advanced data segments are combined to obtain the compressed data of the target field.
17. The method according to claim 16, characterized in that, The method further includes: For each target field in the target sub-data group, determine the third total number of bits required for each advanced data segment of the target field, and the fourth total number of bits required for the data of the target field according to a preset encoding method; In response to the fact that the third total number of bits is less than the fourth total number of bits, the step of combining each of the advanced data segments to obtain the compressed data of the target field is performed; In response to the third total number of bits being greater than or equal to the fourth total number of bits, the target field is compressed using the preset encoding method to obtain compressed data corresponding to the target field.
18. The method according to claim 17, characterized in that, The method further includes: In response to the fact that at least some of the target fields in the target sub-data group use different encoding methods, the indication information of the encoding method used by each target field is encoded to obtain a third compression identifier; The compressed data of each target field is combined with each indication information, and the combined data is used as the final compressed data of each target field.
19. The method according to claim 18, characterized in that, The method further includes: A fourth compression identifier, representing the compression method used by the target sub-data group, is added to the target compressed data group to update the target compressed data group.
20. The method according to any one of claims 14 to 19, characterized in that, Compression processing of the sub-data groups other than the target sub-data group in the second original data group includes: Other sub-data groups in each of the data streams in the second original data, except for the target sub-data group, are taken as the data groups to be compressed; For each field in the data group to be compressed, in response to the data group to be compressed being the first data group to be compressed, the first data group to be compressed being the first data group to be compressed in the second original data group, or the data of the field in the first data group to be compressed being different from the data of the field in the previous data group to be compressed; Select one of the data as the baseline data for the field, combine the baseline data and the statistical values of the second data group to be compressed to obtain the compressed data of the field in the first compressed data group, where the second data group to be compressed is a data group in which the data of each field is the same as the baseline data.
21. The method according to any one of claims 1 to 20, characterized in that, Before performing compression processing on at least one set of the original data groups to obtain the corresponding target compressed data group, the method further includes: The receiving management device sends compression rules configured for the multiple data streams, wherein the configured compression rules are determined based on the type of the device that generates the multiple data streams and / or the data change characteristics of each field of the multiple data streams; The step of compressing at least one set of the original data groups to obtain a corresponding target compressed data group includes: At least one set of the original data groups is selected for compression processing according to the configured compression rules.
22. A data processing apparatus, characterized in that, include: The data acquisition module is used to acquire multiple data streams to be transmitted, where each data stream includes data corresponding to multiple fields. A data group construction module is used to construct at least two sets of original data groups based on the degree of change of the data of each field in each of the data streams. Each set of original data groups includes at least one field corresponding to the data in the multiple data streams. Different sets of original data groups contain different fields, and the degree of change of the data corresponding to the fields in each set of original data groups in the multiple data streams is different. A compression module is used to compress at least one set of the original data groups to obtain a corresponding target compressed data group; An assembly module is used to combine the target compressed data groups to obtain target batch data corresponding to the multiple data streams when all the original data groups have been compressed; and to combine the target compressed data groups with other uncompressed original data groups to obtain target batch data corresponding to the multiple data streams when some of the original data groups have not been compressed.
23. An electronic device, characterized in that, It includes a memory and a processor, the processor being configured to execute program instructions stored in the memory to implement the data processing method according to any one of claims 1 to 21.
24. A computer-readable storage medium having program instructions stored thereon, characterized in that, When the program instructions are executed by the processor, they implement the data processing method according to any one of claims 1 to 21.