Storage method, device, electronic device, storage medium and program product for signaling data

By compressing the interaction time in signaling data, and classifying and secondary compression based on user encoding and category encoding, the columnar storage method is finally adopted to solve the problem of high consumption of signaling data storage resources in the prior art, and efficient storage resource utilization is achieved.

CN120029558BActive Publication Date: 2025-06-27WISDOM FOOTPRINT DATA TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510518023.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-24
Publication Date
2025-06-27
Estimated Expiration
2045-04-24

AI Technical Summary

Technical Problem

The prior art is difficult to effectively save storage resources to store large amounts of signaling data, especially when meeting data compliance requirements.

Method used

By compressing the interaction time in the signaling data, classifying and secondary compression of the compressed signaling based on user encoding and category encoding, the columnar storage method is finally used to store the data to be stored.

Benefits of technology

It significantly reduces the storage space usage of signaling data, improves the utilization rate of storage resources, and ensures data recovery.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120029558B_ABST
    Figure CN120029558B_ABST
Patent Text Reader

Abstract

The present invention provides a method, apparatus, electronic device, storage medium and program product for storing signaling data, relating to the field of data storage. After obtaining a number of signaling data within a target area within a specified time, first, the interaction time in each signaling data is compressed to obtain a compressed signaling corresponding to each signaling data, which can avoid occupying a large amount of storage resources for the interaction time. Then, based on the user code and the category code, all the compressed signalings are classified and secondarily compressed to obtain the data to be stored corresponding to each of the multiple category codes corresponding to each user code. Finally, all the data to be stored are stored in columns. After such classification and then secondary compression and combined with columnar storage, storage resources can be further saved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data storage, and in particular, to a method, apparatus, electronic device, storage medium, and program product for storing signaling data. Background Art

[0002] With the rapid development of mobile communication technology and the popularization of smart phones, the signaling interaction frequency between user terminal devices and base stations has increased exponentially. Signaling data, as the key information for recording user locations, network states, and communication behaviors, is an important basic data source for operators to optimize networks, analyze user behaviors, and improve service quality.

[0003] According to statistics, the amount of signaling data generated by a single base station per day can reach the TB (Tera Byte) level, and the data scale that operators need to process daily nationwide has exceeded the EB (Exa Byte) level. Moreover, operators need to retain signaling data for 6 - 12 months to meet data compliance requirements, so the consumption of storage resources is huge. Therefore, how to store a large amount of signaling data while saving storage resources is an urgent problem to be solved. Summary of the Invention

[0004] The purpose of the present invention is to provide a method, apparatus, electronic device, storage medium, and program product for storing signaling data to improve the problems existing in the prior art.

[0005] The embodiments of the present invention can be implemented as follows:

[0006] In a first aspect, the present invention provides a method for storing signaling data, including:

[0007] Obtain a signaling data set within a specified time in a target area; the signaling data set includes a number of signaling data, and each piece of signaling data includes a user code, a category code, an interaction time, and other information;

[0008] Perform compression processing on the interaction time in each piece of signaling data to obtain a compressed signaling corresponding to each piece of signaling data;

[0009] Based on the user code and the category code, perform classification processing and secondary compression processing on all the compressed signals to obtain data to be stored corresponding to each of the multiple category codes corresponding to each user code;

[0010] Column - store all the data to be stored using the user code as an index.

[0011] Optionally, the step of performing compression processing on the interaction time in each piece of signaling data to obtain a compressed signaling corresponding to each piece of signaling data includes:

[0012] For each piece of the signaling data, read the interaction time in the signaling data;

[0013] Convert both the interaction time and the zero moment of the day when the interaction time is located into timestamps, respectively obtaining the interaction timestamp corresponding to the interaction time and the reference timestamp;

[0014] Calculate the difference between the interaction timestamp and the reference timestamp to obtain the reference offset;

[0015] Add the reference offset to the signaling data and delete the interaction time in the signaling data to obtain the compressed signaling.

[0016] Optionally, the compressed signaling includes the reference offset corresponding to the interaction time in the signaling data;

[0017] The step of classifying and secondarily compressing all the compressed signaling based on the user code and the category code to obtain the data to be stored corresponding to each category code of each user code includes:

[0018] Classify all the compressed signaling according to the user code to obtain a user set corresponding to each user code;

[0019] Divide each user set into multiple category subsets according to the category code; wherein, the user codes and category codes of all the compressed signaling in the category subset are the same;

[0020] Sort all the compressed signaling in each category subset in ascending order according to the reference offset to obtain an ordered category subset corresponding to each category subset;

[0021] Use the user code and the category code as the primary index and the secondary index respectively to obtain multiple index combinations; each index combination uniquely corresponds to an ordered category subset;

[0022] For each index combination, serialize all the reference offsets and other information in the ordered category subset corresponding to the index combination to obtain the data to be stored corresponding to each index combination.

[0023] Optionally, the other information includes data corresponding to at least one information field;

[0024] The step of serializing all the reference offsets and other information in the ordered category subset corresponding to the index combination includes:

[0025] Determine the number of signaling in the ordered category subset corresponding to the index combination;

[0026] Serialize the reference offsets of all the compressed signaling in the ordered category subset corresponding to the index combination to obtain an offset list;

[0027] Using the run-length encoding rule, serialize the data corresponding to each information field of all the compressed signaling in the ordered category subset corresponding to the index combination to obtain a compressed list corresponding to each information field;

[0028] Combine the signaling quantity, the offset list, and the compressed list corresponding to each information field to obtain the data to be stored corresponding to the index combination.

[0029] Optionally, the step of serializing the reference offsets of all the compressed signaling in the ordered category subset corresponding to the index combination to obtain an offset list includes:

[0030] Take the reference offset of the first compressed signaling in the ordered category subset as the reference offset;

[0031] Calculate the difference between the reference offset of the i-th compressed signaling in the ordered category subset and the reference offset to obtain the compressed timestamp corresponding to the i-th compressed signaling; where, , K represents the number of signaling in the ordered category subset;

[0032] Combine the reference offset and the compressed timestamps corresponding to the K compressed signaling in the ordered category subset in sequence to obtain the offset list.

[0033] Optionally, the step of storing all the data to be stored column by column with the user encoding as the index includes:

[0034] Based on the offset list and the Schema information of each compressed list, store the offset list and each compressed list in the data to be stored corresponding to each index combination in the form of a singly linked list column by column to obtain a singly linked list storage structure corresponding to each index combination;

[0035] Based on the Schema information of the main index and the signaling quantity respectively, store each main index and the signaling quantity corresponding to each index combination in the form of a double-pointer linked list column by column; the Schema information reflects the field name, the data type to which the field belongs, the data structure, and the data organization rule;

[0036] Among them, in the data block where the main index is located, the first pointer is the memory address storing the number of signaling messages corresponding to the first index combination where the main index is located. If the main index is not the last one, the second pointer is the memory address storing the next main index of the main index. If the main index is the last one, the second pointer is a null pointer;

[0037] In the data block where the number of signaling messages corresponding to the index combination is located, the first pointer is the memory address of the first data block in the singly-linked list storage structure corresponding to the index combination. If the index combination is not the last one among the multiple index combinations with the same user code, the second pointer is the memory address storing the number of signaling messages corresponding to the next index combination of the current index combination. If the index combination is the last one among the multiple index combinations with the same user code, the second pointer is a null pointer.

[0038] In a second aspect, the present invention provides a signaling data storage device, including:

[0039] An acquisition module, configured to acquire a signaling data set within a target area within a specified time; the signaling data set includes a number of signaling data, and each piece of the signaling data includes a user code, a category code, an interaction time, and other information;

[0040] A time compression module, configured to perform compression processing on the interaction time in each piece of the signaling data to obtain a compressed signaling corresponding to each piece of the signaling data;

[0041] A data compression module, configured to perform classification processing and secondary compression processing on all the compressed signaling based on the user code and the category code to obtain data to be stored corresponding to each of the multiple category codes corresponding to each user code;

[0042] A data writing module, configured to perform columnar storage on all the data to be stored.

[0043] In a third aspect, the present invention provides an electronic device, including: a memory and a processor, where the memory stores a software program, and when the electronic device runs, the processor executes the software program to implement the method as described in the first aspect above.

[0044] In a fourth aspect, the present invention provides a computer-readable storage medium, where the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method as described in the first aspect above is implemented.

[0045] In a fifth aspect, the present invention provides a program product, and when the program product is executed by a processor, the method as described in the first aspect above is implemented.

[0046] Compared with the prior art, the embodiment of the present invention provides a method, device, electronic device, storage medium and program product for storing signaling data. After obtaining a number of signaling data within a specified time in a target area, first, the interaction time in each signaling data is compressed to obtain a compressed signaling corresponding to each signaling data, which can avoid occupying a large amount of storage resources for the interaction time. Then, based on the user code and category code, all the compressed signals are classified and secondarily compressed to obtain the data to be stored corresponding to each of the multiple category codes corresponding to each user code. Finally, all the data to be stored is stored in a columnar manner. After classification and secondary compression and then combined with columnar storage, it can further save storage resources. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required in the embodiments. It should be understood that the following drawings only show some embodiments of the present invention and should not be regarded as limiting the scope. For those of ordinary skill in the art, other related drawings can be obtained based on these drawings without creative efforts.

[0048] Figure 1 FIG. 1 is one of the flow diagrams of a method for storing signaling data provided by an embodiment of the present invention.

[0049] Figure 2 FIG. 2 is another flow diagram of a method for storing signaling data provided by an embodiment of the present invention.

[0050] Figure 3 FIG. 3 is an example diagram of a plurality of ordered category subsets provided by an embodiment of the present invention.

[0051] Figure 4 FIG. 4 is an example diagram of the data to be stored corresponding to a plurality of index combinations provided by an embodiment of the present invention.

[0052] Figure 5 FIG. 5 is an example diagram of a single-linked list storage structure provided by an embodiment of the present invention.

[0053] Figure 6 FIG. 6 is an example diagram of a double-linked list storage structure provided by an embodiment of the present invention.

[0054] Figure 7 FIG. 7 is a structural diagram of a device for storing signaling data provided by an embodiment of the present invention.

[0055] Figure 8 FIG. 8 is a structural diagram of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0056] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. Components of the embodiments of the present invention generally described and illustrated in the figures herein may be arranged and designed in a variety of different configurations.

[0057] Therefore, the following detailed description of the embodiments of the present invention provided in the drawings is not intended to limit the scope of the claimed invention, but is merely representative of selected embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.

[0058] It should be noted that like reference numerals and letters denote like items in the following figures, and thus, once an item is defined in one figure, it need not be further defined and explained in subsequent figures.

[0059] In the description of the present invention, it should be noted that if terms such as "upper", "lower", "inner", "outer", etc. are used to indicate the orientation or positional relationship, it is based on the orientation or positional relationship shown in the drawings or the orientation or positional relationship in which the product of the present invention is usually placed during use. It is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and thus should not be construed as a limitation of the present invention.

[0060] In addition, terms such as "first", "second", etc. are only used for descriptive distinction and should not be construed as indicating or implying relative importance.

[0061] It should be noted that the features in the embodiments of the present invention may be combined with each other without conflict.

[0062] Currently, mobile communication operators mainly use traditional distributed storage systems to store signaling data, and its technical implementation usually includes the following steps:

[0063] (1) The base station uploads signaling data such as user connection requests, handover instructions, and location updates generated in real time to the data center;

[0064] (2) The data center indexes the signaling data by time series or user ID. The method of writing the signaling data into a distributed database or file system is as follows: after converting the time in the signaling data into a timestamp, converting the data content in the signaling data into a binary sequence, and then performing a series of transformations on the binary sequence to achieve data compression, and finally storing it.

[0065] However, the data type of the timestamp is long, which requires 8 bytes of storage. That is, in a signaling data containing a timestamp, the space occupied by the timestamp exceeds 56% of the space occupied by the entire signaling data. Therefore, even after compression by converting the timestamp to binary, the occupied space is still relatively large.

[0066] Based on the discovery of the above technical problems, the inventor believes that the timestamp can be further compressed to reduce the space occupied by time. And through long-term observation and research, the inventor found that: First, signaling data has characteristics such as high-frequency generation, strong content repeatability (such as periodic location updates), and high field redundancy. Therefore, the existing method of directly storing single signaling data has a low storage resource utilization rate; Second, the existing technology lacks in-depth exploration of the spatio-temporal correlation of signaling data. For example, multiple signaling generated by the same user within adjacent time periods often contain repeated base station identifiers or location information, and the existing storage architecture fails to effectively reduce the space occupied by such repeated content.

[0067] In view of this, an embodiment of the present invention provides a method for storing signaling data. On the one hand, it can initially compress the interaction time to obtain compressed signaling, and on the other hand, based on user coding and category coding, classify all compressed signaling and then perform secondary compression. In this way, the repeated information appearing in all compressed signaling corresponding to each category coding of a user coding is compressed and then stored in a columnar format, greatly reducing the storage space occupation. The following will be described in detail through embodiments and in conjunction with the accompanying drawings.

[0068] Please refer to Figure 1 , Figure 1 which is a schematic flowchart of a method for storing signaling data provided by an embodiment of the present invention. The execution subject of this method can be, but is not limited to, computing devices, servers, server clusters, etc. This method includes the following steps S101 to S104.

[0069] S101. Obtain a signaling data set within a specified time in a target area.

[0070] In this embodiment, the signaling data set includes several signaling data. Each signaling data includes a user coding (representing the user terminal used by the user), a category coding (reflecting the event category of signaling interaction, such as attachment request, handover instruction, tracking area update, service request, etc.), an interaction time, and other information. The other information includes data corresponding to multiple information fields.

[0071] Optionally, the target area can be the whole country or a city, such as Beijing, Chengdu, etc., and the specified time period can be one day, multiple days, one week, one month, etc. In the corresponding example, the signaling data set can include all the signaling data generated in Chengdu in April 25. This example is only for illustration, and the embodiments of the present invention are not limited thereto.

[0072] S102. Compress the interaction time in each signaling data to obtain the compressed signaling corresponding to each signaling data.

[0073] In this embodiment, compressing the interaction time in each signaling data can reduce the storage resource occupation of the time information in the subsequent signaling.

[0074] S103. Based on the user code and category code, classify and secondarily compress all the compressed signals to obtain the data to be stored corresponding to each of the multiple category codes corresponding to each user code.

[0075] In an optional example, for 50 compressed signals with the same user code, if there are 4 different category codes among the 50 compressed signals, then the 50 compressed signals will be compressed into 4 data to be stored. This example is only for illustration and is not limited herein.

[0076] S104. Store all the data to be stored column by column with the user code as the index.

[0077] In this embodiment, compared with row storage, column storage can reduce the storage resource occupation.

[0078] For the signaling data storage method provided by the embodiments of the present invention, after obtaining a number of signaling data within the specified time in the target area, first compress the interaction time in each signaling data to obtain the compressed signaling corresponding to each signaling data, which can avoid a large amount of storage resource occupation of the interaction time. Then, based on the user code and category code, classify and secondarily compress all the compressed signals to obtain the data to be stored corresponding to each of the multiple category codes corresponding to each user code. Finally, store all the data to be stored column by column. After such classification and secondary compression and combined with column storage, it can further save storage resources.

[0079] Because the definition of the timestamp is the total number of seconds from Greenwich Mean Time 00:00:00 on January 1, 1970 (i.e., 08:00:00 on January 1, 1970 in Beijing time) to the present. For example, for Beijing time 2025-04-03 12:00:00, the corresponding timestamp is 1743652800, a total of 10 digits. And there are only 86400 seconds in a day, so the maximum timestamp difference within the same day is 86400, a total of 5 digits.

[0080] Therefore, in the above step S102, the zero point of the day when the interaction time is located can be used as the reference time, and then the interaction time and the reference time are both converted into timestamps and the difference is calculated, so as to realize the preliminary compression of the interaction time, and the interaction time can be compressed to a 5-digit decimal number. And such calculation operations are simple addition and subtraction operations, and the algorithm complexity is only O(n), that is, the calculation complexity will not be increased.

[0081] That is, in the above step S102, the process of "compressing the interaction time in each signaling data to obtain the compressed signaling corresponding to each signaling data" may include the following sub-steps S1021 to S1024.

[0082] S1021. For each signaling data, read the interaction time in the signaling data.

[0083] S1022. Convert both the interaction time and the zero point of the day when the interaction time is located into timestamps, and respectively obtain the interaction timestamp corresponding to the interaction time and the reference timestamp.

[0084] In this embodiment, the timestamp corresponding to the interaction time is the interaction timestamp, the zero point of the day when the interaction time is located is the reference time, and the timestamp corresponding to the reference time is the reference timestamp.

[0085] S102. Calculate the difference between the interaction timestamp and the reference timestamp to obtain the reference offset.

[0086] S1024. Add the reference offset to the signaling data and delete the interaction time in the signaling data to obtain the compressed signaling.

[0087] In this embodiment, for each signaling data in the signaling dataset, steps S1021 to S1024 are executed, and the time in each signaling data can be preliminarily compressed.

[0088] In an optional example, the data type of the reference offset can be int type. Assume that the interaction time is 12:00:00 on April 3, 2025, Beijing time, and the corresponding interaction timestamp is 1743652800. Then the reference time is 00:00:00 on April 3, 2025, Beijing time, and the corresponding reference timestamp is 1743609600. Then the reference offset is 43200. In this way, compared with the existing technology that requires 8 bytes to store the interaction timestamp, through the preliminary compression of the interaction time, the reference offset only occupies 4 bytes, that is, the preliminary compression method can greatly reduce the storage resource occupation of the interaction time. This example is only for illustration, and the interaction time can also be accurate to milliseconds. The embodiments of the present invention do not limit the accuracy of the interaction time.

[0089] In an alternative implementation, since signaling data has characteristics such as high-frequency generation, strong content repeatability (such as periodic location updates), and high field redundancy, secondary compression can be achieved by using a serialization process after classification. Figure 1 Based on this, please refer to Figure 2 , for the process of "classifying and secondarily compressing all compressed signaling based on user coding and category coding to obtain the data to be stored corresponding to each category coding for each user coding" in step S103 above, it may include the following sub-steps S1031 to S1035.

[0090] S1031. Classify all compressed signaling according to the user coding to obtain a user set corresponding to each user coding.

[0091] S1032. Divide each user set into multiple category subsets according to the category coding.

[0092] In this embodiment, in a category subset, the user coding and category coding of all compressed signaling are the same.

[0093] S1033. Sort all compressed signaling in each category subset in ascending order of the reference offset to obtain an ordered category subset corresponding to each category subset.

[0094] In an alternative example, it is assumed that other information may include: base station cell coding, base station location coding, flag value, service information, etc., that is, the information fields include: Cid field, Ghash field, Flag field, Amto field, etc.

[0095] Please refer to Figure 3 , Figure 3 Taking the compressed signaling including Uid field (i.e., Tid field), Offset field, Cid field, Ghash field, and Flag field as an example, multiple ordered category subsets are shown, Figure 3 Those with the same serial number color in are an ordered category subset.

[0096] It should be noted that Figure 3 the serial number column in is only for facilitating the display of the number of signaling in each subset, Figure 3 using numbers such as 0, 1, 2, etc. to identify different category codings in is only an example and is not limited here.

[0097] S1034. Use the user coding and category coding as the primary index and secondary index respectively to obtain multiple index combinations.

[0098] In this embodiment, an index combination includes a primary index and a secondary index, so each index combination corresponds uniquely to an ordered category subset.

[0099] If the signaling data set involves 3 user codes, and all the signaling data corresponding to the 3 user codes involve 4 category codes, then 12 index combinations can be determined. For example Figure 3 There are 8 index combinations. This example is only for illustration and is not limited here.

[0100] S1035. For each index combination, serialize all the reference offsets and other information in the ordered category subset corresponding to the index combination to obtain the data to be stored corresponding to each index combination.

[0101] In this embodiment, it is necessary to serialize the reference offsets of all the compressed signaling and the data of each information field in each ordered category subset respectively, so as to reduce duplicate data.

[0102] Optionally, for each index combination, in step S1035, the process of "serializing all the reference offsets and other information in the ordered category subset corresponding to the index combination" may include the following sub-steps S10351 to S10354.

[0103] S10351. Determine the number of signaling in the ordered category subset corresponding to the index combination.

[0104] S10352. Serialize the reference offsets of all the compressed signaling in the ordered category subset corresponding to the index combination to obtain an offset list.

[0105] Optionally, taking the reference offset of the first compressed signaling in the ordered category subset as a benchmark, and taking the difference again can reduce the space occupied by time information. For the ordered category subset corresponding to an index combination, the process of obtaining the offset list may include steps S001 to S003:

[0106] S001. Take the reference offset of the first compressed signaling in the ordered category subset as the benchmark offset;

[0107] S002. Calculate the difference between the reference offset of the i-th compressed signaling in the ordered category subset and the benchmark offset to obtain the compressed timestamp corresponding to the i-th compressed signaling; where , K represents the number of signaling in the ordered category subset;

[0108] S003. Combine the benchmark offset and the compressed timestamps corresponding to the K compressed signaling in the ordered category subset in sequence to obtain the offset list.

[0109] In this embodiment, after ordering the K reference offsets in the ordered category subset corresponding to an index combination according to steps S001 to S003, the obtained offset list includes K + 1 values.

[0110] S10353. Using the run-length encoding rule, serialize the data corresponding to each information field of all the compressed signaling in the ordered category subset corresponding to the index combination, respectively, to obtain a compressed list corresponding to each information field.

[0111] In this embodiment, since most of the information fields are of character type rather than numerical type, for each information field of other information: the run-length encoding rule can be used to process the data corresponding to all the compressed signaling in the ordered category subset corresponding to an index combination into a compressed list.

[0112] Among them, the run-length encoding rule, namely Run-Length Encoding, also known as the RLE encoding rule, is a lossless data compression algorithm. The core idea is to shorten the data length by recording consecutive repeated characters (or symbols) and their occurrence times. Therefore, the specific process of using the run-length encoding rule in the present invention will not be introduced.

[0113] S10354. Combine the signaling quantity, the offset list, and the compressed list corresponding to each information field to obtain the data to be stored corresponding to the index combination.

[0114] In this embodiment, the data to be stored corresponding to a so combination includes the signaling quantity, the offset list, and the compressed lists corresponding to each information field.

[0115] In an optional example, please continue to combine Figure 3 , if the three Uids of Figure 3 "728669839005169607", "311553764816986473", "675786698390053333" are respectively denoted as X1, X2, X3, then Figure 3 the data to be stored obtained by converting the 8 ordered category subsets in Figure 4 is as shown in

[0116] In Figure 4 , the compressed lists corresponding to information fields such as the Cid field, the Ghash field, and the Flag field are Cid_list, Ghash_list, and Flag_list respectively, and according to the above steps S10351 to S10354 for Figure 3 the first ordered category subset (the row with the green serial number) and the second ordered category subset (the row with the red serial number) in Figure 4The content in the green area and the content in the red area.

[0117] Comparison Figure 3 and Figure 4 It can be intuitively seen that by separately ordering the reference offsets (i.e., the Offset field) for each ordered category subset and each information field, data ordered compression is achieved, which can reduce the storage resources required for storage.

[0118] It should be noted that Figure 3 The Uid shown is only an example. Figure 3 , Figure 4 Each information field shown is only an example, and the specific content in the signaling data corresponding to the embodiments of the present invention is not limited.

[0119] In an optional implementation manner, the sub-steps of the above step S104 may include:

[0120] S1041. Based on the Schema information of the offset list and each compression list respectively, store the offset list and each compression list in the data to be stored corresponding to each index combination in a columnar form as a singly linked list, to obtain a singly linked list storage structure corresponding to each index combination.

[0121] In this embodiment, the Schema information can reflect the field name, the data type to which the field belongs, the data structure, and the data organization rule. And the Schema information also needs to be stored, and after storage, it is metadata.

[0122] It can be understood that in the singly linked list storage structure, multiple data in the list are each located in a data block. Taking the offset list as an example, in the singly linked list storage structure, the base offset of the offset list and the K compressed timestamps are respectively located in K + 1 data blocks.

[0123] And the data block is divided into a data domain and a pointer domain. The pointer domain of each data block in the singly linked list storage structure only includes one pointer, which is used to store the memory address of the next data block.

[0124] S1042. Based on the Schema information of the main index and the signaling quantity respectively, store each main index and the signaling quantity corresponding to each index combination in a columnar form as a double-pointer linked list.

[0125] In this embodiment, each main index and the signaling quantity corresponding to each index combination each occupy a data block. By storing each main index and the signaling quantity corresponding to each index combination in a columnar form as a double-pointer linked list, a double-linked list storage structure can be obtained.

[0126] Assume that a total of M user codes are involved in the signaling dataset, that is, there are M main indexes. Then, for the m-th main index (m ∈ [1, M]), the pointer field in the data block where it is located includes 2 pointers. The first pointer is the memory address where the signaling quantity corresponding to the first index combination where the main index is located is stored. The second pointer is divided into the following two cases:

[0127] (1) Among the M main indexes, if the m-th main index is not the last one (i.e., m ≠ M), then the second pointer is the memory address where the next main index of the m-th main index is stored, that is, the second pointer is the memory address where the (m + 1)-th main index is stored;

[0128] (2) Among the M main indexes, if the m-th main index is the last one (i.e., m = M), then the second pointer is a null pointer.

[0129] Assume that N index combinations are obtained in step S1034. Then, for the n-th index combination (n ∈ [1, N]), the pointer field in the data block where the corresponding signaling quantity is located includes 2 pointers. The first pointer is the memory address of the first data block in the singly-linked list storage structure corresponding to the n-th index combination. Assume that among the N index combinations, there are Q index combinations with the same user code as that in the n-th index combination. Then, the second pointer is divided into the following two cases:

[0130] (1) If the n-th index combination is not the last one among the Q index combinations with the same user code, then the second pointer is the memory address where the signaling quantity corresponding to the next index combination of the n-th index combination is stored, that is, the second pointer is the memory address where the signaling quantity corresponding to the (n + 1)-th index combination is stored;

[0131] (2) If the n-th index combination is the last one among the Q index combinations with the same user code, then the second pointer is a null pointer.

[0132] To understand the above-mentioned singly-linked list storage structure and doubly-linked list storage structure, the following is illustrated with examples.

[0133] In the optional example, please combine Figure 4 , Figure 4 The 8 index combinations and their corresponding signaling quantities are shown in Table 1 below. And Figure 4 In the green area shown, the 4 lists of Offset_list, Cid_list, Ghash_list, and Flag_list corresponding to the index combination (X1,0) are shown in Table 2 below.

[0134] Table 1

[0135]

[0136] Table 2. Four lists corresponding to (X1,0)

[0137]

[0138] First, the four lists in Table 2 are stored column by column in sequence, and the singly linked list storage structure corresponding to (X1,0) is obtained as Figure 5 shown. And Figure 4 the singly linked list storage structures corresponding to the other seven index combinations are similar to Figure 5 this, and will not be repeated here.

[0139] Next, the primary index and the signaling quantity shown in Table 1 are stored, and the obtained doubly linked list storage structure is as Figure 6 shown.

[0140] In Figure 5 and Figure 6 , the direction of the arrow represents the location where the memory address of the pointer is located, and NULL represents a null pointer.

[0141] Combined with Figure 6 it can be seen that the primary index uses a double pointer during storage, which ensures that the primary index can perform downward addressing to read other primary indexes, and can also perform rightward addressing to read the corresponding signaling quantity. And the signaling quantity uses a double pointer during storage, which ensures that the signaling quantity can perform downward addressing to read the next signaling quantity with the same primary index, and can also perform rightward addressing to read the singly linked list storage structure.

[0142] Combined with Figure 5 it can be seen that the four lists corresponding to (X1,0) achieve continuous addressing through pointers during storage to ensure the continuous readability of data.

[0143] Combined with Figure 3 , Figure 4 , Figure 5 , Figure 6 it can be known that due to the relatively small time stamp difference on the same day, the present invention compresses the interaction time twice, thereby compressing the K interaction times corresponding to the same index combination into one Offset_list, greatly reducing the storage space occupied by time data.

[0144] Meanwhile, since there is strong content repetition in multiple signaling data with close interaction times and the same user encoding, the present invention also compresses the corresponding data of each information field in the K signaling data corresponding to the same index combination into a compressed list based on the run-length encoding rule, thereby reducing the space occupied by duplicate data. The present invention uses the user encoding as the main index and the category encoding as the secondary index to determine multiple index combinations, and stores the main index of each index combination, the number of signaling corresponding to each index combination, the offset list, and the compressed list corresponding to each information field in a columnar storage manner, and uses pointers to achieve addressing in the storage structure to ensure data readability.

[0145] The above content describes the compression storage of a signaling data set to obtain a doubly linked list storage structure and a singly linked list storage structure corresponding to each index combination.

[0146] Based on the doubly linked list storage structure and the singly linked list storage structure corresponding to each index combination, the process of restoring a signaling data set is the reverse process of compression storage. Taking a specified day as the specified time as an example, the restoration process is briefly introduced as follows:

[0147] (1) First, read all the data from the doubly linked list storage structure and the singly linked list storage structure corresponding to each index combination, and add the secondary index (i.e., the category encoding) according to the order of the number of signaling corresponding to each main index;

[0148] (2) For the offset list (including K + 1 values) corresponding to each index combination, starting from the second value in the list, successively add each data to the first value in the list, so as to restore K reference offsets corresponding to an index combination. Then, add the K reference offsets to the time stamp at the zero moment of the specified day to obtain K interaction time stamps corresponding to an index combination, and further convert them into K interaction times corresponding to an index combination;

[0149] (3) For the compressed list of each information field corresponding to each index combination, based on the run-length encoding rule, restore each compressed list to obtain K original field values corresponding to each information field corresponding to an index combination;

[0150] (4) For each index combination, combine the user encoding, category encoding, corresponding K interaction times, and K original field values corresponding to each information field in the index combination to obtain K signaling data with the same user encoding and category encoding.

[0151] Verified by the inventor, for 3.166 billion signaling data, the storage occupancy is 37.2 GB. After compression processing using the method of the present invention, the storage occupancy is only 19.7 GB, which is nearly half of the original. And the signaling data can be restored through the restoration process introduced above. Therefore, the present invention can reduce the storage space occupied by a large amount of signaling data on the premise of lossless compression.

[0152] To execute the corresponding steps in the above method embodiments and each possible implementation manner, an implementation manner of a storage device for signaling data is given below.

[0153] Please refer to Figure 7 , Figure 7 which shows a schematic structural diagram of a storage device for signaling data provided by an embodiment of the present invention. The storage device 200 for signaling data includes: an acquisition module 210, a time compression module 220, a data compression module 230, and a data writing module 240.

[0154] The acquisition module 210 is configured to acquire a signaling data set within a specified time in a target area; the signaling data set includes a plurality of signaling data, and each signaling data includes a user code, a category code, an interaction time, and other information;

[0155] The time compression module 220 is configured to perform compression processing on the interaction time in each signaling data to obtain a compressed signaling corresponding to each signaling data;

[0156] The data compression module 230 is configured to perform classification processing and secondary compression processing on all the compressed signaling based on the user code and the category code to obtain data to be stored corresponding to each of the multiple category codes corresponding to each user code;

[0157] The data writing module 240 is configured to perform columnar storage on all the data to be stored.

[0158] Optionally, the time compression module 220 may specifically be configured to: for each signaling data, read the interaction time in the signaling data; convert the interaction time and the zero moment of the day when the interaction time is located into timestamps, respectively obtaining an interaction timestamp corresponding to the interaction time and a reference timestamp; calculate the difference between the interaction timestamp and the reference timestamp to obtain a reference offset; add the reference offset to the signaling data and delete the interaction time in the signaling data to obtain the compressed signaling.

[0159] Optionally, the compressed signaling includes a reference offset corresponding to the interaction time in the signaling data. The data compression module 230 can specifically be used to: classify all the compressed signaling according to the user code to obtain a user set corresponding to each user code; divide each user set into multiple category subsets according to the category code; where all the user codes and category codes of the compressed signaling in the category subset are the same; sort all the compressed signaling in each category subset in ascending order of the reference offset to obtain an ordered category subset corresponding to each category subset; use the user code and the category code as the primary index and the secondary index respectively to obtain multiple index combinations; each index combination corresponds uniquely to an ordered category subset; for each index combination, serialize all the reference offsets and other information in the ordered category subset corresponding to the index combination to obtain the data to be stored corresponding to each index combination.

[0160] Optionally, the other information includes the data corresponding to at least one information field. In the process of the data compression module 230 being used to serialize all the reference offsets and other information in the ordered category subset corresponding to the index combination, it can specifically be used to: determine the number of signaling in the ordered category subset corresponding to the index combination; serialize the reference offsets of all the compressed signaling in the ordered category subset corresponding to the index combination to obtain an offset list; use the run-length encoding rule to serialize the data corresponding to each information field of all the compressed signaling in the ordered category subset corresponding to the index combination respectively to obtain a compressed list corresponding to each information field; combine the number of signaling, the offset list, and the compressed list corresponding to each information field to obtain the data to be stored corresponding to the index combination.

[0161] Optionally, in the process of the data compression module 230 being used to serialize the reference offsets of all the compressed signaling in the ordered category subset corresponding to the index combination to obtain an offset list, it can specifically be used to: use the reference offset of the first compressed signaling in the ordered category subset as the reference offset; calculate the difference between the reference offset of the i-th compressed signaling in the ordered category subset and the reference offset to obtain the compressed timestamp corresponding to the i-th compressed signaling; where , K represents the number of signaling in the ordered category subset; combine the reference offset and the compressed timestamps corresponding to the K compressed signaling in the ordered category subset in sequence to obtain the offset list.

[0162] Optionally, the data writing module 240 may specifically be configured to: based on the offset list and the Schema information of each compression list, store the offset list and each compression list in the data to be stored corresponding to each index combination in a columnar manner in the form of a singly linked list, to obtain a singly linked list storage structure corresponding to each index combination; based on the Schema information of the main index and the signaling quantity respectively, store each main index and the signaling quantity corresponding to each index combination in a columnar manner in the form of a double-pointer linked list; the Schema information reflects the field name, the data type to which the field belongs, the data structure, and the data organization rule.

[0163] Among them, in the data block where the main index is located, the first pointer is the memory address where the signaling quantity corresponding to the first index combination where the main index is located is stored. If the main index is not the last one, the second pointer is the memory address where the next main index of the main index is stored. If the main index is the last one, the second pointer is a null pointer.

[0164] Among them, in the data block where the signaling quantity corresponding to the index combination is located, the first pointer is the memory address of the first data block in the singly linked list storage structure corresponding to the index combination. If the index combination is not the last one among the multiple index combinations with the same user code, the second pointer is the memory address where the signaling quantity corresponding to the next index combination of the current index combination is stored. If the index combination is the last one among the multiple index combinations with the same user code, the second pointer is a null pointer.

[0165] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working process of the signaling data storage device 200 described above can refer to the corresponding process in the foregoing method embodiment, and will not be elaborated here.

[0166] Please refer to Figure 8 , Figure 8 which is a schematic structural diagram of an electronic device provided by an embodiment of the present invention. The electronic device 300 includes a processor 310, a memory 320, and a bus 330. The processor 310 is connected to the memory 320 through the bus 330.

[0167] The memory 320 can be used to store software programs. For example, the software program corresponding to the signaling data storage device 200 provided by the embodiment of the present invention. The processor 310 executes various functional applications and data processing by running the software program stored in the memory 320, so as to implement the signaling data storage method provided by the embodiment of the present invention.

[0168] Among them, the memory 320 can be but is not limited to: RAM (Random Access Memory), ROM (Read Only Memory), FLASH (Flash Memory), PROM (Programmable Read-Only Memory), EPROM (Erasable Programmable Read-Only Memory), EEPROM (Electric Erasable Programmable Read-Only Memory), etc.

[0169] The processor 310 can be an integrated circuit chip with signal processing capabilities. The processor 310 can be a general-purpose processor, including: CPU (Central Processing Unit), NP (Network Processor), SoC (System on Chip), etc.; it can also be: DSP (Digital Signal Processing), ASIC (Application Specific Integrated Circuit), FPGA (Field-Programmable Gate Array), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.

[0170] It can be understood that Figure 8 The structure shown is only for illustration, and the electronic device 300 can also include more or fewer components than those shown Figure 8 in the figure, or have a different configuration from that shown Figure 8 in the figure. Figure 8 Each component shown in the figure can be implemented by hardware, software, or a combination thereof.

[0171] An embodiment of the present invention also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is run by a processor, it implements the signaling data storage method disclosed in the above embodiment. The computer-readable storage medium can be but is not limited to: USB flash drives, mobile hard disks, ROM, RAM, PROM, EPROM, EEPROM, FLASH magnetic disks, or optical discs, etc., various media that can store program codes.

[0172] An embodiment of the present invention also provides a program product, which implements the signaling data storage method disclosed in the above embodiment when run by a processor.

[0173] In summary, the embodiments of the present invention provide a method, apparatus, electronic device, storage medium, and program product for storing signaling data. After obtaining a number of signaling data within a specified time in a target area, first, the interaction time in each signaling data is compressed to obtain a compressed signaling corresponding to each signaling data, which can avoid occupying a large amount of storage resources for the interaction time. Then, based on the user code and category code, all the compressed signals are classified and secondarily compressed to obtain the data to be stored corresponding to each of the multiple category codes corresponding to each user code. Finally, all the data to be stored is stored in a columnar manner. After such classification and secondary compression, combined with columnar storage, it can further save storage resources.

[0174] The above is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed by the present invention should be covered by the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.

Claims

1. A method for storing signaling data, characterized in that: include: Obtain the signaling data set within the target area within the specified time; The signaling data set includes a plurality of signaling data, each of which includes a user code, a category code, an interaction time and other information; Compressing the interaction time in each piece of the signaling data to obtain compressed signaling corresponding to each piece of the signaling data; Based on the user code and the category code, all the compressed signaling are classified and compressed twice to obtain data to be stored corresponding to each of the multiple category codes corresponding to each user code; Using the user code as an index, all the data to be stored are stored in column format; The compressed signaling includes a reference offset corresponding to the interaction time in the signaling data; the step of classifying and secondary compressing all the compressed signaling based on the user code and the category code to obtain the data to be stored corresponding to each of the multiple category codes corresponding to each user code includes: Classifying all the compressed signaling according to the user codes to obtain a user set corresponding to each user code; Dividing each of the user sets into a plurality of category subsets according to the category code; wherein the user codes and category codes of all compressed signaling in the category subsets are the same; Sorting all compressed signaling in each of the category subsets in ascending order of the reference offsets to obtain an ordered category subset corresponding to each of the category subsets; The user code and the category code are used as the primary index and the secondary index respectively to obtain a plurality of index combinations; each of the index combinations uniquely corresponds to one of the ordered category subsets; For each of the index combinations, all reference offsets and other information in the ordered category subset corresponding to the index combination are serialized to obtain the data to be stored corresponding to each of the index combinations.

2. The method according to claim 1, characterized in that The step of compressing the interaction time in each piece of the signaling data to obtain compressed signaling corresponding to each piece of the signaling data includes: For each piece of the signaling data, reading the interaction time in the signaling data; Convert the interaction time and the zero o'clock time of the day on which the interaction time falls into timestamps, and obtain an interaction timestamp and a reference timestamp corresponding to the interaction time respectively; Calculate the difference between the interaction timestamp and the reference timestamp to obtain a reference offset; The reference offset is added to the signaling data, and the interaction time in the signaling data is deleted to obtain the compressed signaling.

3. The method according to claim 1, characterized in that The other information includes data corresponding to at least one information field; The step of serializing all reference offsets and other information in the ordered category subset corresponding to the index combination includes: Determining the number of signaling in the ordered category subset corresponding to the index combination; Serializing the reference offsets of all compressed signaling in the ordered category subset corresponding to the index combination to obtain an offset list; Using the run-length encoding rule, serializing the data corresponding to each information field of all compressed signaling in the ordered category subset corresponding to the index combination, to obtain a compression list corresponding to each information field; The signaling quantity, the offset list, and the compression list corresponding to each of the information fields are combined to obtain the data to be stored corresponding to the index combination.

4. The method according to claim 3, characterized in that The step of serializing the reference offsets of all compressed signaling in the ordered category subset corresponding to the index combination to obtain an offset list includes: Using the reference offset of the first compressed signaling in the ordered category subset as the reference offset; Calculate the difference between the reference offset of the i-th compressed signaling in the ordered category subset and the reference offset to obtain the compressed timestamp corresponding to the i-th compressed signaling; wherein, , K represents the number of signals in the ordered category subset; The reference offset and the compressed timestamps corresponding to the K compressed signalings in the ordered category subset are sequentially combined to obtain the offset list.

5. The method according to claim 3, characterized in that: The step of storing all the data to be stored in column format using the user code as an index includes: Based on the Schema information of the offset list and each compression list, the offset list and each compression list in the data to be stored corresponding to each index combination are stored in a columnar manner in a single linked list to obtain a single linked list storage structure corresponding to each index combination; Based on the respective Schema information of the primary index and the signaling quantity, the signaling quantity corresponding to each primary index and each index combination is stored in a columnar form in the form of a double pointer linked list; the Schema information reflects the field name, the data type to which the field belongs, the data structure and the data organization rule; Among them, in the data block where the main index is located, the first pointer is the memory address where the signaling quantity corresponding to the first index combination where the main index is located is stored. If the main index is not the last one, the second pointer is the memory address where the next main index of the main index is stored. If the main index is the last one, the second pointer is a null pointer; In the data block where the signaling quantity corresponding to the index combination is located, the first pointer is the memory address of the first data block in the single linked list storage structure corresponding to the index combination; if the index combination is not the last of the multiple index combinations with the same user code, the second pointer is the memory address where the signaling quantity corresponding to the next index combination of the current index combination is stored; if the index combination is the last of the multiple index combinations with the same user code, the second pointer is a null pointer.

6. A storage device for signaling data, characterized in that: include: An acquisition module is used to acquire a signaling data set in a target area within a specified time; The signaling data set includes a plurality of signaling data, each of which includes a user code, a category code, an interaction time and other information; A time compression module, used to compress the interaction time in each piece of the signaling data to obtain a compressed signaling corresponding to each piece of the signaling data; A data compression module, configured to classify and perform secondary compression on all the compressed signaling based on the user code and the category code, to obtain data to be stored corresponding to each of the multiple category codes corresponding to each user code; A data writing module, used for storing all the data to be stored in column format; The compressed signaling includes a reference offset corresponding to the interaction time in the signaling data; and the data compression module is specifically used to: Classifying all the compressed signaling according to the user codes to obtain a user set corresponding to each user code; Dividing each of the user sets into a plurality of category subsets according to the category code; wherein the user codes and category codes of all compressed signaling in the category subsets are the same; Sorting all compressed signaling in each of the category subsets in ascending order of the reference offsets to obtain an ordered category subset corresponding to each of the category subsets; The user code and the category code are used as the primary index and the secondary index respectively to obtain a plurality of index combinations; each of the index combinations uniquely corresponds to one of the ordered category subsets; For each of the index combinations, all reference offsets and other information in the ordered category subset corresponding to the index combination are serialized to obtain the data to be stored corresponding to each of the index combinations.

7. An electronic device, characterized in that: include: A memory and a processor, wherein the memory stores a software program, and when the electronic device is running, the processor executes the software program to implement the method according to any one of claims 1 to 5.

8. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 5 is implemented.

9. A program product, characterized in that When the program product is executed by a processor, the method according to any one of claims 1 to 5 is implemented.

Citation Information

Patent Citations

  • Log compression method and device, log decompression method and device and storage medium

    CN110851409A