Signaling data storage method and device, electronic equipment, storage medium and program product
By initially compressing the interaction time in the signaling data, and classifying and secondary compression based on user encoding and category encoding, the data is finally stored in columnar storage, which solves the problem of high consumption of signaling data storage resources in the prior art, and an efficient storage solution is realized.
Patent Information
- Application Number
- CN202510518023.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-24
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2045-04-24
AI Technical Summary
The prior art is difficult to effectively save storage resources to store a large amount of signaling data, especially when operators need to retain signaling data for a long time.
By initially compressing the interaction time in the signaling data, compressed signaling is obtained; then, based on user encoding and category encoding, all compressed signaling is classified and secondary compressed to obtain the data to be stored; finally, the data to be stored is stored column-typed based on user encoding and category encoding.
It effectively reduces the storage space of signaling data and achieves the goal of efficiently storing signaling data while saving storage resources.
Smart Images

Figure CN120029558A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data storage, and in particular to a method, device, electronic device, storage medium and program product for storing signaling data. Background Art
[0002] With the rapid development of mobile communication technology and the popularity of smart phones, the frequency of signaling interactions between user terminal devices and base stations has increased exponentially. Signaling data, as key information that records user location, network status, and communication behavior, is an important basic data source for operators to optimize networks, analyze user behavior, and improve service quality.
[0003] According to statistics, the amount of signaling data generated by a single base station in a single day can reach TB (Tera Byte), and the amount of data that operators across the country need to process every day has exceeded EB (Exa Byte). In addition, operators need to retain signaling data for 6-12 months to meet data compliance requirements, which consumes a huge amount of storage resources. Therefore, how to store a large amount of signaling data while saving storage resources is an urgent problem to be solved. Summary of the invention
[0004] The purpose of the present invention is to provide a method, device, electronic device, storage medium and program product for storing signaling data to improve the problems existing in the prior art.
[0005] The embodiments of the present invention can be implemented as follows: In a first aspect, the present invention provides a method for storing signaling data, comprising: Acquire a signaling data set in a target area within a specified time; the signaling data set includes a plurality of signaling data, each of which includes a user code, a category code, an interaction time, and other information; Compressing the interaction time in each piece of the signaling data to obtain compressed signaling corresponding to each piece of the signaling data; Based on the user code and the category code, all the compressed signaling are classified and compressed twice to obtain data to be stored corresponding to each of the multiple category codes corresponding to each user code; All the data to be stored are stored in column format using the user code as an index.
[0006] Optionally, the step of compressing the interaction time in each piece of the signaling data to obtain compressed signaling corresponding to each piece of the signaling data includes: For each piece of the signaling data, reading the interaction time in the signaling data; Convert the interaction time and the zero o'clock time of the day on which the interaction time falls into timestamps, and obtain an interaction timestamp and a reference timestamp corresponding to the interaction time respectively; Calculate the difference between the interaction timestamp and the reference timestamp to obtain a reference offset; The reference offset is added to the signaling data, and the interaction time in the signaling data is deleted to obtain the compressed signaling.
[0007] Optionally, the compressed signaling includes a reference offset corresponding to the interaction time in the signaling data; The step of classifying and recompressing all the compressed signaling based on the user code and the category code to obtain the data to be stored corresponding to each of the multiple category codes corresponding to each user code includes: Classifying all the compressed signaling according to the user codes to obtain a user set corresponding to each user code; Dividing each of the user sets into a plurality of category subsets according to the category code; wherein the user codes and category codes of all compressed signaling in the category subsets are the same; Sorting all compressed signaling in each of the category subsets in ascending order of the reference offsets to obtain an ordered category subset corresponding to each of the category subsets; The user code and the category code are used as the primary index and the secondary index respectively to obtain a plurality of index combinations; each of the index combinations uniquely corresponds to one of the ordered category subsets; For each of the index combinations, all reference offsets and other information in the ordered category subset corresponding to the index combination are serialized to obtain the data to be stored corresponding to each of the index combinations.
[0008] Optionally, the other information includes data corresponding to at least one information field; The step of serializing all reference offsets and other information in the ordered category subset corresponding to the index combination includes: Determining the number of signaling in the ordered category subset corresponding to the index combination; Serializing the reference offsets of all compressed signaling in the ordered category subset corresponding to the index combination to obtain an offset list; Using the run-length encoding rule, serializing the data corresponding to each information field of all compressed signaling in the ordered category subset corresponding to the index combination, to obtain a compression list corresponding to each information field; The signaling quantity, the offset list, and the compression list corresponding to each of the information fields are combined to obtain the data to be stored corresponding to the index combination.
[0009] Optionally, the step of serializing the reference offsets of all compressed signaling in the ordered category subset corresponding to the index combination to obtain an offset list includes: Using the reference offset of the first compressed signaling in the ordered category subset as the reference offset; Calculate the difference between the reference offset of the i-th compressed signaling in the ordered category subset and the reference offset to obtain the compressed timestamp corresponding to the i-th compressed signaling; wherein, , K represents the number of signals in the ordered category subset; The reference offset and the compressed timestamps corresponding to the K compressed signalings in the ordered category subset are sequentially combined to obtain the offset list.
[0010] Optionally, the step of storing all the data to be stored in column format using the user code as an index includes: Based on the Schema information of the offset list and each compression list, the offset list and each compression list in the data to be stored corresponding to each index combination are stored in a columnar manner in a single linked list to obtain a single linked list storage structure corresponding to each index combination; Based on the respective Schema information of the primary index and the signaling quantity, the signaling quantity corresponding to each primary index and each index combination is stored in a columnar form in the form of a double pointer linked list; the Schema information reflects the field name, the data type to which the field belongs, the data structure and the data organization rule; Among them, in the data block where the main index is located, the first pointer is the memory address where the signaling quantity corresponding to the first index combination where the main index is located is stored. If the main index is not the last one, the second pointer is the memory address where the next main index of the main index is stored. If the main index is the last one, the second pointer is a null pointer; In the data block where the signaling quantity corresponding to the index combination is located, the first pointer is the memory address of the first data block in the single linked list storage structure corresponding to the index combination; if the index combination is not the last of the multiple index combinations with the same user code, the second pointer is the memory address where the signaling quantity corresponding to the next index combination of the current index combination is stored; if the index combination is the last of the multiple index combinations with the same user code, the second pointer is a null pointer.
[0011] In a second aspect, the present invention provides a storage device for signaling data, comprising: An acquisition module is used to acquire a signaling data set in a target area within a specified time; the signaling data set includes a plurality of signaling data, each of which includes a user code, a category code, an interaction time and other information; A time compression module, used to compress the interaction time in each piece of the signaling data to obtain a compressed signaling corresponding to each piece of the signaling data; A data compression module, configured to classify and perform secondary compression on all the compressed signaling based on the user code and the category code, to obtain data to be stored corresponding to each of the multiple category codes corresponding to each user code; The data writing module is used to store all the data to be stored in column format.
[0012] In a third aspect, the present invention provides an electronic device, comprising: a memory and a processor, wherein the memory stores a software program, and when the electronic device is running, the processor executes the software program to implement the method described in the first aspect.
[0013] In a fourth aspect, the present invention provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method described in the first aspect is implemented.
[0014] In a fifth aspect, the present invention provides a program product, which, when executed by a processor, implements the method described in the first aspect.
[0015] Compared with the prior art, the embodiments of the present invention provide a method, device, electronic device, storage medium and program product for storing signaling data. After obtaining a number of signaling data within a specified time in a target area, the interaction time in each signaling data is first compressed to obtain the compressed signaling corresponding to each signaling data, thereby avoiding the large amount of storage resources occupied by the interaction time. Then, based on the user code and category code, all compressed signaling are classified and compressed twice to obtain the data to be stored corresponding to multiple category codes corresponding to each user code. Finally, all the data to be stored are stored in columnar form. In this way, secondary compression after classification and columnar storage can further save storage resources. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for use in the embodiments are briefly introduced below. It should be understood that the following drawings only show certain embodiments of the present invention and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other related drawings can be obtained based on these drawings without creative work.
[0017] Figure 1 One of the flow charts of a method for storing signaling data provided by an embodiment of the present invention.
[0018] Figure 2 A second flow chart of a method for storing signaling data provided in an embodiment of the present invention.
[0019] Figure 3 An exemplary diagram of multiple ordered category subsets provided by an embodiment of the present invention.
[0020] Figure 4 An example diagram of data to be stored corresponding to a plurality of index combinations provided in an embodiment of the present invention.
[0021] Figure 5 An example diagram of a single linked list storage structure provided in an embodiment of the present invention.
[0022] Figure 6 An example diagram of a double linked list storage structure provided in an embodiment of the present invention.
[0023] Figure 7 A schematic diagram of the structure of a storage device for signaling data provided by an embodiment of the present invention.
[0024] Figure 8 A schematic diagram of the structure of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0025] In order to make the purpose, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Generally, the components of the embodiments of the present invention described and shown in the drawings here can be arranged and designed in various different configurations.
[0026] Therefore, the following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the invention claimed for protection, but merely represents selected embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0027] It should be noted that similar reference numerals and letters denote similar items in the following drawings, and therefore, once an item is defined in one drawing, further definition and explanation thereof is not required in subsequent drawings.
[0028] In the description of the present invention, it should be noted that if the terms "upper", "lower", "inside", "outside", etc. appear to indicate an orientation or position relationship, they are based on the orientation or position relationship shown in the accompanying drawings, or are the orientation or position relationship in which the product of the invention is usually placed when used. They are only for the convenience of describing the present invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as a limitation on the present invention.
[0029] In addition, the terms “first”, “second”, etc., if used, are merely used to distinguish between the descriptions and should not be understood as indicating or implying relative importance.
[0030] It should be noted that, in the absence of conflict, the features in the embodiments of the present invention may be combined with each other.
[0031] At present, mobile communication operators mainly use traditional distributed storage systems to store signaling data. Its technical implementation usually includes the following steps: (1) The base station uploads the real-time generated user connection request, handover instruction, location update and other signaling data to the data center; (2) The data center uses time series or user ID as an index to write signaling data into a distributed database or file system in the following way: after converting the time in the signaling data into a timestamp, the data content in the signaling data is converted into a binary sequence, and then a series of transformations are performed on the binary sequence to achieve data compression and finally storage.
[0032] However, the data type of the timestamp is long type, which takes up 8 bytes. That is, in a piece of signaling data containing a timestamp, the space occupied by the timestamp exceeds 56% of the space occupied by the entire signaling data. Therefore, even if the timestamp is compressed by converting it into binary, the space occupied by it is still large.
[0033] Based on the discovery of the above technical problems, the inventor believes that the timestamp can be further compressed to reduce the space occupied by time. And through long-term observation and research, the inventor found that: first, signaling data has the characteristics of high-frequency generation, strong content repetitiveness (such as periodic location updates), high field redundancy, etc., so the existing technology directly stores a single signaling data, and the storage resource utilization rate is low; secondly, the existing technology lacks in-depth mining of the spatiotemporal correlation of signaling data. For example, multiple signals generated by the same user in adjacent time periods often contain repeated base station identifiers or location information, and the existing storage architecture fails to effectively reduce the space occupied by such repeated content.
[0034] In view of this, an embodiment of the present invention provides a method for storing signaling data, which can, on the one hand, preliminarily compress the interaction time to obtain compressed signaling, and on the other hand, classify all compressed signaling based on user codes and category codes and then perform secondary compression, so that the repeated information appearing in all compressed signaling of each category code corresponding to a user code is compressed and then stored in a columnar manner, which greatly reduces the storage space occupied. The following is a detailed description through embodiments and in conjunction with the attached drawings.
[0035] Please refer to Figure 1 , Figure 1 The present invention provides a flowchart of a method for storing signaling data, which may be performed by, but not limited to, a computing device, a server, a server cluster, etc. The method includes the following steps S101 to S104.
[0036] S101. Obtain a signaling data set within a target area within a specified time.
[0037] In this embodiment, the signaling data set includes a number of signaling data, each signaling data includes a user code (representing a user terminal used by a user), a category code (reflecting the event category of the signaling interaction, such as an attachment request, a switching instruction, a tracking area update, a service request, etc.), an interaction time and other information, and the other information includes data corresponding to multiple information fields.
[0038] Optionally, the target area may be the whole country or a city, such as Beijing, Chengdu, etc., and the specified period may be one day, multiple days, one week, one month, etc. In the corresponding example, the signaling data set may include all signaling data generated by Chengdu in April 25. This example is only an example, and the embodiments of the present invention are not limited to this.
[0039] S102: compress the interaction time in each piece of signaling data to obtain compressed signaling corresponding to each piece of signaling data.
[0040] In this embodiment, the interaction time in each piece of signaling data is compressed, which can reduce the storage resource occupation of the time information in subsequent signaling.
[0041] S103: Based on the user code and the category code, all compressed signaling is classified and compressed twice to obtain data to be stored corresponding to multiple category codes corresponding to each user code.
[0042] In an optional example, for 50 compressed signalings with the same user code, if the 50 compressed signalings have 4 different category codes, then the 50 compressed signalings will be compressed into 4 data to be stored. This example is only an example and is not limited here.
[0043] S104: Using the user code as an index, all data to be stored are stored in column format.
[0044] In this embodiment, compared with row-based storage, column-based storage can reduce storage resource usage.
[0045] The signaling data storage method provided by the embodiment of the present invention, after acquiring a number of signaling data within a specified time in a target area, first compresses the interaction time in each signaling data to obtain the compressed signaling corresponding to each signaling data, so as to avoid a large amount of storage resources occupied by the interaction time, and then classifies and recompresses all compressed signaling based on user codes and category codes to obtain the data to be stored corresponding to multiple category codes corresponding to each user code, and finally stores all the data to be stored in columnar form, so that the classification process followed by secondary compression combined with columnar storage can further save storage resources.
[0046] Because the timestamp is defined as the total number of seconds from 00:00:00 Greenwich Mean Time, January 1, 1970 (i.e. 08:00:00 Beijing Time, January 1, 1970) to the present, for example, Beijing Time 2025-04-03 12:00:00, the corresponding timestamp is 1743652800, a total of 10 digits. There are only 86400 seconds in a day, so the maximum difference in timestamps within the same day is 86400, a total of 5 digits.
[0047] Therefore, in the above step S102, the zero o'clock of the day when the interaction time is located can be used as the reference time, and then the interaction time and the reference time are converted into timestamps and the difference is calculated, so as to achieve preliminary compression of the interaction time, and the interaction time can be compressed to a 5-digit decimal number. And this type of calculation operation is a simple addition and subtraction operation, and its algorithm complexity is only O(n), that is, it does not increase the calculation complexity.
[0048] That is, in the above step S102, the process of "compressing the interaction time in each signaling data to obtain compressed signaling corresponding to each signaling data" may include the following sub-steps S1021~S1024.
[0049] S1021. For each piece of signaling data, read the interaction time in the signaling data.
[0050] S1022: Convert the interaction time and the zero o'clock time of the day when the interaction time occurs into timestamps, and obtain an interaction timestamp and a reference timestamp corresponding to the interaction time respectively.
[0051] In this embodiment, the timestamp corresponding to the interaction time is the interaction timestamp, and the zero o'clock of the day when the interaction time occurs is the reference time, and the timestamp corresponding to the reference time is the reference timestamp.
[0052] S102: Calculate the difference between the interaction timestamp and the reference timestamp to obtain a reference offset.
[0053] S1024. Add the reference offset to the signaling data, and delete the interaction time in the signaling data to obtain compressed signaling.
[0054] In this embodiment, for each piece of signaling data in the signaling data set, steps S1021 to S1024 are executed to complete preliminary compression of the time in each piece of signaling data.
[0055] In an optional example, the data type of the reference offset can be an int type. Assuming that the interaction time is 2025-04-03 12:00:00 Beijing time, and the corresponding interaction timestamp is 1743652800, then the reference time is 2025-04-03 00:00:00 Beijing time, and the corresponding reference timestamp is 1743609600, then the reference offset is 43200. Compared with the prior art that requires 8 bytes to store the interaction timestamp, through the preliminary compression of the interaction time, the reference offset only occupies 4 bytes, that is, the preliminary compression method can greatly reduce the storage resource occupation of the interaction time. This example is only an example, and the interaction time can also be accurate to milliseconds. The embodiment of the present invention does not limit the accuracy of the interaction time.
[0056] In an optional implementation, since signaling data has the characteristics of high frequency generation, strong content repetitiveness (such as periodic location update), high field redundancy, etc., it can be compressed twice by serialization after classification. Figure 1 Based on Figure 2 For the above step S103, the process of "classifying and recompressing all compressed signaling based on user codes and category codes to obtain the data to be stored corresponding to multiple category codes corresponding to each user code" may include the following sub-steps S1031~S1035.
[0057] S1031. Classify all compressed signaling according to user codes to obtain a user set corresponding to each user code.
[0058] S1032. Divide each user set into multiple category subsets according to the category code.
[0059] In this embodiment, in a category subset, the user codes and category codes of all compressed signaling are the same.
[0060] S1033. Sort all compressed signaling in each category subset in ascending order of reference offsets to obtain an ordered category subset corresponding to each category subset.
[0061] In the optional example, it is assumed that other information may include: base station cell code, base station location code, flag value, service information, etc., that is, the information field includes: Cid field, Ghash field, Flag field, Amto field, etc.
[0062] Please refer to Figure 3 , Figure 3 Taking the compressed signaling including Uid field, i.e. Tid field), Offset field, Cid field, Ghash field, and Flag field as an example, multiple ordered category subsets are shown. Figure 3 The ones with the same serial number color are an ordered category subset.
[0063] It should be noted that Figure 3 The serial number column is only for the convenience of showing the number of signaling in each subset. Figure 3 The use of numbers such as 0, 1, and 2 to identify different category codes is only an example and is not limited here.
[0064] S1034. Use the user code and the category code as the primary index and the secondary index respectively to obtain multiple index combinations.
[0065] In this embodiment, an index combination includes a primary index and a secondary index, so each index combination uniquely corresponds to an ordered category subset.
[0066] If the signaling data set involves 3 user codes, and all signaling data corresponding to the 3 user codes involve 4 category codes, then 12 index combinations can be determined. Figure 3 There are 8 index combinations, and this example is only an example and is not limited here.
[0067] S1035. For each index combination, serialize all reference offsets and other information in the ordered category subset corresponding to the index combination to obtain the data to be stored corresponding to each index combination.
[0068] In this embodiment, it is necessary to serialize the reference offsets of all compressed signaling and the data of each information field in each ordered category subset, so as to reduce duplicate data.
[0069] Optionally, for each index combination, in step S1035, the process of "serializing all reference offsets and other information in the ordered category subset corresponding to the index combination" may include the following sub-steps S10351~S10354.
[0070] S10351. Determine the number of signalings in the ordered category subset corresponding to the index combination.
[0071] S10352. Serialize the reference offsets of all compressed signaling in the ordered category subset corresponding to the index combination to obtain an offset list.
[0072] Optionally, the reference offset of the first compressed signaling in the ordered category subset may be used as a reference, and then subtracted again to reduce the space occupied by the time information. For an ordered category subset corresponding to an index combination, the process of obtaining the offset list may include steps S001 to S003: S001, taking the reference offset of the first compressed signaling in the ordered category subset as the reference offset; S002. Calculate the difference between the reference offset and the base offset of the i-th compressed signaling in the ordered category subset to obtain the compressed timestamp corresponding to the i-th compressed signaling; wherein, , K represents the number of signals in the ordered category subset; S003. Combine the reference offset and the compressed timestamps corresponding to K compressed signalings in the ordered category subset in sequence to obtain an offset list.
[0073] In this embodiment, after the K reference offsets in the ordered category subset corresponding to an index combination are ordered according to steps S001 to S003, the obtained offset list includes K+1 values.
[0074] S10353. Using the run-length coding rule, serialize the data corresponding to each information field of all compressed signaling in the ordered category subset corresponding to the index combination to obtain a compression list corresponding to each information field.
[0075] In this embodiment, since most information fields are of character type rather than numeric type, for each information field of other information: the run-length encoding rule can be used to process all compressed signaling in the ordered category subset corresponding to an index combination into a compressed list corresponding to the data in the information field.
[0076] Among them, Run-Length Encoding, also known as RLE encoding, is a lossless data compression algorithm. The core idea is to shorten the data length by recording consecutive repeated characters (or symbols) and their occurrence times. Therefore, the present invention does not introduce the specific process of using the Run-Length Encoding rule.
[0077] S10354. Combine the signaling quantity, the offset list, and the compression list corresponding to each information field to obtain the data to be stored corresponding to the index combination.
[0078] In this embodiment, the data to be stored corresponding to a combination includes the signaling quantity, the offset list, and the compression list corresponding to each information field.
[0079] In the optional example, please continue to combine Figure 3 , if Figure 3 The three UIDs "728669839005169607", "311553764816986473", and "675786698390053333" are recorded as X1, X2, and X3 respectively. Figure 3 The data to be stored obtained by converting the 8 ordered category subsets in Figure 4 shown.
[0080] exist Figure 4 In the example, the compressed lists corresponding to the information fields such as the Cid field, the Ghash field, and the Flag field are Cid_list, Ghash_list, and Flag_list, and according to the above steps S10351 to S10354, Figure 3 The data to be stored for the first ordered category subset (rows with green serial numbers) and the second ordered category subset (rows with red serial numbers) are Figure 4 Content in the green area and content in the red area.
[0081] contrast Figure 3 and Figure 4 It can be seen intuitively that the reference offset (i.e., Offset field) of each ordered category subset and each information field are ordered separately, so that the ordered data compression is achieved, which can reduce the storage resources required for storage.
[0082] It should be noted that Figure 3 The Uid shown is just an example. Figure 3 , Figure 4 The various information fields shown are only examples, and the embodiments of the present invention do not limit the specific content in the corresponding signaling data.
[0083] In an optional implementation, the sub-steps of step S104 above may include: S1041. Based on the Schema information of the offset list and each compression list, the offset list and each compression list in the data to be stored corresponding to each index combination are stored in columnar form in the form of a single linked list to obtain a single linked list storage structure corresponding to each index combination.
[0084] In this embodiment, the Schema information may reflect the field name, the data type to which the field belongs, the data structure, and the data organization rule. The Schema information also needs to be stored, and after being stored, it becomes metadata.
[0085] It can be understood that in a single linked list storage structure, multiple data in the list are each located in a data block. Taking the offset list as an example, in a single linked list storage structure, the base offset and K compressed timestamps of the offset list are respectively located in K+1 data blocks.
[0086] The data block is divided into a data field and a pointer field. The pointer field of each data block in the single linked list storage structure includes only one pointer, which is used to store the memory address of the next data block.
[0087] S1042. Based on the Schema information of the primary index and the signaling quantity, the signaling quantity corresponding to each primary index and each index combination is stored in a columnar form in the form of a double-pointer linked list.
[0088] In this embodiment, the signaling quantity corresponding to each main index and each index combination occupies one data block respectively. The signaling quantity corresponding to each main index and each index combination is stored in a columnar manner in the form of a double-pointer linked list, and a double-linked list storage structure can be obtained.
[0089] Assuming that there are a total of M user codes involved in the signaling data set, that is, there are M primary indexes, then for the mth primary index (m∈[1,M]), the pointer field in the data block where it is located includes 2 pointers, where the first pointer is the memory address where the signaling quantity corresponding to the first index combination where the primary index is located is stored, and the second pointer is divided into the following two cases: (1) Among the M primary indexes, if the mth primary index is not the last one (i.e., m≠M), the second pointer is the memory address of the next primary index of the mth primary index, i.e., the second pointer is the memory address of the m+1th primary index; (2) Among M primary indexes, if the mth primary index is the last one (i.e., m=M), the second pointer is a null pointer.
[0090] Assuming that N index combinations are obtained in step S1034, then for the nth index combination (n∈[1, N]), the pointer field in the data block where the corresponding signaling quantity is located includes 2 pointers, wherein the first pointer is the memory address of the first data block in the single linked list storage structure corresponding to the nth index combination, and assuming that among the N index combinations, there are Q index combinations that have the same user code as that in the nth index combination, then the second pointer is divided into the following two cases: (1) If the nth index combination is not the last one of the Q index combinations with the same user code, the second pointer is the memory address where the signaling quantity corresponding to the next index combination of the nth index combination is stored, that is, the second pointer is the memory address where the signaling quantity corresponding to the n+1th index combination is stored; (2) If the nth index combination is the last one of the Q index combinations with the same user code, the second pointer is a null pointer.
[0091] In order to understand the above-mentioned single linked list storage structure and double linked list storage structure, the following examples are used for explanation.
[0092] In the optional example, please combine Figure 4 , Figure 4 The 8 index combinations and their corresponding signaling quantities are shown in Table 1. Figure 4 In the green area shown, the four lists corresponding to the index combination (X1,0) are the offset list (Offset_list), Cid_list, Ghash_list, and Flag_list, as shown in Table 2 below.
[0093] Table 1
[0094] Table 2 4 lists corresponding to (X1,0)
[0095] First, the four lists in Table 2 are stored in column format, and the single linked list storage structure corresponding to (X1,0) is obtained as follows: Figure 5 As shown. Figure 4 The single linked list storage structure corresponding to the other 7 index combinations in Figure 5 Similar, no repetition here.
[0096] Next, the main index and signaling quantity shown in Table 1 are stored, and the obtained double linked list storage structure is as follows: Figure 6 shown.
[0097] exist Figure 5 and Figure 6 In the figure, the arrow points to the memory address of the pointer, and NULL represents a null pointer.
[0098] Combination Figure 6 It can be seen that the main index uses double pointers when storing, which ensures that the main index can address downward to read other main indexes, and can also address to the right to read the corresponding signaling quantity, while the signaling quantity uses double pointers when storing, which ensures that the signaling quantity can address downward to read the next signaling quantity of the same main index, and can also address to the right to read the single linked list storage structure.
[0099] Combination Figure 5 It can be seen that the four lists corresponding to (X1,0) are stored using pointers to achieve continuous addressing to ensure continuous readability of the data.
[0100] Combination Figure 3 , Figure 4 , Figure 5 , Figure 6 It can be seen that, due to the small difference in timestamps on the same day, the present invention compresses the interaction time twice, thereby compressing the K interaction times corresponding to the same index combination into an Offset_list, greatly reducing the storage space occupied by time data.
[0101] At the same time, due to the strong repetitiveness of the contents in multiple signaling data with similar interaction time and the same user code, the present invention also compresses the data corresponding to each information field in the K signaling data corresponding to the same index combination into a compressed list based on the run-length coding rule, thereby reducing the space occupied by repeated data. The present invention uses the user code as the main index and the category code as the secondary index to determine multiple index combinations, and uses column storage to store the main index of each index combination and the signaling quantity, offset list and compression list corresponding to each information field corresponding to each index combination, and uses pointers in the storage structure to achieve addressing to ensure data readability.
[0102] The above content introduces the compressed storage of a signaling data set to obtain a double linked list storage structure and a single linked list storage structure corresponding to each index combination.
[0103] Based on the double linked list storage structure and the single linked list storage structure corresponding to each index combination, the process of recovering a signaling data set is the reverse process of compression storage. The following takes the specified time as a specified day as an example to briefly introduce the recovery process: (1) First, read all data from the double-linked list storage structure and the single-linked list storage structure corresponding to each index combination, and add secondary indexes (i.e., category codes) according to the order of the number of signalings corresponding to each primary index; (2) For each index combination corresponding to the offset list (including K+1 values), starting from the second value in the list, each data is added to the first value in the list in turn, so that the K reference offsets corresponding to an index combination can be restored. Then, the K reference offsets are added to the timestamp of the specified time at zero o'clock on the day to obtain the K interaction timestamps corresponding to the index combination, and then the K interaction times corresponding to the index combination can be converted. (3) For the compressed lists of each information field corresponding to each index combination, based on the run-length encoding rule, each compressed list is restored to obtain K original field values corresponding to each information field corresponding to an index combination; (4) For each index combination, the user code, category code, corresponding K interaction times, and K original field values corresponding to each information field in the index combination are combined to obtain K pieces of signaling data with the same user code and category code.
[0104] The inventor has verified that for 3.166 billion pieces of signaling data, the storage space occupied is 37.2 GB. After compression processing by the method of the present invention, the storage space occupied is only 19.7 GB, which is nearly half of the original space. In addition, the signaling data can be restored through the recovery process described above. Therefore, the present invention can reduce the storage space occupied by a large amount of signaling data under the premise of lossless compression.
[0105] In order to execute the corresponding steps in the above method embodiment and various possible implementation modes, an implementation method of a storage device for signaling data is provided below.
[0106] See also Figure 7 , Figure 7 The schematic diagram of the structure of the storage device of signaling data provided by the embodiment of the present invention is shown. The storage device 200 of signaling data comprises: an acquisition module 210, a time compression module 220, a data compression module 230 and a data writing module 240.
[0107] The acquisition module 210 is used to acquire a signaling data set in a target area within a specified time; the signaling data set includes a plurality of signaling data, each of which includes a user code, a category code, an interaction time and other information; The time compression module 220 is used to compress the interaction time in each signaling data to obtain a compressed signaling corresponding to each signaling data; The data compression module 230 is used to classify and re-compress all compressed signaling based on the user code and the category code to obtain the data to be stored corresponding to each of the multiple category codes corresponding to each user code; The data writing module 240 is used to store all the data to be stored in columns.
[0108] Optionally, the time compression module 220 can be specifically used to: for each signaling data, read the interaction time in the signaling data; convert the interaction time and the zero o'clock time of the day when the interaction time occurs into timestamps, and respectively obtain the interaction timestamp and reference timestamp corresponding to the interaction time; calculate the difference between the interaction timestamp and the reference timestamp to obtain a reference offset; add the reference offset to the signaling data, and delete the interaction time in the signaling data to obtain compressed signaling.
[0109] Optionally, the compressed signaling includes a reference offset corresponding to the interaction time in the signaling data. The data compression module 230 can be specifically used to: classify all compressed signaling according to user codes to obtain a user set corresponding to each user code; divide each user set into multiple category subsets according to category codes; wherein the user codes and category codes of all compressed signaling in the category subsets are the same; sort all compressed signaling in each category subset in ascending order of reference offsets to obtain an ordered category subset corresponding to each category subset; use the user code and category code as the primary index and secondary index respectively to obtain multiple index combinations; each index combination uniquely corresponds to an ordered category subset; for each index combination, serialize all reference offsets and other information in the ordered category subset corresponding to the index combination to obtain the data to be stored corresponding to each index combination.
[0110] Optionally, the other information includes data corresponding to at least one information field. In the process of serializing all reference offsets and other information in the ordered category subset corresponding to the index combination, the data compression module 230 can be specifically used to: determine the number of signaling in the ordered category subset corresponding to the index combination; serialize the reference offsets of all compressed signaling in the ordered category subset corresponding to the index combination to obtain an offset list; use the run-length encoding rule to serialize the data corresponding to each information field of all compressed signaling in the ordered category subset corresponding to the index combination to obtain a compression list corresponding to each information field; combine the number of signaling, the offset list and the compression list corresponding to each information field to obtain the data to be stored corresponding to the index combination.
[0111] Optionally, in the process of serializing the reference offsets of all compressed signalings in the ordered category subset corresponding to the index combination to obtain the offset list, the data compression module 230 can be specifically used to: use the reference offset of the first compressed signaling in the ordered category subset as the reference offset; calculate the difference between the reference offset and the reference offset of the i-th compressed signaling in the ordered category subset to obtain the compression timestamp corresponding to the i-th compressed signaling; wherein, , K represents the number of signals in the ordered category subset; the base offset and the compressed timestamps corresponding to the K compressed signals in the ordered category subset are sequentially combined to obtain an offset list.
[0112] Optionally, the data writing module 240 can be specifically used to: based on the Schema information of the offset list and each compression list, store the offset list and each compression list in the data to be stored corresponding to each index combination in a columnar form in the form of a single linked list, and obtain a single linked list storage structure corresponding to each index combination; based on the Schema information of the primary index and the signaling quantity, store the signaling quantity corresponding to each primary index and each index combination in a columnar form in the form of a double pointer linked list; the Schema information reflects the field name, the data type to which the field belongs, the data structure, and the data organization rule; Among them, in the data block where the primary index is located, the first pointer is the memory address where the signaling quantity corresponding to the first index combination where the primary index is located is stored. If the primary index is not the last one, the second pointer is the memory address where the next primary index of the primary index is stored. If the primary index is the last one, the second pointer is a null pointer; In the data block where the signaling quantity corresponding to the index combination is located, the first pointer is the memory address of the first data block in the single linked list storage structure corresponding to the index combination. If the index combination is not the last one of multiple index combinations with the same user code, the second pointer is the memory address where the signaling quantity corresponding to the next index combination of the current index combination is stored. If the index combination is the last one of multiple index combinations with the same user code, the second pointer is a null pointer.
[0113] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working process of the signaling data storage device 200 described above can refer to the corresponding process in the aforementioned method embodiment, and will not be repeated here.
[0114] See also Figure 8 , Figure 8 The electronic device 300 includes a processor 310 , a memory 320 , and a bus 330 , wherein the processor 310 is connected to the memory 320 via the bus 330 .
[0115] The memory 320 may be used to store software programs, for example, software programs corresponding to the storage device 200 for signaling data provided in the embodiment of the present invention. The processor 310 executes various functional applications and data processing to implement the signaling data storage method provided in the embodiment of the present invention by running the software programs stored in the memory 320.
[0116] Among them, the memory 320 can be but is not limited to: RAM (Random Access Memory), ROM (Read Only Memory), FLASH (Flash Memory), PROM (Programmable Read-Only Memory), EPROM (Erasable Programmable Read-Only Memory), EEPROM (Electric Erasable Programmable Read-Only Memory), etc.
[0117] The processor 310 may be an integrated circuit chip with signal processing capability. The processor 310 may be a general-purpose processor, including: CPU (Central Processing Unit), NP (Network Processor), SoC (System on Chip), etc.; it may also be: DSP (Digital Signal Processing), ASIC (Application Specific Integrated Circuit), FPGA (Field-Programmable Gate Array) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.
[0118] Understandably, Figure 8 The structure shown is for illustration only. The electronic device 300 may also include Figure 8 More or fewer components as shown, or with Figure 8 Different configurations are shown. Figure 8 Each component shown in the figure can be implemented by hardware, software or a combination thereof.
[0119] The embodiment of the present invention further provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the signaling data storage method disclosed in the above embodiment is implemented. The computer-readable storage medium can be, but is not limited to, various media that can store program codes, such as a USB flash drive, a mobile hard disk, a ROM, a RAM, a PROM, an EPROM, an EEPROM, a FLASH disk, or an optical disk.
[0120] An embodiment of the present invention further provides a program product, which, when executed by a processor, implements the signaling data storage method disclosed in the above embodiment.
[0121] In summary, the embodiments of the present invention provide a method, device, electronic device, storage medium and program product for storing signaling data. After obtaining a number of signaling data within a specified time in a target area, the interaction time in each signaling data is first compressed to obtain the compressed signaling corresponding to each signaling data, thereby avoiding the large amount of storage resources occupied by the interaction time. Then, based on the user code and category code, all compressed signaling are classified and secondary compressed to obtain the data to be stored corresponding to multiple category codes corresponding to each user code. Finally, all the data to be stored are stored in columnar form. In this way, secondary compression after classification and columnar storage can further save storage resources.
[0122] The above are only specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily thought of by a person skilled in the art within the technical scope disclosed by the present invention should be included in the protection scope of the present invention. Therefore, the protection scope of the present invention shall be subject to the protection scope of the claims.
Claims
1. A method for storing signaling data, characterized in that: include: Obtain the signaling data set within the target area within the specified time; The signaling data set includes a plurality of signaling data, each of which includes a user code, a category code, an interaction time and other information; Compressing the interaction time in each piece of the signaling data to obtain compressed signaling corresponding to each piece of the signaling data; Based on the user code and the category code, all the compressed signaling are classified and compressed twice to obtain data to be stored corresponding to each of the multiple category codes corresponding to each user code; All the data to be stored are stored in column format using the user code as an index.
2. The method according to claim 1, characterized in that The step of compressing the interaction time in each piece of the signaling data to obtain compressed signaling corresponding to each piece of the signaling data includes: For each piece of the signaling data, reading the interaction time in the signaling data; Convert the interaction time and the zero o'clock time of the day on which the interaction time falls into timestamps, and obtain an interaction timestamp and a reference timestamp corresponding to the interaction time respectively; Calculate the difference between the interaction timestamp and the reference timestamp to obtain a reference offset; The reference offset is added to the signaling data, and the interaction time in the signaling data is deleted to obtain the compressed signaling.
3. The method according to claim 1, characterized in that The compressed signaling includes a reference offset corresponding to an interaction time in the signaling data; The step of classifying and recompressing all the compressed signaling based on the user code and the category code to obtain the data to be stored corresponding to each of the multiple category codes corresponding to each user code includes: Classifying all the compressed signaling according to the user codes to obtain a user set corresponding to each user code; Dividing each of the user sets into a plurality of category subsets according to the category code; wherein the user codes and category codes of all compressed signaling in the category subsets are the same; Sorting all compressed signaling in each of the category subsets in ascending order of the reference offsets to obtain an ordered category subset corresponding to each of the category subsets; The user code and the category code are used as the primary index and the secondary index respectively to obtain a plurality of index combinations; each of the index combinations uniquely corresponds to one of the ordered category subsets; For each of the index combinations, all reference offsets and other information in the ordered category subset corresponding to the index combination are serialized to obtain the data to be stored corresponding to each of the index combinations.
4. The method according to claim 3, characterized in that The other information includes data corresponding to at least one information field; The step of serializing all reference offsets and other information in the ordered category subset corresponding to the index combination includes: Determining the number of signaling in the ordered category subset corresponding to the index combination; Serializing the reference offsets of all compressed signaling in the ordered category subset corresponding to the index combination to obtain an offset list; Using the run-length encoding rule, serializing the data corresponding to each information field of all compressed signaling in the ordered category subset corresponding to the index combination, to obtain a compression list corresponding to each information field; The signaling quantity, the offset list, and the compression list corresponding to each of the information fields are combined to obtain the data to be stored corresponding to the index combination.
5. The method according to claim 4, characterized in that The step of serializing the reference offsets of all compressed signaling in the ordered category subset corresponding to the index combination to obtain an offset list includes: Using the reference offset of the first compressed signaling in the ordered category subset as the reference offset; Calculate the difference between the reference offset of the i-th compressed signaling in the ordered category subset and the reference offset to obtain the compressed timestamp corresponding to the i-th compressed signaling; wherein, , K represents the number of signals in the ordered category subset; The reference offset and the compressed timestamps corresponding to the K compressed signalings in the ordered category subset are sequentially combined to obtain the offset list.
6. The method according to claim 4, characterized in that The step of storing all the data to be stored in column format using the user code as an index includes: Based on the Schema information of the offset list and each compression list, the offset list and each compression list in the data to be stored corresponding to each index combination are stored in a columnar manner in a single linked list to obtain a single linked list storage structure corresponding to each index combination; Based on the respective Schema information of the primary index and the signaling quantity, the signaling quantity corresponding to each primary index and each index combination is stored in a columnar form in the form of a double pointer linked list; the Schema information reflects the field name, the data type to which the field belongs, the data structure and the data organization rule; Among them, in the data block where the main index is located, the first pointer is the memory address where the signaling quantity corresponding to the first index combination where the main index is located is stored. If the main index is not the last one, the second pointer is the memory address where the next main index of the main index is stored. If the main index is the last one, the second pointer is a null pointer; In the data block where the signaling quantity corresponding to the index combination is located, the first pointer is the memory address of the first data block in the single linked list storage structure corresponding to the index combination; if the index combination is not the last of the multiple index combinations with the same user code, the second pointer is the memory address where the signaling quantity corresponding to the next index combination of the current index combination is stored; if the index combination is the last of the multiple index combinations with the same user code, the second pointer is a null pointer.
7. A storage device for signaling data, characterized in that: include: An acquisition module is used to acquire a signaling data set in a target area within a specified time; The signaling data set includes a plurality of signaling data, each of which includes a user code, a category code, an interaction time and other information; A time compression module, used to compress the interaction time in each piece of the signaling data to obtain a compressed signaling corresponding to each piece of the signaling data; A data compression module, configured to classify and perform secondary compression on all the compressed signaling based on the user code and the category code, to obtain data to be stored corresponding to each of the multiple category codes corresponding to each user code; The data writing module is used to store all the data to be stored in column format.
8. An electronic device, characterized in that: include: A memory and a processor, wherein the memory stores a software program, and when the electronic device is running, the processor executes the software program to implement the method according to any one of claims 1 to 6.
9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 6 is implemented.
10. A program product, characterized in that When the program product is executed by a processor, the method according to any one of claims 1 to 6 is implemented.
Citation Information
Patent Citations
Data storage method and device, and computer equipment
CN108233942A
Log compression method and device, log decompression method and device and storage medium
CN110851409A
Data processing method and device, electronic equipment and storage medium
CN111835700A
Signaling compression method and system based on space-time coding
CN112234995A
Data compression method, electronic equipment and storage medium
CN113746485A
Cited By
Data storage method and device, electronic equipment and storage medium
CN118981555A
Data storage method and device, electronic equipment and storage medium
CN118981555B